cavendish

Lab app · Scorecard · Q3 2026

The track record.

Lead time measures whether the lab is distinctive. Answer rate measures whether it is useful. A lab needs both, and only one of them will be asked about in a budget review.

Two clocks on this page. Lead time, prediction calibration and cost per validated recommendation are measurements over closed periods — they do not move because somebody concluded a run this afternoon, and they should not. What the lab is doing about the plan and the mix of work in flight are statements about right now, and they read the board and the library live.

Median lead time
0mo0· last crossing

11 fields reached mainstream · the running median as each crossed

Answer rate
0%32 pts· vs last month

share of questions the library answers well · by month

Cost per validated recommendation
$0.0k

947 attributed person-days, 11% of it work that produces none

Top-50 freshness
0%1 overdue

1 library items overdue for review

Fields
30
Candidates
10
Experiments
20
Validated
3
Killed
16
Recommendations
12
Corrections
1
Claims
208

Lead time · per field

When we opened it versus when it went mainstream.

Negative lead time is recorded honestly. An honest negative is the only thing that makes a positive one believable.

Calibration

What we said versus what happened.

Predictions are dated, confidence-scored and resolvable. A perfectly calibrated lab sits on the diagonal. We are over-confident at the top and under-confident at the bottom, which is the usual shape.

bucketpredictedactualn
90%+
92%
83%
6
70–90%
79%
71%
14
50–70%
60%
58%
19
30–50%
41%
47%
11
<30%
22%
30%
8

Brier 0.19 across 58 resolved predictions. Deferred to a spreadsheet for the first year, as the design says.

Against the plan · last 90 days

What the lab did about the firm's objectives.

The rest of this page measures whether the lab is any good. This answers the question the firm actually asks. Every row landed on somebody else — published, concluded, or handed over — because starting things is free and a scorecard that counted activity would produce activity.

Objectives in the plan
30
The lab is on
11
Something landed
8
Nothing landed
3

Being on an objective is not the same as having done something about it, and the gap between those two rows is the only part of this worth arguing about.

20 published3 runs concluded1 handed over
  1. C.1

    Shared AI infrastructure is built and maintained, including governed assets. Every team builds on it. No team starts from zero.

    7 landed

    Standing coverage: Agentic Memory System · AI Gateway · Open Weight Models

  2. 3.1

    We have a standard approach to build, deploy and manage agents and workbenches, faster, more consistently, and to a higher quality bar than building bespoke each time.

    6 landed

    Standing coverage: Agentic Memory System · Auth Broker · AI-SDLC

  3. C.4

    An AI risk, ethics, security, quality and compliance framework is operational, applied across all AI workloads – internally and for clients. Externally credible, not just internally compliant.

    3 landed

    Standing coverage: Auth Broker · Eval Harnesses · Cyber Cold War

  4. 1.2

    Every team has reimagined how we do our work and built the capability to keep doing so. Productivity is visible at function and enterprise level.

    2 landed

    Standing coverage: AI-SDLC

  5. 5.2

    A repeatable methodology underpins how we sell and deliver. All teams take a consistent story to market, and each engagement deepens the playbook for the next.

    2 landed

    Standing coverage: Deciding Table Stakes

  6. A.2

    Financial performance, productivity gains and AI investment impact are tracked and forecasted at company, division and programme level, with forward scenarios available – so we can maximise the impact of our resources.

    2 landed

    Standing coverage: Cost redux on tokens

  7. 1.1

    Everyone at Quantium uses AI to do their job better, every day, with the best enterprise-wide tools.

    1 landed

    Standing coverage: Cost redux on tokens · Citizen Developers and Org Slop

  8. 5.4

    AI transformation engagements deliver material, measurable value for clients. We define and agree impact upfront, build the muscle to have that conversation consistently, and grow our share of transformations in our markets.

    1 landed

    Standing coverage: ROI

Handovers double as evidence for C.3's own KPI — insights documented as having influenced a product decision, an internal transformation initiative or a client conversation. That is the plan's wording, with a date and a counterpart attached, which is stronger than any number the lab could invent for itself.

The mix · this quarter

What kind of work it was.

Stage says where something is and tier says how far to trust it. Neither says what kind of thing it is, and a quarter of maintenance and a quarter of paired experiments read identically on every other view here. Rows overlap — work is several of these at once — so they do not sum and there is no total.

Wet lab16 of 27 · 10 liveThe full thing: preregistered, paired, with a kill condition declared before anybody saw a result. The expensive kind, and the only kind that produces a tested claim.
Alongside a team7 of 27 · 5 liveWork with a counterpart somewhere else in the firm — procurement, platform, a delivery squad. A deliverable and a handover rather than a hypothesis.
Somebody else's, in the plan7 of 27 · 5 liveThe plan names the lab as a dependency on an objective held elsewhere. Real work, and the credit belongs to the objective's owner.
Ours in the plan9 of 27 · 4 liveServes an objective the lab is directly accountable for. If this slips, the plan slips and it is the lab's name on it.
Rogue4 of 27 · 3 liveNobody asked for it and it serves no objective in the plan. The lab put its hand up. This is meant to exist, and it is meant to be a minority.
Keeping it running3 of 27 · 3 liveMaintenance: infrastructure, enablement and access, with nothing being measured. Necessary, invisible in every other view, and the thing that quietly eats a quarter.
Replication1 of 27 · 1 liveRunning somebody else's result again, ours included. Unglamorous and the reason anything here can be trusted twice.
Quick test3 of 27 · 0 liveA probe. Half a day to three days, one person, cheap enough to be wrong. Most of these are supposed to die.

Every mark is computed from the record rather than applied by hand, so the mix cannot be improved by relabelling. The place to change it is the work itself.

Instrument health · weekly

The system applies its own decay model to itself.

Cavendish tells the firm when its knowledge is stale. It also has to know when it is degrading — the observability nobody would accept omitting from a client system and everybody omits from their own.

Source pool diversity

38% ML research

largest single community share; drift +4pts this quarter

Source yield distribution

23 of 61

sources that have ever produced an elected or validated result

Claim extraction quality

0.87

sampled human agreement with dalton-0.4 output, n=120

Extraction cost per signal

$0.041

against a $0.06 per-source cap

Cluster separability

6%

share of clusters failing the separability test — the blob failure mode

Ranking quality

2.6×

acceptance of top-ranked candidates vs a random sample from the pool

Diff signal-to-noise

0.71

share of claim changes attributable to source changes; two fields carry a degraded marker

Dream journal acceptance

31%

rolling four weeks; a feed perceived as noise is abandoned

Cost drift

1 field

voice-and-vision 18% over cap after the latency bench

Human decision load

17 / wk

against the ~20 target

Human vs detector origin

67%

elected fields originating from people rather than detectors. If this runs heavily human, the automation is a filing system, not a discovery engine — still valuable, worth knowing.

Human drop share

60%

signals with at least one human drop attached. The human route stays first-class permanently.

Experiments requested from outside

9

the strongest measure available from the first cycle. It needs no instrumentation and measures whether the firm finds the lab useful.

16 graveyard entries This week’s dream journal Annual adversarial review scheduled November, under Asilomar.