Lab app · Scorecard · Q3 2026
The track record.
Lead time measures whether the lab is distinctive. Answer rate measures whether it is useful. A lab needs both, and only one of them will be asked about in a budget review.
Two clocks on this page. Lead time, prediction calibration and cost per validated recommendation are measurements over closed periods — they do not move because somebody concluded a run this afternoon, and they should not. What the lab is doing about the plan and the mix of work in flight are statements about right now, and they read the board and the library live.
11 fields reached mainstream · the running median as each crossed
share of questions the library answers well · by month
947 attributed person-days, 11% of it work that produces none
1 library items overdue for review
Lead time · per field
When we opened it versus when it went mainstream.
Negative lead time is recorded honestly. An honest negative is the only thing that makes a positive one believable.
- AI-SDLC+6mo
- SLM / Edge / Tuning+4mo
- Open Weight Models+4mo
- Agentic Memory System+4mo
- Fully Agentic QA+4mo
- Eval Harnesses+4mo
- Cost redux on tokens−3mo
- ROI−4mo
- AI Gateway−4mo
- Voice and Vision−4mo
- Quantum Encryption−20mo
Calibration
What we said versus what happened.
Predictions are dated, confidence-scored and resolvable. A perfectly calibrated lab sits on the diagonal. We are over-confident at the top and under-confident at the bottom, which is the usual shape.
Brier 0.19 across 58 resolved predictions. Deferred to a spreadsheet for the first year, as the design says.
Against the plan · last 90 days
What the lab did about the firm's objectives.
The rest of this page measures whether the lab is any good. This answers the question the firm actually asks. Every row landed on somebody else — published, concluded, or handed over — because starting things is free and a scorecard that counted activity would produce activity.
- Objectives in the plan
- 30
- The lab is on
- 11
- Something landed
- 8
- Nothing landed
- 3
Being on an objective is not the same as having done something about it, and the gap between those two rows is the only part of this worth arguing about.
- C.1
Shared AI infrastructure is built and maintained, including governed assets. Every team builds on it. No team starts from zero.
7 landedStanding coverage: Agentic Memory System · AI Gateway · Open Weight Models
- PublishedWhich model for structured extraction? 2026-08-27
- PublishedOpen-weight models for classification and extraction; frontier for agentic loops 2026-08-27
- Run concludedMemory bake-off validated2026-08-21
- PublishedWhich memory layer should a new agent use? 2026-08-21
- PublishedUse a structured episodic store with summarised recall, not a raw vector memory 2026-08-21
- PublishedRoute through a gateway you control; do not standardise on a vendor's garden 2026-08-14
- Run concludedRight model, right task validated2026-06-30
- 3.1
We have a standard approach to build, deploy and manage agents and workbenches, faster, more consistently, and to a higher quality bar than building bespoke each time.
6 landedStanding coverage: Agentic Memory System · Auth Broker · AI-SDLC
- Run concludedMemory bake-off validated2026-08-21
- PublishedWhich memory layer should a new agent use? 2026-08-21
- PublishedUse a structured episodic store with summarised recall, not a raw vector memory 2026-08-21
- PublishedThe direction of agentic development 2026-08-18
- PublishedAgentic delivery works when the spec is the artifact; do not start with the code 2026-08-18
- PublishedAgents act under delegated, scoped, expiring authority — never a service account 2026-08-07
- C.4
An AI risk, ethics, security, quality and compliance framework is operational, applied across all AI workloads – internally and for clients. Externally credible, not just internally compliant.
3 landedStanding coverage: Auth Broker · Eval Harnesses · Cyber Cold War
- PublishedEvery LLM judge ships with a human agreement score or does not ship 2026-08-10
- PublishedAgents act under delegated, scoped, expiring authority — never a service account 2026-08-07
- PublishedWhat eval tooling do we use? 2026-08-05
- 1.2
Every team has reimagined how we do our work and built the capability to keep doing so. Productivity is visible at function and enterprise level.
2 landedStanding coverage: AI-SDLC
- PublishedThe direction of agentic development 2026-08-18
- PublishedAgentic delivery works when the spec is the artifact; do not start with the code 2026-08-18
- 5.2
A repeatable methodology underpins how we sell and deliver. All teams take a consistent story to market, and each engagement deepens the playbook for the next.
2 landedStanding coverage: Deciding Table Stakes
- PublishedTable stakes, not moats: what a right to play costs in 2026 2026-08-25
- PublishedIs the vendor's 'agentic' claim real? 2026-08-14
- A.2
Financial performance, productivity gains and AI investment impact are tracked and forecasted at company, division and programme level, with forward scenarios available – so we can maximise the impact of our resources.
2 landedStanding coverage: Cost redux on tokens
- PublishedPrompt caching: use for stable prefixes over 2k tokens; expect 30–45%, not 60% 2026-08-28
- Handed overThe AI cost model Finance2026-08-22
- 1.1
Everyone at Quantium uses AI to do their job better, every day, with the best enterprise-wide tools.
1 landedStanding coverage: Cost redux on tokens · Citizen Developers and Org Slop
- PublishedPrompt caching: use for stable prefixes over 2k tokens; expect 30–45%, not 60% 2026-08-28
- 5.4
AI transformation engagements deliver material, measurable value for clients. We define and agree impact upfront, build the muscle to have that conversation consistently, and grow our share of transformations in our markets.
1 landedStanding coverage: ROI
- PublishedMeasure ROI from telemetry and cycle time, not surveys 2026-08-19
Handovers double as evidence for C.3's own KPI — insights documented as having influenced a product decision, an internal transformation initiative or a client conversation. That is the plan's wording, with a date and a counterpart attached, which is stronger than any number the lab could invent for itself.
The mix · this quarter
What kind of work it was.
Stage says where something is and tier says how far to trust it. Neither says what kind of thing it is, and a quarter of maintenance and a quarter of paired experiments read identically on every other view here. Rows overlap — work is several of these at once — so they do not sum and there is no total.
Every mark is computed from the record rather than applied by hand, so the mix cannot be improved by relabelling. The place to change it is the work itself.
Instrument health · weekly
The system applies its own decay model to itself.
Cavendish tells the firm when its knowledge is stale. It also has to know when it is degrading — the observability nobody would accept omitting from a client system and everybody omits from their own.
Source pool diversity
largest single community share; drift +4pts this quarter
Source yield distribution
sources that have ever produced an elected or validated result
Claim extraction quality
sampled human agreement with dalton-0.4 output, n=120
Extraction cost per signal
against a $0.06 per-source cap
Cluster separability
share of clusters failing the separability test — the blob failure mode
Ranking quality
acceptance of top-ranked candidates vs a random sample from the pool
Diff signal-to-noise
share of claim changes attributable to source changes; two fields carry a degraded marker
Dream journal acceptance
rolling four weeks; a feed perceived as noise is abandoned
Cost drift
voice-and-vision 18% over cap after the latency bench
Human decision load
against the ~20 target
Human vs detector origin
elected fields originating from people rather than detectors. If this runs heavily human, the automation is a filing system, not a discovery engine — still valuable, worth knowing.
Human drop share
signals with at least one human drop attached. The human route stays first-class permanently.
Experiments requested from outside
the strongest measure available from the first cycle. It needs no instrumentation and measures whether the firm finds the lab useful.