Eval Harnesses
An eval harness is five instruments, not one — capability, task, release canary, judge calibration, cost and latency — and the one that pays first is the canary; a judge without a human agreement score is not an eval, it is an opinion with a decimal point.
Experiment run, measured result. The only tier that becomes a recommendation.
Confidence
84%human-committedExpiry
42duntil review · 15 Oct 2026Lead time
+4moahead of mainstream awarenessOwnership
MLMichelle Lamweekly cadenceWhere it is
unattributedTuring was built first, on conviction, before the graph existed, and it is the field that has converged. The release canary has fired on three foundation-model releases and caught a nine-point regression on structured extraction within six hours of one of them; no client pipeline caught it first. Judge calibration against a human panel put the best judge at kappa 0.71 on rubric-scored task evals and 0.38 on open-ended quality, and showed an uncalibrated judge over-scoring the newer model by eleven points — the finding that sent 'judge without a human score' to the graveyard. Public benchmarks saturate inside nine months and stop discriminating; task evals on the firm's own data are the only durable instrument. The gate is skills: the tooling is open and adequate, foundation labs now ship eval products, and the constraint is that eval engineering is a job title the market is only now creating. Competitors are hiring for it.
Why a Quantium decision hinges on it
unattributedEvery recommendation the lab publishes rests on an eval; every client asks 'how do we know the new model didn't break it'. The canary is the answer to the second and the calibration score is what keeps the first honest. Evals are also the handover artefact — the thing the firm keeps if the lab stops — and the table stake most likely to be a brief moat before competitors staff up. Health clients face a software-as-medical-device question where a calibrated eval is the difference between a regulated product and an unregulated one.
What it actually is
composed from the recordsAn eval harness is five instruments, not one — capability, task, release canary, judge calibration, cost and latency — and the one that pays first is the canary; a judge without a human agreement score is not an eval, it is an opinion with a decimal point. That is the lab's one-line position on it, which is not the same as an explanation.
The shape the field is converging on, from the most authoritative source in it: LLM judges prefer same-family outputs by 6–12 points unless anchored to references.signal
This is the section a page most needs a person for, and the one composition is worst at. Nobody has written the plain-language version — what the idea is, in words that assume nothing — and it is the first thing a reader who has never met the term needs.
Why now
composed from the recordsThe lab opened this field on 2025-09-15 and it reached mainstream awareness on 2026-01-15. The gap between those two dates is the lead time the lab is measured on.
What shipped: Foundation lab ships a hosted evals product with a frontier model as default grader (OpenAI, 2025-10-02). Tooling arriving is what moves a field from argument to something a team could try.signal
Demand is rising on it rather than steady — “Show us the control evidence for this AI system, mapped to our framework.” — which is the difference between a field worth watching and one worth doing something about.demand
What it changes in a system
composed from the recordsWhat changes, concretely: Release canary v2 fired on three foundation releases and caught a 9-point regression on structured extraction within six hours; no client pipeline caught it first (x-eval-canary).signalsignal
There is already a shipped default — judge.humanAgreement = required · judge.rubric = reference-anchored · canary.trigger = model-release · skill: cavendish/judge-calibration — so a team adopting this is changing a setting rather than starting a project.
What is in the way
composed from the recordsThe binding constraint is skills: it is operable and nobody is staffed to run it. Everything upstream of that is solved and everything downstream of it is waiting.
Workforce readiness is medium: Delivery can run the canary; writing and calibrating a rubric still needs the lab. Agent-estimated. A recommendation needing skills the firm does not hold is an aspiration rather than an action, and it routes to the enablement agenda instead of the delivery one.
Something here has already been killed: LLM-as-judge without human calibration, on against a five-person human panel on 600 paired items, an uncalibrated frontier judge reached cohen's κ of 0.41 to 0.58 across three task families, below the 0.8 the hypothesis required and below the 0.6 kill line on two of them.graveyard
The argued case against it is the red team's, further down this page, and it is deliberately one-sided — this section is what stands in the way mechanically, not what somebody thinks of it.
2 of 6 explanatory sections are written; the rest are composed until somebody takes them.
Business priority
Plan criticalClients are asking
- “Show us the control evidence for this AI system, mapped to our framework.”6 engagements · $1M–5M · rising
Priority orders what you see. It never changes what the evidence says — a plan-critical field with nothing tested is still signal tier.
Field attributes
What people have written
Write oneNothing yet. The person who knows a claim is wrong is usually not the person who wrote it.
A note never travels further than the thing it is written on.
Position
What is demonstrated, what is hype, what would have to be true.
The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.
- 01Release canary v2 fired on three foundation releases and caught a 9-point regression on structured extraction within six hours; no client pipeline caught it first (x-eval-canary).
- 02Judge calibration against a twelve-person human panel: best judge kappa 0.71 on rubric-scored task evals, 0.38 on open-ended quality; an uncalibrated judge over-scored the newer same-family model by 11 points (x-judge-calibration).
- 03Rubric-with-references judging raised agreement from 0.58 to 0.71 on the same tasks; free-form judging did not improve with a stronger judge model.
- 04Public benchmarks in our watchlist saturated (>95% top score) a median of nine months after release; task evals on our own data have not.
- 01'We use LLM-as-judge.' Without a human agreement score it measures the judge's preferences, including its preference for its own family.
- 02Leaderboard positions. A model's rank on a saturated public benchmark says nothing about its behaviour on a claims-triage task.
- 03Eval platforms as a product category. The platform is the easy part; the rubric, the references and the panel are the work.
- 01Judge agreement above 0.6 on open-ended quality, which no judge configuration we have tried reaches; until then open-ended evals are human-scored.
- 02A canary that fires on gateway routing changes as well as model releases; today it fires on release announcements only.
- 03Eval engineering as a delivery skill, not a lab skill: a rubric a sector team can write and calibrate without Turing's owner in the room.
- 01Keep r-judge-calibration as a hard rule: every judge ships with a human agreement score or does not ship.
- 02Wire the canary to the gateway so it fires on any model or routing change, not just announcements; this is the Argus teeth.
- 03Move the field toward dissolution by absorption: the tooling is practice now, and the lab's remaining job is calibration and the skills gap.
Signals · 10 in this cluster
What the cluster is made of.
Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

Judge calibration against a human panel: kappa 0.71 on rubric tasks, 0.38 open-ended; uncalibrated judge over-scored the newer model by 11 points
Twelve-person panel scored 600 items across four task evals and two open-ended quality sets. Four judge configurations compared. Rubric-with-references reached kappa 0.71; free-form judging stayed under 0.4 regardless of judge model, and the free-form judge preferred the newer same-family model by 11 points against the panel.
extracted claimJudge reliability is a property of the rubric and references, not the judge model, and uncalibrated judges favour their own family.

'AI Evaluation Engineer' openings in Australia double in six months
Argus hiring scan across consultancies, banks and two foundation-lab AU offices. Forty-one open roles; the title barely existed in January. Inference: the skills gate is being closed by the market, and the lab's advantage on evals has a shelf life.

Release canary v2 caught a 9-point extraction regression within six hours of a model release
Canary fires on foundation-lab release announcements and runs the task-eval suite against the new default. On the third release it flagged a nine-point drop on structured extraction; the client pipeline running the same model noticed four days later from support tickets.
extracted claimA release canary on own-data task evals catches material regressions before client pipelines do.

Logged from Claude Code: switching the judge to a rubric with reference answers lifted agreement on the merchandising eval from 0.55 to 0.7
Product engineer re-ran the retail pilot's eval with reference-anchored rubrics and a small human sample. Agreement moved from 0.55 to 0.70. One eval, one person; tried tier, and consistent with the calibration experiment.


Judges prefer their own: self- and family-preference bias in LLM evaluation
Measures judge preference for outputs from the judge's own model family across five families. Finds a consistent 6–12 point preference that survives prompt-level debiasing and disappears only with reference-anchored rubrics. Our calibration reproduced it.

Watchlist benchmark saturation: median nine months from release to a >95% top score
Argus tracking of fourteen public benchmarks in the capability matrix. Median time from release to saturation was nine months; three saturated within four. The reason the capability matrix is scored on our own task evals.

'Your eval is measuring your prompt'
Argues that task evals overfit to the prompt and rubric that produced them and that a good score is a tautology. Partly right — it is why the panel exists — and kept as the sharpest critique of the field's own instrument.

'How do we know the new model version didn't break our pipeline?'
Asked by a head of model risk after a provider deprecated a model mid-engagement. Logged as unanswered at the time; it is now the canary's job description and the most-asked question in the banking pipeline.

Foundation lab ships a hosted evals product with a frontier model as default grader
Managed eval runs with the vendor's own model as the default judge and no calibration step. Convenient, and the design choice our calibration result argues against. Retained as the source of the 'judge is good enough' claim.
Claims · 5 supporting, 1 refuting
The atoms.
A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.
LLM-judge agreement with humans is task-dependent: around 0.7 kappa on rubric-scored tasks and below 0.4 on open-ended quality.
Release canaries catch material regressions on structured tasks within hours of a model release; no client pipeline has caught one first.
Uncalibrated judges systematically favour newer and same-family models.
Public benchmarks saturate within about nine months of release and stop discriminating; task evals on own data are the only durable instrument.
Eval engineering is becoming a job title; the skills gate is closing on the market's timetable, not the lab's.
A frontier model as judge is reliable enough without human calibration for most enterprise uses.
Position history · the diff is the product
4 validation runs against a fixed brief. Confidence 50% → 84%.
Both experiments concluded. Canary caught a real regression before any client did; judge calibration gives 0.71 on rubric tasks and 0.38 on open-ended. Unchecked judge to the graveyard; recommendation and standing answer current. Field converged; skills is the gate.
- LLM-judge agreement with humans is task-dependent: around 0.7 kappa on rubric-scored tasks and below 0.4 on open-ended quality.
- Release canaries catch material regressions on structured tasks within hours of a model release; no client pipeline has caught one first.
- Uncalibrated judges systematically favour newer and same-family models.
- Eval engineering is becoming a job title; the skills gate is closing on the market's timetable, not the lab's.
- c-eval-harnesses-5 ↓ 0.30 → 0.18
Scoring · ordinal bands
Agents propose. A named human commits.
Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.
Impact
committed · MLEvery recommendation and every canary rests on it; it is also the handover artefact.
Timeline
committed · MLIn production. The remaining timeline is the skills gap in delivery.
Cost
committed · MLThe human panel is the expense: twelve people, two days per calibration cycle.
Demand
committed · AH'Did the new model break it' is the most-asked question in banking steering committees.
TAM
agent-estimatedAgent-estimated from eval-tooling and AI-QA spend. Small market; the value is in avoided regressions. Uncommitted.
Workforce readiness
agent-estimatedDelivery can run the canary; writing and calibrating a rubric still needs the lab. Agent-estimated.
Cost of being wrong
committed · TBAn uncalibrated judge passing a regressed model into a claims process is an assurance failure with a name on it.
Relevance · per vertical
Why it matters here, or explicitly does not.
Ranking is per vertical, not global. Sector owners commit notes against agent drafts.
Model changes in a servicing or credit process need a documented regression check; the canary is that check.
Mechanism · Canary suite on the bank's task evals, fired on every model or routing change; calibration score in the model-risk file.
Claims triage evals are the closest thing to a conduct test for an agent; the judge calibration score is what makes them defensible.
Mechanism · Rubric-with-references judge, human panel drawn from claims assessors, agreement score published with the eval. Agent draft.
Merchandising and product-classification evals already run in the Woolworths pilots; the canary caught a regression there first.
Mechanism · Task evals on category data; canary on the pilot's model endpoint.
A calibrated eval is what a TGA software-as-medical-device conversation needs; an uncalibrated one is what it fears.
Mechanism · Clinician panel for calibration; eval record as part of the quality-management evidence.
Red team · the strongest case against
The strongest case against: the field has converged on measuring what is easy to measure. Rubric-scored tasks give good kappa because rubrics make humans agree with each other, not because the judge understands quality. The 0.38 on open-ended work is the honest number, and open-ended work is most of what clients want. Meanwhile the canary catches regressions on the tasks we thought to write evals for and is blind to the ones we did not. A converged field can be a field that stopped asking the hard question.
- —Kappa on rubric tasks measures rubric quality. The judge may be learning the rubric's surface features, which is overfitting with a human agreement score attached.
- —Three canary firings is a small sample; the base rate of material regressions per release is unknown and the false-negative rate is unmeasured.
- —Twelve panellists from the firm are not the client's assessors; calibration may not transfer across panels.
- —Public benchmark saturation is a claim about the benchmarks we watch; the labs' private evals may not saturate and we cannot see them.
Source diversity
- ML research25%
- Open-source tooling15%
- Foundation labs and vendors15%
- Competitors / Argus10%
- Internal / Engel35%
A field supported by one epistemic community is a flag, not a finding.
Cross-pollination · typed joins
Connected, not merely similar.
Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.
Evals are the table stake most likely to be a brief moat; this field's timing sets that field's.
Parity claims for open-weight models are only claims until they run on the same task evals.
The canary needs to fire on gateway routing changes, not just release announcements; the gateway is the trigger.
Agent-written tests are graded by evals; without calibrated judges the near field cannot be assessed.
Share graph
Provenance running forward.
Discovery, not accountability. No counts, no rankings, no rollups to managers.
Convergence · who else is here
- AHAmber Hall · Sector owner · Banking1 drop
- MLMichelle Lam · Analytics Lead · evals1 drop
- ?Anonymous · Anonymous drop1 drop
- OVOliver Vu · Analyst · product engineering1 drop
Several people’s drops meet here. An informal working group already exists and probably does not know it.
Lineage
What this field produced, and what it killed.
Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.
Every LLM judge ships with a human agreement score or does not ship
strength strong · 29 citations · review 8 Dec 2026
What eval tooling do we use?
strength moderate · 27 citations · review 19 Sep 2026
The judge on trial
An uncalibrated frontier LLM judge, as used in the harness through Q1, agrees with a five-person human panel at Cohen's κ of at least 0.8 across extraction, summarisation and agentic-trace grading.
The canary suite
A 120-item canary suite with calibrated judges catches at least 80% of the behaviour regressions that the last six vendor model releases introduced on our patterns, at under 15 minutes wall-clock per run.
LLM-as-judge without human calibration
“Agreed with itself, mostly.” · lived 4 months
Open questions · return to the pile
Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.
- 01Is there any judge configuration that reaches 0.6 agreement on open-ended quality, or is that work permanently human-scored?
- 02What is the canary's false-negative rate — how many regressions has it missed that clients found?
- 03When does the field dissolve into practice, and what does the lab keep?
Tools in this space · 5
What you could actually buy.
Products aimed at this field, with what the lab has behind each one. Scored on six axes and never summed — “which is better” is not a question anyone has, and the constraint that decides it is named beside every assessment.
Win condition: Declared 21 August. A tool wins on: the same eval suite running unchanged in CI and in a review surface; a self-host path that a security review would pass; agreement with human labels within five points on our calibration set; and a delivery engineer wiring it in under a day without help.
Open-source tracing and evaluation with a self-host path that actually works. The pragmatic answer where client data cannot leave the estate.
SaaS or self-hostMIT core, commercial cloudseen 2025-10-06Self-hosted in the lab's own harness.
Config-driven eval and red-teaming that runs in CI. Not a platform, which is exactly why it survives contact with a client's build pipeline.
Open sourceMITseen 2025-10-20Runs in the lab's own CI.
Eval, logging and prompt iteration in one loop, aimed at teams shipping continuously. Strong on the workflow, opinionated about where the data lives.
SaaS onlyCommercialseen 2025-11-25No commercial relationship.
A rigorous evaluation framework from a public institute, built for capability and safety evals rather than product iteration. Heavier than a delivery team needs and the right reference for what a defensible eval looks like.
Open sourceMITseen 2026-01-06No commercial relationship.
Tracing and evaluation tightly coupled to the LangChain ecosystem. The coupling is the whole trade: excellent inside it, awkward outside.
SaaS or self-hostCommercialseen 2025-09-15No commercial relationship.
Listing is not recommending. Most of Eval Harnesses sits at signal tier — in the space, nothing behind it — and a tool only reaches tested when a run stands behind it. Vendor pricing and capability move monthly, so these carry the shortest half-life in the library, and anyone who cited one gets told when it moves.
Notes · anyone in the firm
What people have written on this.
The person who knows a claim is wrong is usually not the person who wrote it. Corrections, objections and questions are owed an answer and stay open until the field owner says what they did; context and use notes stand as they are.
Notes · 0
Anything here reaches at most client-safe — a note cannot travel further than what it is written on.Nothing written on this yet. The useful notes are the ones from people who are not in the lab — that is where the correction usually comes from.