cavendish
TestedConvergedgate · SkillsNow · 0–12 months×6 sightings

Eval Harnesses

An eval harness is five instruments, not one — capability, task, release canary, judge calibration, cost and latency — and the one that pays first is the canary; a judge without a human agreement score is not an eval, it is an opinion with a decimal point.

Experiment run, measured result. The only tier that becomes a recommendation.

Join with…

Confidence

84%human-committed

Expiry

42duntil review · 15 Oct 2026

Lead time

+4moahead of mainstream awareness

Ownership

MLMichelle Lamweekly cadence

Where it is

unattributed

Turing was built first, on conviction, before the graph existed, and it is the field that has converged. The release canary has fired on three foundation-model releases and caught a nine-point regression on structured extraction within six hours of one of them; no client pipeline caught it first. Judge calibration against a human panel put the best judge at kappa 0.71 on rubric-scored task evals and 0.38 on open-ended quality, and showed an uncalibrated judge over-scoring the newer model by eleven points — the finding that sent 'judge without a human score' to the graveyard. Public benchmarks saturate inside nine months and stop discriminating; task evals on the firm's own data are the only durable instrument. The gate is skills: the tooling is open and adequate, foundation labs now ship eval products, and the constraint is that eval engineering is a job title the market is only now creating. Competitors are hiring for it.

Written by hand and carrying nobody's name. Editing it puts yours on it.

Why a Quantium decision hinges on it

unattributed

Every recommendation the lab publishes rests on an eval; every client asks 'how do we know the new model didn't break it'. The canary is the answer to the second and the calibration score is what keeps the first honest. Evals are also the handover artefact — the thing the firm keeps if the lab stops — and the table stake most likely to be a brief moat before competitors staff up. Health clients face a software-as-medical-device question where a calibrated eval is the difference between a regulated product and an unregulated one.

Written by hand and carrying nobody's name. Editing it puts yours on it.

What it actually is

composed from the records

An eval harness is five instruments, not one — capability, task, release canary, judge calibration, cost and latency — and the one that pays first is the canary; a judge without a human agreement score is not an eval, it is an opinion with a decimal point. That is the lab's one-line position on it, which is not the same as an explanation.

The shape the field is converging on, from the most authoritative source in it: LLM judges prefer same-family outputs by 6–12 points unless anchored to references.signal

This is the section a page most needs a person for, and the one composition is worst at. Nobody has written the plain-language version — what the idea is, in words that assume nothing — and it is the first thing a reader who has never met the term needs.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

Why now

composed from the records

The lab opened this field on 2025-09-15 and it reached mainstream awareness on 2026-01-15. The gap between those two dates is the lead time the lab is measured on.

What shipped: Foundation lab ships a hosted evals product with a frontier model as default grader (OpenAI, 2025-10-02). Tooling arriving is what moves a field from argument to something a team could try.signal

Demand is rising on it rather than steady — “Show us the control evidence for this AI system, mapped to our framework.” — which is the difference between a field worth watching and one worth doing something about.demand

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

What it changes in a system

composed from the records

What changes, concretely: Release canary v2 fired on three foundation releases and caught a 9-point regression on structured extraction within six hours; no client pipeline caught it first (x-eval-canary).signalsignal

There is already a shipped default — judge.humanAgreement = required · judge.rubric = reference-anchored · canary.trigger = model-release · skill: cavendish/judge-calibration — so a team adopting this is changing a setting rather than starting a project.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

What is in the way

composed from the records

The binding constraint is skills: it is operable and nobody is staffed to run it. Everything upstream of that is solved and everything downstream of it is waiting.

Workforce readiness is medium: Delivery can run the canary; writing and calibrating a rubric still needs the lab. Agent-estimated. A recommendation needing skills the firm does not hold is an aspiration rather than an action, and it routes to the enablement agenda instead of the delivery one.

Something here has already been killed: LLM-as-judge without human calibration, on against a five-person human panel on 600 paired items, an uncalibrated frontier judge reached cohen's κ of 0.41 to 0.58 across three task families, below the 0.8 the hypothesis required and below the 0.6 kill line on two of them.graveyard

The argued case against it is the red team's, further down this page, and it is deliberately one-sided — this section is what stands in the way mechanically, not what somebody thinks of it.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

2 of 6 explanatory sections are written; the rest are composed until somebody takes them.

Business priority

Plan critical
  • A control with a test behind it is the difference between an assurance claim and a policy statement.

    TB Travis Boastcommitted
  • Turing is the handover artifact — the thing a delivery team keeps after the lab leaves.

    AW Adam Witanowskicommitted

Clients are asking

  • “Show us the control evidence for this AI system, mapped to our framework.”6 engagements · $1M–5M · rising

Priority orders what you see. It never changes what the evidence says — a plan-critical field with nothing tested is still signal tier.

Field attributes

StateConverged
GateSkills · operable, not yet staffed
OriginConviction
Measurablefull
Written forpractice
Reach · TLPClient-safe · TLP:GREEN
Horizonnow
Opened15 Sep 2025
Mainstream15 Jan 2026
Last validated31 Aug 2026
Sightings6

What people have written

Write one

Nothing yet. The person who knows a claim is wrong is usually not the person who wrote it.

A note never travels further than the thing it is written on.

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01Release canary v2 fired on three foundation releases and caught a 9-point regression on structured extraction within six hours; no client pipeline caught it first (x-eval-canary).
  • 02Judge calibration against a twelve-person human panel: best judge kappa 0.71 on rubric-scored task evals, 0.38 on open-ended quality; an uncalibrated judge over-scored the newer same-family model by 11 points (x-judge-calibration).
  • 03Rubric-with-references judging raised agreement from 0.58 to 0.71 on the same tasks; free-form judging did not improve with a stronger judge model.
  • 04Public benchmarks in our watchlist saturated (>95% top score) a median of nine months after release; task evals on our own data have not.
What is hype
  • 01'We use LLM-as-judge.' Without a human agreement score it measures the judge's preferences, including its preference for its own family.
  • 02Leaderboard positions. A model's rank on a saturated public benchmark says nothing about its behaviour on a claims-triage task.
  • 03Eval platforms as a product category. The platform is the easy part; the rubric, the references and the panel are the work.
What would have to be true
  • 01Judge agreement above 0.6 on open-ended quality, which no judge configuration we have tried reaches; until then open-ended evals are human-scored.
  • 02A canary that fires on gateway routing changes as well as model releases; today it fires on release announcements only.
  • 03Eval engineering as a delivery skill, not a lab skill: a rubric a sector team can write and calibrate without Turing's owner in the room.
What we would do
  • 01Keep r-judge-calibration as a hard rule: every judge ships with a human agreement score or does not ship.
  • 02Wire the canary to the gateway so it fires on any model or routing change, not just announcements; this is the Argus teeth.
  • 03Move the field toward dissolution by absorption: the tooling is practice now, and the lab's remaining job is calibration and the skills gap.

Signals · 10 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
0.71
judge–human kappa (rubric)
Finding·band 1Tested

Judge calibration against a human panel: kappa 0.71 on rubric tasks, 0.38 open-ended; uncalibrated judge over-scored the newer model by 11 points

Twelve-person panel scored 600 items across four task evals and two open-ended quality sets. Four judge configurations compared. Rubric-with-references reached kappa 0.71; free-form judging stayed under 0.4 regardless of judge model, and the free-form judge preferred the newer same-family model by 11 points against the panel.

extracted claimJudge reliability is a property of the rubric and references, not the judge model, and uncalibrated judges favour their own family.
Lab · x-judge-calibration · Michelle Lam31 Aug 2026
detector · bleeding edge
41
openings
Job posting·band 3Signal

'AI Evaluation Engineer' openings in Australia double in six months

Argus hiring scan across consultancies, banks and two foundation-lab AU offices. Forty-one open roles; the title barely existed in January. Inference: the skills gate is being closed by the market, and the lab's advantage on evals has a shelf life.

SEEK / competitor careers pages11 Aug 2026
detector · demand
3 of 3 releases
regressions caught
Finding·band 1Tested

Release canary v2 caught a 9-point extraction regression within six hours of a model release

Canary fires on foundation-lab release announcements and runs the task-eval suite against the new default. On the third release it flagged a nine-point drop on structured extraction; the client pipeline running the same model noticed four days later from support tickets.

extracted claimA release canary on own-data task evals catches material regressions before client pipelines do.
Lab · x-eval-canary · Michelle Lam23 Jul 2026
detector · bleeding edge 2
0.55 → 0.70
agreement
Finding·band 1Tried

Logged from Claude Code: switching the judge to a rubric with reference answers lifted agreement on the merchandising eval from 0.55 to 0.7

Product engineer re-ran the retail pilot's eval with reference-anchored rubrics and a small human sample. Agreement moved from 0.55 to 0.70. One eval, one person; tried tier, and consistent with the calibration experiment.

MCP · log_finding · Oliver Vu19 Jun 2026
OVdropped
28k
stars
Repository·band 2Tried

Open eval framework passes 28k stars; adds calibration and human-panel modules

The framework Turing is built on. The calibration module landed after we filed the issue; the tooling gate is closed, which is why the field's gate is skills.

github.com19 May 2026
MLdropped 2
6–12 pts
same-family preference
Paper·band 1Signal

Judges prefer their own: self- and family-preference bias in LLM evaluation

Measures judge preference for outputs from the judge's own model family across five families. Finds a consistent 6–12 point preference that survives prompt-level debiasing and disappears only with reference-anchored rubrics. Our calibration reproduced it.

arxiv.org · Iqbal, Sørensen et al.12 Mar 2026
detector · bleeding edge 3
9 months
median time to saturation
Benchmark·band 2Signal

Watchlist benchmark saturation: median nine months from release to a >95% top score

Argus tracking of fourteen public benchmarks in the capability matrix. Median time from release to saturation was nine months; three saturated within four. The reason the capability matrix is scored on our own task evals.

Public leaderboards14 Jan 2026
detector · early adoption 2
Post·band 2Signal

'Your eval is measuring your prompt'

Argues that task evals overfit to the prompt and rubric that produced them and that a good score is a tautology. Partly right — it is why the panel exists — and kept as the sharpest critique of the field's own instrument.

Substack · An applied-ML voice9 Jan 2026
?dropped 3
Client question·band 3Signal

'How do we know the new model version didn't break our pipeline?'

Asked by a head of model risk after a provider deprecated a model mid-engagement. Logged as unanswered at the time; it is now the canary's job description and the most-asked question in the banking pipeline.

Engel · banking engagement10 Dec 2025
AHdropped 5
Release·band 1Signal

Foundation lab ships a hosted evals product with a frontier model as default grader

Managed eval runs with the vendor's own model as the default judge and no calibration step. Convenient, and the design choice our calibration result argues against. Retained as the source of the 'judge is good enough' claim.

OpenAI2 Oct 2025
detector · bleeding edge 3
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 5 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

LLM-judge agreement with humans is task-dependent: around 0.7 kappa on rubric-scored tasks and below 0.4 on open-ended quality.

Testedc-eval-harnesses-1dalton-0.431 Aug 2026Lab · x-judge-calibration, MCP · log_finding
88%

Release canaries catch material regressions on structured tasks within hours of a model release; no client pipeline has caught one first.

Testedc-eval-harnesses-2dalton-0.431 Aug 2026Lab · x-eval-canary, Engel · banking engagement
84%

Uncalibrated judges systematically favour newer and same-family models.

Testedc-eval-harnesses-3dalton-0.44 Jun 2026Lab · x-judge-calibration, arxiv.org
76%

Public benchmarks saturate within about nine months of release and stop discriminating; task evals on own data are the only durable instrument.

Assessedc-eval-harnesses-4dalton-0.321 Jan 2026Public leaderboards, Substack
70%

Eval engineering is becoming a job title; the skills gate is closing on the market's timetable, not the lab's.

Assessedc-eval-harnesses-6dalton-0.431 Aug 2026SEEK / competitor careers pages, github.com
60%

A frontier model as judge is reliable enough without human calibration for most enterprise uses.

Testedc-eval-harnesses-5dalton-0.38 Oct 2025OpenAI, Lab · x-judge-calibration
18%

Position history · the diff is the product

4 validation runs against a fixed brief. Confidence 50% → 84%.

runs compare claim sets, never prose
What we said · run 4

Both experiments concluded. Canary caught a real regression before any client did; judge calibration gives 0.71 on rubric tasks and 0.38 on open-ended. Unchecked judge to the graveyard; recommendation and standing answer current. Field converged; skills is the gate.

84%
Changed since run 3
  • LLM-judge agreement with humans is task-dependent: around 0.7 kappa on rubric-scored tasks and below 0.4 on open-ended quality.
  • Release canaries catch material regressions on structured tasks within hours of a model release; no client pipeline has caught one first.
  • Uncalibrated judges systematically favour newer and same-family models.
  • Eval engineering is becoming a job title; the skills gate is closing on the market's timetable, not the lab's.
  • c-eval-harnesses-5 ↓ 0.30 → 0.18
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · ML
high

Every recommendation and every canary rests on it; it is also the handover artefact.

Timeline

committed · ML
0–18mo

In production. The remaining timeline is the skills gap in delivery.

Cost

committed · ML
medium

The human panel is the expense: twelve people, two days per calibration cycle.

Demand

committed · AH
high

'Did the new model break it' is the most-asked question in banking steering committees.

TAM

agent-estimated
$100M–1B

Agent-estimated from eval-tooling and AI-QA spend. Small market; the value is in avoided regressions. Uncommitted.

Workforce readiness

agent-estimated
medium

Delivery can run the canary; writing and calibrating a rubric still needs the lab. Agent-estimated.

Cost of being wrong

committed · TB
high

An uncalibrated judge passing a regressed model into a claims process is an assurance failure with a name on it.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Banking
relevant

Model changes in a servicing or credit process need a documented regression check; the canary is that check.

Mechanism · Canary suite on the bank's task evals, fired on every model or routing change; calibration score in the model-risk file.

AH committed by Amber Hallcommitted · AH
Insurance
relevant

Claims triage evals are the closest thing to a conduct test for an agent; the judge calibration score is what makes them defensible.

Mechanism · Rubric-with-references judge, human panel drawn from claims assessors, agreement score published with the eval. Agent draft.

Agent draft · awaiting a sector owneragent-estimated
Retail & FMCG
relevant

Merchandising and product-classification evals already run in the Woolworths pilots; the canary caught a regression there first.

Mechanism · Task evals on category data; canary on the pilot's model endpoint.

DB committed by Dillon Blakecommitted · DB
Health
relevant

A calibrated eval is what a TGA software-as-medical-device conversation needs; an uncalibrated one is what it fears.

Mechanism · Clinician panel for calibration; eval record as part of the quality-management evidence.

SL committed by Sylvia Liucommitted · SL

Red team · the strongest case against

The strongest case against: the field has converged on measuring what is easy to measure. Rubric-scored tasks give good kappa because rubrics make humans agree with each other, not because the judge understands quality. The 0.38 on open-ended work is the honest number, and open-ended work is most of what clients want. Meanwhile the canary catches regressions on the tasks we thought to write evals for and is blind to the ones we did not. A converged field can be a field that stopped asking the hard question.

  • —Kappa on rubric tasks measures rubric quality. The judge may be learning the rubric's surface features, which is overfitting with a human agreement score attached.
  • —Three canary firings is a small sample; the base rate of material regressions per release is unknown and the false-negative rate is unmeasured.
  • —Twelve panellists from the firm are not the client's assessors; calibration may not transfer across panels.
  • —Public benchmark saturation is a claim about the benchmarks we watch; the labs' private evals may not saturate and we cannot see them.
Run by an agent briefed to argue the field is nothing — sources here are correlated, and without a deliberate adversary synthesis converges on consensus and calls it insight. Kept as a dated pass rather than overwritten. Nobody has answered it yet, and a challenge nobody answers is a disclaimer.thesis holds

Source diversity

  • ML research25%
  • Open-source tooling15%
  • Foundation labs and vendors15%
  • Competitors / Argus10%
  • Internal / Engel35%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsMLAWATTBOVAHSL

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01Is there any judge configuration that reaches 0.6 agreement on open-ended quality, or is that work permanently human-scored?
  2. 02What is the canary's false-negative rate — how many regressions has it missed that clients found?
  3. 03When does the field dissolve into practice, and what does the lab keep?

Tools in this space · 5

What you could actually buy.

Products aimed at this field, with what the lab has behind each one. Scored on six axes and never summed — “which is better” is not a question anyone has, and the constraint that decides it is named beside every assessment.

Bake-off runningDelivery teams are each picking their own eval tool. Should there be a default, and does it survive a client who will not let data leave the estate?proposed by Michelle Lam

Win condition: Declared 21 August. A tool wins on: the same eval suite running unchanged in CI and in a review surface; a self-host path that a security review would pass; agreement with human labels within five points on our calibration set; and a delivery engineer wiring it in under a day without help.

  • LangfuseLangfuseTestedRecommended

    Open-source tracing and evaluation with a self-host path that actually works. The pragmatic answer where client data cannot leave the estate.

    SaaS or self-hostMIT core, commercial cloudseen 2025-10-06

    Self-hosted in the lab's own harness.

  • PromptfoopromptfooTestedRecommended

    Config-driven eval and red-teaming that runs in CI. Not a platform, which is exactly why it survives contact with a client's build pipeline.

    Open sourceMITseen 2025-10-20

    Runs in the lab's own CI.

  • BraintrustBraintrustAssessedAssessed

    Eval, logging and prompt iteration in one loop, aimed at teams shipping continuously. Strong on the workflow, opinionated about where the data lives.

    SaaS onlyCommercialseen 2025-11-25

    No commercial relationship.

  • InspectUK AI Security InstituteSignalWatching

    A rigorous evaluation framework from a public institute, built for capability and safety evals rather than product iteration. Heavier than a delivery team needs and the right reference for what a defensible eval looks like.

    Open sourceMITseen 2026-01-06

    No commercial relationship.

  • LangSmithLangChainSignalWatching

    Tracing and evaluation tightly coupled to the LangChain ecosystem. The coupling is the whole trade: excellent inside it, awkward outside.

    SaaS or self-hostCommercialseen 2025-09-15

    No commercial relationship.

Listing is not recommending. Most of Eval Harnesses sits at signal tier — in the space, nothing behind it — and a tool only reaches tested when a run stands behind it. Vendor pricing and capability move monthly, so these carry the shortest half-life in the library, and anyone who cited one gets told when it moves.

Notes · anyone in the firm

What people have written on this.

The person who knows a claim is wrong is usually not the person who wrote it. Corrections, objections and questions are owed an answer and stay open until the field owner says what they did; context and use notes stand as they are.

Notes · 0

Anything here reaches at most client-safe — a note cannot travel further than what it is written on.

    Nothing written on this yet. The useful notes are the ones from people who are not in the lab — that is where the correction usually comes from.