cavendish

The Wire · continuously watched, ranked, cited

Keep up without reading everything.

External signal and the lab's own output in one stream, ranked by velocity rather than volume, with a synthesis composed over whatever you have in scope. Configure it once and it is remembered — including whose drops you want to follow.

In the wire
0
This fortnight
0
Lab output
0
First time here. Nothing is marked new — everything would be, and that is not information.

Composed for you · 439 of 439 in scope

Rendered from the graph. Nothing here is written by hand.

What moved in the field

Hardware vendor: tens of logical qubits at 1e-4 error are now demonstrated, not projected.1

It did not move alone. Agentic Memory System, Eval Harnesses and Personal Wiki all took a step this fortnight, and they are at three different stages — which is the more useful thing about them than the fact that they are busy.23

Agentic Memory System is the most settled of them: tested and citable, inside a year. Long-running agents need a memory layer that separates what happened from what was learned; raw vector recall over transcripts is the wrong abstraction and will be replaced by structured episodic stores with summarised recall. The lab's own run: above ~50k tokens of history, retrieval latency — not storage — is the dominant cost of an agent memory layer.45

Eval Harnesses sits between them — tested and citable, inside a year. An eval harness is five instruments, not one — capability, task, release canary, judge calibration, cost and latency — and the one that pays first is the canary; a judge without a human agreement score is not an eval, it is an opinion with a decimal point. The lab's own run: judge reliability is a property of the rubric and references, not the judge model, and uncalibrated judges favour their own family.26

Personal Wiki is the least settled: in somebody's harness, not yet a finding, one to three years out. The durable output of an agent's experience is a human-readable wiki of skills and knowledge that the next agent — and the next person — can read and revise, not a vector store or a weight update; compiling experience into that form is the consolidation step the memory field left hand-rolled. Anthropic: skill storage and retrieval are becoming a first-party platform primitive.37

Something also stopped being true. Vector-store memory as the agent's long-term memory is in the graveyard: the memory bake-off answered the question with a different design: structured episodic store with summarised recall beat raw vector recall on precision and latency above 50k tokens of…8

The next twelve months in Agentic Memory System turn on one thing: A consolidation step that a delivery team can configure without a research engineer — currently the step we hand-roll every time. Nothing in this fortnight moved it.

Moving

Today on the wire

Composed from the records underneath. Every clause carries its sources.
48 @ 1e-4
logical qubits
ConvergingQuantum Compute General Tried12 Aug 2026· 3 min read

Tens of logical qubits at 1e-4 error are now demonstrated, not projected

7 sources arrived at the same thing without seeing each other, which is a strength signal rather than a duplicate.

The quick read

Largest error-corrected demonstration to date; numbers were independently checked by a university group within the month. This is the datapoint on the tracker. Still 20× short on count and 100× short on error rate for cryptographic relevance.1

It is not one source: 4 others in the window say adjacent things — client question, paper, job posting — and 3 of them were seen independently more than once.234

Clients have already raised it: 2 band-3 signals come from the engagement record rather than from the field moving.25

Nothing here is a lab position yet. The field sits at signal tier with confidence 55%, and cannot be cited as a finding.

0.71
judge–human kappa (rubric)
MeasuredEval Harnesses Tested31 Aug 2026· 4 min read

Eval engineering is becoming a market job title (inference, from hiring)

Eval Harnesses moved from argument to number. Here is what was run, what it cost, and what it does not prove.

The quick read

Twelve-person panel scored 600 items across four task evals and two open-ended quality sets. Four judge configurations compared. Rubric-with-references reached kappa 0.71; free-form judging stayed under 0.4 regardless of judge model, and the free-form judge preferred the newer same-family model by 11 points against the panel.1

It is not one source: 4 others in the window say adjacent things — job posting, finding, repository — and 2 of them were seen independently more than once.234

Clients have already raised it: 1 band-3 signal comes from the engagement record rather than from the field moving.2

The lab has published on it: Every LLM judge ships with a human agreement score or does not ship — tested tier, strength strong, review by 2026-12-08.5

Naming eventPersonal Wiki Tried28 Aug 2026· 4 min read

The wiki cuts repeat-task time and produces wrong pages agents trust; both effects are large

The term is doing work it was not doing a quarter ago — 10 sightings across 10 sources, and the vocabulary is settling before the capability has.

The quick read

Demand-band signal: the word 'skills' in the sense of reusable agent capability appeared in eleven talk titles. Cross-band ignition with the August primitive release and the paper.1

It is not one source: 4 others in the window say adjacent things — finding, job posting, release — and 1 of them were seen independently more than once.234

Clients have already raised it: 2 band-3 signals come from the engagement record rather than from the field moving.51

Nothing here is a lab position yet. The field sits at tried tier with confidence 53%, and cannot be cited as a finding.

$12 / merged PR
price
Naming eventDark Harness Tried24 Aug 2026· 4 min read

Unattended software production is commercially ready (vendor claim, priced per merged PR)

The term is doing work it was not doing a quarter ago — 9 sightings across 8 sources, and the vocabulary is settling before the capability has.

The quick read

Pricing by merged PR rather than by seat — a business-model tell that the vendors expect volume from unattended runs. Neither publishes rework or escape rates. Naming event: 'factory' replaced 'copilot' in both.1

It is not one source: 4 others in the window say adjacent things — finding, job posting, client question — and 0 of them were seen independently more than once.234

Clients have already raised it: 2 band-3 signals come from the engagement record rather than from the field moving.56

Nothing here is a lab position yet. The field sits at tried tier with confidence 47%, and cannot be cited as a finding.

41% vs 18%
survey vs telemetry
MeasuredROI Tested28 Aug 2026· 4 min read

Self-reported coding-assistant productivity gains overstate telemetry-measured gains by two to three times

ROI moved from argument to number. Here is what was run, what it cost, and what it does not prove.

The quick read

Versioned measurement run over git and CI metadata plus assistant session logs, twelve-month baseline. Median PR cycle time fell 23%, review load rose 14%, net delivery gain 10–18%. The same engineers surveyed the same fortnight reported 41%. Query stored; re-runnable in November.1

It is not one source: 4 others in the window say adjacent things — post, job posting, dataset — and 3 of them were seen independently more than once.234

Clients have already raised it: 2 band-3 signals come from the engagement record rather than from the field moving.23

The lab has published on it: Measure ROI from telemetry and cycle time, not surveys — tested tier, strength moderate, review by 2026-12-17.5

Everything in scope

The raw records, ranked. Signal sits beside the lab's own output.
Post·band 3×2.0Signal

'Consulting moats were never capabilities'

Capability parity is irrelevant because consulting is won on relationships and distribution.

LinkedIn · A former big-four partner12 Aug 2026
? 422 days old · 4 independent sightings · demand signal — a client asked
Standing answer×2.0Tested

What does inference actually cost right now?

Blended across our six instrumented patterns, 1 September: $0.9–1.6 per thousand agent turns on the open-weight tier, $4–11 on frontier mid-tier. Auto-refreshed from the ledger; not yet…

Cost redux on tokens1 Sep 2026
4landed this fortnight · 4 independent sightings · measured result
RecommendationTested

Measure ROI from telemetry and cycle time, not surveys

Do not report AI productivity gains from self-reported surveys. Instrument the workflow and measure cycle time, throughput and rework before and after.

ROI19 Aug 2026
415 days old · 4 independent sightings · measured result
RecommendationTested

Put backpressure on agent fan-out before you put it on the model

Bound the number of in-flight sub-agents and tool calls at the orchestrator with a queue that applies backpressure. Model-side rate limits are the wrong place to discover you have a…

Backpressure25 Aug 2026
3landed this fortnight · 3 independent sightings · measured result
Standing answerTested

Which model for structured extraction?

Qwen 3.5 32B through the gateway's extraction tier for flat and moderately nested schemas; Claude Sonnet 5 for schemas over ~40 fields or with cross-field constraints. As of 27 August.

Open Weight Models27 Aug 2026
4landed this fortnight · 4 independent sightings · measured result
Position×2.0Assessed

Table stakes, not moats: what a right to play costs in 2026

Most of what consultancies sold as AI differentiation in 2024 is now table stakes: a gateway, an eval harness, a calibrated judge, a memory default, delegated auth. The moat is not in…

Deciding Table Stakes25 Aug 2026
4landed this fortnight · 4 independent sightings
6
Recommendation×2.0Tested

Every LLM judge ships with a human agreement score or does not ship

Do not report an eval number produced by an LLM judge unless the judge has a published agreement rate against a human panel on the same task. Below 0.8 Cohen's kappa the judge is not a…

Eval Harnesses10 Aug 2026
424 days old · 4 independent sightings · measured result
Standing answerTested

Which memory layer should a new agent use?

Under ten sessions of history: none, just the context window. Over ten: the lab's episodic store with summarised recall via the memory adapter. Not a vendor memory product yet.

Agentic Memory System21 Aug 2026
4landed this fortnight · 4 independent sightings · measured result
PositionAssessed

The direction of agentic development

Agentic development is real, the throughput gain is real, and the bottleneck has moved from writing code to specifying and verifying it. The firms that win will be the ones that…

AI-SDLC18 Aug 2026
416 days old · 4 independent sightings

Showing the top 60 of 439. Narrow the configuration rather than scrolling.