Graveyard · killed by Cavendish
What we killed.
With names and causes.
Refuted, abandoned, superseded, rejected — each with a bylined author, a plainly stated cause, and a flag if the constraint that killed it might move. Abandoned is never dressed as refuted, and every entry keeps its full lineage.
The same records, filtered to killed outcomes, with lineage intact.
- Killed
- 0
- Resurrectable
- 0
- Returned to pile
- 0
Mean lifespan 100 days. Corrections counted on the scorecard, without a target.
- Superseded† 21 Aug 2026
Vector-store memory as the agent's long-term memory
“Remembered everything, retrieved nothing.”
10 Feb 2026 — 21 Aug 2026Tested6 monthsCause · The memory bake-off answered the question with a different design: structured episodic store with summarised recall beat raw vector recall on precision and latency above 50k tokens of history.
DDDavid DiazAgentic Memory System - Abandoned† 15 Jul 2026
Realtime voice agents replace tier-1 IVR
“Hung up before the answer.”
7 Apr 2026 — 15 Jul 2026Tried3 monthsCause · The telco partner withdrew the test trunk on day 24 of a planned 40-call measurement and the pair was reassigned; 11 calls were measured, which is not enough to answer the hypothesis either way.
DYDylan Desmarchelier resurrectableVoice and Vision - Refuted† 14 Jul 2026
Self-reported productivity surveys as ROI evidence
“Everyone said it was working.”
14 Oct 2025 — 14 Jul 2026Tried9 monthsCause · In the first arm of the Nightingale run, self-reported time saved per developer correlated at 0.11 with measured cycle-time change over the same ten weeks, and over-stated it by a median factor of 3.4.
ATAndrew TranROI - Rejected† 10 Jul 2026
World-model simulation for retail demand planning
“Simulated a world that had no shelves in it.”
16 Jun 2026 — 10 Jul 2026Signal24 daysCause · Not selected: no published world model operates on tabular demand series, the proposed run would have been a conventional forecasting comparison with a new name, and the field is 'next' with a stated breakthrough gate that has not moved.
AWAdam Witanowski resurrectableWorld Models - Superseded† 30 Jun 2026
One commercial gateway over every model garden
“Standardised on someone else's roadmap.”
2 Dec 2025 — 30 Jun 2026Tested7 monthsCause · The routing experiment showed the value sits in the routing policy and the cost ledger, both of which we could not get out of the vendor gateway, so the recommendation became a gateway we control.
DYDylan DesmarchelierAI Gateway - Abandoned† 26 Jun 2026
Ambient Slack agent that acts unprompted
“Waited for permission. Still waiting.”
12 May 2026 — 26 Jun 2026Tried1 monthCause · The security review needed for a bot with read access to firm channels was not completed inside the run window, so the agent never went live and nothing was observed.
DDDavid Diaz resurrectableAmbient Agents - Refuted† 5 Jun 2026
Template guardrails stop org slop
“Guarded the template. Not the fork.”
24 Feb 2026 — 5 Jun 2026Tried3 monthsCause · After three months with mandated project templates, the count of unowned AI-built internal tools rose from 41 to 67; builders forked the template once and never took the updates.
MMMichael MenacheCitizen Developers and Org Slop - Superseded† 22 May 2026
On-prem H100 cluster pays back inside 18 months
“Payback period outran the price list.”
18 Nov 2025 — 22 May 2026Assessed6 monthsCause · Two AU-region hosted price cuts in Q1 and Q2 moved the breakeven to over four years before we finished the model; the standing answer now carries the calculation and its refresh date.
DYDylan DesmarchelierOn-Prem Inference - Abandoned† 19 May 2026
Plain OAuth scopes are enough for agent delegation
“Never got a tenant to fail in.”
24 Mar 2026 — 19 May 2026Tried2 monthsCause · No test tenant was made available in the run window, so the claim that native identity-provider scopes can express per-task agent delegation was never exercised against a real directory.
TBTravis Boast resurrectableAuth Broker - Refuted† 8 May 2026
LLM-as-judge without human calibration
“Agreed with itself, mostly.”
13 Jan 2026 — 8 May 2026Tested4 monthsCause · Against a five-person human panel on 600 paired items, an uncalibrated frontier judge reached Cohen's κ of 0.41 to 0.58 across three task families, below the 0.8 the hypothesis required and below the 0.6 kill line on two of them.
MLMichelle LamEval Harnesses - Rejected† 28 Apr 2026
Agent-written test suites replace QA engineers
“Could not be wrong, so could not run.”
14 Apr 2026 — 28 Apr 2026Signal14 daysCause · Not selected: the hypothesis as written is a headcount claim with no measurable kill condition, and the narrower pilot that could be preregistered (x-agentic-qa-pilot) was selected in its place.
AWAdam WitanowskiFully Agentic QA - Refuted† 3 Apr 2026
Unattended agent PRs merged on green CI
“Green CI, red incident.”
10 Mar 2026 — 3 Apr 2026Tried24 daysCause · In a three-week trial on the lab's own tooling repo, 4 of 31 unattended merges passed CI and broke behaviour the suite did not cover, including one that silently changed a cost calculation.
MMMichael Menache resurrectableAI-SDLC - Refuted† 27 Mar 2026
Prompt caching as a universal 60% cost cut
“Sixty percent, on the vendor's workload.”
4 Nov 2025 — 27 Mar 2026Tested5 monthsCause · Across six client patterns the measured saving from prompt caching ranged from 4% to 44%, and the 60% figure only appeared on workloads with a stable prefix over 2k tokens and a call rate high enough to keep the cache warm.
DYDylan DesmarchelierCost redux on tokens - Rejected† 17 Mar 2026
Sell a quantum-readiness audit this year
“Right question, wrong decade, wrong department.”
3 Mar 2026 — 17 Mar 2026Signal14 daysCause · Killed on inspection: no client has asked, the ASD's published migration timeline gives a 2030 horizon for most in-scope systems, and the substance of the audit is a crypto inventory that the firm's security practice already sells.
AWAdam WitanowskiQuantum Encryption - Abandoned† 31 Jan 2026
Multi-agent debate improves reasoning on our tasks
“Adjourned, not decided.”
20 Jan 2026 — 31 Jan 2026Tried11 daysCause · A Type 2 run stopped on day 11 when the owner was pulled onto the memory bench and the claims-triage eval it was using was re-baselined under a new extractor, leaving the two completed conditions incomparable with the planned third.
DDDavid DiazSelf-Organising Agents - Refuted† 16 Jan 2026
Fine-tuned 7B beats frontier on banking classification
“Beat the baseline it was allowed to pick.”
21 Oct 2025 — 16 Jan 2026Tried3 monthsCause · On a like-for-like eval with a properly prompted frontier baseline, the fine-tuned 7B trailed by 6 F1 points; the earlier 'win' had been against a frontier model with a one-line prompt.
MLMichelle Lam resurrectableSLM / Edge / Tuning