cavendish
SignalCandidategate · BreakthroughDistant · 7+ years

AGI

still in the foothills

AGI is not a capability the lab can watch for directly; it is a set of specific capabilities — sustained learning in deployment, reliable long-horizon agency, transfer across unrelated professional domains — each of which the lab tracks elsewhere, and 'AGI' is what we call the state where all of them have fired.

Clustered only. No lab work behind it. Cannot be cited.

Join with…

Confidence

20%unresearched

Expiry

84duntil review · 26 Nov 2026

Lead time

36moopened after mainstream — recorded honestly

Ownership

HBHarley Barnessponsor · prospective vertical

Where it is

unattributed

The field exists because the exec sponsor asked the question and the lab needed a defensible answer that was not a lab's press release. The honest position: definitions vary so much that lab timelines cannot be compared, and the evidence available to us is our own canary suite, which shows frontier models clearing more of our professional task families each release and failing in the same three ways — they do not learn from the last task, they lose the thread over long horizons, and they cannot yet be trusted to be wrong in a predictable way. Lab timelines have moved earlier every year; our capability tracker has moved steadily but not exponentially. We hold the field at Distant because the specific things we would need to see have not been seen, and we record what they are so the field is watchable rather than a mood.

Written by hand and carrying nobody's name. Editing it puts yours on it.

Why a Quantium decision hinges on it

unattributed

Every board conversation the firm's executives have with clients now contains the question, and the answer shapes ten-year decisions: what to build, what to hire, what to defer. The value of holding the field is not predicting the date; it is being able to say what evidence would move us and to show that the evidence has not arrived. That is a more useful answer than any date.

Written by hand and carrying nobody's name. Editing it puts yours on it.

What it actually is

composed from the records

AGI is not a capability the lab can watch for directly; it is a set of specific capabilities — sustained learning in deployment, reliable long-horizon agency, transfer across unrelated professional domains — each of which the lab tracks elsewhere, and 'AGI' is what we call the state where all of them have fired. That is the lab's one-line position on it, which is not the same as an explanation.

The shape the field is converging on, from the most authoritative source in it: The task horizon at which agent success halves is doubling roughly every seven months.signal

This is the section a page most needs a person for, and the one composition is worst at. Nobody has written the plain-language version — what the idea is, in words that assume nothing — and it is the first thing a reader who has never met the term needs.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

Why now

composed from the records

The lab opened this field on 2026-03-02 and it reached mainstream awareness on 2023-03-14. The gap between those two dates is the lead time the lab is measured on.

What shipped: Frontier lab CEO: 'AGI by 2028' in an investor letter (Foundation lab, 2026-02-10). Tooling arriving is what moves a field from argument to something a team could try.signal

And the rule moved: APRA letter to ADIs on AI concentration and dependency risk In a regulated vertical that usually decides the timing more than the technology does.signal

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

What it changes in a system

composed from the records

What changes, concretely: Frontier models clear an increasing share of the lab's professional task families each release: 41% of families at expert level in 2025-09, 58% in 2026-08 (canary suite).signal

Nothing is shipped as a default yet, so adopting this is a piece of work rather than a configuration change. That is usually the difference between a field being interesting and being used.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

What is in the way

composed from the records

The binding constraint is breakthrough: it is not yet technically possible. Everything upstream of that is solved and everything downstream of it is waiting.

It is also gated on another field: continuous-learning — A deployed frontier model improves on the lab's held-out task families across at least four weeks without a release — weight-level or otherwise. Detected by re-running the canary suite on a fixed model version monthly and seeing the score move.; eval-harnesses — A frontier release scores at expert level on three professional task families the lab wrote after the model's training cutoff and never published, with failures a human expert would make; the canary suite is extended with sealed families for exactly this.; self-organising-agents — An agent or collective completes a multi-day professional task from the lab's suite with an error distribution the lab can bound in advance — the long-horizon reliability trigger.. Until that trigger fires, effort here compounds slowly.clusterclustercluster

Workforce readiness is low: Not a readiness question on this horizon. Agent-estimated. A recommendation needing skills the firm does not hold is an aspiration rather than an action, and it routes to the enablement agenda instead of the delivery one.

The argued case against it is the red team's, further down this page, and it is deliberately one-sided — this section is what stands in the way mechanically, not what somebody thinks of it.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

2 of 6 explanatory sections are written; the rest are composed until somebody takes them.

Business priority

Off-plan

The plan has not asked for this.

No objective names it, so it is here on the lab's judgement alone. That is legitimate for a long-horizon field and is exactly how optionality is meant to look — but it is reviewed each quarter, because drift looks identical from the outside.

Priority orders what you see. It never changes what the evidence says — a plan-critical field with nothing tested is still signal tier.

Field attributes

StateCandidate
GateBreakthrough · not yet technically possible
OriginQuestion
Measurablepartial
Written forexec
Reach · TLPThe firm · TLP:AMBER
Horizondistant
Opened2 Mar 2026
Mainstream14 Mar 2023
Last validated26 Aug 2026
Sightings1

What people have written

Write one

Nothing yet. The person who knows a claim is wrong is usually not the person who wrote it.

A note never travels further than the thing it is written on.

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01Frontier models clear an increasing share of the lab's professional task families each release: 41% of families at expert level in 2025-09, 58% in 2026-08 (canary suite).
  • 02The three consistent failure modes across every release: no learning from the previous task, thread loss over multi-day horizons, unpredictable error distribution.
  • 03Lab-published timelines have moved earlier each year since 2023; our own capability tracker shows a steady, not accelerating, slope.
What is hype
  • 01Timelines announced by people whose valuation depends on them.
  • 02Benchmark saturation presented as generality; every saturated benchmark was written by humans who could not imagine the failure mode.
  • 03'AGI achieved internally' as a genre of post.
What would have to be true
  • 01A model that improves on the lab's held-out task families across weeks of deployment without a release — the Continuous Learning trigger.
  • 02Reliable agency over a multi-day task with a measured error distribution the lab can bound — the Self-Organising Agents and Learning Agents triggers.
  • 03Expert-level performance on three professional task families the model was not trained toward and the lab did not publish, with the failures being the kind an expert makes.
What we would do
  • 01Answer the sponsor's question with the three triggers and the canary trend, quarterly, in one page.
  • 02Never issue a date. Issue what would move us.
  • 03If two of three triggers fire inside a year: elect the field, assign an owner, and treat every Now-horizon default as up for review.

Signals · 9 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
58%
families at expert level
Finding·band 1Assessed

Canary suite, 2026-08 release: 58% of professional task families at expert level, up from 41% a year ago

The lab's release canary across 34 professional task families. Coverage at expert level rose from 41% to 58% over twelve months on a steady slope. The three failure modes recur in every family that has not cleared. The only AGI evidence the lab generates itself.

extracted claimFrontier models clear more professional task families each release on a steady, not accelerating, slope.
Lab · Turing canary suite · Michelle Lam22 Aug 2026
detector · bleeding edge
Regulatory·band 3Signal

APRA letter to ADIs on AI concentration and dependency risk

Asks banks to assess dependency on a small number of model providers. Not about AGI, but the regulatory shadow of the same question: what happens to the system if the capability jumps. Logged for the banking brief.

APRA5 Aug 2026
AHdropped
Post·band 2Signal

'We have achieved AGI internally' — a thread

The genre. No evidence, very widely shared, arrived at the exec sponsor within a day. Logged so the lab's brief can name it.

Social · A frontier-lab employee30 Jul 2026
HB?dropped 5
Finding·band 1Tried

Logged from Claude Code: sealed task family (APRA reporting reconciliation) — expert on the parts it had seen, novice on the part it had not

Red team ran a task family written after the model's cutoff and never published. Expert-level on sub-tasks resembling public material; failed the novel reconciliation step in a way no trained analyst would. The transfer test in miniature. Tried tier.

MCP · log_finding · Travis Boast6 Jul 2026
TBdropped
Paper·band 1Signal

Task Horizon and Failure: How Long Can an Agent Hold the Thread?

Measures the task duration at which agents' success rate halves; it has doubled roughly every seven months for two years and currently sits around one working day. An exponential in a metric that matters, which the red team notes and the position does not yet reflect.

~7 monthshorizon doubling
arxiv.org · An evals organisation19 Jun 2026
DDAWdropped 4
Benchmark·band 2Signal

Public expert-level benchmark saturated; successor announced within a month

The third 'final exam' style benchmark to saturate in eighteen months, followed by a harder one. The pattern the hype list refers to. The lab's canary uses sealed families to avoid it.

Benchmark leaderboard28 May 2026
MLdropped 2
Paper·band 2Signal

Forty Definitions of AGI and Why Their Timelines Cannot Be Compared

Catalogues published definitions from labs, academics and forecasters. Finds no two labs use a testable definition in common. The paper behind the lab's refusal to compare timelines.

extracted claimNo two frontier labs share a testable definition of AGI.
40definitions catalogued
arxiv.org · Philosophers and an evals researcher16 Apr 2026
MLTBdropped 2
Client question·band 3Signal

'What do we do if it happens in 2028?'

Asked by a bank director after the investor letter. The question that opened the field. Answered with the three triggers; the director asked for the quarterly page.

Engel · board briefing, banking20 Mar 2026
HBdropped
Announcement·band 1Signal

Frontier lab CEO: 'AGI by 2028' in an investor letter

The latest in a series of earlier-every-year timelines. The definition used is 'can do most economically valuable cognitive work', which is not testable. Argus positioning delta: the same lab said 2030 in 2024 and 2029 in 2025.

extracted claimAGI on the lab's own definition arrives by 2028.
Foundation lab10 Feb 2026
HB?MAdropped 6
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 4 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

Published lab definitions of AGI differ enough that their timelines cannot be compared to each other or to our tracker.

Assessedc-agi-4dalton-0.330 Apr 2026arxiv.org, Foundation lab
82%

Frontier models clear more of the lab's professional task families each release on a steady, not accelerating, slope.

Assessedc-agi-1dalton-0.426 Aug 2026Lab · Turing canary suite, Benchmark leaderboard
78%

Three failure modes persist across every release: no learning from the last task, long-horizon thread loss, unpredictable error distribution.

Assessedc-agi-2dalton-0.426 Aug 2026Lab · Turing canary suite, MCP · log_finding, arxiv.org
74%

Transfer to an unpublished professional task family is the most discriminating test available to the lab, and no model has passed it at expert level.

Triedc-agi-5dalton-0.48 Jul 2026MCP · log_finding, Lab · Turing canary suite
60%

AGI on any published lab definition arrives before 2030.

Signalc-agi-3dalton-0.426 Aug 2026Foundation lab, Social
20%

Position history · the diff is the product

2 validation runs against a fixed brief. Confidence 18% → 20%.

runs compare claim sets, never prose
What we said · run 2

Canary coverage up to 58% of families at expert level, steady slope. Three failure modes unchanged in kind. Red team argues the tracker is capped by design; unanswered. Still Distant.

20%
Changed since run 1
  • Frontier models clear more of the lab's professional task families each release on a steady, not accelerating, slope.
  • Three failure modes persist across every release: no learning from the last task, long-horizon thread loss, unpredictable error distribution.
  • Transfer to an unpublished professional task family is the most discriminating test available to the lab, and no model has passed it at expert level.
  • c-agi-3 ↑ 0.15 → 0.2
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · AW
high

Nothing else in the graph makes every other field moot.

Timeline

committed · AW
4yr+

On our triggers, none of which has fired. Not a date.

Cost

committed · ML
low

The canary suite already runs on every release; the field costs one page a quarter.

Cost of being wrong

agent-estimated
high

Agent-estimated in both directions: early means misallocated decade; late means the firm's defaults are wrong all at once.

Demand

agent-estimated
not-measurable-here

The question arrives from boards, not engagements; Engel does not see it. Agent-drafted.

Workforce readiness

agent-estimated
low

Not a readiness question on this horizon. Agent-estimated.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Cross-sector
watch

The field is a board question in every vertical and an engagement question in none.

Mechanism · Quarterly one-page brief from the triggers; no delivery mechanism.

HB committed by Harley Barnescommitted · HB
Banking
watch

Bank boards ask it most; APRA has started asking about AI concentration risk, which is the regulatory shadow of the same question.

Mechanism · Brief reused for board-level conversations; nothing in delivery.

AH committed by Amber Hallcommitted · AH
Health
not-relevant

Health clients ask about specific clinical capabilities, which are tracked in their own fields; the general question does not arise.

Mechanism · None.

Agent draft · awaiting a sector owneragent-estimated

Red team · the strongest case against

The strongest case against our position: the lab's canary suite is a lagging indicator by construction — it tests what we already know how to test — and the three 'persistent' failure modes are each being worked on directly at frontier labs with results the lab has already logged (the continual-learning paper, the long-horizon agent results). A steady slope on our tracker with an accelerating slope on theirs is the signature of a tracker that measures the wrong thing. Also: our refusal to issue a date is itself a position, and it is the position of every institution that was late.

  • The canary suite saturates at expert level per family and cannot see beyond it; the slope is capped by our test design, not the models.
  • Each of the three failure modes has a credible lab result against it this year; three separate research programmes converging is exactly what 'before 2030' would look like from here.
  • Holding at Distant with no owner is the cheapest possible position and cheap positions are usually wrong on the things that matter most.
Run by an agent briefed to argue the field is nothing — sources here are correlated, and without a deliberate adversary synthesis converges on consensus and calls it insight. Kept as a dated pass rather than overwritten. Nobody has answered it yet, and a challenge nobody answers is a disclaimer.thesis weakened

Source diversity

  • Evals research35%
  • Frontier labs20%
  • Commentary / social15%
  • Internal / Engel30%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Trigger · A deployed frontier model improves on the lab's held-out task families across at least four weeks without a release — weight-level or otherwise. Detected by re-running the canary suite on a fixed model version monthly and seeing the score move.

When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.

depends onEval Harnesses

Trigger · A frontier release scores at expert level on three professional task families the lab wrote after the model's training cutoff and never published, with failures a human expert would make; the canary suite is extended with sealed families for exactly this.

When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.

Trigger · An agent or collective completes a multi-day professional task from the lab's suite with an error distribution the lab can bound in advance — the long-horizon reliability trigger.

When the trigger fires, this field is resurfaced automatically. Watchable rather than parked.

enablingNon-Weight-Bound Continuous Learning

Learning in deployment is the first of the three triggers and is tracked there.

enablingSelf-Organising Agents

Reliable long-horizon agency is the second trigger; the collective form is tracked there.

compoundingEval Harnesses

The canary suite is the only evidence the lab generates itself on this question.

compoundingPost Economy

The macro half of that field is what this one produces if it fires.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

No experiments, recommendations or graveyard entries yet. That is what a candidate looks like.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01Is the canary suite capped by design, and would a sealed-family extension change the slope?
  2. 02If the task-horizon doubling continues, at what date does it cross the multi-day trigger, and is that inside the Next horizon?
  3. 03What does the firm actually do differently on the day two of three triggers fire?

Notes · anyone in the firm

What people have written on this.

The person who knows a claim is wrong is usually not the person who wrote it. Corrections, objections and questions are owed an answer and stay open until the field owner says what they did; context and use notes stand as they are.

Notes · 0

Anything here reaches at most the firm — a note cannot travel further than what it is written on.

    Nothing written on this yet. The useful notes are the ones from people who are not in the lab — that is where the correction usually comes from.