cavendish
AssessedEmerginggate · SkillsNow · 0–12 months×3 sightings

Deciding Table Stakes

right to play, as opposed to moats — i.e. skills

By 2026 skills, harnesses, evals and agent delivery are the price of being in the room — every competitor claims them and every client assumes them — so none of them is a moat; the only defensible position is measured results on the client's own data and the credibility to say no with evidence.

A validation run. Researched position, no experiment.

Join with…

Confidence

56%human-committed

Expiry

27duntil review · 30 Sep 2026

Lead time

—not yet mainstream · opened 2 Jun 2026

Ownership

AWAdam Witanowskimonthly cadence

Where it is

unattributed

Opened on conviction by the lab director in June after an RFP debrief in which an insurer said every firm they had spoken to 'does agents'. Argus positioning deltas across Accenture, Deloitte, QuantumBlack and Palantir since then show 'agentic' on all four AU practice pages, 'evals' on two, and a measured client outcome on none. A competitor's 'proprietary orchestration' pitch mapped to an open-source harness when Dalton read it; the director rebuilt a public 'agent factory' demo in three hours with stock tooling. A foundation lab's open skills spec has made skill libraries a commodity in one release. The skills gate is the honest one: the firm has to have all of this, cannot differentiate on any of it, and has to build it while pretending it is special. The disconfirming case — that consulting moats were always relationships and distribution and capability was never the point — is carried at moderate confidence and is the reason the field has an exec sponsor.

Written by hand and carrying nobody's name. Editing it puts yours on it.

Why a Quantium decision hinges on it

unattributed

This is the field that decides what the lab is for. If agents, evals and harnesses are table stakes, the lab's job is to get the firm to the table fast and cheaply, and the differentiated work is Nightingale measurement and the graveyard — being the firm that can show what worked and what did not on real data. If they are moats, the lab should be building proprietary tooling. The RFP scoring evidence says the first; the sponsor's conviction says the second is what wins pitches this year. The answer sets next year's budget split between capability and measurement.

Written by hand and carrying nobody's name. Editing it puts yours on it.

What it actually is

composed from the records

By 2026 skills, harnesses, evals and agent delivery are the price of being in the room — every competitor claims them and every client assumes them — so none of them is a moat; the only defensible position is measured results on the client's own data and the credibility to say no with evidence. That is the lab's one-line position on it, which is not the same as an explanation.

The shape the field is converging on, from the most authoritative source in it: A shared skills spec removes skill libraries as a differentiator.signal

This is the section a page most needs a person for, and the one composition is worst at. Nobody has written the plain-language version — what the idea is, in words that assume nothing — and it is the first thing a reader who has never met the term needs.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

Why now

composed from the records

The lab opened this field on 2026-06-02, and it has not reached mainstream awareness yet. Everything below is what has moved since.

A term arrived: Foundation lab publishes an open skills specification and marketplace That is the earliest reliable signal of a field being born — the vocabulary settles before the capability does.signal

What shipped: Foundation lab publishes an open skills specification and marketplace (Anthropic, 2026-04-22) and Accenture announces multi-thousand-headcount 'agentic services' practice and an agent factory (Accenture press release, 2026-03-10). Tooling arriving is what moves a field from argument to something a team could try.signalsignal

Demand is rising on it rather than steady — “Your competitor says agents will halve the delivery timeline. Is that true?” — which is the difference between a field worth watching and one worth doing something about.demand

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

What it changes in a system

composed from the records

What changes, concretely: Argus positioning delta, June 2026: 'agentic' present on 4 of 4 competitor AU practice pages; 'evaluation' on 2 of 4; a named, measured client outcome on 0 of 4.signal

Nothing is shipped as a default yet, so adopting this is a piece of work rather than a configuration change. That is usually the difference between a field being interesting and being used.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

What is in the way

composed from the records

The binding constraint is skills: it is operable and nobody is staffed to run it. Everything upstream of that is solved and everything downstream of it is waiting.

Workforce readiness is low: The firm has the skills at lab depth, not delivery depth. That is the gate. Agent-estimated. A recommendation needing skills the firm does not hold is an aspiration rather than an action, and it routes to the enablement agenda instead of the delivery one.

The argued case against it is the red team's, further down this page, and it is deliberately one-sided — this section is what stands in the way mechanically, not what somebody thinks of it.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

2 of 6 explanatory sections are written; the rest are composed until somebody takes them.

Business priority

Plan critical
  • Knowing what is merely a right to play, versus what is defensible, decides where the practice spends its capability budget.

    AH Amber Hallcommitted
  • WL-2Contributes toAgentic delivery at scale

    Competitor positions on agentic delivery are the best read on where the market thinks the boundary sits.

    AW Adam Witanowskicommitted

Clients are asking

  • “Your competitor says agents will halve the delivery timeline. Is that true?”4 engagements · $1M–5M · rising

Priority orders what you see. It never changes what the evidence says — a plan-critical field with nothing tested is still signal tier.

Field attributes

StateEmerging
GateSkills · operable, not yet staffed
OriginConviction
Measurablepartial
Written forexec
Reach · TLPPublic · TLP:CLEAR
Horizonnow
Opened2 Jun 2026
Mainstreamnot yet
Last validated26 Aug 2026
Sightings3

What people have written

Write one

Nothing yet. The person who knows a claim is wrong is usually not the person who wrote it.

A note never travels further than the thing it is written on.

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01Argus positioning delta, June 2026: 'agentic' present on 4 of 4 competitor AU practice pages; 'evaluation' on 2 of 4; a named, measured client outcome on 0 of 4.
  • 02A competitor's 'proprietary agent orchestration' pitch language mapped by Dalton to an open-source harness with a naming layer; the director rebuilt the public demo in three hours with stock tooling (tried tier).
  • 03A foundation lab published an open skills specification and marketplace in April; three competitors' 'skill libraries' now implement it.
  • 04Competitor hiring for 'AI evals engineer' roles rose from near zero to 23 open AU roles in six months — evals moved from differentiator to job description.
What is hype
  • 01'Proprietary agent platform' claims from consultancies. Every one we have examined is a harness plus a brand.
  • 02'Thousands of agents deployed.' Counts of things built, not things running, and never a measured outcome.
  • 03The lab's own temptation: that a good skills library is a moat. It was, for about a quarter.
What would have to be true
  • 01RFP scoring that rewards measured outcomes over capability claims. Currently procurement panels score the claims, and a measured 'no' can lose to an unmeasured 'yes'.
  • 02A Nightingale phase-2 result on the firm's own book — lead time, win rate on measured pitches — that shows measurement wins work. That is the evidence the sponsor's conviction is waiting for.
  • 03At least one competitor conceding the capability layer is a commodity in public. When the first one says it, the market moves.
What we would do
  • 01Publish pos-table-stakes as the firm's position and keep sa-vendor-claim-agents current for every sales conversation that starts with 'the vendor says'.
  • 02Fund capability building to parity, not beyond: skills, harness, evals, agent delivery, at the cheapest credible level. Spend the difference on measurement.
  • 03Run a Type 2 each quarter: rebuild the most-cited competitor demo with stock tooling and log the hours. The gap is the moat, measured.

Signals · 10 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
4 / 4
'agentic' on practice pages
Analyst·band 3Assessed

Argus positioning delta: 'agentic' on 4 of 4 competitor AU practice pages; measured outcome on 0

Computed from the claim layer across Accenture, Deloitte, QuantumBlack and Palantir AU practice pages, quarterly since 2025-09. 'Agentic' went from 1 of 4 to 4 of 4 in three quarters; 'evaluation' from 0 to 2; a named client with a measured outcome stayed at 0 throughout. The delta that opened the field.

extracted claimAgentic delivery claims have reached saturation across the AU competitor set inside three quarters, and none is backed by a published measured outcome.
Argus · positioning delta · Argus watchlist12 Jun 2026
detector · demand 2
Post·band 3Signal

'Consulting moats were never capabilities'

Argues that every consulting moat in history has been trust, distribution and incumbency; that capability claims are theatre both sides perform; and that a firm owned by a retailer has the only moat that matters in retail. Strongest disconfirming voice and the red team's spine.

LinkedIn · A former big-four partner12 Aug 2026
?dropped 4
Talk·band 3Signal

'The model is not the moat. Neither is the harness.' — competitor keynote

A QuantumBlack partner says on stage that the capability layer is a commodity and the differentiation is 'outcomes on your data'. The first competitor concession we have recorded; it is also exactly our position, which either confirms it or makes it table stakes too.

Industry conference, Sydney6 Aug 2026
detector · demand 2
23
openings
Job posting·band 2Signal

'AI Evaluation Engineer' — 23 open roles across AU competitor firms

Argus hiring scan. Near zero in January; 23 in July across four firms. Inference: evals are being staffed as a delivery function, which puts them on the table-stakes side within a year.

SEEK / competitor careers pages14 Jul 2026
detector · early adoption
3
hours to rebuild
Finding·band 1Tried

Logged from Claude Code: rebuilt a competitor's public 'agent factory' demo in three hours with a stock harness and open skills

The director reproduced the flow shown in a competitor's public demo video — intake, triage, drafting, handoff — using a stock harness, the open skills spec and no bespoke code. Three hours including the video. One run, tried tier, and a demo is not a programme.

MCP · log_finding · Adam Witanowski3 Jul 2026
AWdropped
Drop·band 3Signal

Competitor pitch excerpt: 'proprietary agent orchestration layer' — mapped by Dalton to an open-source harness

The sponsor dropped a de-identified excerpt from a competitor's pitch to a shared client. Dalton matched the described features one-for-one to an open-source harness's documentation. Aggregate only; no engagement named.

Slack drop25 Jun 2026
HBdropped 2
Client question·band 3Signal

'Every firm we spoke to says they do agents. What do you do that they don't?'

Asked by a COO at a debrief the firm lost. The sponsor logged it verbatim. The question the field exists to answer, and the reason it has an exec sponsor.

Engel · insurance RFP debrief28 May 2026
HBdropped 3
Post·band 2Signal

'Skills are the new prompts — and just as easy to copy'

Argues that skill libraries, like prompt libraries before them, have a shelf life of one open spec. Written three weeks after the open skills release; correct so far.

Personal blog · A former lab researcher2 May 2026
MMdropped 3
Release·band 1Signal

Foundation lab publishes an open skills specification and marketplace

Portable skill packaging with a public directory. Within two months three competitors' 'skill libraries' were implementations of it. The release that turned the lab's own skills library from an asset into a table stake.

Anthropic22 Apr 2026
detector · bleeding edge 3
Announcement·band 3Signal

Accenture announces multi-thousand-headcount 'agentic services' practice and an agent factory

Headcount, a factory metaphor and a partnership list. No outcome figures. Argus records the move from 'AI' to 'agentic' in the firm's naming; the largest competitor now uses the same word as everyone else.

extracted claimThe largest consultancy has rebranded its AI practice around agents at scale.
Accenture press release10 Mar 2026
detector · demand 4
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 4 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

Every top-tier competitor in Australia now publicly claims agentic delivery; the claim no longer discriminates in an RFP.

Assessedc-table-stakes-1dalton-0.416 Jun 2026Argus · positioning delta, Accenture press release, Engel · insurance RFP debrief
78%

Evals are claimed by half the competitor set and demonstrated publicly by none; they are six to twelve months from table stakes.

Assessedc-table-stakes-2dalton-0.421 Jul 2026Argus · positioning delta, SEEK / competitor careers pages
66%

'Proprietary orchestration' claims in competitor pitches map to open-source harnesses in the majority of cases examined.

Triedc-table-stakes-3dalton-0.421 Jul 2026Slack drop, MCP · log_finding, Anthropic
60%

The only capability a client cannot obtain from any competitor is a measured outcome on their own data, published with what did not work.

Assessedc-table-stakes-4dalton-0.426 Aug 2026Industry conference, Sydney, Personal blog
55%

Capability is irrelevant to consulting moats; relationships and distribution decide, and the firm's moat is the Woolworths relationship.

Assessedc-table-stakes-5dalton-0.426 Aug 2026LinkedIn
38%

Position history · the diff is the product

3 validation runs against a fixed brief. Confidence 45% → 56%.

runs compare claim sets, never prose
What we said · run 3

Position drafted: parity on capability, differentiation on measurement and the graveyard. Relationship counter-case rated higher than expected; verdict in doubt until Engel can show measured pitches win. Position and standing answer published with the doubt stated.

56%
Changed since run 2
  • The only capability a client cannot obtain from any competitor is a measured outcome on their own data, published with what did not work.
  • Capability is irrelevant to consulting moats; relationships and distribution decide, and the firm's moat is the Woolworths relationship.
  • c-table-stakes-3 ↑ 0.52 → 0.60
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · HB
high

Sets the capability-versus-measurement budget split for next year.

Timeline

committed · AW
0–18mo

The skills gate closes on the firm's own build pace, not the market's.

Cost of being wrong

committed · AW
high

Wrong one way funds a proprietary platform nobody buys; wrong the other way loses pitches to firms that claim more.

Demand

committed · HB
high

Every pitch debrief this half has had a 'what do you do that they don't' moment.

TAM

agent-estimated
>$10B

AU AI-services market; not a meaningful axis for a positioning field. Agent-estimated, uncommitted.

Workforce readiness

agent-estimated
low

The firm has the skills at lab depth, not delivery depth. That is the gate. Agent-estimated.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Cross-sector
relevant

The positioning question is the same in every sector; the RFP language differs only in which vendor is being quoted.

Mechanism · Position paper plus the standing answer on vendor agentic claims, used in every pitch.

HB committed by Harley Barnescommitted · HB
Banking
relevant

Bank procurement panels score capability claims on a grid; a measured 'not yet' scores below an unmeasured 'yes'.

Mechanism · Lead with the graveyard and the measured result; make the panel score evidence.

AH committed by Amber Hallcommitted · AH
Government
watch

Panel arrangements reward breadth of claimed capability; the measured-outcome position may not score until the buying guidance changes.

Mechanism · Track whether the digital sourcing framework adds evidence-of-outcome criteria.

AV committed by Aadhithyanarayanan V Acommitted · AV
Retail & FMCG
relevant

The Woolworths relationship is the disconfirming case's strongest example, so it is where the thesis gets tested hardest.

Mechanism · Compare win rate on measured pitches against relationship pitches once Engel data allows. Agent draft.

Agent draft · awaiting a sector owneragent-estimated

Red team · the strongest case against

The strongest case against: this is a lab justifying itself. 'Capability is table stakes, measurement is the moat' is exactly what a measurement-led lab would conclude, and the evidence is four practice pages and one rebuilt demo. Consulting has never been won on capability. It is won on who the CEO trusts, which is why a firm owned by the country's largest retailer has a moat no eval will move. The field's origin is conviction and its sponsor is the person whose conviction it is.

  • —Four competitor practice pages are marketing, not capability. Their claims being undifferentiated says nothing about whether their delivery is.
  • —One RFP debrief and one rebuilt demo. Three hours to rebuild a demo is not three hours to deliver a programme.
  • —Measured outcomes are a moat only if clients pay for them. Engel does not yet show that a measured pitch wins more often than a confident one.
  • —The field's own conclusion allocates budget to the lab that wrote it. That is a conflict, and it is not addressed by naming it.
Run by an agent briefed to argue the field is nothing — sources here are correlated, and without a deliberate adversary synthesis converges on consensus and calls it insight. Kept as a dated pass rather than overwritten. Nobody has answered it yet, and a challenge nobody answers is a disclaimer.thesis in doubt

Source diversity

  • Competitors / Argus40%
  • Foundation labs15%
  • Practitioner and commentary20%
  • Internal / Engel25%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsAWHBMMMLAHAV

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01Does a measured pitch win more often than a confident one? Engel can answer this and has not yet.
  2. 02How long is the window between a capability being claimed by one competitor and by all of them — and is it shrinking?
  3. 03If measurement is the moat, what stops a competitor with a bigger client base measuring more than we can?

Notes · anyone in the firm

What people have written on this.

The person who knows a claim is wrong is usually not the person who wrote it. Corrections, objections and questions are owed an answer and stay open until the field owner says what they did; context and use notes stand as they are.

Notes · 0

Anything here reaches at most public — a note cannot travel further than what it is written on.

    Nothing written on this yet. The useful notes are the ones from people who are not in the lab — that is where the correction usually comes from.