cavendish
SignalCandidategate · AdoptionNear · 1–3 years×2 sightings

Education

AI tutoring has crossed from promising to evidenced in narrow, well-instrumented settings; the field that matters to the firm is not the schools market but professional upskilling — including the firm's own — where the same efficacy evidence applies and the buyer already exists.

Clustered only. No lab work behind it. Cannot be cited.

Join with…

Confidence

42%unresearched

Expiry

14doverdue for review

Lead time

7moopened after mainstream — recorded honestly

Ownership

Unownedcandidate — a named human elects

Where it is

unattributed

The evidence base changed in the last year. Two large randomised trials — one in higher education, one in corporate technical training — showed AI tutoring producing learning gains of a third to half a standard deviation over instructor-led baselines, with the effect concentrated in structured, assessable skills. Assessment integrity is the counterweight: universities are moving to in-person and oral assessment because take-home work no longer evidences anything. Quantium is not an education company and has no sector owner for it. What it has is a 400-person upskilling problem in agentic delivery, a government practice whose clients include education departments, and three clients who have asked whether the firm's own AI capability-building programme is something they could buy. Candidate; not elected; no owner.

Written by hand and carrying nobody's name. Editing it puts yours on it.

Why a Quantium decision hinges on it

unattributed

Two decisions hinge on it. The first is internal: the firm's ability to execute on half the fields on this board is gated by skills, and AI tutoring is the only upskilling approach with trial evidence at the scale the firm needs. The second is whether 'we upskilled ourselves and can do it for you' is a product. The schools and university market is a distraction the firm should explicitly decline.

Written by hand and carrying nobody's name. Editing it puts yours on it.

What it actually is

composed from the records

AI tutoring has crossed from promising to evidenced in narrow, well-instrumented settings; the field that matters to the firm is not the schools market but professional upskilling — including the firm's own — where the same efficacy evidence applies and the buyer already exists. That is the lab's one-line position on it, which is not the same as an explanation.

The shape the field is converging on, from the most authoritative source in it: AI tutoring produces large learning gains on structured, assessable skills and none on open-ended ones.signal

This is the section a page most needs a person for, and the one composition is worst at. Nobody has written the plain-language version — what the idea is, in words that assume nothing — and it is the first thing a reader who has never met the term needs.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

Why now

composed from the records

The lab opened this field on 2026-05-28 and it reached mainstream awareness on 2025-11-04. The gap between those two dates is the lead time the lab is measured on.

What shipped: Foundation lab ships a 'learning mode' that withholds answers and asks Socratic questions (Anthropic, 2025-11-04). Tooling arriving is what moves a field from argument to something a team could try.signal

And the rule moved: TEQSA guidance: assessment must evidence learning in a way that is 'secure against generative AI' In a regulated vertical that usually decides the timing more than the technology does.signal

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

What it changes in a system

composed from the records

What changes, concretely: Randomised trials in higher education and corporate technical training show AI tutoring gains of 0.3–0.5 SD on structured, assessable skills over instructor-led baselines (band-1 papers; not lab work).

Nothing is shipped as a default yet, so adopting this is a piece of work rather than a configuration change. That is usually the difference between a field being interesting and being used.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

What is in the way

composed from the records

The binding constraint is adoption: it is ready, and trust, regulation, procurement or change capacity are what is left. Everything upstream of that is solved and everything downstream of it is waiting.

Workforce readiness is low: Nobody in the firm has learning-science training; the lab can read the trials but not design one. Agent-estimated. A recommendation needing skills the firm does not hold is an aspiration rather than an action, and it routes to the enablement agenda instead of the delivery one.

The argued case against it is the red team's, further down this page, and it is deliberately one-sided — this section is what stands in the way mechanically, not what somebody thinks of it.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

2 of 6 explanatory sections are written; the rest are composed until somebody takes them.

Business priority

Loosely aligned
  • Capability building across the firm is downstream of the objective, though nobody has committed to owning it.

    Agent draft — no owner has committed this alignmentagent-estimated

Priority orders what you see. It never changes what the evidence says — a plan-critical field with nothing tested is still signal tier.

Field attributes

StateCandidate
GateAdoption · ready — blocked by trust, regulation, procurement or change capacity
OriginSignal
Measurablepartial
Written forexec
Reach · TLPThe firm · TLP:AMBER
Horizonnear
Opened28 May 2026
Mainstream4 Nov 2025
Last validated16 Jul 2026
Sightings2

What people have written

Write one

Nothing yet. The person who knows a claim is wrong is usually not the person who wrote it.

A note never travels further than the thing it is written on.

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01Randomised trials in higher education and corporate technical training show AI tutoring gains of 0.3–0.5 SD on structured, assessable skills over instructor-led baselines (band-1 papers; not lab work).
  • 02Take-home assessment no longer evidences learning; universities are moving to in-person, oral and process-based assessment at scale.
  • 03The firm's own agentic-delivery upskilling cohort (40 people, Q2) used an AI tutor over the lab's material; completion 88% against 51% for the previous cohort's video course (tried, no control).
What is hype
  • 01'Personalised learning for every child.' The trial evidence is on structured, assessable skills in motivated adult learners; the extrapolation to schools is unsupported.
  • 02AI tutors as a replacement for instructors. The trials with the largest effects kept the instructor and changed what they did.
  • 03Detection tools for AI-written assessment. Every one tested has false-positive rates that make it unusable for a decision about a student.
What would have to be true
  • 01A client paying for the firm's capability programme rather than treating it as a pre-sales conversation.
  • 02The internal cohort result holding under a control, which means running the next cohort as a Type 2 with a comparison arm.
  • 03A named owner. Nobody in the lab or the firm owns education, and a candidate without one dies quietly.
What we would do
  • 01Keep it a candidate. Run the next internal upskilling cohort as a Type 2 with a comparison arm so the firm has its own evidence.
  • 02Decline the schools and university market explicitly and record the rejection.
  • 03Ask the government sector owner to log any education-department demand in Engel so the field has a demand score in a quarter.

Signals · 9 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
0.41 SD
effect size
Paper·band 1Signal

AI Tutoring at Scale: A Pre-Registered Randomised Trial Across 11,400 University Students

Pre-registered RCT across four universities and three disciplines. AI tutor plus instructor beat instructor alone by 0.41 SD on end-of-course assessment in quantitative subjects; no significant effect in essay-based subjects. Effect strongest for students in the bottom tercile at baseline.

extracted claimAI tutoring produces large learning gains on structured, assessable skills and none on open-ended ones.
arxiv.org · A multi-university learning-sciences consortium9 Mar 2026
AWMLdropped 3
88% vs 51%
completion
Finding·band 1Tried

Internal: Q2 agentic-delivery cohort with an AI tutor over lab material — 88% completion

Forty people, six weeks, AI tutor built on the lab's own agentic-delivery material. Completion 88% against 51% for the prior cohort's video course. No comparison arm, no post-test; tried tier, novelty effect likely.

extracted claimAn AI tutor over the firm's own material raises upskilling completion sharply (uncontrolled).
Lab · capability programme · Michelle Lam10 Jul 2026
MLdropped
Client question·band 3Signal

'Could we buy the programme you ran for your own people?'

Asked by a retail client's chief data officer after a delivery review where the firm's upskilling cohort came up. Logged informally by the retail sector owner. The third such question this half; none has become a proposal.

Engel · retail engagement30 Jun 2026
DBdropped 3
Post·band 2Signal

'The tutor taught them SQL. It did not teach them what to ask.'

Argues the RCT effects are confined to skills with a right answer and that judgement-heavy skills show no gain because the tutor cannot assess them. Kept as the strongest disconfirming voice on transfer.

Substack · A learning-science researcher14 Jun 2026
AWdropped 2
80
respondents
Analyst·band 3Signal

Analyst: AU corporate L&D budgets shifting 20% from content licences to AI tutoring platforms by 2027

Survey of 80 AU L&D leaders. Directional; vendor-sponsored. Supports the buyer-exists half of the thesis and nothing else.

Forrester26 May 2026
HBdropped
−34%
time to certification
Paper·band 1Signal

Does It Work at Work? AI Tutoring for Corporate Technical Upskilling: A Field Experiment

Field experiment in a 2,000-person technology firm's data-engineering upskilling programme. AI-tutored arm reached certification 34% faster with a 0.29 SD higher practical assessment score. The instructor's role shifted to review and unblocking.

extracted claimAI tutoring transfers to adult technical upskilling with a smaller but material effect.
arxiv.org · Vasquez, Ng et al.12 May 2026
MLdropped 2
8–24%
false positives
Benchmark·band 2Signal

AI-writing detector evaluation: false-positive rates on non-native English writers

Open evaluation of seven commercial detectors. False-positive rates of 8–24% on human-written text by non-native English speakers. No detector clears a threshold that would be defensible for an individual academic-misconduct finding.

github.com22 Apr 2026
detector · early adoption 2
Regulatory·band 2Signal

TEQSA guidance: assessment must evidence learning in a way that is 'secure against generative AI'

Australian higher-education regulator's request-for-action. Universities are responding with in-person, oral and process-based assessment. The spend has moved from detection to redesign.

TEQSA18 Feb 2026
AVdropped 2
Release·band 1Signal

Foundation lab ships a 'learning mode' that withholds answers and asks Socratic questions

Product mode designed for education use, released after the higher-ed RCT's pre-registration became public. Marks the point the tutoring use case became mainstream; the field was opened seven months later, which is a negative lead time and recorded as one.

Anthropic4 Nov 2025
detector · bleeding edge 2
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 4 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

AI-writing detection tools have false-positive rates that make them unusable for individual assessment decisions.

Signalc-education-5dalton-0.330 Apr 2026github.com
80%

Take-home assessment no longer evidences learning; the assessment redesign, not the tutor, is where institutions are spending.

Signalc-education-3dalton-0.422 Jun 2026TEQSA, Forrester
75%

AI tutoring produces learning gains of 0.3–0.5 SD over instructor-led baselines on structured, assessable skills in adult learners.

Signalc-education-1dalton-0.416 Jul 2026arxiv.org, arxiv.org
72%

The firm's own upskilling is the tractable application; the external education market is not one Quantium should enter.

Signalc-education-4dalton-0.416 Jul 2026Lab · capability programme, Engel · retail engagement
60%

The tutoring effect transfers to open-ended, judgement-heavy skills at the same magnitude.

Signalc-education-2dalton-0.416 Jul 2026arxiv.org, Substack
18%

Position history · the diff is the product

2 validation runs against a fixed brief. Confidence 35% → 42%.

runs compare claim sets, never prose
What we said · run 2

Internal upskilling cohort result logged. The firm's own capability building is the tractable application; the external market is not. Remain a candidate; run the next cohort with a control.

42%
Changed since run 1
  • The tutoring effect transfers to open-ended, judgement-heavy skills at the same magnitude.
  • The firm's own upskilling is the tractable application; the external education market is not one Quantium should enter.
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

agent-estimated
medium

High for the firm's own capability; low as a market. Agent-estimated on the internal case.

Timeline

committed · AW
0–18mo

The internal upskilling need is this year; the trial evidence is already published.

TAM

agent-estimated
>$10B

Agent-estimated global education and corporate-training spend. Almost none addressable by the firm. Uncommitted.

Cost

committed · ML
low

A Type 2 on the next cohort with a comparison arm is a fortnight of measurement work.

Demand

committed · AV
low

One education-department mention in Engel this half, about assessment policy, not tutoring. Three clients asked about the firm's own programme informally.

Workforce readiness

agent-estimated
low

Nobody in the firm has learning-science training; the lab can read the trials but not design one. Agent-estimated.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Government
watch

State education departments are clients of the government practice for other things. Assessment-integrity policy is live; tutoring procurement is not.

Mechanism · Would need an education department asking for evidence synthesis or policy analytics, which the practice already does in other domains.

AV committed by Aadhithyanarayanan V Acommitted · AV
Cross-sector
relevant

Every client's AI programme is gated by skills. The firm's own capability programme is the artefact, and the tutoring evidence says how to run it.

Mechanism · AI tutor over the lab's material for client teams as part of delivery; instructor role redesigned around it.

Agent draft · awaiting a sector owneragent-estimated
Banking
not-relevant

Banks run their own L&D at scale and buy content, not tutoring platforms. Nothing here changes a bank's decision.

Mechanism · None.

AH committed by Amber Hallcommitted · AH

Red team · the strongest case against

The strongest case against: this is a candidate because it is interesting, not because the firm can act on it. The internal cohort result has no control and a strong novelty effect; the trial evidence is on skills the firm does not primarily need to teach; and 'we can upskill you' is a sentence every consultancy says and none is paid for. Meanwhile the real education market is a policy and procurement domain the firm has no standing in.

  • Completion rate is not learning. The internal cohort measured who finished, not what they could do afterwards.
  • The 0.3–0.5 SD gains are on structured, assessable skills. Agentic delivery judgement — the skill the firm is actually short of — is the kind the trials show no effect on.
  • No owner, no sector, no demand score above low. Every mechanism in the system says this field should not be elected, and it is right.
Run by an agent briefed to argue the field is nothing — sources here are correlated, and without a deliberate adversary synthesis converges on consensus and calls it insight. Kept as a dated pass rather than overwritten. Nobody has answered it yet, and a challenge nobody answers is a disclaimer.thesis in doubt

Source diversity

  • Learning-sciences research40%
  • Regulators / institutions15%
  • Foundation lab / vendor / analyst20%
  • Internal / Engel25%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsAWMLAV

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

No experiments, recommendations or graveyard entries yet. That is what a candidate looks like.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01Does the internal cohort's completion gain survive a comparison arm and a practical post-test?
  2. 02Is there a version of the firm's capability programme a client would pay for, or is it permanently pre-sales?
  3. 03What does the lab need to teach that the tutoring evidence says tutors cannot — and how is that taught instead?

Notes · anyone in the firm

What people have written on this.

The person who knows a claim is wrong is usually not the person who wrote it. Corrections, objections and questions are owed an answer and stay open until the field owner says what they did; context and use notes stand as they are.

Notes · 0

Anything here reaches at most the firm — a note cannot travel further than what it is written on.

    Nothing written on this yet. The useful notes are the ones from people who are not in the lab — that is where the correction usually comes from.