cavendish
TriedEmerginggate · AdoptionNear · 1–3 years

Effective Assistants

Whether an assistant is still being used in week three is decided by four properties — proactivity, calibration, refusal honesty and integration depth — and none of them is a model property; they are product decisions that most deployments get wrong by default.

Someone ran it in their own harness. Artifact, no protocol. Decays fast.

Join with…

Confidence

60%human-committed

Expiry

14doverdue for review

Lead time

not yet mainstream · opened 3 Feb 2026

Ownership

MMMichael Menachemonthly cadence

Where it is

unattributed

This field exists because clients keep asking the same question in different words: we deployed an assistant, usage collapsed after a fortnight, why? Engel has eleven engagements where that question was asked in the last two quarters. The evidence, such as it is, comes from usage telemetry rather than papers: assistants that raise something before being asked, that say how sure they are, that refuse visibly rather than confabulate, and that can act inside the system where the work happens are the ones people keep opening. Assistants that answer questions well and do nothing else decay to single-digit weekly use within a month. Model quality barely shows up in the retention data. The lab has not run an experiment; what we hold is telemetry from three engagements, one product engineer's observation, and a set of vendor and academic signals that point the same way.

Written by hand and carrying nobody's name. Editing it puts yours on it.

Why a Quantium decision hinges on it

unattributed

Quantium has sold, built or advised on more than a dozen assistant deployments and the firm's reputation is attached to whether they are used. A model recommendation does not move week-three retention; a product recommendation does. If the four properties hold up under measurement, they become the checklist for every assistant engagement and a standing answer for the question clients actually ask. The gate is adoption: the properties are cheap to build and organisations still do not ask for them, because the procurement conversation is about the model.

Written by hand and carrying nobody's name. Editing it puts yours on it.

What it actually is

composed from the records

Whether an assistant is still being used in week three is decided by four properties — proactivity, calibration, refusal honesty and integration depth — and none of them is a model property; they are product decisions that most deployments get wrong by default. That is the lab's one-line position on it, which is not the same as an explanation.

The shape the field is converging on, from the most authoritative source in it: Calibrated uncertainty sustains use; miscalibrated uncertainty destroys it faster than overconfidence does.signal

This is the section a page most needs a person for, and the one composition is worst at. Nobody has written the plain-language version — what the idea is, in words that assume nothing — and it is the first thing a reader who has never met the term needs.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

Why now

composed from the records

The lab opened this field on 2026-02-03, and it has not reached mainstream awareness yet. Everything below is what has moved since.

What shipped: Enterprise assistant vendor publishes its own retention cohort analysis (Vendor blog, 2026-06-05). Tooling arriving is what moves a field from argument to something a team could try.signal

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

What it changes in a system

composed from the records

What changes, concretely: Across three engagements' telemetry, assistants with at least one proactive surface retained 3.1× the weekly active users at week eight of assistants that only answered when asked.

Nothing is shipped as a default yet, so adopting this is a piece of work rather than a configuration change. That is usually the difference between a field being interesting and being used.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

What is in the way

composed from the records

The binding constraint is adoption: it is ready, and trust, regulation, procurement or change capacity are what is left. Everything upstream of that is solved and everything downstream of it is waiting.

Workforce readiness is medium: Delivery teams can build all four properties; they are not asked to. A recommendation needing skills the firm does not hold is an aspiration rather than an action, and it routes to the enablement agenda instead of the delivery one.

The argued case against it is the red team's, further down this page, and it is deliberately one-sided — this section is what stands in the way mechanically, not what somebody thinks of it.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

2 of 6 explanatory sections are written; the rest are composed until somebody takes them.

Business priority

On the plan
  • Internal adoption of shipped defaults is a named measure, and week-three abandonment is what defeats it.

    HB Harley Barnescommitted

Clients are asking

  • Can you build us a chatbot for customer service?2 engagements · <$250k · falling

Priority orders what you see. It never changes what the evidence says — a plan-critical field with nothing tested is still signal tier.

Field attributes

StateEmerging
GateAdoption · ready — blocked by trust, regulation, procurement or change capacity
OriginQuestion
Measurablefull
Written forexec
Reach · TLPThe firm · TLP:AMBER
Horizonnear
Opened3 Feb 2026
Mainstreamnot yet
Last validated26 Aug 2026
Sightings1

What people have written

Write one

Nothing yet. The person who knows a claim is wrong is usually not the person who wrote it.

A note never travels further than the thing it is written on.

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01Across three engagements' telemetry, assistants with at least one proactive surface retained 3.1× the weekly active users at week eight of assistants that only answered when asked.
  • 02A product engineer added a visible 'I am not confident about this' state to an internal assistant and weekly use rose 40% over the following month; one assistant, self-reported (s-effective-assistants-finding-ollie).
  • 03Integration depth predicts retention better than model choice in every dataset we have seen, including a vendor's own published cohort analysis.
What is hype
  • 01Model upgrades as the fix for retention. Two engagements upgraded to a better model with no measurable change in week-eight use.
  • 02'Copilot everywhere' — an assistant in every screen. Breadth without depth in any one workflow is the pattern most likely to decay.
  • 03Engagement metrics that count opens. An assistant people open, ask, and distrust is not being used.
What would have to be true
  • 01The four properties have to hold in a controlled comparison, not just in telemetry from engagements that differ in twenty other ways.
  • 02Proactivity has to be tunable without becoming noise; the sensing-agent evidence says unwanted interruptions get a tool switched off.
  • 03Calibration has to be honest, which needs the assistant to know when it is wrong — a capability that is partial at best in current models.
What we would do
  • 01Write the four properties into a standing answer for 'why did usage collapse' and use it on the next three assistant engagements.
  • 02Propose a Type 3 experiment with two matched assistants differing only in the four properties, retention at week eight as the measure.
  • 03Move the field to assessed once a validation run has gone through the academic and vendor evidence properly; it is running on telemetry and instinct.

Signals · 9 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
3.1×
week-8 WAU
Finding·band 1Tried

Week-eight retention across three assistant engagements: proactive and integrated wins 3.1×

Pooled usage telemetry from three engagements with client permission. Assistants with a proactive surface and write access to the work system held 3.1× the weekly active users at week eight. Uncontrolled; the engagements differ in team and budget. Tried tier because no protocol.

extracted claimProactivity and integration depth are the strongest correlates of week-eight assistant retention in the telemetry we hold.
Lab · Nightingale measurement (Engel telemetry) · Andrew Tran31 Jul 2026
ATdropped
1,200
respondents
Paper·band 1Signal

Calibrated Refusal and Sustained Use of Workplace Language Assistants

Field study across 1,200 knowledge workers. Assistants that expressed calibrated uncertainty and refused visibly retained users better than confident-always variants at twelve weeks; the effect reversed when the uncertainty was miscalibrated.

extracted claimCalibrated uncertainty sustains use; miscalibrated uncertainty destroys it faster than overconfidence does.
arxiv.org · Whitfield, Nkemelu et al.14 Aug 2026
MMdropped 2
~2 / day
ceiling
Paper·band 1Signal

The Interruption Ceiling: Proactive Assistant Behaviour and Mute Rates

Varies proactive-message frequency and measures mute rates. Retention rises to about two unsolicited messages a day and falls sharply beyond; the curve is not monotonic.

arxiv.org · Sørensen, Achebe et al.24 Jul 2026
detector · bleeding edge
Drop·band 3Signal

Slack drop: 'the RFP scores the model on 40 points and the integration on 5'

A sector owner's note from an RFP evaluation. The weighting is the adoption gate in one line: clients procure for the thing that does not predict retention.

Slack drop10 Jul 2026
AHdropped
+40%
weekly use
Finding·band 1Tried

Logged from Claude Code: visible 'not confident' state raised weekly use 40%

Product engineer added an explicit low-confidence state to an internal assistant and watched weekly use rise 40% over the next month. One assistant, one team, no control. Attributed, tried, decays fast.

MCP · log_finding · Oliver Vu27 Jun 2026
OVdropped
Release·band 2Signal

Enterprise assistant vendor publishes its own retention cohort analysis

Vendor shows week-twelve retention split by integration depth and finds the same ordering we do. Self-interested — it sells integrations — but the cohort sizes are large and the direction matches independent data.

Vendor blog5 Jun 2026
detector · early adoption 2
Post·band 3Signal

'We upgraded the model and nothing changed'

Short, widely shared post. Upgraded an internal assistant to a frontier model; usage did not move. The comments are a catalogue of the same experience. Kept as the demand-band echo of the telemetry.

LinkedIn · A head of digital at an AU insurer2 May 2026
MA?dropped 3
Client question·band 3Signal

'We rolled it out to 4,000 people and 180 still use it. What did we do wrong?'

The question that opened the field, in the words of a chief operating officer. Eleven engagements have asked a version of it since. The assistant in question answered well, raised nothing, and lived in a separate tab.

Engel · banking engagement9 Apr 2026
AHSLdropped 4
61%
below 10% WAU
Analyst·band 3Signal

Enterprise copilot survey: 61% of deployments below 10% weekly active use at six months

Survey of 400 enterprises. Establishes the base rate for decay. The report attributes it to 'change management'; the telemetry says product properties.

Forrester18 Mar 2026
HBdropped 2
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 5 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

Organisations do not procure for the four properties because the procurement conversation is about the model; the gate is adoption, not capability.

Triedc-effective-assistants-5dalton-0.431 Jul 2026Engel · banking engagement, Slack drop
70%

Integration depth in one workflow predicts retention better than model choice or breadth of deployment.

Triedc-effective-assistants-3dalton-0.431 Jul 2026Lab · Nightingale measurement (Engel telemetry), Vendor blog, Engel · banking engagement
66%

Assistants with a proactive surface retain roughly three times the weekly active users at week eight of assistants that only respond to prompts.

Triedc-effective-assistants-1dalton-0.426 Aug 2026Lab · Nightingale measurement (Engel telemetry), Forrester
62%

Visible uncertainty and honest refusal increase use rather than reduce it; people return to an assistant that tells them when it does not know.

Triedc-effective-assistants-2dalton-0.426 Aug 2026MCP · log_finding, arxiv.org
58%

Proactivity above a low threshold becomes interruption and drives the assistant to be muted; the curve is not monotonic.

Signalc-effective-assistants-6dalton-0.426 Aug 2026arxiv.org
50%

Upgrading to a stronger model materially improves week-eight retention.

Triedc-effective-assistants-4dalton-0.314 May 2026Lab · Nightingale measurement (Engel telemetry), LinkedIn
20%

Position history · the diff is the product

4 validation runs against a fixed brief. Confidence 38% → 60%.

runs compare claim sets, never prose
What we said · run 4

Calibration claim strengthened by a paper and a second engagement. Proactivity has a ceiling; above it the assistant gets muted. Standing answer drafted; experiment proposed.

60%
Changed since run 3
  • Visible uncertainty and honest refusal increase use rather than reduce it; people return to an assistant that tells them when it does not know.
  • Proactivity above a low threshold becomes interruption and drives the assistant to be muted; the curve is not monotonic.
  • c-effective-assistants-3 ↑ 0.58 → 0.66
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · AW
high

Every assistant engagement the firm has sold is exposed to week-three decay.

Timeline

committed · MM
0–18mo

Nothing here needs a capability that does not exist; it needs a checklist and a will to use it.

Demand

committed · AH
high

Eleven engagements asked why usage collapsed. The most-asked question in Engel this half.

Cost of being wrong

agent-estimated
medium

A wrong checklist costs an engagement's credibility, not a client's business. Agent-estimated.

TAM

agent-estimated
>$10B

Agent-estimated from enterprise assistant spend; the field is about capturing value already being spent. Uncommitted.

Workforce readiness

committed · MM
medium

Delivery teams can build all four properties; they are not asked to.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Banking
relevant

Two internal knowledge assistants in banking clients decayed to under 5% weekly use; both answered well and did nothing else.

Mechanism · Proactive surfacing of policy changes; visible confidence on regulatory answers; write access to the case system.

AH committed by Amber Hallcommitted · AH
Health
relevant

Clinical admin assistants are the highest-stakes case for refusal honesty; a confabulated answer about a medication is a safety event.

Mechanism · Refusal-first design with a named human fallback; calibration shown on every clinical answer.

SL committed by Sylvia Liucommitted · SL
Retail & FMCG
relevant

Merchandising copilots that live outside the planning tool are not opened; the one inside it is.

Mechanism · Integration into the range-planning workflow rather than a separate chat surface.

Agent draft · awaiting a sector owneragent-estimated
Government
watch

Proactivity in a public-sector assistant raises a records and accountability question — who authorised what it raised — that has not been answered.

Mechanism · Would need an accountability model for unsolicited output.

AV committed by Aadhithyanarayanan V Acommitted · AV

Red team · the strongest case against

The strongest case against: the four properties are a post-hoc story fitted to three engagements' telemetry. Assistants with proactive surfaces and deep integration are also the assistants that had better product teams, more budget and executive sponsorship, and those explain retention on their own. The field may be describing good product management and calling it a research finding.

  • No controlled comparison exists. Every retention difference we cite is confounded by team quality, budget and sponsorship.
  • The calibration result rests on one assistant and one engineer's report; the academic evidence on calibration and trust is mixed and some of it points the other way.
  • Retention is not value. An assistant people keep opening because it interrupts them usefully may be costing more attention than it saves; nobody has measured that.
Run by an agent briefed to argue the field is nothing — sources here are correlated, and without a deliberate adversary synthesis converges on consensus and calls it insight. Kept as a dated pass rather than overwritten. Nobody has answered it yet, and a challenge nobody answers is a disclaimer.thesis weakened

Source diversity

  • ML / HCI research30%
  • Vendor10%
  • Practitioner / LinkedIn15%
  • Analyst10%
  • Internal / Engel35%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

No experiments, recommendations or graveyard entries yet. That is what a candidate looks like.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01In a controlled comparison, how much of the retention gap survives once team quality and sponsorship are held constant?
  2. 02What is the interruption ceiling for a clinical or banking user, and is it the same as for a knowledge worker?
  3. 03Is retention the right measure, or is it attention cost per useful action?

Notes · anyone in the firm

What people have written on this.

The person who knows a claim is wrong is usually not the person who wrote it. Corrections, objections and questions are owed an answer and stay open until the field owner says what they did; context and use notes stand as they are.

Notes · 0

Anything here reaches at most the firm — a note cannot travel further than what it is written on.

    Nothing written on this yet. The useful notes are the ones from people who are not in the lab — that is where the correction usually comes from.