cavendish
TestedContestedgate · ReliabilityNow · 0–12 months×4 sightings

Voice and Vision

Realtime voice agents in Australian contact centres are gated by reliability, not capability: a round-trip floor from Sydney to US-hosted realtime models, accent-driven recognition errors, barge-in failures and hallucinated confirmations mean they hold for tier-1 triage and fail above it; vision in field operations and retail is further along because latency is not the product.

Experiment run, measured result. The only tier that becomes a recommendation.

Join with…

Confidence

63%human-committed

Expiry

28duntil review · 1 Oct 2026

Lead time

−4moopened after mainstream — recorded honestly

Ownership

DDDavid Diazfortnightly cadence

Where it is

unattributed

The field is contested because the two halves are moving at different speeds. Vision on structured imagery — meter photos, shelf gaps, claim forms — is at production accuracy on frontier models and on a distilled small model, and the retail sector owner is already using it. Voice is not. Our latency run for an AU telco measured a p95 turn latency of 2.6 seconds via US-hosted realtime endpoints, well past the point at which callers talk over the agent; a Sydney-region endpoint launched in July halves that but does not fix barge-in, which failed on 18% of interruptions, or hallucinated confirmations, which occurred in 4 of 50 test calls. Vendors claim 70% containment; the analyst line says 40% of AU volume by 2027. Both may be true for tier-1 queues and neither is true for anything transactional. The recommendation is 'not yet' above triage, and the graveyard holds the IVR-replacement thesis.

Written by hand and carrying nobody's name. Editing it puts yours on it.

Why a Quantium decision hinges on it

unattributed

Telco and insurance clients are being pitched voice agents by every contact-centre vendor, and the pitch numbers are containment rates measured on US English in US data centres. Quantium's edge is a measured AU floor: the latency, the accent error rate and the confirmation-hallucination rate on real queues. That number is the difference between a triage deployment that works and a payments deployment that generates complaints to the TIO. Vision is the opposite story — a quiet production capability in field ops and retail that nobody is asking the lab about because it already works.

Written by hand and carrying nobody's name. Editing it puts yours on it.

What it actually is

composed from the records

Realtime voice agents in Australian contact centres are gated by reliability, not capability: a round-trip floor from Sydney to US-hosted realtime models, accent-driven recognition errors, barge-in failures and hallucinated confirmations mean they hold for tier-1 triage and fail above it; vision in field operations and retail is further along because latency is not the product. That is the lab's one-line position on it, which is not the same as an explanation.

The shape the field is converging on, from the most authoritative source in it: Recognition error on Australian-accented English is roughly twice the US baseline on current realtime stacks.signal

This is the section a page most needs a person for, and the one composition is worst at. Nobody has written the plain-language version — what the idea is, in words that assume nothing — and it is the first thing a reader who has never met the term needs.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

Why now

composed from the records

The lab opened this field on 2026-03-12 and it reached mainstream awareness on 2025-11-01. The gap between those two dates is the lead time the lab is measured on.

What shipped: Frontier lab launches Sydney-region realtime voice endpoint (OpenAI, 2026-07-15) and Contact-centre platform vendor claims 70% containment with voice agents (Vendor press release, 2026-05-06). Tooling arriving is what moves a field from argument to something a team could try.signalsignal

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

What it changes in a system

composed from the records

What changes, concretely: AU voice latency floor: p50 1.4 s and p95 2.6 s turn latency via US-hosted realtime endpoints; 0.78 s p50 via the Sydney-region endpoint launched in July (x-voice-latency).signalsignal

There is already a shipped default — voice.scope = tier1-triage · voice.endpoint = au-region · voice.readback = grounded-to-tool-output · handoff.on = transactional-intent · skill: cavendish/voice-guardrails — so a team adopting this is changing a setting rather than starting a project.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

What is in the way

composed from the records

The binding constraint is reliability: it works and is not yet dependable enough to put near a customer. Everything upstream of that is solved and everything downstream of it is waiting.

Workforce readiness is low: Nobody in delivery has built a barge-in-safe voice loop; the vendor stacks hide it until it fails. Agent-estimated. A recommendation needing skills the firm does not hold is an aspiration rather than an action, and it routes to the enablement agenda instead of the delivery one.

Something here has already been killed: Realtime voice agents replace tier-1 IVR, on the telco partner withdrew the test trunk on day 24 of a planned 40-call measurement and the pair was reassigned; 11 calls were measured, which is not enough to answer the hypothesis either way. It is marked resurrectable on: A telco or contact-centre partner lending a SIP trunk for three uninterrupted weeks, or an AU-region realtime endpoint from a second vendor so the comparison can run without one..graveyard

The argued case against it is the red team's, further down this page, and it is deliberately one-sided — this section is what stands in the way mechanically, not what somebody thinks of it.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

2 of 6 explanatory sections are written; the rest are composed until somebody takes them.

Business priority

Plan critical

Clients are asking

  • “Can you build us a chatbot for customer service?”2 engagements · <$250k · falling

Priority orders what you see. It never changes what the evidence says — a plan-critical field with nothing tested is still signal tier.

Field attributes

StateContested
GateReliability · possible, not yet dependable enough
OriginQuestion
Measurablepartial
Written forpractice
Reach · TLPClient-safe · TLP:GREEN
Horizonnow
Opened12 Mar 2026
Mainstream1 Nov 2025
Last validated19 Aug 2026
Sightings4

What people have written

Write one

Nothing yet. The person who knows a claim is wrong is usually not the person who wrote it.

A note never travels further than the thing it is written on.

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01AU voice latency floor: p50 1.4 s and p95 2.6 s turn latency via US-hosted realtime endpoints; 0.78 s p50 via the Sydney-region endpoint launched in July (x-voice-latency).
  • 02Barge-in failed on 18% of caller interruptions across three stacks; abandonment correlated with barge-in failure more strongly than with recognition errors.
  • 03Hallucinated confirmations — the agent reading back a reference number, amount or date that did not exist — in 4 of 50 scripted test calls. Rules out unsupervised transactional use.
  • 04Vision on structured field imagery (meter dials, shelf gaps, handwritten forms) at 96% task accuracy on frontier models; a distilled 3B model reached 94% on shelf gaps on in-store hardware.
What is hype
  • 01'70% containment' from contact-centre vendors. Measured on tier-1 US-English queues with generous definitions of contained; nobody publishes the AU number.
  • 02'Voice AI handles 40% of AU volume by 2027.' Possible for triage; the transactional share of volume does not move on current reliability.
  • 03'Human-level' ASR. True for a US newsreader; error rates on Australian-accented and non-native English were 2.1× the US baseline across the stacks we tested.
What would have to be true
  • 01A regional realtime endpoint with barge-in handled at the edge, holding p95 under 1.2 s from Sydney on a real queue rather than a test rig.
  • 02A confirmation-grounding pattern — the agent may only read back values it retrieved, never generated — that survives a full call without a latency cost.
  • 03AU-accent recognition error within 1.2× of the US baseline; currently 2.1×.
What we would do
  • 01Keep r-voice-not-yet as the position for anything above tier-1 triage; revisit in Q4 with the regional endpoint on a live queue.
  • 02Ship vision on structured imagery as a default pattern for field ops and store audit; it does not need the lab any more.
  • 03Run the latency bench again against the regional endpoint with barge-in at the edge, and add the confirmation-grounding pattern as the intervention.

Signals · 10 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
2.6 s
p95 turn latency (US-hosted)
Finding·band 1Tested

Voice latency floor for AU telco: p95 2.6 s via US endpoints, barge-in fails 18%

Three realtime stacks driven from a Sydney telephony rig with scripted callers and timed interruptions. US-hosted endpoints gave p50 1.4 s and p95 2.6 s turn latency; the regional endpoint gave p50 0.78 s. Barge-in failed on 18% of interruptions and predicted abandonment better than recognition errors did.

extracted claimFrom Sydney, US-hosted realtime voice exceeds the talk-over threshold at p95; barge-in failure, not recognition, drives abandonment.
Lab · x-voice-latency · David Diaz19 Aug 2026
detector · bleeding edge
0.78 s
p50 turn latency (regional)
Release·band 1Tested

Frontier lab launches Sydney-region realtime voice endpoint

Realtime speech-to-speech API available in an AU region with data residency. Arrived mid-experiment; we added it as a fourth arm. Halves the p50 and brings p95 under 1.5 s on the rig. Does not change barge-in behaviour.

extracted claimA regional realtime endpoint removes roughly half the turn latency from Sydney.
OpenAI15 Jul 2026
detector · bleeding edge 3
4 / 50
hallucinated confirmations
Finding·band 1Tried

Logged from Claude Code: voice agent read back a booking reference that did not exist in 4 of 50 calls

During harness setup the agent confidently confirmed reference numbers and amounts it had not retrieved. Grounding the read-back to tool output removed it in a follow-up run of twenty calls. Logged as tried; became a bench arm.

MCP · log_finding · David Diaz24 Jun 2026
DDdropped
Post·band 2Signal

'The hardest part of voice is knowing when the human stopped talking'

Practitioner post on endpointing and barge-in: the model is rarely the problem; the turn-detection layer is. Matches our bench, where barge-in failure predicted abandonment. Widely shared in the infra community.

Engineering blog · A voice-infrastructure engineer2 Jun 2026
OVdropped 3
70%
claimed containment
Announcement·band 3Signal

Contact-centre platform vendor claims 70% containment with voice agents

Headline containment figure from US-English tier-1 deployments; 'contained' includes calls that ended after the agent read the account balance. No AU customer named. Carried as the claim our bench contradicts.

Vendor press release6 May 2026
detector · demand 4
Client question·band 3Signal

'Can the voice bot take a payment over the phone without a human on the line?'

Asked by a contact-centre GM after a vendor demo. The delivery lead did not have an answer and logged it. Became the transactional-intent arm of the bench and the reason the recommendation draws its line at triage.

Engel · telco engagement30 Apr 2026
MAdropped 3
96%
meter-read accuracy
Benchmark·band 2Signal

Field-imagery benchmark: meter dials, forms and asset photos

Public multimodal benchmark of structured field imagery. Frontier models above 96% task accuracy on meter reads and handwritten forms; the energy sector's use case is solved at the model level and blocked on integration only.

huggingface.co9 Apr 2026
detector · early adoption 2
2.1×
WER vs US baseline
Paper·band 1Signal

Accent-conditioned error rates in realtime speech models: an Australian and non-native English study

Measures word error rate across three realtime stacks on Australian-accented, Indian-English and Mandarin-accented English against a US baseline. AU-accented error rate was 2.1× baseline; the gap narrowed but did not close with vendor accent options.

arxiv.org · Whitmore, Nguyen et al.20 Mar 2026
detector · bleeding edge 2
6.1k
stars
Repository·band 2Tried

shelfsight — open vision-language model for shelf-gap and planogram detection

Small open VLM fine-tuned on retail shelf imagery. The retail sector owner ran it on store photos during a Woolworths pilot and it matched the frontier model on gap detection. Basis for the edge classifier work in the SLM field.

github.com11 Feb 2026
DBdropped 2
40%
forecast share by 2027
Analyst·band 3Signal

'Voice AI will handle 40% of Australian contact-centre volume by 2027'

Demand-band forecast that appears in every vendor deck since. Plausible for triage volume; no reliability caveat. Included as the mainstream framing.

IDC28 Jan 2026
detector · demand 5
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 4 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

From Sydney, p95 turn latency on US-hosted realtime models exceeds the 1.5 s threshold at which callers talk over the agent; a regional endpoint roughly halves it.

Testedc-voice-and-vision-1dalton-0.419 Aug 2026Lab · x-voice-latency, OpenAI
86%

Vision on structured field imagery is production-ready now; the reliability gate is specific to voice.

Assessedc-voice-and-vision-4dalton-0.314 May 2026huggingface.co, github.com
79%

Barge-in failure, not recognition accuracy, is the dominant cause of caller abandonment in AU voice-agent trials.

Testedc-voice-and-vision-2dalton-0.419 Aug 2026Lab · x-voice-latency, Engineering blog
74%

Voice agents hallucinate confirmations at a rate (5–8% of calls) that rules out unsupervised transactional use.

Testedc-voice-and-vision-3dalton-0.42 Jul 2026MCP · log_finding, Lab · x-voice-latency
70%

Vendor-reported 70% containment generalises to Australian queues once the accent model is tuned.

Assessedc-voice-and-vision-5dalton-0.42 Jul 2026Vendor press release, arxiv.org
25%

Position history · the diff is the product

4 validation runs against a fixed brief. Confidence 42% → 63%.

runs compare claim sets, never prose
What we said · run 4

Latency bench concluded. US-hosted floor is unusable; regional endpoint halves it; barge-in is the abandonment driver. Recommendation published; IVR-replacement thesis to the graveyard. Field stays contested because the floor is moving.

63%
Changed since run 3
  • From Sydney, p95 turn latency on US-hosted realtime models exceeds the 1.5 s threshold at which callers talk over the agent; a regional endpoint roughly halves it.
  • Barge-in failure, not recognition accuracy, is the dominant cause of caller abandonment in AU voice-agent trials.
  • c-voice-and-vision-3 ↑ 0.62 → 0.70
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · MA
high

Contact-centre cost is the largest AI line item in every telco and insurance pipeline we have.

Timeline

committed · DD
0–18mo

Triage is deployable now; transactional depends on grounding and regional endpoints, both inside the window.

TAM

agent-estimated
$1B–10B

Agent-estimated from AU contact-centre labour spend addressable by tier-1 automation. Uncommitted.

Demand

committed · MA
high

Every telco engagement this half has asked; two have live vendor pilots we are being asked to assess.

Cost

committed · DD
medium

The bench needs a telephony rig and scripted callers; two engineers, three weeks per cycle.

Cost of being wrong

committed · TB
high

A hallucinated payment confirmation on a regulated queue is a complaint to the TIO or AFCA, not a bug.

Workforce readiness

agent-estimated
low

Nobody in delivery has built a barge-in-safe voice loop; the vendor stacks hide it until it fails. Agent-estimated.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Telco
relevant

The largest queues in the country and the most vendor pressure. The AU latency and accent numbers are the assessment they are asking us for.

Mechanism · Tier-1 triage with a regional endpoint and human handoff on any transactional intent; measure barge-in and confirmation errors on the live queue.

MA committed by Mario Attardcommitted · MA
Retail & FMCG
relevant

Vision, not voice. Shelf-gap and planogram compliance from store photos is already in a Woolworths pilot.

Mechanism · Distilled vision model on in-store hardware; frontier model for exception review.

DB committed by Dillon Blakecommitted · DB
Energy & Utilities
relevant

Meter-reading and asset-inspection photos are structured imagery where frontier accuracy is already sufficient.

Mechanism · Batch vision over field-app uploads; no latency constraint, so no gate. Agent draft.

Agent draft · awaiting a sector owneragent-estimated
Insurance
watch

Claims intake by voice is attractive and the confirmation-hallucination rate makes it dangerous; vision on claim photos is fine.

Mechanism · Vision now; voice only after the grounding pattern is measured. Agent draft.

Agent draft · awaiting a sector owneragent-estimated

Red team · the strongest case against

The strongest case against: we measured a floor that is already moving under us. The regional endpoint arrived mid-experiment, barge-in is a stack-engineering problem the vendors will fix within a release or two, and confirmation grounding is a prompt-and-tool pattern, not a research problem. A 'not yet' that expires in a quarter is a 'yes' with extra steps, and clients who wait on our advice will be a quarter behind the ones who did not.

  • —The 2.6 s p95 was measured on US endpoints that no serious AU deployment would use after July. Our headline number is already historical.
  • —Fifty scripted test calls is a small sample for a 4-in-50 hallucination rate; the confidence interval includes rates that are acceptable with a human check.
  • —Contained tier-1 calls are most of the volume in most queues. 'Holds for triage' may be 80% of the business case, which makes 'not yet' the wrong headline.
  • —The vision half is not contested by anyone and does not need a field; keeping it here inflates the field's confidence.
Run by an agent briefed to argue the field is nothing — sources here are correlated, and without a deliberate adversary synthesis converges on consensus and calls it insight. Kept as a dated pass rather than overwritten. Nobody has answered it yet, and a challenge nobody answers is a disclaimer.thesis weakened

Source diversity

  • Speech and ML research25%
  • Voice infra practitioners20%
  • Vendor and analyst25%
  • Internal / Engel30%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Convergence · who else is here

Several people’s drops meet here. An informal working group already exists and probably does not know it.

ContributorsDDDYMADBOVTB

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01What is the p95 on a live queue, not a rig, with the regional endpoint and barge-in at the edge?
  2. 02Does confirmation grounding survive a full call without adding a turn of latency?
  3. 03Should vision leave this field and go to practice now, given nobody contests it?

Tools in this space · 4

What you could actually buy.

Products aimed at this field, with what the lab has behind each one. Scored on six axes and never summed — “which is better” is not a question anyone has, and the constraint that decides it is named beside every assessment.

  • LiveKit AgentsLiveKitAssessedAssessed

    Realtime media infrastructure with an agent framework on top, self-hostable. The answer when the round trip has to stay in-country.

    SaaS or self-hostApache-2.0 core, commercial cloudseen 2026-01-20

    No commercial relationship.

  • DeepgramDeepgramSignalWatching

    Speech recognition with the latency profile realtime work needs, and a self-host option that most competitors do not offer.

    SaaS or self-hostCommercialseen 2025-12-15

    No commercial relationship.

  • ElevenLabsElevenLabsSignalWatching

    Speech synthesis good enough that quality has stopped being the constraint. Which moves the argument to latency, cost and consent, where it belongs.

    SaaS onlyCommercialseen 2025-10-27

    No commercial relationship.

  • VapiVapiSignalWatching

    Orchestration for realtime voice agents — turn-taking, interruption, telephony. The part everyone underestimates until they build it themselves.

    SaaS onlyCommercialseen 2026-02-24

    No commercial relationship.

Listing is not recommending. Most of Voice and Vision sits at signal tier — in the space, nothing behind it — and a tool only reaches tested when a run stands behind it. Vendor pricing and capability move monthly, so these carry the shortest half-life in the library, and anyone who cited one gets told when it moves.

Notes · anyone in the firm

What people have written on this.

The person who knows a claim is wrong is usually not the person who wrote it. Corrections, objections and questions are owed an answer and stay open until the field owner says what they did; context and use notes stand as they are.

Notes · 0

Anything here reaches at most client-safe — a note cannot travel further than what it is written on.

    Nothing written on this yet. The useful notes are the ones from people who are not in the lab — that is where the correction usually comes from.