Voice and Vision
Realtime voice agents in Australian contact centres are gated by reliability, not capability: a round-trip floor from Sydney to US-hosted realtime models, accent-driven recognition errors, barge-in failures and hallucinated confirmations mean they hold for tier-1 triage and fail above it; vision in field operations and retail is further along because latency is not the product.
Experiment run, measured result. The only tier that becomes a recommendation.
Confidence
63%human-committedExpiry
28duntil review · 1 Oct 2026Lead time
−4moopened after mainstream — recorded honestlyOwnership
DDDavid Diazfortnightly cadenceWhere it is
unattributedThe field is contested because the two halves are moving at different speeds. Vision on structured imagery — meter photos, shelf gaps, claim forms — is at production accuracy on frontier models and on a distilled small model, and the retail sector owner is already using it. Voice is not. Our latency run for an AU telco measured a p95 turn latency of 2.6 seconds via US-hosted realtime endpoints, well past the point at which callers talk over the agent; a Sydney-region endpoint launched in July halves that but does not fix barge-in, which failed on 18% of interruptions, or hallucinated confirmations, which occurred in 4 of 50 test calls. Vendors claim 70% containment; the analyst line says 40% of AU volume by 2027. Both may be true for tier-1 queues and neither is true for anything transactional. The recommendation is 'not yet' above triage, and the graveyard holds the IVR-replacement thesis.
Why a Quantium decision hinges on it
unattributedTelco and insurance clients are being pitched voice agents by every contact-centre vendor, and the pitch numbers are containment rates measured on US English in US data centres. Quantium's edge is a measured AU floor: the latency, the accent error rate and the confirmation-hallucination rate on real queues. That number is the difference between a triage deployment that works and a payments deployment that generates complaints to the TIO. Vision is the opposite story — a quiet production capability in field ops and retail that nobody is asking the lab about because it already works.
What it actually is
composed from the recordsRealtime voice agents in Australian contact centres are gated by reliability, not capability: a round-trip floor from Sydney to US-hosted realtime models, accent-driven recognition errors, barge-in failures and hallucinated confirmations mean they hold for tier-1 triage and fail above it; vision in field operations and retail is further along because latency is not the product. That is the lab's one-line position on it, which is not the same as an explanation.
The shape the field is converging on, from the most authoritative source in it: Recognition error on Australian-accented English is roughly twice the US baseline on current realtime stacks.signal
This is the section a page most needs a person for, and the one composition is worst at. Nobody has written the plain-language version — what the idea is, in words that assume nothing — and it is the first thing a reader who has never met the term needs.
Why now
composed from the recordsThe lab opened this field on 2026-03-12 and it reached mainstream awareness on 2025-11-01. The gap between those two dates is the lead time the lab is measured on.
What shipped: Frontier lab launches Sydney-region realtime voice endpoint (OpenAI, 2026-07-15) and Contact-centre platform vendor claims 70% containment with voice agents (Vendor press release, 2026-05-06). Tooling arriving is what moves a field from argument to something a team could try.signalsignal
What it changes in a system
composed from the recordsWhat changes, concretely: AU voice latency floor: p50 1.4 s and p95 2.6 s turn latency via US-hosted realtime endpoints; 0.78 s p50 via the Sydney-region endpoint launched in July (x-voice-latency).signalsignal
There is already a shipped default — voice.scope = tier1-triage · voice.endpoint = au-region · voice.readback = grounded-to-tool-output · handoff.on = transactional-intent · skill: cavendish/voice-guardrails — so a team adopting this is changing a setting rather than starting a project.
What is in the way
composed from the recordsThe binding constraint is reliability: it works and is not yet dependable enough to put near a customer. Everything upstream of that is solved and everything downstream of it is waiting.
Workforce readiness is low: Nobody in delivery has built a barge-in-safe voice loop; the vendor stacks hide it until it fails. Agent-estimated. A recommendation needing skills the firm does not hold is an aspiration rather than an action, and it routes to the enablement agenda instead of the delivery one.
Something here has already been killed: Realtime voice agents replace tier-1 IVR, on the telco partner withdrew the test trunk on day 24 of a planned 40-call measurement and the pair was reassigned; 11 calls were measured, which is not enough to answer the hypothesis either way. It is marked resurrectable on: A telco or contact-centre partner lending a SIP trunk for three uninterrupted weeks, or an AU-region realtime endpoint from a second vendor so the comparison can run without one..graveyard
The argued case against it is the red team's, further down this page, and it is deliberately one-sided — this section is what stands in the way mechanically, not what somebody thinks of it.
2 of 6 explanatory sections are written; the rest are composed until somebody takes them.
Business priority
Plan criticalClients are asking
- “Can you build us a chatbot for customer service?”2 engagements · <$250k · falling
Priority orders what you see. It never changes what the evidence says — a plan-critical field with nothing tested is still signal tier.
Field attributes
What people have written
Write oneNothing yet. The person who knows a claim is wrong is usually not the person who wrote it.
A note never travels further than the thing it is written on.
Position
What is demonstrated, what is hype, what would have to be true.
The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.
- 01AU voice latency floor: p50 1.4 s and p95 2.6 s turn latency via US-hosted realtime endpoints; 0.78 s p50 via the Sydney-region endpoint launched in July (x-voice-latency).
- 02Barge-in failed on 18% of caller interruptions across three stacks; abandonment correlated with barge-in failure more strongly than with recognition errors.
- 03Hallucinated confirmations — the agent reading back a reference number, amount or date that did not exist — in 4 of 50 scripted test calls. Rules out unsupervised transactional use.
- 04Vision on structured field imagery (meter dials, shelf gaps, handwritten forms) at 96% task accuracy on frontier models; a distilled 3B model reached 94% on shelf gaps on in-store hardware.
- 01'70% containment' from contact-centre vendors. Measured on tier-1 US-English queues with generous definitions of contained; nobody publishes the AU number.
- 02'Voice AI handles 40% of AU volume by 2027.' Possible for triage; the transactional share of volume does not move on current reliability.
- 03'Human-level' ASR. True for a US newsreader; error rates on Australian-accented and non-native English were 2.1× the US baseline across the stacks we tested.
- 01A regional realtime endpoint with barge-in handled at the edge, holding p95 under 1.2 s from Sydney on a real queue rather than a test rig.
- 02A confirmation-grounding pattern — the agent may only read back values it retrieved, never generated — that survives a full call without a latency cost.
- 03AU-accent recognition error within 1.2× of the US baseline; currently 2.1×.
- 01Keep r-voice-not-yet as the position for anything above tier-1 triage; revisit in Q4 with the regional endpoint on a live queue.
- 02Ship vision on structured imagery as a default pattern for field ops and store audit; it does not need the lab any more.
- 03Run the latency bench again against the regional endpoint with barge-in at the edge, and add the confirmation-grounding pattern as the intervention.
Signals · 10 in this cluster
What the cluster is made of.
Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

Voice latency floor for AU telco: p95 2.6 s via US endpoints, barge-in fails 18%
Three realtime stacks driven from a Sydney telephony rig with scripted callers and timed interruptions. US-hosted endpoints gave p50 1.4 s and p95 2.6 s turn latency; the regional endpoint gave p50 0.78 s. Barge-in failed on 18% of interruptions and predicted abandonment better than recognition errors did.
extracted claimFrom Sydney, US-hosted realtime voice exceeds the talk-over threshold at p95; barge-in failure, not recognition, drives abandonment.

Frontier lab launches Sydney-region realtime voice endpoint
Realtime speech-to-speech API available in an AU region with data residency. Arrived mid-experiment; we added it as a fourth arm. Halves the p50 and brings p95 under 1.5 s on the rig. Does not change barge-in behaviour.
extracted claimA regional realtime endpoint removes roughly half the turn latency from Sydney.

Logged from Claude Code: voice agent read back a booking reference that did not exist in 4 of 50 calls
During harness setup the agent confidently confirmed reference numbers and amounts it had not retrieved. Grounding the read-back to tool output removed it in a follow-up run of twenty calls. Logged as tried; became a bench arm.

'The hardest part of voice is knowing when the human stopped talking'
Practitioner post on endpointing and barge-in: the model is rarely the problem; the turn-detection layer is. Matches our bench, where barge-in failure predicted abandonment. Widely shared in the infra community.

Contact-centre platform vendor claims 70% containment with voice agents
Headline containment figure from US-English tier-1 deployments; 'contained' includes calls that ended after the agent read the account balance. No AU customer named. Carried as the claim our bench contradicts.

'Can the voice bot take a payment over the phone without a human on the line?'
Asked by a contact-centre GM after a vendor demo. The delivery lead did not have an answer and logged it. Became the transactional-intent arm of the bench and the reason the recommendation draws its line at triage.

Field-imagery benchmark: meter dials, forms and asset photos
Public multimodal benchmark of structured field imagery. Frontier models above 96% task accuracy on meter reads and handwritten forms; the energy sector's use case is solved at the model level and blocked on integration only.

Accent-conditioned error rates in realtime speech models: an Australian and non-native English study
Measures word error rate across three realtime stacks on Australian-accented, Indian-English and Mandarin-accented English against a US baseline. AU-accented error rate was 2.1× baseline; the gap narrowed but did not close with vendor accent options.

shelfsight — open vision-language model for shelf-gap and planogram detection
Small open VLM fine-tuned on retail shelf imagery. The retail sector owner ran it on store photos during a Woolworths pilot and it matched the frontier model on gap detection. Basis for the edge classifier work in the SLM field.

'Voice AI will handle 40% of Australian contact-centre volume by 2027'
Demand-band forecast that appears in every vendor deck since. Plausible for triage volume; no reliability caveat. Included as the mainstream framing.
Claims · 4 supporting, 1 refuting
The atoms.
A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.
From Sydney, p95 turn latency on US-hosted realtime models exceeds the 1.5 s threshold at which callers talk over the agent; a regional endpoint roughly halves it.
Vision on structured field imagery is production-ready now; the reliability gate is specific to voice.
Barge-in failure, not recognition accuracy, is the dominant cause of caller abandonment in AU voice-agent trials.
Voice agents hallucinate confirmations at a rate (5–8% of calls) that rules out unsupervised transactional use.
Vendor-reported 70% containment generalises to Australian queues once the accent model is tuned.
Position history · the diff is the product
4 validation runs against a fixed brief. Confidence 42% → 63%.
Latency bench concluded. US-hosted floor is unusable; regional endpoint halves it; barge-in is the abandonment driver. Recommendation published; IVR-replacement thesis to the graveyard. Field stays contested because the floor is moving.
- From Sydney, p95 turn latency on US-hosted realtime models exceeds the 1.5 s threshold at which callers talk over the agent; a regional endpoint roughly halves it.
- Barge-in failure, not recognition accuracy, is the dominant cause of caller abandonment in AU voice-agent trials.
- c-voice-and-vision-3 ↑ 0.62 → 0.70
Scoring · ordinal bands
Agents propose. A named human commits.
Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.
Impact
committed · MAContact-centre cost is the largest AI line item in every telco and insurance pipeline we have.
Timeline
committed · DDTriage is deployable now; transactional depends on grounding and regional endpoints, both inside the window.
TAM
agent-estimatedAgent-estimated from AU contact-centre labour spend addressable by tier-1 automation. Uncommitted.
Demand
committed · MAEvery telco engagement this half has asked; two have live vendor pilots we are being asked to assess.
Cost
committed · DDThe bench needs a telephony rig and scripted callers; two engineers, three weeks per cycle.
Cost of being wrong
committed · TBA hallucinated payment confirmation on a regulated queue is a complaint to the TIO or AFCA, not a bug.
Workforce readiness
agent-estimatedNobody in delivery has built a barge-in-safe voice loop; the vendor stacks hide it until it fails. Agent-estimated.
Relevance · per vertical
Why it matters here, or explicitly does not.
Ranking is per vertical, not global. Sector owners commit notes against agent drafts.
The largest queues in the country and the most vendor pressure. The AU latency and accent numbers are the assessment they are asking us for.
Mechanism · Tier-1 triage with a regional endpoint and human handoff on any transactional intent; measure barge-in and confirmation errors on the live queue.
Vision, not voice. Shelf-gap and planogram compliance from store photos is already in a Woolworths pilot.
Mechanism · Distilled vision model on in-store hardware; frontier model for exception review.
Meter-reading and asset-inspection photos are structured imagery where frontier accuracy is already sufficient.
Mechanism · Batch vision over field-app uploads; no latency constraint, so no gate. Agent draft.
Claims intake by voice is attractive and the confirmation-hallucination rate makes it dangerous; vision on claim photos is fine.
Mechanism · Vision now; voice only after the grounding pattern is measured. Agent draft.
Red team · the strongest case against
The strongest case against: we measured a floor that is already moving under us. The regional endpoint arrived mid-experiment, barge-in is a stack-engineering problem the vendors will fix within a release or two, and confirmation grounding is a prompt-and-tool pattern, not a research problem. A 'not yet' that expires in a quarter is a 'yes' with extra steps, and clients who wait on our advice will be a quarter behind the ones who did not.
- —The 2.6 s p95 was measured on US endpoints that no serious AU deployment would use after July. Our headline number is already historical.
- —Fifty scripted test calls is a small sample for a 4-in-50 hallucination rate; the confidence interval includes rates that are acceptable with a human check.
- —Contained tier-1 calls are most of the volume in most queues. 'Holds for triage' may be 80% of the business case, which makes 'not yet' the wrong headline.
- —The vision half is not contested by anyone and does not need a field; keeping it here inflates the field's confidence.
Source diversity
- Speech and ML research25%
- Voice infra practitioners20%
- Vendor and analyst25%
- Internal / Engel30%
A field supported by one epistemic community is a flag, not a finding.
Cross-pollination · typed joins
Connected, not merely similar.
Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.
On-device recognition and vision are where the latency floor and the residency question both disappear.
Regional or on-prem realtime inference is the only path under the 1.2 s p95 target from Sydney.
Audio tokens are priced an order of magnitude above text; the cost ledger needs a voice column.
Voice is the interface most assistants will eventually need; the reliability lessons transfer.
Share graph
Provenance running forward.
Discovery, not accountability. No counts, no rankings, no rollups to managers.
Convergence · who else is here
- MAMario Attard · Senior Analyst · Telco delivery1 drop
- DDDavid Diaz · Lead Analytics Specialist · agents1 drop
- DBDillon Blake · Senior Analyst · Retail1 drop
- OVOliver Vu · Analyst · product engineering1 drop
Several people’s drops meet here. An informal working group already exists and probably does not know it.
Lineage
What this field produced, and what it killed.
Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.
Not yet: realtime voice for AU contact centres above tier-1 triage
strength strong · 18 citations · review 29 Oct 2026
How fast is fast enough
At least one realtime voice endpoint reachable from Sydney holds median turn latency under 800ms across 40 scripted tier-1 telco triage calls over a partner SIP trunk.
Realtime voice agents replace tier-1 IVR
“Hung up before the answer.” · lived 3 months
Open questions · return to the pile
Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.
- 01What is the p95 on a live queue, not a rig, with the regional endpoint and barge-in at the edge?
- 02Does confirmation grounding survive a full call without adding a turn of latency?
- 03Should vision leave this field and go to practice now, given nobody contests it?
Tools in this space · 4
What you could actually buy.
Products aimed at this field, with what the lab has behind each one. Scored on six axes and never summed — “which is better” is not a question anyone has, and the constraint that decides it is named beside every assessment.
Realtime media infrastructure with an agent framework on top, self-hostable. The answer when the round trip has to stay in-country.
SaaS or self-hostApache-2.0 core, commercial cloudseen 2026-01-20No commercial relationship.
Speech recognition with the latency profile realtime work needs, and a self-host option that most competitors do not offer.
SaaS or self-hostCommercialseen 2025-12-15No commercial relationship.
Speech synthesis good enough that quality has stopped being the constraint. Which moves the argument to latency, cost and consent, where it belongs.
SaaS onlyCommercialseen 2025-10-27No commercial relationship.
Orchestration for realtime voice agents — turn-taking, interruption, telephony. The part everyone underestimates until they build it themselves.
SaaS onlyCommercialseen 2026-02-24No commercial relationship.
Listing is not recommending. Most of Voice and Vision sits at signal tier — in the space, nothing behind it — and a tool only reaches tested when a run stands behind it. Vendor pricing and capability move monthly, so these carry the shortest half-life in the library, and anyone who cited one gets told when it moves.
Notes · anyone in the firm
What people have written on this.
The person who knows a claim is wrong is usually not the person who wrote it. Corrections, objections and questions are owed an answer and stay open until the field owner says what they did; context and use notes stand as they are.
Notes · 0
Anything here reaches at most client-safe — a note cannot travel further than what it is written on.Nothing written on this yet. The useful notes are the ones from people who are not in the lab — that is where the correction usually comes from.