AI Gateway
over all model gardens
One gateway the client controls — routing, observability and cost attribution over Bedrock, Vertex, Azure, direct APIs and on-prem — is table stakes; standardising on a vendor's gateway trades a small build for a large lock-in and loses the on-prem and open-weight legs.
Experiment run, measured result. The only tier that becomes a recommendation.
Confidence
81%human-committedExpiry
63duntil review · 5 Nov 2026Lead time
−4moopened after mainstream — recorded honestlyOwnership
DYDylan Desmarcheliermonthly cadenceWhere it is
unattributedThe field converged in Q2. Every delivery pattern now routes through a gateway, the open-source gateways reached parity with the commercial ones on provider coverage, and our routing experiment found that static task-class rules capture nearly all the cost saving that learned routers promise. What clients actually buy the gateway for is cost attribution per use case, not routing. The remaining work is operational: adapter drift on streaming and tool-call formats recurs every provider release, and that is a tooling problem, not a research one. We expect to dissolve this field into practice by Q4.
Why a Quantium decision hinges on it
unattributedTelco and banking clients ask for one view of spend across two or three model gardens before they ask anything about models. Without a gateway they control, a client cannot route to an open-weight or on-prem model when sovereignty or cost demands it, and cannot answer 'what does use case X cost' at all. A hyperscaler gateway answers that question for the hyperscaler's garden only. The choice is made once per client and is expensive to reverse.
What it actually is
composed from the recordsOne gateway the client controls — routing, observability and cost attribution over Bedrock, Vertex, Azure, direct APIs and on-prem — is table stakes; standardising on a vendor's gateway trades a small build for a large lock-in and loses the on-prem and open-weight legs. That is the lab's one-line position on it, which is not the same as an explanation.
The shape the field is converging on, from the most authoritative source in it: Learned routers underperform static task-class rules outside their training distribution.signal
This is the section a page most needs a person for, and the one composition is worst at. Nobody has written the plain-language version — what the idea is, in words that assume nothing — and it is the first thing a reader who has never met the term needs.
Why now
composed from the recordsThe lab opened this field on 2025-10-06 and it reached mainstream awareness on 2025-06-10. The gap between those two dates is the lead time the lab is measured on.
What shipped: Hyperscaler gateway adds 'cross-provider routing' — to models inside its own garden only (Cloud provider changelog, 2026-04-15). Tooling arriving is what moves a field from argument to something a team could try.signal
What it changes in a system
composed from the recordsWhat changes, concretely: Static task-class routing across Bedrock, Vertex and a direct API cut cost 34% at equal eval score on four of six task classes (x-gateway-routing). A learned router added 2 points on top and lost 9 outside its training distribution.signal
Measured rather than argued: Validated. Routed configuration cut blended cost per task by 34% with no family dropping more than 1 point of pass rate. The vendor router comparison failed our extraction eval on 38% of calls. Published as r-gateway-default; g-single-gateway-vendor superseded.experiment
There is already a shipped default — gateway.routing.default = task-class-static · gateway.tags = use-case,cost-centre · skill: cavendish/gateway-routing — so a team adopting this is changing a setting rather than starting a project.
What is in the way
composed from the recordsThe binding constraint is tooling: it works and is affordable, and nobody can operate it yet. Everything upstream of that is solved and everything downstream of it is waiting.
Workforce readiness is medium: Delivery teams can stand it up; adapter drift still lands on the lab. Agent-estimated. A recommendation needing skills the firm does not hold is an aspiration rather than an action, and it routes to the enablement agenda instead of the delivery one.
Something here has already been killed: One commercial gateway over every model garden, on the routing experiment showed the value sits in the routing policy and the cost ledger, both of which we could not get out of the vendor gateway, so the recommendation became a gateway we control.graveyard
The argued case against it is the red team's, further down this page, and it is deliberately one-sided — this section is what stands in the way mechanically, not what somebody thinks of it.
2 of 6 explanatory sections are written; the rest are composed until somebody takes them.
Business priority
Plan criticalClients are asking
- “Which model should we standardise on?”11 engagements · not priced · steady
Priority orders what you see. It never changes what the evidence says — a plan-critical field with nothing tested is still signal tier.
Field attributes
What people have written
Write oneNothing yet. The person who knows a claim is wrong is usually not the person who wrote it.
A note never travels further than the thing it is written on.
Position
What is demonstrated, what is hype, what would have to be true.
The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.
- 01Static task-class routing across Bedrock, Vertex and a direct API cut cost 34% at equal eval score on four of six task classes (x-gateway-routing). A learned router added 2 points on top and lost 9 outside its training distribution.
- 02Per-use-case cost attribution through the gateway answered the telco client's spend question in one dashboard; the hyperscaler console could not.
- 03One open-source gateway covered all five legs we needed, including on-prem vLLM. The commercial gateway we trialled covered three.
- 01Learned routers. Outside their training distribution they underperform a lookup table, and the table is auditable.
- 02'Semantic caching' as a headline feature. Hit rates on our workloads were under 6%; prompt caching at the provider does the real work.
- 03Analyst claims of a discrete 'AI gateway market'. It is a feature of a platform, and most of it will be free.
- 01Provider tool-call and streaming formats stabilising enough that adapter drift stops being a monthly incident.
- 02Hyperscaler gateways routing outside their own garden, which none of them does today and none has announced.
- 03Cost attribution surviving agentic loops, where one task fans out across models and the per-call attribution stops meaning anything.
- 01Keep r-gateway-default as the delivery default and hand the field to practice in Q4.
- 02Stop re-running routing benchmarks; the answer is stable. Track adapter drift as an instrument-health metric instead.
- 03Fold the on-prem and open-weight legs into the standard gateway build so those fields do not need their own plumbing.
Signals · 10 in this cluster
What the cluster is made of.
Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

Routing experiment: static task-class rules cut cost 34% at equal eval score; learned router adds 2 points and loses 9 out of distribution
Six task classes routed across Bedrock, Vertex and a direct API through an open-source gateway. Static rules keyed on task class captured a 34% saving on four classes; learned routing added marginally in-distribution and underperformed the rules outside it. Agentic loops gained nothing.
extracted claimStatic task-class routing captures the cost saving; learned routers do not generalise.

Logged from Claude Code: gateway upgrade broke tool-call streaming on one provider for four days
Product engineer logged from a session: a provider changed its streamed tool-call delta format; the gateway adapter silently dropped arguments. Found by a failing eval, not by monitoring. Tried tier, one harness.

RouteBench: Learned LLM Routers Under Distribution Shift
Evaluates eight learned routers against static rules across shifted task mixes. Learned routers beat rules in-distribution by 1–3 points and lose by 6–12 under shift. Matches our experiment almost exactly.
extracted claimLearned routers underperform static task-class rules outside their training distribution.

Three AU banks post 'LLM Gateway Engineer' roles in the same fortnight
All three job descriptions lead with cost attribution and showback; routing appears in one. Argus inference: banks are building their own gateways and treating them as finance tooling.

'Your gateway is your lock-in'
Argues the gateway is the most durable lock-in in the stack because every prompt, log and cost centre tag lives in its schema. Agrees with our position on owning it; disagrees on whether open source is cheaper. Kept for the second half.

Hyperscaler gateway adds 'cross-provider routing' — to models inside its own garden only
Marketed as multi-model routing. Reads the fine print: every route terminates inside the provider's own catalogue. No direct-API leg, no on-prem leg. Confirms the pattern from the commercial trial.

Analyst note: '70% of enterprises will route inference through a gateway by 2027'
Demand-band signal with a discrete-market framing we do not share. Useful as evidence the pattern is mainstream; not useful on build versus buy, where its three-year TCO figure assumes zero adapter maintenance for the commercial option.

Commercial gateway trial ended: no on-prem leg, opaque cost attribution
Six-week trial of a commercial gateway on the telco pattern. Three of five legs covered; cost attribution exported as a monthly CSV with no use-case dimension. Became g-single-gateway-vendor.

Open-source gateway passes 30k stars; adds Bedrock, Vertex, Azure and vLLM parity in one release
The release that closed the provider-coverage gap with the commercial gateways. Streaming and tool-call adapters are where the issue tracker lives; a third of open issues are format drift after a provider release.
extracted claimOpen-source gateways reached provider-coverage parity with commercial ones in early 2026.

'Can we see cost per use case across Bedrock and Vertex in one place?'
Asked by a telco CFO's office, not the engineering team. The question that reframed the field from routing to attribution. Logged unanswered; the gateway build that followed answered it in one dashboard.
Claims · 5 supporting, 1 refuting
The atoms.
A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.
Static task-class routing rules capture most of the achievable cost saving (30–35%); learned routers add little and lose outside their training distribution.
Vendor gateways cover their own garden well and every other garden badly; the on-prem and open-weight legs are where they fail.
Cost attribution per use case is the feature clients buy the gateway for; routing is secondary.
Routing gains vanish on agentic loops because the loop needs one model's tool-call semantics end to end.
Adapter drift on streaming and tool-call formats is the recurring operational cost of a gateway; it re-appears with every provider release.
A commercial gateway is cheaper over three years than an owned open-source one once staffing is included.
Position history · the diff is the product
4 validation runs against a fixed brief. Confidence 50% → 81%.
Routing experiment concluded: static rules capture the saving; learned routers do not generalise; agentic loops do not benefit. Field converged. Recommendation published; hand to practice in Q4.
- Static task-class routing rules capture most of the achievable cost saving (30–35%); learned routers add little and lose outside their training distribution.
- Cost attribution per use case is the feature clients buy the gateway for; routing is secondary.
- Adapter drift on streaming and tool-call formats is the recurring operational cost of a gateway; it re-appears with every provider release.
- Routing gains vanish on agentic loops because the loop needs one model's tool-call semantics end to end.
- c-ai-gateway-3 ↑ 0.7 → 0.77
Scoring · ordinal bands
Agents propose. A named human commits.
Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.
Impact
committed · DYEvery inference call in every pattern passes through it. Made once per client.
Timeline
committed · AWAlready mainstream. We opened this field late.
TAM
agent-estimatedAgent-estimated from gateway vendor revenue and platform spend; most of it will be bundled. Uncommitted.
Cost
committed · DYOne engineer, one week for the standard build; drift maintenance thereafter.
Demand
committed · MAAsked in every telco engagement this year, always as a cost question.
Cost of being wrong
committed · TBWrong gateway is a migration, not an incident.
Workforce readiness
agent-estimatedDelivery teams can stand it up; adapter drift still lands on the lab. Agent-estimated.
Relevance · per vertical
Why it matters here, or explicitly does not.
Ranking is per vertical, not global. Sector owners commit notes against agent drafts.
Two model gardens under separate contracts and a finance team that wants spend per use case. The gateway is the only place that number exists.
Mechanism · Gateway tags every call with use case and cost centre; showback runs off the gateway log.
Banks need an on-prem or sovereign leg for some workloads and a frontier leg for others; the gateway is where the policy lives.
Mechanism · Routing rules by data classification; sovereign leg for PII-bearing calls.
Health clients hit residency constraints first and cost second; both are gateway questions.
Mechanism · Residency-aware routing with the on-prem leg as default for clinical text.
Agencies procure gardens through panels one at a time; multi-garden routing is not yet a question they can ask.
Mechanism · Would apply once a second garden lands on a whole-of-government arrangement.
Red team · the strongest case against
The strongest case against: the hyperscalers will make the gateway free and native, and 'a gateway you control' is a maintenance burden a twelve-person lab is telling delivery teams to carry forever. Most clients have one cloud contract, and multi-garden routing solves a problem they do not have.
- —Adapter drift is a permanent tax. We measured it at roughly one incident per provider release; over five providers that is a part-time engineer per client, which the cost score does not include.
- —Cost attribution is a logging feature. Any provider console could ship it next quarter and remove the main reason to own the gateway.
- —Routing gains of 34% were measured on classification and extraction. On agentic loops — where spend is growing fastest — the gain was zero.
- —A converged field with a low-priority score is a field that should already be in practice. Keeping it open flatters the lab's coverage.
Source diversity
- Open-source infra35%
- ML research15%
- Vendor20%
- Analyst10%
- Internal / Engel20%
A field supported by one epistemic community is a flag, not a finding.
Cross-pollination · typed joins
Connected, not merely similar.
Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.
Open weights become a routable leg only through the gateway; task-class parity is what makes routing to them worth doing.
Routing and caching both live in the gateway; the ledger that measures them is the gateway log.
The delegation broker needs one choke point for tool calls and the gateway already is one for model calls.
An on-prem leg is only usable in delivery when the gateway can route to it by data classification.
Share graph
Provenance running forward.
Discovery, not accountability. No counts, no rankings, no rollups to managers.
Convergence · who else is here
- OVOliver Vu · Analyst · product engineering2 drops
- DYDylan Desmarchelier · Senior Platform Analyst · inference2 drops
- MAMario Attard · Senior Analyst · Telco delivery1 drop
- MLMichelle Lam · Analytics Lead · evals1 drop
Several people’s drops meet here. An informal working group already exists and probably does not know it.
Lineage
What this field produced, and what it killed.
Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.
Route through a gateway you control; do not standardise on a vendor's garden
strength strong · 33 citations · review 11 Jan 2027
Right model, right task
A routing policy we own, keyed on our per-task eval pass rate, cuts blended cost per task by at least 25% against a fixed frontier default with no drop in eval pass rate.
One commercial gateway over every model garden
“Standardised on someone else's roadmap.” · lived 7 months
Open questions · return to the pile
Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.
- 01Does per-call cost attribution mean anything on an agentic loop that fans out across three models?
- 02What is the true adapter-drift cost per client per year, measured rather than estimated?
- 03When a hyperscaler first routes outside its own garden, does the own-the-gateway recommendation survive?
Tools in this space · 5
What you could actually buy.
Products aimed at this field, with what the lab has behind each one. Scored on six axes and never summed — “which is better” is not a question anyone has, and the constraint that decides it is named beside every assessment.
Win condition: To be declared at preregistration. The proposal is that a tool wins on cost of exit, provider-drift resilience measured against the last two breaking changes, and whether a delivery team can operate it without a platform engineer.
A thin translation layer over every provider API, plus a proxy with keys, budgets and logging. Boring in the way infrastructure should be.
Open sourceMIT, commercial enterprise tierseen 2025-09-15Runs in the lab's own gateway deployment.
A hosted marketplace and single endpoint across many models. Excellent for evaluation and a hard sell for regulated production — the data path is somebody else's.
SaaS onlyCommercialseen 2025-09-22Used by the lab for evaluation runs, on lab accounts only.
Caching, rate limiting and analytics in front of provider APIs, at edge. Cheapest possible answer if the estate is already on Cloudflare and an odd one if it is not.
SaaS onlyCommercialseen 2025-10-27Cloudflare is a Quantium technology partner.
AI routing as plugins on an existing API gateway. The right answer for an estate that already governs every other API through Kong and a heavy one otherwise.
SaaS or self-hostOpen-source core, commercialseen 2026-01-13No commercial relationship.
Gateway with routing, caching, guardrails and observability as one product. More opinionated than LiteLLM and correspondingly harder to leave.
SaaS or self-hostOpen-source core, commercial cloudseen 2025-11-08No commercial relationship.
Listing is not recommending. Most of AI Gateway sits at signal tier — in the space, nothing behind it — and a tool only reaches tested when a run stands behind it. Vendor pricing and capability move monthly, so these carry the shortest half-life in the library, and anyone who cited one gets told when it moves.
Notes · anyone in the firm
What people have written on this.
The person who knows a claim is wrong is usually not the person who wrote it. Corrections, objections and questions are owed an answer and stay open until the field owner says what they did; context and use notes stand as they are.
Notes · 0
Anything here reaches at most client-safe — a note cannot travel further than what it is written on.Nothing written on this yet. The useful notes are the ones from people who are not in the lab — that is where the correction usually comes from.