cavendish
AssessedValidatinggate · EconomicsNear · 1–3 years

On-Prem Inference

Owned inference hardware pays back only above a sustained utilisation floor that most enterprises never reach; below it, sovereignty and residency are better bought as a hosted AU-region service than as a GPU cluster, and the crossover point moves against on-prem every time a hosted price drops.

A validation run. Researched position, no experiment.

Join with…

Confidence

74%human-committed

Expiry

42duntil review · 15 Oct 2026

Lead time

not yet mainstream · opened 4 Nov 2025

Ownership

DYDylan Desmarcheliermonthly cadence

Where it is

unattributed

Validating, because the position has been tested against real numbers three times and held each time. The pressure came from two clients — a bank and a state government agency — who wanted to know whether to buy a GPU cluster. Our model puts the cost crossover for a 70B-class open-weight model at roughly 55–65% sustained utilisation on H100-class hardware at AU power and colocation prices, and almost no enterprise workload we have measured runs above 40%. Below the crossover, an AU-region hosted service with contractual data residency is cheaper and carries the same sovereignty story for every regulator we have checked. The exceptions are real: air-gapped defence and some health workloads, and any client with a GPU allocation they already own. Our own 18-month payback claim for an H100 cluster is in the graveyard. GPU supply eased in 2026 and hosted prices fell about 40% in twelve months, which moved the crossover further from most clients, not closer.

Written by hand and carrying nobody's name. Editing it puts yours on it.

Why a Quantium decision hinges on it

unattributed

This is the most expensive single decision a client can make in AI infrastructure and the one where vendor pressure is strongest. A bank that buys a $12M cluster at 30% utilisation has bought sovereignty at three times the hosted price and will run it for five years to justify it. Quantium gets asked directly and answers with a standing answer that has to be right. It also sets the floor for TEE and Secure Inference: if hosted AU-region inference is the default, confidential computing is what makes it acceptable for regulated data, and the two fields move together.

Written by hand and carrying nobody's name. Editing it puts yours on it.

What it actually is

composed from the records

Owned inference hardware pays back only above a sustained utilisation floor that most enterprises never reach; below it, sovereignty and residency are better bought as a hosted AU-region service than as a GPU cluster, and the crossover point moves against on-prem every time a hosted price drops. That is the lab's one-line position on it, which is not the same as an explanation.

The shape the field is converging on, from the most authoritative source in it: AU regulators accept contractual residency for the data classes clients run inference on.signal

This is the section a page most needs a person for, and the one composition is worst at. Nobody has written the plain-language version — what the idea is, in words that assume nothing — and it is the first thing a reader who has never met the term needs.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

Why now

composed from the records

The lab opened this field on 2025-11-04, and it has not reached mainstream awareness yet. Everything below is what has moved since.

What shipped: Two AU-region hosted inference providers cut 70B-class prices for the third time in a year (Provider pricing pages, 2026-08-04) and Two frontier labs announce AU-region inference with contractual residency (Lab announcements, 2026-06-24). Tooling arriving is what moves a field from argument to something a team could try.signalsignal

And the rule moved: OAIC and DTA guidance: contractual AU-region residency satisfies APP 8 and hosting policy for PROTECTED and below In a regulated vertical that usually decides the timing more than the technology does.signal

Demand is rising on it rather than steady — “Can this run entirely on-shore, and can you prove it?” — which is the difference between a field worth watching and one worth doing something about.demand

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

What it changes in a system

composed from the records

What changes, concretely: A 70B open-weight model on an 8×H100 node at AU colocation and power prices costs more per million tokens than the cheapest AU-region hosted equivalent until sustained utilisation passes about 60% (validation run 4, re-baselined).signal

There is already a shipped default — inference.host = hosted-au-region · exceptions: owned-allocation | air-gapped | util>60% · standing answer: sa-onprem-when — so a team adopting this is changing a setting rather than starting a project.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

What is in the way

composed from the records

The binding constraint is economics: it works and does not yet pay. Everything upstream of that is solved and everything downstream of it is waiting.

Workforce readiness is high: The position is a spreadsheet and a set of exceptions; any sector owner can deliver it. Agent-estimated. A recommendation needing skills the firm does not hold is an aspiration rather than an action, and it routes to the enablement agenda instead of the delivery one.

Something here has already been killed: On-prem H100 cluster pays back inside 18 months, on two au-region hosted price cuts in q1 and q2 moved the breakeven to over four years before we finished the model; the standing answer now carries the calculation and its refresh date.graveyard

The argued case against it is the red team's, further down this page, and it is deliberately one-sided — this section is what stands in the way mechanically, not what somebody thinks of it.

Correct and traceable, and nobody's judgement yet. The first person to write it gets the byline.

2 of 6 explanatory sections are written; the rest are composed until somebody takes them.

Business priority

On the plan
  • This is the watch item's central question, and the answer is currently 'do not buy the cluster'.

    HB Harley Barnescommitted
  • Only becomes a cost lever above a volume threshold almost no engagement reaches.

    Agent draft — no owner has committed this alignmentagent-estimated

Clients are asking

  • Can this run entirely on-shore, and can you prove it?5 engagements · $1M–5M · rising

Priority orders what you see. It never changes what the evidence says — a plan-critical field with nothing tested is still signal tier.

Field attributes

StateValidating
GateEconomics · possible, not yet affordable
OriginPressure
Measurablefull
Written forexec
Reach · TLPClient-safe · TLP:GREEN
Horizonnear
Opened4 Nov 2025
Mainstreamnot yet
Last validated20 Aug 2026
Sightings1

What people have written

Write one

Nothing yet. The person who knows a claim is wrong is usually not the person who wrote it.

A note never travels further than the thing it is written on.

Position

What is demonstrated, what is hype, what would have to be true.

The shape every position request answers. Signal-tier fields carry a draft; assessed and tested fields carry a validated one.

What is demonstrated
  • 01A 70B open-weight model on an 8×H100 node at AU colocation and power prices costs more per million tokens than the cheapest AU-region hosted equivalent until sustained utilisation passes about 60% (validation run 4, re-baselined).
  • 02Measured utilisation on three client inference clusters ranged from 18% to 41% over a quarter; none approached the crossover.
  • 03Every regulator we checked — APRA, the OAIC, the Digital Transformation Agency — accepts contractual AU-region residency for the data classes our clients run; none requires owned hardware outside defence.
What is hype
  • 01'Sovereign AI' as a reason to buy hardware. Sovereignty is a residency and control question that a contract can answer for most workloads.
  • 02GPU scarcity as a reason to buy now. Supply eased in 2026 and secondary prices fell; the scarcity argument aged badly.
  • 03Payback models that assume 80% utilisation. Nobody we have measured runs there, and the vendor models never show the number.
What would have to be true
  • 01For on-prem to be the default, hosted AU-region prices would have to stop falling, or a regulator would have to require owned hardware for a common data class. Neither is happening.
  • 02For the crossover to move toward on-prem, an inference-efficiency step on owned hardware (e.g. better batching or a cheaper accelerator) would have to outpace the hosted price curve.
  • 03For our position to be wrong, client utilisation would have to be far higher than we measured; a fourth measurement on a busier client is the test.
What we would do
  • 01Keep sa-onprem-when current and re-run the cost model quarterly; the answer changes with hosted prices, not with client sentiment.
  • 02Name the exceptions explicitly in every engagement: air-gapped, already-owned allocation, and utilisation demonstrably above 60%.
  • 03Pair with TEE and Secure Inference so that the hosted default has a regulated-data story; without it the on-prem argument comes back through the risk function.

Signals · 10 in this cluster

What the cluster is made of.

Every item carries its source, tier and sightings. Detector-found signal sits beside human drops; downstream they are indistinguishable except by provenance.

band 1 · bleeding edgeband 2 · early adoptionband 3 · demand
55–65% util
crossover
Finding·band 1Assessed

Cost crossover model v3: owned H100 vs AU-region hosted, 70B open-weight, AU power and colo

Rebuilt cost model using quoted Sydney colocation, NEM power prices and current hosted $/Mtok. Crossover at 55–65% sustained utilisation depending on batch efficiency. Spreadsheet and assumptions published internally; the number moves with hosted prices.

extracted claimOwned H100-class inference beats AU-region hosted only above roughly 60% sustained utilisation.
Lab · validation run 4 · Dylan Desmarchelier20 Aug 2026
detector · bleeding edge
$0.55 / Mtok
70B hosted
Release·band 1Signal

Two AU-region hosted inference providers cut 70B-class prices for the third time in a year

Cumulative drop of about 40% on 70B-class open-weight inference in AU regions since August 2025. Each cut moves the crossover further from typical client utilisation.

extracted claimAU-region hosted inference prices are falling faster than owned-hardware costs.
Provider pricing pages4 Aug 2026
DYdropped 2
41%
max utilisation
Finding·band 1Tried

Three client clusters measured: 18%, 31%, 41% sustained utilisation over a quarter

Utilisation telemetry from three client inference clusters, with permission. None approached the crossover. The 41% cluster is the one with an existing allocation, which is the exception the position names.

extracted claimEnterprise inference clusters run well below the cost crossover.
Lab · Nightingale measurement · Andrew Tran29 Jul 2026
ATdropped
3
openings
Job posting·band 2Signal

AU hyperscaler region hiring 'Inference Capacity Planner' ×3

Three openings for AU-region inference capacity. Argus inference: hosted AU capacity is expanding, which pushes prices down further. Carried as inference.

Careers page21 Jul 2026
detector · early adoption
Announcement·band 1Signal

Two frontier labs announce AU-region inference with contractual residency

Frontier-model inference in Sydney regions with data-residency terms. Removes the last practical reason for owned hardware for clients who wanted frontier capability with residency.

Lab announcements24 Jun 2026
detector · bleeding edge 2
Regulatory·band 3Signal

OAIC and DTA guidance: contractual AU-region residency satisfies APP 8 and hosting policy for PROTECTED and below

Consolidated guidance we checked against for the regulator survey. No requirement for owned hardware outside classified workloads. A change here is the thing that would flip the position.

OAIC / DTA13 May 2026
AVdropped
−35% YoY
H100 secondary
Post·band 2Signal

'The GPU shortage is over and nobody told procurement'

Charts secondary-market H100 prices and cloud spot availability through early 2026. Supply eased; the scarcity argument for buying is retired here and in our position.

Substack · A well-followed infra voice21 Apr 2026
AWdropped 2
Drop·band 3Signal

Slack drop: 'the uni has an idle DGX and the health department can borrow it'

A sector owner's note about an existing allocation. Became the 'already-owned' exception in the position: marginal cost is power, so use it.

Slack drop26 Feb 2026
SLdropped
27%
median utilisation
Analyst·band 3Signal

Survey: median enterprise GPU utilisation 27%

Survey of 220 enterprises with owned accelerators. Median sustained utilisation 27%; top decile 58%. Consistent with our three measurements and with the crossover being rarely reached.

Forrester10 Feb 2026
detector · demand
Client question·band 3Signal

'The board wants to know if we should buy GPUs before they run out'

The question that opened the field, from a bank's CIO. Supply and sovereignty were the stated reasons; utilisation was not mentioned. Answered with the standing answer three months later.

Engel · banking engagement19 Nov 2025
AHHBdropped 3
Seen something that belongs here?Under fifteen seconds, or it will not be used.

Claims · 5 supporting, 1 refuting

The atoms.

A document cannot go stale; an assertion can. Claims are immutable and stamped with the extractor that produced them, so staleness, diffs and the graveyard operate at claim level.

Hosted AU-region inference prices fell about 40% in twelve months, moving the crossover away from on-prem.

Assessedc-on-prem-inference-5dalton-0.420 Aug 2026Provider pricing pages, Lab announcements
82%

The cost crossover between owned H100-class hardware and AU-region hosted inference for a 70B open-weight model sits at roughly 55–65% sustained utilisation at 2026 prices.

Assessedc-on-prem-inference-1dalton-0.420 Aug 2026Lab · validation run 4, Provider pricing pages
80%

AU regulators accept contractual AU-region residency for the data classes our clients run; owned hardware is required only for defence-classified workloads.

Assessedc-on-prem-inference-3dalton-0.49 Jun 2026OAIC / DTA, Engel · banking engagement
78%

Enterprise inference clusters run at 20–40% sustained utilisation; the crossover is not reached in practice.

Assessedc-on-prem-inference-2dalton-0.420 Aug 2026Lab · Nightingale measurement, Forrester
76%

Clients with an existing GPU allocation should run inference on it regardless of the crossover; the marginal cost is power alone.

Assessedc-on-prem-inference-6dalton-0.312 Mar 2026Lab · Nightingale measurement, Slack drop
70%

GPU supply constraints make owning hardware necessary to guarantee capacity.

Assessedc-on-prem-inference-4dalton-0.49 Jun 2026Substack, Provider pricing pages
20%

Position history · the diff is the product

4 validation runs against a fixed brief. Confidence 50% → 74%.

runs compare claim sets, never prose
What we said · run 4

Hosted AU-region prices down about 40% year on year; crossover moved further from typical utilisation. Position held on a third client measurement. Standing answer refreshed.

74%
Changed since run 3
  • The cost crossover between owned H100-class hardware and AU-region hosted inference for a 70B open-weight model sits at roughly 55–65% sustained utilisation at 2026 prices.
  • Enterprise inference clusters run at 20–40% sustained utilisation; the crossover is not reached in practice.
  • Hosted AU-region inference prices fell about 40% in twelve months, moving the crossover away from on-prem.
  • c-on-prem-inference-4 ↓ 0.28 → 0.20
Positions are superseded, never edited. The prediction record is worthless if it can be quietly revised.Crystal ball

Scoring · ordinal bands

Agents propose. A named human commits.

Uncommitted scores are visibly marked and never leave the building. Bands, not point estimates — false precision is the tell that a number was generated rather than derived.

Impact

committed · AW
high

Largest single infrastructure decision a client makes in AI; the firm's answer is on the record.

Timeline

committed · DY
0–18mo

Decisions are being made now; the field is about getting them right, not about a future capability.

TAM

agent-estimated
$1B–10B

Agent-estimated from AU enterprise AI infrastructure spend. Uncommitted.

Cost of being wrong

committed · TB
high

A wrong 'buy' is a five-year eight-figure mistake; a wrong 'don't' is a residency incident. Both are bad.

Demand

committed · AV
high

Government asks every quarter; banking asked twice this half. The standing answer is the second most cited in the library.

Cost

committed · DY
low

A cost model and quarterly re-run. No hardware bought.

Workforce readiness

agent-estimated
high

The position is a spreadsheet and a set of exceptions; any sector owner can deliver it. Agent-estimated.

Relevance · per vertical

Why it matters here, or explicitly does not.

Ranking is per vertical, not global. Sector owners commit notes against agent drafts.

Banking
relevant

A major bank asked whether to buy a cluster; the utilisation measurement on their pilot was 22%. The answer was no, with the exceptions named.

Mechanism · Hosted AU-region inference with contractual residency and, for PII classes, a TEE-backed tier.

AH committed by Amber Hallcommitted · AH
Government
relevant

State agencies are under sovereignty pressure and vendor pressure at once; the hosted-with-residency answer holds for PROTECTED and below.

Mechanism · IRAP-assessed AU-region hosted inference; owned hardware only for classified workloads.

AV committed by Aadhithyanarayanan V Acommitted · AV
Defenceprospective
watch

Air-gapped requirements make on-prem the only option, but it is a prospective vertical and the economics are not the deciding factor there.

Mechanism · Sponsor holds the question of whether to enter; the field's position does not apply.

Agent draft · awaiting a sector owneragent-estimated
Health
relevant

Health data residency is the strictest of the served verticals and the one where risk functions most often default to 'own it'; the TEE path is what makes hosted acceptable.

Mechanism · Hosted AU-region with attestation; owned hardware for the small set of workloads a state health department classifies above PROTECTED.

SL committed by Sylvia Liucommitted · SL

Red team · the strongest case against

The strongest case against: the cost model prices today's hosted market, which is subsidised by hyperscaler capital and will not stay at these prices once the market consolidates. A client who buys hardware is buying price certainty for five years, and the model does not value that. Utilisation is also a choice: a client who owns a cluster fills it, and the low utilisation we measured is on clients who did not commit. And the regulatory position is a snapshot; one adverse OAIC determination on offshore-controlled AU-region infrastructure would change the answer overnight.

  • Hosted prices are set by companies losing money on inference. The crossover assumes the discount persists.
  • Utilisation is endogenous. Measuring it on clients who chose hosted tells us little about clients who would commit to owned hardware and consolidate workloads onto it.
  • Residency by contract is control by contract, and the Privacy Act reforms in progress may not treat them as equivalent for all data classes.
Run by an agent briefed to argue the field is nothing — sources here are correlated, and without a deliberate adversary synthesis converges on consensus and calls it insight. Kept as a dated pass rather than overwritten. Nobody has answered it yet, and a challenge nobody answers is a disclaimer.thesis holds

Source diversity

  • Infra / market30%
  • Vendor / provider20%
  • Regulator15%
  • Analyst10%
  • Internal / Engel25%

A field supported by one epistemic community is a flag, not a finding.

Cross-pollination · typed joins

Connected, not merely similar.

Enabling, compounding, substituting, blocking. A satisfied dependency trigger is a far stronger signal than semantic proximity.

Share graph

Provenance running forward.

Discovery, not accountability. No counts, no rankings, no rollups to managers.

Lineage

What this field produced, and what it killed.

Experiments, recommendations and graveyard entries stay attached. The reasoning that killed a claim is the reusable asset.

Open questions · return to the pile

Every run leaves a record. Separately, its question either closes or returns to the pile with notes — which is what the next person proposing the same thing will see.

  1. 01What happens to the crossover if hosted prices stop falling — and what is the early signal that they have?
  2. 02Does a client who commits to owned hardware actually reach 60% by consolidating, or does utilisation stay low because workloads are bursty?
  3. 03Will the Privacy Act reforms treat contractual residency as equivalent to control for health and financial data?

Notes · anyone in the firm

What people have written on this.

The person who knows a claim is wrong is usually not the person who wrote it. Corrections, objections and questions are owed an answer and stay open until the field owner says what they did; context and use notes stand as they are.

Notes · 0

Anything here reaches at most client-safe — a note cannot travel further than what it is written on.

    Nothing written on this yet. The useful notes are the ones from people who are not in the lab — that is where the correction usually comes from.