Sizing · Workload shape

Training and inference need different machines.

Buying one estate to do both is the most expensive common mistake in AI infrastructure. The requirements diverge on almost every axis that costs money.

Tell us what you need See live inventoryAcknowledged within one hour, first sourcing pass within one business day.
At a glance

What decides this requirement.

Training
Fabric-bound, single location, bursty, density-hungry
Inference
Memory-bound, geographically distributed, steady, latency-bound
Overlap
Less than most buyers assume
Cost of ignoring it
Usually overpaying substantially on the inference half
Where they diverge

Four axes, four different answers.

Interconnect. Training synchronises gradients constantly across nodes and lives or dies on fabric quality. Most inference fits within a single node and gains almost nothing from expensive east-west networking.

Geography. Training is a batch job indifferent to where it runs; put it where power is cheap. Inference serves users who are somewhere specific, so latency makes distribution a requirement rather than an option.

Utilisation. Training is bursty around runs; inference is a steady service. That difference alone argues for different commercial structures, not just different hardware.

Memory versus compute. Training is usually compute and fabric bound. Inference is usually memory bound, with the KV cache at target concurrency deciding the card.

A dense training cluster aisle beside a single compact edge cabinet
What it means commercially

Two requirements, priced separately.

The practical design most serious operations converge on: a reserved training cluster where power is cheapest, and geographically distributed inference capacity sized on memory and latency, connected privately.

That is more thought than buying one big estate, and it is reliably cheaper. The single-estate approach usually means paying training-grade fabric and density premiums on hardware doing inference work that never uses them.

We price the two separately and then look at the whole, which is the only way to see where the overspend would have been. Details on the training and inference pages.

Where to next

Start from what is verified.

Current verified lines with quantities, lead times and indicative pricing are public on the live inventory. Anything not listed becomes a sourcing requirement with a first pass inside one business day.

Tell us what you need See live inventoryAcknowledged within one hour, first sourcing pass within one business day.
Straight answers

Asked first, answered straight.

Can one cluster do both?

Physically yes, economically rarely well. You end up paying training-grade fabric and density premiums on hardware doing inference that never uses them, and accepting inference-grade geography for training that did not need it.

What is the biggest single difference?

Interconnect. Training depends on it completely; most inference barely uses it. That one difference accounts for a large share of the cost gap between the two estates.

Should inference use older GPUs?

Frequently yes. Previous-generation parts with adequate memory serve production inference well at a fraction of flagship cost, and the workload rarely notices.

How do we split the buying?

Reserve the training cluster where power is cheap, distribute inference where the users are, size each on its own constraint. We price them separately and then together.

Talk to the desk

Working through this on a real requirement?

Twenty minutes with the desk, no pitch and no quote at the end of it. Tell us roughly what you need and we will come back within one business day.

Acknowledged within one hour, first sourcing pass within one business day.

Sent. We are on it.

Your enquiry has landed with the desk. Acknowledged within one hour.