Buying one estate to do both is the most expensive common mistake in AI infrastructure. The requirements diverge on almost every axis that costs money.
Interconnect. Training synchronises gradients constantly across nodes and lives or dies on fabric quality. Most inference fits within a single node and gains almost nothing from expensive east-west networking.
Geography. Training is a batch job indifferent to where it runs; put it where power is cheap. Inference serves users who are somewhere specific, so latency makes distribution a requirement rather than an option.
Utilisation. Training is bursty around runs; inference is a steady service. That difference alone argues for different commercial structures, not just different hardware.
Memory versus compute. Training is usually compute and fabric bound. Inference is usually memory bound, with the KV cache at target concurrency deciding the card.

The practical design most serious operations converge on: a reserved training cluster where power is cheapest, and geographically distributed inference capacity sized on memory and latency, connected privately.
That is more thought than buying one big estate, and it is reliably cheaper. The single-estate approach usually means paying training-grade fabric and density premiums on hardware doing inference work that never uses them.
We price the two separately and then look at the whole, which is the only way to see where the overspend would have been. Details on the training and inference pages.
Current verified lines with quantities, lead times and indicative pricing are public on the live inventory. Anything not listed becomes a sourcing requirement with a first pass inside one business day.
Physically yes, economically rarely well. You end up paying training-grade fabric and density premiums on hardware doing inference that never uses them, and accepting inference-grade geography for training that did not need it.
Interconnect. Training depends on it completely; most inference barely uses it. That one difference accounts for a large share of the cost gap between the two estates.
Frequently yes. Previous-generation parts with adequate memory serve production inference well at a fraction of flagship cost, and the workload rarely notices.
Reserve the training cluster where power is cheap, distribute inference where the users are, size each on its own constraint. We price them separately and then together.
Twenty minutes with the desk, no pitch and no quote at the end of it. Tell us roughly what you need and we will come back within one business day.
Your enquiry has landed with the desk. Acknowledged within one hour.