For teams serving models in production: the desk sizes and sources the estate around concurrency and memory, not benchmarks. Inference is a completely different problem from training, and buying it as though it were training is the most expensive habit in the market. It wants memory, geography and steady utilisation, not fabric.
Training is a batch job that runs flat out for weeks across many nodes in one place. Inference is a continuous service, latency-sensitive, usually fitting within a single node, and serving users who are somewhere specific. Almost every design decision follows differently from that.
Fabric quality, which dominates training, is largely irrelevant to single-node inference. Geography, which barely matters for training, becomes the primary constraint. And utilisation is steady rather than bursty, which changes the buying structure entirely.
The practical consequence: flagship silicon bought for an inference workload is the most common overspend we encounter. A great deal of production serving runs beautifully on previous-generation or right-sized parts at a fraction of the cost.

Model weights have to fit in GPU memory, and quantisation changes that arithmetic significantly. Once weights fit, the KV cache for concurrent requests consumes the remainder, which is why concurrency targets drive memory sizing as much as model size does.
Then geography. Latency to users is a product decision as much as a technical one, and it usually means several smaller deployments rather than one large one. Our regional pages cover where capacity actually exists.
Because utilisation is steady and predictable, inference is the strongest possible case for reserved capacity over on-demand pricing. The comparison is in cloud versus bare metal.
Current verified lines with quantities, lead times and indicative pricing are public on the live inventory. Anything not listed becomes a sourcing requirement with a first pass inside one business day.
Enough memory for the model weights plus the KV cache at your concurrency target. That is frequently far less card than teams assume, and previous-generation or right-sized parts often serve production perfectly.
Usually not. Training wants fabric and density in one location; inference wants memory and geographic distribution. Buying one estate for both typically means overpaying for the inference half.
Rarely. Most inference fits within a single node, so the expensive east-west fabric that training depends on does very little for it.
Almost always, because utilisation is steady. On-demand pricing exists to buy flexibility, and a service running continuously is not using that flexibility while paying for it.
Twenty minutes with the desk, no pitch and no quote at the end of it. Tell us roughly what you need and we will come back within one business day.
Your enquiry has landed with the desk. Acknowledged within one hour.