Use case · Inference and serving

GPU infrastructure for inference and serving.

For teams serving models in production: the desk sizes and sources the estate around concurrency and memory, not benchmarks. Inference is a completely different problem from training, and buying it as though it were training is the most expensive habit in the market. It wants memory, geography and steady utilisation, not fabric.

Size an inference deployment See live inventoryAcknowledged within one hour, first sourcing pass within one business day.
At a glance

What decides this requirement.

Deciding factor
Memory capacity and latency to users
Typical shape
Right-sized cards, distributed geographically, high utilisation
Fabric
Far less critical than for training; most inference is single-node
Economics
Steady utilisation makes reserved capacity strongly favourable
Inference is not small training

The requirements barely overlap.

Training is a batch job that runs flat out for weeks across many nodes in one place. Inference is a continuous service, latency-sensitive, usually fitting within a single node, and serving users who are somewhere specific. Almost every design decision follows differently from that.

Fabric quality, which dominates training, is largely irrelevant to single-node inference. Geography, which barely matters for training, becomes the primary constraint. And utilisation is steady rather than bursty, which changes the buying structure entirely.

The practical consequence: flagship silicon bought for an inference workload is the most common overspend we encounter. A great deal of production serving runs beautifully on previous-generation or right-sized parts at a fraction of the cost.

A compact edge data cabinet standing alone in a clean modern facility
What actually sizes an inference deployment

Memory, then concurrency, then location.

Model weights have to fit in GPU memory, and quantisation changes that arithmetic significantly. Once weights fit, the KV cache for concurrent requests consumes the remainder, which is why concurrency targets drive memory sizing as much as model size does.

Then geography. Latency to users is a product decision as much as a technical one, and it usually means several smaller deployments rather than one large one. Our regional pages cover where capacity actually exists.

Because utilisation is steady and predictable, inference is the strongest possible case for reserved capacity over on-demand pricing. The comparison is in cloud versus bare metal.

Where to next

Start from what is verified.

Current verified lines with quantities, lead times and indicative pricing are public on the live inventory. Anything not listed becomes a sourcing requirement with a first pass inside one business day.

Tell us what you need See live inventoryAcknowledged within one hour, first sourcing pass within one business day.
Straight answers

Asked first, answered straight.

What GPU do we need for inference?

Enough memory for the model weights plus the KV cache at your concurrency target. That is frequently far less card than teams assume, and previous-generation or right-sized parts often serve production perfectly.

Should inference run on the same hardware as training?

Usually not. Training wants fabric and density in one location; inference wants memory and geographic distribution. Buying one estate for both typically means overpaying for the inference half.

Does inference need InfiniBand?

Rarely. Most inference fits within a single node, so the expensive east-west fabric that training depends on does very little for it.

Is reserved capacity right for inference?

Almost always, because utilisation is steady. On-demand pricing exists to buy flexibility, and a service running continuously is not using that flexibility while paying for it.

Talk to the desk

Working through this on a real requirement?

Twenty minutes with the desk, no pitch and no quote at the end of it. Tell us roughly what you need and we will come back within one business day.

Acknowledged within one hour, first sourcing pass within one business day.

Sent. We are on it.

Your enquiry has landed with the desk. Acknowledged within one hour.