Integration · Serving

GPU infrastructure for vLLM and model serving.

Modern serving frameworks changed what inference hardware needs to look like. Continuous batching and paged attention made memory management far more efficient, which means the right card is frequently smaller than teams assume.

Tell us what you need See live inventoryAcknowledged within one hour, first sourcing pass within one business day.
At a glance

What decides this requirement.

Where it fits
Production LLM serving at scale
Key constraint
Memory for weights plus KV cache at target concurrency
Changed the maths
Efficient KV cache management raised achievable concurrency
We handle
Right-sized cards and geographic distribution
Serving frameworks changed the sizing

Concurrency per GPU went up.

Techniques such as continuous batching and paged attention substantially improved how efficiently GPU memory is used during serving, raising the concurrency a given card can sustain. That is a genuine change in the hardware requirement, not a marginal optimisation.

The practical consequence for a buyer: hardware recommendations from before these techniques became standard tend to over-specify. Teams that size on older assumptions buy more card than they need, then run it well below capacity.

Memory remains the binding constraint, but it is now weights plus a more efficiently managed KV cache. Sizing should be done against your actual concurrency and context length targets rather than a generic recommendation.

A GPU server photographed from above with its lid removed
Geography beats raw specification

Latency is a product decision.

Once a card can serve your concurrency, the next question is where it sits. Users experience latency, not FLOPS, and that usually argues for several smaller deployments near users rather than one large one in a cheap location.

That is the opposite of the training answer, and buying one estate for both is the expensive mistake covered on training versus inference hardware.

Where capacity actually exists by region is on the regional pages. Steady serving utilisation also makes reserved capacity strongly favourable over on-demand pricing.

Where to next

Start from what is verified.

Current verified lines with quantities, lead times and indicative pricing are public on the live inventory. Anything not listed becomes a sourcing requirement with a first pass inside one business day.

Tell us what you need See live inventoryAcknowledged within one hour, first sourcing pass within one business day.
Straight answers

Asked first, answered straight.

What GPU do we need for vLLM?

Enough memory for weights plus KV cache at your concurrency and context targets. Modern serving techniques raised achievable concurrency per card, so the answer is frequently smaller than older guidance suggests.

Does quantisation help serving?

It reduces weight memory substantially, leaving more for KV cache and concurrency. The quality cost is workload-dependent and worth measuring against your own evaluations.

Should we centralise or distribute serving?

Distribute, usually. Users experience latency, so several deployments near them typically beat one large one somewhere cheap. Training is the opposite case.

Reserved or on-demand for serving?

Reserved, almost always. Serving utilisation is steady, and on-demand pricing exists to buy flexibility you are not using while paying for it.

Talk to the desk

Working through this on a real requirement?

Twenty minutes with the desk, no pitch and no quote at the end of it. Tell us roughly what you need and we will come back within one business day.

Acknowledged within one hour, first sourcing pass within one business day.

Sent. We are on it.

Your enquiry has landed with the desk. Acknowledged within one hour.