Modern serving frameworks changed what inference hardware needs to look like. Continuous batching and paged attention made memory management far more efficient, which means the right card is frequently smaller than teams assume.
Techniques such as continuous batching and paged attention substantially improved how efficiently GPU memory is used during serving, raising the concurrency a given card can sustain. That is a genuine change in the hardware requirement, not a marginal optimisation.
The practical consequence for a buyer: hardware recommendations from before these techniques became standard tend to over-specify. Teams that size on older assumptions buy more card than they need, then run it well below capacity.
Memory remains the binding constraint, but it is now weights plus a more efficiently managed KV cache. Sizing should be done against your actual concurrency and context length targets rather than a generic recommendation.

Once a card can serve your concurrency, the next question is where it sits. Users experience latency, not FLOPS, and that usually argues for several smaller deployments near users rather than one large one in a cheap location.
That is the opposite of the training answer, and buying one estate for both is the expensive mistake covered on training versus inference hardware.
Where capacity actually exists by region is on the regional pages. Steady serving utilisation also makes reserved capacity strongly favourable over on-demand pricing.
Current verified lines with quantities, lead times and indicative pricing are public on the live inventory. Anything not listed becomes a sourcing requirement with a first pass inside one business day.
Enough memory for weights plus KV cache at your concurrency and context targets. Modern serving techniques raised achievable concurrency per card, so the answer is frequently smaller than older guidance suggests.
It reduces weight memory substantially, leaving more for KV cache and concurrency. The quality cost is workload-dependent and worth measuring against your own evaluations.
Distribute, usually. Users experience latency, so several deployments near them typically beat one large one somewhere cheap. Training is the opposite case.
Reserved, almost always. Serving utilisation is steady, and on-demand pricing exists to buy flexibility you are not using while paying for it.
Twenty minutes with the desk, no pitch and no quote at the end of it. Tell us roughly what you need and we will come back within one business day.
Your enquiry has landed with the desk. Acknowledged within one hour.