The most useful hour in AI infrastructure planning happens before the first supplier call. These pages cover the arithmetic that changes every conversation afterwards.
Almost every expensive mistake in AI infrastructure is made before a single quote is requested. A number gets chosen early, usually because it sounds substantial or because it matches someone else's announcement, and then every subsequent decision is bent to justify it. Facilities are shortlisted against it. Budgets are approved against it. By the time the arithmetic gets done properly, changing the answer means restarting the procurement.
Sizing done first inverts that. You arrive at supplier conversations with a specification rather than an aspiration, and the conversations get dramatically shorter. You can tell immediately when a quote has quietly dropped a line, because you already know which lines exist. Most usefully, you stop negotiating hardest on the number that matters least.
None of this requires specialist tooling. It requires an evening, a calculator and a willingness to write down the assumptions you are actually making about utilisation, model size and time to result.
The order matters. Begin with memory requirements, because a model that does not fit cannot run at any budget, and memory is what eliminates most of the catalogue immediately. Then move to how many GPUs a given model needs, which is where time-to-result enters and where most published estimates quietly assume utilisation nobody achieves.
With a GPU count in hand, convert it into nodes, racks and megawatts. This is the step that turns a hardware question into a facility question, and it is where a surprising number of projects discover their preferred site cannot host the design. Power and cooling then tells you which halls remain in contention.
Only at that point does cost become meaningful. The full bill sets out every line in the order the costs actually arrive, and lead times covers the schedule risk that turns a good decision into a late one. If you are sizing for serving rather than training, read training versus inference hardware before anything else, because the two requirements diverge earlier than most people expect.
Start from memory, not compute. Establish whether the model fits, then work backwards from your target time to result at a realistic utilisation figure rather than a theoretical one. The count that falls out of that is the honest starting point.
It is the highest-return hour in the project. Buyers who arrive with a specification get shorter, more specific conversations and can spot an incomplete quote immediately, because they already know which lines a complete one contains.
Assuming utilisation they will not achieve, and forgetting that the accelerator is one of five cost drivers. Fabric, power and cooling, storage and the freight and import path routinely add up to more than the GPUs.
Memory arithmetic, time-to-train and why there is no single number.
Nodes not cards, and the five dials that set the real cost.
Weights, optimiser state, KV cache and what it means for SKU choice.
The conversion that turns a hardware question into a facility question.
What a cluster actually draws, and which halls can host it.
Why one estate for both is the expensive default.
Every line on the bill, in the order the costs arrive.
What is realistic, and why quoted timelines slip.
Twenty minutes with the desk, no pitch and no quote at the end of it. Tell us roughly what you need and we will come back within one business day.
Your enquiry has landed with the desk. Acknowledged within one hour.
Run this before the first supplier call and every later conversation gets shorter.
The conversion that turns a number into something a supplier can price