The unit the market actually uses
Serious AI hardware moves in nodes: a server carrying (typically) eight GPUs plus the CPUs, memory, storage and network cards around them. Four hundred GPUs is fifty HGX nodes. That translation sounds trivial, and everything expensive hides inside it.
A node is not eight GPUs at the GPU price. It is a machine with its own cost structure, and it drags obligations behind it: rack space at a density most halls cannot supply, cooling matched to the generation, a fabric connecting it to its neighbours, and freight at weights that surprise people. Budget in GPUs and every one of those lines lands later, as an overrun.
The five dials
When I size a cluster, five choices set the real number, and the GPU rate everyone negotiates hardest is reliably the smallest lever of the five:
- The SKU. Obvious, and often chosen wrong: flagship silicon for a workload a cheaper generation runs happily.
- The node configuration. CPU, memory and storage around the GPUs. Under-spec it and the GPUs starve; over-spec it and you bought servers you did not need.
- The fabric. How nodes talk to each other. For distributed training this decides whether fifty nodes behave like a cluster or like fifty computers.
- The facility. Power and cooling per rack decide which buildings can even host the machine, and the qualifying list is shorter than the market maps pretend.
- The term. How long you commit, on which structure: rent, reserve or own. This dial moves more money than any hardware discount ever will.
Budget in nodes and the whole system they need, not in GPUs. The card price is the headline; the node is the bill.
A worked translation
Someone says: we need four hundred H200s. In node terms that is fifty HGX machines, megawatt-class power once cooling overhead is counted, hundreds of transceivers and cables for the fabric, freight measured in tonnes, and a facility shortlist you can count on one hand in some countries. None of that appears in a per-GPU quote, and all of it appears in the project.
The buyers who get this right run the translation before they shop. The ones who do not, meet it during delivery, which is the most expensive classroom in the industry.
Running the translation yourself
You do not need a desk to do the first pass; you need an evening and honesty. Take your GPU number and divide by eight: that is your node count. Multiply nodes by the generation's real power draw with cooling overhead, and you have the megawatt conversation you are about to have with a facility. Sketch the fabric: every node wants multiple high-speed ports, every port wants a transceiver at each end, and the total will be a number that surprises you. Add freight at roughly a tonne per handful of nodes, then ask which buildings in your target country can actually take the density. That one exercise, done before any supplier call, changes every conversation that follows.
It also changes the internal ones. Boards approve GPU budgets and then meet the node bill in instalments, which is how AI projects earn their reputation for overruns. Presenting the node-complete number once, up front, is briefly uncomfortable and permanently cheaper. The teams I rate highest walk into their first sourcing call already speaking in nodes, megawatts and weeks. It takes an evening to get there, and it is the highest-return evening in the entire project.

