Sizing · Method

How to size a GPU cluster.

Most cluster sizing starts with a GPU count and works outward. That is backwards, and it is why so many deployments arrive over budget and under-performing.

Tell us what you need See live inventoryAcknowledged within one hour, first sourcing pass within one business day.
At a glance

What decides this requirement.

Unit
Nodes, not GPUs. Four hundred GPUs is fifty HGX nodes
Five dials
SKU, node configuration, fabric, facility, term
Smallest lever
The GPU rate everyone negotiates hardest
Largest lever
The term you commit to
Think in nodes

The unit the market actually sells.

Requirements arrive in GPUs and get delivered in nodes: a server carrying typically eight accelerators plus CPUs, memory, storage and network cards. Four hundred GPUs is fifty machines, and fifty machines is a power conversation, a cooling conversation and a freight plan.

Everything expensive hides in that translation. A node drags obligations behind it that a card does not, and a budget built on card prices meets those obligations later, as overruns.

The full argument is on the think in nodes article, and the practical arithmetic below.

A partially built GPU cluster with racks installed and one bay still empty
The five dials

Turn them in the right order.

Unfamiliar with any of the terms here? The glossary defines them the way the desk actually uses them, and the blog works through the reasoning behind each dial.

The SKU. Chosen wrong more often than any other dial, usually by picking flagship silicon for a workload a cheaper generation runs happily.

The node configuration. CPU, memory and storage around the accelerators. Under-spec and the GPUs starve; over-spec and you bought servers you did not need.

The fabric. For distributed training this decides whether the cluster behaves like one machine or many. For inference it barely matters. Sizing it identically for both is a common and costly error.

The facility. Power and cooling per rack decide which buildings can host the machine at all, and that list is shorter than the market implies.

The term. The largest lever of the five, and the one negotiated least carefully.

Do the first pass yourself

One evening changes every later conversation.

Divide your GPU count by eight for nodes. Multiply nodes by the generation's real draw plus cooling overhead for the power conversation. Count fabric ports, then transceivers at both ends of each. Add freight at roughly a tonne per handful of nodes. Then ask which buildings in your target market can take that density.

That exercise, done before any supplier call, changes every conversation that follows. It also produces the node-complete number to take to your board once, rather than meeting it in instalments.

Bring the result to us and we will price it properly, including the layers most quotes leave out. The cost calculator models the commercial structures on top.

Where to next

Start from what is verified.

Current verified lines with quantities, lead times and indicative pricing are public on the live inventory. Anything not listed becomes a sourcing requirement with a first pass inside one business day.

Tell us what you need See live inventoryAcknowledged within one hour, first sourcing pass within one business day.
Straight answers

Asked first, answered straight.

Why size in nodes rather than GPUs?

Because nodes are what the market sells and what facilities host. A node carries CPUs, memory, storage and networking alongside the accelerators, and drags power, cooling and freight obligations that a card price does not reflect.

Which decision has the biggest cost impact?

The term you commit to, by a wide margin, followed by the facility. The GPU rate that receives the most negotiation attention is reliably the smallest of the five levers.

How much power will our cluster need?

Multiply node count by the generation's real draw and add cooling overhead. Even modest deployments become megawatt-class conversations quickly, which is why we qualify power before hardware.

Can you size it for us?

Yes, and it takes about twenty minutes with your workload, timeline and constraints. There is no cost and it does not end in a quote unless you ask for one.

Talk to the desk

Working through this on a real requirement?

Twenty minutes with the desk, no pitch and no quote at the end of it. Tell us roughly what you need and we will come back within one business day.

Acknowledged within one hour, first sourcing pass within one business day.

Sent. We are on it.

Your enquiry has landed with the desk. Acknowledged within one hour.