Use case · LLM training

GPU clusters for training large language models.

For teams training or continuing to pre-train large models: the desk sources the cluster, fabric and facility as one file. Training a frontier or mid-size language model is the most demanding thing you can ask of infrastructure. It is also where sizing mistakes cost the most, because an undersized fabric wastes the GPUs you fought hardest to get.

Send your training requirement See live inventoryAcknowledged within one hour, first sourcing pass within one business day.
At a glance

What decides this requirement.

Deciding factor
Interconnect quality, not raw GPU count
Typical shape
Multi-node HGX clusters with high-speed east-west fabric
Memory
Model and optimiser state per GPU usually sets the SKU
Term
Match the commitment to the training roadmap, not the vendor's quarter
What actually decides the cluster

Fabric first, then memory, then count.

Distributed training is a communication problem wearing a compute costume. Gradient synchronisation across nodes runs constantly, and a fabric that cannot keep up leaves expensive accelerators waiting. This is why two clusters with identical GPU counts can differ enormously in effective throughput, and why we size the interconnect before the node count.

Memory per GPU is the second gate. Model parameters, optimiser state and activations have to fit, and the moment they do not you are into sharding strategies that cost throughput. That is frequently the argument for H200 over H100, or B300 over B200: not raw speed, but avoiding the need to buy more nodes than the compute requires.

Only then does node count matter, and it should follow from a target time-to-train rather than a budget line. We work backwards from the run you actually need to complete.

A vast AI training hall with long parallel rows of GPU racks
Storage and the parts people forget

A starved cluster is an idle cluster.

Training pipelines read enormous datasets continuously and write checkpoints periodically, both of which will expose a storage layer sized by guesswork. Checkpoint cadence in particular deserves explicit design: too frequent and you burn throughput, too rare and a failure costs days of progress.

Then the unglamorous remainder: optics and cabling in quantities that surprise first-time builders, power and cooling matched to the generation, and a facility that can genuinely carry the density. Each is covered on the networking desk and the space and power desk.

We price the whole configuration rather than the accelerators alone, because a cluster quoted on GPUs is a cluster that will surprise you twice.

Buying structure

Term the commitment to the roadmap.

For teams whose demand is bursty rather than sustained, the AI training and inference desk covers the commercial shape, and hardware and capacity the landed side.

Training demand is lumpy: intense during a run, quieter between them. That shape argues against paying on-demand rates for a sustained baseline and equally against a five-year commitment made against an eighteen-month roadmap.

The structures that work are covered on the leasing desk and compared honestly in buy versus lease. Financing routes exist for venture-funded teams whose runway is shorter than the useful life of the hardware.

Tell us the model size, target time-to-train and runway, and we will size the cluster and the structure together rather than separately.

Where to next

Start from what is verified.

Current verified lines with quantities, lead times and indicative pricing are public on the live inventory. Anything not listed becomes a sourcing requirement with a first pass inside one business day.

Tell us what you need See live inventoryAcknowledged within one hour, first sourcing pass within one business day.
Straight answers

Asked first, answered straight.

How many GPUs do I need to train a model?

It depends on parameter count, token budget and how long you are willing to wait. We work backwards from a target time-to-train rather than starting from a GPU number. The sizing pages walk through the arithmetic.

Does the interconnect really matter that much?

For multi-node training, yes, more than almost any other choice. Gradient synchronisation runs constantly, and a slow fabric leaves expensive GPUs idle. It is the most common and most costly sizing mistake we see.

H100, H200 or B300 for training?

Memory per GPU usually decides it. If the model, optimiser state and activations fit comfortably in 80GB, H100 is hard to beat on price. If they do not, the larger-memory parts avoid buying extra nodes to solve a memory problem.

Should we buy or lease a training cluster?

It follows from utilisation and runway rather than preference. Sustained, funded, multi-year training favours owning; uncertain roadmaps favour leasing. We run both models on your numbers.

Talk to the desk

Working through this on a real requirement?

Twenty minutes with the desk, no pitch and no quote at the end of it. Tell us roughly what you need and we will come back within one business day.

Acknowledged within one hour, first sourcing pass within one business day.

Sent. We are on it.

Your enquiry has landed with the desk. Acknowledged within one hour.