Training, inference, fine-tuning, HPC and rendering want genuinely different machines. Buying one estate for all of them is the most common way to overspend.
Two teams can ask for four hundred accelerators and need almost nothing in common. One is training a large model from scratch and needs tight multi-node interconnect, sustained power and a fabric that keeps expensive silicon from waiting on itself. The other is serving inference at scale and would be better off with more, smaller, geographically distributed units and a fraction of the networking budget.
Quoted as a GPU count, those two requirements look identical. Priced properly, they diverge by a wide margin, and the wrong shape costs far more than the wrong vendor ever will. This is why the desk starts from what the workload does rather than from what hardware is available.
If you are training or continuing to pre-train a large model, LLM training covers the interconnect, power and commitment questions that dominate that case. If you are adapting an existing model, fine tuning is a substantially smaller footprint than most teams assume, and the method you choose changes the answer more than the model size does.
For production serving, LLM inference deals with concurrency, memory headroom and why inference is not simply training at a smaller scale. For traditional scientific and engineering computing, HPC begins from precision requirements, which immediately point at different silicon. And for graphics, simulation and visual production, rendering and visualisation covers why HGX-class hardware is usually the wrong purchase.
If your requirement spans more than one of these, that is normal and worth saying out loud early. A single estate serving both training and inference is the expensive default, and splitting the requirement is frequently the cheaper answer.
Materially, yes. Training rewards interconnect bandwidth and sustained power. Inference rewards memory headroom, distribution and cost per request. Buying one estate to do both well is usually the most expensive option available.
You can, and it is frequently the wrong economic choice. The two have different utilisation profiles and different failure costs, and sizing a single estate to satisfy both means over-buying for one of them.
Ask what the hardware will spend most of its hours doing. Building a model from scratch points at training, adapting one points at fine tuning, and answering user requests points at inference. Most teams eventually need two of the three.
Fabric first, then memory, then node count. Sized from a target time-to-train.
Memory and geography, not fabric. Steady utilisation favours reserved capacity.
Bursty demand and method-dependent footprints. The clearest case for a hybrid structure.
Double precision, MPI topologies and institutional procurement processes.
Frame-parallel work that rarely needs HGX platforms or exotic fabric.
Twenty minutes with the desk, no pitch and no quote at the end of it. Tell us roughly what you need and we will come back within one business day.
Your enquiry has landed with the desk. Acknowledged within one hour.
The hours, not the count, price the deal. That is why the first question is what the hardware is for.
Two identical requirements walk in, five different estates walk out