Use case · Fine-tuning

GPU capacity for fine-tuning and post-training.

For teams adapting existing models: the footprint is smaller than most expect, and the desk prices it honestly. Fine-tuning sits between training and inference in every dimension: smaller than a pre-training run, heavier than serving, and far more bursty than either. That shape rewards a different buying structure.

Scope a fine-tuning setup See live inventoryAcknowledged within one hour, first sourcing pass within one business day.
At a glance

What decides this requirement.

Deciding factor
Memory headroom and burst availability
Typical shape
Single node to a handful of nodes, used intensively then idle
Method matters
Parameter-efficient methods cut the footprint dramatically
Structure
The strongest case on the site for mixing reserved and on-demand
Method decides the footprint

LoRA and full fine-tuning are different purchases.

Full fine-tuning updates every parameter and needs memory for weights, gradients and optimiser state, which puts it close to pre-training requirements per unit of model. Parameter-efficient approaches such as LoRA update a small fraction and cut the memory footprint dramatically, frequently bringing a model that needed a multi-node cluster onto a single node.

That difference is worth establishing before anyone quotes hardware, because it can change the requirement by an order of magnitude. We ask about method before we ask about budget.

Quantised and mixed-precision approaches shift the arithmetic again. The right answer is workload-specific and worth ten minutes of conversation rather than a guess.

A small cluster of GPU servers on a lab bench with one chassis open
Bursty demand wants a hybrid

Reserve the floor, burst the peaks.

Fine-tuning is intensive for days then quiet for weeks. Paying reserved rates for peak capacity that sits idle is waste; paying on-demand rates for a continuous baseline is also waste. The correct answer for most teams is both.

We size the reserved floor against what you genuinely run continuously, and leave the peaks to on-demand or GPUaaS capacity. The structures are on the leasing desk and GPUaaS page.

This is the clearest case on the whole site for not buying one structure. Teams that force fine-tuning into a single commitment shape almost always overpay on one side of it.

Where to next

Start from what is verified.

Current verified lines with quantities, lead times and indicative pricing are public on the live inventory. Anything not listed becomes a sourcing requirement with a first pass inside one business day.

Tell us what you need See live inventoryAcknowledged within one hour, first sourcing pass within one business day.
Straight answers

Asked first, answered straight.

Does fine-tuning need a full cluster?

It depends entirely on method. Parameter-efficient approaches like LoRA can bring a model that would otherwise need multiple nodes onto a single one. Full fine-tuning is much closer to pre-training in its requirements.

Should we reserve or use on-demand?

For most fine-tuning teams, both. Reserve what you run continuously, burst the peaks. Forcing bursty demand into one commitment shape means overpaying on one side of it.

What memory do we need?

More than inference, less than full pre-training, and heavily dependent on method and precision. Tell us the base model and the approach and we will size it properly.

Can you source single nodes?

Yes. Single-node and small multi-node requirements are entirely normal and we source them on the same terms as larger deployments.

Talk to the desk

Working through this on a real requirement?

Twenty minutes with the desk, no pitch and no quote at the end of it. Tell us roughly what you need and we will come back within one business day.

Acknowledged within one hour, first sourcing pass within one business day.

Sent. We are on it.

Your enquiry has landed with the desk. Acknowledged within one hour.