Most GPU clusters are planned backwards. Someone gets a quote for a round number of servers, and the workload is fitted to whatever arrives. The result is familiar: a training run that takes twice the promised time because nobody budgeted for failures, an inference fleet sized for average load that falls over at the daily peak, or a storage system that stalls a thousand GPUs every time a checkpoint is written.

This article plans in the other direction. It starts from what the workload needs, in floating-point operations, bytes of memory and tokens per second, converts that into a GPU count, and then sizes the parts people forget: the network fabric, checkpoint bandwidth, the failure budget and the lead time. Every number in the worked example comes from a small Python model you can run and change. Which GPU to buy is a separate decision, covered in GPU selection in depth; here the GPU type is an input.

Advertisement

The planning chain

Workload demandtokens, models, SLOsComputeFLOPs / (peak x MFU)Memorystates, KV cacheGPU countmax of the twoFabricNIC per GPU, fat treeStoragecheckpoint GB/s, dataFailure budgetMTBF, spares, goodputFacility envelopekW per rack, MW, coolingPhasing and lead timeorders, burn-in, growthCapacity plancount, topology, budget, dates
The planning chain: demand sets compute and memory, the larger of the two sets the GPU count, and the count drives the fabric, storage and failure budget, which then meet the facility envelope and lead times.

A capacity plan is a chain of conversions, and each link has one dominant input. Training demand is measured in total floating-point operations. Inference demand is measured in tokens per second at a latency target. Memory sets a floor that compute cannot go below. Once you have a GPU count, it fixes how many network ports you need, how fast storage must absorb a checkpoint, and how often the job will be interrupted. Only then does the plan meet the building: power per rack, total megawatts and cooling, which the datacenter power article covers in detail.

Write every link down with its assumptions. A plan whose numbers cannot be traced back to an input is a guess, and guesses are how clusters end up 40 percent idle or permanently short.

Training demand: FLOPs, MFU and wall-clock

For a dense transformer, one training token costs about 6 times the parameter count in floating-point operations: 2N for the forward pass and 4N for the backward pass. A run over D tokens therefore needs about 6ND operations. This ignores attention's quadratic term, which matters at long context, so add a margin when sequences are long relative to model width.

The GPU's datasheet peak is not what you get. Model FLOPs utilisation (MFU) is the fraction of peak spent on useful model maths. Well-tuned large dense runs usually land somewhere around 35 to 50 percent; small models, long pipelines, heavy communication or mixture-of-experts routing push it lower. Use a figure you have measured on your own stack if you have one. Use the dense figure for peak, never the structured-sparsity figure: an H100 SXM is rated at 989 dense BF16 TFLOPS, and the sparse number is double that and irrelevant to training.

So the ideal wall-clock is 6ND divided by (GPUs times peak times MFU). The word ideal matters, because real runs lose time to failures, restarts and checkpoints, which the failure budget below adds back.

Advertisement

Memory is a floor, not the answer

Mixed-precision training with Adam holds about 16 bytes per parameter: 2 for BF16 weights, 2 for gradients, 4 for an FP32 master copy and 8 for the two Adam moments. A 70-billion-parameter model needs about 1.1 TB for that state alone, before activations. With full sharding of states across data-parallel ranks (see ZeRO sharding), 1.1 TB spread over even 64 GPUs is under 20 GB each, so memory only sets a floor of a few dozen 80 GB GPUs. For almost every serious pre-training run, compute, not memory, decides the count. Memory decides the parallelism layout inside the count.

Inference flips this. Serving is usually memory-bound, and the KV cache is the variable part. Per token it costs 2 (keys and values) times layers times KV heads times head dimension times bytes per element. For a Llama-3-70B-shaped model, 80 layers, 8 KV heads and a head dimension of 128 in BF16, that is 327,680 bytes, about 320 KB per token. On an 8-GPU node with 640 GB of HBM, 140 GB of BF16 weights and a 10 percent reserve leave about 436 GB, roughly 1.33 million cached tokens, or about 325 concurrent sequences at 4,096 tokens each. That number, not peak FLOPS, bounds how many users one node can hold.

A planning model you can run

The model below turns the training assumptions into wall-clock days with failures included. The failure rate is derived from a public data point: Meta's Llama 3 paper reports 419 unexpected interruptions in a 54-day snapshot on 16,384 H100 GPUs, which is roughly one interruption per 50,000 GPU-hours. A job fails when any of its GPUs, hosts or links fails, so the job's mean time between failures is that figure divided by the GPU count. The checkpoint interval uses the Young/Daly first-order optimum, the square root of twice the checkpoint stall times the job MTBF.

import math

def training_plan(params, tokens, gpus, peak_flops, mfu, mtbf_gpu_hours,
                  ckpt_stall_s, restart_s):
    """Wall-clock days for a dense-transformer pre-training run, failures included."""
    flops = 6 * params * tokens                        # forward + backward
    ideal_s = flops / (gpus * peak_flops * mfu)
    mtbf_s = mtbf_gpu_hours * 3600 / gpus              # any GPU failing stops the job
    interval_s = math.sqrt(2 * ckpt_stall_s * mtbf_s)  # Young/Daly optimum
    # time lost per unit of work: checkpoint stalls + half an interval redone + restart
    overhead = ckpt_stall_s / interval_s + (interval_s / 2 + restart_s) / mtbf_s
    wall_s = ideal_s / (1 - overhead)
    return dict(gpu_hours=ideal_s * gpus / 3600, ideal_days=ideal_s / 86400,
                job_mtbf_h=mtbf_s / 3600, ckpt_every_min=interval_s / 60,
                goodput=1 - overhead, wall_days=wall_s / 86400)

mtbf = 54 * 24 * 16384 / 419          # about 50,700 GPU-hours per interruption
print(training_plan(params=70e9, tokens=2e12, gpus=1024, peak_flops=989e12,
                    mfu=0.40, mtbf_gpu_hours=mtbf, ckpt_stall_s=60, restart_s=900))

The inputs are deliberately explicit. Change the GPU type by changing peak_flops, change the software stack by changing mfu, and change your reliability assumption by replacing the Llama 3 figure with your own fleet's interruption rate once you have one.

Worked example: a 70B run on 1,024 GPUs

Take a 70-billion-parameter dense model trained on 2 trillion tokens, on 1,024 H100 SXM GPUs at 40 percent MFU. Running the model above gives:

QuantityValueHow it was derived
Training compute8.4 x 10^23 FLOPs6 x 70e9 x 2e12
GPU-hours at 40% MFUabout 590,000FLOPs / (989 TFLOPS x 0.40)
Ideal wall-clock24.0 daysGPU-hours / 1,024
Job MTBFabout 49.5 hours50,700 GPU-hours / 1,024
Checkpoint intervalabout 77 minutessqrt(2 x 60 s x MTBF)
Goodputabout 96.9%stalls + rework + 15-minute restarts
Planned wall-clockabout 24.8 daysideal / goodput

Two lessons fall out. First, a job this size will be interrupted roughly every two days, so the restart path is part of the production system, not an edge case. Second, the 15-minute restart assumption matters as much as the checkpoint interval: if restarts take an hour because a scheduler has to find and drain a replacement node, goodput drops by several points. Keep hot spares, typically a few percent of the fleet, so a failed node is swapped rather than repaired before the job can resume.

The model's 96.9 percent is an upper bound. Real runs also lose time to stragglers, data-loader stalls and evaluation pauses. Plan against 90 percent unless your own history says otherwise, which turns 24 ideal days into about 27.

Inference demand: measure, then add headroom

Inference throughput depends too much on the serving engine, batch policy, prompt mix and latency target to derive from datasheets. The reliable method is empirical. Replay a realistic prompt and output-length distribution against one node, raise the load until the p95 time to first token or inter-token latency breaks your SLO, and record the output tokens per second at the last load that met it. That is the node's capacity at your SLO, and it is often well below its throughput with no latency target at all.

Then size for peak, not average. Suppose a benchmark shows 12,000 output tokens per second per 8-GPU node at the SLO (an illustrative figure; measure yours) and forecast peak demand is 60,000 tokens per second. Five nodes would be saturated, so plan to run at 70 percent at peak to absorb bursts and slow requests: 60,000 / (12,000 x 0.7) = 7.1, rounded up to 8. Add at least one node for failure and rolling upgrades, making 9 nodes, 72 GPUs. If traffic is regional, apply the calculation per region, since spare capacity on another continent does not help a latency SLO.

Fabric, storage and the facility

Training fabric is sized from the GPU count. Current large-cluster designs give each GPU its own high-speed NIC (400 Gb/s per GPU is common for H100-class systems) and build a non-blocking fat tree, often rail-optimised so GPU i of every node lands on the same leaf group. With switches of radix k, a two-tier fat tree connects up to k squared over 2 endpoints: 64-port switches give 2,048. For 1,024 GPUs, 32 leaves each use 32 ports down and 32 up, and 16 spines of 64 ports absorb the 1,024 uplinks. Past the two-tier limit you add a third tier, more optics, more latency and a harder failure domain; the InfiniBand article covers the transport side.

Storage has two jobs. Training data needs steady read bandwidth, modest for text and much higher for images and video. Checkpoints need bursts: the 70B example writes about 980 GB per checkpoint (14 bytes per parameter for weights, master copy and Adam moments), so a 60-second stall needs about 16 GB/s aggregate. Asynchronous checkpointing, which copies state to host memory and writes in the background, shrinks the stall and changes the requirement from peak burst to sustained throughput between checkpoints.

Finally, the facility. Take system power from the vendor's server specification, not the GPU TDP: GPU TDP alone (700 W for an H100 SXM) substantially understates what a server draws once CPUs, memory, NICs and fans are counted. Multiply by nodes, add network and storage, and check the result against per-rack and hall limits before any order is placed.

Phasing, lead time and trade-offs

Hardware arrives in phases, and each phase needs burn-in before it carries production work: run stress tests and collective benchmarks on every node, and pull anything that underperforms. A node that passes power-on but runs its links at reduced speed will slow every collective it joins. Plan the schedule backwards from when capacity must be usable, including delivery, racking, cabling, burn-in and scheduler integration.

DecisionOption AOption BWhat decides it
Own vs rentOwn: lowest unit cost at high utilisationRent: no lead time, no stranded capacityExpected utilisation over 3 years
One big cluster vs severalOne: largest possible jobsSeveral: smaller failure domainsLargest single job you must run
SparesMore: faster restartsFewer: more GPUs workingRestart time vs repair time
Fabric oversubscriptionNon-blocking: predictable collectivesOversubscribed: cheaperTraining vs inference-only use

Failure modes in capacity planning

  • Planning from peak TFLOPS. Using the sparse figure, or assuming 100 percent MFU, overstates capacity by a factor of two to five.
  • No failure budget. The run is promised at ideal wall-clock, then slips every week as interruptions accumulate.
  • Sizing inference for average load. The fleet is fine at noon and saturated at the daily peak, and latency SLOs fail first.
  • Checkpoint storage sized for capacity, not bandwidth. Petabytes are provisioned but writes stall every GPU for minutes.
  • Ignoring lead time. The plan is correct and the hardware arrives two quarters after the run was due to start.
  • Plan never revisited. Actual MFU and failure rates are known a month in; feed them back and re-forecast.

What to do next

  1. Write demand down in units: total training FLOPs per run and peak output tokens per second per region, each with its source.
  2. Measure MFU on your stack at a small scale, and measure per-node inference capacity at your latency SLO.
  3. Run the planning model with your numbers; keep the inputs in version control next to the plan.
  4. Derive the fabric from the GPU count and switch radix, and the checkpoint bandwidth from state size and acceptable stall.
  5. Budget failures explicitly: job MTBF, checkpoint interval, restart time and a hot-spare pool.
  6. Check power and cooling against vendor system figures, then schedule orders backwards from the date capacity must be usable.
  7. After the first month, replace every assumption with measured values and re-forecast.
Key takeaway: Plan GPU infrastructure from the workload down. Training needs about 6ND FLOPs divided by peak times a measured MFU; inference needs measured per-node throughput at the latency SLO, sized for peak with headroom and a spare. Memory is a floor, not the count. The GPU count then fixes the fabric, the checkpoint bandwidth and the failure budget: at 1,024 GPUs, public failure rates imply an interruption about every two days, so checkpoint intervals, restart time and spares are part of the plan. Keep the model runnable and replace every assumption with measurements as soon as you have them.