A GPU fleet is expensive to have and expensive to lack. Too few GPUs and training runs queue for days while inference misses its latency targets. Too many and you pay for silicon that sits idle, a cost that does not shrink because nobody is using it. Capacity planning is the practice of turning measured demand into a dated decision: how many GPUs, in what shape, bought or reserved by when.
This page covers the running fleet: the weekly and quarterly loop for a shared cluster with several teams on it. Sizing a single new build from a model's FLOPs is a different job, covered in GPU infrastructure planning. Here you will build a demand ledger, measure how job shapes waste GPUs, size a spare pool from failure data, size an inference fleet with Little's law, and set a buy trigger. Every number in the worked example comes from the code on this page. Failure rates and lead times are inputs you replace with your own.
Units: GPU-hours, utilisation targets and three fleet metrics
Plan in GPU-hours of useful work per week, per team or workload class. Avoid "GPUs requested": teams ask for peaks and keep allocations they are not using. GPU-hours come straight from scheduler accounting (Slurm's sacct, Kubernetes allocation metrics or your own job database).
Every workload class also gets a target utilisation: the fraction of its allocated hours that should be doing work. The target differs by class for a good reason. A pretraining job runs for weeks on a fixed allocation and can sit near 90%. An inference fleet needs headroom for traffic spikes and replica failures, so 60% is often healthy. Research and notebook work is bursty and interactive, so 50% is realistic. Dividing demand by utilisation converts useful work into the GPUs you must hold.
Keep three fleet-wide measures distinct, because each one answers a different question:
| Metric | Definition | What a bad value means |
|---|---|---|
| Allocated | GPUs assigned to a job or reservation | low: demand is lower than you planned for |
| Busy | GPUs whose SMs are active (for example DCGM's SM-active metric) | much lower than allocated: jobs hold GPUs they do not use |
| Stranded | free GPUs that no waiting job can use because of their shape | high: you have a packing problem, not a supply problem |
The demand ledger
The ledger is a small table you can review in a meeting: weekly GPU-hours, expected growth per quarter and target utilisation, one row per class. The code turns it into GPUs needed per quarter. Growth rates should come from each team's roadmap and be checked against the last four quarters of actuals.
HOURS = 168 # hours in a week
teams = { # class: (gpu_hours_per_week, growth_per_quarter, target_utilisation)
"pretrain": (60_000, 0.10, 0.90),
"finetune": (18_000, 0.25, 0.70),
"inference": (35_000, 0.30, 0.60),
"research": (12_000, 0.15, 0.50),
}
def gpus_needed(quarters_ahead):
total = 0
for gh, growth, util in teams.values():
demand = gh * (1 + growth) ** quarters_ahead
total += demand / (HOURS * util)
return total
for q in (0, 2, 4):
print(q, round(gpus_needed(q)))
# 0 1040
# 2 1495
# 4 2196Today's demand needs 1,040 GPUs. A year out it needs 2,196, roughly twice as many, because the fast-growing classes are also the ones with low utilisation targets. A 30% growth rate on inference held at 60% utilisation adds GPUs faster than the same rate on pretraining at 90%. That is the first non-obvious lesson: raising utilisation in a low-target class is often worth more than buying GPUs.
Shape: why a half-empty fleet cannot start a job
A ledger counts GPUs as if they were interchangeable. They are not. A distributed training job needs whole 8-GPU nodes on one fast fabric. A single-GPU fine-tune can land anywhere. If small jobs are scattered across many nodes, every node has a few free GPUs and none is completely free, so an 8-GPU or 16-GPU job waits while the fleet looks half empty. Those free GPUs are stranded.
The simulation below sends a stream of jobs of 1, 2, 4, 8 and 16 GPUs to 128 eight-GPU nodes. Jobs of 8 or more GPUs need whole nodes. The only thing that changes between runs is where small jobs go. Pack puts them on the fullest node that fits, and spread puts them on the emptiest. Jobs that cannot be placed are dropped, which keeps the model short; a real scheduler queues them.
import heapq, random
def simulate(policy, n_nodes=128, steps=20_000, seed=0):
rng = random.Random(seed)
free, ends = [8] * n_nodes, []
used_sum = stranded = 0
for t in range(steps):
while ends and ends[0][0] <= t:
for node, g in heapq.heappop(ends)[2]:
free[node] += g
size = rng.choices([1, 2, 4, 8, 16], weights=[30, 20, 20, 20, 10])[0]
if size >= 8: # multi-node jobs need whole nodes
whole = [i for i, f in enumerate(free) if f == 8][: size // 8]
alloc = [(i, 8) for i in whole] if len(whole) == size // 8 else []
else:
fits = [i for i, f in enumerate(free) if f >= size]
pick = min if policy == "pack" else max
alloc = [(pick(fits, key=lambda i: free[i]), size)] if fits else []
if not alloc:
stranded += sum(free) >= size # enough GPUs exist, wrong shape
else:
for node, g in alloc:
free[node] -= g
heapq.heappush(ends, (t + rng.randint(20, 400), t, alloc))
used_sum += 8 * n_nodes - sum(free)
return used_sum / (steps * 8 * n_nodes), stranded
# pack utilisation=0.902 stranded_rejections=451
# spread utilisation=0.504 stranded_rejections=3845Same hardware, same jobs, and utilisation drops from 90% to 50% just by changing where small jobs go. Under spread, 3,845 jobs were rejected even though enough GPUs were free in total. A capacity plan built on the spread fleet's numbers would order nearly twice the hardware it needs. Before you buy, fix placement. Pack small jobs (in Kubernetes, the NodeResourcesFit plugin with the MostAllocated scoring strategy does this; check your scheduler's equivalent), keep a dedicated pool of nodes for small jobs, or split GPUs with MIG so a notebook takes a slice instead of a whole H100.
Failures and the spare pool
Nodes fail. The Llama 3 paper gives one public data point: over a 54-day snapshot of pretraining on 16,384 H100 GPUs (2,048 eight-GPU nodes), the job saw 466 interruptions, 419 of them unexpected, and about 78% of the unexpected ones were attributed to confirmed or suspected hardware problems. Faulty GPUs alone caused 30.1% and HBM3 memory 17.2%. That works out to about 0.004 unexpected interruptions per node per day. If a failed node takes about a day to drain, diagnose and return, then roughly 0.4% of nodes are out of service at any moment. Treat that as a starting input, not a constant. Your rate depends on hardware age, burn-in and cooling.
A large synchronous job needs all its nodes, so a fleet must hold hot spares it can swap in. The spare count is a binomial tail: the smallest number s such that the chance of more than s nodes being down at once is below your risk budget.
import math
def spares_needed(nodes, p_down, confidence=0.99):
cdf, s = 0.0, 0
while True:
cdf += math.comb(nodes, s) * p_down**s * (1 - p_down)**(nodes - s)
if cdf >= confidence:
return s
s += 1
for n in (16, 128, 512):
print(n, spares_needed(n, 0.004))
# 16 1
# 128 3
# 512 6Spares grow more slowly than the job. A 16-node job needs one spare node (6% overhead) and a 512-node job needs six (about 1.2%). The model assumes independent failures. Correlated failures, such as a top-of-rack switch, a power feed or a bad firmware rollout, break that assumption, so place spares in different failure domains and run a separate check for losing one whole rack. Spares need not sit idle: let preemptible research jobs use them and evict those jobs when a spare is claimed.
Inference capacity with Little&#x27;s law
Inference capacity is a queueing problem. Little's law says the average number of requests in the system equals the arrival rate times the average time each spends there. At 40 requests per second, with replies of 400 output tokens streamed at 40 tokens per second to each user, each request lives about 10 seconds, so about 400 streams are in flight. What you must measure is how many output tokens per second one replica can produce while still meeting your latency targets. It is never the peak a benchmark prints with latency ignored.
def serving_replicas(peak_rps, tokens_out, tok_per_s_per_replica, headroom=0.3):
tok_rate = peak_rps * tokens_out # tokens/s the fleet must emit
return math.ceil(tok_rate / tok_per_s_per_replica / (1 - headroom))
print(serving_replicas(40, 400, 2_500)) # 10 replicas16,000 tokens per second at a measured 2,500 per replica is 6.4 replicas of pure work. With 30% headroom for spikes and a replica failure, that becomes 10. If each replica is one 8-GPU node, that is 80 GPUs at peak. Off-peak, the same fleet can run batch inference or evaluation jobs, which is how you defend a 60% average utilisation target without wasting the trough.
Supply and the buy trigger
Demand gives you GPUs needed per quarter. Supply comes in three forms with different prices and different commitments. Owned hardware has a long lead time and the lowest unit cost at high utilisation. Reserved cloud capacity has a fixed term and a discount; see reserved capacity for break-even maths. On-demand capacity is flexible but expensive, and at times unavailable for popular SKUs. A common pattern is to cover the steady base with owned or reserved capacity and the growth band with shorter commitments.
The buy trigger is the rule that turns the plan into an order. Find the first quarter where forecast need, plus spares, exceeds committed supply. If that quarter falls within your procurement lead time (order to racked, burned-in and accepted), place the order now. If it falls after, put it on the watch list and recompute next month. Lead times for current-generation accelerators have swung widely in recent years, so measure your own with vendors instead of using a published figure.
Worked example: one quarterly review
A shared fleet owns 1,280 GPUs (160 nodes). The largest pretraining job spans 40 nodes, so its partition holds 2 spare nodes, and the other 120 nodes share 3 more: 40 spare GPUs in all. Today the requirement is 1,040 plus 40, or 1,080, against 1,280 owned, a 200-GPU margin. That looks comfortable until you apply the buy trigger quarter by quarter.
OWNED, SPARE_GPUS, LEAD_TIME_Q = 1280, 40, 2
for q in range(5):
gap = gpus_needed(q) + SPARE_GPUS - OWNED
flag = "ORDER" if gap > 0 and q <= LEAD_TIME_Q else ""
print(q, round(gap), flag)
# 0 -200
# 1 4 ORDER
# 2 255 ORDER
# 3 567
# 4 956The first gap appears next quarter, inside a two-quarter lead time, so the order is already late. Hardware ordered today lands in quarter 2 at the earliest. Quarter 1 must be covered another way, and the cheapest way is the ledger itself. Raising inference from 60% to 70% utilisation, by moving batch evaluation into off-peak hours, saves about 64 GPUs in quarter 1. Raising research from 50% to 65%, by packing notebooks onto MIG slices instead of whole GPUs, saves about 38 more. Quarter 1 returns to a 99-GPU margin. With both levers in place, the quarter-2 gap falls from 255 to about 128 GPUs, or 16 nodes. Order those now, and keep a short reservation or on-demand budget as insurance in case delivery slips. The fleet is invented, but every figure comes from the code above, so you can rerun it with your own ledger.
Failure modes
- Planning from requests. Allocations reflect what teams asked for, not what they used. Always reconcile allocated with busy before forecasting.
- Ignoring shape. The fleet is 60% free on paper, but no 16-GPU job can start. The stranded metric catches this, and placement fixes are cheaper than hardware.
- Average-only inference sizing. Sizing for average traffic means peak traffic misses the SLO. Size for peak, then fill the trough with batch work.
- Spares that live in one rack. A single switch failure takes out the job and all its spares together.
- Lead time measured from the order, not acceptance. Burn-in and acceptance testing catch early failures, and they take time. Count them.
Trade-offs
High utilisation versus queue time. Above about 85-90% utilisation, queue waits grow sharply because there is no slack to absorb bursts. Decide which classes may queue (research) and which may not (production inference).
Packing versus isolation. Packing small jobs together raises utilisation but puts more jobs in each node's failure and noisy-neighbour domain. MIG gives hardware isolation at the cost of fixed slice sizes.
Owning versus renting. Owned hardware wins at high, steady utilisation. Renting wins when demand is uncertain. Model both with a total cost of ownership model rather than a list price.
Quotas versus a shared pool. Hard per-team quotas are predictable but strand capacity when a team is quiet. Fair-share scheduling, for example via Slurm priorities with preemption, lends idle quota to others.
What to do next
- Export last quarter's job accounting and compute GPU-hours per workload class.
- Add allocated, busy and stranded GPUs to the fleet dashboard. If you do not yet measure stranded GPUs, start there.
- Agree a target utilisation per class with each team and write it into the ledger.
- Run the placement simulation with your real job-size mix. If spread-style placement is costing you capacity, fix the scheduler before forecasting.
- Measure your node failure and repair rate and size spares per failure domain.
- Benchmark tokens per second per replica at your latency SLO and size inference for peak.
- Ask procurement for real order-to-acceptance lead times and encode the buy trigger.
- Rerun the ledger monthly and review the gap and trigger quarterly.