Most writing about reserved GPU capacity is about fleets: a serving cluster that runs all year, a demand trace, a break-even utilization, a discount to optimize. A large share of real GPU decisions look different. A team has one training run to do. It needs 32 eight-GPU nodes at once, for roughly three weeks, starting soon. The question is not what fraction of a year the capacity will be busy. It is whether to prepay for a fixed window of guaranteed capacity, how long that window should be when nobody knows exactly how long the run will take, and what happens if the run outlives the window.

This article models the run duration as a random variable, prices the two kinds of mistake (slack days and running out of window), shows why on-demand behaves differently for a gang of nodes, and ends with a placement table by workload. The fleet-level arithmetic is covered in LLM Reserved Capacity and the discount side in LLM Committed Use Discounts; this page assumes neither and focuses on the run.

Fleet decisions and run decisions

At fleet scale the argument is about utilization. At run scale utilization is close to 100 percent by construction: a training job keeps every GPU busy until the final checkpoint. What varies is two other things.

First, obtainability. A single on-demand GPU instance is usually available somewhere. Thirty-two of them, in one location, on one network fabric, at the same moment, often are not. The run cannot start on 20 nodes and wait politely for the other 12; a data-parallel job with a fixed world size needs all of them.

Second, duration uncertainty. A reserved window has a hard end. On AWS EC2 Capacity Blocks for ML, for example, the documentation states that blocks end at 11:30 UTC, that termination of running instances begins at 11:00 UTC on the final day, and that cancellations are not allowed. Extensions exist but are only possible if capacity is available, which is not guaranteed. So the window has to be chosen before the run starts, from an estimate, and both directions of error cost money.

QuestionFleet viewRun view
What decides itUtilization over a yearObtainability and duration uncertainty
Unit of purchaseNode-hours over a termA window of N nodes for D days
Main risk of reservingIdle reserved hoursSlack days at the end, or running out of window
Main risk of on-demandPaying full price for the baseNever assembling the gang, or losing it mid-run
Sizing ruleQuantile of hourly demandQuantile of run duration

The architecture

The pieces of a run-level capacity plan are few. A run plan states the work: tokens to process, model size, node count and a measured throughput. A duration model turns that into a distribution of run lengths. The window choice converts the distribution into a reservation length. A controller watches the end of the window and either extends, falls back to on-demand, or takes a final checkpoint and drains. The checkpoint store is the only component that must outlive every window, because it is what lets the run continue on whatever capacity comes next.

One training run: from a run plan to a reservation window and a fallbackRun plantokens, nodes, throughputDuration modelMonte Carlo of run daysWindow choicecost of slack vs overrunReserved blockprepaid, fixed end, gang-placedOn-demand gang requestall N nodes or nothing usefuloverflow / overrunEnd-of-window controllerextend, checkpoint, drainCheckpoint storeoutlives every windowfinal savelaunch failures and wait timeextension is not guaranteed
A run-level capacity plan. The reservation is sized from a duration distribution, not a point estimate, and the checkpoint store is what makes any fallback possible.

On-demand for a gang of nodes

Why does on-demand behave so differently for a gang? Suppose each launch attempt for one node in a given zone succeeds independently with probability a. The probability of getting all N in one attempt is a to the power N. With a = 0.98 and N = 32 that is about 0.52; with a = 0.95 it is about 0.19. Real failures are not independent, which makes this worse: when a zone is short of a GPU type, every request fails together for hours.

Teams respond by acquiring incrementally: launch what is available, keep it running, and retry for the rest. That is rational, but price it honestly. Every node held while waiting for the gang is billed at the full on-demand rate and does no useful work. If the gang takes three days to assemble and nodes arrive at a roughly steady rate, the team pays for about half the gang for three days, which is 1.5 node-days per node of the gang before the first training step. Scale that by 32 nodes and the waiting alone is comparable to a slack day or two on a reservation.

The same reasoning explains hoarding: never releasing on-demand GPUs between runs because they might not come back. That is an unofficial reservation at full price, sometimes right for a week, rarely for a month, and a signal that the base load should be reserved properly.

Placement matters too. AWS documents that Capacity Block instances are placed close together inside an EC2 UltraCluster; an on-demand gang assembled over days may not be, so run a short all-reduce benchmark on the actual gang before trusting planned throughput.

Run length is a distribution

A run length is not a number; it is a distribution with three main sources of spread. Throughput varies by a few percent between gangs and software versions. Hardware and software failures stop the job, and each failure costs restart time plus the work done since the last checkpoint. Overheads such as checkpoint writes and evaluation passes add a roughly fixed percentage. The sketch below draws run lengths by simulation; replace each parameter with your own measurements from a short pilot run.

import numpy as np

rng = np.random.default_rng(7)

def sample_run_days(n, tokens=1.2e12, gpus=256, nodes=32,
                    tps_mean=3000, tps_sd=150,            # tokens/s per GPU, from a pilot
                    fail_per_node_day=0.02,               # interruptions per node per day
                    hours_lost_per_fail=1.5,              # restart + recompute since checkpoint
                    ckpt_overhead=0.03):                  # checkpoint and eval time fraction
    tps = rng.normal(tps_mean, tps_sd, n)
    compute_h = tokens / (tps * gpus) / 3600 * (1 + ckpt_overhead)
    fails = rng.poisson(fail_per_node_day * nodes * compute_h / 24)
    return (compute_h + fails * hours_lost_per_fail) / 24

days = sample_run_days(100_000)
print(np.percentile(days, [50, 90, 95, 99]))

With these inputs the run needs a median of about 19.4 days, the 90th percentile is about 20.7 days, the 95th about 21.1 and the 99th about 22.0. Notice how the failure term shifts the whole distribution: 32 nodes at 0.02 interruptions per node per day is roughly 12 interruptions over the run, about 18 hours of lost time. Halving the checkpoint interval halves the recompute part of that loss, which is why checkpoint cadence is a capacity decision and not only a reliability one. GPU Training Checkpointing covers how to make frequent checkpoints cheap.

Sizing the window

Now price the window. Let a reserved node-day cost b, and let a day of overrun cost o: the price of finishing the run some other way after the window ends. The marginal reasoning is the same as any newsvendor problem. Adding one more day to the window costs b for certain. It saves o only if the run would otherwise overrun that day, which happens with probability P(D greater than W). So keep extending the window while P(D greater than W) times o is greater than b. The optimal window is the run-length quantile at which the overrun probability equals b divided by o.

The value of o is where the real decision lives. If extensions or an on-demand gang are easy to get, o is roughly the on-demand rate and the ratio b / o is close to the reserved-to-on-demand price ratio, so you reserve near the median. If an overrun means the gang falls apart and the team waits a week for the next window, o includes the cost of idle researchers and a slipped milestone, the ratio is small, and you reserve deep into the tail.

def expected_cost(days, window_days, block_per_day, overrun_per_day):
    overrun = np.clip(days - window_days, 0, None)
    return window_days * block_per_day + (overrun * overrun_per_day).mean()

nodes = 32
block_per_day = 70 * 24 * nodes                       # illustrative: 70 units per node-hour
for overrun_per_day in (100 * 24 * nodes + 40_000,    # on-demand fallback works
                        400_000):                     # fallback fails, team stalls
    costs = {w: expected_cost(days, w, block_per_day, overrun_per_day) for w in range(18, 24)}
    best = min(costs, key=costs.get)
    print(overrun_per_day, best, round(costs[best]))

Worked example: a 32-node run

Run the numbers for the 32-node job. In the illustrative units used here, on-demand costs 100 per node-hour and the reservation 70, so a reserved day for the gang costs 53,760. Two scenarios bound the overrun cost.

Fallback works. An overrun day costs the on-demand gang (76,800) plus 40,000 for the disruption of moving the job, 116,800 in all. The ratio b / o is about 0.46. Expected total cost is about 1.137 million for an 18-day window, 1.096 million for 19 days, 1.097 million for 20 days and 1.133 million for 21 days. The optimum is flat across 19 and 20 days, and the run overruns a 20-day window about 27 percent of the time.

Fallback fails. If an overrun day costs 400,000 because the team stalls, the expected cost is 1.276 million for 19 days, 1.148 million for 20 and 1.143 million for 21, so the window moves out to 21 days and the overrun probability drops to about 6 percent.

For comparison, running the whole job on on-demand capacity at the same expected duration costs about 1.49 million, before counting days spent assembling the gang, so the reservation wins by roughly a quarter. The window moves by only a day or two between the scenarios; ignoring the overrun cost is the expensive mistake.

Two adjustments follow. AWS termination starts half an hour before the block ends, so subtract that plus your final checkpoint and drain time from the window. And because extensions can be requested well before the end, you can decide on one once the remaining duration is far less uncertain; since availability is not guaranteed, use that to trim the tail, not replace it.

Placing each workload

Most teams run more than one kind of GPU work. Each kind has a different tolerance for interruption and a different demand shape, and each belongs on different capacity.

WorkloadShapeInterruption toleranceBest fit
Large pretraining or long fine-tuneFixed gang, weeksLowTime-boxed reservation sized from the duration quantile
Production inference base loadContinuousNoneLong-term reservation or commitment for the base
Inference peaksDaily or weekly swingLowOn-demand above the reserved base
Hyperparameter sweepsMany small jobsHigh with checkpointsSpot, or idle reserved capacity
Evaluation and batch inferenceBursty, deadline in hoursHighSpot or idle reserved capacity, on-demand near the deadline
Interactive developmentBusiness hoursMediumOn-demand single nodes with idle shutdown

The common pattern is to reserve for the run and the base load, fill idle reserved hours with preemptible work through a queue such as Kueue, and send overflow to on-demand or spot capacity.

Operating a reserved window

Reservations fail operationally more often than financially.

  • Pre-flight the window. Before the block starts, stage the container image, dataset shards and the latest checkpoint in the same region, and test-launch one node into the reservation. On AWS an instance must target the reservation ID explicitly to land in a Capacity Block; a launch that misses the target runs on-demand next to an idle block.
  • Measure goodput daily. Track tokens processed per reserved GPU-hour, not only step time. A run that loses two hours a day to restarts burns the window faster than the plan assumed, and the duration model should be rerun with the observed numbers.
  • Decide extensions on a schedule. Re-forecast the remaining duration at fixed points, such as 50 and 75 percent of the window, and request an extension when the updated overrun probability crosses the ratio from the window calculation. On AWS an extension stays payment-pending until paid and fails if payment does not clear 35 minutes before the block ends, so do not leave it to the last hour.
  • Own the end. Schedule a final checkpoint well before termination starts, and make the trainer handle a termination signal by saving and exiting cleanly.

Failure modes

  • The median trap. The window is sized from the planned duration with no allowance for failures. The run overruns, no extension is available, and the job is killed mid-epoch.
  • The pilot that lied. Throughput was measured on two nodes and assumed to scale. At 32 nodes communication overhead lowers tokens per GPU and the run takes longer than every simulated draw.
  • Late start. Image pulls, data copies and NCCL debugging eat the first prepaid day.
  • Lost final checkpoint. The last save started after termination began and was cut off, so the next window resumes from a checkpoint hours older.
  • Phantom gang. On-demand nodes are held for days while the team waits for the last few, and the waiting bill exceeds what a reservation would have cost.

Trade-offs

A reservation buys certainty and pays in rigidity: no cancellation, a fixed end, and a location your data must move to. On-demand buys flexibility and pays in price and, for gangs, in the risk of never starting. Extensions and fallback are options on the tail that lose value exactly when capacity is tight. For a single run of known size, a reservation sized to a duration quantile usually wins; for small, short or interruptible work, on-demand and spot win; for anything that repeats month after month, move to the fleet view and the commitment math.

What to do next

  1. Run a pilot on at least a quarter of the target node count and record tokens per GPU per second, interruptions per node-day and restart time.
  2. Put those numbers into the duration simulation and read off the 50th, 90th and 95th percentiles.
  3. Write down the cost of an overrun day for this run, including people and deadlines, and compute the ratio of reserved day cost to overrun cost.
  4. Choose the window at the matching duration quantile, after subtracting termination lead time and drain time.
  5. Stage images, data and checkpoints before the window opens and test-launch one node into the reservation.
  6. Re-forecast at half and three quarters of the window and extend early if the overrun probability crosses your ratio.
Key takeaway: For a finite training run the choice is not about utilization but about getting the whole gang and about how long the run will really take. Model the duration as a distribution, price an overrun day honestly, reserve to the quantile where the overrun probability equals the cost ratio, and keep on-demand for the work that tolerates waiting.