Reserved capacity is the part of GPU procurement that decides whether your job starts at all. For CPU workloads you can usually assume that asking a cloud for an instance gets you one. For the high-end accelerators that LLM training and serving use, that assumption fails often enough that teams build around it: a launch request for a block of eight-GPU nodes can come back with an insufficient-capacity error for hours or days. A reservation converts that uncertainty into a bill. You pay for the GPUs whether or not you use them, and in return they are there when you ask.
This article explains reserved capacity from the software side. It separates the two things people conflate, a capacity guarantee and a price discount, then shows the arithmetic that decides how many GPUs to reserve, the scheduler configuration that keeps a reserved pool busy, the controller that handles the moment a time-boxed reservation ends, and the failure modes that turn a reservation into idle spend. Provider terms change quickly, so the specifics quoted here are the ones published at the time of writing; check the current documentation before you sign anything.
Two promises: capacity and price
A capacity reservation holds physical machines for you in a specific location for a period. A billing commitment (savings plans, committed use discounts, reserved instance pricing) lowers the price you pay for usage but, on its own, does not guarantee that any machine is available. The two are sold separately, stack in some combinations, and fail in different ways. A commitment without a reservation is a discount on capacity you may not get. A reservation without a commitment is guaranteed capacity at full price.
| Mechanism | Guarantees capacity? | Discount? | Shape |
|---|---|---|---|
| AWS On-Demand Capacity Reservation | Yes, in one Availability Zone | No by itself; savings plans or regional reserved instances can apply | Open-ended until cancelled; billed at the on-demand rate whether instances run or not |
| AWS EC2 Capacity Blocks for ML | Yes, for a fixed window | Priced per block, paid up front | 1 to 14 days in 1-day steps, longer in 7-day steps up to 182 days; 1, 2, 4, 8, 16, 32 or 64 instances; bookable up to 8 weeks ahead |
| Google Cloud DWS calendar mode | Yes, from a chosen start date | Priced per reservation | Fixed-duration future reservations; launched with 7- and 14-day durations, bookable up to 8 weeks ahead |
| Google Cloud DWS flex-start | No, but the request waits in a queue | Discounted versus on-demand | Starts when capacity appears, runs for the requested duration |
| Savings plans, committed use discounts | No | Yes | 1 or 3 year spend or resource commitment |
| Spot / preemptible | No; reclaimable | Large | Any time, any duration until reclaimed |
Two practical consequences follow. First, a reservation is zonal or cluster-scoped, so your data, storage and network fabric must live where the reservation lives. Second, a reservation only pays off if your launches actually land in it. On AWS an On-Demand Capacity Reservation can be open, matching any instance with the same attributes, or targeted, matching only launches that name it; a targeted reservation with no launch template pointing at it sits idle while you pay on-demand prices next door. The spot alternative, with its own interruption mechanics, is covered in GPU spot instances.
The architecture
Procurement produces labelled node pools; the queue layer knows each team's entitlement and who may borrow idle capacity; workloads are split by how badly they suffer from interruption; and the controller exists because time-boxed reservations end on schedule whether or not your job has finished.
Break-even and sizing from first principles
Start with one GPU node. Let the on-demand price be p_od per hour and the effective reserved price p_res per hour, paid for every hour of the term. If the node is busy a fraction u of the time, reservation costs p_res per wall-clock hour while on-demand costs u * p_od. Reserving is cheaper when u > p_res / p_od. That ratio is the break-even utilization. A reservation priced at 60 percent of on-demand breaks even at 60 percent busy.
The ratio ignores the reason most teams reserve: on-demand may not be available. If a launch fails with probability q at the moment you need it, and a failed launch costs you L per hour in missed revenue, idle researchers or a slipped training milestone, the expected hourly cost of on-demand is u * (p_od + q * L). For a serving fleet where an unavailable GPU means rejected requests, that term dominates and reserving the base load is cheap insurance.
Now size a fleet. Take an hourly trace of how many nodes you needed over a representative month. Consider the n-th reserved node. It is used in exactly the hours where demand is at least n, so its utilization is the fraction of hours with demand at least n. Keep adding reserved nodes while that fraction exceeds the break-even ratio. The optimal reserved count is the demand level exceeded in a fraction p_res / p_od of hours; everything above it should be on-demand, flex or spot.
import numpy as np
def optimal_reserved(demand_nodes, p_res, p_od):
# demand_nodes: hourly node demand over a representative window
d = np.asarray(demand_nodes)
ratio = p_res / p_od
best = 0
for n in range(1, int(d.max()) + 1):
# utilization of the n-th reserved node
if (d >= n).mean() > ratio:
best = n
else:
break
return best
def fleet_cost(demand_nodes, n_res, p_res, p_od):
d = np.asarray(demand_nodes)
hours = len(d)
reserved = n_res * p_res * hours
overflow = np.clip(d - n_res, 0, None).sum() * p_od
return reserved + overflowThis is a newsvendor problem: the answer is only as good as the demand trace, so include committed growth.
Worked example: a serving fleet
A team serves a 70B-parameter model on eight-GPU nodes. Its hourly demand over a month is 10 nodes at night, 22 at the daily peak and 14 on average. The numbers below are illustrative, not quotes: assume on-demand at 100 units per node-hour and a reservation at 62 units per node-hour, so break-even is 62 percent.
Run the sizing rule. Nodes 1 to 10 are busy every hour, utilization 100 percent. Node 12 is busy in about 80 percent of hours, node 14 in about 55 percent, node 18 in about 25 percent. The last node above the 62 percent line is node 13, so the team reserves 13 nodes and serves the remaining peak with on-demand. Monthly cost with 730 hours: 13 reserved nodes cost 13 x 62 x 730, about 588,000 units; overflow above 13 nodes averages roughly 2.5 node-hours per hour, about 183,000 units on-demand. Total about 771,000, against 1,022,000 for all on-demand at 14 average nodes. Reserving all 22 nodes for peak would cost 996,000 units and buy nothing but idle hours.
Then apply the availability adjustment. If on-demand launches in this zone fail during peaks, the overflow hours are exactly the ones at risk. The team either moves the break-even line down and reserves a few more nodes, or keeps 13 and adds admission control so overflow degrades gracefully; see GPU admission control for the shedding side.
Filling the reserved pool
A reservation sized for the daily peak is idle at night by construction. The way to recover that money is to let lower-priority work borrow the idle reserved nodes and give them back the moment the owner needs them. On Kubernetes, Kueue expresses this with two object types and a cohort name: a ResourceFlavor that maps to the reserved node label, and ClusterQueues that hold each team's nominal quota and share a cohort so they can borrow unused quota. Jobs are submitted through a namespaced LocalQueue that points at a ClusterQueue.
apiVersion: kueue.x-k8s.io/v1beta1
kind: ResourceFlavor
metadata: {name: reserved-h100}
spec:
nodeLabels: {pool: reserved}
---
apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata: {name: serving}
spec:
cohort: reserved-gpus
namespaceSelector: {} # unset selects no namespaces
preemption:
reclaimWithinCohort: Any # take lent GPUs back immediately
resourceGroups:
- coveredResources: ["nvidia.com/gpu"]
flavors:
- name: reserved-h100
resources:
- {name: "nvidia.com/gpu", nominalQuota: 104} # 13 nodes x 8
---
apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata: {name: batch-fill}
spec:
cohort: reserved-gpus
namespaceSelector: {}
resourceGroups:
- coveredResources: ["nvidia.com/gpu"]
flavors:
- name: reserved-h100
resources:
- {name: "nvidia.com/gpu", nominalQuota: 0, borrowingLimit: 104}The batch queue owns nothing and borrows everything. When serving needs a node back, Kueue evicts borrowed workloads to reclaim quota. That only works if the borrowed work is genuinely interruptible: offline evaluation, embedding backfills, synthetic data generation, or training jobs that checkpoint every few minutes. Anything that loses hours of progress on eviction does not belong in the fill queue. For the broader picture of placement and queueing see GPU scheduling.
Scale serving on a signal that leads demand, so reclaim starts before latency degrades, and give batch pods a grace period long enough to flush a checkpoint but no longer.
When the reservation ends
Time-boxed reservations such as Capacity Blocks end at a fixed time and the provider reclaims the hardware for the next customer. AWS documents that it begins terminating instances in a Capacity Block 30 minutes before the block's end time (60 minutes for UltraServer types) and emits an EventBridge event 10 minutes before termination begins. Your usable window is therefore shorter than the window you paid for, and a training job that ignores this loses everything since its last checkpoint.
Treat the end time as a deadline your own controller enforces, with margin for the checkpoint itself:
import datetime as dt, time
TERMINATION_LEAD = dt.timedelta(minutes=30) # provider begins terminating here
CKPT_DURATION = dt.timedelta(minutes=12) # measured, p99, for this model size
SAFETY = dt.timedelta(minutes=10)
def watch(block_end, trainer):
deadline = block_end - TERMINATION_LEAD - CKPT_DURATION - SAFETY
while True:
now = dt.datetime.now(dt.timezone.utc)
if now >= deadline:
trainer.request_stop() # finish current step
trainer.save_checkpoint(final=True)
trainer.upload_and_verify() # checksum against object store
trainer.mark_resumable()
return
time.sleep(30)The checkpoint must land somewhere that outlives the block, usually object storage in the same region, and must be verified before the controller reports success. Decide in advance what happens next: renew with another block, which needs to be booked before the current one ends because blocks are sold ahead of time; fall back to a smaller on-demand or flex footprint with a reshaped parallelism plan; or pause. Elastic training frameworks that can resume on a different world size make the fallback cheap; fixed-topology jobs make it a manual rebuild.
Failure modes
- Launches miss the reservation. A targeted reservation with no launch template or node pool referencing it stays empty. Alert on reserved-but-unused instance count, not only on cost.
- Wrong zone for the data. Reservation in one zone, file system and data cache in another: training crawls and cross-zone traffic adds its own bill.
- Allocated is not busy. Kubernetes reports the GPUs as allocated, but the pods are waiting on input, stuck in a crash loop, or holding GPUs for a notebook nobody uses. Measure busy time from device metrics such as DCGM's SM activity, not from the scheduler's allocation view.
- Hardware failure inside the pool. A failed node reduces the pool until the provider replaces it. Whether and how fast replacement happens depends on the provider's terms; budget a spare node for large training jobs rather than assuming the full count for the whole term.
- Fill work that cannot be interrupted. A long uncheckpointed job borrows reserved nodes, serving reclaims them, and the job restarts from zero, over and over. It burns the idle hours the fill queue was meant to recover.
- The end date surprise. A block expires mid-run because nobody owned renewal. Put the end time in the job's metadata and in the on-call calendar.
- Committing to a generation. A long term outlives the GPU generation's lead in tokens per dollar; match term length to the hardware's useful life for the workload.
Operating reservations
Run reservations as a product with a dashboard, not a purchase order. The core metrics are reserved GPU-hours, allocated GPU-hours, busy GPU-hours from device telemetry, and the split of busy hours between owners and borrowers. The ratio of busy to reserved is the number that tells you whether the break-even assumption still holds. Review it monthly against the demand trace that justified the purchase, and resize at renewal rather than letting the original number roll over.
Write down a procurement rule: reserved fraction of base load, which workloads may use spot, block booking lead time, and who approves exceptions. Capacity planning for the whole stack, from FLOPs to facility, is covered in GPU infrastructure planning; the full list of cost levers, of which commitment mix is only one, is in GPU cost optimization.
Trade-offs
| Choice | You gain | You give up |
|---|---|---|
| Reserve base load only | Lowest cost per busy hour for steady demand | Exposure to on-demand shortages at peak |
| Reserve for peak | Peak is always served | Many idle hours unless you run a fill queue |
| Time-boxed blocks for training | Guaranteed start date, no long-term commitment | Hard stop, booking lead time, premium price |
| Long-term reservation | Predictable capacity and price | Locked to one generation, zone and term |
| Flex or queued capacity | Lower price, no idle cost | Uncertain start time |
What to do next
- Export an hourly trace of GPU node demand for the last month, split by serving, training and batch.
- Measure busy time from device telemetry and compare it with allocation; fix idle allocations before buying anything.
- Compute the break-even ratio from current quotes and run the sizing rule on the serving trace.
- Check that every reservation is referenced by a node pool or launch template, and alert on unused reserved instances.
- Set up a cohort with a preemptible fill queue and move one checkpointed batch workload into it.
- For every time-boxed block, write the end-time controller, measure checkpoint duration, and test resume on a smaller fallback footprint.
- Put renewal dates and the procurement rule in a document with an owner, and review utilization monthly.