A committed use discount is a trade: you promise to pay for a quantity of compute for a fixed term, usually one or three years, and the provider charges a lower rate for it. For CPU fleets the trade is routine. For LLM fleets it is the single largest financial decision the platform team makes, because GPU-hours dominate the bill, the discount is large, and the promise cannot be undone. Google states plainly that you are billed for a resource-based commitment "regardless of whether or not you use those resources" and that you "can't cancel a commitment after its purchase". AWS says the terms of a Savings Plan "can't be changed after purchase".
This article treats commitments as a portfolio: the instruments, the two health metrics, a first-principles sizing rule, laddered terms, and the GPU-specific risks. The companion article on reserved capacity covers capacity guarantees, Capacity Blocks and break-even; this one is about the discount side and how to manage it over years.
The instruments and what they bind
Every commitment answers three questions: what is committed (a resource such as a GPU type, an instance family, or a dollar amount per hour), where it applies (a zone, a region, or the whole billing account), and whether it also secures capacity. The answers trade discount depth against flexibility, and the right mix depends on how confident you are about each dimension.
On Google Cloud, resource-based commitments cover vCPUs, memory, GPUs, local SSD and sole-tenant nodes in a specific region. Google's overview lists hardware discounts of up to 70% for memory-optimized series and up to 55% for other series. They can be purchased with or without attached reservations; only the attached form also holds capacity. Compute flexible CUDs are spend-based: you commit to an hourly amount of eligible spend across the billing account, regardless of project or region. The catch for LLM teams is eligibility. Google's documentation says that for GPUs "you can purchase compute flexible commitments only for GPUs that are used with G2 and G4 machine series"; for other series only vCPUs, memory and local SSD qualify. Flexible commitments also do not support attached reservations. For any other GPU type, confirm resource-based eligibility for your type and region before planning around it.
On AWS, Savings Plans are hourly dollar commitments. Compute Savings Plans apply to EC2 usage regardless of instance family, size, region, OS or tenancy, with prices up to 66% off on-demand. EC2 Instance Savings Plans bind you to one family in one region and go up to 72% off. SageMaker AI Savings Plans cover SageMaker instance usage, up to 64% off. AWS also notes that Savings Plans do not provide capacity reservations; capacity is bought separately, for example through On-Demand Capacity Reservations. Check the eligibility list for the specific accelerated instance types you run rather than assuming coverage.
Azure has reservations and a compute savings plan along similar lines; read its current exchange rules before relying on them. Everywhere, headline percentages are maximums; use the published rate for your GPU SKU.
How the discount is matched each hour
You cannot reason about utilization until you know how the discount is matched to usage. On AWS, Savings Plans are applied hour by hour to the usage with the highest savings percentage first, then the next highest, until the hourly commitment is used up; anything left over is billed at on-demand, and any unused commitment in that hour is simply lost. Two consequences follow. First, an hour of idle commitment cannot be carried forward to a busy hour. Second, a Compute Savings Plan bought to cover GPU instances may quietly be consumed by other eligible usage with a higher discount percentage in the same hour, so the GPU line can show less coverage than you expected while the plan itself reports full utilization.
Google's flexible commitments work similarly in dollars, with overage billed at the normal rate; resource-based commitments bill whether or not anything is running.
The practical rule: model the matching per hour, not per month. An average of 100 GPUs against a 100-GPU commitment looks perfect, yet if usage sits at 60 for half the day and 140 for the other half, a fifth of the commitment is wasted.
Utilization and coverage
Two numbers describe a commitment portfolio, and they pull in opposite directions. Utilization is the fraction of committed hours that were matched to real usage: committed-and-used divided by committed. Coverage is the fraction of eligible usage that was billed at a committed rate: committed-and-used divided by total eligible usage. Buying more raises coverage and lowers utilization. Neither alone is a target. 100% utilization usually means you are under-committed and paying on-demand for a large baseline; 100% coverage usually means you are paying for idle commitment at night.
Track both per hour, per instrument and per GPU type, plus a third number that ties them to money: effective savings rate, the total bill against what the same usage would have cost entirely on-demand. That is the figure finance cares about and the one the sizing rule below maximizes.
Sizing from first principles: the quantile rule
Take one GPU type in one region, an on-demand price P per GPU-hour, and a discount d. If you commit to c GPUs, every hour costs c x P x (1 - d) for the commitment plus P for each GPU above c. Now ask what the next committed GPU is worth. It costs P x (1 - d) every hour. It saves P only in the hours when usage exceeds c. So it pays for itself exactly when the fraction of hours with usage above c is greater than 1 - d. The optimum is where those are equal: commit to the level that usage exceeds in a fraction 1 - d of hours, which is the d-th quantile of hourly usage. With a 40% discount, commit to the 40th percentile of hourly demand, not the peak and not the minimum.
Committing to the minimum leaves money on the table when d is large; committing to the average overshoots when demand is spiky. A 60% discount justifies the 60th percentile.
Worked example: an inference fleet
The script below builds four weeks of synthetic hourly demand for an inference fleet: a floor of 96 GPUs, a daily swing up to 224, lighter weekends. Prices are illustrative, not quotes: $10 per GPU-hour on-demand, 40% off with a commitment.
import math
PRICE, DISCOUNT = 10.0, 0.40 # illustrative, not a quote
usage = [] # 4 weeks of hourly GPU demand
for h in range(24 * 28):
hod, dow = h % 24, (h // 24) % 7
day = 64 * (1 + math.sin((hod - 8) / 24 * 2 * math.pi))
usage.append(round(96 + day * (0.6 if dow >= 5 else 1.0)))
def cost(c, u=usage):
return sum(c * PRICE * (1 - DISCOUNT) + max(0, x - c) * PRICE for x in u)
def best_commit(u, d): # the d-th quantile of hourly usage
s = sorted(u)
return s[min(len(s) - 1, int(d * len(s)))]
od = cost(0)
for c in (96, 120, best_commit(usage, DISCOUNT), 160, 200, 224):
used = sum(min(x, c) for x in usage)
print(c, f"saving {1 - cost(c) / od:.1%}",
f"utilization {used / (c * len(usage)):.1%}",
f"coverage {used / sum(usage):.1%}")| Commit (GPUs) | 4-week bill | Saving vs on-demand | Utilization | Coverage |
|---|---|---|---|---|
| 0 | $1,025,840 | 0.0% | - | 0.0% |
| 96 (the floor) | $767,792 | 25.2% | 100.0% | 62.9% |
| 120 | $736,400 | 28.2% | 95.9% | 75.4% |
| 134 (40th percentile) | $732,448 | 28.6% | 92.6% | 81.3% |
| 160 | $746,720 | 27.2% | 86.0% | 90.1% |
| 200 | $828,000 | 19.3% | 74.7% | 97.9% |
| 224 (the peak) | $903,168 | 12.0% | 68.1% | 100.0% |
The quantile rule picks 134 GPUs, and it is the cheapest row. Notice how flat the curve is near the optimum: anything from 120 to 160 is within about 1.5 points. In the tails the penalties differ: below the floor each missing GPU forgoes P x d per hour, above the peak each extra one wastes P x (1 - d), so at 40% off overshooting costs 1.5 times as much. The real asymmetry is time: demand can fall during a term you cannot cancel, as the next paragraph shows.
Now suppose that six months in, a quantization and batching project cuts demand by 30%. Keeping the 134-GPU commitment on the smaller fleet saves 21.5% against on-demand; a 96-GPU commitment, which is the new 40th percentile, would have saved 28.6%. The efficiency win is real but the commitment gives up a quarter of the achievable saving. LLM fleets get efficiency wins like this regularly, which is why the forecast you size against should be the demand you expect after planned optimizations land, not today's.
Ladders and layers
Buying the whole commitment on one day means one forecast decides one to three years of spend, and the whole position comes up for renewal at once. A ladder splits the target into tranches bought at intervals. With four one-year tranches staggered by a quarter, a quarter of the portfolio is reconsidered every three months, each purchase uses the freshest data, and a demand drop is absorbed by not renewing the next tranche rather than by carrying idle commitment for a year.
Layer instruments by confidence. The floor you are nearly certain of for three years, such as a production model that will be served on some GPU regardless, can take the deepest, least flexible instrument. The band above it goes on one-year terms or flexible spend-based plans. Everything above the quantile stays on-demand or on spot. Write the policy down: for each layer, the instrument, the term, and the quantile it targets.
Risks specific to GPUs
Three risks are much larger for GPUs than for general compute.
- Generation turnover. A resource-based commitment names a GPU type. New accelerator generations arrive on a cadence shorter than a three-year term and often give better price-performance per token. If you migrate, the old commitment keeps billing. Before a three-year resource commitment, write down what you will run on that hardware in year three.
- Efficiency drift. Quantization, speculative decoding, better batching and smaller distilled models routinely cut GPU-hours per token. The worked example shows how a 30% win becomes a 7-point drop in savings under a fixed commitment. Size against post-roadmap demand.
- Training and inference move independently. A training run can end early or be cancelled; inference demand can spike with a launch. Keep separate commitment layers for the two, so a finished training program does not leave inference paying for its idle hours.
Amortizing commitments internally
Commitments distort internal cost reporting unless you decide how to spread them. The common pattern is to amortize: compute an effective blended rate per GPU-hour (commitment fees plus on-demand overage, divided by GPU-hours used) and charge every team that rate. Idle commitment then shows up as a platform cost rather than landing on whichever team happened to run in the busy hour. Report the unused commitment as its own line; it is the cleanest signal that the portfolio is oversized. Cost allocation by team and feature is covered in LLM cost attribution, and the wider operating model in LLM FinOps.
Operating the portfolio
Run the portfolio as a monthly review with a fixed agenda, fed by a job that pulls hourly billing exports per SKU and region:
- Utilization and coverage per instrument for the last month, hourly, with the worst day shown.
- Forecast for the next four quarters by GPU type, adjusted for the efficiency roadmap and any planned generation migration.
- Recomputed optimal commitment per layer, using the quantile rule on the forecast.
- Tranches expiring in the next quarter, with a renew, resize or drop decision for each.
Failure modes
- Sizing to the monthly average. Hourly matching wastes commitment in every trough. Use hourly data.
- Assuming the discount reserves GPUs. Savings Plans and flexible CUDs do not hold capacity. If you need guaranteed GPUs, buy a reservation as well.
- Region migration under a regional commitment. Moving serving to another region for latency or residency leaves a resource-based commitment idle in the old one.
- Renewal by default. Auto-renewing an expiring tranche at the same size skips the one moment when you could correct an oversized position.
Trade-offs
| Choice | Gains | Costs |
|---|---|---|
| 3-year over 1-year | Deeper discount | Locks in a GPU generation and a demand forecast for longer |
| Resource-based over spend-based | Deeper discount; can attach capacity | Idle if you change GPU type or region |
| Commit at the d-th quantile | Maximum expected saving | Coverage stays well below 100%; finance may read that as waste |
| Laddered tranches | Small, frequent, fresh decisions | More purchases to manage; slightly less discount if mixing shorter terms |
| Leave the top band on-demand or spot | No idle commitment | Higher marginal price at peak |
What to do next
- Export a month of hourly GPU usage per GPU type and region from your billing data.
- For each, compute the d-th quantile using the actual discount for that SKU and term, and compare it with what you have committed today.
- Add utilization, coverage and effective savings rate to your cost dashboard, hourly.
- Write down the efficiency roadmap for the next year and recompute the quantile against the demand you expect after it lands.
- Split any new purchase into at least four staggered tranches and record the layer, instrument, term and target quantile for each.
- Confirm whether each instrument holds capacity; if a workload needs guaranteed GPUs, read GPU capacity planning and buy the reservation explicitly.