A multi-year GPU commitment is a promise to pay for a fixed fleet of one generation for two to five years, usually take-or-pay, in exchange for a lower hourly rate and, often more importantly, the guarantee that the capacity exists at all. Sizing such a promise one year at a time is covered in committed use discounts, and what a reservation guarantees in reserved capacity. This article is about what changes when the term is long enough that the hardware market moves underneath you.

The central fact is simple. Your rate is fixed; the market price of the work your GPUs do is not. Each new generation does more training or inference per dollar, and supply of the older generation grows as buyers move on. A discount that looks like 33% on signing day can be a loss by year three. The rest of the article turns that into a model you can run, a sizing rule under uncertain demand, a list of contract clauses that change the answer, and a plan for keeping the fleet useful for its whole life. Every price here is illustrative; plug in your own quotes.

What a multi-year commitment buys

Long commitments come in a few shapes. A take-or-pay reservation charges for every GPU-hour in the term whether used or not. A spend commitment promises a dollar amount across a provider's services and discounts whatever you run. A prepayment pays part of the term up front for a further discount. Dedicated cluster deals from GPU cloud providers often add specifics the hyperscaler instruments do not: a named generation, a cluster size in one interconnect fabric, a data centre, a delivery date and a ramp schedule.

What you are really buying is three things bundled together: a price, a capacity guarantee and a topology. For training, the topology is often the scarce part: a few thousand GPUs on one non-blocking fabric cannot be assembled from on-demand instances scattered across regions. For inference, which scales out across many small replicas, the price matters more and the topology less. Knowing which of the three you need tells you how much term risk is worth taking.

Prepayment deserves its own arithmetic. Paying a year up front is a loan to the provider: if the extra discount for prepaying is smaller than your cost of capital over the same period, or than the value of keeping cash for a better deal next year, it is a bad trade even if the hourly number looks better. Prepayment also deepens the loss if the provider fails to deliver, so weigh it against counterparty risk, especially with younger GPU clouds.

The term against the generations

A 36-month commitment against the generations that ship during itmonth 0122436Committed fleet: one generation, fixed rate, take-or-payNext generation shipsGeneration after thatMarket price for the same work falls as better GPUs shipyour ratemarketYear 1: frontier trainingfull fabric, newest kernelsYear 2: fine-tuning, smaller runspartitioned jobsYear 3: inference, batch, evalfits in memory, older formatsThe commitment pays off only if the work it carries in year 3 is still worth more than the falling market rate
The rate is flat for the term while the market price of equivalent work falls with each generation; the fleet moves from frontier training to lighter work as it ages.

The diagram shows the shape of the problem. Early in the term your rate is below the market and the commitment is saving money. As newer GPUs ship, renting enough of them to do the same work gets cheaper, and the crossing point may arrive before the term ends. Meanwhile the work suited to your fleet changes: a cluster that trains the frontier model in year one is more likely serving, fine-tuning and evaluating in year three.

A present-value model in code

To compare fairly, price the alternative in units of work, not GPU-hours. If a newer GPU does the same job in half the time at 1.4 times the hourly price, the market price of that work has fallen by 30%. Call the annual fall in the price of your fleet's work d. The model below discounts monthly cash flows, charges the commitment for every committed GPU, buys overflow demand at market, and compares against buying all demand at market.

def commitment_npv(n_gpus, months, commit_rate, demand, market0, decline,
                   disc=0.08, hours=730):
    """Present cost of a take-or-pay commitment versus buying the same work at market.
    demand[t]: GPUs of this generation the workloads can use in month t.
    market0:   today's market price per GPU-hour for this generation's work.
    decline:   annual fall in that market price (newer GPUs, more supply)."""
    cost_commit = cost_market = 0.0
    for t in range(months):
        df = (1 + disc) ** (-t / 12)                   # discount factor
        market = market0 * (1 - decline) ** (t / 12)
        overflow = max(demand[t] - n_gpus, 0)          # beyond the commitment
        cost_commit += df * hours * (n_gpus * commit_rate + overflow * market)
        cost_market += df * hours * demand[t] * market
    return cost_commit, cost_market

Two simplifications are deliberate. Unused committed hours are worth nothing here; if your contract lets you resell them, add that value back. And the market alternative assumes you could actually get the capacity; if you could not, the commitment's value is the project it makes possible, which no price model captures.

Worked example: break-even and term length

Take 1,000 GPUs, demand flat at 1,000, a market price of 3.00 USD per GPU-hour for this generation's work today, an 8% discount rate and a 36-month commitment at 2.00. The commitment costs 47.1 million in present value. Buying the same work at market costs 70.6 million if prices never fall, 52.6 million if they fall 20% a year and 41.3 million at 35% a year. The 33% headline discount becomes a 10.5% saving at 20% decline and a 13.9% loss at 35%. Bisection on the model puts break-even at a 27.1% annual decline.

Longer terms usually come with deeper rates, so compare terms with their own rates. With illustrative rates of 2.60, 2.30, 2.00 and 1.85 for 12, 24, 36 and 48 months:

TermRate (USD/h)Saving at 10%/yr declineAt 20%/yrAt 30%/yr
12 months2.609.1%4.3%-1.3%
24 months2.3015.6%6.4%-4.7%
36 months2.0023.1%10.5%-4.9%
48 months1.8525.6%9.7%-10.2%

Read each column as one forecast of decline. At slow decline the longest term wins; at 20% the 36-month term edges out 48; at 30% every term loses and shorter loses least. Your decision is therefore mostly a forecast of d for the specific work you will run, and that forecast should come from measured throughput of your own workloads on newer hardware, not vendor peak numbers. The three-year TCO model covers the owned-hardware version of the same calculation.

Sizing under demand uncertainty

Demand is not flat. The model can be run over simulated demand paths to choose the committed size. With 2,000 paths starting at 1,000 GPUs, no drift, 5% monthly volatility and 20% annual decline, the present-value saving in millions of USD was:

Committed GPUsMean saving5th percentile95th percentile
6003.293.193.33
8003.730.204.44
9003.21-3.044.99
1,0001.58-6.555.55
1,200-4.95-16.404.86

Committing to the expected demand of 1,000 has a lower mean saving than 800 and a bad tail; 800 has the best mean and a positive 5th percentile. This is the multi-year version of a rule worth memorizing: commit to the floor of demand you are confident in, not the forecast, and buy the rest flexibly. The capacity planning article shows how to build the demand ledger that produces that floor, and spot capacity is one way to fill the gap above it.

The simulation treats demand and price decline as independent, and in reality they are not. The arrival of a new generation is exactly when your own teams want to move their largest jobs off the old fleet, so demand for committed capacity tends to fall at the same moment its market value falls. Stress the model with a scenario where demand drops by a third in the month a new generation becomes available; if the 5th percentile is still acceptable, the size is robust.

Contract clauses that move the numbers

Terms that look like legal detail often move the economics more than the rate does. Negotiate them with the model open.

ClauseWhat it protectsWhat to ask for
Delivery and ramp schedulePaying for GPUs that are not yet usableBilling starts at acceptance, credits for late delivery
Acceptance testingAccepting a cluster that cannot run your jobsBurn-in and collective-bandwidth tests you define
Availability SLANodes down for weeksPer-node replacement time, credits, spare pool
Topology guaranteeFabric split or oversubscribedNamed fabric, bandwidth floor, placement
Generation swapBeing locked to aging hardwareRight to move to a newer generation at a set ratio
Resale or assignmentPaying for idle hoursRight to sublet capacity or assign the contract
Step-down or exitDemand collapseReduce committed size after year one for a fee
Price escalatorsPower and cost pass-throughCaps on escalation, defined indices

Swap and resale rights directly reduce the generation risk in the model: a swap right flattens the market curve you are competing against, and a resale right gives unused hours a value. Put a number on each and compare it with the rate concession the provider asks for in return.

Keeping the fleet useful for the whole term

A long commitment is profitable only if the fleet stays busy through its last year, and that is a software problem. Plan three workload lanes and move work between them as the fleet ages.

Frontier training needs the full fabric, the newest low-precision formats and the most memory. It leaves first when a better generation arrives, because a run that takes weeks on new hardware is worth moving. Fine-tuning and smaller pretraining can run on partitions of the cluster and tolerate older kernels. Inference, batch scoring and evaluation need the model to fit in memory and throughput per dollar to stay competitive, and they can absorb whatever capacity is left.

Hardware features decide which lane a generation can serve. FP8 tensor-core support arrived with Hopper and Ada, and Blackwell added FP4; an older fleet runs models trained in newer formats at a higher precision, with more memory and less throughput. Per-GPU memory limits which models fit on one node for serving. Keep the software stack portable across formats and parallelism layouts from day one, so that moving a workload to the older fleet is a configuration change rather than a port.

Operationally, run acceptance as code: hardware diagnostics such as NVIDIA DCGM's diagnostic runs, collective bandwidth tests such as the all-reduce benchmarks in nccl-tests, and a short real training job, with results recorded per node. Re-run the same suite after every hardware replacement and before every renewal discussion.

Failure modes

  • Comparing GPU-hours instead of work. The commitment looks cheap against the old generation's on-demand price while newer GPUs do the job for less.
  • Committing to the forecast. Demand lands below plan and idle take-or-pay hours erase the discount.
  • Billing before acceptance. Months of payments for a cluster still failing burn-in or delivered in pieces.
  • No lane for year three. The training team moves to new hardware and nobody has an inference or batch stack ready for the old fleet.
  • Software tied to one generation. Kernels and checkpoints that assume a format or memory size the older fleet lacks make migration a rewrite.
  • Ignoring the topology clause. A cluster split across fabrics cannot run the large synchronous jobs it was bought for.
  • Treating the deal as finance-only. Engineering never sees the clauses, so swap and resale rights go unused.

Trade-offs

Longer terms buy lower rates and guaranteed capacity at the price of generation risk and demand risk. Larger commitments buy topology and leverage at the price of a fatter tail of idle hours. Swap, resale and step-down rights cost rate but cap the downside. A mixed portfolio, a committed floor on a long term, a middle layer on one-year terms and the peak on flexible capacity, gives up a little expected saving for a much better worst case, which is usually the right trade for anyone whose demand forecast is younger than the contract.

What to do next

  1. Measure throughput per dollar of your real workloads on your current and the newest available generation.
  2. Estimate the annual decline in the price of that work, with a low and a high case.
  3. Run the commitment model across terms and rates you are quoted and find the break-even decline.
  4. Simulate demand paths and commit to the size whose 5th-percentile saving is still acceptable.
  5. Price swap, resale, step-down and acceptance clauses and negotiate them alongside the rate.
  6. Write the acceptance suite before signing and make billing start on passing it.
  7. Plan the three workload lanes and keep the software stack portable across formats.
  8. Review the model against actual market prices every quarter and before any renewal.
Key takeaway: A multi-year GPU commitment is a bet that your locked rate stays below the falling market price of the same work: model it in units of work, commit to a demand floor, buy swap and resale rights, and plan the workloads that will keep the fleet busy in its last year.