A trained model is a capital asset with a short life. The GPU-hours that produced it are spent before it serves a single request, and they are only recovered through the tokens it serves before a better model replaces it. Model lifetime utility is the accounting that connects those two facts: how much useful work a model version does over the months it is in service, compared with every GPU-hour it consumed, from the training run to the last replica you switch off.
Teams that skip this accounting price a model by its serving cost alone, hiding a fixed cost that can exceed all serving spend, or keep an old version alive long after its tokens cost more than a successor's. This article builds the ledger in GPU-hours, turns it into a small simulator, works a hypothetical example, and derives the rule for when to refresh, distil or retire a version.
Every figure here is hypothetical and chosen to make the arithmetic visible. Replace each with your own measurements before using the conclusions; the structure carries over, the numbers do not.
The lifetime ledger
Start by separating the GPU-hours a model version consumes into two kinds. Fixed hours are spent once per version regardless of how much it is used. Recurring hours scale with time in service and with traffic.
| Item | Kind | What drives it | How to measure it |
|---|---|---|---|
| Pre-training (or its share) | Fixed | Parameters times training tokens | Scheduler accounting for the run, including failed restarts |
| Post-training: SFT, preference tuning | Fixed | Data volume, number of iterations | Job GPU-hours tagged with the version |
| Evaluation and red-teaming before launch | Fixed | Suite size, number of candidates | Eval cluster hours per release candidate |
| Serving replicas | Recurring | Tokens served, utilisation, replica floor | Fleet GPU-hours per version per month |
| Upkeep: eval reruns, safety refresh, canaries | Recurring | Release cadence of the surrounding system | Hours per month tagged with the version |
| Migration away from the version | Fixed, at the end | Prompt rework, customer re-validation | Engineering time plus eval hours |
If the model is a fine-tune of a base you did not train, the base is not free: the provider has priced its training into licence or API terms, and if you trained the base yourself for several products, allocate a share of its hours to each. The ledger only works if every GPU-hour carries a model-version tag; the most common reason lifetime numbers are wrong is that serving clusters are shared and nobody attributed their hours to versions.
The basic identity is simple. Over a service life of T months, the lifetime cost per token is the fixed cost plus the sum of each month's recurring cost, divided by the sum of tokens served. The marginal cost per token in a given month is only that month's recurring cost divided by that month's tokens. The two numbers answer different questions: the lifetime figure tells you whether the version was worth building, and the marginal figure tells you whether it is worth keeping.
Demand over a service life
Tokens served do not arrive evenly. A typical version ramps up as clients migrate to it, holds a plateau while it is the default, and then decays as traffic moves to its successor, with a long tail of clients that pinned the version and have not migrated. A workable model for planning has three parameters: the ramp length, the plateau length and the half-life of the decay. Fit them from your own history: look at the previous two versions, plot monthly tokens per version, and read the half-life off the log-scale slope of the tail.
The tail interacts badly with how serving fleets are sized. You cannot run a fraction of a replica, and you usually keep a minimum number running for availability across zones, even at three in the morning on the quietest day. Call that the replica floor. While demand is high the floor does not matter, because traffic needs more GPUs than the floor anyway. Once the tail falls below what the floor can serve, you pay for the floor whatever the traffic, and cost per token rises month after month even though nothing about the model changed.
Utilisation is the second hidden variable. A replica that could produce a given number of tokens per GPU-hour at full batch rarely averages that: traffic is diurnal, you hold headroom for spikes, and batches are smaller off-peak. Measure the ratio of achieved to achievable throughput per version per month. In the example below it is set to 40 percent, which is a planning assumption, not a benchmark.
The ledger as code
The simulator below is deliberately small: a plan with the fixed hours, the price of a GPU-hour, measured throughput, target utilisation, the replica floor and monthly upkeep, and a demand function. It prints, for each month, the tokens served, the marginal cost per million tokens and the lifetime cost per million tokens to date.
from dataclasses import dataclass
@dataclass
class Plan:
fixed_gpu_hours: float # pre-training share + post-training + evals + red-team
price: float # fully loaded $ per GPU-hour
tok_per_gpu_hour: float # measured throughput at your batch sizes
target_util: float # average fraction of that throughput you achieve
floor_gpus: int # replicas you keep up regardless of demand
upkeep_gpu_hours: float # monthly eval reruns, safety refresh, canaries
hours_per_month: float = 730.0
def demand(month, peak=50e9, ramp=1, plateau_end=6, half_life=4.0):
"""Tokens served in a month: ramp, plateau, then exponential decay."""
if month <= ramp:
return peak * month / (ramp + 1)
if month <= plateau_end:
return peak
return peak * 0.5 ** ((month - plateau_end) / half_life)
def ledger(plan, months):
rows, cum_tokens, cum_cost = [], 0.0, plan.fixed_gpu_hours * plan.price
floor_hours = plan.floor_gpus * plan.hours_per_month
for m in range(1, months + 1):
tokens = demand(m)
needed = tokens / (plan.tok_per_gpu_hour * plan.target_util)
serve_hours = max(needed, floor_hours)
month_cost = (serve_hours + plan.upkeep_gpu_hours) * plan.price
cum_tokens += tokens
cum_cost += month_cost
rows.append(dict(month=m, tokens_b=tokens / 1e9,
marginal_per_m=month_cost / (tokens / 1e6),
lifetime_per_m=cum_cost / (cum_tokens / 1e6)))
return rows
plan = Plan(fixed_gpu_hours=230_000, price=2.50, tok_per_gpu_hour=7.2e6,
target_util=0.40, floor_gpus=4, upkeep_gpu_hours=600)
for r in ledger(plan, 24):
print(r)Three choices in this code matter more than they look. Serving hours are the larger of what demand needs and what the floor forces, which is where the tail cost comes from. Upkeep is charged every month the version is alive, because evaluation reruns and safety refreshes do not stop when traffic falls. And the price is a fully loaded GPU-hour, including power, networking and the idle capacity you reserve, not the headline rental rate.
Worked example
Take a hypothetical fine-tuned model whose fixed cost is 230,000 GPU-hours: its share of a base model plus post-training and evaluation. At 2.50 dollars per fully loaded GPU-hour that is 575,000 dollars before launch. Serving measures 7.2 million output tokens per GPU-hour at full batch (2,000 tokens per second); at 40 percent average utilisation the fleet produces 2.88 million tokens per GPU-hour. Demand ramps to 50 billion tokens a month, holds for five months and then halves every four months. The floor is four GPUs, and upkeep is 600 GPU-hours a month. Running the simulator gives:
| Month | Tokens served | Marginal $ per M tokens | Lifetime $ per M tokens to date |
|---|---|---|---|
| 1 | 25.0B | 0.93 | 23.93 |
| 3 | 50.0B | 0.90 | 5.50 |
| 6 | 50.0B | 0.90 | 2.99 |
| 9 | 29.7B | 0.92 | 2.41 |
| 12 | 17.7B | 0.95 | 2.20 |
| 16 | 8.8B | 1.04 | 2.08 |
| 18 | 6.2B | 1.41 | 2.06 |
| 20 | 4.4B | 1.99 | 2.06 |
| 21 | 3.7B | 2.37 | 2.06 |
| 24 | 2.2B | 3.98 | 2.08 |
Read the table in three ways. First, the fixed cost dominates early: after the first month the version has cost nearly 24 dollars per million tokens, and even at the end of the plateau, month 6, it stands at 2.99, more than three times the serving cost. A team that priced this model at its marginal serving cost of about 0.90 would believe it was cheap when, over its life, it costs more than twice that.
Second, the tail is where amortisation finishes. Between months 9 and 18 the lifetime figure falls from 2.41 to 2.06 even though traffic is shrinking, because each extra token still costs less at the margin than the running average. Switching the version off at month 9 to save fleet hours would have left it at 2.41.
Third, the floor eventually wins. From about month 17 the four-GPU floor serves more than demand needs, the marginal cost climbs, and around month 20 it crosses the lifetime average. After that every additional month makes the version as a whole more expensive per token: month 21 costs 2.37 per million against an average of 2.06, and by month 24 the average itself has started rising again.
The retirement rule: marginal against average
The crossing in the example is not a coincidence; it is the general rule for any average. The lifetime average cost falls as long as the marginal cost of the next month is below it and rises once the marginal cost exceeds it, so the average is minimised at the month where the two meet. That gives a retirement test you can compute every month from the ledger, with no forecast needed: when the marginal cost per token of keeping the version exceeds its lifetime average, the version has done its best work and further service only dilutes it.
Cost alone is not the whole decision, because the alternative is not zero. Three comparisons settle it:
- Retire and migrate. Move the tail to the current default. If the successor serves those tokens at a marginal cost below the old version's marginal cost, and the one-off migration cost (re-validating prompts, notifying pinned clients, rerunning their acceptance tests) is recovered within a few months of the difference, retire.
- Shrink the floor. If you cannot retire because contracts pin the version, attack the floor instead: pack the old version onto shared replicas with multi-model serving or adapter swapping, or serve it from a smaller quantised build whose evals you have rerun. Halving the floor to two GPUs in the example moves the crossing from about month 20 to about month 23.
- Refresh rather than replace. A light fine-tune on recent data costs a small fraction of the original fixed hours and can lift the plateau or slow the decay. Model it as a new fixed cost added mid-life and a new demand curve; it is worth doing when the extra tokens at the margin repay the extra hours before the successor arrives.
Distillation and quantisation change the other side of the ledger: they add a fixed cost but cut the recurring cost per token for the rest of the life. They pay best on the plateau, when the tokens are many; applied in the tail, they rarely recover their fixed hours. The distillation cost ledger walks through that one-time calculation in detail.
Using expected lifetime to size the next model
Expected lifetime tokens are the denominator of everything above, so they should shape what you build. A smaller model trained on more tokens costs more per parameter to train but less per token served; the more tokens you expect to serve, the further the optimum moves toward small and over-trained, as derived in inference-optimal training. Feed that derivation a lifetime estimate taken from the decay curves of your previous versions, not the plateau rate times an optimistic number of months. And if the expected life is short, fine-tuning an existing base spreads someone else's fixed cost over far more tokens than you will serve; the hardware side of owning the full stack is the three-year TCO model.
Telemetry and operations
The ledger needs five monthly series per version: tokens served, split by client so you can see who is pinned; serving GPU-hours attributed by version even on shared fleets; achieved against maximum throughput; upkeep hours; and the fixed hours recorded once in the model registry. With those, the simulator becomes a monthly report: plot marginal and lifetime cost for every live version and review those near the crossing when you plan the next release. LLM cost analysis shows how to get a defensible cost per million tokens from a GPU-hour.
Failure modes
- Averaging over the plateau only. Quoting cost per token from the best months hides both the launch amortisation and the tail. Always report the lifetime figure next to the current marginal one.
- Untagged shared fleets. When versions share replicas without attribution, the old version looks free and is never retired. Attribute by tokens served if nothing better exists.
- Forgetting the floor. Planning the tail as if GPU-hours scale smoothly with traffic understates tail cost; at low volume the floor is the cost.
- Sunk-cost retention. Keeping a version because it was expensive to build confuses the lifetime question with the marginal one. Past fixed hours do not change the decision to keep serving; only future marginal costs and alternatives do.
- Optimistic life estimates. Sizing the next model on a life of two years when the last three versions lasted under one inflates the expected token count and justifies a bigger fixed cost than the version will ever repay.
Trade-offs
Longer lives amortise fixed hours better but keep users on older quality and more versions to operate. Shorter lives ship improvements quickly but force each release to repay its fixed cost over fewer tokens. A few long-lived versions with light refreshes is often the middle path.
What to do next
- Tag every training, eval and serving GPU-hour with a model version, starting with the version you are about to ship.
- Record the fixed GPU-hours of each live version in the model registry.
- Fit ramp, plateau and half-life from the monthly token history of your last two versions.
- Run the simulator with your own throughput, utilisation, floor and price, and find the month where marginal cost crosses the lifetime average.
- For each live version past that month, price retirement, floor reduction and a light refresh, and pick the cheapest.
- Feed realistic lifetime token estimates into the sizing decision for the next model.