Cost per token trained is the number that turns a training plan into a budget and lets you compare runs, clusters and model sizes on one axis. It sounds simple: dollars divided by tokens. In practice teams quote figures that differ by a factor of three for the same run, because they disagree about the numerator (only the final run, or everything it took to get there) and the denominator (tokens processed, or tokens that ended up in the released weights).
This article derives the marginal cost from first principles, works a full example, then builds the fully loaded figure that finance and planning actually need. It shows how to compute the realised number live from training logs and how it relates to inference cost per token. The FLOP counting itself, MFU and the GPU-hour ledger are covered in depth in the FLOPS budget for LLM training; here we use those results and focus on the unit economics.
The marginal cost from first principles
Training a dense transformer costs about 6N floating-point operations per token, where N is the parameter count: 2N for the forward pass and 4N for the backward pass. Attention adds a term that grows with sequence length, which matters at long context and is ignored below for clarity. A GPU delivers its peak FLOP/s multiplied by model FLOPs utilisation (MFU), the fraction of peak spent on useful model arithmetic, and by goodput, the fraction of wall-clock time spent making forward progress rather than restarting or waiting. So:
tokens_per_gpu_second = peak_flops * mfu * goodput / (6 * N)
cost_per_token = price_per_gpu_hour / (3600 * tokens_per_gpu_second)
= 6 * N * price_per_gpu_hour / (3600 * peak_flops * mfu * goodput)Three consequences follow directly. Cost per token is linear in model size: doubling N doubles every token's cost. It is inversely proportional to MFU, so an engineering improvement from 30 to 40 percent MFU is a 25 percent price cut. And the GPU price only matters in ratio to delivered FLOP/s: a GPU that costs 1.8 times as much per hour but delivers 2.2 times the useful FLOP/s is cheaper per token.
Worked example: 7B parameters, 2T tokens
Take a 7B-parameter model trained on 2 trillion tokens on H100 GPUs. Assume a dense BF16 peak of 989 TFLOP/s, 40 percent MFU, 90 percent goodput and an illustrative loaded price of $2.50 per GPU-hour.
| Quantity | Calculation | Value |
|---|---|---|
| FLOPs per token | 6 x 7e9 | 4.2e10 |
| Delivered FLOP/s per GPU | 989e12 x 0.40 x 0.90 | 3.56e14 |
| Tokens per GPU-second | 3.56e14 / 4.2e10 | about 8,480 |
| Tokens per GPU-hour | x 3600 | about 30.5 million |
| Cost per million tokens | $2.50 / 30.5 | about $0.082 |
| GPU-hours for 2T tokens | 2e12 / 30.5e6 | about 65,500 |
| Marginal cost of the run | 65,500 x $2.50 | about $164,000 |
That $164,000 is the number people quote. It is a real lower bound: it is what the final run costs if nothing else happens. It is not what the model cost.
The same arithmetic compares hardware. Suppose a newer GPU type is quoted at $4.50 per hour, 1.8 times the price, with 2.2 times the dense BF16 peak, but your first trial reaches only 35 percent MFU because kernels for it are less mature. Delivered FLOP/s per dollar is 2.2 x 0.35 / 1.8 = 0.43 against 0.40 / 1.0 = 0.40 for the incumbent, so it is about 7 percent cheaper per token today, and cheaper still once MFU recovers. Run that comparison with measured MFU from a short full-scale trial on each candidate, never with peak numbers alone, because the MFU gap between a mature and an immature software stack is often larger than the price gap between two GPU generations.
Notice how total cost scales. Total compute is 6ND, so at a fixed tokens-per-parameter ratio the cost grows with the square of model size. The common habit of training smaller models on many more tokens than a compute-optimal ratio suggests lowers the cost per token while raising the token count; teams do it because the smaller model is cheaper to serve for its whole life.
The fully loaded cost
The fully loaded figure divides everything spent to produce the model by the tokens in the final training run. The overheads that sit on top of the marginal run are:
- Experiments. Small-scale ablations, data-mix searches, hyperparameter sweeps and proxy runs used to choose the recipe. On a new architecture or data mix this can rival the final run.
- Lost work. Steps recomputed after hardware failures, rollbacks after loss spikes and abandoned runs. Goodput above covers only the final run's own restarts; abandoned runs are pure overhead. Checkpoint interval choices that drive this are in training checkpointing.
- Data. Acquisition and licensing, crawling, filtering, deduplication, classifier passes and tokenisation, which are often CPU and storage heavy.
- Evaluation. Periodic benchmark runs during training and the final evaluation suite.
- Capacity you hold but do not use. Reserved clusters idle between runs, and spare nodes kept for fast replacement.
- People and storage. The team, checkpoints and dataset copies.
A useful way to carry these is as multipliers on the marginal compute, with the assumptions written down. The values below are illustrative for a team training a new model family, not industry constants; replace them with your own records.
| Component | Assumed multiplier on final-run compute | Added cost |
|---|---|---|
| Final run (marginal) | 1.00 | $164,000 |
| Experiments and ablations | 0.60 | $98,400 |
| Abandoned runs and rollbacks | 0.15 | $24,600 |
| Evaluation compute | 0.05 | $8,200 |
| Idle reserved capacity | 0.20 | $32,800 |
| Data pipeline and licensing | fixed | $60,000 |
| Total | $388,000 |
With the same 2 trillion tokens as the denominator, the loaded cost is about $0.19 per million tokens, roughly 2.4 times the marginal figure. People cost would push it further. The gap is the point: a plan built on the marginal number runs out of money.
Which tokens to count
The denominator needs as much care as the numerator. Tokens processed counts every token the optimizer saw, including those recomputed after a rollback, so it flatters the figure. Tokens advanced counts only forward progress, the training step counter times tokens per step at the end of the run. Unique tokens counts distinct data, which matters when data is repeated over several epochs: four epochs over 500 billion unique tokens is 2 trillion advanced tokens but only 500 billion unique, and data cost should be spread over unique tokens.
Pick one definition per use and write it next to the number. For comparing hardware or engineering efficiency use marginal cost per advanced token. For budgeting and model pricing use loaded cost per advanced token. For data purchasing decisions use data cost per unique token.
A cost model as code
The model is short enough to keep in a repository and review like code, so that changing an assumption is a visible diff:
from dataclasses import dataclass, field
@dataclass
class TrainingCost:
params: float # N
tokens: float # advanced tokens in the final run
peak_flops: float # per GPU, dense, at the training precision
mfu: float
goodput: float
usd_per_gpu_hour: float # loaded hourly price
overhead: dict = field(default_factory=dict) # multipliers on final-run compute
fixed_usd: float = 0.0 # data, licensing, other non-compute spend
def gpu_hours(self):
tok_per_s = self.peak_flops * self.mfu * self.goodput / (6 * self.params)
return self.tokens / tok_per_s / 3600
def marginal_usd(self):
return self.gpu_hours() * self.usd_per_gpu_hour
def loaded_usd(self):
return self.marginal_usd() * (1 + sum(self.overhead.values())) + self.fixed_usd
def per_million(self):
return {"marginal": self.marginal_usd() / self.tokens * 1e6,
"loaded": self.loaded_usd() / self.tokens * 1e6}
run = TrainingCost(7e9, 2e12, 989e12, 0.40, 0.90, 2.50,
overhead={"experiments": 0.60, "lost": 0.15, "eval": 0.05, "idle": 0.20},
fixed_usd=60_000)
print(round(run.gpu_hours()), run.per_million())
for mfu in (0.30, 0.40, 0.50): # sensitivity: which input moves the answer most
r = TrainingCost(**{**run.__dict__, "mfu": mfu})
print(mfu, round(r.per_million()["loaded"], 3))The fully loaded GPU-hour that feeds usd_per_gpu_hour deserves its own model, built from capital, power and facility; see GPU infrastructure cost. If you train on preemptible capacity, the price is lower but goodput falls; GPU spot instances shows how to price the useful hour.
Tracking the realised cost during a run
The planned number is a forecast. The realised number comes from the run itself and should be on the same dashboard as loss. Every step log already has what you need: step number, tokens per step, wall-clock time and the number of GPUs allocated. Allocated, not busy: you pay for a GPU that waits on a straggler.
def realised_cost(steps, tokens_per_step, usd_per_gpu_hour):
# steps: list of dicts {"step", "t_wall", "gpus_allocated"} in log order, restarts included
gpu_seconds = 0.0
for prev, cur in zip(steps, steps[1:]):
gpu_seconds += (cur["t_wall"] - prev["t_wall"]) * prev["gpus_allocated"]
advanced = (max(s["step"] for s in steps) - steps[0]["step"]) * tokens_per_step
usd = gpu_seconds / 3600 * usd_per_gpu_hour
return {"usd": usd, "advanced_tokens": advanced, "usd_per_m": usd / advanced * 1e6}Because the wall-clock gaps include restarts and the step counter only counts forward progress, rollbacks automatically raise the realised cost. Alert when the realised figure drifts more than 10 percent above plan for a day: the usual causes are a slow node, a data-loader stall or a silent MFU regression after a code change.
Training tokens versus inference tokens
It is tempting to compare a training token with an inference token. A training token costs about 6N FLOPs; a generated inference token about 2N plus attention over the cache. But inference decode is usually limited by memory bandwidth, not arithmetic, so its cost per token depends on batch size and latency targets far more than on FLOPs; see LLM cost analysis for that side.
The comparison that does matter is amortisation. Divide loaded training cost by the number of tokens you expect to serve over the model's life, and add it to the serving cost per token. In the example, $388,000 spread over 500 billion served tokens adds about $0.78 per million served tokens; spread over 50 trillion it adds less than one cent. That is why heavily used models justify long training runs and rarely used ones do not.
Failure modes
- Quoting marginal as total. Budgets built on the final run alone overrun by the overhead multiplier.
- List price instead of loaded price. Hourly rates without power, networking, storage and support understate cost.
- Counting processed tokens. Recomputed tokens after rollbacks inflate the denominator.
- Ignoring attention at long context. 6N alone undercounts when sequence length is large.
- Peak at the wrong precision. Using a sparse or FP8 peak for a dense BF16 run makes MFU look worse and plans look better than reality.
- Busy GPUs instead of allocated GPUs. Idle time behind stragglers is still paid for.
Trade-offs
Lowering cost per token usually trades against something else. Larger batches raise MFU but can hurt sample efficiency, so cost per token falls while cost per unit of model quality may not. Preemptible capacity lowers price but lowers goodput and lengthens the calendar. Fewer ablations cut the loaded figure but raise the risk that the final run is the wrong one, which is the most expensive outcome of all. Cost per token is the right unit for efficiency and budgeting; it is the wrong objective to minimise on its own.
What to do next
- Write down your denominator definition and use advanced tokens for planning.
- Build the TrainingCost model with your loaded GPU-hour and review it like code.
- Measure MFU and goodput on a short run at full scale before committing the budget.
- Pull last quarter's experiment, eval and idle spend to set real overhead multipliers.
- Add realised cost per million tokens to the training dashboard with a drift alert.
- Amortise the loaded training cost over expected served tokens when pricing the model.