Cost per token trained is the number that turns a training plan into a budget and lets you compare runs, clusters and model sizes on one axis. It sounds simple: dollars divided by tokens. In practice teams quote figures that differ by a factor of three for the same run, because they disagree about the numerator (only the final run, or everything it took to get there) and the denominator (tokens processed, or tokens that ended up in the released weights).

This article derives the marginal cost from first principles, works a full example, then builds the fully loaded figure that finance and planning actually need. It shows how to compute the realised number live from training logs and how it relates to inference cost per token. The FLOP counting itself, MFU and the GPU-hour ledger are covered in depth in the FLOPS budget for LLM training; here we use those results and focus on the unit economics.

The marginal cost from first principles

Training a dense transformer costs about 6N floating-point operations per token, where N is the parameter count: 2N for the forward pass and 4N for the backward pass. Attention adds a term that grows with sequence length, which matters at long context and is ignored below for clarity. A GPU delivers its peak FLOP/s multiplied by model FLOPs utilisation (MFU), the fraction of peak spent on useful model arithmetic, and by goodput, the fraction of wall-clock time spent making forward progress rather than restarting or waiting. So:

tokens_per_gpu_second = peak_flops * mfu * goodput / (6 * N)
cost_per_token        = price_per_gpu_hour / (3600 * tokens_per_gpu_second)
                      = 6 * N * price_per_gpu_hour / (3600 * peak_flops * mfu * goodput)

Three consequences follow directly. Cost per token is linear in model size: doubling N doubles every token's cost. It is inversely proportional to MFU, so an engineering improvement from 30 to 40 percent MFU is a 25 percent price cut. And the GPU price only matters in ratio to delivered FLOP/s: a GPU that costs 1.8 times as much per hour but delivers 2.2 times the useful FLOP/s is cheaper per token.

Worked example: 7B parameters, 2T tokens

Take a 7B-parameter model trained on 2 trillion tokens on H100 GPUs. Assume a dense BF16 peak of 989 TFLOP/s, 40 percent MFU, 90 percent goodput and an illustrative loaded price of $2.50 per GPU-hour.

QuantityCalculationValue
FLOPs per token6 x 7e94.2e10
Delivered FLOP/s per GPU989e12 x 0.40 x 0.903.56e14
Tokens per GPU-second3.56e14 / 4.2e10about 8,480
Tokens per GPU-hourx 3600about 30.5 million
Cost per million tokens$2.50 / 30.5about $0.082
GPU-hours for 2T tokens2e12 / 30.5e6about 65,500
Marginal cost of the run65,500 x $2.50about $164,000

That $164,000 is the number people quote. It is a real lower bound: it is what the final run costs if nothing else happens. It is not what the model cost.

The same arithmetic compares hardware. Suppose a newer GPU type is quoted at $4.50 per hour, 1.8 times the price, with 2.2 times the dense BF16 peak, but your first trial reaches only 35 percent MFU because kernels for it are less mature. Delivered FLOP/s per dollar is 2.2 x 0.35 / 1.8 = 0.43 against 0.40 / 1.0 = 0.40 for the incumbent, so it is about 7 percent cheaper per token today, and cheaper still once MFU recovers. Run that comparison with measured MFU from a short full-scale trial on each candidate, never with peak numbers alone, because the MFU gap between a mature and an immature software stack is often larger than the price gap between two GPU generations.

Notice how total cost scales. Total compute is 6ND, so at a fixed tokens-per-parameter ratio the cost grows with the square of model size. The common habit of training smaller models on many more tokens than a compute-optimal ratio suggests lowers the cost per token while raising the token count; teams do it because the smaller model is cheaper to serve for its whole life.

The fully loaded cost

The fully loaded figure divides everything spent to produce the model by the tokens in the final training run. The overheads that sit on top of the marginal run are:

  • Experiments. Small-scale ablations, data-mix searches, hyperparameter sweeps and proxy runs used to choose the recipe. On a new architecture or data mix this can rival the final run.
  • Lost work. Steps recomputed after hardware failures, rollbacks after loss spikes and abandoned runs. Goodput above covers only the final run's own restarts; abandoned runs are pure overhead. Checkpoint interval choices that drive this are in training checkpointing.
  • Data. Acquisition and licensing, crawling, filtering, deduplication, classifier passes and tokenisation, which are often CPU and storage heavy.
  • Evaluation. Periodic benchmark runs during training and the final evaluation suite.
  • Capacity you hold but do not use. Reserved clusters idle between runs, and spare nodes kept for fast replacement.
  • People and storage. The team, checkpoints and dataset copies.

From a marginal compute cost to a fully loaded cost per useful tokenFLOPs per tokenabout 6N, plus attentionDelivered FLOP/s per GPUpeak x MFU x goodputPrice per GPU-hourloaded, not listMarginal cost per tokenfinal run onlyExperimentsablations, sweepsLost workrestarts, rollbacksDataacquire, filter, tokeniseEval, storage, peopleLoaded cost per useful tokenprogram spend / tokens in final model
The marginal figure uses only the top row; the loaded figure adds every overhead and divides by the tokens that reached the final model.

A useful way to carry these is as multipliers on the marginal compute, with the assumptions written down. The values below are illustrative for a team training a new model family, not industry constants; replace them with your own records.

ComponentAssumed multiplier on final-run computeAdded cost
Final run (marginal)1.00$164,000
Experiments and ablations0.60$98,400
Abandoned runs and rollbacks0.15$24,600
Evaluation compute0.05$8,200
Idle reserved capacity0.20$32,800
Data pipeline and licensingfixed$60,000
Total$388,000

With the same 2 trillion tokens as the denominator, the loaded cost is about $0.19 per million tokens, roughly 2.4 times the marginal figure. People cost would push it further. The gap is the point: a plan built on the marginal number runs out of money.

Which tokens to count

The denominator needs as much care as the numerator. Tokens processed counts every token the optimizer saw, including those recomputed after a rollback, so it flatters the figure. Tokens advanced counts only forward progress, the training step counter times tokens per step at the end of the run. Unique tokens counts distinct data, which matters when data is repeated over several epochs: four epochs over 500 billion unique tokens is 2 trillion advanced tokens but only 500 billion unique, and data cost should be spread over unique tokens.

Pick one definition per use and write it next to the number. For comparing hardware or engineering efficiency use marginal cost per advanced token. For budgeting and model pricing use loaded cost per advanced token. For data purchasing decisions use data cost per unique token.

A cost model as code

The model is short enough to keep in a repository and review like code, so that changing an assumption is a visible diff:

from dataclasses import dataclass, field

@dataclass
class TrainingCost:
    params: float                 # N
    tokens: float                 # advanced tokens in the final run
    peak_flops: float             # per GPU, dense, at the training precision
    mfu: float
    goodput: float
    usd_per_gpu_hour: float       # loaded hourly price
    overhead: dict = field(default_factory=dict)    # multipliers on final-run compute
    fixed_usd: float = 0.0        # data, licensing, other non-compute spend

    def gpu_hours(self):
        tok_per_s = self.peak_flops * self.mfu * self.goodput / (6 * self.params)
        return self.tokens / tok_per_s / 3600

    def marginal_usd(self):
        return self.gpu_hours() * self.usd_per_gpu_hour

    def loaded_usd(self):
        return self.marginal_usd() * (1 + sum(self.overhead.values())) + self.fixed_usd

    def per_million(self):
        return {"marginal": self.marginal_usd() / self.tokens * 1e6,
                "loaded": self.loaded_usd() / self.tokens * 1e6}

run = TrainingCost(7e9, 2e12, 989e12, 0.40, 0.90, 2.50,
                   overhead={"experiments": 0.60, "lost": 0.15, "eval": 0.05, "idle": 0.20},
                   fixed_usd=60_000)
print(round(run.gpu_hours()), run.per_million())

for mfu in (0.30, 0.40, 0.50):            # sensitivity: which input moves the answer most
    r = TrainingCost(**{**run.__dict__, "mfu": mfu})
    print(mfu, round(r.per_million()["loaded"], 3))

The fully loaded GPU-hour that feeds usd_per_gpu_hour deserves its own model, built from capital, power and facility; see GPU infrastructure cost. If you train on preemptible capacity, the price is lower but goodput falls; GPU spot instances shows how to price the useful hour.

Tracking the realised cost during a run

The planned number is a forecast. The realised number comes from the run itself and should be on the same dashboard as loss. Every step log already has what you need: step number, tokens per step, wall-clock time and the number of GPUs allocated. Allocated, not busy: you pay for a GPU that waits on a straggler.

def realised_cost(steps, tokens_per_step, usd_per_gpu_hour):
    # steps: list of dicts {"step", "t_wall", "gpus_allocated"} in log order, restarts included
    gpu_seconds = 0.0
    for prev, cur in zip(steps, steps[1:]):
        gpu_seconds += (cur["t_wall"] - prev["t_wall"]) * prev["gpus_allocated"]
    advanced = (max(s["step"] for s in steps) - steps[0]["step"]) * tokens_per_step
    usd = gpu_seconds / 3600 * usd_per_gpu_hour
    return {"usd": usd, "advanced_tokens": advanced, "usd_per_m": usd / advanced * 1e6}

Because the wall-clock gaps include restarts and the step counter only counts forward progress, rollbacks automatically raise the realised cost. Alert when the realised figure drifts more than 10 percent above plan for a day: the usual causes are a slow node, a data-loader stall or a silent MFU regression after a code change.

Training tokens versus inference tokens

It is tempting to compare a training token with an inference token. A training token costs about 6N FLOPs; a generated inference token about 2N plus attention over the cache. But inference decode is usually limited by memory bandwidth, not arithmetic, so its cost per token depends on batch size and latency targets far more than on FLOPs; see LLM cost analysis for that side.

The comparison that does matter is amortisation. Divide loaded training cost by the number of tokens you expect to serve over the model's life, and add it to the serving cost per token. In the example, $388,000 spread over 500 billion served tokens adds about $0.78 per million served tokens; spread over 50 trillion it adds less than one cent. That is why heavily used models justify long training runs and rarely used ones do not.

Failure modes

  • Quoting marginal as total. Budgets built on the final run alone overrun by the overhead multiplier.
  • List price instead of loaded price. Hourly rates without power, networking, storage and support understate cost.
  • Counting processed tokens. Recomputed tokens after rollbacks inflate the denominator.
  • Ignoring attention at long context. 6N alone undercounts when sequence length is large.
  • Peak at the wrong precision. Using a sparse or FP8 peak for a dense BF16 run makes MFU look worse and plans look better than reality.
  • Busy GPUs instead of allocated GPUs. Idle time behind stragglers is still paid for.

Trade-offs

Lowering cost per token usually trades against something else. Larger batches raise MFU but can hurt sample efficiency, so cost per token falls while cost per unit of model quality may not. Preemptible capacity lowers price but lowers goodput and lengthens the calendar. Fewer ablations cut the loaded figure but raise the risk that the final run is the wrong one, which is the most expensive outcome of all. Cost per token is the right unit for efficiency and budgeting; it is the wrong objective to minimise on its own.

What to do next

  1. Write down your denominator definition and use advanced tokens for planning.
  2. Build the TrainingCost model with your loaded GPU-hour and review it like code.
  3. Measure MFU and goodput on a short run at full scale before committing the budget.
  4. Pull last quarter's experiment, eval and idle spend to set real overhead multipliers.
  5. Add realised cost per million tokens to the training dashboard with a drift alert.
  6. Amortise the loaded training cost over expected served tokens when pricing the model.
Key takeaway: Cost per token trained is 6N times the loaded GPU-hour price divided by delivered FLOP/s, and that marginal figure is only the floor. Carry experiments, lost work, data and idle capacity as explicit multipliers, count advanced tokens, track the realised number from step logs, and amortise the result over the tokens the model will serve.