Everyone who trains a transformer stares at a loss curve, and most of the advice about it is about reading shapes: what a spike looks like, what overfitting looks like, what a learning rate that is too high looks like. That advice assumes the curve is measured correctly and that two curves on the same chart are comparable. Often neither is true. A loss averaged the wrong way across micro-batches, a step axis that hides a batch-size change, or two runs with different tokenizers plotted against each other can produce a confident conclusion that is simply false.
This article is about making the curve trustworthy: computing the number, choosing the axes, building the evaluation sets, comparing runs across tokenizers, smoothing without hiding problems, forecasting the final loss and turning the curve into automated alerts. For the gallery of curve shapes and what each one means, read diagnosing loss curves alongside this page.
What the number is
The training loss of a causal language model is the mean negative log-likelihood of the correct next token, in nats, over every target token in the optimizer step. If the model assigns probability 0.1 to the right token, that token contributes ln(1/0.1) = 2.30 nats. A freshly initialized model spreads probability roughly evenly over the vocabulary, so its loss starts near ln(V): about 10.4 for a 32,000-token vocabulary and about 11.8 for 128,000. A first logged loss far from that value usually means a bug in the labels or the logits, not an interesting model.
Two quantities get called the loss curve. The training loss is measured on the batch the optimizer just used: it is free, noisy and slightly optimistic once data repeats. The evaluation loss is measured on a fixed held-out set at intervals: it costs compute but it is the number you compare and forecast. Perplexity is just exp(loss), and it magnifies differences, so keep dashboards in nats and convert only for reporting.
Getting the average right
Batches contain different numbers of target tokens, because of padding, masked prompt tokens in fine-tuning, and variable-length documents. The correct loss for an optimizer step is the total loss over all target tokens divided by the total number of target tokens. The tempting shortcut, averaging the mean loss of each micro-batch, gives short sequences the same weight as long ones. This changes both the logged curve and the gradient.
This is not hypothetical. In October 2024 a widely shared report showed that the Hugging Face Trainer's loss did not match between runs with and without gradient accumulation, for exactly this reason. The fix was to pass the total number of target tokens across all micro-batches, called num_items_in_batch, into the loss so that it divides by the global count. Any hand-written loop can have the same bug, so here is the correct pattern for gradient accumulation under data parallelism.
import torch
import torch.distributed as dist
import torch.nn.functional as F
IGNORE = -100
def train_step(model, micro_batches, optimizer):
# Count every target token in the whole optimizer step first, across all ranks.
n_local = sum((mb["labels"][:, 1:] != IGNORE).sum() for mb in micro_batches) # count only predicted targets
n_global = n_local.clone()
dist.all_reduce(n_global) # total target tokens in the global batch
loss_sum_local = torch.zeros((), device=n_local.device)
for mb in micro_batches: # all but the last can run inside model.no_sync() to skip redundant all-reduces
logits = model(mb["input_ids"]).float()
tok_loss = F.cross_entropy(
logits[:, :-1].reshape(-1, logits.size(-1)),
mb["labels"][:, 1:].reshape(-1),
ignore_index=IGNORE, reduction="sum")
# Divide by the GLOBAL count; DDP averages gradients over R ranks, so scale by R.
(tok_loss * dist.get_world_size() / n_global).backward()
loss_sum_local += tok_loss.detach()
optimizer.step(); optimizer.zero_grad(set_to_none=True)
loss_sum = loss_sum_local.clone()
dist.all_reduce(loss_sum)
return (loss_sum / n_global).item(), n_global.item() # nats per token, tokens this stepThe extra world-size factor exists because distributed data parallel averages gradients across ranks. Each rank's contribution, divided by the global token count, must be multiplied back up by the rank count so that the average produces the global sum. A quick test catches mistakes here: run the same data with accumulation steps of 1 and 8 at the same global batch. The two losses should agree to within numerical noise. The gradient accumulation article covers the gradient side in more detail.
Which tokens count
- Padding must carry label -100 (or whatever your ignore index is), and the count must exclude it. A loss that falls when you increase padding is counting pad tokens.
- Prompt tokens in supervised fine-tuning are usually masked, so the loss measures response tokens only. Changing the masking policy changes the level of the curve, so record it with the run.
- Packed sequences join several documents into one row. Without a document-level attention mask, tokens attend across document boundaries, and the first tokens of each document get unrealistically easy context. The loss looks better than the model is.
- The shift: labels must be shifted one position relative to logits. Forgetting it produces a loss that collapses towards zero within a few hundred steps, because the model learns to copy its input.
Put tokens on the x-axis
Steps are a poor x-axis. A step at a global batch of 1 million tokens is four times a step at 256,000 tokens, so two runs with different batch sizes plotted against steps look as if the larger batch learns faster per step, which is true and uninformative. Batch-size warmup, sequence-length curricula and resumed runs with changed settings all make step counts incomparable even within one run. Log tokens seen, cumulatively, as the primary x-axis, and keep steps and wall-clock as secondary axes for debugging throughput.
Plot the learning rate on a shared x-axis below the loss. Many features of the curve come from the schedule rather than the model. A cosine schedule makes the loss fall faster near the end as the rate decays. A warmup-stable-decay schedule produces a visible drop during the final decay phase. Without the schedule in view, people attribute those changes to data or architecture.
Designing the evaluation sets
The evaluation loss is only as good as the set behind it. Fix the set at the start of a project and never change it in place; when you must change it, create a new version and report both for a while. Make it large enough that its noise is well below the differences you care about: a few million tokens is typical for pretraining-scale comparisons. Deduplicate it against the training data, including near duplicates, or the evaluation loss rewards memorization.
Report evaluation loss per domain, not only in aggregate: web text, code, maths, each language, and your target domain. Aggregate loss can improve while one domain gets worse, especially after a data-mixture change. Per-domain curves are also where contamination shows up, as a domain whose loss falls suspiciously fast.
Comparing runs with different tokenizers
Loss in nats per token cannot be compared across tokenizers, because tokens are different sizes. A tokenizer with longer tokens makes each prediction harder and the per-token loss higher, even if the model is better. Normalize by bytes of text instead. Bits per byte is the total loss in bits divided by the number of UTF-8 bytes in the evaluated text: bpb = (loss in nats per token) × (tokens) / (ln 2 × bytes).
A worked example. Run A uses a 32,000-token vocabulary, reaches an evaluation loss of 2.40 nats per token, and averages 4.1 bytes per token on the evaluation set. Its bits per byte is 2.40 / (0.693 × 4.1) = 0.845. Run B uses a 128,000-token vocabulary with 4.9 bytes per token, and reaches 2.85 nats per token. Its bits per byte is 2.85 / (0.693 × 4.9) = 0.839. Run B looked 0.45 nats worse on the usual chart and is in fact slightly better at modelling the same text. Compute bytes per token on the evaluation set itself, not on a different corpus.
Smoothing without hiding problems
Training loss is noisy, and a smoothed line is easier to read. An exponential moving average works well, but it needs bias correction for the first steps, or the curve starts at zero and climbs. Keep the raw values: smoothing exists for human eyes and for comparison, and alerts should look at raw values against the smoothed baseline.
class DebiasedEMA:
"""Exponential moving average with Adam-style bias correction for early steps."""
def __init__(self, beta=0.98):
self.beta, self.value, self.t = beta, 0.0, 0
def update(self, x):
self.t += 1
self.value = self.beta * self.value + (1 - self.beta) * x
return self.value / (1 - self.beta ** self.t)
def check_step(loss, grad_norm, ema, state, spike_ratio=1.5, patience=3):
"""Return an alert string or None. Feed the RAW loss; compare it against the EMA."""
if loss != loss or loss == float("inf"):
return "non-finite loss"
smoothed = ema.value / (1 - ema.beta ** ema.t) if ema.t else loss
if ema.t > 100 and loss > spike_ratio * smoothed:
state["spikes"] += 1
if state["spikes"] >= patience:
return f"loss {loss:.3f} above {spike_ratio}x EMA {smoothed:.3f} for {patience} steps"
else:
state["spikes"] = 0
ema.update(loss)
return NoneThe alert compares each raw loss with the debiased average and fires only after several consecutive spikes, because isolated single-step spikes are common and usually harmless. Pair it with a gradient-norm alert, since the norm often moves before the loss. The gradient clipping article explains how clipping and skip-step logic respond automatically.
Forecasting the final loss
During the constant-learning-rate part of training, evaluation loss as a function of tokens D is well approximated by a power law with a floor: L(D) = E + B·D^-β. Fitting that curve to the tail of a run lets you forecast where it will end and decide early whether a variant is worth finishing.
import numpy as np
from scipy.optimize import curve_fit
def power_law(d, E, B, beta):
return E + B * np.power(d, -beta)
# illustrative numbers: tokens (billions) and eval loss during a constant-LR phase
d = np.array([2, 4, 6, 8, 10, 12, 14, 16], dtype=float)
loss = np.array([3.05, 2.86, 2.77, 2.71, 2.67, 2.64, 2.615, 2.595])
params, cov = curve_fit(power_law, d, loss, p0=(2.0, 1.5, 0.5), maxfev=20000)
E, B, beta = params
print(f"E={E:.3f} B={B:.3f} beta={beta:.3f}")
print("forecast at 40B tokens:", power_law(40.0, *params))
print("param std devs:", np.sqrt(np.diag(cov))) # wide intervals = do not trust the forecastThree cautions. Fit only on a region with a constant or slowly changing learning rate, since a decaying schedule bends the curve in a way the power law does not model. Drop the first few percent of training, which is dominated by warmup. And read the parameter uncertainties: if E has a standard deviation comparable to the differences between your variants, the forecast cannot separate them, and you need more tokens before deciding.
Failure modes in measurement
- Train loss averaged per micro-batch rather than per token, so it shifts when batch composition changes.
- Evaluation run with dropout on, or with a different sequence length from training, which shifts the level of the curve.
- A resumed run that replays or skips data, producing a visible step in the curve at the resume point. Log the data-loader position with each checkpoint.
- Mixed precision computing the loss in bf16. Cast logits to float32 before cross-entropy, as the example does; see mixed precision training.
- Comparing runs at equal steps rather than equal tokens, or at equal tokens but different schedules, and attributing the difference to the change under test.
- An evaluation set silently regenerated with a new sampling seed, so old and new curves are measured on different data.
What to do next
- Check your step-zero loss against ln(vocabulary size).
- Run the same data with gradient accumulation of 1 and 8 at the same global batch and confirm the logged losses agree.
- Switch the primary x-axis to tokens seen, and plot the learning rate beneath the loss.
- Freeze and version a deduplicated evaluation set, split by domain.
- Add bits per byte to every evaluation report so runs with different tokenizers can be compared.
- Log raw and debiased EMA loss, and add NaN, spike and gradient-norm alerts on the raw values.
- Fit a power law to the constant-learning-rate tail before deciding whether to stop a variant, and check the parameter uncertainty.