A large language model training run lasts days or weeks on hundreds or thousands of GPUs, is interrupted by hardware failures many times, and produces a model whose quality is judged by small differences in loss and evaluation scores. Three questions follow from that. Can you resume after a failure and get the run you would have had without it? When a new run looks worse than an old one, is it a real regression or noise? And when a run diverges, can you find out why?

Those are reproducibility questions, but at the scale of a whole run rather than a single training step. The kernel-level causes of non-determinism, such as floating-point summation order, atomics and the PyTorch determinism switches, are covered in ML reproducibility on GPUs. This article builds on them and covers what an LLM team needs on top: a run manifest, checkpoints complete enough for exact resume, a data order that survives restarts and resharding, an understanding of why changing the parallel layout changes the numbers, statistical comparison across seeds, and a procedure for bisecting divergent runs and replaying loss spikes.

Four levels of guarantee

Be precise about the guarantee you want, because each level costs more than the one before. Statistical reproducibility means a rerun with a different seed lands within the known spread of final metrics. It is what research claims need and is always achievable. Resume exactness means that stopping at step N and resuming continues the same run as running straight through would have. It is achievable with discipline and is the most valuable property in operations, because it makes failures invisible in the results. It has two strengths. State-exact resume means every restored tensor, RNG stream and data cursor is bitwise what was saved, so no sample is repeated or skipped; it is cheap and should be the production standard. Bit-identical continuation additionally requires the resumed steps to compute exactly what the uninterrupted run computed, which needs deterministic kernels and an unchanged layout, because non-deterministic split-K GEMMs, atomics or collectives make two executions of the same step differ even from a perfect checkpoint. Bitwise rerun means two independent runs from step 0 match exactly. It requires deterministic kernels and identical hardware, costs throughput, and is mostly useful for debugging. Cross-layout equality, identical numbers on a different number of GPUs or a different parallel layout, is generally not achievable at reasonable cost, and you should not promise it.

Everything a run depends on

A run's result is a function of more inputs than its config file. Every one of them must be pinned or recorded.

InputHow it changes resultsHow to pin it
CodeAny change in model, loss or optimizerGit SHA; refuse to launch from a dirty tree
ConfigHyperparameters, schedulesCanonical serialised config and its hash
ContainerPyTorch, CUDA, cuBLAS, cuDNN, NCCL versions select different kernelsImage digest, not a tag
DataContent and order of tokensPer-shard content hashes, tokenizer hash, mixture weights
SeedsInit, dropout, data shufflingOne root seed, derived per purpose and rank
LayoutWorld size and TP, PP, DP degrees change reduction orderRecorded and checked on resume
HardwareGPU model and driver can change kernel choiceRecorded; alert on mixed pools

The run manifest

Write a manifest at step 0, store it with the checkpoints, and verify it on every resume. The point is not documentation but enforcement: a resume that silently picks up a different container or a re-tokenised dataset produces a run nobody can explain.

import hashlib, json, os, subprocess, torch

def sha(b: bytes) -> str:
    return hashlib.sha256(b).hexdigest()[:16]

def build_manifest(cfg: dict, shard_paths: list[str], tokenizer_path: str, layout: dict) -> dict:
    dirty = subprocess.run(["git", "status", "--porcelain"], capture_output=True, text=True).stdout
    if dirty.strip():
        raise RuntimeError("refusing to launch from a dirty working tree")
    return {
        "git_sha": subprocess.check_output(["git", "rev-parse", "HEAD"], text=True).strip(),
        "config_hash": sha(json.dumps(cfg, sort_keys=True).encode()),
        "image_digest": os.environ.get("IMAGE_DIGEST", "unknown"),
        "torch": torch.__version__, "cuda": torch.version.cuda,
        "cudnn": torch.backends.cudnn.version(), "nccl": ".".join(map(str, torch.cuda.nccl.version())),
        "gpu": torch.cuda.get_device_name(0),
        "shards": {p: sha(open(p, "rb").read(1 << 20)) for p in shard_paths},  # head hash; full hash offline
        "tokenizer": sha(open(tokenizer_path, "rb").read()),
        "layout": layout,               # {"world": 256, "tp": 8, "pp": 4, "dp": 8}
        "root_seed": cfg["seed"],
    }

def verify_on_resume(saved: dict, current: dict, allow: set[str] = frozenset()) -> None:
    diffs = {k for k in saved if saved[k] != current.get(k)} - allow
    if diffs:
        raise RuntimeError(f"manifest mismatch on resume: {sorted(diffs)}")

The allow set makes deliberate changes explicit, for example a container patch after a driver bug, and the decision is logged with the run.

What must be captured so a run can be rerun, resumed and comparedCode + configgit SHA, config hashEnvironmentimage digest, CUDA, NCCLDatashard hashes, tokenizerLayoutworld size, TP/PP/DPRun manifestwritten at step 0Checkpointweights, optimizer, RNG, cursorRerun / resumestate-exact; bitwise needs determinismCompare runsseed bands, metric toleranceBisect divergencereplay window, diff tensorsWithout the manifest and a complete checkpoint, none of the three right-hand activities is possible.
The manifest and a complete checkpoint make rerun, resume, comparison and bisection possible.

Checkpoints complete enough for exact resume

Most resumed runs are not exact because the checkpoint holds weights and optimizer moments but not the rest of the state. A complete checkpoint contains:

  • Model parameters in their training precision, and master weights if mixed precision keeps FP32 copies.
  • Optimizer state: both Adam moments and the step counter, which bias correction uses.
  • LR scheduler state and the global step, tokens seen and samples consumed.
  • Loss scaler state for FP16 training (current scale and growth tracker).
  • RNG state for every rank: Python, NumPy, torch CPU and every CUDA device, plus any tensor-parallel RNG tracker used for dropout in sharded layers.
  • The data cursor: enough to recompute exactly which samples come next.
  • The manifest, so a resume can check it is resuming the same run.
import random, numpy as np, torch

def rng_state():
    return {"py": random.getstate(), "np": np.random.get_state(),
            "cpu": torch.get_rng_state(), "cuda": torch.cuda.get_rng_state_all()}

def set_rng_state(s):
    random.setstate(s["py"]); np.random.set_state(s["np"])
    torch.set_rng_state(s["cpu"]); torch.cuda.set_rng_state_all(s["cuda"])

def resume_exactness_test(train_steps, init_state, n=200, k=100):
    # run n steps straight through; separately run k, save, reload, run n-k; compare
    a = train_steps(init_state(), n)
    mid = train_steps(init_state(), k, save_at_end=True)
    b = train_steps(mid.reload(), n - k)
    for (name, x), (_, y) in zip(a.named_parameters(), b.named_parameters()):
        if not torch.equal(x, y):
            raise AssertionError(f"resume not exact at {name}")

The test checks bit-identical continuation, so run it with deterministic kernels enabled (torch.use_deterministic_algorithms(True) and CUBLAS_WORKSPACE_CONFIG=:4096:8); without them it fails on a correct checkpoint. On a small model the cost is negligible. Run the resume test in CI on a small model and the production layout's shape, because each new feature (a new optimizer, a curriculum, a data mixer) tends to add state that nobody saves. Sharded checkpoint formats and save costs are covered in the GPU checkpointing deep dive.

Data order that survives resume and resharding

Data loaders that hold shuffled iterators in worker processes cannot resume exactly unless their internal state is saved; torchdata's StatefulDataLoader exists for this. A simpler design for pretraining avoids the problem: make the sample order a pure function of the root seed and a global sample index, then derive each rank's slice from the index. Resuming means recomputing from the saved index, and changing the data-parallel degree changes which rank reads a sample but not the global order.

import numpy as np

class GlobalOrder:
    # epoch-wise permutation of sample ids, a pure function of (seed, epoch)
    def __init__(self, n_samples: int, seed: int):
        self.n, self.seed, self._cache = n_samples, seed, {}

    def sample_at(self, g: int) -> int:
        epoch, pos = divmod(g, self.n)
        if epoch not in self._cache:
            self._cache = {epoch: np.random.default_rng([self.seed, epoch]).permutation(self.n)}
        return int(self._cache[epoch][pos])

def batch_for(order: GlobalOrder, step: int, global_batch: int, dp_rank: int, dp_size: int):
    assert global_batch % dp_size == 0, "global batch must divide evenly across DP ranks"
    start = step * global_batch
    per_rank = global_batch // dp_size
    lo = start + dp_rank * per_rank
    return [order.sample_at(g) for g in range(lo, lo + per_rank)]

Two warnings. If the global batch size changes mid-run, step numbers no longer map to the same samples, so store the global sample index, not the step. And data mixtures sampled stochastically from several sources need the same treatment for the mixture draw, or the mixture will differ after resume.

Why the parallel layout changes the numbers

Floating-point addition is not associative, so any change that regroups a sum changes the result in its last bits. Changing the data-parallel degree changes how gradients are reduced across ranks; NCCL may also choose a different algorithm, ring or tree, for a different topology or message size. Changing the tensor-parallel degree splits matrix multiplications differently, so partial sums are combined in a different order. Changing micro-batch count changes gradient accumulation order. Each difference starts at the level of rounding error, and training dynamics amplify it: after a few thousand steps two runs that started identically differ the way two different seeds do. That is expected and is not a bug. The practical rules: record the layout, treat a layout change as a new seed for comparison purposes, and when a node failure forces a smaller world size, prefer waiting for replacement capacity on runs where exact continuation matters.

Telling a regression from noise

Because layouts and kernels inject noise anyway, deciding whether a change helped is a statistical question. Measure the seed spread once per model scale: train the same configuration with several seeds and record the spread of the metrics you make decisions on. Then compare a change against that spread, not against a single baseline run.

import statistics as st

def compare(baseline: list[float], candidate: list[float], min_effect: float):
    mb, mc = st.mean(baseline), st.mean(candidate)
    se = (st.variance(baseline) / len(baseline) + st.variance(candidate) / len(candidate)) ** 0.5
    z = (mb - mc) / se if se else float("inf")
    verdict = "improved" if z > 2 and (mb - mc) >= min_effect else "inconclusive"
    return {"delta": mb - mc, "z": round(z, 2), "verdict": verdict}

# final validation losses from 4 seeds each
print(compare([2.412, 2.418, 2.409, 2.416], [2.405, 2.411, 2.404, 2.409], min_effect=0.003))

Treat z above 2 as a rough screen: with four seeds per arm the t-distribution puts the bar nearer 2.4, so add seeds before acting on a borderline result. Seeds are expensive at large scale, so measure the spread on smaller proxies and on short runs, and use evaluation sets large enough that their own sampling noise is small compared with the effect you care about.

Bisecting divergence and replaying spikes

When a new run diverges from a reference, bisect the cause instead of guessing. First diff the manifests; most explanations stop there. Then load both runs from the same checkpoint, on the same layout, with deterministic kernels enabled, and feed the same batch. Compare per-layer activations, then gradients, then parameters after the optimizer step, and report the first tensor whose difference exceeds a tolerance. Ignoring tiny differences at the bit level is essential; use a relative tolerance near the precision of the dtype, and look for the first tensor where differences grow by orders of magnitude rather than the first that differs at all.

Loss spikes call for a related procedure: restore the checkpoint before the spike and replay the window with the same data. If the spike recurs, inspect the batches in the window, since one bad document or a long run of repeated tokens is a common cause, and the remedy may be to skip those samples. If it does not recur, the cause was non-deterministic, and hardware is a suspect; see GPU hardware faults.

Worked example

A team trains a 7-billion-parameter model on 256 GPUs with TP 8, PP 4, DP 8. At step 40,000 one 8-GPU node fails. Each data-parallel replica spans TP 8 x PP 4 = 32 GPUs, so losing a node removes a whole replica: the job resumes on 224 GPUs at DP 7, with the surviving 24 GPUs of the broken replica idle, and the gradient-accumulation count changes so the global batch stays divisible by 7. Two weeks later an ablation that differs only in warm-up shows a validation loss 0.004 lower, and the team wants to adopt it. The manifest shows the reference run had a layout change at step 40,000, so it is a different sample from the seed distribution, not a matched control. The seed spread measured on a 1-billion proxy is 0.003 standard deviation. One run each cannot separate a 0.004 difference from noise. The team runs three seeds of each configuration on the proxy, finds a mean improvement of 0.005 with a z score above 3, and adopts the change. Meanwhile the resume test is added to CI, and the runbook says to wait up to two hours for replacement nodes before shrinking.

Failure modes

  • Missing RNG or cursor state: resumed runs repeat or skip data and drift.
  • Container tag drift: a rebuilt latest image picks new kernels mid-run.
  • Data re-tokenised in place: shard names unchanged, content different.
  • Single-run comparisons: decisions made inside the noise band.
  • Determinism in production: deterministic kernels left on, costing throughput for no benefit.
  • Silent layout change: elastic resizing alters numerics without a record.

Trade-offs

State-exact resume is cheap once built and should be standard. Bit-identical continuation and bitwise rerun need deterministic kernels that can cost noticeable throughput; enable them for debugging and short CI runs, not for production pretraining. Multiple seeds are the only honest basis for small-effect decisions but multiply compute, so spend them on proxies. Waiting for capacity preserves the layout and costs idle time; shrinking keeps the run moving and costs comparability.

What to do next

  1. Write a manifest at step 0 and fail any resume whose manifest differs.
  2. Add RNG, scheduler, scaler and data-cursor state to checkpoints.
  3. Add a resume-exactness test to CI on a small model.
  4. Make sample order a function of seed and global sample index.
  5. Measure seed spread on a proxy and require it for every comparison.
  6. Write a divergence-bisect script before you need it.
Key takeaway: Reproducibility for LLM training means exact resume on a fixed layout, recorded inputs for every run, and statistical comparison everywhere else. Pin and verify a manifest, checkpoint every piece of state, derive data order from a global index, and compare changes against measured seed spread. Related reading: <a href="gpu_mixed_precision_training.html">mixed precision training</a> and <a href="gpu_nccl_collectives.html">NCCL collectives</a>.