Two training runs with the same code, the same data and the same seed often end with different weights. The gap starts in the last bits of a float and grows: after a few thousand steps the loss curves separate, and an evaluation score can move by more than the improvement you were trying to measure. Reproducibility on GPUs is not one switch. It is a set of decisions about which sources of variation you remove, which you record, and which you accept because removing them costs too much speed.
This article explains where nondeterminism comes from, starting from the fact that floating-point addition is not associative. It then walks through the PyTorch controls that remove most of it, a seeded data pipeline, the extra problems of multi-GPU training and batched inference, and a bitwise replay test you can put in CI. By the end you should be able to say which level of reproducibility your project needs and prove you have it.
Three levels of reproducibility
Before turning anything on, decide what you are promising. There are three useful levels, and they cost very different amounts.
| Level | Promise | Typical use | Cost |
|---|---|---|---|
| Bitwise, same setup | Same weights to the last bit on the same GPU model, driver, library versions and world size | Debugging, regression tests, bisecting a bad commit | Deterministic kernels, often 5 to 30 percent slower on affected ops |
| Bitwise across resume | Stopping and resuming from a checkpoint gives the same result as never stopping | Preemptible clusters, long runs | Saving every RNG and data-loader position |
| Statistical | Metric distribution over seeds is stable; single runs may differ | Research claims, model comparisons | Several seeds per configuration |
Bitwise reproducibility across different GPU models or library versions is generally not on offer: a new cuBLAS release may pick a different kernel for the same matrix shape, and a different GPU has different tile sizes. Treat the hardware and software stack as part of the experiment and record it, rather than expecting results to transfer exactly.
Why the same code gives different numbers
Every source of GPU nondeterminism comes back to one property: (a + b) + c is not always equal to a + (b + c) in floating point, because each addition rounds. A sum over a million values gives slightly different results depending on the order the values are combined in. On a CPU loop that order is fixed. On a GPU it often is not.
import torch
x = torch.randn(1_000_000, dtype=torch.float32)
a = x.sum()
b = x[torch.randperm(x.numel())].sum() # same values, different order
print(a.item(), b.item(), (a - b).item()) # difference is usually nonzeroThe order changes for four main reasons:
- Atomic accumulation. Kernels such as
index_add_,scatter_add_and the backward pass of embedding lookups let many threads add into the same output withatomicAdd. The hardware decides which thread lands first, so the order differs on every run. - Algorithm selection. With
torch.backends.cudnn.benchmark = True, cuDNN times several convolution algorithms on the first call and keeps the fastest. Timing noise can pick a different winner on the next run, and different algorithms reduce in different orders. - Split reductions. A matrix multiply may split its inner dimension across thread blocks (split-K) and add the partial results. Which split is chosen can depend on shape, workspace size and concurrent streams.
- Reduced precision. TF32 and BF16 round inputs earlier, so the same order produces a different answer than FP32. That is a deterministic difference, but it makes runs with different settings look irreproducible.
Where a training step can diverge
A training step passes through several of these points. The diagram marks which stages are bookkeeping that you can make exact and which are numerical ordering that needs deterministic kernels or a fixed topology.
The PyTorch determinism controls
PyTorch gathers the controls in a few places. Put them in one function that runs before the model, the optimizer or any CUDA tensor is created, because some of them are read when the CUDA context initialises.
import os, random
import numpy as np
import torch
def make_deterministic(seed: int) -> None:
# cuBLAS needs a fixed workspace configuration to be deterministic (CUDA 10.2+).
# Must be set before the first cuBLAS call; ":16:8" uses less memory, ":4096:8" is faster.
os.environ.setdefault("CUBLAS_WORKSPACE_CONFIG", ":4096:8")
random.seed(seed)
np.random.seed(seed)
torch.manual_seed(seed) # seeds the CPU and all CUDA generators
torch.backends.cudnn.benchmark = False # no timing-based algorithm choice
torch.backends.cudnn.deterministic = True # only deterministic cuDNN algorithms
torch.use_deterministic_algorithms(True) # raise on ops with no deterministic version
# Make precision an explicit choice instead of an inherited default.
torch.backends.cuda.matmul.allow_tf32 = False
torch.backends.cudnn.allow_tf32 = Falsetorch.use_deterministic_algorithms(True) is the important one. It switches operations to deterministic implementations where they exist and raises a RuntimeError when an operation is known to be nondeterministic and has no alternative, so a hidden source of variation becomes a visible failure. Passing warn_only=True turns those errors into warnings, which is useful for a first audit of a large codebase. With the flag on, torch.utils.deterministic.fill_uninitialized_memory is true by default and fills memory from torch.empty with a known value, so a bug that reads uninitialised memory cannot make runs differ.
Two notes on precision. Matmul TF32 has been off by default since PyTorch 1.12, but cudnn.allow_tf32 defaults to on, so convolutions use TF32 on Ampere and later unless you say otherwise. Recent docs mark the allow_tf32 flags for deprecation in favour of fp32_precision settings; use whichever your version documents. Turning TF32 off does not make anything deterministic; it stops a silent precision change between machines.
A seeded data pipeline that survives resume
The data pipeline is pure bookkeeping, so it can be made exact. Three things need seeds: the sampler that shuffles indices, each worker process that runs random augmentation, and any library that keeps its own global generator. PyTorch's documented pattern derives each worker's seed from the loader's generator:
def seed_worker(worker_id: int) -> None:
worker_seed = torch.initial_seed() % 2**32 # already distinct per worker
np.random.seed(worker_seed)
random.seed(worker_seed)
g = torch.Generator()
g.manual_seed(seed)
loader = torch.utils.data.DataLoader(
train_ds, batch_size=64, shuffle=True, num_workers=8,
worker_init_fn=seed_worker, generator=g, persistent_workers=True,
)With a DistributedSampler, call sampler.set_epoch(epoch) at the start of every epoch; without it each epoch reuses the same permutation, which is reproducible but wrong. Resume is where most pipelines break. Restarting a loader at epoch 7 does not put it at batch 3,120 of epoch 7. Either save the sampler position and skip that many indices on resume, or use a stateful loader that can save and restore its own position. Also store the generator states in the checkpoint:
ckpt = {
"model": model.state_dict(),
"optim": optimizer.state_dict(),
"step": step, "epoch": epoch, "batch_in_epoch": batch_idx,
"rng": {
"python": random.getstate(),
"numpy": np.random.get_state(),
"torch": torch.get_rng_state(),
"cuda": torch.cuda.get_rng_state_all(),
"loader": g.get_state(),
},
}Dropout uses the CUDA generator, so if torch.cuda.get_rng_state_all() is not restored, the first step after resume draws different masks and the runs part ways immediately.
Multi-GPU training and reduction order
Data-parallel training adds an all-reduce of gradients after every backward pass. NCCL picks an algorithm (ring or tree, among others), a protocol and a chunking based on message size, world size and topology. Each choice adds partial sums in a different order. On a fixed job (same nodes, same GPU count, same NCCL version and same environment) the choice is usually stable, but NCCL does not promise bitwise-identical results across different configurations, and you should not assume it.
Practical consequences:
- Changing the number of GPUs changes results, even with the global batch held constant through gradient accumulation. Eight GPUs at 32 per GPU and four GPUs at 64 per GPU sum the same gradients in different orders.
- Pinning
NCCL_ALGOandNCCL_PROTOin the job environment removes one source of variation between runs on the same cluster. Record whatever you set, and check the NCCL documentation for your version before relying on a particular value. - Gradient bucketing in DistributedDataParallel depends on parameter registration order and bucket size. Keep
bucket_cap_mbfixed and do not reorder modules between runs you want to compare. - Sharded optimizers (FSDP, ZeRO) use reduce-scatter instead of all-reduce. The same reasoning applies: fixed world size and fixed sharding give stable order; changing either does not.
Inference and batch invariance
Inference has its own trap. A server that batches requests dynamically sends your prompt through kernels whose reduction strategy can depend on how many other requests share the batch. Horace He's post at Thinking Machines Lab, "Defeating Nondeterminism in LLM Inference" (September 2025), argued that this lack of batch invariance, not thread-level concurrency, is the main reason greedy decoding at temperature zero still returns different text, and showed that batch-invariant matmul, attention and normalisation kernels make repeated completions identical at some cost in throughput.
For training and evaluation the lesson is the same: when you compare two checkpoints, evaluate them with the same batch size and the same padding layout, or the comparison includes noise from the harness itself. If you need exact replay of a served answer, log the batch composition along with the prompt, or run a batch-invariant mode where your serving stack offers one.
Proving it: a bitwise replay test
Claims about determinism should be tested, not configured and hoped for. The test below trains a small model for a few steps twice in fresh processes and compares a hash of the weights. A second test interrupts at step k, saves, reloads in a new process and finishes. All three hashes must match.
# repro_check.py -- run: python repro_check.py
import hashlib, subprocess, sys, torch
def weights_hash(model) -> str:
h = hashlib.sha256()
for name, t in sorted(model.state_dict().items()):
h.update(name.encode())
h.update(t.detach().cpu().contiguous().numpy().tobytes())
return h.hexdigest()
def run(mode: str) -> str:
out = subprocess.run([sys.executable, "train_small.py", "--steps", "50", "--mode", mode],
capture_output=True, text=True, check=True)
return out.stdout.strip().splitlines()[-1] # script prints weights_hash(model) last
full_a = run("full")
full_b = run("full")
resumed = run("stop-at-20-then-resume")
assert full_a == full_b, "two clean runs differ: a numerical source is still open"
assert full_a == resumed, "resume differs: some RNG or loader position is not checkpointed"
print("bitwise reproducible:", full_a[:16])Use separate processes so cached cuDNN choices and allocator state do not leak from one run to the next. Hash every parameter and buffer, not just the loss: a loss can match to six digits while a batch-norm running statistic has already drifted. When the first assertion fails, bisect by step: dump hashes every step and find the first step that differs, then the first layer whose gradient differs. That usually points straight at one atomic kernel.
Worked example: is the new augmentation better?
Consider a hypothetical team that fine-tunes a vision transformer on 8 GPUs and sees validation accuracy of 81.4 and 80.9 from two runs with the same seed. Is a new augmentation, which scored 81.6, an improvement?
- They add
make_deterministicand the seeded loader. The repro check now raises insideindex_add_used by a custom pooling layer. Replacing it with a sorted segment sum fixes it; the step is 4 percent slower. - The two-run hash now matches, but resume fails. The checkpoint did not include the CUDA generator, so dropout masks diverge at the first step after reload. Adding it closes the gap.
- With bitwise reproducibility proven, they still run five seeds for each configuration, because the question is about the augmentation, not one seed. Baseline: 81.1 plus or minus 0.3. New: 81.3 plus or minus 0.3. The difference is inside seed noise, so they do not claim an improvement.
Determinism made debugging possible; multiple seeds answered the scientific question. You need both.
Failure modes
- Seeding after model creation. Weights were already initialised from an unseeded generator. Seed first.
- Setting
CUBLAS_WORKSPACE_CONFIGtoo late. Set it in the launcher environment or at the very top of the entry script, before any CUDA work. - Library-level randomness. Augmentation libraries, tokeniser sampling or
hash()based sharding withPYTHONHASHSEEDunset all add variation the PyTorch seed never touches. - Non-deterministic file order.
os.listdirand object-store listings do not guarantee order. Sort file lists before building the dataset. - Comparing across stacks. A driver or cuDNN upgrade changes results. Pin the container image and record
torch.__version__,torch.version.cuda,torch.backends.cudnn.version()and the GPU name with each run. - Leaving determinism on in production training. Some deterministic paths are much slower. Measure, and use the strict mode for debugging and regression tests if the cost is too high for full runs.
Trade-offs and related reading
Determinism trades throughput for debuggability. The cost is concentrated: most matmuls and convolutions have deterministic versions at little cost, while scatter-heavy operations, some backward passes and benchmarked cuDNN choices can lose much more. A reasonable policy is strict determinism in unit tests, the repro check and short debugging runs; in long runs, the bookkeeping fixes (seeds, loader state, RNG in checkpoints, pinned image) always and the kernel flags only where the measured slowdown is acceptable. For claims about model quality, statistical reproducibility over several seeds is required no matter how deterministic each run is.
For more context on the pieces involved, see mixed precision training for how BF16 and TF32 change numerics, NCCL collectives for how all-reduce algorithms order their sums, the GPU data pipeline for loader workers and sharding, and CUDA streams and graphs for how concurrency is scheduled on the device.
What to do next
- Write down which level you need: bitwise same setup, bitwise across resume, or statistical.
- Add a single
make_deterministic(seed)call at the top of the entry point and setCUBLAS_WORKSPACE_CONFIGin the launcher. - Seed the loader with a generator and
worker_init_fn, sort file lists, and callset_epochon distributed samplers. - Save every RNG state and the loader position in checkpoints, and restore them on resume.
- Add the three-hash repro check to CI on one GPU type, with a pinned container image.
- Record GPU model, driver, CUDA, cuDNN, PyTorch version, world size and NCCL settings with every run.
- Report model comparisons over at least three to five seeds, with the spread.