A single GPU is a reliable device. Ten thousand GPUs in one synchronous training job are not, because the job stops when any one of them stops. Meta's Llama 3 paper describes the result: in a 54-day snapshot of pre-training on 16,384 H100 GPUs there were 466 job interruptions, 47 planned and 419 unexpected. The paper attributes 148 interruptions (30.1%) to faulty GPUs and 72 (17.2%) to HBM3 memory, and it reports more than 90% effective training time despite them.
That last number is the goal. Hardware will fail. The software around it decides whether a failure costs ten minutes or ten hours. This article explains what the faults are, how each one reaches your training code, and how to build a job and a cluster that absorb them.
Failure arithmetic: why rare becomes routine
If each GPU independently has a mean time between failures of M hours, a job that needs all N of them has a mean time between interruptions of roughly M / N. Take an illustrative M of 50,000 hours, about five and a half years per device. A 1,000-GPU job then sees an interruption about every 50 hours. At 16,000 GPUs it is about every three hours. The per-device number did not change; the job size did.
Two consequences follow. First, recovery time matters as much as failure rate: if each interruption costs 40 minutes of restart and lost work, a three-hour failure interval wastes a fifth of the cluster. Second, detection has to be automatic. Nobody can triage an interruption every three hours by reading logs.
A taxonomy by what software sees
The useful way to classify faults is by the signal they produce, because that determines what your automation can do. The driver reports many hardware events as Xid messages in the kernel log. NVIDIA's Xid catalog gives each one a description and a recommended resolution label. The table quotes those labels. RESTART_APP means restart the application, RESET_GPU means reset the GPU, RESTART_BM means restart the bare-metal host, and WORKFLOW_* points to a multi-step procedure.
| Fault | Xid (catalog description) | Catalog resolution | What the job sees |
|---|---|---|---|
| Correctable memory error | 92 if the rate is high (High single-bit ECC error rate) | IGNORE | Nothing; counters rise |
| Row remap recorded | 63 (GPU memory remapping event) | IGNORE | Nothing now; remap applies at next reset |
| Row remap failed | 64 (GPU memory remapping failure) | RESET_GPU | Possibly nothing yet; node needs service |
| Uncorrectable, contained | 94 (Contained memory error) | RESTART_APP | The affected process gets a CUDA error |
| Uncorrectable, uncontained | 95 (Uncontained memory error) | RESET_GPU | All processes on the GPU fail |
| Double-bit ECC | 48 (Double Bit ECC Error) | WORKFLOW_XID_48 | CUDA error, usually fatal to the process |
| Interconnect | 74 (NVLINK Error) | WORKFLOW_NVLINK_ERR | Collective errors or hangs |
| Device lost | 79 (GPU has fallen off the bus) | RESTART_BM | Every call fails; the device is gone |
| Firmware | 119 (GSP RPC Timeout), 120 (GSP Error) | RESET_GPU | Hangs or errors on the next call |
| Software, often yours | 13 (Graphics Engine Exception), 31 (GPU memory page fault) | RESTART_APP | Illegal address or kernel fault |
Xid 13 and 31 deserve emphasis. They usually come from an out-of-bounds access in a kernel, a framework bug or a bad custom op, not from broken hardware. Treating them as hardware faults sends good machines to repair and leaves the real bug in place.
Two classes produce no Xid at all. Thermal and power throttling make a GPU slow, and in a synchronous job one slow GPU slows every step. Silent data corruption makes a GPU wrong, and nothing reports it. Both are covered below. For reading the counters behind this table, see nvidia-smi and DCGM.
ECC and row remapping, from first principles
HBM stores extra check bits next to the data. The ECC scheme corrects a single flipped bit in a word and detects, but cannot fix, two. A correctable error is free in the moment but is a health signal: a rising count on one device often comes before an uncorrectable one. HBM covers the memory itself.
When a memory row keeps producing errors, the GPU can retire it. On A100 and newer parts this is done by row remapping: the faulty row is replaced with a spare row in the same bank. The replacement is recorded when the error happens (Xid 63) but only takes effect at the next GPU reset. Until then, nvidia-smi's row-remapper section shows a pending remap. Some platforms also require a reboot before the pending flag clears, so check your provider's procedure. If a bank has no spare rows left, the remap fails (Xid 64). The catalog says reset; in practice, take the GPU out of service through your vendor's replacement process.
Operationally this gives a simple rule. A pending remap means drain the node at the next safe point and reset it. A remap failure, or an uncorrectable error that keeps coming back, means remove the GPU from service.
How a fault travels to your training loop
Follow an uncorrectable memory error through a data-parallel job. The driver logs the Xid. The next CUDA call in the affected process returns an error, which PyTorch raises as a RuntimeError mentioning the CUDA error. That rank exits. Every other rank is now waiting in an all-reduce for a participant that will never arrive. They do not crash, they hang, until the NCCL watchdog's timeout fires and aborts the communicator. The launcher sees a failed worker and tears the job down.
Three practical points follow. Set the process-group timeout deliberately: long enough to cover your slowest legitimate collective and checkpoint barrier, short enough that a dead peer costs minutes rather than an hour. Log the hostname and rank in every fatal error, because the first failing rank is the evidence and the hundred timeouts after it are noise. And make the training script exit with codes that tell the scheduler what happened, as in the loop below.
# A training loop that classifies failures so the scheduler can react correctly.
import sys, math, socket, datetime, torch, torch.distributed as dist
EXIT_HW_SUSPECT = 75 # convention agreed with the launcher: exclude this node and requeue
EXIT_NUMERIC = 76 # loss went non-finite: roll back further, do not blame hardware yet
EXIT_PEER_FAILED = 77 # a collective timed out: some other rank died, requeue without excluding
dist.init_process_group("nccl", timeout=datetime.timedelta(minutes=10))
step = load_latest_checkpoint(model, optimizer) # your checkpoint code
try:
while step < total_steps:
loss = train_step(next(batches))
if not math.isfinite(loss.item()):
sys.exit(EXIT_NUMERIC)
step += 1
if step % ckpt_every == 0:
save_checkpoint(model, optimizer, step) # async writer, atomic rename on finish
if step % verify_every == 0:
verify_replicas(model) # see the SDC section
except RuntimeError as e:
msg = str(e).lower()
print(f"rank {dist.get_rank()} host {socket.gethostname()}: {e}", flush=True)
if "ecc" in msg or "uncorrectable" in msg:
sys.exit(EXIT_HW_SUSPECT)
if "nccl" in msg or "watchdog" in msg:
sys.exit(EXIT_PEER_FAILED)
raise # illegal address and the rest: likely a bug, fail loudlyThe string matching is crude; tune it to your stack's messages. The point is the split. A memory error excludes this node and requeues. A collective timeout means a peer died, so it requeues without excluding. A numerical failure rolls back further. Anything else, including illegal-address errors, fails loudly for a person. Collective hangs and topology issues are covered in NCCL collectives.
Silent data corruption
Some faults produce wrong results with no error at all: a marginal arithmetic unit, a bit flip in an unprotected structure, a defective part that fails only under particular voltage and temperature. Large CPU fleet studies from Meta (Silent Data Corruptions at Scale) and Google (Cores that Don't Count) showed such defects occur at measurable rates across large fleets.
In training, corruption usually shows up in one of three ways. Sometimes it is harmless noise. Sometimes it is a sudden loss spike or a non-finite value. Sometimes it is slow divergence that is hard to tell apart from an optimisation problem. You have three tools against it. Watch loss and gradient norms with rollback on anomalies. Periodically compare replicas, which in pure data parallelism must hold identical weights. And when a node is suspected, re-run a fixed computation on it and compare against a known answer.
# Data-parallel replicas must hold identical weights after each optimizer step.
# Compare an exact integer checksum of the raw bits across the group; a mismatch names the odd rank.
import torch, torch.distributed as dist
def verify_replicas(model, group=None):
with torch.no_grad():
fp = torch.zeros(1, dtype=torch.int64, device="cuda")
for p in model.parameters():
word = {1: torch.uint8, 2: torch.int16, 4: torch.int32, 8: torch.int64}[p.element_size()]
fp += p.detach().contiguous().view(word).sum(dtype=torch.int64) # any single bit flip changes it
world = dist.get_world_size(group)
all_fp = [torch.zeros_like(fp) for _ in range(world)]
dist.all_gather(all_fp, fp, group=group)
vals = [t.item() for t in all_fp] # exact: replicas must match bit for bit
majority = max(set(vals), key=vals.count)
bad = [r for r, v in enumerate(vals) if v != majority]
if bad:
raise RuntimeError(f"replica divergence (SDC suspect) on ranks {bad}")The replica check costs one tiny all-gather and a pass over the weights, so running it every few hundred steps is cheap. Under sharded schemes such as FSDP or ZeRO, each rank holds different shards, so compare fingerprints of the same shard across the replica group instead. The majority vote names the node to quarantine.
Checkpointing: how often
Every interruption loses the work since the last checkpoint plus the restart time. Checkpointing more often loses less but costs time on every save. The classic first-order answer, Young's approximation, puts the best interval at roughly the square root of 2 x C x M, where C is the time to write a checkpoint and M is the job's mean time between interruptions.
With C = 1 minute and M = 180 minutes, the interval is the square root of 360, about 19 minutes. Asynchronous checkpointing, where the step continues while a background thread writes, cuts the effective C and pushes the best interval down. Faster restart attacks the other term. Warm spare nodes, pre-pulled images and a fast checkpoint tier all help.
Node lifecycle: quarantine before you trust
Jobs should only land on nodes that recently proved healthy, and a node should leave the pool the moment it looks suspect. A workable state machine has five states. Healthy: schedulable. Suspect: an Xid, a fatal-error exit code or a straggler signal; finish or drain current work, schedule nothing new. Quarantined: run DCGM diagnostics, check row-remapper state, reset the GPUs. Burn-in: a short stress run plus a known-answer computation and an all-reduce bandwidth test against neighbours. Healthy again, or out for repair.
Run a short preflight on every node a job is about to use, and keep a per-node history. A node that fails three times in a month goes to repair even if every individual diagnostic passes, because intermittent faults are exactly what diagnostics miss.
Worked example: a 2,048-GPU run
Assume, for illustration, 256 nodes of 8 GPUs and an interruption every 6 hours across the job. Checkpoints take 2 minutes synchronously, and a restart (detect, exclude, reschedule, load) takes 25 minutes.
Young's interval is the square root of 2 x 2 x 360, about 38 minutes. Each interruption then costs on average half an interval of lost work, 19 minutes, plus 25 minutes of restart: 44 minutes per 6 hours, about 12% of wall time. Checkpoint writes cost another 2 / 38, about 5%. Goodput is roughly 83%.
Now improve recovery. Asynchronous saves cut C to 20 seconds, which shortens the interval to about 15 minutes and the average loss to about 8 minutes. Hot spares and a cached image cut restart to 8 minutes. Interruptions now cost 16 minutes per 6 hours, about 4.5%, plus about 2% for saves. Goodput rises to roughly 93%, with no change in hardware reliability at all.
Failure modes in the fault handling itself
- Restart onto the same bad node. The requeue does not exclude the failing host, so the job dies again within minutes.
- Blaming the hundredth rank. Triage reads the last error, an NCCL timeout, instead of the first one from the rank that actually died.
- Hardware tickets for software bugs. Xid 13 or 31 from a new custom kernel sends healthy nodes to repair.
- Corrupted checkpoint. A save written after silent corruption is restored as truth. Keep several checkpoints and validate loss after restore.
- Timeouts set for the happy path. A watchdog shorter than a checkpoint barrier kills healthy jobs. A watchdog of hours turns every fault into an hour of idle cluster.
Trade-offs
| Decision | Cheaper | Safer |
|---|---|---|
| Checkpoint interval | Long: less save overhead | Short: less lost work |
| Preflight depth | Quick checks, faster starts | Full diagnostics, fewer bad starts |
| Replica verification | Off: no overhead | Every few hundred steps: catches divergence early |
| Spare capacity | None: every GPU trains | Hot spares: much faster recovery |
What to do next
- Measure your current interruption rate and mean restart time; compute goodput.
- Ship Xid and DCGM health events into the scheduler so suspect nodes stop receiving work automatically.
- Give your training script distinct exit codes for hardware-suspect, numerical and unknown failures, and map them to requeue rules that exclude the failing host.
- Set the NCCL process-group timeout deliberately and log hostname and rank in every fatal error.
- Set the checkpoint interval from Young's formula, and move to asynchronous saves if C is more than a minute.
- Add a periodic replica fingerprint check and a known-answer test to node burn-in.
- Drain and reset nodes with pending row remaps; remove GPUs with remap failures or recurring uncorrectable errors.
- Keep a per-node failure history and retire repeat offenders even when diagnostics pass.