A single GPU is a reliable device. Ten thousand GPUs in one synchronous training job are not, because the job stops when any one of them stops. Meta's Llama 3 paper describes the result: in a 54-day snapshot of pre-training on 16,384 H100 GPUs there were 466 job interruptions, 47 planned and 419 unexpected. The paper attributes 148 interruptions (30.1%) to faulty GPUs and 72 (17.2%) to HBM3 memory, and it reports more than 90% effective training time despite them.

That last number is the goal. Hardware will fail. The software around it decides whether a failure costs ten minutes or ten hours. This article explains what the faults are, how each one reaches your training code, and how to build a job and a cluster that absorb them.

Advertisement

Failure arithmetic: why rare becomes routine

If each GPU independently has a mean time between failures of M hours, a job that needs all N of them has a mean time between interruptions of roughly M / N. Take an illustrative M of 50,000 hours, about five and a half years per device. A 1,000-GPU job then sees an interruption about every 50 hours. At 16,000 GPUs it is about every three hours. The per-device number did not change; the job size did.

Two consequences follow. First, recovery time matters as much as failure rate: if each interruption costs 40 minutes of restart and lost work, a three-hour failure interval wastes a fifth of the cluster. Second, detection has to be automatic. Nobody can triage an interruption every three hours by reading logs.

A taxonomy by what software sees

The useful way to classify faults is by the signal they produce, because that determines what your automation can do. The driver reports many hardware events as Xid messages in the kernel log. NVIDIA's Xid catalog gives each one a description and a recommended resolution label. The table quotes those labels. RESTART_APP means restart the application, RESET_GPU means reset the GPU, RESTART_BM means restart the bare-metal host, and WORKFLOW_* points to a multi-step procedure.

FaultXid (catalog description)Catalog resolutionWhat the job sees
Correctable memory error92 if the rate is high (High single-bit ECC error rate)IGNORENothing; counters rise
Row remap recorded63 (GPU memory remapping event)IGNORENothing now; remap applies at next reset
Row remap failed64 (GPU memory remapping failure)RESET_GPUPossibly nothing yet; node needs service
Uncorrectable, contained94 (Contained memory error)RESTART_APPThe affected process gets a CUDA error
Uncorrectable, uncontained95 (Uncontained memory error)RESET_GPUAll processes on the GPU fail
Double-bit ECC48 (Double Bit ECC Error)WORKFLOW_XID_48CUDA error, usually fatal to the process
Interconnect74 (NVLINK Error)WORKFLOW_NVLINK_ERRCollective errors or hangs
Device lost79 (GPU has fallen off the bus)RESTART_BMEvery call fails; the device is gone
Firmware119 (GSP RPC Timeout), 120 (GSP Error)RESET_GPUHangs or errors on the next call
Software, often yours13 (Graphics Engine Exception), 31 (GPU memory page fault)RESTART_APPIllegal address or kernel fault

Xid 13 and 31 deserve emphasis. They usually come from an out-of-bounds access in a kernel, a framework bug or a bad custom op, not from broken hardware. Treating them as hardware faults sends good machines to repair and leaves the real bug in place.

Two classes produce no Xid at all. Thermal and power throttling make a GPU slow, and in a synchronous job one slow GPU slows every step. Silent data corruption makes a GPU wrong, and nothing reports it. Both are covered below. For reading the counters behind this table, see nvidia-smi and DCGM.

Advertisement

ECC and row remapping, from first principles

HBM stores extra check bits next to the data. The ECC scheme corrects a single flipped bit in a word and detects, but cannot fix, two. A correctable error is free in the moment but is a health signal: a rising count on one device often comes before an uncorrectable one. HBM covers the memory itself.

When a memory row keeps producing errors, the GPU can retire it. On A100 and newer parts this is done by row remapping: the faulty row is replaced with a spare row in the same bank. The replacement is recorded when the error happens (Xid 63) but only takes effect at the next GPU reset. Until then, nvidia-smi's row-remapper section shows a pending remap. Some platforms also require a reboot before the pending flag clears, so check your provider's procedure. If a bank has no spare rows left, the remap fails (Xid 64). The catalog says reset; in practice, take the GPU out of service through your vendor's replacement process.

Operationally this gives a simple rule. A pending remap means drain the node at the next safe point and reset it. A remap failure, or an uncorrectable error that keeps coming back, means remove the GPU from service.

How a fault travels to your training loop

How one GPU fault becomes a stopped training jobHardware eventHBM, NVLink, PCIe, GSPDriverXid in kernel logCUDA runtimeerror on next callFrameworkrank raises, exitsOther ranksblock in collectiveNCCL watchdogtimeout, abortLauncherjob fails, exit codeSchedulerexclude node, requeueHealth pipelineXid + DCGM to node stateQuarantine and diagnosediag, reset, RMARestart from checkpointlose work since last saveSilent data corruption skips every box above: no Xid, no CUDA error, only wrong numbers.It is caught, if at all, by loss and gradient checks or by comparing replicas.
A fatal fault surfaces first in one rank; everyone else learns about it through a collective that never completes.

Follow an uncorrectable memory error through a data-parallel job. The driver logs the Xid. The next CUDA call in the affected process returns an error, which PyTorch raises as a RuntimeError mentioning the CUDA error. That rank exits. Every other rank is now waiting in an all-reduce for a participant that will never arrive. They do not crash, they hang, until the NCCL watchdog's timeout fires and aborts the communicator. The launcher sees a failed worker and tears the job down.

Three practical points follow. Set the process-group timeout deliberately: long enough to cover your slowest legitimate collective and checkpoint barrier, short enough that a dead peer costs minutes rather than an hour. Log the hostname and rank in every fatal error, because the first failing rank is the evidence and the hundred timeouts after it are noise. And make the training script exit with codes that tell the scheduler what happened, as in the loop below.

# A training loop that classifies failures so the scheduler can react correctly.
import sys, math, socket, datetime, torch, torch.distributed as dist

EXIT_HW_SUSPECT = 75      # convention agreed with the launcher: exclude this node and requeue
EXIT_NUMERIC = 76         # loss went non-finite: roll back further, do not blame hardware yet
EXIT_PEER_FAILED = 77     # a collective timed out: some other rank died, requeue without excluding

dist.init_process_group("nccl", timeout=datetime.timedelta(minutes=10))
step = load_latest_checkpoint(model, optimizer)          # your checkpoint code

try:
    while step < total_steps:
        loss = train_step(next(batches))
        if not math.isfinite(loss.item()):
            sys.exit(EXIT_NUMERIC)
        step += 1
        if step % ckpt_every == 0:
            save_checkpoint(model, optimizer, step)      # async writer, atomic rename on finish
        if step % verify_every == 0:
            verify_replicas(model)                       # see the SDC section
except RuntimeError as e:
    msg = str(e).lower()
    print(f"rank {dist.get_rank()} host {socket.gethostname()}: {e}", flush=True)
    if "ecc" in msg or "uncorrectable" in msg:
        sys.exit(EXIT_HW_SUSPECT)
    if "nccl" in msg or "watchdog" in msg:
        sys.exit(EXIT_PEER_FAILED)
    raise                      # illegal address and the rest: likely a bug, fail loudly

The string matching is crude; tune it to your stack's messages. The point is the split. A memory error excludes this node and requeues. A collective timeout means a peer died, so it requeues without excluding. A numerical failure rolls back further. Anything else, including illegal-address errors, fails loudly for a person. Collective hangs and topology issues are covered in NCCL collectives.

Silent data corruption

Some faults produce wrong results with no error at all: a marginal arithmetic unit, a bit flip in an unprotected structure, a defective part that fails only under particular voltage and temperature. Large CPU fleet studies from Meta (Silent Data Corruptions at Scale) and Google (Cores that Don't Count) showed such defects occur at measurable rates across large fleets.

In training, corruption usually shows up in one of three ways. Sometimes it is harmless noise. Sometimes it is a sudden loss spike or a non-finite value. Sometimes it is slow divergence that is hard to tell apart from an optimisation problem. You have three tools against it. Watch loss and gradient norms with rollback on anomalies. Periodically compare replicas, which in pure data parallelism must hold identical weights. And when a node is suspected, re-run a fixed computation on it and compare against a known answer.

# Data-parallel replicas must hold identical weights after each optimizer step.
# Compare an exact integer checksum of the raw bits across the group; a mismatch names the odd rank.
import torch, torch.distributed as dist

def verify_replicas(model, group=None):
    with torch.no_grad():
        fp = torch.zeros(1, dtype=torch.int64, device="cuda")
        for p in model.parameters():
            word = {1: torch.uint8, 2: torch.int16, 4: torch.int32, 8: torch.int64}[p.element_size()]
            fp += p.detach().contiguous().view(word).sum(dtype=torch.int64)   # any single bit flip changes it
    world = dist.get_world_size(group)
    all_fp = [torch.zeros_like(fp) for _ in range(world)]
    dist.all_gather(all_fp, fp, group=group)
    vals = [t.item() for t in all_fp]                    # exact: replicas must match bit for bit
    majority = max(set(vals), key=vals.count)
    bad = [r for r, v in enumerate(vals) if v != majority]
    if bad:
        raise RuntimeError(f"replica divergence (SDC suspect) on ranks {bad}")

The replica check costs one tiny all-gather and a pass over the weights, so running it every few hundred steps is cheap. Under sharded schemes such as FSDP or ZeRO, each rank holds different shards, so compare fingerprints of the same shard across the replica group instead. The majority vote names the node to quarantine.

Checkpointing: how often

Every interruption loses the work since the last checkpoint plus the restart time. Checkpointing more often loses less but costs time on every save. The classic first-order answer, Young's approximation, puts the best interval at roughly the square root of 2 x C x M, where C is the time to write a checkpoint and M is the job's mean time between interruptions.

With C = 1 minute and M = 180 minutes, the interval is the square root of 360, about 19 minutes. Asynchronous checkpointing, where the step continues while a background thread writes, cuts the effective C and pushes the best interval down. Faster restart attacks the other term. Warm spare nodes, pre-pulled images and a fast checkpoint tier all help.

Node lifecycle: quarantine before you trust

Jobs should only land on nodes that recently proved healthy, and a node should leave the pool the moment it looks suspect. A workable state machine has five states. Healthy: schedulable. Suspect: an Xid, a fatal-error exit code or a straggler signal; finish or drain current work, schedule nothing new. Quarantined: run DCGM diagnostics, check row-remapper state, reset the GPUs. Burn-in: a short stress run plus a known-answer computation and an all-reduce bandwidth test against neighbours. Healthy again, or out for repair.

Run a short preflight on every node a job is about to use, and keep a per-node history. A node that fails three times in a month goes to repair even if every individual diagnostic passes, because intermittent faults are exactly what diagnostics miss.

Worked example: a 2,048-GPU run

Assume, for illustration, 256 nodes of 8 GPUs and an interruption every 6 hours across the job. Checkpoints take 2 minutes synchronously, and a restart (detect, exclude, reschedule, load) takes 25 minutes.

Young's interval is the square root of 2 x 2 x 360, about 38 minutes. Each interruption then costs on average half an interval of lost work, 19 minutes, plus 25 minutes of restart: 44 minutes per 6 hours, about 12% of wall time. Checkpoint writes cost another 2 / 38, about 5%. Goodput is roughly 83%.

Now improve recovery. Asynchronous saves cut C to 20 seconds, which shortens the interval to about 15 minutes and the average loss to about 8 minutes. Hot spares and a cached image cut restart to 8 minutes. Interruptions now cost 16 minutes per 6 hours, about 4.5%, plus about 2% for saves. Goodput rises to roughly 93%, with no change in hardware reliability at all.

Failure modes in the fault handling itself

  • Restart onto the same bad node. The requeue does not exclude the failing host, so the job dies again within minutes.
  • Blaming the hundredth rank. Triage reads the last error, an NCCL timeout, instead of the first one from the rank that actually died.
  • Hardware tickets for software bugs. Xid 13 or 31 from a new custom kernel sends healthy nodes to repair.
  • Corrupted checkpoint. A save written after silent corruption is restored as truth. Keep several checkpoints and validate loss after restore.
  • Timeouts set for the happy path. A watchdog shorter than a checkpoint barrier kills healthy jobs. A watchdog of hours turns every fault into an hour of idle cluster.

Trade-offs

DecisionCheaperSafer
Checkpoint intervalLong: less save overheadShort: less lost work
Preflight depthQuick checks, faster startsFull diagnostics, fewer bad starts
Replica verificationOff: no overheadEvery few hundred steps: catches divergence early
Spare capacityNone: every GPU trainsHot spares: much faster recovery

What to do next

  1. Measure your current interruption rate and mean restart time; compute goodput.
  2. Ship Xid and DCGM health events into the scheduler so suspect nodes stop receiving work automatically.
  3. Give your training script distinct exit codes for hardware-suspect, numerical and unknown failures, and map them to requeue rules that exclude the failing host.
  4. Set the NCCL process-group timeout deliberately and log hostname and rank in every fatal error.
  5. Set the checkpoint interval from Young's formula, and move to asynchronous saves if C is more than a minute.
  6. Add a periodic replica fingerprint check and a known-answer test to node burn-in.
  7. Drain and reset nodes with pending row remaps; remove GPUs with remap failures or recurring uncorrectable errors.
  8. Keep a per-node failure history and retire repeat offenders even when diagnostics pass.
Key takeaway: At scale, GPU faults are a scheduling and software problem. Learn the signals: Xids and their resolution labels for hard faults, throttling for stragglers, nothing at all for silent corruption. Turn each signal into an automatic action: exit codes that exclude the bad node, quarantine with diagnostics and burn-in, replica checks for corruption. Then reduce what each interruption costs with checkpoint intervals sized by Young's formula, asynchronous saves and fast restarts. Hardware reliability barely changes; goodput is what you control.