A training GPU holds tens of gigabytes of high-bandwidth memory and many megabytes of on-die SRAM, and every bit can flip. On one GPU this is rare. On ten thousand GPUs running for months it happens daily, and one flipped exponent bit in a weight can turn a loss curve into NaNs.

Error-correcting codes (ECC) fix most single-bit flips silently, detect double-bit flips, and feed counters that tell you which GPUs are wearing out. This article builds ECC from first principles with code you can run, then covers where it sits on an NVIDIA data-centre GPU, what the driver does with an uncorrectable error, how row remapping spends a finite repair budget, and how to turn counters into drain, reset and replace decisions. For how faults reach a training loop, read GPU Hardware Faults alongside this one.

Why bits flip, and why scale makes it your problem

DRAM stores a bit as charge on a tiny capacitor that leaks and must be refreshed. A cell that leaks too fast, a particle strike, or disturbance from heavy access to neighbouring rows can leave it reading the wrong value. SRAM has no refresh, but a particle strike can still flip a latch.

Transient (soft) errors flip a bit once; rewrite the location and it is fine. Persistent (hard) errors come from a damaged cell or row and keep returning. One correctable error tells you little; the same address failing repeatedly is a part wearing out, which is what the hardware repair mechanisms target.

If one GPU sees an uncorrectable error once a decade, a 4,096-GPU job sees about one a day, and each can kill a rank. ECC behaviour belongs in your training design, not just the datasheet.

SECDED from first principles

The workhorse is SECDED: single-error correct, double-error detect. Start with Hamming's idea. Put data bits and parity bits in numbered positions so that parity bit k covers every position whose binary index has bit k set. When you read the word back, recompute each parity. The failed checks, read as a binary number, are the syndrome, and the syndrome is the position of the flipped bit. Flip it back and the data is correct.

Plain Hamming codes cannot tell one error from two: a double error points at an innocent bit, and the decoder 'corrects' it into a third. SECDED adds an overall parity bit. Odd flips break it; even flips do not. A non-zero syndrome with overall parity intact means two errors: report uncorrectable.

Here is SECDED on 4 data bits, small enough to test exhaustively:

def encode(d):                      # d = [d1, d2, d3, d4]
    d1, d2, d3, d4 = d
    p1 = d1 ^ d2 ^ d4               # covers positions 1,3,5,7
    p2 = d1 ^ d3 ^ d4               # covers positions 2,3,6,7
    p4 = d2 ^ d3 ^ d4               # covers positions 4,5,6,7
    word = [p1, p2, d1, p4, d2, d3, d4]
    p0 = 0
    for b in word:
        p0 ^= b
    return [p0] + word

def decode(cw):
    word = cw[1:]
    syndrome = 0
    for pos in range(1, 8):
        if word[pos - 1]:
            syndrome ^= pos
    overall = 0
    for b in cw:
        overall ^= b
    if syndrome == 0 and overall == 0:
        status = "clean"
    elif overall == 1:              # odd flips: assume one, fix it
        if syndrome:
            word[syndrome - 1] ^= 1
        status = f"corrected bit {syndrome}"
    else:
        return None, "uncorrectable (double)"
    return [word[2], word[4], word[5], word[6]], status

cw = encode([1, 0, 1, 1])           # [0, 0, 1, 1, 0, 0, 1, 1]
cw[5] ^= 1
print(decode(cw))                   # ([1, 0, 1, 1], 'corrected bit 5')
cw[2] ^= 1
print(decode(cw))                   # (None, 'uncorrectable (double)')

Looping over all 16 data values and every single and double flip confirms both guarantees. Three flips are not covered: they can look like one and be miscorrected. Real memory uses wider words, such as 8 check bits over 64 data bits, and DRAM ECC may use symbol-based codes. NVIDIA does not publish the exact HBM code, so treat 'SECDED-class' as the mental model, not a specification.

Where ECC lives on a GPU

HBM carries check bits alongside data. On-die SRAM, such as the register file, L1 and L2, is protected by SECDED or by parity; NVIDIA's replacement policy names both kinds. Parity detects a flip but cannot fix it, so a single-bit parity hit is still an uncorrectable error.

ECC does not cover computation. A marginal tensor core that produces a wrong product writes it into protected memory with valid check bits. That silent data corruption is why large runs add loss-spike detection and periodic burn-in tests on top of ECC.

Where error protection sits on a data-centre GPU, and what it does not coverGPU dieRegister fileECC or parityL1 / shared memECC or parityL2 cache slicesSECDED or parity, per structureALUs, tensor cores, datapathsno ECC: logic faults become silent errorsHBM stacksdata + check bits, SECDED-class codespare rows per DRAM bank (row remap)spare channels on some Blackwell partsmemory controllerInfoROM + drivererror counters, remap entries, Xid logGreen and blue are protected storage. Red is computation: ECC never sees a wrong multiply.
ECC protects storage on the die and in HBM. Datapaths and arithmetic units are outside it, so a compute fault is invisible to the counters.

On cards where ECC is switchable, nvidia-smi -e 0|1 changes it after a reset or reboot, and some GDDR cards lose visible capacity with ECC on; measure memory.total on your model. Leave ECC on for anything that trains or serves production traffic.

Reading the counters

The driver keeps two families of counters. Volatile counters reset when the driver reloads or the GPU resets. Aggregate counters persist in the InfoROM, a small non-volatile store on the board. Export both: volatile shows now, aggregate shows the board's whole history, so a clean volatile count after a reboot proves nothing.

# Current and lifetime counters, one line per GPU
nvidia-smi --query-gpu=index,serial,ecc.mode.current,\
ecc.errors.corrected.volatile.total,ecc.errors.uncorrected.volatile.total,\
ecc.errors.uncorrected.aggregate.total --format=csv

# Row-remapper state (Ampere and later)
nvidia-smi -q -d ROW_REMAPPER
nvidia-smi --help-query-remapped-rows

# Kernel log: the driver reports memory events as Xids
dmesg -T | grep -E "NVRM: Xid"

If you already run DCGM, the same data arrives as fields such as DCGM_FI_DEV_ECC_DBE_VOL_TOTAL, DCGM_FI_DEV_UNCORRECTABLE_REMAPPED_ROWS and DCGM_FI_DEV_ROW_REMAP_FAILURE; the DCGM deep dive covers the exporter. Check the field list your version publishes before writing alert rules.

When an error cannot be corrected

An uncorrectable error means data is wrong and the driver must stop it spreading. The details below come from NVIDIA's GPU memory error management guide.

First the driver logs Xid 48 for an uncorrectable ECC error; the guide also lists Xid 171 for a DRAM location and Xid 172 for SRAM. Then it attempts error containment. If containment works, the log shows Xid 94 and only the processes that touched the bad data are terminated; everything else on the GPU keeps running with correct results, and new work can launch. If containment fails, the log shows Xid 95, and the GPU needs a reset, which ends every process on it. On Volta and earlier, every uncorrectable error behaved like the second case.

Next, dynamic page offlining marks the page holding the faulty location unusable, so no current or future allocation lands on it, without a reset. At the same time the GPU records a row remap in the InfoROM (Xid 63). The remap swaps the faulty DRAM row for a spare row in the same bank, but only takes effect at the next GPU reset. Until then the 'pending' flag is set. After the reset the row is fixed in hardware, the offlined page returns, and there is no hole in the address space.

Uncorrectable HBM error on a GPU with containment (GA100, GH100, GB10x class)UncorrectableXid 48Containment?the driver decidesyes: Xid 94Kill the processothers keep runningno: Xid 95Reset GPUall work on it lostOffline the pagedynamic page offliningRemap pendingXid 63, waits for resetResetrow fixedOr: failure flagspares gone: RMA pathCheckpoint-restart handles the killed process; the scheduler handles drain, reset and replacement.
The response to an uncorrectable HBM error, from NVIDIA's memory error management guide. A contained error costs one process; an uncontained one costs the GPU until reset.

Support differs by chip. Row remapping exists on GA100, GA10x, Ada AD10x, GH100 and GB10x, but containment and page offlining are listed only for GA100, GH100 and GB10x, so plan for whole-GPU resets on the others. Even where supported, rare errors stay uncontained.

The repair budget: remaps, failure flags and RMA

Row remapping replaced page retirement. Older GPUs retired up to 64 pages, permanently, leaving holes. Ampere and later support up to 512 remappings, and remaps caused by correctable errors can be displaced by remaps for uncorrectable errors when a bank runs short of spares. NVIDIA sets the row-remap failure flag if any of these happen:

  • an uncorrectable error needs a remap in a bank that already has eight uncorrectable rows remapped;
  • an uncorrectable error lands on a row that was already remapped, even if the bank has fewer than eight;
  • the GPU has already performed 512 remappings for uncorrectable errors.

Separately, Xid 64 means a remap entry could not be recorded in the InfoROM. Either signal puts the GPU on the replacement path.

On some Blackwell products there is one more step. After two remaps in a bank, the next uncorrectable error there tries to swap the whole DRAM channel, or a failing L2 slice, for a spare. Success logs Xid 160, and the repair needs a reboot to take effect. NVIDIA notes this is not guaranteed on every Blackwell GPU.

nvidia-smi also reports how many banks have Max, High, Partial, Low or None spare rows left; banks drifting into Low are the early warning. NVIDIA's RMA rule for DRAM is that the failure flag is set and confirmed by its field diagnostic; for SRAM, more than two unique uncorrectable events in one address bank of a SECDED-protected structure, or more than four for parity-protected ones.

A node policy you can run

Run a small agent on every node, before placement and on a timer. Pending means 'reset at the next safe point', failure means 'replace', a corrected-error burst means 'watch'.

import csv, io, subprocess

FIELDS = ["index", "serial",
          "ecc.errors.uncorrected.volatile.total",
          "ecc.errors.corrected.volatile.total"]
REMAP = ["gpu_bus_id", "remapped_rows.pending", "remapped_rows.failure"]

def query(flag, fields):
    out = subprocess.run(["nvidia-smi", f"{flag}={','.join(fields)}",
                          "--format=csv,noheader,nounits"],
                         capture_output=True, text=True, check=True).stdout
    return [dict(zip(fields, [v.strip() for v in row]))
            for row in csv.reader(io.StringIO(out))]

def decide(gpu, remap, prev_corrected, burst=1000):
    if remap["remapped_rows.failure"].lower() in ("yes", "1"):
        return "REPLACE"
    if remap["remapped_rows.pending"].lower() in ("yes", "1"):
        return "DRAIN_THEN_RESET"
    if int(gpu["ecc.errors.uncorrected.volatile.total"]) > 0:
        return "DRAIN_THEN_RESET"
    corrected = int(gpu["ecc.errors.corrected.volatile.total"])
    if corrected - prev_corrected.get(gpu["serial"], 0) > burst:
        return "WATCH"
    return "OK"

gpus = query("--query-gpu", FIELDS)
remaps = query("--query-remapped-rows", REMAP)
for g, r in zip(gpus, remaps):
    print(g["index"], g["serial"], decide(g, r, prev_corrected={}))

Confirm the remap field names with --help-query-remapped-rows on your driver, and in production match rows by bus ID. Persist corrected counts per serial, not slot, so a moved card keeps its history. Let the scheduler drain: the agent only sets a taint, and the job manager moves work at a checkpoint boundary.

Worked example: a week on a 512-GPU cluster

Take a 512-GPU H100 cluster, 64 nodes of eight, running one long pre-training job with checkpoints every 30 minutes. Over a week the agent sees the following.

DaySignalMeaningAction
1Node 17: corrected +40Background rateNone
2Node 22: Xid 48, 94, 63Contained, remap pendingRestart; drain, reset, return
4Node 40: corrected +25,000 in an hourA row degradingWATCH
5Node 40: Xid 48, 94, 63, same bankDegradation confirmedRestart; drain, reset, book diagnostics
6Node 9: Xid 95UncontainedRestart; reset GPU
7Node 40: Xid 94 on the row remapped on day 5Failure flag setRestart; remove, RMA

The job restarted four times. With 30-minute checkpoints and about 10 minutes to reload, each restart cost on average 15 minutes of lost progress plus 10 of recovery: 25 minutes times 512 GPUs, about 213 GPU-hours. Had node 40 been pulled for diagnostics on day 4, when the corrected burst appeared, the day 5 and day 7 restarts would not have happened.

Failure modes

  • Alerting on any corrected error. Background corrected errors are expected. Page on uncorrected errors, pending remaps and failure flags; trend corrected counts.
  • Ignoring the pending flag. A GPU with a pending remap still has an offlined page and an unfixed row. Leaving it for weeks wastes the repair and risks a repeat on the same row, which by NVIDIA's rule sets the failure flag.
  • Resetting under a live job. nvidia-smi -r needs every process off the GPU, and on some NVSwitch baseboards GPUs must be reset together; check your platform's procedure. Drain first.
  • Treating ECC as correctness. Compute faults bypass it. Keep loss-spike detection and periodic diagnostics such as DCGM's longer run levels in the node lifecycle.
  • Assuming containment on every SKU. Ada and GA10x cards remap rows but do not contain errors.

Trade-offs

ECC costs check-bit storage and, on some GDDR cards, visible capacity; in return it removes most soft errors and turns hard ones into counted, repairable events. The real decisions are operational. Draining on every uncorrected error costs churn but keeps nodes clean. Deferring resets saves churn but leaves an offlined page and risk. Replacing boards early on a rising trend costs spares but avoids mid-run failures. The vendor's flags are the floor, not the ceiling. For how these signals combine with others into a node score, see LLM System Health Scoring, and for the restart side, the checkpointing deep dive. The HBM article explains the memory these codes protect.

What to do next

  1. Run nvidia-smi -q -d ECC and nvidia-smi -q -d ROW_REMAPPER on every node today and record any pending or failure flags.
  2. Export volatile and aggregate ECC counters and remap state per GPU serial to your metrics system.
  3. Alert on Xid 48, 64, 94, 95 and any pending remap; trend corrected errors per GPU instead of paging on them.
  4. Wire a drain-then-reset path into the scheduler so pending remaps are applied at checkpoint boundaries.
  5. Check which SKUs in your fleet support error containment before promising tenants isolation.
  6. Keep silent-corruption defences that ECC cannot provide: loss-spike checks and periodic diagnostics.
Key takeaway: ECC corrects single flips and detects doubles in GPU memory and on-die SRAM, but it never sees a wrong computation. On Ampere and later, an uncorrectable error is contained where the chip supports it, its page is offlined, and a row remap waits for the next reset. Trend corrected errors, act on pending remaps at checkpoint boundaries, and replace the GPU when the failure flag says its spares are gone.