Most writing on MLOps is about models: experiment tracking, registries, feature stores and deployment pipelines. On a GPU fleet there is a second, lower layer of operations, and it decides whether any of that matters. GPUs fail far more often than CPUs, one bad device can stall a job spread over a thousand others, and the hardware is expensive enough that ten percent of lost time is a line item people ask about.
This article treats GPU operations as a control loop. Signals come from the devices, the interconnect and the jobs. A controller moves each node through a small set of lifecycle states. The jobs are written to survive the failures the controller cannot prevent. We will build the state machine, write the triage classifier, work through the arithmetic of failures at scale, derive a checkpoint interval from first principles and define goodput, the one metric that ties it all together.
What GPU ML Ops owns
The scope is everything between the hardware and the training or serving code. That covers the node image (driver, CUDA user-space libraries, NCCL, the container runtime hooks), health monitoring, the admission gate that decides whether a node can take work, remediation when it breaks, the scheduler's view of which nodes are healthy, and the checkpoint and restart machinery that lets jobs survive a loss. The team that owns it is judged on one question: of the GPU-hours we paid for, how many produced useful training steps or served requests?
Two things make this harder than ordinary server operations. Failures are correlated with work: a GPU that passes idle checks can fail under sustained load or heavy NVLink traffic. And the blast radius is set by the job, not the node: one failed rank aborts the collective for everyone, and one slow rank slows everyone.
The node lifecycle state machine
Give every node exactly one state, store it somewhere authoritative (a Kubernetes node label and taint, a Slurm node state and reason, or a fleet database that drives both), and allow only a fixed set of transitions. The states in the diagram are a good default.
The rules that matter are the edges. Cordoned means no new work, but the current job keeps running. Use it for soft signals such as a rising corrected-error rate or a thermal warning, and let the job finish or reach its next checkpoint. Draining means the job must leave now: signal it to checkpoint if it still can, then terminate it. Repair is the cheap fix list, in order: GPU reset, node reboot, firmware reflash, cable or transceiver reseat. RMA is the exit for hardware that fails repair twice. The non-negotiable rule is that nothing returns to Ready without passing burn-in again. Record every transition with a reason code, because the history of transitions is how you find the node that fails every nine days.
Signals, and why the Xid class matters more than the Xid
Device telemetry comes from DCGM, which the DCGM deep dive covers field by field. For lifecycle decisions you need a small subset: DCGM_FI_DEV_XID_ERRORS for driver-reported faults, DCGM_FI_DEV_ECC_DBE_VOL_TOTAL for uncorrectable memory errors, DCGM_FI_DEV_ROW_REMAP_FAILURE and DCGM_FI_DEV_UNCORRECTABLE_REMAPPED_ROWS for HBM health on recent parts, DCGM_FI_DEV_PCIE_REPLAY_COUNTER for a degrading PCIe link, plus temperature, power and clocks. Add kernel logs, InfiniBand port counters and, most importantly, the job's own per-rank step times.
Xid codes are the driver's fault reports, and the first operational decision is whether the fault belongs to the hardware or to the application. Several common codes are usually caused by the program: Xid 13 (graphics engine exception), 31 (GPU memory page fault), 43 (the GPU stopped processing, often after a user fault) and 45 (preemptive cleanup when a process is killed). Draining a node for these punishes the hardware for a bug in someone's kernel. Others point at the device: 48 is a double-bit ECC error, 79 means the GPU has fallen off the bus, 94 and 95 are the contained and uncontained ECC errors introduced with the A100, and 63 and 64 record page retirement on older parts and row remapping on Ampere and later, with 64 meaning it failed. Treat NVIDIA's Xid catalog as the source of truth for the exact recommended action, because it changes between driver branches.
| Signal | Likely owner | Lifecycle action |
|---|---|---|
| Xid 13, 31, 43, 45 on one job only | Application | No node action; attach to the job's failure record |
| Same app-class Xid across many jobs on one node | Hardware, probably | Cordon, then long diagnostic |
| Xid 48, 79, 95, row-remap failure | Hardware | Drain now, then Repair |
| Xid 94, or remap pending | Hardware, contained | Cordon; reset at the job boundary |
| PCIe replays or IB symbol errors rising | Link or cable | Cordon, reseat at the next window |
| Rank step time persistently above peers | Unknown | Cordon, run the performance burn-in |
The remediation controller
The controller is a loop that reads events, classifies them and applies transitions. Keep it boring and deterministic, rate-limit it, and give humans an override. The core fits on a page.
APP_XIDS = {13, 31, 43, 45}
HARD_XIDS = {48, 79, 95}
CONTAINED_XIDS = {94}
def classify(event, history):
# event: dict(kind, node, gpu, xid=None, job=None)
# history: recent events for this node, newest last
if event["kind"] == "xid":
x = event["xid"]
if x in HARD_XIDS:
return "drain"
if x in CONTAINED_XIDS:
return "cordon"
if x in APP_XIDS:
jobs = {e["job"] for e in history if e.get("xid") in APP_XIDS}
# one job faulting is a bug; many distinct jobs faulting is the node
return "cordon" if len(jobs) >= 3 else "record"
return "cordon" # unknown code: be conservative, never ignore
if event["kind"] in ("row_remap_failure", "dbe"):
return "drain"
if event["kind"] in ("pcie_replay_rate", "ib_symbol_errors", "straggler"):
return "cordon"
return "record"
def reconcile(node, action, fleet):
if action == "record":
return
if fleet.fraction_unavailable() > 0.05:
fleet.page_human(node, "cordon budget exhausted") # stop a bad rule draining the fleet
return
if action == "drain":
fleet.signal_checkpoint(node, deadline_s=120)
fleet.set_state(node, "Draining", reason=action)
else:
fleet.set_state(node, "Cordoned", reason=action)Three details earn their place. The distinct-jobs test separates a buggy kernel from a sick GPU using evidence the controller already has. Unknown codes default to cordon, never to ignore. And the unavailability budget stops a bad rule, or a bad driver, from quietly draining a third of the fleet overnight.
Burn-in: the only door into Ready
Burn-in has three stages, and each catches a different class of fault. First, device diagnostics: dcgmi diag -r 3 is the long level, roughly 15 minutes, which stresses memory, compute, PCIe and power. Level 1 takes seconds and suits a pre-job check. Level 4 is extended, can run for over an hour, and belongs in the repair path. Second, an interconnect test: run NCCL all-reduce across all GPUs in the node, then across the node and a known-good peer, and compare bus bandwidth with the fleet median for that SKU. Third, a performance test: a fixed GEMM or a short training step with known throughput. Admit the node only if every stage is within a few percent of the median.
The performance stage matters because reduced clocks, delayed thermal throttling or a lane-degraded PCIe link all pass functional tests and still run slow. Compare against the fleet median, not the datasheet.
Why big jobs fail often: the arithmetic
Failures are roughly independent across nodes, so a job's mean time between failures is about the per-node MTBF divided by the number of nodes. Suppose, as an assumption for this example, that an 8-GPU node suffers an interruption that kills its job about once every 50 days, which is 1,200 hours. A 16-node job then expects a failure every 75 hours. A 128-node job, 1,024 GPUs, expects one every 1,200 / 128 = 9.4 hours. At 1,024 nodes it is about 70 minutes.
Large training jobs must therefore be designed to fail: frequent checkpoints, fast restart, automatic exclusion of the failed node and warm spares. The step anatomy in the training step walkthrough shows why a single rank stalls everyone: every step ends in a collective that cannot complete until every rank arrives.
Choosing a checkpoint interval
Checkpointing trades two costs. Each checkpoint blocks training for some time, call it d. A failure throws away, on average, half an interval of work, plus a fixed restart cost R for rescheduling, reloading and warm-up. With interval T and job MTBF M, the expected fraction of time wasted is about d/T + T/(2M) + R/M. Minimising the first two terms gives Young's approximation, T = sqrt(2dM). Daly's refinement subtracts d, which matters only when d is not small relative to T.
from math import sqrt
def plan(node_mtbf_h, nodes, ckpt_block_s, restart_min):
M = node_mtbf_h / nodes # job MTBF, hours
d = ckpt_block_s / 3600.0
R = restart_min / 60.0
T = sqrt(2 * d * M) # Young
waste = d / T + T / (2 * M) + R / M
return round(T * 60, 1), round(100 * waste, 1)
print(plan(1200, 128, 120, 15)) # synchronous save: (47.4 min, 11.1 %)
print(plan(1200, 128, 10, 15)) # async save: (13.7 min, 5.1 %)The worked numbers are the lesson. With a two-minute synchronous checkpoint the best interval is about 47 minutes and the job still wastes about 11 percent of its time. Make the save asynchronous so training blocks for only ten seconds while a background thread writes, and the best interval falls to about 14 minutes and waste halves to about 5 percent. Of that, the 15-minute restart now costs 2.7 points, so the next win is restart time, not the checkpoint: cache container images on every node, keep spares warm and load the checkpoint from a nearby tier. Re-run the calculation whenever the job size changes, because M falls as the job grows.
Goodput: the one number to report
GPU utilisation tells you the device was busy. It does not tell you the work was useful. Define goodput as the fraction of allocated GPU-hours that produced training progress which was kept: steps completed and later saved in a checkpoint. Everything else is loss, and it is worth splitting into buckets people can act on.
| Bucket | Example cause | Owner |
|---|---|---|
| Scheduling gap | Nodes idle while a job waits for a full gang | Scheduler |
| Startup | Image pull, data loader warm-up, checkpoint load | Platform |
| Checkpoint blocking | Synchronous save stalls every rank | Training infra |
| Lost work | Steps since the last checkpoint when a node dies | Fleet health plus checkpoint cadence |
| Degraded speed | Straggler rank, throttled GPU, bad link | Fleet health |
Emit these buckets from the job itself, with timestamps for start, first step, each checkpoint and exit, and join them with node lifecycle transitions. A weekly goodput report per team shows whether to spend the next month on faster restarts, better burn-in or a smarter scheduler configuration.
Stragglers and silent slowdowns
A synchronous job's step time is the maximum over its ranks, so one GPU running at 85 percent speed slows 1,023 healthy ones to its pace. Detection has to come from the job. Log per-rank compute time for each step, before the gradient collective, and compare each rank with the median across ranks.
import statistics
def stragglers(step_times, threshold=1.08, window=50):
# step_times: {rank: [compute seconds for the last `window` steps]}
med = {r: statistics.median(ts[-window:]) for r, ts in step_times.items()}
fleet = statistics.median(med.values())
return sorted(r for r, m in med.items() if m > threshold * fleet)Use a median over many steps so data-dependent variation does not trigger it, and map ranks to nodes before emitting a straggler event. Common causes are thermal throttling, a lower power cap, a degraded NVLink or NIC, or a noisy co-tenant. Watch NCCL bandwidth as well: collective performance falls to the slowest link in the ring or tree.
Software drift and staged rollouts
The driver, CUDA user-space libraries, NCCL, the network drivers and firmware form a compatibility matrix, and most fleet-wide incidents come from changing one of them everywhere at once. Pin versions in a baked node image and roll it through stages: a canary pool, then non-critical jobs, then ten percent, then everything. Gate each stage on goodput, Xid rates and NCCL bandwidth against the previous image, and keep that image bootable so rollback is a reboot. The container side of the matrix, which driver a given CUDA toolkit needs, is covered in GPU containers.
Failure modes and trade-offs
- Flapping nodes. A node passes a short test after reboot, rejoins and fails again two days later. Count repairs per node over 30 days and send repeat offenders to the long diagnostic or RMA.
- Over-eager draining. Treating application Xids as hardware faults removes healthy capacity and hides the real bug. Keep the distinct-jobs rule.
- Checkpoint storms. Thousands of ranks writing at once saturate shared storage and lengthen d for every job. Stagger saves, write sharded checkpoints and use a local tier first.
- Spare pool sizing. Spares cost money when idle, but without them a restart waits for a repair. Size the pool from the expected failures per day across all large jobs.
- Metric blindness. High GPU utilisation with low goodput is common: the GPUs are busy redoing lost work. Report goodput first.
What to do next
- Write down your node states and legal transitions, and store the current state with a reason code for every node.
- Build a three-stage burn-in (diagnostic, NCCL bandwidth, performance against the fleet median) and make it the only path into Ready.
- Implement the triage classifier with the distinct-jobs rule and an unavailability budget.
- Measure per-node MTBF and checkpoint blocking time, then compute each large job's checkpoint interval with Young's formula.
- Instrument jobs with start, first-step, checkpoint and exit timestamps and publish a weekly goodput breakdown.
- Add per-rank step-time logging and wire straggler events into the controller.
- Move driver and library changes to staged image rollouts with a bootable rollback.