A GPU fleet that serves and trains large language models fails in a small number of recurring ways: a GPU falls off the PCIe bus, HBM throws uncorrectable ECC errors, a collective hangs until the NCCL watchdog kills the job, the KV cache fills and requests start queueing, an engine upgrade moves time-to-first-token. Each of these has a known diagnosis and a known safe action. The difference between a ten-minute incident and a two-hour one is usually whether the person paged at 3 a.m. finds that knowledge in the first minute.
A runbook index is the system that makes this happen. It is not a wiki folder. It is a catalogue of runbook entries stored as data, keyed by the symptom an operator actually sees, with every alert carrying a pointer to exactly one entry, and with automated checks that fail the build when an alert has no runbook, a runbook has no owner, or an entry has not been reviewed for a quarter. This article designs one for GPU LLM infrastructure: the schema, the alert wiring, the lint, a starter catalogue of GPU symptoms with verified error codes, a worked page, and the habits that stop it rotting.
Runbook, playbook and index
Separate three things that teams tend to blur. A runbook is the procedure for one situation: how to confirm it, how to contain it, how to fix it, and when to escalate. A playbook is the incident process around any situation: who is incident commander, where people talk, how updates go out. The index is the lookup structure that gets a responder from a signal to the right runbook, plus the metadata that keeps the collection healthy: owner, scope, last review, related alerts, related entries.
The index earns its keep in three moments. During an incident, it answers ‘what is this and what do I do’ from the alert itself, without search. During design review, it answers ‘which alerts have no procedure’. During postmortems, it answers ‘was the runbook used, and was it right’, which turns every incident into a small edit.
Architecture: the index as data
Keep the index in git next to the alert rules, not in a separate wiki. That gives you review on every change, history for every procedure, and the ability to lint alerts and runbooks together in one pipeline. Render the entries to whatever your responders read, a static site, the internal developer portal or the chat bot, but treat the repository as the source of truth. A wiki is a fine rendering target but a poor database: nothing stops a page being renamed or orphaned.
The entry schema
Every entry is a small structured header plus a body in a fixed shape. The header is what the tooling reads; the body is what the human reads.
# runbooks/gpu-xid-79-fell-off-bus.yaml
id: gpu-xid-79-fell-off-bus
title: "GPU has fallen off the bus (Xid 79)"
symptoms:
- "dmesg shows NVRM: Xid ... 79"
- "nvidia-smi errors or lists fewer GPUs than the node should have"
- "training rank exits with CUDA error; serving replica crash-loops on one node"
scope: [training, serving]
severity_default: page
owner: team-gpu-fleet
escalation: hardware-vendor-ticket
alerts: [GpuXidCritical]
related: [gpu-ecc-uncorrectable, nccl-collective-timeout]
last_reviewed: 2026-09-14
review_interval_days: 90
safe_actions: [cordon, drain, reboot-node]
forbidden_actions: [reset-gpu-while-job-running]The body follows the same five headings everywhere, because a responder under stress reads by position: Confirm (the commands that prove this is the situation), Contain (stop the damage, usually cordon and drain), Fix (the repair), Verify (how you know it worked) and Escalate (who, with what evidence). The safe_actions and forbidden_actions fields look bureaucratic until an automation bot uses them to decide what it may do unattended, and a human uses them to avoid resetting a GPU under a job that was about to checkpoint.
A symptom-keyed catalogue for GPU fleets
Key entries by what the responder sees, not by which team owns the component. Nobody is paged with ‘NVLink problem’; they are paged with ‘step time doubled on job X’ or ‘Xid 74 on node Y’. A starter catalogue for a GPU LLM fleet looks like this. The Xid meanings follow NVIDIA's Xid documentation and cloud providers' GPU troubleshooting guides; check the catalogue for your driver branch, because codes are added over time.
| Entry | Signal | First action |
|---|---|---|
| GPU fell off bus | Xid 79 | cordon and drain the node; reboot; replace if it recurs |
| Uncorrectable memory error | Xid 48; Xid 94 (contained) or 95 (uncontained) on A100 and later | 94: restart the affected application; 95: drain and reset or reboot |
| Row remapping | Xid 63 or 64; nvidia-smi -q -d ROW_REMAPPER shows pending or failed remaps | pending: reset the GPU when idle; failure: retire the GPU |
| NVLink error | Xid 74 | drain; run dcgmi diag -r 3 before returning the node |
| GSP firmware timeout | Xid 119 or 120 | drain and reboot; check driver version against known issues |
| Collective hang | NCCL watchdog timeout in rank logs | collect flight recorder dumps; find the slow rank; restart from checkpoint |
| Training OOM | CUDA out of memory at a fixed step | check sequence length or batch change; enable activation checkpointing |
| KV cache saturation | cache usage near 100%, preemptions, waiting queue growing | scale replicas; cap max tokens; check for a long-context traffic shift |
| TTFT regression | TTFT SLO burn with normal error rate | compare engine and config versions; check prefill share |
| Throttling | clock drops with thermal or power violation counters rising | check cooling and power caps; move jobs off the node |
Two of these deserve care in alerting. The DCGM field DCGM_FI_DEV_XID_ERRORS holds the value of the last Xid seen, not a count, so alert when it changes rather than when it is non-zero, and read the kernel log for repeats of the same code. And serving engine metric names move between releases (vLLM, for example, has renamed its KV cache usage gauge), so the entry should name the concept and link to the dashboard rather than hard-code a metric name that will go stale.
Wiring alerts to entries
The link from alert to runbook should be data the lint can check, not a URL pasted into free text. Put a stable runbook_id label on every rule and let the alerting layer build the URL. Prometheus itself attaches no meaning to it; runbook_url is a widely used annotation convention that Alertmanager templates and most pager integrations can render.
groups:
- name: gpu-hardware
rules:
- alert: GpuXidCritical
expr: changes(DCGM_FI_DEV_XID_ERRORS[5m]) > 0
and DCGM_FI_DEV_XID_ERRORS == 79
labels:
severity: page
runbook_id: gpu-xid-79-fell-off-bus
annotations:
summary: "Xid 79 on {{ $labels.Hostname }} GPU {{ $labels.gpu }}"
runbook_url: "https://runbooks.internal/gpu/gpu-xid-79-fell-off-bus"Annotation templates only see the labels of the query result, not the rule's own label block, so either write the URL literally as here or build it in the Alertmanager template from .CommonLabels.runbook_id; the lint checks that the two agree. Route one alert per entry where you can. A single ‘GPU unhealthy’ alert that points at a generic page forces the responder to triage before they can act; ten alerts for ten Xid families each pointing at the right entry put the diagnosis in the page itself. Exporter label names (Hostname, gpu) vary by dcgm-exporter version, so copy them from your own series rather than from this example.
The lint that keeps it honest
The lint is what makes this an index rather than a folder. Run it in CI on every change to alert rules or runbooks, and nightly so that review dates expire even when nobody commits.
import datetime as dt, pathlib, sys, yaml
REQUIRED = ["id", "title", "symptoms", "owner", "alerts", "last_reviewed", "review_interval_days"]
BODY_HEADINGS = ["## Confirm", "## Contain", "## Fix", "## Verify", "## Escalate"]
def load_entries(root):
entries = {}
for f in pathlib.Path(root).glob("*.yaml"):
e = yaml.safe_load(f.read_text())
body = f.with_suffix(".md").read_text() if f.with_suffix(".md").exists() else ""
entries[e["id"]] = (e, body)
return entries
def lint(entries, alert_rules, teams, today=None):
today = today or dt.date.today()
errors = []
for rule in alert_rules: # every alert resolves
rid = rule.get("labels", {}).get("runbook_id")
if rid not in entries:
errors.append(f"alert {rule['alert']}: runbook_id {rid!r} not in index")
for rid, (e, body) in entries.items():
missing = [k for k in REQUIRED if k not in e]
if missing:
errors.append(f"{rid}: missing {missing}")
continue
if e["owner"] not in teams:
errors.append(f"{rid}: owner {e.get('owner')!r} is not a current team")
due = e["last_reviewed"] + dt.timedelta(days=e["review_interval_days"])
if due < today:
errors.append(f"{rid}: review overdue since {due}")
errors += [f"{rid}: body lacks {h}" for h in BODY_HEADINGS if h not in body]
return errors
if __name__ == "__main__":
groups = yaml.safe_load(open("alerts.yaml"))["groups"]
rules = [r for g in groups for r in g["rules"]]
errs = lint(load_entries("runbooks"), rules, set(yaml.safe_load(open("teams.yaml"))))
print("\n".join(errs) or "index ok")
sys.exit(1 if errs else 0)Also flag entries listing alerts that no longer exist. Fail hard on unresolved alerts and missing owners; warn on overdue reviews for a week before failing.
Writing an entry that works at 3 a.m.
Write the body for someone competent but unfamiliar with this subsystem, at night. Commands are copy-pasteable, outputs say what good and bad look like, and every destructive step says what it affects. Here is the core of a collective-hang entry for PyTorch training:
## Confirm
- Rank logs show the NCCL watchdog timing out on a collective (look for "Watchdog caught collective
operation timeout" and the op sequence number).
- All ranks stopped advancing at roughly the same step; GPU SM activity dropped to near zero.
## Contain
- Do not kill the job until the dumps below are written; they are the only record of which rank stalled.
## Fix
- The launcher must already set these (check the job spec):
TORCH_NCCL_TRACE_BUFFER_SIZE=2000 # flight recorder ring buffer, in entries
TORCH_NCCL_DUMP_ON_TIMEOUT=1 # write the buffer to disk when the watchdog fires
- Compare the last completed collective per rank; the rank that is behind is the suspect.
- On the suspect node: check dmesg for Xid 74/79, run dcgmi diag -r 3, check NIC error counters.
- Cordon the suspect node and restart from the last checkpoint on the remaining capacity.
## Verify
- Step time back to baseline for 30 minutes; no new watchdog messages.
## Escalate
- Two hangs in 24 hours with no hardware finding: page the training-infra owner with the dumps.Notice the entry depends on configuration that must exist before the incident: if the flight recorder was not enabled, the evidence was never collected. Runbooks expose these preconditions, and the index should list them in a prerequisites field so a nightly check can verify them against job templates.
Worked example: an Xid 79 page
At 02:10 the pager fires GpuXidCritical for one node in a 64-node training cluster. The alert body carries the runbook link, so the responder opens the Xid 79 entry directly. Confirm takes one command: nvidia-smi lists seven GPUs on an eight-GPU node, and the kernel log shows Xid 79. Contain is kubectl cordon then kubectl drain --ignore-daemonsets --delete-emptydir-data on the node; the training job has already failed on the lost rank and the scheduler restarts it from the 01:45 checkpoint on a spare node. Fix is a reboot; the GPU returns, dcgmi diag -r 3 passes, and the node is uncordoned with a label recording the event.
Total time to recovery: 14 minutes, of which 11 were the job reloading its checkpoint. In the review the next morning the team makes two edits: the entry gains a line telling the responder to check whether the same GPU has a prior Xid 79 within 30 days (the second time, it goes to the vendor instead of back into service), and a new alert is added for a node reporting fewer GPUs than its inventory says, with the same runbook_id. Both edits go through CI, which confirms the new alert resolves.
Keeping the index alive
Indexes rot in predictable ways, and each has a counter-measure. Measure coverage: the fraction of paging alerts with a resolving runbook, which should be 100% enforced by the lint. Measure usage: log entry views from alert links, and in each postmortem record whether the runbook was opened and whether it was right. An entry that is never opened during incidents for its alert is either unfindable or untrusted. Measure freshness: the review date, and more importantly whether commands in the entry still run, which you can test by executing the Confirm steps against a healthy node in a scheduled job.
Game days close the loop: inject a fault (kill a rank, fill the KV cache with synthetic long prompts) and have someone who did not write the entry follow it. Where they hesitate, the entry is wrong. Keep the entries few enough to own, and merge near-duplicates.
Trade-offs and failure modes
| Choice | Gain | Cost |
|---|---|---|
| Index in git with lint | reviewed changes, enforceable links | editing feels slower than a wiki during an incident |
| One alert per symptom | diagnosis arrives with the page | more rules to maintain |
| Automation reads safe_actions | auto-cordon in seconds | a wrong entry now acts, not just advises |
| Strict review expiry | stale entries surface | review becomes rubber-stamping if the interval is too short |
The commonest failure is drift: alerts link to a search page and entries cite retired dashboards. The lint and review loop are cheap insurance.
What to do next
- Export your current alert rules and count how many paging alerts carry a runbook link that resolves; that number is your baseline.
- Create the entry schema and move the five most frequent GPU incidents into it first, starting with Xid 79, uncorrectable ECC and NCCL hangs.
- Add the
runbook_idlabel to those alerts and the CI lint to the alert repository. - Wire Xid alerts on change, as described in the DCGM deep dive, and route latency alerts from SLO burn rates to symptom entries.
- Add a runbook-used field to the postmortem template and link the entry from the incident channel bot's opening message.
- Schedule a game day each quarter that follows one entry end to end.