A gameday is a scheduled rehearsal in which you deliberately break part of a production-like system, watch whether detection, automation and people respond the way you believe they will, and write down every place they did not. For an LLM platform the stakes are unusual: a single replica holds tens of gigabytes of weights, a cold start can take minutes rather than seconds, a training step involves hundreds of ranks that all stall if one stalls, and the most common user-visible failure is not an error at all but a slow first token. Those properties mean recovery paths that look fine on a diagram are frequently untested in practice.
This article explains how to design and run gamedays for GPU-backed LLM serving and training: the failure catalogue worth rehearsing, how to express a steady-state hypothesis as code, how to inject each fault safely, a worked serving scenario with the capacity arithmetic, a training scenario built around NCCL timeouts and checkpoint restore, and how to score the day so it changes something. It complements the site's articles on GPU hardware faults and running an LLM incident channel, which cover what fails and how to communicate; this one is about proving your response before you need it.
What a gameday is, and what it is not
Four ingredients separate a gameday from someone pressing buttons in staging. First, a steady-state hypothesis: a measurable statement of normal, such as "p95 time to first token stays under 1.5 s and the error ratio under 0.5% while one replica is lost". Second, a bounded blast radius: which nodes, which tenants, what fraction of traffic, for how long. Third, abort criteria decided in advance, with one named person holding the abort command. Fourth, a scorecard that records time to detect, time to mitigate, whether the right alert fired, and whether the runbook was correct.
What a gameday is not: a load test (that measures capacity, not recovery), a benchmark, or an excuse to discover that nobody knows how to roll back. If you cannot state what you expect to happen, you are not ready to inject the fault; write the expectation first, because the gap between expectation and outcome is the whole product of the exercise.
A failure catalogue for LLM fleets
Pick faults by asking what has actually hurt you or similar fleets, then rank by likelihood times cost of a bad response. A starting catalogue for LLM systems:
| Fault | How to inject | What should detect it | Expected response |
|---|---|---|---|
| Serving pod dies | kubectl delete pod on one replica | Router health checks, replica count | Traffic shifts; replacement ready within cold-start budget |
| GPU node lost | kubectl drain or power off a test node | Node NotReady, DCGM exporter gap | Pods rescheduled; capacity alert if headroom is gone |
| GPU reports errors | dcgmi test --inject on a monitored GPU | DCGM health watch, policy alert | Node cordoned by automation, not by a human |
| KV-cache exhaustion | Replay long-context traffic at 2x share | Queue depth, preemption counter | Admission control sheds or queues; no OOM crash loop |
| Slow upstream model | Add latency with a fault proxy or tc netem | TTFT histogram, timeout counters | Fallback model engaged; retries do not amplify load |
| Training rank hangs | kill -STOP one rank's process | NCCL watchdog timeout | Job aborts, restarts from checkpoint, loses bounded work |
| Checkpoint store slow | Throttle the bucket or NFS path | Checkpoint duration metric | Training continues; alert before the next save overlaps |
Notice the third column. Many gamedays find that the fault was handled but nothing told anyone, which is fine once and fatal when the automation itself breaks. Each row should test both the remedy and the signal.
Architecture of the exercise
Keep the injectors, the checker and the abort switch outside the system under test. If the checker queries the same Prometheus that runs on the GPU nodes you are draining, you may lose your eyes at the moment you need them. The runner holds a clock, logs every injection with a timestamp, and the scorecard is assembled from those timestamps and alert history, not from memory in the retro.
Steady-state hypotheses as code
Write the hypothesis as code so it is evaluated identically every 15 seconds, and so the abort decision is mechanical. The sketch below polls Prometheus. The vllm: names follow vLLM's documented metrics, but they have changed between releases, so read them from your build's /metrics endpoint; router_requests_total stands in for your gateway's own request counter.
import time, requests
PROM = "http://prometheus.ops:9090/api/v1/query"
CHECKS = {
# name: (PromQL, upper limit for steady state)
"ttft_p95_s": ('histogram_quantile(0.95, sum(rate(vllm:time_to_first_token_seconds_bucket[1m])) by (le))', 1.5),
"error_ratio": ('sum(rate(router_requests_total{code=~"5.."}[1m])) / sum(rate(router_requests_total[1m]))', 0.005),
"queue_depth": ('sum(vllm:num_requests_waiting)', 200),
}
ABORT = {"ttft_p95_s": 4.0, "error_ratio": 0.03} # hard limits: stop the experiment
def q(expr):
r = requests.get(PROM, params={"query": expr}, timeout=5).json()
res = r["data"]["result"]
return float(res[0]["value"][1]) if res else float("nan")
def tick(log):
row = {"t": time.time()}
for name, (expr, thr) in CHECKS.items():
v = q(expr)
row[name] = v
row[name + "_ok"] = v < thr # NaN compares False: missing data fails the check
log.append(row)
return any(row[k] >= lim for k, lim in ABORT.items())
def run(log, duration_s=1800):
end = time.time() + duration_s
while time.time() < end:
if tick(log):
print("ABORT threshold crossed; rolling back injection")
return "aborted"
time.sleep(15)
return "completed"Two details matter. Missing data counts as a failed check, because "the exporter died" is a finding, not a pass. And the abort limits are looser than the hypothesis: you want to observe a hypothesis breach and learn from it, but stop before customers are seriously hurt.
Worked example: losing a serving replica
Scenario: an internal chat service runs 8 replicas of a 70B-parameter model, each replica on one 8-GPU node with tensor parallelism of 8. At peak each replica comfortably serves about 30 concurrent streams, and peak demand is around 180 streams, so utilisation is 180 / 240 = 75%. Hypothesis: losing one replica keeps p95 TTFT under 1.5 s, because 7 replicas offer 210 streams of capacity.
The arithmetic already hints at trouble. With 7 replicas the fleet runs at 180 / 210 = 86%, and queueing delay grows sharply as utilisation approaches 1, so the TTFT tail will stretch even though nothing is technically overloaded. The replacement replica must also load its weights: at 16-bit precision 70 billion parameters is about 140 GB. If the node pulls that from object storage at an effective 1 GB/s, loading alone takes over two minutes before CUDA graphs are captured and the engine reports ready. Those are illustrative numbers; measure your own cold start, because it is the single most important figure in this scenario.
Run it at a quiet-but-real hour with 25% of traffic mirrored or routed to the gameday pool. Inject by deleting one pod and record timestamps for: pod gone, router marks it unhealthy, first request retried, replacement scheduled, weights loaded, replacement passes readiness, TTFT back under target. Typical findings in exercises like this: the router kept sending requests to the dead endpoint for a full health-check interval; retries doubled load on survivors; readiness passed before CUDA graph capture finished, so the first real requests on the new pod were slow; and the capacity alert was configured on GPU count rather than serving headroom, so it never fired. Each is a concrete fix.
Training: a hung rank and checkpoint restore
Training gamedays rehearse a different promise: that a stuck or dead rank costs you a bounded amount of work rather than hours of a silently hung job. In PyTorch the collective timeout is set when the process group is created, and the NCCL watchdog can dump a flight-recorder trace of recent collectives when it fires, which tells you which rank stopped participating.
# launch: the elastic agent restarts the whole job on failure, up to 3 times
# TORCH_NCCL_TRACE_BUFFER_SIZE=2000 TORCH_NCCL_DUMP_ON_TIMEOUT=1 \
# torchrun --nnodes=16 --nproc-per-node=8 --max-restarts=3 \
# --rdzv-backend=c10d --rdzv-endpoint=$HEAD:29400 train.py
import datetime, torch.distributed as dist
dist.init_process_group(
backend="nccl",
timeout=datetime.timedelta(minutes=10), # how long a collective may wait before the job is torn down
)
step = load_latest_checkpoint() # must be idempotent and verified
while step < total_steps:
train_one_step()
step += 1
if step % ckpt_every == 0:
save_checkpoint_async(step) # measure duration; alert if it approaches ckpt intervalThe injection is simple: on one node, send kill -STOP to one rank so it stops without exiting. The other ranks block in their next all-reduce. Hypothesis: the watchdog fires within the configured timeout, the flight-recorder dump names the stuck rank, the job restarts and resumes from the last checkpoint, and total lost time is under timeout plus restart plus half the checkpoint interval on average. If you checkpoint every 30 minutes, set a 10-minute timeout and restarts take 8 minutes, expect roughly 33 minutes lost per incident. Gamedays regularly reveal that the timeout was left at a large default, that restart required a human to rejoin rendezvous, or that the latest checkpoint was incomplete because the save was still in flight. See checkpointing in depth for choosing the interval.
Testing the GPU fault pipeline with DCGM injection
Real GPU faults are hard to produce on demand, but the pipeline that reacts to them can be tested. NVIDIA's DCGM includes an error-injection framework: you start the host engine, enable health watches or policies, and inject a field value with dcgmi test --inject. The examples in NVIDIA's error-injection guide are:
dcgmi health -c # check current health status
dcgmi test --inject --gpuid 0 -f 202 -v 99999 # inject a value for field 202 on GPU 0
dcgmi test --inject --gpuid 0 -f 319 -v 4 # inject a value for field 319 on GPU 0Look up what each field ID means in the field-identifier header for your DCGM version rather than copying numbers from blog posts, including this one. Treat the exercise as a test of your detection and remediation path: did the health watch flip, did your node-problem automation cordon the node, did the scheduler stop placing new GPU pods there, and did a ticket open with the right labels? Run it on a node you have already drained. The DCGM article covers the exporter and health-watch setup this depends on.
Running the day
A useful running order for a two-hour session:
- T-7 days: circulate the plan, hypotheses, blast radius, abort criteria and the named abort owner.
- T-1 day: verify the abort path works by injecting and rolling back the smallest fault in staging.
- T-0: announce start in the incident channel; freeze unrelated deploys to the target pool.
- Inject one fault at a time. Never stack faults until each has been run alone and passed.
- Let on-call respond as they would for real; the facilitator answers only questions about scope.
- Stop at the time box or on abort, restore, confirm steady state for 15 minutes, announce end.
- Retro within 48 hours while timestamps and memories are fresh.
Rotate who is on call during gamedays. The point is to test the runbook and the alerts, and an expert who knows the system by heart hides exactly the gaps a newer engineer will hit at 03:00.
Scoring and follow-through
| Measure | Source | Why it matters |
|---|---|---|
| Time to detect | Injection log vs first alert | Automation cannot act on what it cannot see |
| Time to mitigate | Injection log vs steady state restored | The number users feel |
| Right alert, right person | Pager history | Wrong-team pages add minutes |
| Runbook accuracy | Responder notes | Stale commands are common after upgrades |
| Hypothesis held? | Checker log | Yes or no, with the breach curve attached |
| Lost training work | Step counter before and after | Validates checkpoint interval maths |
Each gap becomes a ticket with an owner and a date, and the same scenario is rerun after the fix. A gameday whose findings are not rerun is a story, not a control.
Failure modes of gamedays themselves
- The gameday becomes the incident. Blast radius was not enforced, or the abort path was never tested. Always rehearse rollback first.
- Staging lies. Staging with two replicas and no real traffic does not exhibit queueing collapse or retry storms. Use a production pool with bounded traffic once the basics pass.
- Checker shares fate with the target. Metrics scraped from the drained node vanish, and missing data is read as healthy. Treat gaps as failures.
- Heroics mask gaps. A senior engineer fixes it from memory; the runbook stays wrong.
- Scenarios fossilise. The fleet changes model sizes and engines; rerun the cold-start measurement after every major upgrade.
Trade-offs
Gamedays cost GPU hours, engineer time and some customer risk. The trade is worth it where recovery is slow and rare, which describes most LLM failure paths: you cannot learn weight-loading time or rendezvous behaviour from unit tests. Continuous automated chaos (randomly killing pods daily) gives better coverage for stateless layers, but for expensive operations like full-node loss or training restarts a scheduled, observed exercise is safer and teaches the humans as well as the software. Many teams do both: automated pod kills in serving, quarterly facilitated gamedays for the big scenarios. For multi-region failover patterns, see multi-region LLM serving.
What to do next
- List your five most likely LLM failure modes and write one steady-state hypothesis for each.
- Measure cold start for your largest model end to end: pod scheduled to readiness to first fast token.
- Build the checker script against your real metric names and confirm missing data fails a check.
- Rehearse the abort path in staging, then run a single pod-deletion gameday on 25% of traffic.
- Run a
kill -STOPrank test on a small training job and record lost work against your estimate. - Inject a DCGM error on a drained node and confirm automation cordons it and a ticket opens.
- File every gap with an owner, fix it, and rerun the same scenario within a month.