A gameday is a scheduled rehearsal in which you deliberately break part of a production-like system, watch whether detection, automation and people respond the way you believe they will, and write down every place they did not. For an LLM platform the stakes are unusual: a single replica holds tens of gigabytes of weights, a cold start can take minutes rather than seconds, a training step involves hundreds of ranks that all stall if one stalls, and the most common user-visible failure is not an error at all but a slow first token. Those properties mean recovery paths that look fine on a diagram are frequently untested in practice.

This article explains how to design and run gamedays for GPU-backed LLM serving and training: the failure catalogue worth rehearsing, how to express a steady-state hypothesis as code, how to inject each fault safely, a worked serving scenario with the capacity arithmetic, a training scenario built around NCCL timeouts and checkpoint restore, and how to score the day so it changes something. It complements the site's articles on GPU hardware faults and running an LLM incident channel, which cover what fails and how to communicate; this one is about proving your response before you need it.

What a gameday is, and what it is not

Four ingredients separate a gameday from someone pressing buttons in staging. First, a steady-state hypothesis: a measurable statement of normal, such as "p95 time to first token stays under 1.5 s and the error ratio under 0.5% while one replica is lost". Second, a bounded blast radius: which nodes, which tenants, what fraction of traffic, for how long. Third, abort criteria decided in advance, with one named person holding the abort command. Fourth, a scorecard that records time to detect, time to mitigate, whether the right alert fired, and whether the runbook was correct.

What a gameday is not: a load test (that measures capacity, not recovery), a benchmark, or an excuse to discover that nobody knows how to roll back. If you cannot state what you expect to happen, you are not ready to inject the fault; write the expectation first, because the gap between expectation and outcome is the whole product of the exercise.

A failure catalogue for LLM fleets

Pick faults by asking what has actually hurt you or similar fleets, then rank by likelihood times cost of a bad response. A starting catalogue for LLM systems:

FaultHow to injectWhat should detect itExpected response
Serving pod dieskubectl delete pod on one replicaRouter health checks, replica countTraffic shifts; replacement ready within cold-start budget
GPU node lostkubectl drain or power off a test nodeNode NotReady, DCGM exporter gapPods rescheduled; capacity alert if headroom is gone
GPU reports errorsdcgmi test --inject on a monitored GPUDCGM health watch, policy alertNode cordoned by automation, not by a human
KV-cache exhaustionReplay long-context traffic at 2x shareQueue depth, preemption counterAdmission control sheds or queues; no OOM crash loop
Slow upstream modelAdd latency with a fault proxy or tc netemTTFT histogram, timeout countersFallback model engaged; retries do not amplify load
Training rank hangskill -STOP one rank's processNCCL watchdog timeoutJob aborts, restarts from checkpoint, loses bounded work
Checkpoint store slowThrottle the bucket or NFS pathCheckpoint duration metricTraining continues; alert before the next save overlaps

Notice the third column. Many gamedays find that the fault was handled but nothing told anyone, which is fine once and fatal when the automation itself breaks. Each row should test both the remedy and the signal.

Architecture of the exercise

A gameday control loop for an LLM fleet: inject, observe, compare to hypothesis, abort if neededGameday runnerplan, clock, scorecardAbort switchone command, one ownerFault injectorskubectl, kill, tc, dcgmiRouter / gatewayadmission, retries, fallbackServing replicasvLLM pods on GPU nodesTraining jobtorchrun ranks, NCCLTelemetryDCGM, engine, router metricsSteady-state checkerSLO queries every 15 sAlerting + on-calldid a human get paged?faultmetricsbreach: abortThe checker, not a person watching a dashboard, decides whether the hypothesis held; humans decide what to change afterwards.
Gameday control loop. Injection and observation are separate systems so that a broken target cannot hide its own breakage.

Keep the injectors, the checker and the abort switch outside the system under test. If the checker queries the same Prometheus that runs on the GPU nodes you are draining, you may lose your eyes at the moment you need them. The runner holds a clock, logs every injection with a timestamp, and the scorecard is assembled from those timestamps and alert history, not from memory in the retro.

Steady-state hypotheses as code

Write the hypothesis as code so it is evaluated identically every 15 seconds, and so the abort decision is mechanical. The sketch below polls Prometheus. The vllm: names follow vLLM's documented metrics, but they have changed between releases, so read them from your build's /metrics endpoint; router_requests_total stands in for your gateway's own request counter.

import time, requests

PROM = "http://prometheus.ops:9090/api/v1/query"
CHECKS = {
    # name: (PromQL, upper limit for steady state)
    "ttft_p95_s": ('histogram_quantile(0.95, sum(rate(vllm:time_to_first_token_seconds_bucket[1m])) by (le))', 1.5),
    "error_ratio": ('sum(rate(router_requests_total{code=~"5.."}[1m])) / sum(rate(router_requests_total[1m]))', 0.005),
    "queue_depth": ('sum(vllm:num_requests_waiting)', 200),
}
ABORT = {"ttft_p95_s": 4.0, "error_ratio": 0.03}   # hard limits: stop the experiment

def q(expr):
    r = requests.get(PROM, params={"query": expr}, timeout=5).json()
    res = r["data"]["result"]
    return float(res[0]["value"][1]) if res else float("nan")

def tick(log):
    row = {"t": time.time()}
    for name, (expr, thr) in CHECKS.items():
        v = q(expr)
        row[name] = v
        row[name + "_ok"] = v < thr          # NaN compares False: missing data fails the check
    log.append(row)
    return any(row[k] >= lim for k, lim in ABORT.items())

def run(log, duration_s=1800):
    end = time.time() + duration_s
    while time.time() < end:
        if tick(log):
            print("ABORT threshold crossed; rolling back injection")
            return "aborted"
        time.sleep(15)
    return "completed"

Two details matter. Missing data counts as a failed check, because "the exporter died" is a finding, not a pass. And the abort limits are looser than the hypothesis: you want to observe a hypothesis breach and learn from it, but stop before customers are seriously hurt.

Worked example: losing a serving replica

Scenario: an internal chat service runs 8 replicas of a 70B-parameter model, each replica on one 8-GPU node with tensor parallelism of 8. At peak each replica comfortably serves about 30 concurrent streams, and peak demand is around 180 streams, so utilisation is 180 / 240 = 75%. Hypothesis: losing one replica keeps p95 TTFT under 1.5 s, because 7 replicas offer 210 streams of capacity.

The arithmetic already hints at trouble. With 7 replicas the fleet runs at 180 / 210 = 86%, and queueing delay grows sharply as utilisation approaches 1, so the TTFT tail will stretch even though nothing is technically overloaded. The replacement replica must also load its weights: at 16-bit precision 70 billion parameters is about 140 GB. If the node pulls that from object storage at an effective 1 GB/s, loading alone takes over two minutes before CUDA graphs are captured and the engine reports ready. Those are illustrative numbers; measure your own cold start, because it is the single most important figure in this scenario.

Run it at a quiet-but-real hour with 25% of traffic mirrored or routed to the gameday pool. Inject by deleting one pod and record timestamps for: pod gone, router marks it unhealthy, first request retried, replacement scheduled, weights loaded, replacement passes readiness, TTFT back under target. Typical findings in exercises like this: the router kept sending requests to the dead endpoint for a full health-check interval; retries doubled load on survivors; readiness passed before CUDA graph capture finished, so the first real requests on the new pod were slow; and the capacity alert was configured on GPU count rather than serving headroom, so it never fired. Each is a concrete fix.

Training: a hung rank and checkpoint restore

Training gamedays rehearse a different promise: that a stuck or dead rank costs you a bounded amount of work rather than hours of a silently hung job. In PyTorch the collective timeout is set when the process group is created, and the NCCL watchdog can dump a flight-recorder trace of recent collectives when it fires, which tells you which rank stopped participating.

# launch: the elastic agent restarts the whole job on failure, up to 3 times
#   TORCH_NCCL_TRACE_BUFFER_SIZE=2000 TORCH_NCCL_DUMP_ON_TIMEOUT=1 \
#   torchrun --nnodes=16 --nproc-per-node=8 --max-restarts=3 \
#            --rdzv-backend=c10d --rdzv-endpoint=$HEAD:29400 train.py
import datetime, torch.distributed as dist

dist.init_process_group(
    backend="nccl",
    timeout=datetime.timedelta(minutes=10),   # how long a collective may wait before the job is torn down
)
step = load_latest_checkpoint()               # must be idempotent and verified
while step < total_steps:
    train_one_step()
    step += 1
    if step % ckpt_every == 0:
        save_checkpoint_async(step)           # measure duration; alert if it approaches ckpt interval

The injection is simple: on one node, send kill -STOP to one rank so it stops without exiting. The other ranks block in their next all-reduce. Hypothesis: the watchdog fires within the configured timeout, the flight-recorder dump names the stuck rank, the job restarts and resumes from the last checkpoint, and total lost time is under timeout plus restart plus half the checkpoint interval on average. If you checkpoint every 30 minutes, set a 10-minute timeout and restarts take 8 minutes, expect roughly 33 minutes lost per incident. Gamedays regularly reveal that the timeout was left at a large default, that restart required a human to rejoin rendezvous, or that the latest checkpoint was incomplete because the save was still in flight. See checkpointing in depth for choosing the interval.

Testing the GPU fault pipeline with DCGM injection

Real GPU faults are hard to produce on demand, but the pipeline that reacts to them can be tested. NVIDIA's DCGM includes an error-injection framework: you start the host engine, enable health watches or policies, and inject a field value with dcgmi test --inject. The examples in NVIDIA's error-injection guide are:

dcgmi health -c                                  # check current health status
dcgmi test --inject --gpuid 0 -f 202 -v 99999    # inject a value for field 202 on GPU 0
dcgmi test --inject --gpuid 0 -f 319 -v 4        # inject a value for field 319 on GPU 0

Look up what each field ID means in the field-identifier header for your DCGM version rather than copying numbers from blog posts, including this one. Treat the exercise as a test of your detection and remediation path: did the health watch flip, did your node-problem automation cordon the node, did the scheduler stop placing new GPU pods there, and did a ticket open with the right labels? Run it on a node you have already drained. The DCGM article covers the exporter and health-watch setup this depends on.

Running the day

A useful running order for a two-hour session:

  1. T-7 days: circulate the plan, hypotheses, blast radius, abort criteria and the named abort owner.
  2. T-1 day: verify the abort path works by injecting and rolling back the smallest fault in staging.
  3. T-0: announce start in the incident channel; freeze unrelated deploys to the target pool.
  4. Inject one fault at a time. Never stack faults until each has been run alone and passed.
  5. Let on-call respond as they would for real; the facilitator answers only questions about scope.
  6. Stop at the time box or on abort, restore, confirm steady state for 15 minutes, announce end.
  7. Retro within 48 hours while timestamps and memories are fresh.

Rotate who is on call during gamedays. The point is to test the runbook and the alerts, and an expert who knows the system by heart hides exactly the gaps a newer engineer will hit at 03:00.

Scoring and follow-through

MeasureSourceWhy it matters
Time to detectInjection log vs first alertAutomation cannot act on what it cannot see
Time to mitigateInjection log vs steady state restoredThe number users feel
Right alert, right personPager historyWrong-team pages add minutes
Runbook accuracyResponder notesStale commands are common after upgrades
Hypothesis held?Checker logYes or no, with the breach curve attached
Lost training workStep counter before and afterValidates checkpoint interval maths

Each gap becomes a ticket with an owner and a date, and the same scenario is rerun after the fix. A gameday whose findings are not rerun is a story, not a control.

Failure modes of gamedays themselves

  • The gameday becomes the incident. Blast radius was not enforced, or the abort path was never tested. Always rehearse rollback first.
  • Staging lies. Staging with two replicas and no real traffic does not exhibit queueing collapse or retry storms. Use a production pool with bounded traffic once the basics pass.
  • Checker shares fate with the target. Metrics scraped from the drained node vanish, and missing data is read as healthy. Treat gaps as failures.
  • Heroics mask gaps. A senior engineer fixes it from memory; the runbook stays wrong.
  • Scenarios fossilise. The fleet changes model sizes and engines; rerun the cold-start measurement after every major upgrade.

Trade-offs

Gamedays cost GPU hours, engineer time and some customer risk. The trade is worth it where recovery is slow and rare, which describes most LLM failure paths: you cannot learn weight-loading time or rendezvous behaviour from unit tests. Continuous automated chaos (randomly killing pods daily) gives better coverage for stateless layers, but for expensive operations like full-node loss or training restarts a scheduled, observed exercise is safer and teaches the humans as well as the software. Many teams do both: automated pod kills in serving, quarterly facilitated gamedays for the big scenarios. For multi-region failover patterns, see multi-region LLM serving.

What to do next

  1. List your five most likely LLM failure modes and write one steady-state hypothesis for each.
  2. Measure cold start for your largest model end to end: pod scheduled to readiness to first fast token.
  3. Build the checker script against your real metric names and confirm missing data fails a check.
  4. Rehearse the abort path in staging, then run a single pod-deletion gameday on 25% of traffic.
  5. Run a kill -STOP rank test on a small training job and record lost work against your estimate.
  6. Inject a DCGM error on a drained node and confirm automation cordons it and a ticket opens.
  7. File every gap with an owner, fix it, and rerun the same scenario within a month.
Key takeaway: An LLM gameday turns recovery assumptions into measurements. State a steady-state hypothesis as code, inject one bounded fault at a time from outside the target, let the real on-call respond, and score detection, mitigation and runbook accuracy. Cold start, retry behaviour, NCCL timeouts and checkpoint freshness are where LLM fleets most often surprise their owners.