Being on call for an LLM serving fleet is different from being on call for a web service. The worst incidents often return HTTP 200: responses arrive, just slowly, or arrive quickly and are wrong. Capacity is measured in GPU memory and tokens, not CPU and requests, and adding capacity takes minutes because a model must be pulled and loaded. Hardware faults are routine at fleet scale, and a single bad GPU can stall every rank in a tensor-parallel group. A playbook written for stateless web servers sends the responder to the wrong graphs.

This playbook is about the shift itself: what to do in the first fifteen minutes, how to walk the serving stack from the edge to the model in a fixed order, a script that pulls the key signals in one go, the mitigations ranked by speed and reversibility, a worked example of a night page, and how to hand over and run a rotation that people can sustain. It assumes vLLM-style engines with Prometheus metrics and DCGM on the nodes; the ideas transfer to other stacks. The catalogue of individual failure entries lives in the LLM runbook index, and what to tell people during the incident is in incident communication channels.

The first fifteen minutes

The first fifteen minutes decide most incidents. The order below exists to stop the two classic mistakes: diagnosing for forty minutes while users suffer, and mitigating blindly in a way that destroys the evidence or makes things worse.

  1. Acknowledge within the paging SLA so the escalation does not wake a second person for nothing.
  2. Scope it. Which models, which regions, which tenants, what fraction of traffic? Read the alert's own dashboard link before opening anything else.
  3. Find the start time and line it up with changes: deploys, model version or config pushes, traffic shifts, node pool changes, upstream provider incidents. Most incidents start within minutes of a change.
  4. Mitigate before you diagnose if users are hurt and a safe, reversible action exists. Rolling back a change that lines up with the start time is almost always that action.
  5. Declare if it meets the severity bar or you will need help; declaring early is cheap and declaring late is expensive.
  6. Capture evidence before restarting things: a metrics snapshot, the engine logs of one affected replica, nvidia-smi output from an affected node.

A five-layer triage tree

Walk the stack top-down, because higher layers are cheaper to inspect and to fix, and because a fault low in the stack usually shows up higher as well. At each layer, ask whether the signal is abnormal on all replicas or on a subset; a subset points down the stack to specific nodes, everything points up to config, traffic or a shared dependency.

Triage top-down through five layers; mitigate at the highest layer that works1. Edge: gateway, auth, router5xx and 429 by route, upstream errors, retries2. Engine schedulerrequests waiting vs running, TTFT, queue time3. KV cachekv_cache_usage_perc, preemptions, prefix hit rate4. GPU and nodeXid events, DCGM health, missing GPUs, NCCL hangs5. Model outputcanary answers, refusal and length drift, templateAsk first:scope? since when? what changed?Mitigate:shed, route, roll back, scale
The five layers of an LLM serving stack and the signals that clear or implicate each one. Scope and recent changes are checked before the walk starts.
SymptomCheckUsual cause
TTFT up, inter-token latency flatvllm:num_requests_waiting risingqueueing: traffic above capacity, or a replica dropped out
TTFT and inter-token latency both uppreemption rate, KV usage near 1.0KV pressure: longer prompts or outputs, batch limits raised
Errors from a subset of replicasnode of each failing pod, Xid events, GPU counthardware: e.g. Xid 79, a GPU fallen off the bus
All replicas time out, GPUs busy, no tokensengine logs, NCCL watchdoghung collective in a tensor-parallel group
Fast 200s, wrong or odd answerscanary prompts, output length, refusal ratechat template, tokenizer, quantisation or model version drift
Token spend jumps, latency normaloutput tokens per request, retries per requestclient retry loop, prompt change, max_tokens removed

Two signals are worth special attention. A rising waiting count with a flat running count means the scheduler cannot admit more work, which is almost always KV memory, not compute. And preemptions in vLLM evict a running request's KV blocks and recompute them later, so a burst of preemptions converts memory pressure into extra prefill work and latency for everyone; see paged KV cache for the mechanism.

A triage script instead of hand-typed queries

During an incident, nobody should hand-type PromQL. Keep a triage script in the on-call repository that pulls the layer signals for one model and prints them next to their values from a day earlier. It uses only the Prometheus HTTP query API.

import sys
import time
import requests

PROM = "http://prometheus.monitoring:9090"   # your Prometheus address

QUERIES = {
    "running":      'sum(vllm:num_requests_running{model_name="$M"})',
    "waiting":      'sum(vllm:num_requests_waiting{model_name="$M"})',
    "kv_usage_max": 'max(vllm:kv_cache_usage_perc{model_name="$M"})',
    "preempt_rate": 'sum(rate(vllm:num_preemptions_total{model_name="$M"}[5m]))',
    "ttft_p95":     'histogram_quantile(0.95, sum by (le) (rate('
                    'vllm:time_to_first_token_seconds_bucket{model_name="$M"}[5m])))',
    "itl_p95":      'histogram_quantile(0.95, sum by (le) (rate('
                    'vllm:inter_token_latency_seconds_bucket{model_name="$M"}[5m])))',
    "xid_gpus_1h":  'count(changes(DCGM_FI_DEV_XID_ERRORS[1h]) > 0)',
}

def query(expr, at=None):
    params = {"query": expr}
    if at is not None:
        params["time"] = at                  # evaluate the instant query in the past
    r = requests.get(f"{PROM}/api/v1/query", params=params, timeout=10)
    r.raise_for_status()
    result = r.json()["data"]["result"]
    return float(result[0]["value"][1]) if result else float("nan")

def main(model):
    yesterday = time.time() - 86400
    print(f"{'signal':14} {'now':>10} {'1d ago':>10}")
    for name, q in QUERIES.items():
        expr = q.replace("$M", model)
        now, before = query(expr), query(expr, at=yesterday)
        print(f"{name:14} {now:10.3f} {before:10.3f}")

if __name__ == "__main__":
    main(sys.argv[1])

The comparison uses the time parameter of the instant-query endpoint rather than PromQL offset, so the same expression, histograms included, is evaluated unchanged at both moments. A missing series prints nan rather than zero, which is deliberate: absence is itself a finding. Check every metric name against your engine version: names change between releases, and a panel built on a series that no longer exists shows no data rather than an error, which is easy to misread at night. DCGM_FI_DEV_XID_ERRORS holds the last Xid seen, so the script counts GPUs whose value changed in the last hour rather than any non-zero value; a repeat of the same Xid code does not change it, so confirm on the node with nvidia-smi or the kernel log.

The mitigation ladder

Mitigations differ in how fast they act and how easily they are undone. Prefer the top of this ladder, and climb down only when the higher rungs do not apply.

ActionTime to effectReversible?Use when
Roll back the last changeminutesyesstart time lines up with a deploy or config push
Shift traffic to another pool or regionseconds to minutesyesone pool or region is impaired
Cordon and drain bad nodesminutesyeserrors follow specific nodes or GPUs
Shed load: tighter rate limits, cap max_tokenssecondsyes, but users noticedemand exceeds capacity
Fall back to a smaller or older modelminutesyes, quality dropsprimary model cannot serve at all
Scale out replicasmany minutes: pull and load weightsyessustained demand growth, not a spike
Change engine limits (max_num_seqs, memory fraction)restart per replicariskyonly with a tested value

Scaling out sits low on the ladder for a reason: a replica of a large model must pull tens of gigabytes of weights and warm up before it serves, so it rarely helps in the first fifteen minutes, and a rushed scale-out during a hardware incident can land replicas on the same bad nodes. Changing engine limits such as --max-num-seqs or --gpu-memory-utilization during an incident restarts replicas and can trade a latency problem for out-of-memory crashes; do it only with values already proven in load tests.

Worked example: a 02:10 TTFT page

At 02:10 the burn-rate alert for the chat model's TTFT SLO fires (the alert maths is in SLO burn rate alerts). Error rate is zero. The responder acknowledges at 02:12 and runs the triage script: waiting requests are 340 against a usual 5, running requests are flat, p95 TTFT is 9.8 s against 0.7 s a day earlier, inter-token latency is only slightly up, and xid_gpus_1h is 1.

Scope comes next. The waiting count is high on every replica, which would point up the stack, but the replica count is 15 rather than 16. The Kubernetes events show one pod in a crash loop on a node whose DCGM exporter reports Xid 79, and nvidia-smi on that node lists seven GPUs. One tensor-parallel replica died, its traffic spread over the other fifteen, and those were already near their KV limit, so the scheduler queues instead of admitting.

The responder cordons and drains the node at 02:18 (runbook entry for Xid 79), temporarily lowers the per-tenant rate limit for the two largest batch tenants to cut queued demand, and posts a status update. A replacement replica lands on a healthy node and starts serving at 02:31, after loading weights. The queue drains by 02:36, the rate limits are restored at 02:45, and the incident closes. The handoff note records the node for the hardware queue and one follow-up: the fleet had no headroom to lose a single replica at night, which is a capacity bug, not bad luck.

Handoffs and a rotation people can sustain

A playbook only works if the people running it are rested and informed. Hand over with a short written note at every shift change, not a verbal summary: open incidents and their state, mitigations still in place that must be undone (rate limits, cordoned nodes, traffic shifts, a fallback model), alerts that fired and were judged noise, and changes scheduled for the next shift. Unreverted mitigations are the most common source of the next incident.

Measure the rotation like a service. Track pages per shift, pages outside working hours, and the fraction of pages that needed action; a page that needed no action is a bug in an alert. Keep the rotation large enough that each person is on call at most one week in four or five, follow the sun across time zones if you can, and give the secondary a real role: taking over communication so the primary can work. Every incident with a novel cause should end with a new or corrected runbook entry and a test of the triage script against it.

Failure modes

  • Treating 200s as healthy. Latency and quality incidents return success codes; alert on TTFT, inter-token latency and canary correctness, not only error rate.
  • Restarting before capturing evidence. A restart clears the engine state, the hung collective and the logs that explain it.
  • Scaling into a hardware fault. New replicas scheduled onto the same bad node pool fail the same way and burn the time you had.
  • Dashboards pinned to renamed metrics. An empty panel is easy to read as quiet. Alert on absent series.
  • Leaving mitigations in place. A fallback model or a tight rate limit left on for a week becomes a silent quality or revenue incident.

Trade-offs

Mitigating early trades diagnostic certainty for user impact, and is usually right when the action is reversible. Automated remediation, such as cordoning on Xid 79, shortens incidents but needs guard rails so a buggy detector cannot drain the fleet: cap how many nodes automation may cordon per hour. Headroom costs GPU hours every night but is what lets the fleet lose a replica without paging anyone; the worked example is the classic case where a few percent more capacity would have prevented the incident. See LLM serving reliability for the design side of these choices.

What to do next

  1. Write the first-fifteen-minutes list into your paging tool's alert template.
  2. Build the triage script for your stack, verify every metric name against your engine version, and run it during a quiet hour.
  3. Add an absent-series alert for each metric on the on-call dashboard.
  4. Rank your mitigations into a ladder with time-to-effect, and rehearse the top three.
  5. Check that the fleet can lose its largest replica at peak and still meet the SLO.
  6. Adopt a written handoff template that lists active mitigations.
  7. Review pages per shift and actionable fraction monthly, and delete or fix noisy alerts.
Key takeaway: LLM serving incidents often look like slowness or wrong answers rather than errors, so the responder needs a fixed order: scope it, line it up with changes, mitigate reversibly, then walk the stack from the edge through the scheduler, KV cache and GPUs to the model output. Script the signals, rank mitigations by speed and reversibility, keep capacity headroom so one bad GPU does not page anyone, and hand over every active mitigation in writing.