LLM serving fails differently from ordinary web serving. A request can run for a minute and stream thousands of tokens, so a crash does not lose one request, it loses every stream on that replica. Capacity is not CPU but GPU memory for the KV cache, which fills with request length nobody controls in advance. A GPU can fault while its process stays alive and keeps answering health checks. And a model replica takes minutes, not seconds, to come back, because weights have to be loaded and kernels warmed.
This article is a map of those failure domains and the controls that contain each one, written for an engine such as vLLM behind a load balancer or gateway. It covers health checks that tell the truth, KV pressure and preemption, per-request limits and timeouts, retries that do not amplify an outage, graceful drains, GPU and tensor-parallel faults, and a worked error-budget calculation that shows which failure actually dominates. Alerting on the resulting SLIs is covered in LLM SLO burn-rate alerts, so this article concentrates on the mechanisms that keep those SLIs green.
A map of failure domains
Reliability work starts with the question: when this breaks, how much breaks with it? A bad request should fail alone. An engine overload should shed load, not stall every request. A process crash should cost one replica's in-flight streams and nothing more. A sick GPU should be drained and cordoned before it corrupts more work. A fleet-level loss should be absorbed by spare capacity or another region. The rest of this article takes the domains in that order.
Health checks that tell the truth
vLLM's OpenAI-compatible server exposes GET /health. Per the vLLM API documentation, it returns 200 when the engine is healthy and 503 when it catches an EngineDeadError, the error raised when the engine core process has died and cannot recover. That makes it a good liveness check and a poor proof of service: a live process does not prove the GPU can still run kernels, and after a GPU fault the process may stay up while inference no longer works.
The fix is a deep probe: a tiny real generation, run by a sidecar or the gateway rather than the kubelet, with a short timeout and a failure threshold. It checks the whole path the user depends on: tokenizer, scheduler, GPU kernels, detokenizer.
import time, requests
def deep_probe(base_url, model, timeout_s=10.0):
# One short completion. Fails on HTTP errors, timeouts and empty output.
t0 = time.monotonic()
r = requests.post(f"{base_url}/v1/completions", timeout=timeout_s, json={
"model": model, "prompt": "ping", "max_tokens": 1, "temperature": 0,
})
r.raise_for_status()
if r.json()["usage"]["completion_tokens"] < 1:
raise RuntimeError("empty completion")
return time.monotonic() - t0
# Gateway loop: three consecutive failures -> stop routing, drain, then restart.Wire the three probes to three actions. Liveness on /health restarts the container. Readiness removes the replica from routing without restarting it, and should fail while the model loads, while draining, and when the waiting queue exceeds a threshold. The deep probe, after consecutive failures, triggers a drain and restart and records the node, because repeated deep-probe failures on the same node usually mean hardware. Give the liveness probe a generous startup allowance: large models take minutes to load, and a probe that kills a loading pod creates a crash loop that looks like a bad image.
KV pressure and preemption storms
The engine's scarce resource is KV-cache memory. Each running sequence holds blocks proportional to its current length, and the scheduler admits work while blocks remain. When running sequences grow and the cache fills, vLLM preempts some of them, freeing their blocks and recomputing them later. A few preemptions are normal. A storm, where the engine repeatedly preempts and recomputes, burns GPU time on repeated prefill and makes latency for everyone climb, while throughput falls. Sizing the cache in the first place is covered in vLLM on GPU.
vLLM exposes the signals on /metrics. In the V1 metrics documentation the relevant names are vllm:kv_cache_usage_perc (a gauge, 1 means full), vllm:num_requests_running and vllm:num_requests_waiting (gauges), and vllm:num_preemptions (a counter, exported with the Prometheus _total suffix). Useful alert expressions:
# Sustained preemption: recompute is eating capacity
sum by (pod) (rate(vllm:num_preemptions_total[5m])) > 0.5
# Cache near full while work queues: admission is too generous
max_over_time(vllm:kv_cache_usage_perc[10m]) > 0.95
and on (pod) vllm:num_requests_waiting > 10
# Queue growing faster than it drains
deriv(vllm:num_requests_waiting[10m]) > 0The engine-level controls are --max-num-seqs, which caps concurrent sequences, --max-model-len, which caps context and therefore the worst-case blocks per sequence, and --gpu-memory-utilization, which sets how much memory the engine may use for weights, activations and cache together. Lowering the first two trades peak throughput for predictable latency. The fleet-level control is admission: reject or queue at the gateway, with a clear 429 and a retry-after hint, before the engine is forced to preempt. The policies are described in admission control for LLM serving.
Per-request limits and three timeouts
Most engine stalls start as one unusual request. A 120,000-token prompt monopolises prefill. A request with max_tokens of 32,000 holds blocks for minutes. A structured output schema with a pathological grammar slows every decode step it participates in. A client that disconnects without the server noticing keeps generating for nobody.
Enforce limits at the gateway, where they are cheap and visible, and again in the engine as a backstop: maximum prompt tokens per tier, maximum max_tokens, maximum concurrent requests per tenant, and a rate limit on prompt tokens per tenant, not only on requests. Then set three timeouts, because a single total timeout cannot fit both a chat reply and a long report:
| Timeout | Measures | Typical starting value | Action on expiry |
|---|---|---|---|
| Time to first token | Queue plus prefill | 10 to 30 s | Cancel; safe to retry elsewhere |
| Inter-token idle | Gap between streamed chunks | 5 to 15 s | Cancel; the engine or GPU may be stuck |
| Total | Whole request | Tied to max_tokens and tier | Cancel and return what streamed |
The inter-token idle timeout is the one teams forget, and it is the best client-side detector of a hung GPU. Make sure cancellation reaches the engine: when the client goes away the server must abort the request and free its blocks, or abandoned streams quietly consume capacity.
Retries that do not amplify outages
Retries hide transient failures and multiply real ones. Three rules keep them honest. First, retry only before the first token reaches the client. After that the client has partial output, and a retry produces a different continuation; return an error or a clean truncation instead. Second, use a retry budget, not a retry count: allow retries up to, say, 10 percent of recent successful requests across the gateway, so an outage cannot triple the load on the replicas that are still up. Third, retry on a different replica, and treat 429 and 503 from admission as a signal to back off, not to try harder.
class RetryBudget:
# Token bucket: each success deposits `ratio` tokens, each retry spends one.
def __init__(self, ratio=0.1, cap=100.0):
self.ratio, self.cap, self.tokens = ratio, cap, cap
def on_success(self):
self.tokens = min(self.cap, self.tokens + self.ratio)
def try_spend(self):
if self.tokens >= 1.0:
self.tokens -= 1.0
return True
return False
def send(req, replicas, budget, first_token_timeout):
tried = set()
while True:
rep = pick(replicas, exclude=tried) # None when all were tried
if rep is None:
raise Unavailable("no replica left to try")
tried.add(rep)
try:
stream = rep.open_stream(req, first_token_timeout) # raises before any byte is sent
except (Timeout, Unavailable):
if not budget.try_spend():
raise
continue
budget.on_success()
return stream # after this point: no retriesHedging, sending a delayed second copy, doubles prefill cost for every hedged request. Use it only for short prompts on lightly loaded fleets, and count hedges against the same budget.
Deploys and graceful drains
Rollouts are the most frequent cause of dropped streams, because they restart every replica on purpose. A graceful drain makes them invisible. On SIGTERM: fail readiness immediately so the load balancer stops routing, stop accepting new requests, let in-flight streams finish up to a deadline, then exit. Set the pod's termination grace period above the longest stream you are willing to wait for, plus the load balancer's deregistration delay. Bring new replicas up before old ones go down: with a model that takes four minutes to load, a rollout that removes capacity first runs at reduced capacity for four minutes per batch of pods. Warm new replicas with a few real requests before marking them ready.
GPU faults and hung tensor-parallel groups
GPUs fail more often than CPUs, and their failures are less polite. Uncorrectable ECC errors, Xid events reported by the driver, and NVLink errors can crash the process, or worse, leave it alive and stuck. The details of each fault class are in GPU hardware faults. Two serving-specific consequences matter here.
Tensor parallelism widens the blast radius. A model split across eight GPUs with tensor parallelism needs all eight for every forward pass; one bad GPU stops the whole replica, and a collective operation waiting on it can hang rather than fail. Detect that with a progress watchdog: if vllm:num_requests_running is above zero and the generation token counter has not advanced for a minute, the replica is stuck, whatever /health says. Drain it, restart it, and if it recurs on the same node, cordon the node and open a hardware ticket.
# Stuck replica: work is running but no tokens are being produced
(sum by (pod) (vllm:num_requests_running) > 0)
and on (pod)
(sum by (pod) (increase(vllm:generation_tokens_total[1m])) == 0)Combine hardware signals, engine metrics and probe results into one score per replica so routing can prefer healthy replicas before they fail outright; the scoring approach is in LLM system health scoring.
Worked example: where the error budget goes
How much of the error budget do crashes consume? A crash kills the streams in flight, and by Little's law the fraction of requests killed is roughly crashes per replica per second times the mean stream duration. Duration is the multiplier, so compare two workloads on the same hypothetical fleet: 8 replicas, 40 concurrent streams each on average, an availability SLO of 99.9 percent over a 30-day month, and 10 crashes or deep-probe restarts per replica per month, each taking 6 minutes including model load.
| Quantity | Chat, 20 s streams | Reasoning, 300 s streams |
|---|---|---|
| Arrival rate (320 streams / duration) | 16 req/s | about 1.07 req/s |
| Requests per month | about 41.5 million | about 2.76 million |
| Error budget (0.1 percent) | about 41,500 | about 2,760 |
| Streams killed: 8 x 10 crashes x 40 | 3,200 (about 8 percent of budget) | 3,200 (about 116 percent) |
| 4 undrained rollouts: 4 x 8 x 40 | 1,280 (about 3 percent) | 1,280 (about 46 percent) |
For short chat traffic crashes are a nuisance; overload, admission and timeouts decide the SLO. For long reasoning or report streams the same crash rate blows the whole budget by itself, while the capacity lost to restarts, 480 replica-minutes or about 0.14 percent, stays negligible with one spare replica. The cost of a crash is the in-flight streams, not the downtime. So for long streams: crash less (bounded requests, admission before preemption storms, cordoning bad nodes), detect hung replicas fast, and drain on every planned restart. Halving crashes and draining rollouts brings the reasoning fleet to about 1,600 failures, inside budget with little margin; standby capacity for regional loss is covered in active-passive LLM serving.
Failure modes
- Liveness tied to the deep probe. A busy engine misses a probe deadline, gets restarted, and the restart kills every stream on it. Keep liveness cheap.
- Retry storms. Unbudgeted client and gateway retries turn a one-replica failure into a fleet overload. Budget at every layer that retries.
- One bad node, many restarts. Restarting on a faulty GPU just moves the failure. Track restarts per node and cordon on repeats.
- Metric names drift. Engine releases rename or add metrics. Pin the engine version and test alert queries against a live /metrics after every upgrade.
Trade-offs
| Control | Gains | Costs |
|---|---|---|
| Lower max-num-seqs | Fewer preemptions, stable latency | Lower peak throughput |
| Strict admission at the gateway | No storms, honest 429s | Rejected work during bursts |
| Deep probe | Detects live-but-broken engines | Extra load, risk of false positives |
| Retry before first token, with budget | Hides most transient faults | Added latency on retried requests |
| N+1 or N+2 spare replicas | Absorbs restarts and rollouts | Idle GPU cost |
| Long drain deadlines | No dropped streams on rollout | Slower rollouts |
What to do next
- Draw your failure-domain map and write, for each domain, the control that contains it and the metric that proves it works.
- Keep /health as liveness, add a readiness check that fails while loading and draining, and run a one-token deep probe from the gateway.
- Alert on preemption rate, on KV usage with a waiting queue, and on the stuck-replica query, after checking every metric name against your engine's /metrics.
- Set gateway limits on prompt tokens and max_tokens per tier, plus TTFT, inter-token and total timeouts, and test that cancellation frees engine capacity.
- Implement retry-before-first-token with a fleet-wide retry budget.
- Add SIGTERM draining, surge-first rollouts and warm-up requests, then measure dropped streams during a rollout.
- Redo the worked error-budget calculation with your own crash rate and concurrency.