Active-passive serving runs one pool of inference servers that takes all traffic and a second pool that is kept ready to take over when the first fails. It is the oldest high-availability pattern there is, and for most stateless web services it is a solved problem. For LLM serving it is not, for one reason: a GPU replica takes minutes, not seconds, to become useful. A standby that looks healthy can still be ten minutes from serving its first token, and the scarce resource it needs, GPUs, may be unavailable during exactly the outage that triggered failover.
This article covers the parts that are specific to GPUs and LLM runtimes: how warm the standby must be, where cold-start time actually goes, how to prove a standby can serve, how to promote it without two primaries, and what happens to requests that were streaming when the primary died. Traffic steering by DNS and its timeline are covered in DNS failover for LLM serving; the regional capacity question is covered in multi-region LLM serving. Every latency figure below is illustrative. Measure your own.
The shape of an active-passive deployment
The pattern has five parts. An active pool serves all user traffic. A passive pool runs the same model version and configuration but receives none. A router sends requests to whichever pool currently owns the primary role. A failover controller decides when to switch, and a lease store gives each switch a monotonically increasing epoch so that only one pool can be primary at a time. A synthetic prober sends real generation requests to both pools, because the passive pool has no user traffic to tell you it is broken.
Why choose this over active-active? Active-active spreads traffic across both pools all the time, so both stay proven and failover is just losing capacity. It costs more coordination: both pools must handle routing, session affinity and model rollout together. Active-passive is simpler to reason about, fits workloads tied to one primary site for data or licensing reasons, and lets the standby be cheaper than the primary. The price is that the standby is unproven unless you work to prove it.
How warm is the standby?
'Passive' covers a wide range of readiness. Name the tier explicitly, because each one buys a different recovery time at a different cost.
| Tier | State of the standby | Recovery dominated by | Cost |
|---|---|---|---|
| Hot | Full replica count, weights in GPU memory, probed | Router switch and client retries | Same GPUs as primary |
| Warm (sleeping) | GPUs allocated, process up, weights parked in host memory | Moving weights back to GPU and re-allocating cache | Same GPUs, idle |
| Pilot light | One or two replicas hot, rest scaled to zero | Provisioning GPUs, loading weights on new replicas | A fraction of primary |
| Cold | Images and weights staged, no GPUs | Getting GPUs at all, then full cold start | Storage only |
The warm tier is worth a closer look because recent runtimes support it directly. vLLM has a sleep mode, enabled with --enable-sleep-mode. At level 1 it offloads model weights to CPU memory and discards the KV cache; at level 2 it discards both, and the weights must be reloaded on wake. The HTTP endpoints for it (/sleep, /wake_up, /is_sleeping) are only exposed when the server runs with VLLM_SERVER_DEV_MODE=1, which the vLLM docs say must never be reachable by production users. If you use it, put those endpoints on a private control network that only the controller can reach. Level 1 also needs enough host RAM to hold the weights for every replica on the node.
Pilot light is the tier most teams actually run, and it has a trap. Scaling the standby from two replicas to twenty during a regional incident means requesting eighteen replicas' worth of GPUs at the same moment as everyone else affected by that incident. If capacity is not reserved, recovery time is unbounded.
Where cold-start time goes
Recovery time for any tier below hot is the sum of stages, and each stage has its own fix. Measure them separately on your own stack:
| Stage | What happens | Main lever |
|---|---|---|
| Node | GPU instance or pod is scheduled and drivers come up | Reserved capacity, pre-provisioned nodes |
| Image | Multi-gigabyte serving image is pulled | Pre-pulled images, local registry mirror |
| Weights | Tens to hundreds of GB read from storage into GPU memory | Local NVMe cache, faster formats, parallel reads |
| Warmup | Memory profiling, cache allocation, kernel and graph capture | Runtime flags, keeping warmup caches |
| First token | First real prefill and decode | A warmup request before taking traffic |
The weight stage is usually the largest and it is plain arithmetic. As an illustration, 140 GB of weights read at an effective 2 GB/s from network storage take about 70 seconds; read from local NVMe at several times that rate they take a fraction of it. Your effective rate depends on file format, the number of parallel readers and tensor-parallel sharding, so measure it. This script times the two numbers that matter to users, readiness and first token, from process start:
import time, requests
def time_to_serve(base: str, model: str, limit_s: float = 1800.0) -> dict:
t0 = time.monotonic()
while True:
try:
if requests.get(f"{base}/health", timeout=2).status_code == 200:
break
except requests.RequestException:
pass
if time.monotonic() - t0 > limit_s:
raise TimeoutError("replica never became healthy")
time.sleep(1)
ready = time.monotonic() - t0
r = requests.post(f"{base}/v1/completions",
json={"model": model, "prompt": "ping", "max_tokens": 4},
timeout=300)
r.raise_for_status()
return {"ready_s": round(ready, 1), "first_tokens_s": round(time.monotonic() - t0, 1)}Run it for each tier in a game day and record the breakdown, not just the total. A health endpoint that answers before warmup is finished is common, which is why the script also times a real completion.
Keeping the standby honest
A passive pool fails in ways a liveness check cannot see: a driver update left one node unable to load the model, the standby is still on last month's model version, a config change raised the context length past what its GPUs can hold. Keep it honest with three controls.
- Generating probes. Every minute, send each standby replica a fixed prompt with deterministic sampling and check the output and the latency. A probe that does not produce tokens is not a probe. Signals and scoring for this are covered in LLM system health scoring.
- Parity checks. Compare the model artifact digest, runtime version, tokenizer and serving flags on both pools, and alarm on any drift. Roll new model versions to the standby first.
- Load proof. A standby that answers one probe may still fall over at full traffic. Replay a recorded traffic sample at primary load against it regularly, as in LLM load testing.
Then fail over for real on a schedule. A game day that moves production traffic to the standby for an hour is the only proof that promotion, capacity and client retries work together.
Promotion without two primaries
Promotion must never produce two primaries: two pools writing to the same downstream queue, or a recovered old primary that resumes serving with a stale configuration. Model promotion as a state machine and attach a fencing epoch from a store with compare-and-set semantics such as etcd or a database row with a version column.
class FailoverController:
def __init__(self, lease, router, prober, standby):
self.lease, self.router, self.prober, self.standby = lease, router, prober, standby
self.state = "PRIMARY_OK"
def tick(self):
bad = self.prober.consecutive_failures("primary")
if self.state == "PRIMARY_OK" and bad >= 3:
self.state = "SUSPECT" # avoid flapping on one bad probe
elif self.state == "SUSPECT":
if bad == 0:
self.state = "PRIMARY_OK"
elif bad >= 6 and self.prober.healthy("standby"):
self.promote()
def promote(self):
epoch = self.lease.current_epoch()
if not self.lease.compare_and_set(epoch, epoch + 1, owner="standby"):
return # someone else already moved it
self.state = "PROMOTING"
self.standby.wake_if_sleeping() # private control endpoint only
self.standby.scale_to_primary_size()
self.standby.wait_until_serving(min_fraction=0.8)
self.router.route_to("standby", epoch=epoch + 1) # router rejects lower epochs
self.state = "STANDBY_PRIMARY"Two rules make the epoch effective. The router refuses any route update carrying an epoch lower than the one it holds, so a delayed command from an old controller cannot switch traffic back. And each pool's serving layer checks its lease periodically and stops accepting work when it no longer owns the current epoch, which fences a primary that comes back from a network partition believing it is still in charge. Failback is a separate, deliberate promotion in the other direction, run when the original pool has passed probes for a sustained period, never automatically on first recovery.
What happens to in-flight requests
When the primary dies, three kinds of work are lost. Requests in the queue have not started and can simply be retried on the new primary. Requests mid-stream have sent some tokens to the client. The KV cache for every conversation, including any prefix cache warmed by shared system prompts, exists only in the failed GPUs' memory and is gone. That is why the KV cache appears on the diagram as never shared: replicating it across pools would cost interconnect bandwidth comparable to serving itself.
For streams, pick a contract and tell clients. The simplest is to fail the stream with a retryable error and let the client resend the whole request; the user sees the answer restart. A gentler option is for the gateway to resend the original prompt plus the tokens already delivered as a continuation, which avoids repeating visible text but costs a long prefill and can change the remainder of the answer. Either way, expect a burst of long prefills right after promotion, because every active conversation rebuilds its cache at once. Size the standby for that spike, not for steady state, or apply admission control for the first minutes. Separating prefill from decode, as in disaggregated serving, changes the shape of that spike but not its existence.
Worked example: a five-minute objective
A team serves a 70B-parameter model in 16-bit precision (about 140 GB of weights) on eight replicas of two GPUs each, with a recovery objective of five minutes. Illustrative numbers from their game days:
| Tier | Node | Image | Weights | Warmup | Total |
|---|---|---|---|---|---|
| Hot | 0 | 0 | 0 | 0 | under 1 min (routing) |
| Warm, level 1 sleep | 0 | 0 | host to GPU | cache re-alloc | about 1 min |
| Pilot light, 2 hot | 4-12 min | 2 min | 2 min | 1 min | 9-17 min |
| Cold | unbounded | 2 min | 2 min | 1 min | unbounded |
Pilot light misses the objective because of node provisioning, which they cannot control, not weight loading, which they can. They choose a warm tier for four of the eight replicas and pilot light for the rest: half capacity within about a minute, with admission control shedding low-priority traffic until the remaining replicas arrive. The reserved GPUs for those four replicas are the cost of the objective, and the table makes that cost a decision rather than a surprise.
Failure modes
- Probes that do not generate. A green health check on a replica that cannot load the model. Probe with real completions.
- Silent drift. The standby runs an older model or different flags; failover changes answers. Check artifact digests and flags on both pools.
- Unreserved capacity. Pilot light assumes GPUs will be available during the incident. Reserve or pre-provision what the objective needs.
- Split brain. A recovered primary keeps serving. Use epochs, a router that rejects stale epochs, and self-fencing.
- Prefill storm. Every conversation rebuilds its cache at once and the new primary saturates. Plan admission control for the first minutes.
- Exposed control endpoints. Sleep and wake endpoints reachable by users. Keep them on a private network.
- Automatic failback. Traffic returns to a pool that recovered for a moment, then fails again. Make failback manual or require sustained health.
What to do next
- Write down the recovery objective and pick the standby tier that can meet it, per replica if needed.
- Measure node, image, weight, warmup and first-token times separately on your own stack.
- Probe both pools with real generation and alarm on model, runtime and flag drift.
- Implement promotion as a state machine with an epoch from a compare-and-set store, a router that rejects stale epochs and self-fencing on the serving side.
- Define the client contract for broken streams and size the standby for the post-failover prefill spike.
- Reserve the GPUs your objective depends on, or accept and document a longer recovery.
- Run a game day that moves real traffic to the standby, and repeat it on a schedule.