Disaster recovery for an LLM serving fleet looks easy on paper. Inference servers are stateless, the weights are files, and a load balancer can point anywhere. The hard part is something database DR never faces: the scarce resource is not the data but the accelerators. You cannot buy 32 GPUs in a new region during an outage in the time it takes to update DNS. So for serving, DR is mostly a capacity decision you make months ahead, plus an artifact discipline that keeps the standby region able to run exactly the model you run today.
This article covers the standing posture: the DR tiers costed in GPUs, the replication manifest and the check that proves the DR region is complete, where the DR GPUs come from, a degraded-mode ladder for when you have fewer GPUs than demand, and how to decide that a disaster has happened. Testing that posture is a separate discipline, covered in the LLM DR drill article, which owns per-asset RTO and RPO and the restore path.
What a disaster means for an inference fleet
Start by listing the events you are planning for, because each needs a different answer. A region loss takes compute, storage and networking together. A capacity loss is more common: a zone loses power, a GPU SKU is recalled for a firmware issue, or a provider reclaims capacity, and you keep the region but lose a third of the fleet. A control-plane loss leaves the GPUs running but stops you deploying or scaling. A logical disaster is a bad release (a corrupted adapter, a tokenizer mismatch, a broken chat template) that reaches every region at once. Replication makes this worse, not better.
Then list what actually holds state. The engine process holds only caches. The real state is elsewhere: the artifacts (weights, tokenizer, chat template, LoRA adapters, container images, engine configuration), the policy (routing rules, quotas, guardrail models and their thresholds), and user data (conversation history, uploaded files, asynchronous batch jobs in flight). DR means having all three available somewhere else, plus the GPUs to use them.
DR tiers, priced in GPUs
The classic DR tiers still apply, but price them in GPUs rather than in storage. The table assumes a primary region running 32 GPUs (four 8-GPU nodes) at peak and a DR target of carrying the full load. The RTO column names what dominates recovery time, not a number, because the numbers depend on your stack.
| Tier | Standing GPUs in DR | Recovery is dominated by | Fits |
|---|---|---|---|
| Backup and restore | 0 | Finding GPUs at all, then image pull and weight load | Internal tools, batch-only workloads |
| Pilot light | 0 running, capacity reserved | Node boot, weight load, engine warm-up | Services that can tolerate tens of minutes |
| Warm standby | A floor, for example 8 of 32 | Scale-out from the floor while a degraded mode carries traffic | Most customer-facing chat and APIs |
| Active-active | Full share in every region | Router convergence only | Revenue-critical, strict latency SLOs |
Two points follow. First, backup and restore is really a bet on spot availability during someone else's incident. When a region fails, every tenant reaches for the same SKU in the neighbouring region. Second, active-active is only DR if each region keeps enough headroom to absorb the loss of another. That N+1 arithmetic is worked through in multi-region LLM serving. Warm standby is the usual compromise because it pairs a small always-on floor with the degraded modes described below.
The release manifest is the unit of replication
The most common DR failure is not missing GPUs. It is a DR region that starts the engine and serves a different model: last month's adapter, a tokenizer without the new special tokens, or an engine image built with a different attention kernel. The fix is to make one release manifest the unit of replication. Every artifact is listed by content digest, and the DR region counts as ready only if it holds every digest.
# release-manifest.yaml (one per deployed model version)
release: chat-70b-2026-10-01
weights: {uri: weights/chat-70b/, digest: sha256:9f1c..., bytes: 141_000_000_000}
tokenizer: {digest: sha256:41aa...}
chat_template: {digest: sha256:0b7e...}
adapters: [{name: support-v7, digest: sha256:c3d2...}]
engine_image: {ref: registry/llm-engine@sha256:77e0...}
engine_config: {digest: sha256:aa19...} # tensor parallel, max model len, KV dtype
guardrail_models: [{name: injection-clf-v3, digest: sha256:5e11...}]
routing_policy: {digest: sha256:e902...}
secrets_refs: [kms/dr/llm-api-signing] # references only, resolved per regionThe parity checker runs on a schedule and blocks promotion of a new release until the DR region has caught up. It compares digests, never timestamps or file names.
def parity(manifest: dict, dr_store) -> list[str]:
"""Return a list of problems; an empty list means the DR region is ready."""
problems = []
for kind, entry in iter_artifacts(manifest): # weights, adapters, image, ...
have = dr_store.digest(kind, entry) # None if missing
if have is None:
problems.append(f"{kind}: missing in DR")
elif have != entry["digest"]:
problems.append(f"{kind}: digest {have[:12]} != {entry['digest'][:12]}")
for ref in manifest.get("secrets_refs", []):
if not dr_store.secret_resolvable(ref): # KMS keys are regional
problems.append(f"secret {ref}: not resolvable in DR")
return problemsNote the secrets line. Container registries, KMS keys and identity providers are often regional, and they are the dependencies people forget. A DR region that cannot pull the engine image because the registry lives in the failed region has an RTO of however long it takes to rebuild the registry.
What not to replicate
Some state should not be replicated. The KV cache and the prefix cache are derived from prompts and rebuild on the first request. Copying them across regions costs more bandwidth than it saves, and a stale cache is a correctness risk. Accept the consequence and plan for it: right after failover, time to first token rises because every conversation re-prefills its whole history.
Conversation history is different. It is user data with a real RPO. Replicate it asynchronously and decide on a lag you can explain, for example a few seconds. When a user returns after failover and their last turn is missing, the client should resend it rather than act as if it never happened. In-flight streaming responses are lost; the client retries with the same request ID so the server can deduplicate. Async batch jobs need durable job records in replicated storage so the DR region can resume or restart them, not silently drop them.
Where the DR GPUs come from
The DR GPUs have to come from somewhere, and there are four sources. Most teams combine them.
- Reserved capacity in the DR region, paid for whether or not it is used. The trick is to fill it with preemptible work: evaluation runs, batch inference and embedding backfills. That work is evicted when you declare. See LLM reserved capacity for the break-even maths.
- A different GPU type. If the DR region only has an older or smaller SKU, the engine config changes: more tensor parallelism, a shorter maximum context, or a quantized checkpoint that fits. Each of those is a distinct artifact. It belongs in the manifest and needs its own evaluation run before the disaster, not during it.
- Borrowing from training. A training cluster in the DR region can be checkpointed and paused. Agree the procedure with the training owners in advance, including who can trigger it.
- Another provider. An API from a different vendor can carry part of the traffic if the product tolerates a different model. Routing and error classification for that path are covered in LLM provider failover.
A degraded-mode ladder
Warm standby means that for some minutes, demand exceeds the GPUs you have. Without a plan, queues grow without bound, every request times out, and the region that survived fails too. A degraded-mode ladder decides in advance what you give up, in order. Each rung frees capacity and has a known cost to users.
| Level | Action | What it frees |
|---|---|---|
| 0 | Normal | - |
| 1 | Pause batch and async jobs, queue them durably | All offline load |
| 2 | Cap output length per request | Decode time per request |
| 3 | Cap input context, reject very long prompts | KV memory, prefill time |
| 4 | Route free-tier traffic to a smaller or quantized model | Large-model GPU share |
| 5 | Overflow to a second provider for eligible tenants | Whatever it can carry |
A small controller picks the level from the ratio of available capacity to demand. It uses hysteresis so it does not flap.
UP_AT = [1.00, 0.90, 0.75, 0.60, 0.45] # capacity/demand ratio that triggers level i+1
DOWN_AT = [1.10, 1.00, 0.85, 0.70, 0.55] # must recover past this to step back down
def next_level(level: int, ratio: float, dwell_s: float) -> int:
if level < 5 and ratio < UP_AT[level]:
return level + 1 # degrade immediately
if level > 0 and ratio > DOWN_AT[level - 1] and dwell_s > 300:
return level - 1 # recover slowly, one rung at a time
return levelThe thresholds here are placeholders. Set them from load tests that measure how much each rung actually frees in your traffic mix.
Deciding that it is a disaster
Failing over too early is expensive: cold caches, a degraded experience and a failback later. Failing over too late lets an outage run long. So declaring a disaster should be a written rule, not a judgment made at 3 a.m. A practical rule automates the capacity response but keeps a human in the loop for the region decision.
def assess(sig) -> str:
# sig: health from >= 3 independent vantage points, provider status, error budget burn
region_down = sig.vantage_failures >= 2 and sig.minutes_unhealthy >= 5
capacity_loss = sig.healthy_gpu_fraction < 0.7
if sig.bad_release_suspected:
return "ROLLBACK" # logical disaster: never fail over to a replica of it
if region_down:
return "PAGE_AND_PROPOSE_FAILOVER" # human confirms; router change is one command
if capacity_loss:
return "SCALE_DR_AND_RAISE_LADDER" # automatic
return "OK"Two details matter. Health must be judged from several vantage points, so one broken probe cannot trigger a failover. And the logical-disaster branch comes first: if the last release is suspect, failing over moves you to a region running the same bad release. The answer is to roll back the manifest. DNS-level mechanics, including real failover timing with TTLs, are in DNS failover for LLM serving.
Worked example: a 70B chat service loses its region
Take a chat service running a 70B model in BF16 on 32 GPUs in the primary region. It uses warm standby with an 8-GPU floor in the DR region and reserved capacity for 24 more GPUs, filled with batch embedding jobs. These are the assumptions, so measure your own: weights are about 140 GB (70 billion parameters times 2 bytes); each node reads from the regional object store at about 1.5 GB/s effective; and engine warm-up (kernel compilation, CUDA graph capture and a health check) takes about 2 minutes.
The primary region fails at T+0. The controller sees capacity collapse within a minute and raises the ladder to level 4. The 8-GPU floor now serves paid traffic at capped lengths, and free-tier traffic goes to the quantized fallback. The on-call engineer confirms the region decision at about T+6. Batch jobs on the reserved nodes are evicted. Each node then pulls 140 GB in about 95 seconds (140 / 1.5), and the three nodes pull in parallel because each has its own path to the store. The new replicas are healthy at about T+11. The ladder steps down one rung every five minutes as the capacity ratio recovers, and reaches level 0 at about T+30. Time to first token stays high for several minutes while prefix caches refill.
The parts of that timeline you control before the incident are the manifest (the DR store already holds the weights, so nothing crosses regions during recovery), the reservation (the nodes exist), and the ladder (users get a slower service instead of errors). The lesson from the numbers: weight loading is minutes, but finding GPUs without a reservation can be hours.
Failure modes
- Manifest drift. An adapter is hot-patched in the primary region outside the release process, and the DR region serves the old one. Make parity a deploy gate.
- Regional hidden dependencies. The image registry, KMS keys, feature flags or the guardrail classifier live only in the failed region.
- Thundering herd on the survivor. Every client retries at once, the DR floor drowns, and its health checks fail. Clients need jittered backoff, and the ladder must move before the queue explodes.
- Untested fallback configs. The quantized or higher-parallelism config has never run against the evaluation set, and it fails in a new way during the incident.
- Failing over a logical disaster. Moving traffic to a replica of a bad release.
- Failback stampede. Moving all traffic back at once onto cold caches. Shift it in steps.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Larger warm floor | Shorter degraded period | Idle GPUs every day |
| Reservation filled with batch work | Capacity without waste | Batch SLAs suffer during incidents |
| Different-SKU fallback | Uses whatever exists | More configs to evaluate and keep current |
| Human-confirmed region failover | Fewer false failovers | Minutes added to the RTO |
| Async conversation replication | No write latency cost | A few seconds of RPO |
What to do next
- List your disaster events: region, capacity, control plane and bad release. Write the response to each.
- Turn every release into a manifest of digests, including image, tokenizer, template, adapters, guardrails and secret references.
- Run a parity checker against the DR region and make it a release gate.
- Choose a DR tier in GPUs, then reserve that capacity and fill it with preemptible work.
- Define a degraded-mode ladder, load-test how much each rung frees, and wire up the controller.
- Write the declare rule with a rollback branch first, and rehearse it with the drill.