Disaster recovery for an LLM serving fleet looks easy on paper. Inference servers are stateless, the weights are files, and a load balancer can point anywhere. The hard part is something database DR never faces: the scarce resource is not the data but the accelerators. You cannot buy 32 GPUs in a new region during an outage in the time it takes to update DNS. So for serving, DR is mostly a capacity decision you make months ahead, plus an artifact discipline that keeps the standby region able to run exactly the model you run today.

This article covers the standing posture: the DR tiers costed in GPUs, the replication manifest and the check that proves the DR region is complete, where the DR GPUs come from, a degraded-mode ladder for when you have fewer GPUs than demand, and how to decide that a disaster has happened. Testing that posture is a separate discipline, covered in the LLM DR drill article, which owns per-asset RTO and RPO and the restore path.

What a disaster means for an inference fleet

Start by listing the events you are planning for, because each needs a different answer. A region loss takes compute, storage and networking together. A capacity loss is more common: a zone loses power, a GPU SKU is recalled for a firmware issue, or a provider reclaims capacity, and you keep the region but lose a third of the fleet. A control-plane loss leaves the GPUs running but stops you deploying or scaling. A logical disaster is a bad release (a corrupted adapter, a tokenizer mismatch, a broken chat template) that reaches every region at once. Replication makes this worse, not better.

Then list what actually holds state. The engine process holds only caches. The real state is elsewhere: the artifacts (weights, tokenizer, chat template, LoRA adapters, container images, engine configuration), the policy (routing rules, quotas, guardrail models and their thresholds), and user data (conversation history, uploaded files, asynchronous batch jobs in flight). DR means having all three available somewhere else, plus the GPUs to use them.

DR tiers, priced in GPUs

The classic DR tiers still apply, but price them in GPUs rather than in storage. The table assumes a primary region running 32 GPUs (four 8-GPU nodes) at peak and a DR target of carrying the full load. The RTO column names what dominates recovery time, not a number, because the numbers depend on your stack.

TierStanding GPUs in DRRecovery is dominated byFits
Backup and restore0Finding GPUs at all, then image pull and weight loadInternal tools, batch-only workloads
Pilot light0 running, capacity reservedNode boot, weight load, engine warm-upServices that can tolerate tens of minutes
Warm standbyA floor, for example 8 of 32Scale-out from the floor while a degraded mode carries trafficMost customer-facing chat and APIs
Active-activeFull share in every regionRouter convergence onlyRevenue-critical, strict latency SLOs

Two points follow. First, backup and restore is really a bet on spot availability during someone else's incident. When a region fails, every tenant reaches for the same SKU in the neighbouring region. Second, active-active is only DR if each region keeps enough headroom to absorb the loss of another. That N+1 arithmetic is worked through in multi-region LLM serving. Warm standby is the usual compromise because it pairs a small always-on floor with the degraded modes described below.

The release manifest is the unit of replication

The most common DR failure is not missing GPUs. It is a DR region that starts the engine and serves a different model: last month's adapter, a tokenizer without the new special tokens, or an engine image built with a different attention kernel. The fix is to make one release manifest the unit of replication. Every artifact is listed by content digest, and the DR region counts as ready only if it holds every digest.

LLM serving DR: one manifest, two regions, a ladder in betweenGlobal routerhealth + brownout levelPrimary regionN GPU replicas, live trafficDR regionwarm floor, reserved headroomnormalfailover shareRelease manifestdigests of every artifactArtifact store (primary)weights, images, adaptersArtifact store (DR)replicated by digestreplicateParity checkermanifest vs DR, every hourDeclare controllersignals, quorum, humanset levelConversation storeasync replica, RPO in secondsDegraded modesshed, cap, smaller modelKV and prefix caches are never replicated: they rebuild
One release manifest drives replication. A parity checker proves the DR region holds every digest. The declare controller moves traffic and sets the degraded-mode level. Caches are deliberately left out.

# release-manifest.yaml (one per deployed model version)
release: chat-70b-2026-10-01
weights:          {uri: weights/chat-70b/, digest: sha256:9f1c..., bytes: 141_000_000_000}
tokenizer:        {digest: sha256:41aa...}
chat_template:    {digest: sha256:0b7e...}
adapters:         [{name: support-v7, digest: sha256:c3d2...}]
engine_image:     {ref: registry/llm-engine@sha256:77e0...}
engine_config:    {digest: sha256:aa19...}   # tensor parallel, max model len, KV dtype
guardrail_models: [{name: injection-clf-v3, digest: sha256:5e11...}]
routing_policy:   {digest: sha256:e902...}
secrets_refs:     [kms/dr/llm-api-signing]  # references only, resolved per region

The parity checker runs on a schedule and blocks promotion of a new release until the DR region has caught up. It compares digests, never timestamps or file names.

def parity(manifest: dict, dr_store) -> list[str]:
    """Return a list of problems; an empty list means the DR region is ready."""
    problems = []
    for kind, entry in iter_artifacts(manifest):        # weights, adapters, image, ...
        have = dr_store.digest(kind, entry)              # None if missing
        if have is None:
            problems.append(f"{kind}: missing in DR")
        elif have != entry["digest"]:
            problems.append(f"{kind}: digest {have[:12]} != {entry['digest'][:12]}")
    for ref in manifest.get("secrets_refs", []):
        if not dr_store.secret_resolvable(ref):          # KMS keys are regional
            problems.append(f"secret {ref}: not resolvable in DR")
    return problems

Note the secrets line. Container registries, KMS keys and identity providers are often regional, and they are the dependencies people forget. A DR region that cannot pull the engine image because the registry lives in the failed region has an RTO of however long it takes to rebuild the registry.

What not to replicate

Some state should not be replicated. The KV cache and the prefix cache are derived from prompts and rebuild on the first request. Copying them across regions costs more bandwidth than it saves, and a stale cache is a correctness risk. Accept the consequence and plan for it: right after failover, time to first token rises because every conversation re-prefills its whole history.

Conversation history is different. It is user data with a real RPO. Replicate it asynchronously and decide on a lag you can explain, for example a few seconds. When a user returns after failover and their last turn is missing, the client should resend it rather than act as if it never happened. In-flight streaming responses are lost; the client retries with the same request ID so the server can deduplicate. Async batch jobs need durable job records in replicated storage so the DR region can resume or restart them, not silently drop them.

Where the DR GPUs come from

The DR GPUs have to come from somewhere, and there are four sources. Most teams combine them.

  1. Reserved capacity in the DR region, paid for whether or not it is used. The trick is to fill it with preemptible work: evaluation runs, batch inference and embedding backfills. That work is evicted when you declare. See LLM reserved capacity for the break-even maths.
  2. A different GPU type. If the DR region only has an older or smaller SKU, the engine config changes: more tensor parallelism, a shorter maximum context, or a quantized checkpoint that fits. Each of those is a distinct artifact. It belongs in the manifest and needs its own evaluation run before the disaster, not during it.
  3. Borrowing from training. A training cluster in the DR region can be checkpointed and paused. Agree the procedure with the training owners in advance, including who can trigger it.
  4. Another provider. An API from a different vendor can carry part of the traffic if the product tolerates a different model. Routing and error classification for that path are covered in LLM provider failover.

A degraded-mode ladder

Warm standby means that for some minutes, demand exceeds the GPUs you have. Without a plan, queues grow without bound, every request times out, and the region that survived fails too. A degraded-mode ladder decides in advance what you give up, in order. Each rung frees capacity and has a known cost to users.

LevelActionWhat it frees
0Normal-
1Pause batch and async jobs, queue them durablyAll offline load
2Cap output length per requestDecode time per request
3Cap input context, reject very long promptsKV memory, prefill time
4Route free-tier traffic to a smaller or quantized modelLarge-model GPU share
5Overflow to a second provider for eligible tenantsWhatever it can carry

A small controller picks the level from the ratio of available capacity to demand. It uses hysteresis so it does not flap.

UP_AT   = [1.00, 0.90, 0.75, 0.60, 0.45]   # capacity/demand ratio that triggers level i+1
DOWN_AT = [1.10, 1.00, 0.85, 0.70, 0.55]   # must recover past this to step back down

def next_level(level: int, ratio: float, dwell_s: float) -> int:
    if level < 5 and ratio < UP_AT[level]:
        return level + 1                      # degrade immediately
    if level > 0 and ratio > DOWN_AT[level - 1] and dwell_s > 300:
        return level - 1                      # recover slowly, one rung at a time
    return level

The thresholds here are placeholders. Set them from load tests that measure how much each rung actually frees in your traffic mix.

Deciding that it is a disaster

Failing over too early is expensive: cold caches, a degraded experience and a failback later. Failing over too late lets an outage run long. So declaring a disaster should be a written rule, not a judgment made at 3 a.m. A practical rule automates the capacity response but keeps a human in the loop for the region decision.

def assess(sig) -> str:
    # sig: health from >= 3 independent vantage points, provider status, error budget burn
    region_down = sig.vantage_failures >= 2 and sig.minutes_unhealthy >= 5
    capacity_loss = sig.healthy_gpu_fraction < 0.7
    if sig.bad_release_suspected:
        return "ROLLBACK"           # logical disaster: never fail over to a replica of it
    if region_down:
        return "PAGE_AND_PROPOSE_FAILOVER"   # human confirms; router change is one command
    if capacity_loss:
        return "SCALE_DR_AND_RAISE_LADDER"  # automatic
    return "OK"

Two details matter. Health must be judged from several vantage points, so one broken probe cannot trigger a failover. And the logical-disaster branch comes first: if the last release is suspect, failing over moves you to a region running the same bad release. The answer is to roll back the manifest. DNS-level mechanics, including real failover timing with TTLs, are in DNS failover for LLM serving.

Worked example: a 70B chat service loses its region

Take a chat service running a 70B model in BF16 on 32 GPUs in the primary region. It uses warm standby with an 8-GPU floor in the DR region and reserved capacity for 24 more GPUs, filled with batch embedding jobs. These are the assumptions, so measure your own: weights are about 140 GB (70 billion parameters times 2 bytes); each node reads from the regional object store at about 1.5 GB/s effective; and engine warm-up (kernel compilation, CUDA graph capture and a health check) takes about 2 minutes.

The primary region fails at T+0. The controller sees capacity collapse within a minute and raises the ladder to level 4. The 8-GPU floor now serves paid traffic at capped lengths, and free-tier traffic goes to the quantized fallback. The on-call engineer confirms the region decision at about T+6. Batch jobs on the reserved nodes are evicted. Each node then pulls 140 GB in about 95 seconds (140 / 1.5), and the three nodes pull in parallel because each has its own path to the store. The new replicas are healthy at about T+11. The ladder steps down one rung every five minutes as the capacity ratio recovers, and reaches level 0 at about T+30. Time to first token stays high for several minutes while prefix caches refill.

The parts of that timeline you control before the incident are the manifest (the DR store already holds the weights, so nothing crosses regions during recovery), the reservation (the nodes exist), and the ladder (users get a slower service instead of errors). The lesson from the numbers: weight loading is minutes, but finding GPUs without a reservation can be hours.

Failure modes

  • Manifest drift. An adapter is hot-patched in the primary region outside the release process, and the DR region serves the old one. Make parity a deploy gate.
  • Regional hidden dependencies. The image registry, KMS keys, feature flags or the guardrail classifier live only in the failed region.
  • Thundering herd on the survivor. Every client retries at once, the DR floor drowns, and its health checks fail. Clients need jittered backoff, and the ladder must move before the queue explodes.
  • Untested fallback configs. The quantized or higher-parallelism config has never run against the evaluation set, and it fails in a new way during the incident.
  • Failing over a logical disaster. Moving traffic to a replica of a bad release.
  • Failback stampede. Moving all traffic back at once onto cold caches. Shift it in steps.

Trade-offs

ChoiceGainCost
Larger warm floorShorter degraded periodIdle GPUs every day
Reservation filled with batch workCapacity without wasteBatch SLAs suffer during incidents
Different-SKU fallbackUses whatever existsMore configs to evaluate and keep current
Human-confirmed region failoverFewer false failoversMinutes added to the RTO
Async conversation replicationNo write latency costA few seconds of RPO

What to do next

  1. List your disaster events: region, capacity, control plane and bad release. Write the response to each.
  2. Turn every release into a manifest of digests, including image, tokenizer, template, adapters, guardrails and secret references.
  3. Run a parity checker against the DR region and make it a release gate.
  4. Choose a DR tier in GPUs, then reserve that capacity and fill it with preemptible work.
  5. Define a degraded-mode ladder, load-test how much each rung frees, and wire up the controller.
  6. Write the declare rule with a rollback branch first, and rehearse it with the drill.
Key takeaway: For LLM serving, DR is a capacity commitment plus an artifact discipline. Decide how many GPUs stand by and where the rest will come from, replicate one manifest of digests and prove parity before every release, leave caches out, and prepare a degraded-mode ladder so a short-handed region slows down instead of failing. Check for a bad release before you fail over.