A disaster recovery plan for an LLM service is a claim: if the primary region, cluster or account becomes unusable, we can serve again within a stated time (the recovery time objective, RTO) and lose no more than a stated window of data (the recovery point objective, RPO). Until it has been exercised end to end, the claim is a guess. LLM stacks make the guess especially unreliable, because recovery depends on things classic DR runbooks do not mention: hundreds of gigabytes of weights, GPU capacity that may not exist in the recovery region on the day you need it, and the question of whether the restored model is the same model.

This article is about the drill: the scheduled exercise that rebuilds the service somewhere else and measures how long it took and what was missing. It is broader in scope than a gameday, which injects faults such as a lost replica or a hung rank into a running system. A DR drill assumes the site is gone and asks whether you can start from backups.

Scenarios and per-asset objectives

Start from the scenarios, because each one removes different things. A regional outage takes compute and regional storage but leaves your accounts, global control planes and cross-region replicas intact. A cluster loss from a bad upgrade or a deleted namespace takes the serving stack but not the buckets. An account compromise or destructive deletion is the worst case: anything the compromised credentials could reach, including replicas in other regions of the same account, must be assumed gone. Backups that the production role can delete are not backups for this scenario.

For each scenario write RTO and RPO per asset, not per service. A chat product might accept an RTO of an hour for the model but need conversation history with an RPO of minutes, while a retrieval index can be rebuilt from source documents and tolerate a day.

AssetTypical sizeHow it is recoveredRPO driver
Base model weights~2 bytes per parameter in BF16Copy from replicated object storageChanges only on release
Fine-tuned adapters and checkpointsMB to GB per adapterCopy; checkpoints from the training storeTraining cadence
Serving config, prompts, routing rulesKBFrom version control, redeployedLast merged commit
Retrieval indexGB to TBRestore snapshot or re-embed from sourceSnapshot interval
Conversation and session stateVariesDatabase replica or backupReplication lag
Secrets, API keys, KMS keysKBFrom a separate recovery accountRotation events
Container images and driversSeveral GBPull from a replicated registryImage tag pinning

The restore path and its bottlenecks

What has to exist in the recovery region before the first token is servedModel weightsbase + adaptersServing configengine flags, promptsRetrieval indexvectors + documentsSecrets and keysseparate accountGPU capacity + imagesquota, drivers, engine containerValidated replicachecksums + golden-set evalTraffic shiftDNS or global LB, then scale outThe drill measures each arrow. The slowest one, not the sum of good intentions, is the RTO.
Four groups of assets converge on GPU capacity in the recovery region; only a validated replica receives traffic.

The restore path has a fixed shape. Assets must be present in the recovery region, GPUs must be obtained, the engine container must be pulled, weights loaded into GPU memory, the replica validated, and traffic shifted. Time adds up along the longest chain, and on LLM stacks the longest chain is nearly always weights or GPUs.

Do the arithmetic before the drill. A 70B-parameter model in BF16 is about 140 GB. At a sustained 1 GB/s from object storage it takes about 140 seconds to copy; at 100 MB/s, which is easy to hit with a single-stream download, about 23 minutes. Multiply by the number of nodes that each need a local copy unless they share a filesystem. Those are assumptions to replace with your own measurements, which is exactly what the drill produces. The cheap fix is usually to keep weights pre-staged in the recovery region and to download in parallel ranges, not to buy faster storage.

GPU capacity is the harder problem. Quota in the recovery region is a number in a console; available capacity on the day of a regional outage, when everyone else is failing over too, is not. If your RTO depends on obtaining GPUs on demand, your RTO is unknown. Options in order of cost: a warm standby with a small number of replicas already serving, a capacity reservation held in the recovery region, or an explicitly degraded plan that serves a smaller model on whatever GPUs you can get. Multi-region serving discusses the capacity side of active-active designs in detail.

Validating that it is the same model

A service that answers is not necessarily the right service. Restores fail quietly in ways that produce fluent, wrong output: an older adapter, a tokenizer from a different revision, a quantized artifact where the full-precision one was expected, a missing system prompt, a sampling default that changed with an engine upgrade. Validation therefore has two layers.

First, identity: every artifact is checked against a manifest of SHA-256 digests produced at release time and stored with the backups. Second, behaviour: a golden set of prompts is run with greedy decoding and the outputs compared with recorded reference outputs, plus a small quality eval whose score must fall within a tolerance of the production baseline. Exact-match comparison is too strict across different GPU types or engine builds, because numerics differ; compare at the level of task scores and flag individual large divergences for a human.

import hashlib, json, pathlib

def sha256(path, chunk=1 << 24):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        while block := f.read(chunk):
            h.update(block)
    return h.hexdigest()

def verify_manifest(root, manifest_path):
    manifest = json.loads(pathlib.Path(manifest_path).read_text())
    bad = [rel for rel, digest in manifest["files"].items()
           if sha256(pathlib.Path(root) / rel) != digest]
    if bad:
        raise RuntimeError(f"{len(bad)} artifacts differ from release manifest: {bad[:5]}")

def behaviour_check(client, golden, baseline_score, tolerance=0.02):
    score = sum(case["grade"](client.complete(case["prompt"], temperature=0))
                for case in golden) / len(golden)
    if score < baseline_score - tolerance:
        raise RuntimeError(f"golden-set score {score:.3f} below baseline {baseline_score:.3f}")
    return score

The client, grading functions and tolerance are yours to define; the structure is the point. Both checks are gates in the runbook, and no traffic moves until both pass.

Scripting the drill

Script the drill so that it is repeatable and so that it records its own evidence. Each step is timed, logged with its output, and either passes or stops the drill. Running it by hand from a wiki page measures the operator, not the system.

import time, json

STEPS = [
    ("declare",        lambda ctx: ctx.page_incident("DR drill: primary region unavailable")),
    ("capacity",       lambda ctx: ctx.acquire_gpus(region=ctx.dr_region, replicas=ctx.min_replicas)),
    ("image",          lambda ctx: ctx.pull_image(ctx.engine_image_digest)),
    ("weights",        lambda ctx: ctx.stage_weights(ctx.release_manifest)),
    ("verify",         lambda ctx: verify_manifest(ctx.weights_dir, ctx.release_manifest)),
    ("start",          lambda ctx: ctx.start_replicas(warmup_prompts=8)),
    ("behaviour",      lambda ctx: behaviour_check(ctx.client, ctx.golden, ctx.baseline)),
    ("shift",          lambda ctx: ctx.shift_traffic(weight=ctx.drill_traffic_share)),
]

def run_drill(ctx, rto_minutes):
    t0, record = time.monotonic(), []
    for name, step in STEPS:
        s = time.monotonic()
        try:
            out = step(ctx)
            record.append({"step": name, "ok": True, "sec": round(time.monotonic() - s), "out": str(out)[:200]})
        except Exception as e:
            record.append({"step": name, "ok": False, "sec": round(time.monotonic() - s), "error": repr(e)})
            break
    total = (time.monotonic() - t0) / 60
    ctx.store_evidence(json.dumps({"total_min": round(total, 1), "rto_min": rto_minutes,
                                   "met": total <= rto_minutes and record[-1]["ok"],
                                   "steps": record}))

Steps run sequentially here for clarity; in practice image pull and weight staging run in parallel. The ctx object wraps your cloud, cluster and load-balancer APIs; keep it thin so the drill exercises the same automation an incident would. Every step should also appear in your runbook index so the on-call engineer can run it by hand if the script itself is part of what was lost.

Worked example: a 70B chat service

Worked example (illustrative): first drill timeline, primary region declared lost at T+0Declare + decideT+6 minObtain GPUsT+22 minPull imageT+31 minCopy weightsT+47 minLoad + warmT+55 minValidateT+63 minShift trafficT+72 minWeights copy and GPU acquisition dominate. Image pull ran in parallel with the copy.
An illustrative first-drill timeline for the scenario below. The shape is typical; your numbers come from your own drill.

Consider an illustrative team serving a 70B chat model on eight-GPU nodes, with a stated RTO of 60 minutes and a recovery region holding a reservation for two nodes. The first drill took 72 minutes and missed the target. The evidence file showed where the time went: six minutes to declare, because nobody was sure who could authorize the drill; sixteen minutes to bring up GPU nodes, because the node pool had scaled to zero and the GPU driver installer ran on boot; 25 minutes to copy weights, because the copy was a single-stream download from the primary region's replica bucket; eight minutes to load and warm; eight to validate; nine to shift traffic, mostly waiting for DNS TTLs.

The fixes were mechanical. Weights were pre-staged in the recovery region and downloaded in parallel ranges. One node was kept warm with the driver baked into the image. The decision owner was named in the runbook. Clients were moved to a global load balancer so the shift did not wait on resolvers; the trade-offs of the DNS route are covered in DNS failover for LLM serving. The next drill in this scenario finishes well inside the target. The point is that neither result is knowable without running the drill.

Cadence and failback

A drill that runs once is an audit; a drill that runs on a schedule is a control. A workable cadence: a tabletop walk-through quarterly, a partial drill monthly that restores into the recovery region without shifting user traffic, and a full drill twice a year that shifts a real share of traffic. Run a drill after any change that touches the restore path: a new model release, an engine upgrade, a new GPU type, or a new region.

Failback deserves its own drill. Moving traffic back to the primary region repeats the same steps in reverse, and data written in the recovery region (conversations, feedback, logs) must be reconciled. Teams that drill only the failover discover the failback problems during a real incident's second half. Training state is a separate plan: long runs recover from checkpoints, and how often you save determines their RPO, as covered in GPU checkpointing.

Failure modes

  • Backups reachable by production credentials. An account compromise deletes them along with production. Keep a copy in a separate account with write-once retention.
  • Weights replicated, tokenizer and config not. The model loads and produces garbage or subtly wrong output. Back up the full release bundle under one manifest.
  • Unpinned images. The recovery region pulls a newer engine tag with different defaults. Pin by digest.
  • Quota treated as capacity. The drill passes on a quiet day and fails during a real regional event. Hold a reservation or a warm standby for the RTO-critical path.
  • Secrets in the lost region. KMS keys or API keys scoped to the primary region block decryption of the backups. Test decryption in the drill.
  • Drill shortcuts. Pre-warming nodes the night before, or an expert running steps from memory, produces a time that the incident will not reproduce.

Trade-offs

ChoiceGainsCosts
Warm standbyRTO in minutes; capacity assuredIdle GPUs every hour of the year
Reservation, cold nodesCapacity assured at lower costBoot, pull and load time in the RTO
On-demand in recovery regionCheapestRTO unknown during regional events
Degraded smaller modelServes on scarce GPUsLower quality; must be validated and announced
Full traffic drillsReal evidenceUser-facing risk; needs a rollback path

What to do next

  1. List your assets with size, location, recovery method, RTO and RPO for each of the three scenarios: regional loss, cluster loss and account compromise.
  2. Produce a release manifest with SHA-256 digests for every artifact and store it with the backups.
  3. Measure the weight copy into your recovery region today, with the tool you would actually use.
  4. Decide how GPU capacity in the recovery region is assured, and write down the RTO that decision implies.
  5. Build the golden-set behaviour check and record a baseline score for the current release.
  6. Script the drill with per-step timing and evidence, run a partial drill this month, and schedule the full drill and the failback drill.
Key takeaway: An LLM disaster recovery plan is only as good as its last drill. Set RTO and RPO per asset and per scenario, recognize that weight transfer and GPU capacity dominate recovery time, keep backups out of reach of production credentials, and refuse to shift traffic until checksums and a golden-set eval prove the restored model is the released one. Script the drill, record its evidence, and run it on a schedule and after every change to the restore path.