Teams usually discover what their AI feature depends on during the first regional outage. The model endpoint fails over cleanly to a second region, and then every answer is useless because the vector index in that region is three weeks stale, or the prompt registry was only ever deployed in the primary region, or the tool that looks up orders still points at a database in the region that just went dark. The model was the part everyone planned for, and it was the smallest part of the problem.

This article follows an AI feature through a regional outage from detection to failback. It covers both your own region failing and your model provider's region failing while yours is healthy. The capacity side of running serving in more than one region, including idle GPUs, weight distribution, active-active versus active-passive, and conversations pinned to a region's KV cache, is covered in Multi-Region LLM Serving. Here the focus is recovery: what must be ready in the other region, how the decision is made, how traffic moves, and how it comes back.

Advertisement

An inference request depends on much more than the model

Draw the dependency map for one request, and mark each dependency's location. A typical retrieval-augmented assistant with tools touches an API gateway and authentication, a configuration and prompt registry, an embedding model for the query, a vector index, a document store for the retrieved chunks, the generation model, a safety classifier, one or more tool backends, a conversation store, secrets, and telemetry. Each is either global, regional with a replica elsewhere, or regional with no replica. The last category is the list of things that turn a region outage into a feature outage, regardless of how well the model fails over.

Do this exercise per feature, write it down, and review it whenever a dependency is added. The typical surprises are the embedding model, which was deployed next to the index and forgotten, and tool backends, which belong to other teams whose disaster recovery plans nobody has read.

Failing over the model is not enough: every regional dependency must be ready in the other regionGlobal traffic layerDNS / anycast / clientRegion A (impaired)Region B (takes over)Gateway + authModel endpointor provider regionVector indexsnapshot + lagConversation storeasync replicaTool backendsCRM, paymentsConfig + promptsregistrySafety classifier, secrets, telemetryGateway + authModel endpointwarm, quota reservedVector indexdual-ingestedConversation storeRPO: secondsTool backendsregional or remoteConfig + promptspre-pushedSafety classifier, secrets, telemetrydrainshiftreplicateStatic stability: nothing in Region B's failover path calls Region A's control plane
A regional failover for a retrieval-augmented assistant. Everything the request touches must exist in the takeover region with acceptable staleness, and the failover path must not depend on the impaired region's control plane.

RTO and RPO per AI asset

Recovery time objective, how long until the feature works again, and recovery point objective, how much recent data you may lose, are usually set for databases. An AI feature has several assets with very different answers, and each needs its own:

AssetTypical RPOWhat drives RTOPractice
Model weights and adaptersZero: immutable artifactsDownload and load time into GPU memoryReplicate to every region before a version is deployable
Prompts, configs, guardrail rulesZero: versionedConfig propagationPush to all regions on every change; serve locally
Vector indexIngestion lagRebuild time if missing (hours)Dual-ingest or replicate snapshots; alert on lag
Conversation and agent stateSeconds of replication lagStore failoverAsync replicate; accept small loss, detect gaps
KV and prefix cachesNot replicatedWarm-up under trafficPlan for higher latency after failover
Eval sets and telemetryHours acceptableNot on the request pathReplicate, but do not block failover on it

The vector index usually dominates. If the secondary region's index is rebuilt from source during an outage, recovery time is the rebuild time, often hours for a large corpus, and the embedding model must be the same version as the primary's or query vectors will not match document vectors. Either ingest into both regions continuously, or replicate snapshots with a monitored lag and a stated RPO such as "answers may miss documents added in the last 30 minutes".

Advertisement

Static stability: fail over without the failed region

A failover plan that needs to call the impaired region is not a plan. Static stability means the takeover region can serve the full load using only what it already has: capacity that is provisioned or can be scaled without the primary's control plane, credentials and secrets already present, configuration already pushed, quotas with the model provider already granted, and DNS or load balancer changes that are driven from a global or secondary control plane. Check each step of the runbook with one question: does this step need anything from the region that is down? If you cannot answer, assume yes.

For self-hosted inference, the hardest part is capacity. GPUs may not be available on demand in the secondary region during a large incident, because every other customer of that cloud is failing over too, and loading weights of tens or hundreds of gigabytes into a freshly started server takes minutes. That is why the decision between warm standby and scale-on-failover is really a decision about RTO, as discussed in the capacity section of the multi-region serving article.

Detecting and deciding

Detect region health from the outside. Run synthetic canaries from each region that exercise the full path, including retrieval, generation and a harmless tool call, and a second set from outside the region against its public endpoint. Correlate with the per-route breaker state described in Circuit Breaker Pattern. Then decide who pulls the trigger. Automatic failover is fast but can flap, can fail over for a problem that is actually global, such as a bad prompt deployed everywhere, and can overload the secondary. A common compromise: automate failover of stateless, read-only paths, and require a human decision for paths with state or side effects, with the automation preparing everything and the human confirming.

from dataclasses import dataclass
import time

@dataclass
class RegionHealth:
    canary_success: float      # fraction over the last window, full-path canaries
    breaker_open: bool
    ttft_p95_s: float

class FailoverController:
    """Hysteresis: fail over on sustained failure, fail back only after sustained recovery."""
    def __init__(self, fail_after_s=180, recover_after_s=1800, min_success=0.9):
        self.fail_after_s, self.recover_after_s = fail_after_s, recover_after_s
        self.min_success = min_success
        self.bad_since = None
        self.good_since = None
        self.failed_over = False

    def healthy(self, h: RegionHealth) -> bool:
        return h.canary_success >= self.min_success and not h.breaker_open

    def step(self, primary: RegionHealth, secondary: RegionHealth, now=None):
        now = now or time.time()
        if not self.failed_over:
            if self.healthy(primary):
                self.bad_since = None
            else:
                self.bad_since = self.bad_since or now
                if now - self.bad_since >= self.fail_after_s and self.healthy(secondary):
                    self.failed_over, self.good_since = True, None
                    return "FAIL_OVER"        # stateless paths: execute; stateful: page for confirm
        else:
            if self.healthy(primary):
                self.good_since = self.good_since or now
                if now - self.good_since >= self.recover_after_s:
                    return "BEGIN_FAILBACK"   # gradual, gated ramp; never a flip
            else:
                self.good_since = None
        return "HOLD"

Two checks in that logic matter. Failover only proceeds if the secondary is healthy, so a global problem does not trigger a pointless move. And failback requires a much longer period of good health than failover requires of bad health, so a primary that recovers for two minutes does not pull traffic back into another failure.

Executing the move

Traffic moves through whatever layer you control: DNS records with short TTLs, a global load balancer, anycast, or client-side region selection. DNS is the most common and the least precise, because resolvers and clients cache beyond the TTL; the mechanics are covered in Cloud DNS failover and traffic steering. Whatever the layer, handle three kinds of in-flight work explicitly.

  • Streaming generations in the failed region are lost. Clients should retry with the same request ID; the secondary regenerates from the start. Never try to splice a half-finished answer across regions.
  • Conversations resume from the replicated store, which may be missing the last few turns. Detect the gap using a turn counter the client also holds, and either replay missing turns from the client or tell the user plainly that the last message was not saved.
  • Agent runs with side effects must resume from a durable log with idempotency keys, so a tool call that may have executed in the failed region is reconciled before it is retried. The patterns are in Durable Agent Workflows.

When the provider's region fails and yours does not

With a hosted model, the more common case is that your region is fine and the provider's endpoint serving it is impaired. The options are to call the same provider's endpoint in another region, to switch to another provider, or to degrade. Cross-region calls add network latency, typically tens of milliseconds within a continent and more across oceans, which is minor next to generation time; the real constraints are contractual and capacity-related. Data-residency commitments may forbid sending a region's requests elsewhere, in which case the only honest options are another provider in the same region or a degraded mode. And quota in the other provider region must already exist at the size of the failover load, because quota increases are rarely granted in minutes during a provider incident.

Model-version parity also matters. The same model name may not be available at the same version in every provider region, and a different version can change output format enough to break parsers. Pin explicit versions per region, keep them aligned, and include every region's endpoint in your scheduled evaluation runs, following the fallback discipline in ADK model fallback.

Failback is the dangerous half

Failback happens when everyone is tired and the primary looks healthy, which is exactly when a second incident starts. The primary's prefix and KV caches are cold, its autoscaler may have scaled inference servers down to a minimum, and weights may need to be reloaded. Moving all traffic back at once produces a latency spike that can trip breakers and fail you over again. Ramp instead: 5, 25, 50 and then 100 percent, with each step gated on time to first token, error rate and canary success holding for a fixed period.

Reconcile state written during the outage before or during the ramp. Conversation turns written only in the secondary must reach the primary's store; documents ingested only in the secondary must reach the primary's index; tool actions taken in the secondary must be visible to reconciliation jobs in the primary. Treat failback as a planned migration with its own checklist, not as undoing the failover.

Worked example: a regional failover timeline

An EU assistant runs in two EU regions, with the secondary as warm standby at 30 percent capacity, a dual-ingested index and an asynchronously replicated conversation store. At 09:00 the primary's canaries begin failing on the retrieval step; the model endpoint is healthy but the vector database is unreachable. At 09:03 the controller reports sustained failure with a healthy secondary. Read-only question answering fails over automatically; the on-call engineer confirms failover for the refund workflow at 09:09 after checking that the payments backend is reachable from the secondary.

The secondary autoscaler adds GPU servers; new servers need about 6 minutes to load weights and pass readiness checks, so between 09:03 and 09:12 the service sheds background summarisation and serves interactive traffic with p95 time to first token of 2.4 seconds instead of 0.9. Conversation replication lag at failover was 4 seconds; 31 conversations lost a final turn and were shown a notice. The measured RTO was 3 minutes for read paths and 9 for the refund workflow, against objectives of 5 and 15. The primary recovered at 10:20; failback began at 10:50 after 30 minutes of good health and completed at 11:30 through four gated steps.

Drills and trade-offs

Recovery that has not been rehearsed does not exist. Evacuate a region in business hours on a schedule, per feature, and measure RTO and RPO against the objectives. Include a drill where only one dependency, such as the index or a tool backend, fails, because partial failures are more common than whole-region losses and are harder to detect. Then choose a standby posture with the numbers in hand: active-active costs the most in idle capacity but has the shortest RTO and warm caches; warm standby at partial capacity trades a few minutes of degraded service for much lower cost; cold standby depends on GPU availability during an incident and on weight-loading time, and is only acceptable where an RTO of an hour or more is.

What to do next

  1. Draw the dependency map for each AI feature and mark every regional dependency without a replica.
  2. Set RTO and RPO per asset: weights, prompts and config, vector index, conversation state, caches.
  3. Make the embedding model and index available in every serving region at matching versions, with monitored ingestion lag.
  4. Walk the failover runbook step by step and remove every dependency on the impaired region's control plane.
  5. Deploy full-path canaries per region and a failover controller with hysteresis; decide which paths fail over automatically.
  6. Reserve provider quota in the alternative regions and verify residency rules before routing there.
  7. Write a failback checklist with a gated ramp and state reconciliation.
  8. Run regional evacuation drills, including single-dependency failures, and record measured RTO and RPO.
Key takeaway: Regional recovery for AI features is decided by everything around the model: indexes, embeddings, prompts, conversation state, tools and quotas must already exist in the takeover region. Set RTO and RPO per asset, keep the failover path statically stable, decide with hysteresis, handle streams, conversations and side-effecting agent runs explicitly, fail back gradually with state reconciliation, and prove all of it with drills.