Running in several regions is the standard answer to the question of what happens when a whole cloud region fails. It is also one of the most expensive and error-prone decisions in cloud architecture, because the hard part is not starting servers in a second place. The hard part is the data: where each write is accepted, how fast it is copied, what is lost when a region disappears, and who decides that it has disappeared.

This article is vendor-neutral. The AWS and GCP multi-region articles map these ideas onto each provider's services, and choosing which regions to use is covered in region selection. Here we build the model you need to pick a topology, route writes safely, design the failover control plane, size capacity, and prove that the whole thing works before a real outage tests it for you.

Advertisement

What multi-region protects against, and what it does not

A region is a set of data centres in one metro area, usually split into availability zones with independent power and networking. Zones protect against a building or power failure. A second region protects against events that take out a whole region: a regional control-plane failure, a widespread network problem, a natural disaster, or a provider-side bad deploy that stays inside one region. It also serves users far away with lower latency and helps meet data-residency rules.

It does not protect against your own mistakes, which cause most outages. A bad deploy, a corrupt migration or a wrong configuration pushed to all regions at once takes every region down together, and replication faithfully copies corrupted data within seconds. Multi-region therefore has to come with staged deploys (one region at a time, with bake time), point-in-time backups that replication cannot overwrite, and no global configuration that changes everywhere at once. Without those, a second region adds cost and complexity while protecting against the least likely failure.

The physics you cannot negotiate

Light in optical fibre travels at about 200,000 km per second, which is about 5 microseconds per kilometre one way, or roughly 1 millisecond of round trip for every 100 km of cable. Cable paths are longer than straight lines. Two regions 1,000 km apart by map might have 1,300 km of fibre between them, so the round trip cannot be less than about 13 ms and in practice is higher after routers and queues.

This number decides your consistency options. A write that must be durable in two regions before it is acknowledged costs at least one inter-region round trip, and consensus protocols often need more. On another continent that is tens to over a hundred milliseconds per write. Asynchronous replication avoids the wait but means the second region is always slightly behind, and whatever is not yet copied when the first region fails is lost. That lag is your real recovery point objective (RPO), whatever the design document says.

Advertisement

Four topologies

Every multi-region design is a variant of one of four shapes. They differ in who accepts writes and how data reaches the other region.

TopologyWrites acceptedRPORTOMain cost
Backup and restoreOne regionBackup intervalHoursSlow recovery; restore must be rehearsed
Warm standby (active-passive)One regionReplication lagMinutesIdle standby capacity; failover is a rare, risky event
Home-region partitioned (active-active by tenant)Each region, for its own tenantsReplication lagMinutes, per tenantRouting directory, fencing, re-homing logic
Global consensusAny region, via quorumZero for committed writesSeconds, automaticWrite latency of a quorum round trip; three or more regions

Multi-writer replication, where two regions accept writes to the same record and reconcile later, is deliberately missing from the table as a general answer. It needs a conflict rule, and the usual default, last writer wins, silently discards one of two concurrent updates. It fits data that is naturally mergeable (counters built as CRDTs, append-only logs, user preferences where losing one change is acceptable) and little else. For money, inventory or anything with invariants, pick a topology with one writer per record.

The home-region pattern is often the sweet spot. Each tenant, user or account has one home region that accepts its writes; other regions hold replicas and can serve reads that tolerate staleness. Both regions do useful work every day, so the standby path is exercised constantly, and a regional failure affects only the tenants homed there.

Global traffic layerDNS or anycast, health checksRegion Aserves tenants homed in ARegion Bserves tenants homed in Bhome = Ahome = BPrimary for A's datareplica of B's dataPrimary for B's datareplica of A's dataasync replication, lag = RPOFailover controlleroutside both regionsfence + promoteshift weightsEach region is a primary for its own tenants and a warm standby for the other's.The controller needs no control-plane call into the failed region to act.
Home-region active-active. Each region is primary for its own tenants and holds async replicas of the other's. A controller outside both regions fences, promotes and shifts traffic.

Routing writes safely: home regions and fencing

The danger in any failover is two regions believing they are the writer for the same data, which is split brain. It happens when a region that is only partially failed, or cut off from the controller, keeps accepting writes after its tenants have been moved. The defence is a fencing token: every re-homing increases an epoch number, writes carry the epoch they were routed under, and the storage layer rejects any write whose epoch is older than the one it has recorded. A region that cannot renew its lease stops accepting writes on its own.

class HomeRegionRouter:
    """Route writes to the tenant's home region; refuse writes that carry
    a stale epoch so a demoted region cannot accept them after failover."""

    def __init__(self, directory, local_region):
        self.directory = directory      # replicated, read-mostly: tenant -> (home, epoch)
        self.local = local_region

    def handle_write(self, tenant, request):
        home, epoch = self.directory.lookup(tenant)
        if home != self.local:
            return forward(home, request)          # or redirect the client
        if not self.directory.i_hold_lease(self.local, epoch):
            raise Unavailable("lease lost; not accepting writes")
        # Stamp the epoch into the write so storage rejects it if the tenant
        # has been re-homed since (fencing token).
        return storage.write(tenant, request, fence=epoch,
                             idempotency_key=request.idempotency_key)

Two more details matter. Idempotency keys let clients retry a write safely after a failover, when they cannot know whether the old region committed it before failing. And the directory that maps tenants to home regions must itself be replicated to every region and readable when any one region is down, so it cannot live only in the region it describes.

The failover control plane

Failover is a small distributed system of its own: detect, decide, fence, promote, shift traffic, and later fail back. Detection must use probes from several vantage points, including outside the provider, because a region can look dead from one network and healthy from another. Decision needs hysteresis: a confirmation window so that a one-minute blip does not trigger a failover that costs more than the blip. Many teams keep a human approval for full regional failover and automate everything after the approval.

The controller must follow the principle AWS calls static stability: during the failure, recovery must not depend on anything in the failed region, and ideally on no control-plane actions at all. Promoting replicas and moving traffic weights should be operations on the healthy side. Capacity in the surviving region must already be running, because the provider's control plane in other regions may be overloaded by every other customer trying to scale up at the same moment.

def failover_loop(region, probes, directory, traffic, cfg):
    bad_since = None
    while True:
        healthy = probes.fraction_ok(region, window_s=60)   # from several vantage points
        if healthy < cfg.unhealthy_threshold:
            bad_since = bad_since or now()
        else:
            bad_since = None

        if bad_since and now() - bad_since > cfg.confirm_s:
            if cfg.require_human and not approvals.granted(region):
                page_oncall(region); sleep(10); continue
            # 1. Fence: bump the epoch for every tenant homed in the region.
            moved = directory.rehome_all(region, to=cfg.standby_for[region])
            # 2. Promote replicas; record the replication position we promoted at.
            lost = storage.promote(cfg.standby_for[region], tenants=moved)
            audit.log(region=region, tenants=len(moved), unreplicated_writes=lost)
            # 3. Shift traffic only after the new home accepts writes.
            traffic.set_weight(region, 0)
            return
        sleep(5)

Order matters: fence first, promote second, move traffic last. Moving traffic before promotion sends users to a region that cannot yet write. Traffic steering itself, with DNS time-to-live, anycast and health-checked records, is covered in the DNS failover article. Failback is a second failover in the other direction and deserves the same care, including waiting for the recovered region to catch up on replication before any tenant is moved back.

Worked example: capacity and RPO for a two-region SaaS

A SaaS product serves 6,000 requests per second at peak, split evenly across tenants homed in two regions. If one region fails, the other must carry all 6,000. So each region needs capacity for 6,000 requests per second while normally serving 3,000, which caps normal utilisation at 50 percent. With three regions each normally serves 2,000 and must absorb 3,000 when one fails, a ceiling of 67 percent. In general, with N regions each must be sized for the total load divided by N minus 1, so the normal utilisation ceiling is (N-1)/N. That is the arithmetic behind the cost of multi-region, and it applies to databases, caches and queues as well as application servers.

Now RPO. The team measures replication lag at its 99th percentile as 800 ms, with occasional spikes to 20 seconds during bulk imports. Their stated RPO of one second is therefore false for exactly the moments that matter, because failures are often correlated with load. They add an alert when lag exceeds 5 seconds, throttle bulk imports on lag, and report the number of unreplicated writes at every promotion, as the controller above does. For the few operations that must never be lost, such as payment captures, they use a synchronous write to a quorum spanning regions and accept the extra 30 ms on that path only.

Failure modes

  • Hidden global dependency. Sign-in, secrets, the container registry or the deploy system live in one region, so the second region cannot start or scale when the first is down.
  • Untested standby. The passive region has drifted: different configuration, an expired certificate, schema migrations not applied. It fails on the day it is needed.
  • Split brain. A partitioned region keeps writing after failover because nothing fences it. Merging the diverged data by hand takes days.
  • Replication lag grows silently. RPO is worse than documented and nobody finds out until the promotion report.
  • Retry storm. Clients retry in lockstep against the surviving region, which then fails too. Use backoff with jitter and load shedding.
  • Correlated deploys. A release goes to all regions at once and fails everywhere, turning a regional design into a global outage.
  • Residency violation. Replication copies regulated data to a region where it may not be stored. Decide per data class which regions may hold it.

Testing and operating it

An untested failover plan is a hope, not a plan. Run regional evacuation drills on a schedule: drain one region in production during business hours, measure the real RTO and the number of unreplicated writes, and fix whatever broke. Start with a small tenant set, then grow. Between drills, keep a dashboard with replication lag per data store, capacity headroom per region against the N-1 target, and the time since the last successful drill. Deploy region by region with bake time so that a bad release is caught in one region first. If the service calls AI inference providers, the same thinking applies to them; see multi-region inference and provider outage recovery.

What to do next

  1. Write down the failures you are protecting against and confirm that most outages (bad deploys, bad config) are handled by staged rollouts and backups, not by the second region.
  2. Measure inter-region round-trip time and replication lag at p99 and during bulk jobs; replace the documented RPO with the measured one.
  3. Pick one of the four topologies per data store, and give every record exactly one writer unless the data is genuinely mergeable.
  4. Add epoch fencing and idempotency keys to the write path before you automate failover.
  5. List every dependency the standby region needs at failover time and remove any that live only in the primary.
  6. Size each region for total load divided by N-1 and alert when headroom drops below it.
  7. Schedule a regional evacuation drill and record RTO, lost writes and every manual step.
Key takeaway: Multi-region is a data problem before it is an infrastructure problem. Physics sets the price of synchronous writes, asynchronous lag sets your real RPO, and every record should have exactly one writer, enforced by fencing. A failover controller outside the failed region must fence, promote and then move traffic, using capacity that is already running. Size each region for N-1, stage every deploy, and drill evacuations until the recovery time you measure matches the one you promise.