Running in several regions is the standard answer to the question of what happens when a whole cloud region fails. It is also one of the most expensive and error-prone decisions in cloud architecture, because the hard part is not starting servers in a second place. The hard part is the data: where each write is accepted, how fast it is copied, what is lost when a region disappears, and who decides that it has disappeared.
This article is vendor-neutral. The AWS and GCP multi-region articles map these ideas onto each provider's services, and choosing which regions to use is covered in region selection. Here we build the model you need to pick a topology, route writes safely, design the failover control plane, size capacity, and prove that the whole thing works before a real outage tests it for you.
What multi-region protects against, and what it does not
A region is a set of data centres in one metro area, usually split into availability zones with independent power and networking. Zones protect against a building or power failure. A second region protects against events that take out a whole region: a regional control-plane failure, a widespread network problem, a natural disaster, or a provider-side bad deploy that stays inside one region. It also serves users far away with lower latency and helps meet data-residency rules.
It does not protect against your own mistakes, which cause most outages. A bad deploy, a corrupt migration or a wrong configuration pushed to all regions at once takes every region down together, and replication faithfully copies corrupted data within seconds. Multi-region therefore has to come with staged deploys (one region at a time, with bake time), point-in-time backups that replication cannot overwrite, and no global configuration that changes everywhere at once. Without those, a second region adds cost and complexity while protecting against the least likely failure.
The physics you cannot negotiate
Light in optical fibre travels at about 200,000 km per second, which is about 5 microseconds per kilometre one way, or roughly 1 millisecond of round trip for every 100 km of cable. Cable paths are longer than straight lines. Two regions 1,000 km apart by map might have 1,300 km of fibre between them, so the round trip cannot be less than about 13 ms and in practice is higher after routers and queues.
This number decides your consistency options. A write that must be durable in two regions before it is acknowledged costs at least one inter-region round trip, and consensus protocols often need more. On another continent that is tens to over a hundred milliseconds per write. Asynchronous replication avoids the wait but means the second region is always slightly behind, and whatever is not yet copied when the first region fails is lost. That lag is your real recovery point objective (RPO), whatever the design document says.
Four topologies
Every multi-region design is a variant of one of four shapes. They differ in who accepts writes and how data reaches the other region.
| Topology | Writes accepted | RPO | RTO | Main cost |
|---|---|---|---|---|
| Backup and restore | One region | Backup interval | Hours | Slow recovery; restore must be rehearsed |
| Warm standby (active-passive) | One region | Replication lag | Minutes | Idle standby capacity; failover is a rare, risky event |
| Home-region partitioned (active-active by tenant) | Each region, for its own tenants | Replication lag | Minutes, per tenant | Routing directory, fencing, re-homing logic |
| Global consensus | Any region, via quorum | Zero for committed writes | Seconds, automatic | Write latency of a quorum round trip; three or more regions |
Multi-writer replication, where two regions accept writes to the same record and reconcile later, is deliberately missing from the table as a general answer. It needs a conflict rule, and the usual default, last writer wins, silently discards one of two concurrent updates. It fits data that is naturally mergeable (counters built as CRDTs, append-only logs, user preferences where losing one change is acceptable) and little else. For money, inventory or anything with invariants, pick a topology with one writer per record.
The home-region pattern is often the sweet spot. Each tenant, user or account has one home region that accepts its writes; other regions hold replicas and can serve reads that tolerate staleness. Both regions do useful work every day, so the standby path is exercised constantly, and a regional failure affects only the tenants homed there.
Routing writes safely: home regions and fencing
The danger in any failover is two regions believing they are the writer for the same data, which is split brain. It happens when a region that is only partially failed, or cut off from the controller, keeps accepting writes after its tenants have been moved. The defence is a fencing token: every re-homing increases an epoch number, writes carry the epoch they were routed under, and the storage layer rejects any write whose epoch is older than the one it has recorded. A region that cannot renew its lease stops accepting writes on its own.
class HomeRegionRouter:
"""Route writes to the tenant's home region; refuse writes that carry
a stale epoch so a demoted region cannot accept them after failover."""
def __init__(self, directory, local_region):
self.directory = directory # replicated, read-mostly: tenant -> (home, epoch)
self.local = local_region
def handle_write(self, tenant, request):
home, epoch = self.directory.lookup(tenant)
if home != self.local:
return forward(home, request) # or redirect the client
if not self.directory.i_hold_lease(self.local, epoch):
raise Unavailable("lease lost; not accepting writes")
# Stamp the epoch into the write so storage rejects it if the tenant
# has been re-homed since (fencing token).
return storage.write(tenant, request, fence=epoch,
idempotency_key=request.idempotency_key)Two more details matter. Idempotency keys let clients retry a write safely after a failover, when they cannot know whether the old region committed it before failing. And the directory that maps tenants to home regions must itself be replicated to every region and readable when any one region is down, so it cannot live only in the region it describes.
The failover control plane
Failover is a small distributed system of its own: detect, decide, fence, promote, shift traffic, and later fail back. Detection must use probes from several vantage points, including outside the provider, because a region can look dead from one network and healthy from another. Decision needs hysteresis: a confirmation window so that a one-minute blip does not trigger a failover that costs more than the blip. Many teams keep a human approval for full regional failover and automate everything after the approval.
The controller must follow the principle AWS calls static stability: during the failure, recovery must not depend on anything in the failed region, and ideally on no control-plane actions at all. Promoting replicas and moving traffic weights should be operations on the healthy side. Capacity in the surviving region must already be running, because the provider's control plane in other regions may be overloaded by every other customer trying to scale up at the same moment.
def failover_loop(region, probes, directory, traffic, cfg):
bad_since = None
while True:
healthy = probes.fraction_ok(region, window_s=60) # from several vantage points
if healthy < cfg.unhealthy_threshold:
bad_since = bad_since or now()
else:
bad_since = None
if bad_since and now() - bad_since > cfg.confirm_s:
if cfg.require_human and not approvals.granted(region):
page_oncall(region); sleep(10); continue
# 1. Fence: bump the epoch for every tenant homed in the region.
moved = directory.rehome_all(region, to=cfg.standby_for[region])
# 2. Promote replicas; record the replication position we promoted at.
lost = storage.promote(cfg.standby_for[region], tenants=moved)
audit.log(region=region, tenants=len(moved), unreplicated_writes=lost)
# 3. Shift traffic only after the new home accepts writes.
traffic.set_weight(region, 0)
return
sleep(5)Order matters: fence first, promote second, move traffic last. Moving traffic before promotion sends users to a region that cannot yet write. Traffic steering itself, with DNS time-to-live, anycast and health-checked records, is covered in the DNS failover article. Failback is a second failover in the other direction and deserves the same care, including waiting for the recovered region to catch up on replication before any tenant is moved back.
Worked example: capacity and RPO for a two-region SaaS
A SaaS product serves 6,000 requests per second at peak, split evenly across tenants homed in two regions. If one region fails, the other must carry all 6,000. So each region needs capacity for 6,000 requests per second while normally serving 3,000, which caps normal utilisation at 50 percent. With three regions each normally serves 2,000 and must absorb 3,000 when one fails, a ceiling of 67 percent. In general, with N regions each must be sized for the total load divided by N minus 1, so the normal utilisation ceiling is (N-1)/N. That is the arithmetic behind the cost of multi-region, and it applies to databases, caches and queues as well as application servers.
Now RPO. The team measures replication lag at its 99th percentile as 800 ms, with occasional spikes to 20 seconds during bulk imports. Their stated RPO of one second is therefore false for exactly the moments that matter, because failures are often correlated with load. They add an alert when lag exceeds 5 seconds, throttle bulk imports on lag, and report the number of unreplicated writes at every promotion, as the controller above does. For the few operations that must never be lost, such as payment captures, they use a synchronous write to a quorum spanning regions and accept the extra 30 ms on that path only.
Failure modes
- Hidden global dependency. Sign-in, secrets, the container registry or the deploy system live in one region, so the second region cannot start or scale when the first is down.
- Untested standby. The passive region has drifted: different configuration, an expired certificate, schema migrations not applied. It fails on the day it is needed.
- Split brain. A partitioned region keeps writing after failover because nothing fences it. Merging the diverged data by hand takes days.
- Replication lag grows silently. RPO is worse than documented and nobody finds out until the promotion report.
- Retry storm. Clients retry in lockstep against the surviving region, which then fails too. Use backoff with jitter and load shedding.
- Correlated deploys. A release goes to all regions at once and fails everywhere, turning a regional design into a global outage.
- Residency violation. Replication copies regulated data to a region where it may not be stored. Decide per data class which regions may hold it.
Testing and operating it
An untested failover plan is a hope, not a plan. Run regional evacuation drills on a schedule: drain one region in production during business hours, measure the real RTO and the number of unreplicated writes, and fix whatever broke. Start with a small tenant set, then grow. Between drills, keep a dashboard with replication lag per data store, capacity headroom per region against the N-1 target, and the time since the last successful drill. Deploy region by region with bake time so that a bad release is caught in one region first. If the service calls AI inference providers, the same thinking applies to them; see multi-region inference and provider outage recovery.
What to do next
- Write down the failures you are protecting against and confirm that most outages (bad deploys, bad config) are handled by staged rollouts and backups, not by the second region.
- Measure inter-region round-trip time and replication lag at p99 and during bulk jobs; replace the documented RPO with the measured one.
- Pick one of the four topologies per data store, and give every record exactly one writer unless the data is genuinely mergeable.
- Add epoch fencing and idempotency keys to the write path before you automate failover.
- List every dependency the standby region needs at failover time and remove any that live only in the primary.
- Size each region for total load divided by N-1 and alert when headroom drops below it.
- Schedule a regional evacuation drill and record RTO, lost writes and every manual step.