Choosing a cloud region looks like a dropdown on the first day of a project and turns out to be one of the stickiest architecture decisions you make. Data accumulates where it is written, encryption keys, IP allowlists and private network links are regional, customers sign contracts that name a jurisdiction, and moving a production database across a continent later is a project measured in months. Many teams end up in the console's default region simply because nobody decided otherwise.

This page gives you a repeatable method. It separates hard constraints, which eliminate regions outright, from soft factors, which you score; explains the physics of network latency and how to measure it from where your users actually are; shows how to check capacity and service availability with real CLI commands; covers cost, failure domains and hidden cross-region dependencies; and ends with a scoring script, a worked example and a checklist. It is provider-agnostic, with AWS, Azure and Google Cloud commands where they help.

Advertisement

Why the region decision is architecture, not configuration

A region is a geographic cluster of data centres with its own control plane, typically split into several availability zones (AZs) with independent power, cooling and networking a few kilometres to tens of kilometres apart. Almost everything you create is scoped to one region: VPCs and subnets, managed databases, KMS keys, object storage buckets, container registries, private endpoints. Cross-region traffic crosses a billing boundary and adds tens to hundreds of milliseconds of latency.

That scoping creates data gravity. Once terabytes of data and the services that read them sit in one region, every new component is pulled there too, because putting it elsewhere adds latency and egress charges. The decision therefore deserves the same rigour as choosing a database: written requirements, measured inputs, a recorded rationale and defined triggers for revisiting it.

The method at a glance

Treat region selection as a pipeline. Hard filters come first and are binary; only survivors are measured and scored; the result is a primary region, a secondary for disaster recovery, and a list of conditions that would reopen the decision.

requirementsusers, data classes, SLOs, budgethard filtersresidency, services, capacitysurvivorsmeasurementsRTT from users, prices, quotaweighted scorelatency, cost, risk, carbonprimary regionplus secondary, 2+ AZsIaC configregion as a variabledependency mapSaaS, model APIs, KMS, DNSreview triggersnew market, law, outage, pricere-run the whole pipelinefilters first, scores second: a great score cannot fix a failed filter
Region selection as a pipeline: requirements feed binary filters, survivors are measured and scored, the chosen primary and secondary become IaC variables, dependencies are mapped, and explicit triggers re-run the process.
Advertisement

Step 1: hard filters

Apply these before looking at a single latency number, because each one can eliminate a region regardless of how well it scores elsewhere.

  • Data residency and sovereignty. Laws, regulators and customer contracts may require personal, health, financial or government data to stay in a jurisdiction. Classify your data first; the rule may apply to one data class, not the whole system. Remember that residency covers backups, logs, analytics copies and support tooling, not only the primary database.
  • Isolated partitions. Government and sovereign workloads may need a separate partition, such as AWS GovCloud (US) or Azure Government, with its own accounts, endpoints and service catalogue.
  • Service availability. New managed services, GPU instance families and AI model endpoints launch in a subset of regions first. Check each provider's regional services list for every managed service in your design, not just the headline ones.
  • Region enablement. On AWS, regions launched after 20 March 2019 are opt-in and must be enabled per account; organisation policies may block them. Know this before you promise a launch date.
  • Capacity for your scale. Offering a GPU type is not the same as having capacity for your fleet. Quotas and real availability are separate checks, covered below.

Step 2: latency from first principles

Light in optical fibre travels at about two-thirds of its speed in vacuum, roughly 200 km per millisecond. A round trip therefore costs at least 1 ms per 100 km of fibre route, and real routes are longer than the great-circle distance and pass through routers, so measured round-trip times are commonly well above the theoretical minimum. London to Frankfurt is about 640 km in a straight line: at least 6.4 ms round trip in theory, in practice more. Across the Atlantic or from Europe to South Asia you are in the tens to low hundreds of milliseconds, and no amount of engineering removes that.

What users feel is round trips multiplied by protocol chattiness. A new HTTPS connection needs a TCP handshake plus a TLS handshake before the first byte of the request; a page whose backend makes ten sequential calls to a database in another region pays ten cross-region round trips. So the latency question has two halves: how far are users from the region, and how far are the region's own dependencies from each other. Content delivery networks and edge termination hide the first half for static and cacheable content, not for writes or personalised reads.

Step 3: measure from where users are

Measure RTT from your users' networks, not from your office. The best source is real-user monitoring from an existing product; if you have none, run probes from cloud instances or third-party vantage points in each major user location. A minimal probe measures TCP connect time to a stable regional endpoint:

import socket, statistics, time

ENDPOINTS = {                         # regional service endpoints; any stable host works
    "eu-central-1": "ec2.eu-central-1.amazonaws.com",
    "eu-west-1":    "ec2.eu-west-1.amazonaws.com",
    "eu-west-3":    "ec2.eu-west-3.amazonaws.com",
}

def tcp_rtt_ms(host, port=443, samples=20):
    out = []
    for _ in range(samples):
        t0 = time.perf_counter()
        with socket.create_connection((host, port), timeout=3):
            pass                      # TCP handshake is about one round trip
        out.append((time.perf_counter() - t0) * 1000)
        time.sleep(0.2)
    out.sort()
    return statistics.median(out), out[int(0.9 * len(out)) - 1]

for region, host in ENDPOINTS.items():
    p50, p90 = tcp_rtt_ms(host)
    print(f"{region:14s} p50={p50:6.1f} ms  p90={p90:6.1f} ms")

Collect medians and a high percentile per user location and region, then weight them by where your traffic actually comes from. A region that is 5 ms better for 10% of users and 20 ms worse for 60% is a bad trade. Re-run the probes at several times of day; peering congestion is real and time-dependent.

Step 4: capacity, quotas and service parity

For CPU workloads, capacity is rarely the deciding factor in large regions. For GPUs and newer instance families it often is. Check which zones offer the types you need, then request quota, then test real launches at the size you plan to run.

# AWS: which AZs in a region offer an instance type (offering, not free capacity)
aws ec2 describe-instance-type-offerings --region eu-central-1 \
    --location-type availability-zone \
    --filters Name=instance-type,Values=g6.xlarge --output table

# Azure: SKUs in a location, including zone and subscription restrictions
az vm list-skus --location westeurope --size Standard_NC --output table

# Google Cloud: accelerator types per zone
gcloud compute accelerator-types list --filter="zone ~ europe-west4"

These commands show what is offered, not what is free right now. Diversifying across instance types and zones, and using flexible purchasing, is covered in spot capacity architecture. Also check parity for the boring services: a region that lacks the managed queue or the specific database engine version you rely on forces either a redesign or a cross-region call on your hot path.

Step 5: cost

Prices for compute, storage and managed services differ between regions of the same provider, and the differences can be material for large fleets. Pull them from the provider's pricing API or calculator for your actual bill of materials rather than comparing one instance type. Then add the costs that region choice creates indirectly: cross-region replication traffic, internet egress to users, inter-AZ traffic inside the region, and the price of any service you must reach in another region. For many data-heavy systems, transfer costs outweigh the compute price difference between neighbouring regions.

Step 6: failure domains and hidden dependencies

Use at least two, preferably three, AZs within the primary region for anything that must stay up, because AZ failures are the common case. Region-wide events are rarer but real, so choose a secondary region for backups and, if your recovery objectives demand it, warm or active standby; multi-region inference outage recovery walks through failover design.

Choosing the secondary is a trade-off. Close regions give low replication lag and allow near-synchronous replication, but are more exposed to correlated events such as a regional power-grid or natural disaster. Distant regions are more independent but force asynchronous replication and a non-zero recovery point. Azure publishes region pairs with sequenced platform updates, although some newer Azure regions have no pair; AWS and Google Cloud leave the choice to you. Stay inside the same residency boundary unless your data classification allows otherwise.

Finally, map dependencies that are not in your region at all: identity providers, SaaS APIs, third-party model endpoints, DNS and global services whose control planes may be homed in a single region, and CI/CD systems. An application can be perfectly multi-AZ and still stop deploying when a remote dependency does.

Worked example: an EU SaaS with GPU inference

A B2B product has customers in Germany, France and the Netherlands, stores customer personal data that contracts require to remain in the EU, and serves a model on mid-range GPUs. The hard filters remove every non-EU region, including a cheaper non-EU region with excellent latency. Of the EU regions, the team keeps the three whose service lists include every managed service in the design and whose zone-offering queries list the GPU type. Probes from customer networks give median RTTs in the low teens of milliseconds for all three, with one consistently a few milliseconds faster for the German majority. Quota requests succeed in two of them immediately.

The scoring script below encodes the result. The numbers are illustrative, chosen for the example rather than measured, and the weights reflect this team's priorities: latency matters most because the product is interactive, capacity next because GPU scaling is the main operational risk.

WEIGHTS = {"latency": 0.40, "cost": 0.25, "capacity": 0.20, "carbon": 0.05, "ops": 0.10}

candidates = [
    # scores are 0..10, higher is better; residency and service flags are hard filters
    dict(name="region-A", residency=True, services=True,
         latency=9, cost=6, capacity=8, carbon=5, ops=9),
    dict(name="region-B", residency=True, services=True,
         latency=7, cost=8, capacity=6, carbon=8, ops=8),
    dict(name="region-C", residency=False, services=True,
         latency=10, cost=9, capacity=9, carbon=7, ops=9),
]

def score(c):
    return sum(WEIGHTS[k] * c[k] for k in WEIGHTS)

eligible = [c for c in candidates if c["residency"] and c["services"]]
for c in sorted(eligible, key=score, reverse=True):
    print(f'{c["name"]}: {score(c):.2f}')
# region-C never appears: a failed filter is not a low score

The output ranks region-A first, with region-B as the secondary, because it is a different region inside the EU with acceptable latency and better carbon intensity. The record of the decision includes the filters applied, the raw probe data, the weights and the triggers for revisiting it. Six months later, when the team wins a large customer in the Nordics, they re-run the same pipeline rather than arguing from memory.

Sustainability and other soft factors

Grid carbon intensity varies widely between regions. Google Cloud publishes a carbon-free energy percentage per region, and the other providers publish sustainability data at varying levels of detail; if your organisation reports emissions, include it with a modest weight. Other soft factors worth scoring: support coverage and time zone of your operations team, the provider's history of incidents in the region, and whether your partners and customers already run there, which can enable private connectivity instead of internet paths.

Failure modes

  • Default-region drift. Resources land in whatever region a console or CLI profile defaults to, and a second, accidental footprint appears. Enforce allowed regions with organisation policies.
  • Measuring from the wrong place. An office on a corporate network is not a user on a mobile carrier in another country.
  • Residency leaks. The database is in-region, but logs ship to a global observability vendor, backups copy to another continent, or a model API processes prompts elsewhere.
  • Cross-region hot paths. A single dependency left in the old region adds a round trip to every request after a migration.
  • Assuming capacity. A GPU type listed in a region does not mean you can launch hundreds of them next week.
  • Hard-coded region strings. Region names embedded in code, bucket names and ARNs make every future move expensive; keep the region a variable in infrastructure-as-code from day one, as recommended in landing zone architecture.

Operating the decision

Write the decision down as a short record: requirements, filters, candidates, measurements, weights, the chosen primary and secondary, and review triggers. Typical triggers are entering a new market, a change in regulation or a customer contract, a provider launching a service or instance family you need, a sustained price change, or a major regional incident. Keep latency probes running continuously so that when a trigger fires you already have data. And keep migrations cheap: region as configuration, replication-friendly storage, and no region-specific identifiers in application code.

What to do next

  1. Classify your data and write down which classes carry residency requirements, including backups and logs.
  2. List every managed service, instance type and external dependency in your design, and check regional availability for each.
  3. Deploy the RTT probe, or pull real-user data, for your top user locations against every candidate region.
  4. Query zone offerings and request quota for your capacity-critical instance types in the finalists; test real launches.
  5. Score the survivors with explicit weights, choose a primary and a secondary within your residency boundary, and record the decision.
  6. Enforce allowed regions with policy, parameterise region in IaC, and set calendar and event-based review triggers.
Key takeaway: Pick regions with a pipeline, not a dropdown: eliminate regions that fail residency, service, partition or capacity requirements; measure round-trip latency from real users and account for protocol chattiness; compare full bills including transfer; use multiple AZs and a deliberately chosen secondary region; map dependencies outside the region; then score survivors with written weights and record the decision with triggers to revisit it. The physics of about 1 ms per 100 km of fibre and the gravity of data make the first choice expensive to undo, so make it on purpose.