Most platforms make you choose a region and then hide the network behind a load balancer that lives in that region. Fly.io turns that around. Every app gets addresses that are announced from many locations at once, so a user in Sydney and a user in Frankfurt connect to the same IP and land on different Fly edges. From there the Fly Proxy forwards the connection to a Machine running your code, preferring one that is close and not overloaded. Deploying to a new region is a scaling command, not a new architecture.

That convenience hides real decisions. Requests are routed to compute, but your data usually lives in one place, so a globally spread app that writes to a single Postgres primary can be slower than a single-region app. This article explains how anycast and the proxy actually route traffic, how regions and the primary region are configured, how to send writes to the primary with the fly-replay header, how Machines in different regions find each other, and the failure modes that catch teams in their first multi-region month. Details were checked against the Fly.io documentation in October 2026; platform behaviour changes, so confirm limits against the current docs before relying on them.

From anycast IP to Machine

Anycast means one IP address is advertised over BGP from many points of presence at the same time. Internet routers pick the path they consider shortest, which is usually, but not always, a geographically nearby site. The client never learns which site answered; it just sees a TCP handshake complete quickly. Fly uses this for the public addresses of every app, so you do not run a DNS-based geo steering layer to get users to a nearby entry point. The same principle is behind anycast DNS resolvers, covered in Anycast DNS explained.

User in SydneyUser in Parissame anycast IPsame anycast IPEdge: sydTLS, Fly ProxyEdge: cdgTLS, Fly ProxybackhaulbackhaulMachines in sydread replica nearbyMachines in amsprimary_region, writesfly-replay: region=amsBGP delivers each user to a nearby edge; the proxy picks the least loaded close Machine;the app itself redirects writes to the primary region with fly-replay.
Two users, one address. Each lands on a nearby edge, the proxy forwards to a Machine, and requests that need the primary database are replayed to the primary region.

At the edge, the Fly Proxy terminates TLS for HTTP services (or passes raw TCP through, depending on the handlers you configure) and decides which Machine should receive the request. That Machine may be in the same region as the edge or in another one; traffic travels between Fly servers over the platform's private backbone. The documentation describes the choice as the least loaded, closest Machine, where closeness is the measured round-trip time between the edge that received the request and the worker host running the Machine. Note what is not in that sentence: the client. If BGP sends a user to an unusual edge, closeness is computed from there.

Two consequences follow. First, an app with Machines in one region still benefits from edges everywhere: TLS handshakes complete at a nearby edge and the slow leg rides the backbone. Second, adding a region changes where requests execute, not where data lives: a page making five sequential queries to a database 150 ms away costs most of a second, however close the Machine is to the user.

Regions and the primary region

A region is a Fly location identified by a three-letter code such as ams, iad or syd. The regions page listed 17 of them when this was written; run fly platform regions for the live list rather than trusting a copy. Every Machine is placed in exactly one region and learns which through the FLY_REGION environment variable. The app also has a primary region, set as primary_region in fly.toml; deploys use it to decide where new Machines are created, and it is exported to every Machine as PRIMARY_REGION. Code compares the two to decide whether it can write locally.

app = "orders-web"
primary_region = "ams"

[http_service]
  internal_port = 8080
  force_https = true
  auto_stop_machines = "stop"
  auto_start_machines = true
  min_machines_running = 1          # counted in the primary region only

  [http_service.concurrency]
    type = "requests"
    soft_limit = 40
    hard_limit = 60

  [[http_service.checks]]
    interval = "15s"
    timeout = "2s"
    grace_period = "10s"
    method = "GET"
    path = "/healthz"

Adding capacity elsewhere is done with the scaling commands, for example fly scale count 2 --region syd or by cloning an existing Machine with fly machine clone <id> --region syd; fly scale show confirms the result per region. Machines in non-primary regions still run the same image and config. Treat the region list as part of your deployment review, because each region you add multiplies the number of places a bad release, a missing secret or an exhausted connection pool can show up.

How the proxy chooses a Machine

Load, in the proxy's model, is concurrency against the limits in [http_service.concurrency]. With type = "requests" the proxy counts in-flight HTTP requests, which suits most web services; with "connections" it counts open connections, which suits long-lived TCP or WebSocket traffic. Below soft_limit a Machine is a preferred target. Between the soft and hard limit it only receives traffic when every other candidate is also past its soft limit. At hard_limit it receives nothing new.

The rule that surprises people concerns regions. Per the load-balancing reference, cross-region routing only happens when all Machines in the local region are unhealthy or at their hard limit. Soft limits do not push traffic abroad. So a region with one busy Machine will queue requests up to the hard limit before spilling to an idle region next door. Set the hard limit to what the process can actually serve with acceptable latency, measured under load, not to a large number chosen to avoid rejections.

Autostop interacts with all of this. With auto_stop_machines = "stop" (or "suspend", where supported) the proxy stops idle Machines, and with auto_start_machines = true it starts one when a request arrives and nothing suitable is running. min_machines_running keeps a floor, but only in the primary region. A secondary region with every Machine stopped will either cold-start one or, if the proxy finds a running Machine elsewhere first, serve from farther away. If you want warm capacity in Sydney, autostop is the wrong default for Sydney.

Sending writes home with fly-replay

The standard multi-region pattern on Fly is a single write primary with read replicas: the database primary lives in primary_region, replicas live in the others, and app Machines in every region read locally. Writes must reach the primary. Instead of opening a cross-region database connection, the app answers the request with a response header, fly-replay, and the proxy discards that response and replays the original request to the target. The documented fields include region (one or a comma-separated list), instance and prefer_instance for a specific Machine, app for another app, elsewhere=true to exclude the replying Machine, state to pass a note that arrives in the fly-replay-src header, plus timeout and fallback.

import os
from flask import Flask, request, Response

app = Flask(__name__)
HERE = os.environ.get("FLY_REGION")
PRIMARY = os.environ.get("PRIMARY_REGION")
WRITE_METHODS = {"POST", "PUT", "PATCH", "DELETE"}

@app.before_request
def send_writes_to_primary():
    if request.method in WRITE_METHODS and HERE != PRIMARY:
        # The proxy throws this response away and re-sends the original request.
        return Response(status=409, headers={"fly-replay": f"region={PRIMARY}"})

@app.after_request
def mark_write(resp):
    # current_lsn() stands for your database's WAL or binlog position after the commit.
    if request.method in WRITE_METHODS:
        resp.set_cookie("last_write", str(current_lsn()), max_age=10, httponly=True)
    return resp

Two limits shape the design. Requests larger than 1 MB cannot be replayed, so uploads should go straight to object storage with a signed URL, or be accepted in any region and then handed to the primary through a queue. And replay only moves the request; it does not fix read-your-writes. After a write in Amsterdam, the user's next read may hit a Sydney replica that has not caught up. The cookie in the sketch records the write position; a read handler that sees a recent cookie can either wait until the local replica passes that position or replay itself to the primary. Fly also documents replay caching, keyed by path or by a session cookie or header, which lets the proxy remember a decision for a minimum of 10 seconds. Cached replays cannot carry state, timeout or fallback, so use them for stable routing such as tenant pinning, not per-request logic.

Clients can express a preference too. Request headers fly-prefer-region and fly-force-region ask the proxy to route to a region, with or without falling back to the nearest one, and fly-prefer-instance-id and fly-force-instance-id do the same for a specific Machine. They are handy for testing a region from your laptop and for debugging a single bad Machine.

Private networking across regions

Machines in an organisation share a private IPv6 network called 6PN, and Fly runs an internal DNS server that answers names under .internal. <app>.internal returns the private addresses of every started Machine of the app, <region>.<app>.internal narrows that to one region, top1.nearest.of.<app>.internal returns the closest, and regions.<app>.internal lists the regions where Machines are running. Only started Machines appear, so a stopped Machine is invisible to DNS and will not be woken by a direct connection.

That last point is why Flycast exists. A Flycast address is a private address that goes through the Fly Proxy, so internal callers get proxy features such as load balancing and autostart. Use .internal names when you want to address a specific region or Machine directly, for example a replica in your own region, and Flycast when you want the proxy to choose and to wake Machines on demand. Mixing them up produces a familiar bug: a worker app that scales to zero and a caller that connects through .internal and finds nothing.

Worked example: one app, three regions

Take an order-tracking app with users in Europe, the US east coast and Australia. The team picks ams as the primary region because most writes come from Europe, puts the Postgres primary there, and adds read replicas in iad and syd. App Machines run in all three regions, two per region, with autostop disabled in the secondary regions so a warm Machine is always there.

A read from Sydney lands on the Sydney edge, the proxy picks a Sydney Machine below its soft limit, and the handler queries syd.orders-db.internal: two short local hops. A status update from Sydney reaches the same Machine, which returns fly-replay: region=ams; the proxy replays the request to Amsterdam, where the write commits locally and the response travels back. The write costs one long round trip, roughly what any client would pay to reach Europe, instead of several long round trips for each statement in a transaction. The follow-up read carries the last-write cookie; the Sydney handler checks the replica's replay position, waits up to 200 ms, and replays itself to ams if the replica is still behind.

Before launch the team measures read and write p99 and replica lag per region; if lag regularly exceeds the wait, they route reads for recent writers to the primary.

Failure modes

  • Chatty code in the wrong region. App Machines are spread out but every query goes to one primary. Symptom: secondary regions are slower than before the expansion. Fix: replay writes, read from local replicas, and batch round trips.
  • Expecting soft limits to shed load across regions. Traffic stays local until every local Machine is unhealthy or at its hard limit, so one region saturates while another idles. Fix: size each region for its own peak and set hard limits from load tests.
  • Cold starts in secondary regions. min_machines_running only counts the primary region, so autostopped secondary regions start from zero. Fix: disable autostop there or keep Machines running with a scheduled warm-up.
  • Replaying large requests. Bodies over 1 MB cannot be replayed and the write fails in non-primary regions only, which looks intermittent. Fix: upload to object storage and replay a small reference.
  • Replay loops. A handler that replays to the primary region while the primary itself is unhealthy can bounce requests. Fix: replay only when FLY_REGION differs from PRIMARY_REGION, set a replay timeout, and alert on replay rate.
  • Stale reads after writes. Users see their change vanish on refresh. Fix: track the write position per client and wait or replay as in the worked example.

Trade-offs

Anycast removes a whole layer, DNS-based geo steering with its TTL delays, from the problem of getting users to a nearby entry point, and failover between edges happens at BGP speed. What it does not give you is control over the client-to-edge mapping or visibility into why a given user landed where they did. If you need hard data residency per user, enforce it in the app with replay and region checks, not by trusting the network path. For broader patterns, see Multi-region cloud architecture and the DNS-based alternative in DNS failover and traffic steering.

The single-primary model is simple and correct but puts a long round trip in every write from far away. Multi-primary designs remove that at the price of conflict handling, and most apps do not need them. Spreading Machines widely improves latency for read-heavy, cache-friendly workloads, and buys little for write-heavy ones; a single region with edges everywhere is a strong baseline. Compare with a classic regional balancer in Cloud load balancers and with code at the edge in Edge compute.

What to do next

  1. Draw the request path for your top three endpoints: which region executes them and where every network dependency lives.
  2. Set primary_region to where your write database lives, and read FLY_REGION and PRIMARY_REGION in code.
  3. Load-test one Machine and set soft_limit and hard_limit from the concurrency at which p99 latency degrades.
  4. Add a second region with fly scale count or fly machine clone, and decide explicitly whether autostop is on there.
  5. Implement write replay with fly-replay and a read-your-writes rule, and keep request bodies under 1 MB or move uploads to object storage.
  6. Switch internal callers to <region>.<app>.internal for local replicas and to Flycast where you rely on autostart.
  7. Dashboard p99 by region, replay rate and replica lag, and test a region outage by stopping its Machines.
Key takeaway: Anycast gets every user to a nearby Fly edge, and the proxy sends each request to the least loaded close Machine, crossing regions only when the local ones are unhealthy or at their hard limit. Regions move compute, not data, so keep writes in the primary region with fly-replay, read from local replicas, handle read-your-writes explicitly, size each region for its own peak, and remember that min_machines_running only protects the primary region.