Every global service has to answer the same question for each user: which of my locations should handle this connection? The two families of answers are anycast, where many sites announce the same IP prefix and Internet routing delivers packets to one of them, and DNS steering, where an authoritative name server returns a different address depending on where it believes the user is. Each has a reputation ("anycast is always nearest", "DNS failover is slow") that is only partly true.

This article treats edge routing as an engineering decision with numbers attached. It separates the problem into layers, explains what signal each mechanism really uses to choose a site, works through failover time for both, and shows how production systems combine them. It assumes you know what BGP and DNS are; for the mechanics of announcements and catchments, read anycast architecture first.

Advertisement

Three layers, three decisions

Edge routing is easier to reason about as three independent choices that happen in sequence:

  1. Which IP address does the client use? Decided by DNS: the authoritative server's answer, filtered through resolver caches.
  2. Which physical site receives packets for that address? Decided by Internet routing. With unicast there is one answer; with anycast, BGP on each network between the user and you picks one of several sites announcing the prefix.
  3. Which origin or region serves the request? Decided by your own proxies at the edge, which can forward over a private backbone to any region.
Clientbrowser / appRecursive resolverISP or public DNSAuthoritative DNSgeo / latency mapqueryquery + ECS?answer, TTLcachedHealth + RUMper prefix, per regionLayer 1: which IP address (DNS steering)Clientconnects to IPInternet routingBGP picks a pathEdge site AEdge site BanycastLayer 2: which site receives packets for that IP (anycast or unicast)L7 proxyat the edgeRegion us-eastRegion eu-westLayer 3: which origin serves the request(proxy policy, backhaul over private network)
Layer 1 is controlled by your authoritative DNS and limited by resolver caching. Layer 2 is controlled by BGP announcements and limited by other networks' routing policy. Layer 3 is fully under your control.

Most confusion comes from comparing a layer-1 mechanism with a layer-2 one as if they were alternatives. They are not; many large deployments use both. The real question is how much steering you need at each layer and what control and failover speed each gives you.

How anycast chooses a site

With anycast you announce the same prefix, say 203.0.113.0/24, from every site. Each network that hears several announcements picks one using its own BGP policy: local preference (often favouring customers over peers over transit), then shortest AS path, then a series of tie-breakers. Geography and latency are not inputs. In practice, shortest AS path correlates with proximity, which is why anycast usually works well, but the set of users routed to each site, the catchment, is decided by other networks' commercial relationships.

Strengths follow directly. There is no dependency on resolvers, so steering applies to the very first packet and to clients with hard-coded IPs. Attack traffic is split across sites in proportion to catchments, which is why DNS and DDoS mitigation are almost always anycast. Removing a site is a BGP withdrawal; routers re-converge onto the next-best announcement without any client action.

Weaknesses follow just as directly. You cannot say "send this ISP's users to Frankfurt"; you can only influence routing with AS-path prepending, BGP communities your transit providers honour, selective announcements and better peering. When routing changes mid-connection, packets for an established TCP or QUIC flow arrive at a site that has no state for it, and the connection resets. For short HTTP requests this is invisible; for long downloads, WebSockets and video sessions it is a measurable error rate during every routing event. See BGP for the selection process in detail.

Advertisement

How DNS steering chooses a site

DNS steering returns different A or AAAA records depending on who is asking. The authoritative server never sees the user; it sees the recursive resolver that asks on the user's behalf. Geolocation steering maps the resolver's address to a country or region. Latency-based steering maps it to the region with the lowest measured latency from that network. Managed DNS services expose these as named policies; Amazon Route 53, for example, offers simple, failover, geolocation, geoproximity, latency, IP-based, multivalue answer and weighted routing.

The resolver problem is the main source of error. A user in Singapore using a public resolver whose nearest instance is in Hong Kong is steered as if they were in Hong Kong; a corporate network that sends all DNS through a head office in another continent is steered to the wrong continent entirely. EDNS Client Subnet (ECS, RFC 7871) lets a resolver pass a truncated client prefix, typically a /24 for IPv4, so the authoritative server can steer on the user's network rather than the resolver's. It is not universal: some large public resolvers deliberately do not send it for privacy reasons, and caching per subnet fragments resolver caches.

The second limit is caching. The answer is cached for its TTL by the resolver, and often longer by operating systems, language runtimes and long-lived connection pools that never re-resolve. DNS steering is therefore precise in space (you choose per prefix) but imprecise in time. The DNS architecture article covers TTL behaviour and resolver caching in depth.

Failover arithmetic

Failover speed is the number most teams guess wrong. Work it out explicitly for your configuration.

StepDNS failover (TTL 60 s)Anycast withdrawal
Detect failureHealth check every 10 s, 3 failures: about 30 sLocal health check: a few seconds
Change the answer or routeAuthoritative update: secondsWithdraw announcement: immediate
PropagateUp to 60 s for resolvers honouring TTLBGP re-convergence: typically seconds, occasionally minutes
Clients switchOn next resolution; pooled connections may hold the old IP until they failNext packet; established flows to the dead site break
Typical totalAbout 1.5 to 2 minutes for most users, a tail of much longerWell under a minute for most users

Two caveats keep the table honest. The DNS tail is real: resolvers that clamp TTLs upward, a JVM with a long or infinite DNS cache, and HTTP clients that keep connections open for hours all extend it, so measure your own clients rather than trusting the TTL. And anycast's speed only helps if the failure is visible to BGP. A site whose routers are fine but whose application is broken keeps attracting its full catchment and black-holes it. Tie announcements to application health, not just to the router being up.

Hybrid designs that work

Production systems mix the mechanisms so each layer does what it is good at:

  • Anycast DNS, steered answers. Your authoritative servers are anycast, so DNS itself is fast and attack-resistant, and they return geo or latency-steered unicast addresses for the application.
  • Anycast edge, private backhaul. One anycast address for every user lands traffic at the nearest edge site; an L7 proxy there terminates TLS and forwards to the best healthy region over your backbone. Region failover becomes a proxy configuration change with no DNS or BGP involvement. This is the model of most CDNs; see CDN architecture.
  • Regional anycast prefixes chosen by DNS. Announce one prefix per continent from the sites in that continent, and use DNS to pick the continent. You keep DNS control over coarse placement and anycast's fast failover inside each continent, and you limit how far a catchment surprise can send traffic.
  • Weighted DNS for migrations. Shift a small percentage of answers to a new region, watch error rates, then increase. This is DNS's strongest use and anycast has no clean equivalent.

Steering from real-user measurements

Static geography is a weak proxy for latency: an undersea cable cut or a poorly peered ISP makes the nearest region the slowest. The robust approach is to measure from real users and steer on the measurements. A small fraction of page loads fetch a tiny object from each candidate region and report timings, keyed by client prefix. A batch job turns those beacons into a steering map:

from collections import defaultdict
from statistics import quantiles

def build_map(beacons, min_samples=200):
    """beacons: iterable of (client_prefix, region, rtt_ms) from real-user measurement.
    Returns prefix -> regions ordered by p75 latency, best first."""
    by_prefix = defaultdict(lambda: defaultdict(list))
    for prefix, region, rtt in beacons:
        by_prefix[prefix][region].append(rtt)
    steering = {}
    for prefix, per_region in by_prefix.items():
        scored = {r: quantiles(v, n=4)[2] for r, v in per_region.items()
                  if len(v) >= min_samples}
        if scored:
            steering[prefix] = sorted(scored, key=scored.get)
    return steering

def answer(steering, ecs_prefix, resolver_prefix, healthy, capacity_ok, default_order):
    order = steering.get(ecs_prefix) or steering.get(resolver_prefix) or default_order
    for region in order:
        if healthy[region] and capacity_ok[region]:
            return region
    return next(r for r in default_order if healthy[r])   # last resort: any healthy region

The p75 rather than the mean keeps a few slow beacons from dominating. The minimum-sample rule stops a prefix with ten measurements from being steered on noise. The capacity check matters as much as health: if the best region for half your users is full, sending them there anyway turns a latency optimisation into an outage. The same measurements diagnose anycast: if users in one prefix show high latency to your anycast address but low latency to a specific regional unicast address, their network is routing them to the wrong site.

A worked example

A team runs an API in three regions, us-east, eu-west and ap-southeast, and has a 99.95% availability target. Users are 50% North America, 30% Europe, 20% Asia-Pacific. Many clients are mobile apps that hold connections open.

Option one is latency-based DNS with a 60 s TTL and health checks. Placement is good and controllable, but the failover estimate is about two minutes for most users and longer for apps that pool connections. A regional outage of ten minutes a quarter fits the target; frequent short failures would not.

Option two is an anycast edge with L7 proxies forwarding to regions. Failover between regions becomes a proxy decision measured in seconds, users always land at the nearest edge site, and TLS handshakes complete close to the user. The costs are running or buying an edge network, and connection resets when BGP shifts catchments, which the mobile app handles with retries on idempotent requests.

The team chooses option two through a CDN-style provider, keeps a latency-steered DNS name as an emergency bypass, and uses weighted DNS on that bypass name during region migrations. They publish RUM beacons to verify that each continent's users actually land in their own continent, and they add a client retry policy so a reset during a routing change costs one retry, not a failed checkout.

Failure modes

  • Black-hole site. An anycast site announces while its application fails. Withdraw on application health checks, from several vantage points.
  • Cascading withdrawal. A loaded site withdraws, its catchment moves to a neighbour, which overloads and withdraws in turn. Cap how many sites may withdraw automatically, and prefer prepending or partial withdrawal under load.
  • Catchment surprise. A large network prefers a transit path that sends its users to another continent. Detect it with RUM per network, fix it with communities, prepends or direct peering.
  • Stale DNS in long-lived clients. Pools never re-resolve, so DNS failover never reaches them. Set connection maximum lifetimes and runtime DNS cache TTLs explicitly.
  • Resolver mis-location. Users behind distant resolvers are steered badly. Use ECS where available, RUM per resolver, and anycast for the first hop if precision matters.
  • Health checks that test the wrong thing. A check that only fetches a static page passes while the database behind the region is down. Check a path that exercises the real dependencies.

Operational guidance

Keep TTLs as low as your DNS provider's query costs allow for names you might fail over, typically 30 to 60 seconds, and long for names you never move. Rehearse failover on a schedule: withdraw a site or drain a region during business hours and measure how long until traffic moves and how many errors users saw. Record per site and per region: request share, p50 and p95 latency by client network, connection reset rate, and health-check state over time. When a single region serves both reads and writes, remember that edge routing only moves traffic; whether the destination can serve it depends on your data layer, which geo-distributed systems covers.

What to do next

  1. Draw your three layers and write down, for each, what mechanism decides and who controls it.
  2. Compute your failover time with the table above using your real health-check interval, TTL and client connection lifetimes.
  3. Measure actual failover by draining one site or region during a planned window, and compare against the estimate.
  4. Instrument real-user latency beacons per client prefix and per region, and check that users land where you think they do.
  5. Tie anycast announcements or DNS health to application-level checks, and cap automatic withdrawals.
  6. Set explicit DNS cache TTLs and connection maximum lifetimes in your own clients and SDKs.
  7. Choose one hybrid pattern above and document why, including the emergency bypass path.
Key takeaway: Edge routing is three decisions: which IP (DNS), which site receives it (BGP, unicast or anycast) and which origin serves it (your proxies). DNS steering is precise per network but slow to change because of resolver and client caching; anycast fails over quickly and absorbs attacks but steers by other networks' policies and can reset long-lived flows. Most robust designs combine anycast at the edge with health- and measurement-driven steering behind it, and verify the result with real-user measurements.