Anycast means giving the same IP address to many machines in many places and letting the routing system deliver each packet to one of them. The client does nothing special. It sends a packet to 192.0.2.53, and some router along the way decides which copy of that address is closest by its own rules. DNS root servers, public resolvers, CDN front doors and DDoS scrubbing networks all work this way.

The architecture view is covered in the anycast architecture article, and the choice between anycast and DNS-based steering is covered in anycast versus geo-DNS. This article is the hands-on manual. It follows a packet from a client to a socket, builds a health-gated announcement with ExaBGP, explains why hashing inside a site matters as much as routing between sites, and shows how to prove which site actually answered. If BGP itself is unfamiliar, read BGP basics for developers first.

Advertisement

From first principles: why the same address can live in many places

A router forwards a packet by looking up the destination in its forwarding table and picking the most specific matching prefix. The table does not say whether a prefix lives in one building or fifty. If three sites all announce 192.0.2.0/24 into BGP, every network on the internet receives up to three paths to that prefix and keeps the one its policy prefers. That choice is usually the shortest AS path after local preference, so it is topologically near, not necessarily geographically near or lowest latency.

The set of client networks that end up at a given site is called its catchment. Catchments are decided by other people's routing policy, which is why anycast is easy to start and hard to steer precisely. Three properties fall out of this model and drive every design decision below:

  • Selection is per packet, not per connection. Routers keep no session state. If a route changes mid-connection, the next packet can arrive at a different site that has never heard of the connection.
  • Failover is removal. A site stops receiving traffic only when it stops announcing the prefix. Anything that keeps the announcement up while the service is broken turns that site into a black hole for its whole catchment.
  • Every site must be able to serve every request. Anycast gives you no control over which client reaches which site, so configuration, data and certificates must be identical everywhere.
One prefix, three sites: BGP picks the site, ECMP picks the L4 node, a consistent hash picks the serverClient in ISP Adst 192.0.2.53Client in ISP Bdst 192.0.2.53Internet BGPbest path per ASSite FRAAS 64500Site AMSAS 64500Site SINAS 64500catchmentBorder routerECMP over L4 nodesL4 balancersMaglev hash + conntrackServersVIP on lo, serve locally5-tuple hashsame flow, same serverHealth checker + ExaBGPannounce while healthy, withdraw on drainBGPInside every site (expanded view of FRA)
Three layers of selection. BGP picks the site, the border router uses ECMP to pick an L4 balancer, and the balancer uses a consistent hash so that every packet of a flow reaches the same server.

What you need before the first announcement

RequirementWhyPractical note
A prefix of at least /24 (IPv4) or /48 (IPv6)Most networks filter longer prefixes, so a /25 will not propagateYou cannot anycast a single /32 on the public internet; you anycast a /24 and use addresses inside it
An origin ASThe prefix needs an origin in every announcementOne AS at all sites is simplest; mixed origins need a ROA per origin
ROAs in RPKINetworks doing route origin validation drop announcements whose origin or length does not matchCreate a ROA for each origin AS, with maxLength equal to the length you announce
Upstreams at each siteTransit and peering decide the size of each catchmentSimilar upstream mixes at every site give more even catchments
Identical service stacksAny client can land anywhereSame software version, data and TLS certificates; config pushed from one source

A common RPKI mistake is a ROA for a /22 with maxLength 22 while announcing /24s from it. Validating networks drop those /24s, so some catchments silently vanish. Check validity externally after every prefix or origin change.

Advertisement

Health-gated announcements with ExaBGP

The core control loop is simple: a server announces the service route only while it can serve. A common design runs a BGP speaker on each server that announces the service /32 to the site router, and the site router announces the covering /24 to the internet as long as at least one server is announcing. Draining a server is then a matter of withdrawing its /32. Draining a site means all of its servers withdraw, or the router stops exporting the /24.

ExaBGP is convenient here because it runs a helper process and turns lines that process prints into BGP updates. The service address is configured on the loopback interface of every server so the kernel accepts packets for it. BIRD can do the same job by exporting a static or direct route that a health script adds and removes; the logic is the same either way.

# /etc/exabgp/exabgp.conf on each server (ExaBGP 4.x)
process health {
    run /usr/local/bin/anycast-health.py;
    encoder text;
}

neighbor 10.20.0.1 {                 # the site's router or L4 tier
    router-id 10.20.0.11;
    local-address 10.20.0.11;
    local-as 64512;                  # private ASN per server
    peer-as 64500;
    api {
        processes [ health ];
    }
}
#!/usr/bin/env python3
"""Announce 192.0.2.53/32 only while the local DNS server answers real queries."""
import os, subprocess, sys, time

ROUTE = "192.0.2.53/32"
RISE, FALL, INTERVAL = 3, 3, 2.0       # 3 good checks to announce, 3 bad to withdraw
DRAIN_FILE = "/etc/anycast/drain"      # touch this to drain the server on purpose

def healthy():
    if os.path.exists(DRAIN_FILE):
        return False
    try:
        out = subprocess.run(
            ["dig", "@127.0.0.1", "health.example.net", "A", "+short", "+time=1", "+tries=1"],
            capture_output=True, text=True, timeout=2).stdout.strip()
        return out == "198.51.100.7"   # a known answer, not just "the port is open"
    except subprocess.TimeoutExpired:
        return False

def say(command):
    sys.stdout.write(command + "\n")
    sys.stdout.flush()                 # ExaBGP reads our stdout line by line

announced, good_run, bad_run = False, 0, 0
while True:
    if healthy():
        good_run, bad_run = good_run + 1, 0
    else:
        good_run, bad_run = 0, bad_run + 1
    if not announced and good_run >= RISE:
        say(f"announce route {ROUTE} next-hop self")
        announced = True
    elif announced and bad_run >= FALL:
        say(f"withdraw route {ROUTE} next-hop self")
        announced = False
    time.sleep(INTERVAL)

Three details carry most of the value. The check compares a real answer: a port check keeps announcing while the server returns SERVFAIL, creating a zombie site. The rise and fall counters add hysteresis; without them an intermittently failing server flaps its route, and route flap damping elsewhere may suppress your prefix long after the fault. The drain file takes a server out deliberately while in-flight queries finish.

Inside a site: why ECMP alone breaks TCP

Inside a site, the router usually has several equal-cost next hops for the service address, one per server or per L4 balancer, and spreads flows across them with ECMP. It hashes the five-tuple (source and destination address and port, plus protocol) and takes the result modulo the number of next hops. That keeps every packet of a flow on one machine, until the number of next hops changes.

With plain modulo hashing, adding or removing one of N next hops remaps most flows, not one Nth of them, and every remapped TCP connection lands on a server with no state for it and is reset. Single-packet DNS queries do not care; HTTPS and long-lived connections do. Some switches offer resilient hashing, but behaviour varies by platform.

The common fix is an L4 balancing tier, as in Google's Maglev and Meta's Katran. Each balancer picks a backend with a consistent hash, so all balancers agree without coordination, and keeps a connection table so existing flows stay put when backends change. Maglev fills its lookup table simply: each backend gets a pseudo-random permutation of slots, and backends take turns claiming their next free slot.

import hashlib

def _h(value, seed):
    digest = hashlib.sha256(f"{seed}:{value}".encode()).digest()
    return int.from_bytes(digest[:8], "big")

def maglev_table(backends, m=65537):          # m must be prime and much larger than len(backends)
    offsets = [_h(b, "offset") % m for b in backends]
    skips = [_h(b, "skip") % (m - 1) + 1 for b in backends]
    nxt = [0] * len(backends)
    table, filled = [None] * m, 0
    while True:
        for i, b in enumerate(backends):      # backends take turns claiming their next free slot
            while True:
                slot = (offsets[i] + nxt[i] * skips[i]) % m
                nxt[i] += 1
                if table[slot] is None:
                    break
            table[slot] = b
            filled += 1
            if filled == m:
                return table

def pick(table, src_ip, src_port, dst_ip, dst_port, proto):
    return table[_h((src_ip, src_port, dst_ip, dst_port, proto), "flow") % len(table)]

before = maglev_table([f"srv{i}" for i in range(10)])
after = maglev_table([f"srv{i}" for i in range(10) if i != 3])
moved = sum(1 for a, b in zip(before, after) if a != b and a != "srv3")
print(f"slots moved that did not belong to srv3: {moved / len(before):.2%}")   # small, not ~90%

Run it and the fraction of slots that move, other than the removed server's own, is small. That is the property you need: a change to one backend should disturb only the flows that belonged to it.

The traps that only appear with anycast

  • Path MTU discovery. When a router cannot forward a large packet with Don't Fragment set, it sends an ICMP Packet Too Big message to the source address, which is your anycast address. That ICMP packet can reach a different site, or within a site a different server, because its five-tuple differs from the flow it refers to. The real sender never lowers its packet size and the connection stalls on large responses. Mitigate by clamping the TCP MSS at the edge, using conservative MTUs for UDP responses, and making the balancer hash ICMP errors by the embedded original header.
  • Mid-connection route changes. Rare per client, constant across millions. Long downloads and WebSockets will occasionally reset, so clients must retry cleanly. QUIC connection IDs help routing inside a site, not across sites, because another site lacks the connection keys.
  • Asymmetric return paths. Responses leave via the receiving site's upstreams, so stateful firewalls expecting both directions drop traffic.
  • Uneven catchments. One large ISP's policy can send a whole country to the wrong continent.

Finding out which site answered

Every anycast investigation starts with the same question: where did this client actually land? Build the answer into the service from day one. For DNS, the NSID option and the CHAOS-class identity queries return a server-chosen identifier. For HTTP, make every site add a response header with its site and host name.

# Which DNS site answered? (needs the server to support NSID or CHAOS identity)
dig @192.0.2.53 example.net SOA +nsid          # look for "NSID:" in the OPT section
dig @192.0.2.53 id.server CH TXT +short        # many servers answer with a site/host name
dig @192.0.2.53 hostname.bind CH TXT +short    # older BIND-style equivalent

# Which path did packets take, and through which networks?
mtr -z -w -c 20 192.0.2.53                     # -z adds AS numbers per hop

# HTTP over anycast: have every site add a response header naming itself
curl -sI https://api.example.net/ | grep -i x-served-by

One test from your laptop shows one catchment. Run the same query from many vantage points, such as RIPE Atlas probes or client telemetry, and aggregate the site identifier by client AS and country. The resulting catchment map drives capacity planning and is the first thing to compare after any routing change.

Worked example: draining a site without dropping queries

Suppose an authoritative DNS service on 192.0.2.53 runs in three sites, each with four servers, using the configuration above. Monitoring from 200 vantage points shows 41 percent of probes landing in FRA, 37 percent in AMS and 22 percent in SIN. FRA needs a kernel upgrade.

  1. Check headroom first. If most of FRA's catchment shifts to AMS, AMS must absorb roughly 78 percent of traffic. If it cannot, stop and add capacity.
  2. Touch the drain file on one FRA server. After three failed checks, about 6 seconds, the script withdraws and the site router removes that next hop; connection tracking keeps existing flows in place.
  3. Repeat for the remaining servers. When the last withdraws, the router stops exporting the /24 from FRA and the internet converges on AMS and SIN, typically within seconds to minutes.
  4. Re-run the probes. FRA's share should be zero; any probe still reporting FRA points to a stale route.
  5. Upgrade, remove the drain files, and confirm the catchment returns to its old shape. A different shape means some network changed its routing policy meanwhile.

Failure modes and trade-offs

FailureSymptomPrevention
Shallow health checkOne site serves errors to its whole catchmentCheck with a real request and a known answer
No hysteresisRoute flaps, then damping suppresses the prefix elsewhereRise and fall counters; alert on announcement changes per hour
Modulo ECMP on backend changeConnection resets on every deployConsistent hashing plus connection tracking in an L4 tier
ROA mismatchSome networks cannot reach one or all sitesOne ROA per origin with exact maxLength; external validation
Lost ICMP Packet Too BigLarge responses stallMSS clamping and ICMP-aware balancing
Config drift between sitesBugs that only some users seeSame build and config everywhere, verified per site

The central trade-off is control versus simplicity. Anycast deletes the global traffic-direction tier and gives you fast failover and natural DDoS dilution, but it gives up the per-client control that DNS steering offers. Most large operators combine the two, as discussed in the routing comparison. For the routing policy levers, such as prepending and communities, see the BGP deep dive.

What to do next

  1. Confirm you have a /24 or /48 with ROAs matching every origin AS and announced length.
  2. Configure the service address on loopback and run a BGP speaker with a deep health check, rise and fall counters and a drain file.
  3. Put a consistent-hashing L4 tier with connection tracking in front of any TCP service.
  4. Add a site identifier to every response: NSID or CHAOS identity for DNS, a header for HTTP.
  5. Build a catchment map from many vantage points and plan capacity per site from it.
  6. Clamp MSS at the edge and test large responses from several networks.
  7. Rehearse a full-site drain each quarter and compare catchments before and after.
Key takeaway: Anycast works because routers choose among identical announcements by their own policy, so one address can be served from many sites with no client changes. The hard parts are elsewhere: announce only while a real health check passes, add hysteresis so routes do not flap, use consistent hashing with connection tracking inside each site so TCP survives backend changes, handle ICMP and route changes deliberately, and always be able to prove which site answered. Measure catchments, plan capacity from them and rehearse drains.