Anycast means giving the same IP address to many machines in many places and letting the routing system deliver each packet to one of them. The client does nothing special. It sends a packet to 192.0.2.53, and some router along the way decides which copy of that address is closest by its own rules. DNS root servers, public resolvers, CDN front doors and DDoS scrubbing networks all work this way.
The architecture view is covered in the anycast architecture article, and the choice between anycast and DNS-based steering is covered in anycast versus geo-DNS. This article is the hands-on manual. It follows a packet from a client to a socket, builds a health-gated announcement with ExaBGP, explains why hashing inside a site matters as much as routing between sites, and shows how to prove which site actually answered. If BGP itself is unfamiliar, read BGP basics for developers first.
From first principles: why the same address can live in many places
A router forwards a packet by looking up the destination in its forwarding table and picking the most specific matching prefix. The table does not say whether a prefix lives in one building or fifty. If three sites all announce 192.0.2.0/24 into BGP, every network on the internet receives up to three paths to that prefix and keeps the one its policy prefers. That choice is usually the shortest AS path after local preference, so it is topologically near, not necessarily geographically near or lowest latency.
The set of client networks that end up at a given site is called its catchment. Catchments are decided by other people's routing policy, which is why anycast is easy to start and hard to steer precisely. Three properties fall out of this model and drive every design decision below:
- Selection is per packet, not per connection. Routers keep no session state. If a route changes mid-connection, the next packet can arrive at a different site that has never heard of the connection.
- Failover is removal. A site stops receiving traffic only when it stops announcing the prefix. Anything that keeps the announcement up while the service is broken turns that site into a black hole for its whole catchment.
- Every site must be able to serve every request. Anycast gives you no control over which client reaches which site, so configuration, data and certificates must be identical everywhere.
What you need before the first announcement
| Requirement | Why | Practical note |
|---|---|---|
| A prefix of at least /24 (IPv4) or /48 (IPv6) | Most networks filter longer prefixes, so a /25 will not propagate | You cannot anycast a single /32 on the public internet; you anycast a /24 and use addresses inside it |
| An origin AS | The prefix needs an origin in every announcement | One AS at all sites is simplest; mixed origins need a ROA per origin |
| ROAs in RPKI | Networks doing route origin validation drop announcements whose origin or length does not match | Create a ROA for each origin AS, with maxLength equal to the length you announce |
| Upstreams at each site | Transit and peering decide the size of each catchment | Similar upstream mixes at every site give more even catchments |
| Identical service stacks | Any client can land anywhere | Same software version, data and TLS certificates; config pushed from one source |
A common RPKI mistake is a ROA for a /22 with maxLength 22 while announcing /24s from it. Validating networks drop those /24s, so some catchments silently vanish. Check validity externally after every prefix or origin change.
Health-gated announcements with ExaBGP
The core control loop is simple: a server announces the service route only while it can serve. A common design runs a BGP speaker on each server that announces the service /32 to the site router, and the site router announces the covering /24 to the internet as long as at least one server is announcing. Draining a server is then a matter of withdrawing its /32. Draining a site means all of its servers withdraw, or the router stops exporting the /24.
ExaBGP is convenient here because it runs a helper process and turns lines that process prints into BGP updates. The service address is configured on the loopback interface of every server so the kernel accepts packets for it. BIRD can do the same job by exporting a static or direct route that a health script adds and removes; the logic is the same either way.
# /etc/exabgp/exabgp.conf on each server (ExaBGP 4.x)
process health {
run /usr/local/bin/anycast-health.py;
encoder text;
}
neighbor 10.20.0.1 { # the site's router or L4 tier
router-id 10.20.0.11;
local-address 10.20.0.11;
local-as 64512; # private ASN per server
peer-as 64500;
api {
processes [ health ];
}
}#!/usr/bin/env python3
"""Announce 192.0.2.53/32 only while the local DNS server answers real queries."""
import os, subprocess, sys, time
ROUTE = "192.0.2.53/32"
RISE, FALL, INTERVAL = 3, 3, 2.0 # 3 good checks to announce, 3 bad to withdraw
DRAIN_FILE = "/etc/anycast/drain" # touch this to drain the server on purpose
def healthy():
if os.path.exists(DRAIN_FILE):
return False
try:
out = subprocess.run(
["dig", "@127.0.0.1", "health.example.net", "A", "+short", "+time=1", "+tries=1"],
capture_output=True, text=True, timeout=2).stdout.strip()
return out == "198.51.100.7" # a known answer, not just "the port is open"
except subprocess.TimeoutExpired:
return False
def say(command):
sys.stdout.write(command + "\n")
sys.stdout.flush() # ExaBGP reads our stdout line by line
announced, good_run, bad_run = False, 0, 0
while True:
if healthy():
good_run, bad_run = good_run + 1, 0
else:
good_run, bad_run = 0, bad_run + 1
if not announced and good_run >= RISE:
say(f"announce route {ROUTE} next-hop self")
announced = True
elif announced and bad_run >= FALL:
say(f"withdraw route {ROUTE} next-hop self")
announced = False
time.sleep(INTERVAL)Three details carry most of the value. The check compares a real answer: a port check keeps announcing while the server returns SERVFAIL, creating a zombie site. The rise and fall counters add hysteresis; without them an intermittently failing server flaps its route, and route flap damping elsewhere may suppress your prefix long after the fault. The drain file takes a server out deliberately while in-flight queries finish.
Inside a site: why ECMP alone breaks TCP
Inside a site, the router usually has several equal-cost next hops for the service address, one per server or per L4 balancer, and spreads flows across them with ECMP. It hashes the five-tuple (source and destination address and port, plus protocol) and takes the result modulo the number of next hops. That keeps every packet of a flow on one machine, until the number of next hops changes.
With plain modulo hashing, adding or removing one of N next hops remaps most flows, not one Nth of them, and every remapped TCP connection lands on a server with no state for it and is reset. Single-packet DNS queries do not care; HTTPS and long-lived connections do. Some switches offer resilient hashing, but behaviour varies by platform.
The common fix is an L4 balancing tier, as in Google's Maglev and Meta's Katran. Each balancer picks a backend with a consistent hash, so all balancers agree without coordination, and keeps a connection table so existing flows stay put when backends change. Maglev fills its lookup table simply: each backend gets a pseudo-random permutation of slots, and backends take turns claiming their next free slot.
import hashlib
def _h(value, seed):
digest = hashlib.sha256(f"{seed}:{value}".encode()).digest()
return int.from_bytes(digest[:8], "big")
def maglev_table(backends, m=65537): # m must be prime and much larger than len(backends)
offsets = [_h(b, "offset") % m for b in backends]
skips = [_h(b, "skip") % (m - 1) + 1 for b in backends]
nxt = [0] * len(backends)
table, filled = [None] * m, 0
while True:
for i, b in enumerate(backends): # backends take turns claiming their next free slot
while True:
slot = (offsets[i] + nxt[i] * skips[i]) % m
nxt[i] += 1
if table[slot] is None:
break
table[slot] = b
filled += 1
if filled == m:
return table
def pick(table, src_ip, src_port, dst_ip, dst_port, proto):
return table[_h((src_ip, src_port, dst_ip, dst_port, proto), "flow") % len(table)]
before = maglev_table([f"srv{i}" for i in range(10)])
after = maglev_table([f"srv{i}" for i in range(10) if i != 3])
moved = sum(1 for a, b in zip(before, after) if a != b and a != "srv3")
print(f"slots moved that did not belong to srv3: {moved / len(before):.2%}") # small, not ~90%Run it and the fraction of slots that move, other than the removed server's own, is small. That is the property you need: a change to one backend should disturb only the flows that belonged to it.
The traps that only appear with anycast
- Path MTU discovery. When a router cannot forward a large packet with Don't Fragment set, it sends an ICMP Packet Too Big message to the source address, which is your anycast address. That ICMP packet can reach a different site, or within a site a different server, because its five-tuple differs from the flow it refers to. The real sender never lowers its packet size and the connection stalls on large responses. Mitigate by clamping the TCP MSS at the edge, using conservative MTUs for UDP responses, and making the balancer hash ICMP errors by the embedded original header.
- Mid-connection route changes. Rare per client, constant across millions. Long downloads and WebSockets will occasionally reset, so clients must retry cleanly. QUIC connection IDs help routing inside a site, not across sites, because another site lacks the connection keys.
- Asymmetric return paths. Responses leave via the receiving site's upstreams, so stateful firewalls expecting both directions drop traffic.
- Uneven catchments. One large ISP's policy can send a whole country to the wrong continent.
Finding out which site answered
Every anycast investigation starts with the same question: where did this client actually land? Build the answer into the service from day one. For DNS, the NSID option and the CHAOS-class identity queries return a server-chosen identifier. For HTTP, make every site add a response header with its site and host name.
# Which DNS site answered? (needs the server to support NSID or CHAOS identity)
dig @192.0.2.53 example.net SOA +nsid # look for "NSID:" in the OPT section
dig @192.0.2.53 id.server CH TXT +short # many servers answer with a site/host name
dig @192.0.2.53 hostname.bind CH TXT +short # older BIND-style equivalent
# Which path did packets take, and through which networks?
mtr -z -w -c 20 192.0.2.53 # -z adds AS numbers per hop
# HTTP over anycast: have every site add a response header naming itself
curl -sI https://api.example.net/ | grep -i x-served-byOne test from your laptop shows one catchment. Run the same query from many vantage points, such as RIPE Atlas probes or client telemetry, and aggregate the site identifier by client AS and country. The resulting catchment map drives capacity planning and is the first thing to compare after any routing change.
Worked example: draining a site without dropping queries
Suppose an authoritative DNS service on 192.0.2.53 runs in three sites, each with four servers, using the configuration above. Monitoring from 200 vantage points shows 41 percent of probes landing in FRA, 37 percent in AMS and 22 percent in SIN. FRA needs a kernel upgrade.
- Check headroom first. If most of FRA's catchment shifts to AMS, AMS must absorb roughly 78 percent of traffic. If it cannot, stop and add capacity.
- Touch the drain file on one FRA server. After three failed checks, about 6 seconds, the script withdraws and the site router removes that next hop; connection tracking keeps existing flows in place.
- Repeat for the remaining servers. When the last withdraws, the router stops exporting the /24 from FRA and the internet converges on AMS and SIN, typically within seconds to minutes.
- Re-run the probes. FRA's share should be zero; any probe still reporting FRA points to a stale route.
- Upgrade, remove the drain files, and confirm the catchment returns to its old shape. A different shape means some network changed its routing policy meanwhile.
Failure modes and trade-offs
| Failure | Symptom | Prevention |
|---|---|---|
| Shallow health check | One site serves errors to its whole catchment | Check with a real request and a known answer |
| No hysteresis | Route flaps, then damping suppresses the prefix elsewhere | Rise and fall counters; alert on announcement changes per hour |
| Modulo ECMP on backend change | Connection resets on every deploy | Consistent hashing plus connection tracking in an L4 tier |
| ROA mismatch | Some networks cannot reach one or all sites | One ROA per origin with exact maxLength; external validation |
| Lost ICMP Packet Too Big | Large responses stall | MSS clamping and ICMP-aware balancing |
| Config drift between sites | Bugs that only some users see | Same build and config everywhere, verified per site |
The central trade-off is control versus simplicity. Anycast deletes the global traffic-direction tier and gives you fast failover and natural DDoS dilution, but it gives up the per-client control that DNS steering offers. Most large operators combine the two, as discussed in the routing comparison. For the routing policy levers, such as prepending and communities, see the BGP deep dive.
What to do next
- Confirm you have a /24 or /48 with ROAs matching every origin AS and announced length.
- Configure the service address on loopback and run a BGP speaker with a deep health check, rise and fall counters and a drain file.
- Put a consistent-hashing L4 tier with connection tracking in front of any TCP service.
- Add a site identifier to every response: NSID or CHAOS identity for DNS, a header for HTTP.
- Build a catchment map from many vantage points and plan capacity per site from it.
- Clamp MSS at the edge and test large responses from several networks.
- Rehearse a full-site drain each quarter and compare catchments before and after.