Most of the DNS you use every day is anycast. The root servers, the large top-level domains, managed authoritative providers and the well-known public resolvers all announce the same IP address from many places, and the Internet's routing delivers each query to one of them. Anycast gives DNS low latency, capacity that grows by adding sites, and a way to soak up attacks, without any client changing anything.
It also adds failure modes that unicast DNS does not have: a site that answers wrongly only for the networks that route to it, health checks that can switch a whole service off, and routing that ignores geography. This article explains anycast DNS from the routing up and then focuses on operating it. Anycast routing in general, including drains and DDoS absorption for any service, is covered in the anycast architecture article, and the choice between anycast and DNS-based steering for your own application is in edge routing.
One address, many servers
With unicast, an IP address identifies one machine. With anycast, several sites announce the same prefix into BGP, the Internet's inter-domain routing protocol. Every network that hears more than one announcement picks a single best path by its BGP policy, usually preferring customer routes, then the shortest AS path, then other tie-breakers. The set of networks whose best path leads to a given site is that site's catchment.
Two facts follow. First, the site you reach is the one your network's routing prefers, which correlates with geography but is not the same: a European resolver whose ISP buys transit from a carrier that only peers with your Virginia site will be sent across the Atlantic. Second, nothing in the packet or the protocol says which site answered; you have to ask, which is why the identification options later in this article matter. The routing mechanics are in the BGP article; for a developer-level introduction start with BGP basics.
Why DNS fits anycast so well
Anycast has a known weakness: routing can change between two packets of the same conversation and deliver the second to a different site that knows nothing of the first. Classic DNS barely cares. A typical query is a single UDP datagram and a single UDP reply, with no connection state; if a route changes between two queries, the second just goes to a different server with the same zone. RFC 4786 (BCP 126), on operating anycast services, and RFC 7094, on its architectural considerations, both note that short stateless exchanges are the best case for anycast.
The exceptions are the ones to plan for. Responses larger than the client's advertised EDNS buffer are truncated and retried over TCP, and DNSSEC, zone transfers and encrypted transports such as DNS over TLS all use connections. A route flap mid-connection resets it. In practice routes are stable over the seconds a DNS TCP connection lasts, so this is rare, but it means you should keep responses small where you can, set a conservative EDNS buffer size such as 1232 bytes to avoid fragmentation, and test TCP from many vantage points, not only UDP.
Inside one site
A site is more than a server with an extra address. It has at least two authoritative servers behind a local router or load balancer, each with the anycast address on a loopback interface and a separate unicast management address. A BGP speaker, such as BIRD, FRR or ExaBGP, announces the service prefix to the site's upstream routers. Zone data arrives from a hidden primary that is never itself queried by the public, via NOTIFY and incremental zone transfer, or via a database the servers read.
Use a prefix that will propagate: in practice networks filter IPv4 announcements longer than a /24 and IPv6 longer than a /48, so the service address lives in a prefix of that size announced only by your DNS sites. Never put the anycast address on anything that is not a DNS server behind a health check.
Health checks that drive BGP
The core loop of anycast operations is simple: a site announces the prefix only while it is healthy. If it fails, it withdraws, and BGP moves its catchment to other sites, typically within seconds to a few minutes depending on how far the withdrawal must propagate. The health check must test what users get, an actual DNS answer of the right content, not just whether a process is running.
#!/usr/bin/env python3
# ExaBGP "process" helper: ExaBGP reads announce/withdraw commands from our stdout.
import subprocess, sys, time
PREFIX = "198.51.100.0/24"
RISE, FALL, INTERVAL = 3, 2, 2 # hysteresis: 3 passes to announce, 2 fails to withdraw
def local_ok():
# Ask the local server, on its unicast address, for a canary record and the SOA serial.
try:
out = subprocess.run(
["dig", "+short", "+time=1", "+tries=1", "@127.0.0.1", "canary.example.net", "TXT"],
capture_output=True, text=True, timeout=2).stdout
return "anycast-canary" in out and serial_lag_ok()
except subprocess.TimeoutExpired:
return False
def serial_lag_ok():
return True # TODO: compare local SOA serial with the hidden primary; False if behind > 5 minutes
def global_outage_suspected():
return False # TODO: query peer sites; True if they fail too (shared dependency, not a local fault)
announced, streak = False, 0
while True:
ok = local_ok()
streak = streak + 1 if ok != announced else 0
if not announced and ok and streak >= RISE:
sys.stdout.write(f"announce route {PREFIX} next-hop self\n"); announced = True; streak = 0
elif announced and not ok and streak >= FALL and not global_outage_suspected():
sys.stdout.write(f"withdraw route {PREFIX} next-hop self\n"); announced = False; streak = 0
sys.stdout.flush()
time.sleep(INTERVAL)Three details in this helper matter. Hysteresis, several passes before announcing and several failures before withdrawing, stops a flapping check from flapping routes, which other networks may penalize with route flap damping. The check includes zone freshness, because a site that answers from an old zone is worse than one that answers nothing. And there is a guard against withdrawing during a global problem.
That guard comes from the best-known anycast DNS failure. On 4 October 2021, a configuration change disconnected Facebook's backbone. Their DNS sites were designed to withdraw their BGP announcements when they could not reach the data centers, on the reasoning that an unhealthy site should step aside. Every site saw the same condition at once, so every site withdrew, and Facebook's authoritative DNS vanished from the Internet, which in turn slowed recovery. The lesson: a health check must distinguish 'I am broken' from 'something we all depend on is broken'. Check peer sites, keep a minimum number of announcing sites, or keep a last-resort announcement with a heavily prepended path.
Which instance answered? Debugging catchments
Because all sites share an address, a user report of a wrong answer is useless until you know which site gave it. Configure every server to identify itself, and teach your support team the commands.
# Which instance answered me? NSID (RFC 5001), if the server is configured to return it
dig @198.51.100.53 example.net SOA +nsid
# The older CHAOS-class identity queries (RFC 4892): id.server, or hostname.bind on BIND
dig @198.51.100.53 id.server TXT CH +short
dig @198.51.100.53 hostname.bind TXT CH +short
# Is the path stable, and where does it end?
mtr -u -P 53 -c 20 198.51.100.53
# Did a large answer fall back to TCP, and does TCP work from here?
dig @198.51.100.53 example.net DNSKEY +dnssec +bufsize=1232
dig @198.51.100.53 example.net DNSKEY +dnssec +tcpNSID is an EDNS option defined in RFC 5001: the client sends it empty and the server returns an operator-chosen identifier. The CHAOS-class id.server and hostname.bind queries, described in RFC 4892, predate it and are still widely supported. Use opaque identifiers such as fra-2 rather than internal hostnames. To see catchments at scale, measurement platforms such as RIPE Atlas can run NSID queries from many networks at once, which shows whether one site is attracting traffic from far away. Fix imbalances with BGP tools: prepend the AS path at an over-attracting site, use upstream communities to limit where a site's announcement propagates, or add peering where users are poorly served.
Keeping every site consistent
The subtle anycast bug is a site that is up but serving an old zone. Its catchment gets stale answers, everyone else gets fresh ones, and tests from your office, which route to a different site, pass. Monitor the SOA serial of every site through its unicast management address, alert on any site that lags, and fold the lag into the health check as above.
import dns.edns, dns.message, dns.query, dns.rdatatype
ZONE = "example.net."
SITES = {"FRA": "192.0.2.11", "IAD": "192.0.2.21", "SIN": "192.0.2.31"} # unicast mgmt IPs
ANYCAST = "198.51.100.53"
NSID = 3 # EDNS option code
def soa_serial(ip, want_nsid=False):
opts = [dns.edns.GenericOption(NSID, b"")] if want_nsid else []
q = dns.message.make_query(ZONE, dns.rdatatype.SOA, use_edns=0, options=opts)
r = dns.query.udp(q, ip, timeout=2)
serial = r.answer[0][0].serial
nsid = b""
for o in r.options:
if o.otype == NSID:
nsid = getattr(o, "nsid", None) or getattr(o, "data", b"")
return serial, nsid.decode(errors="replace")
serials = {name: soa_serial(ip)[0] for name, ip in SITES.items()}
newest = max(serials.values())
for name, s in serials.items():
print(f"{name}: {s}{' <-- STALE' if s != newest else ''}")
print("anycast from here:", soa_serial(ANYCAST, want_nsid=True))Run this from several locations every minute. The last line also shows which site your vantage point reaches, so you can correlate a user's complaint with the site that served it. For record changes that must be visible everywhere at once, such as a migration, lower the TTL in advance and wait until every site shows the new serial before relying on the change; DNS caching itself is explained in the DNS architecture article.
Authoritative and recursive anycast differ
Anycast authoritative servers all hold the same data, so it does not matter which one answers. Anycast recursive resolvers are different: each site has its own cache, so a cache miss at one site is not helped by a hit at another, and popular names are warm everywhere while rare ones are cold. They also hide the client's location from authoritative servers that tailor answers by geography, such as CDNs; the EDNS Client Subnet option in RFC 7871 passes a truncated client prefix to restore that, at a cost in cache efficiency and privacy, and some public resolvers do not send it.
Resilience beyond a single anycast cloud
Anycast protects against site failures, not against the failure of the system that runs every site: a bad configuration push, a software bug triggered by one query, a mistaken withdrawal, or an attack on one provider. The October 2016 attack on the managed DNS provider Dyn took many large sites offline for exactly that reason: their NS sets pointed only at one provider.
- List NS names on at least two independent anycast prefixes, ideally announced by different autonomous systems or providers.
- Keep both providers fed from one source of truth, with serial monitoring on each.
- Stage configuration and software changes site by site, never to all sites at once.
- Use response rate limiting so your servers are not useful as reflection amplifiers, and accept that under a large attack one catchment may degrade; withdrawing that site moves the attack, it does not stop it.
Trade-offs
Anycast buys low latency for most users, capacity you add one site at a time, and attack traffic diluted across many catchments, all behind a single address that never changes. The price is operational. Running it yourself needs your own autonomous system number, a dedicated /24 or /48, BGP skills on call and the fixed cost of every site. You also give up control over which site serves whom, since other networks' policies decide, and with it some ease of debugging, because every report starts with finding the instance. Connection-based traffic, DNS over TCP and TLS, is exposed to route changes that UDP shrugs off. For most organizations the right answer is to buy anycast authoritative DNS from two providers rather than build it, and to use DNS-based steering when the goal is choosing among your own application regions rather than serving DNS itself.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Some users get old answers | One site serves a stale zone | Per-site serial monitoring; serial lag in health check |
| Whole service disappears | Every site withdraws on a shared failure | Peer-aware health check, minimum announcing sites |
| Users routed to a distant site | BGP policy, not distance, picks the path | Prepending, communities, more peering; measure with NSID |
| DNSSEC or large answers fail for some | Fragmented UDP dropped, or TCP blocked or reset | EDNS buffer 1232, test TCP from many networks |
| Routes flap, then are suppressed | Health check without hysteresis | Rise and fall thresholds |
| Outage of the whole provider | All NS on one anycast system | Two providers, two prefixes |
What to do next
- Configure NSID or id.server on every authoritative instance with an opaque site identifier, and add the dig commands to your runbook.
- Make each site's BGP announcement depend on a real DNS answer and zone freshness, with rise and fall hysteresis.
- Add a guard so sites do not all withdraw at once on a shared dependency failure, and test it by breaking that dependency in staging.
- Monitor SOA serials per site through unicast addresses from several vantage points and alert on any lag.
- Measure catchments with NSID queries from many networks and correct the worst routing with prepending or peering.
- Put your NS set on two independent anycast providers or prefixes, and read the anycast architecture article for drains and DDoS absorption.