Anycast means announcing the same IP prefix from many locations and letting the internet's routing deliver each client to one of them. CDNs and public DNS resolvers run on it, and it is an attractive front door for an LLM API. You get one address worldwide, no DNS TTL to wait on, a nearby place to terminate TLS, and attack traffic spread across every site. It also fits LLM traffic badly in two ways. A token stream lasts tens of seconds to minutes, and anycast gives no promise that a flow's packets keep reaching the same site. The GPUs that actually produce tokens live in a few regions, not at every point of presence.

This article explains how anycast routing really works, why a naive design breaks long streams, and the standard fix: anycast in front, unicast behind. It then covers keeping streams alive across route changes, health-driven announcements with code, gentler traffic engineering than withdrawal, measuring where clients land, and a worked failover example.

How anycast actually routes

Each site runs routers that announce the prefix, say 203.0.113.0/24, over BGP to its transit providers and peers. Every network on the internet then picks a best path to the prefix using its own policy: local preference first (customers before peers before transit, usually), then AS-path length, then several tie-breakers. The set of clients that ends up at a given site is that site's catchment. Note what BGP does not consider: latency, packet loss, and how loaded your site is. A Madrid client can land in Frankfurt because its ISP peers with your Frankfurt transit and not your Madrid one. RFC 4786 (BCP 126) and RFC 7094 describe the operational and architectural consequences.

Inside a site, routers spread flows across servers with equal-cost multipath (ECMP), hashing the 5-tuple (addresses, ports, protocol). Two events move a flow. A BGP change somewhere upstream can shift the client's catchment to another site. A change in the ECMP set, such as a server added or removed, can rehash flows onto different servers. TCP state lives on one server, so in either case the new destination has never heard of the connection and answers with a reset.

Why LLM traffic is awkward

For a web page fetched in 200 ms, route changes almost never land mid-flow. LLM streams are different. A 1,500-token answer at 50 tokens per second is a 30-second stream. Agent sessions and long reasoning traces hold connections for minutes. A rough estimate: if a client's path changes on average r times per hour and a stream lasts L hours, the chance a given stream sees a change is about r times L. At an illustrative r of 0.05 per hour, a 30-second stream breaks once in about 2,400 streams. A 10-minute agent session breaks once in 120. At millions of streams a day, both are real numbers.

Breaks also cost more than on the web. A retried request repeats prefill, so the GPU work already done is wasted. The user sees a half-finished answer restart. Any tool calls the model already made may run twice. Then there is placement. GPUs sit in two to five regions chosen for power and capacity. Announcing the anycast prefix directly from GPU regions would make BGP, which knows nothing about load, the arbiter of where expensive inference runs.

Anycast in front, unicast behind

The design that works separates the two concerns. Anycast runs only at the PoPs, many small sites close to users. They terminate TCP or QUIC and TLS, authenticate, rate-limit, and then proxy each request over unicast to a GPU region picked by a load-aware policy. The GPU regions keep ordinary unicast addresses that are never anycast. The client-to-PoP leg is short, so its exposure to route changes is small. The PoP-to-region leg runs over your backbone or stable unicast paths that you control.

Anycast front door, unicast GPU regionsClient ABerlinClient BMadridClient CMumbaiPoP FRAannounces 203.0.113.0/24PoP MADannounces 203.0.113.0/24PoP BOMannounces 203.0.113.0/24BGPBGPBGPL4 + L7 at PoPTLS ends hereGPU eu-westunicastGPU us-eastunicastGPU ap-southunicastbackbonespillvia BOM L7Health controller per PoP: announce, prepend or withdraw
Clients reach the nearest PoP by BGP. The PoP terminates TLS and chooses a GPU region by load and policy over unicast. A per-PoP controller manages announcements.

Region selection at the PoP is a separate problem with its own trade-offs: residency filters, TTFT scoring and cache affinity. It is covered in geographic routing for LLM serving. The PoP's L7 layer is also the natural home for the policy functions of an AI gateway, such as keys, budgets and token-based rate limits. The PoP does not need GPUs. If you want small models at the edge too, that is a different design, discussed in edge PoP serving.

Keeping a stream alive

Use three layers of defence, each covering a different event.

Within a PoP: consistent hashing with connection tracking. Put an L4 load balancer between routers and proxies. It should use a consistent-hash table, as in Google's Maglev design, plus a connection table, so that adding or removing a proxy remaps only a small share of new flows and no existing ones. Drain proxies before removing them: stop new flows and wait for the longest streams to finish.

Across PoPs: QUIC connection IDs. QUIC (RFC 9000) identifies connections by connection ID rather than the 5-tuple. A server can encode routing information into the IDs it issues, so a load balancer can find the owning server even after a client address change. The IETF QUIC-LB draft that proposes such encodings has not been published as an RFC (check the datatracker for its current state), so treat any encoding as implementation-specific. Even with routable IDs, a packet arriving at the wrong PoP has to be forwarded to the right one, which needs inter-PoP forwarding that few teams build. Plain TCP connections cannot survive a cross-PoP move at all.

At the application layer: resumable streams. This is the layer that always works. Generation runs in the GPU region and does not depend on the client connection. The region buffers tokens under a stream id with sequence numbers. If the client reconnects, possibly through a different PoP, it asks to resume from the last sequence it received. Server-sent events already define this through event ids and the Last-Event-ID request header.

# GPU-region side: decouple generation from the client connection.
async def generate(stream_id, request):
    buf = buffers.create(stream_id, ttl_s=300)          # bounded, short-lived
    async for seq, token in enumerate_tokens(model.stream(request)):
        buf.append(seq, token)
    buf.close()

# Edge-facing SSE handler: works for first connect and for resume.
async def sse(req):
    stream_id = req.query["stream_id"]
    start = int(req.headers.get("Last-Event-ID", "-1")) + 1
    buf = await buffers.locate(stream_id)               # region lookup by id prefix
    if buf is None:
        return http_error(410, "stream expired, resubmit")   # client restarts explicitly
    async for seq, token in buf.read_from(start):
        yield f"id: {seq}\ndata: {token}\n\n"

Encode the owning region in the stream id so any PoP can route a resume without a global lookup. Keep buffers short-lived and bounded, because they hold user content. Generation that keeps running after a disconnect costs GPU time, so cancel streams that are not resumed within a short grace period.

Health-driven announcements

Each PoP announces the prefix only while it can serve. A small controller probes the local stack end to end and tells the BGP speaker what to do. With ExaBGP, a helper process writes commands such as announce route 203.0.113.0/24 next-hop self to standard output, and ExaBGP turns them into BGP updates. BIRD and FRR can be driven the same way through their own interfaces.

import sys, time

PREFIX = "203.0.113.0/24"
UP_AFTER, DOWN_AFTER = 5, 3        # hysteresis: consecutive results needed to flip
state, streak = "withdrawn", 0

def emit(cmd):
    sys.stdout.write(cmd + "\n"); sys.stdout.flush()

while True:
    ok = probe_local_stack()   # TLS handshake + authenticated tiny completion via this PoP
    want = "announced" if ok else "withdrawn"
    streak = streak + 1 if want != state else 0
    if want != state and streak >= (UP_AFTER if ok else DOWN_AFTER):
        emit(("announce" if ok else "withdraw") + f" route {PREFIX} next-hop self")
        state, streak = want, 0
    time.sleep(2)

Two rules matter more than the code. First, the probe must test the PoP, not the GPU fleet. If GPU regions are saturated, every PoP would fail together and withdraw the prefix worldwide. The PoP should keep accepting traffic and shed or queue requests at the application layer. Withdrawal is for 'this site cannot terminate connections'. Second, flapping is punished. Some networks apply route flap damping (RFC 2439) and suppress a prefix that changes too often, which can make a site unreachable from them for far longer than the outage. Hysteresis, and a cap on transitions per hour, are required.

Gentler traffic engineering

Withdrawal is a blunt tool. It moves the whole catchment at once and resets every TCP connection that does not resume. You have gentler options:

LeverEffectUse it for
AS-path prependingMakes the path look longer, so some networks prefer other sitesPre-drain before maintenance; trimming an overloaded catchment
BGP communitiesAsk a transit or peer to lower preference or stop exportingFine-grained shifts per provider, where the provider supports it
Multiple anycast prefixesClients get addresses in prefix A and B, announced from different site setsMoving a share of traffic, staged migrations
More-specific prefixA longer prefix wins over a shorter one everywherePinning traffic to a site. Mind the /24 (IPv4) and /48 (IPv6) filtering limits

The pattern for planned work is: prepend, wait for in-flight streams to drain (at least your p99 stream length), then withdraw. Prepending moves new connections while existing flows mostly keep their paths.

Measuring catchments

Catchments cannot be predicted from a map, so measure them. Return a header naming the serving PoP on every response, and log it with the client's ASN and coarse location. Compare each client network's actual PoP and RTT with the best PoP for it. Large gaps point to peering problems you can fix by adding a peer or transit at the right site. Active measurement platforms such as RIPE Atlas can map catchments from thousands of vantage points before a launch. Track three numbers per PoP: the share of traffic that is 'misrouted' (RTT more than 30 ms worse than the best site), the stream reset rate, and the resume success rate.

Worked example: patching a busy PoP

All numbers here are illustrative. An API runs 18 PoPs and three GPU regions. The Frankfurt PoP holds 40,000 concurrent streams at the evening peak, and p99 stream duration is 90 seconds. The kernel on its proxies needs patching.

  1. Day before: catchment data shows Frankfurt's clients would mostly fall back to Amsterdam and Paris. Both have proxy headroom for 25,000 extra streams, so the plan fits.
  2. T+0: the controller prepends Frankfurt's announcement three times. New connections drift to Amsterdam and Paris over a few minutes as routes converge. Frankfurt's concurrency falls as streams finish.
  3. T+3 min: concurrency is under 2,000, mostly long agent sessions. The prefix is withdrawn.
  4. About 1,800 connections reset. Resume handles about 95% of them because their buffers are still in eu-west. Roughly 90 requests restart from scratch.
  5. GPU load is unchanged throughout. Amsterdam and Paris proxy to the same eu-west region, so moving the anycast catchment did not move inference.

Without prepending and resume, all 40,000 streams would have reset and been retried at once. That would have repeated their prefill and spiked eu-west, which is how a routine patch becomes an incident.

Failure modes

  • Global withdrawal from a shared dependency. A probe that checks GPU capacity makes every PoP withdraw together.
  • Flap damping. A marginal PoP that flips every minute is suppressed by some networks for an hour.
  • Catchment surprise. A new peering session drags a distant country's traffic into a small PoP.
  • ECMP rehash during deploys. Removing proxies without draining resets long streams.
  • Duplicate side effects. A client retries a broken stream that had already executed tool calls. Use idempotency keys.
  • Orphaned generation. Disconnected streams keep burning GPU time with no reader.

Trade-offs

OptionStrengthsWeaknessesFits
Anycast front door + unicast regionsOne IP, fast failover, DDoS spread, TLS near usersNeeds your own ASN/prefix or a provider, plus BGP operationsGlobal APIs with many PoPs
GeoDNS to regionsSimple, load-aware answersResolver location, TTL lagFew regions, smaller teams. See DNS failover
Anycast directly on GPU regionsFewest hopsBGP decides GPU placement, streams break on shiftsRarely
Cloud global load balancerAnycast without running BGPProvider lock-in, less controlTeams without a network team

What to do next

  1. Measure stream duration distribution (p50, p99) and reset rate today.
  2. Make generation independent of the client connection, and add resumable streams with sequence ids.
  3. Keep GPU regions on unicast. Anycast only at PoPs that terminate TLS.
  4. Write a PoP health probe that excludes GPU capacity, with hysteresis and a transition cap.
  5. Script the maintenance pattern: prepend, drain to p99, withdraw, restore.
  6. Add a serving-PoP header and build per-ASN catchment and misrouting dashboards.
  7. Rehearse a PoP failure and confirm neighbouring PoPs and the GPU region absorb it.
Key takeaway: Anycast is a good front door for an LLM API and a poor way to place GPU work. Announce at PoPs that terminate connections, proxy over unicast to load-aware GPU regions, make streams resumable so route changes cost a reconnect instead of a repeated prefill, gate announcements on PoP health with hysteresis, and drain with prepending before you ever withdraw.