DNS failover is the oldest way to move traffic between regions, and for LLM serving it is still the most common last line of defence. If the primary region's endpoint goes unhealthy, the authoritative DNS server answers with the standby region's address instead. No proxy sits in the data path, and it works for any client that resolves a hostname.
That simplicity hides a timeline that is easy to get wrong. Detection takes tens of seconds, resolvers keep the old answer for the TTL, runtimes cache it again, and clients holding pooled or streaming connections never ask DNS at all. For a GPU-backed service there are two more twists: a health check cannot afford to run a real generation inside the checker's time budget, and a standby region that has not loaded its weights is not a standby. This article walks through each piece and a worked two-region design.
When a second region is worth paying for, and what happens to a stream cut mid-response, is covered in multi-region LLM serving. Here the subject is the mechanism.
What a failover record does
A failover record pair is two answers for one name. In Route 53 terms they are two records with the same name and type, one marked PRIMARY and one SECONDARY, each with its own SetIdentifier. The primary carries a health check. While that check is healthy, queries get the primary's address; when it is unhealthy, they get the secondary's. Cloud DNS expresses the same idea as a failover routing policy with an active set and a backup set, and adds a trickle ratio, a fraction from 0 to 1 of traffic sent to the backup even while the primary is healthy, a cheap way to keep the standby path exercised.
Three things follow. First, DNS fails over a whole region, not a request. Per-request decisions belong to a router or gateway, described in LLM routing strategies and the AI gateway overview. Second, DNS only influences new lookups. Third, the decision is only as good as the health signal.
The failover timeline
The time from the primary failing to the last client moving is the sum of four stages.
- Detection. Route 53 health checkers probe every 10 or 30 seconds, from several locations that do not coordinate. Each checker marks the endpoint unhealthy after the number of consecutive failures you set as the failure threshold. Route 53 treats the endpoint as healthy while more than 18% of checkers report it healthy. With a 10 second interval and a threshold of 3, plan on roughly 30 to 40 seconds.
- Authoritative change. Once the status flips, the authoritative servers answer with the secondary. This is quick compared with the other stages.
- Resolver caching. Recursive resolvers keep the old answer until its TTL expires. A 60 second TTL means up to 60 more seconds for some clients. Some resolvers clamp very low TTLs upwards, so do not assume a 5 second TTL is honoured everywhere.
- Client caching and connection reuse. The JDK caches successful lookups for 30 seconds by default when no security manager is installed, controlled by
networkaddress.cache.ttl. More important, an HTTP client with a keep-alive pool does not resolve again until it opens a new connection, and a server-sent-events stream holds its connection for the whole response.
Typically that is about 35 seconds of detection, up to 60 of resolver TTL and up to 30 of JVM cache gives roughly two minutes before the last well-behaved client resolves again. Clients holding healthy-looking connections to a degraded region can stay longer, because nothing forces a reconnect. A dead region at least breaks connections; the dangerous case is partial failure, where TCP works and the GPUs do not.
Health checks that mean something for an LLM
A health check for an LLM endpoint has to answer the question users care about: can this region produce tokens at an acceptable speed right now? A TCP or liveness check does not: inference servers fail with the port open, from a wedged CUDA engine, a full KV cache queueing requests for minutes, or the wrong weights loaded.
The checker's own limits rule out the obvious fix. An HTTP or HTTPS Route 53 check must connect within 4 seconds and receive a 2xx or 3xx status within 2 seconds after connecting. A real generation can exceed that under load, and many checkers would add steady GPU load. So separate the probe from the check: a background task runs a small synthetic generation every few seconds, and the health endpoint returns the cached verdict instantly.
# healthz.py: the checker hits /healthz and never waits on the GPU.
import json, threading, time, urllib.request
from http.server import BaseHTTPRequestHandler, HTTPServer
ENGINE = "http://llm.internal:8000" # regional internal LB in front of the replicas
MODEL = "chat-model" # the model name this region serves
PROBE_EVERY_S = 5
MAX_PROBE_LATENCY_S = 4.0 # end-to-end budget for a 16-token answer
FAILS_TO_DRAIN = 3 # consecutive bad probes before we go unhealthy
GOOD_TO_RECOVER = 12 # consecutive good probes before healthy again
MAX_AGE_S = 20 # a stale verdict counts as unhealthy
state = {"ok": False, "fails": 0, "good": 0, "at": 0.0, "latency": None, "drain": False}
def probe_once():
body = json.dumps({"model": MODEL, "prompt": "Reply with the word ready.",
"max_tokens": 16, "temperature": 0}).encode()
req = urllib.request.Request(ENGINE + "/v1/completions", body,
{"Content-Type": "application/json"})
t0 = time.monotonic()
with urllib.request.urlopen(req, timeout=MAX_PROBE_LATENCY_S) as r:
out = json.load(r)
dt = time.monotonic() - t0
text = out["choices"][0]["text"].strip().lower()
return dt <= MAX_PROBE_LATENCY_S and "ready" in text, dt
def probe_loop():
while True:
try:
good, dt = probe_once()
except Exception:
good, dt = False, None
state["fails"] = 0 if good else state["fails"] + 1
state["good"] = state["good"] + 1 if good else 0
if state["fails"] >= FAILS_TO_DRAIN:
state["ok"] = False
elif state["good"] >= GOOD_TO_RECOVER:
state["ok"] = True
state["at"], state["latency"] = time.time(), dt
time.sleep(PROBE_EVERY_S)
class Health(BaseHTTPRequestHandler):
def do_GET(self):
fresh = time.time() - state["at"] < MAX_AGE_S
healthy = state["ok"] and fresh and not state["drain"]
payload = json.dumps({"status": "pass" if healthy else "fail",
"latency_s": state["latency"]}).encode()
self.send_response(200 if healthy else 503)
self.send_header("Content-Type", "application/json")
self.end_headers()
self.wfile.write(payload)
threading.Thread(target=probe_loop, daemon=True).start()
HTTPServer(("0.0.0.0", 8081), Health).serve_forever()Three details matter. The verdict expires, so a hung probe thread reads as unhealthy rather than frozen healthy. The drain flag gives operators a manual switch for planned failovers. And the probe asserts on content, which catches a region serving the wrong model. It runs as a small regional service probing through the regional load balancer, so it judges the region rather than one replica. Route 53 HTTPS checks do not validate certificates, so add a separate certificate-expiry alert.
The records
With the health endpoint in place, the records are short. The Route 53 change below creates the pair with a 60 second TTL.
{
"Comment": "chat API: primary us-east, standby us-west",
"Changes": [
{"Action": "UPSERT", "ResourceRecordSet": {
"Name": "api.example.com", "Type": "A", "TTL": 60,
"SetIdentifier": "primary-us-east", "Failover": "PRIMARY",
"HealthCheckId": "HC-PRIMARY-ID",
"ResourceRecords": [{"Value": "203.0.113.10"}]}},
{"Action": "UPSERT", "ResourceRecordSet": {
"Name": "api.example.com", "Type": "A", "TTL": 60,
"SetIdentifier": "standby-us-west", "Failover": "SECONDARY",
"HealthCheckId": "HC-STANDBY-ID",
"ResourceRecords": [{"Value": "198.51.100.20"}]}}
]
}Give the secondary a health check too, and check what your provider returns when every record is unhealthy; you want the standby's health on a dashboard long before you need it. On Google Cloud the equivalent is a Cloud DNS failover policy; public zones can health-check external endpoints at intervals between 30 and 300 seconds, and private zones can health-check internal load balancers. Record management and TTL planning on Cloud DNS are covered in Cloud DNS in practice.
Going much below 60 seconds buys little because detection and clients dominate. Lower a TTL days before a planned move, because the old TTL is still in caches.
Clients that resolve again
Most of the long tail lives in clients you control. Make them re-resolve on purpose.
- Bound connection lifetime. Recycle pooled connections after a fixed age, for example 60 seconds, so every connection resolves again within a minute even if it never errors.
- Reconnect on the right errors. Treat 502, 503, connection resets and first-token timeouts as signals to drop the pool and resolve again before retrying.
- Set the runtime cache deliberately. On the JVM, set
networkaddress.cache.ttlin the security properties to match your record TTL instead of relying on the default. - Respect streams. A stream in progress is not moved by DNS. Decide whether a broken stream is retried from the start or resumed, and make that idempotent; see the mid-stream discussion in the multi-region article.
# A client wrapper that bounds connection age and resolves again on failure.
import time, httpx
class FailoverAwareClient:
def __init__(self, base_url, max_conn_age_s=60, timeout_s=30):
self.base_url, self.max_age, self.timeout = base_url, max_conn_age_s, timeout_s
self._new()
def _new(self):
# A fresh client means a fresh pool, so the next request resolves the name again.
self.client = httpx.Client(base_url=self.base_url, timeout=self.timeout)
self.born = time.monotonic()
def post(self, path, payload, attempts=3):
for i in range(attempts):
if time.monotonic() - self.born > self.max_age:
self.client.close(); self._new()
try:
r = self.client.post(path, json=payload)
if r.status_code not in (502, 503, 504):
return r
except (httpx.ConnectError, httpx.ReadTimeout, httpx.RemoteProtocolError):
pass
self.client.close(); self._new() # drop pool, resolve again
time.sleep(min(2 ** i, 8))
raise RuntimeError("all attempts failed")
A standby that can take the load
For a GPU service the standby region is where DNS failover most often fails in practice. Three properties make a standby real.
- Weights resident. Loading a 70B-parameter model from object storage and warming the engine takes minutes, much longer than the DNS timeline. A warm standby keeps a floor of replicas loaded and serving.
- Capacity to absorb. If region B normally runs at 30% of region A's capacity, failing over sends all traffic into a region that will queue and time out, and its own health check then fails. Decide in advance what the standby must carry: full load, a degraded tier with a smaller model, or priority customers only, enforced by the gateway.
- Parity. Same model version, same tokenizer, same system prompts and the same safety configuration. A probe that asserts on content catches gross mismatches; a release process that deploys both regions together prevents the rest.
On-demand GPUs may not exist at the moment of an incident, so reserve the floor.
Worked example: a two-region chat API
A chat API serves 40 requests per second at peak from region A; each replica sustains about 4, so A runs 10 plus 2 for headroom, with region B as standby. The team decides B must carry full peak for up to an hour, but can start at half and scale.
B runs 5 warm replicas and gets a 5% canary trickle so its path stays exercised. Records use a 60 second TTL. Health checks run at a 10 second interval with a threshold of 3 against the cached probe, which itself drains after three bad probes 5 seconds apart. The gateway recycles connections every 60 seconds and on any 503.
Timeline for a wedged engine fleet in A: probes fail from t=0 and the endpoint returns 503 at about t=15. Checkers mark it unhealthy around t=45. Resolvers serve the old answer until about t=105 at worst, and the gateway's pool recycles within another 60 seconds, so traffic has moved by roughly t=165. During that window B scales from 5 toward 10 replicas; if weights load in four minutes, B is under-provisioned for a few minutes and the gateway sheds free-tier traffic. A quarterly drill measures each stage.
Failback without flapping
Failback is automatic with health-checked records: when A's check passes again, new lookups return A. If A recovers cold, with empty caches and replicas still loading, it attracts the full load and fails again, and traffic flaps between regions every few minutes. Prevent it with hysteresis in the probe, like the 12 consecutive good probes (a minute) the code above requires,, or keep the primary drained by hand until an operator confirms capacity. A calculated health check can also require a capacity signal.
Failure modes
- Shallow checks. A TCP or process check stays green while the engine is wedged. Check generation, through the real load balancer. Route 53 also treats a new check as healthy until it has data.
- Check that pins healthy. A probe thread that dies leaves a stale verdict; expire it.
- Both regions red. A shared dependency (identity, a model registry, a config service) fails both checks and DNS has nothing to fail over to.
- Thundering standby. Full traffic lands on a small warm pool. Size the floor and shed load.
- Long-lived connections. Streams and pools keep talking to a degraded region. Bound connection age.
- Flapping failback. Add hysteresis and a manual hold.
Trade-offs
DNS failover is cheap, provider-neutral and has no component in the request path. Its costs are speed and precision: failover takes minutes rather than seconds, it moves whole regions, and it depends on client behaviour you may not control. An anycast or global load balancer moves traffic faster and can fail over per backend, at the price of a managed proxy in the path and, often, a single provider; see Google Cloud load balancing. A client-side router with several endpoints moves per request but needs every client to run it. Many teams combine them: a gateway retries across regions per request, and DNS failover protects the gateway itself.
What to do next
- Write down your four-stage failover timeline with real numbers for detection, TTL, runtime cache and connection reuse.
- Replace liveness-only checks with a cached synthetic-generation probe that asserts on content and expires stale verdicts.
- Put health checks on both regions and alert on the standby's status.
- Decide what the standby must carry, keep that floor of replicas warm, and send it a trickle of real or synthetic traffic.
- Bound connection age in your gateway and SDKs, and drop pools on 502, 503 and connect errors.
- Add hysteresis or a manual hold to failback.
- Run a failover drill, measure each stage, and repeat it every quarter.