For an ordinary web service, geographic routing means sending each user to the nearest healthy region, and that is usually right because network round-trip time dominates the response. For LLM inference it is often wrong. A round trip across a continent costs 50 to 150 milliseconds. The queue in front of a saturated GPU pool, or a long prefill, can cost seconds. GPUs are scarce and unevenly loaded across regions. Some traffic is legally pinned to a jurisdiction, and a multi-turn conversation leaves a prefix cache in the region that served its last turn. "Nearest" is one input among several.
This article builds the region-selection function: the code that, for each request, filters regions on hard constraints and then scores the rest by estimated time to first token, cache affinity and cost. It does not repeat the mechanics of anycast and DNS steering, covered in Edge Routing: anycast versus GeoDNS, or the case for adding regions at all, covered in Multi-Region LLM Serving.
Where the time to first token goes
Start from the user's experience. For a streamed chat response, time to first token (TTFT) is roughly the connection set-up, plus one round trip for the request, plus queue wait, plus prefill. On a new connection, TCP and TLS 1.3 each add about one round trip before the request can be sent, so a cold request pays about three round trips. A warm, reused connection pays about one. After the first token, tokens stream over the open connection, so distance adds little to the time between tokens. That time is set by the decode step on the GPU.
def ttft_ms(rtt_ms, queue_ms, prompt_tokens, prefill_tps, warm_conn=True):
rtts = 1 if warm_conn else 3 # request only, or TCP + TLS 1.3 + request
prefill_ms = 1000.0 * prompt_tokens / prefill_tps
return rtts * rtt_ms + queue_ms + prefill_msPlug in illustrative numbers for a user in Mumbai sending a 4,000-token prompt to a replica that prefills 8,000 tokens per second (500 ms of prefill). At an evening peak the local region has 900 ms of queue. A region 60 ms away has 80 ms. The nearby region gives about 1,415 ms on a warm connection and the distant one about 640 ms. Distance lost by 45 ms and queueing won by 820 ms. The opposite is also true: off-peak, with both queues near zero, the local region wins by the RTT difference. A static nearest-region rule is wrong half the day.
Three layers, three decisions
Production designs split routing into three layers, each making one decision. The edge layer (anycast or GeoDNS) gets the user's packets to a nearby point of presence and terminates TLS close to them, which also removes the cold-connection penalty from the long-haul path. The global router, running at the edge or in a regional front door, chooses the serving region. The regional router chooses a replica within that region, ideally the one that already holds the conversation's KV blocks, as described in LLM Routing Strategies.
Keeping these layers separate matters operationally. DNS records are cached for their TTL and cannot react to load within seconds. An L7 router can make that decision per request. Use DNS and anycast for reachability and coarse failover, as in DNS Failover for LLM Serving, and use the global router for load-aware placement.
Hard filters before any scoring
Scoring only happens after filtering. Some constraints are not trade-offs and must never be outweighed by a good latency score.
- Data residency. A tenant whose data must stay in the EU can only be served by EU regions, including during an outage. The fallback list is part of the residency policy, not the routing code. See Geo-Partitioning and Data Residency.
- Model availability. Not every model, version or context length is deployed in every region. A request for a 128k-context variant can only go where that variant is running.
- Tenant and compliance tier. Dedicated-capacity customers, or traffic requiring a particular certification, may be restricted to specific pools.
- Health. A region failing its GPU-aware health checks is removed. A simple "HTTP 200" probe is not a GPU-aware check.
If filtering leaves no region, fail the request with a clear, retryable error. Never fall back to an unfiltered list. A residency breach caused by a routing fallback is still a breach.
Load signals and the herd problem
To estimate queue wait you need per-region load signals. Count of requests in flight is a weak signal, because requests differ by orders of magnitude in cost. Better signals come from the serving engines themselves. Note the two throughputs in the code below: queue wait divides by the whole pool's prefill rate, while a request's own prefill runs at one replica's rate.
| Signal | What it tells you | Caveat |
|---|---|---|
| Waiting prefill tokens | Work queued before the next request can start | Divide by measured prefill throughput to get milliseconds |
| KV cache utilisation | How close the pool is to preempting or rejecting | High usage is normal under good batching; watch the trend and preemptions |
| Observed TTFT p50 and p95 | Ground truth for recent requests | Lags the present by the measurement window |
| Free replicas or autoscaler state | Capacity arriving soon | Cold starts can take minutes when weights must load |
Signals reach the global router every one to five seconds. In that window, every router sees the same quiet region and sends it traffic, the region fills, and the next report sends everyone away again. This herd effect is the main failure of load-aware geo-routing. Two standard defences work. First, do not always pick the best region. Sample in proportion to a weight derived from the score, or choose the better of two randomly sampled candidates. Second, smooth the signals and add each router's own recent sends to the reported queue before the next report arrives.
A region-scoring function
The function below combines the pieces. It takes a request, the region table and the latest signals, and returns a region. It filters on hard constraints, estimates TTFT, applies an affinity credit and a cost term, and makes a weighted random choice.
import math, random
def choose_region(req, regions, sig, sticky_region=None, temp_ms=150.0):
cands = [r for r in regions
if r.name in req.allowed_regions # residency + tenant tier
and req.model in r.models # model and context variant
and sig[r.name].healthy]
if not cands:
raise NoEligibleRegion(req.tenant) # retryable; never widen the filter
scores = {}
for r in cands:
s = sig[r.name]
queue_ms = 1000.0 * (s.waiting_prefill_tokens + s.local_inflight_tokens) / s.pool_prefill_tps
uncached = req.prompt_tokens
if r.name == sticky_region:
uncached -= req.cached_prefix_tokens # prefix cache lives here
est = ttft_ms(req.rtt_ms[r.name], queue_ms, uncached, s.replica_prefill_tps, req.warm[r.name])
est += r.cost_ms_equiv * req.expected_tokens / 1000.0 # price as latency
scores[r.name] = est
best = min(scores.values())
weights = {n: math.exp(-(v - best) / temp_ms) for n, v in scores.items()}
pick = random.choices(list(weights), weights=list(weights.values()))[0]
sig[pick].local_inflight_tokens += req.prompt_tokens # count our own sends
return pickThe temperature controls how strongly the router prefers the best score. At 150 ms, a region 150 ms worse than the best receives about 37 percent of the best region's weight, which spreads load enough to blunt the herd effect while still favouring the better region. Expressing cost as milliseconds of latency is a deliberate simplification. It makes the trade-off a single tunable number that product and finance can argue about.
Conversation affinity and the prefix cache
In a multi-turn conversation, the prompt for turn five contains turns one to four. If the region that served turn four still holds those KV blocks, it prefills only the new message. Any other region must prefill the whole history again. That is why the code subtracts cached tokens only for the sticky region. With a 30,000-token history and the same 8,000 tokens per second, moving the conversation costs 3.75 seconds of prefill. A region would need to be almost four seconds less congested to be worth switching to. That rarely happens, so in practice conversations stay where they are until the region becomes unhealthy or very overloaded. How cached blocks are kept and evicted is covered in Prefix Caching in Depth. The affinity credit should decay with idle time, because an eviction-prone cache may no longer hold a conversation that has been idle for twenty minutes.
Spillover with hysteresis
Spillover is the policy for moving new traffic out of a hot region. Without hysteresis it oscillates. A workable pattern has four parts.
- Two thresholds. Start spilling when estimated queue wait exceeds, for example, 400 ms. Stop only when it falls below 150 ms.
- Ramped weights. Move at most 10 percent of new sessions per interval, never the whole region at once, and never existing streams.
- Receiver headroom. Only spill into a region whose own estimate is below its start threshold after accounting for the incoming load. A 20 percent spill can turn a comfortable neighbour into a second hot region.
- A spill cap. Limit cross-region traffic to the amount the receiving regions were provisioned to absorb. Beyond that, shedding or queueing with an honest wait message is better than overloading everyone.
Worked example: evening peak in Mumbai
A consumer assistant serves India and Southeast Asia from ap-south and ap-southeast, with an EU region for EU tenants and global overflow. At 21:00 IST, ap-south reports 432,000 waiting prefill tokens. Its 60 replicas each prefill 8,000 tokens per second, so the pool clears 480,000 tokens per second and a new request waits about 900 ms. This triggers spillover. The global router starts moving 10 percent of new non-EU sessions per minute to ap-southeast, where the estimate is 80 ms plus 45 ms of extra RTT. Existing conversations stay put because their affinity credit outweighs the gap. EU-pinned tenants never enter the candidate list for either Asian region. After twenty minutes, ap-southeast's estimate reaches 300 ms and the spill stops growing. At 23:30, ap-south drops below 150 ms and the weights ramp back. TTFT p95 for Indian users during the peak fell from 2.4 s under the nearest-region rule to 1.1 s, at the cost of about 6 percent more cross-region traffic. These numbers are an illustrative model, not a benchmark. Measure your own before you tune thresholds.
Failure modes
- Herding. Every router picks the same quiet region from stale signals. Fix it with weighted choice, local send accounting and smoothing.
- Oscillation. Spillover without hysteresis flips traffic every reporting period.
- Residency leak by fallback. An error handler retries "anywhere" during an outage. Make the eligible-region list the only list the retry path can see.
- Healthy but saturated. Probes return 200 while the queue is minutes long. Base health on serving-engine signals and a synthetic generation, not on the HTTP layer.
- Retry storms across regions. Client retries combine with router fallbacks and multiply load on the surviving regions. Use retry budgets and a single owner for failover.
- Cold capacity counted as warm. An autoscaled region reports free replicas whose weights are still loading.
Trade-offs
Nearest-region routing is simple, predictable and good enough when every region has ample headroom. Load-aware routing improves tail TTFT and lets you run regions hotter, but it adds a control loop that can misbehave, and it makes per-region load harder to forecast. Strong affinity saves prefill compute and latency but concentrates long conversations in one region. A cost term saves money but will sometimes route users to a slower region by design. Agree that trade-off explicitly rather than discovering it in a latency review.
What to do next
- Break your current TTFT into RTT, queue and prefill per region from logs; check whether queueing or distance dominates.
- Write the residency and tenant constraints as a declarative eligible-region list, and make every retry path use it.
- Export waiting prefill tokens, prefill throughput and preemption counts from each serving pool to the router.
- Implement scoring with weighted random choice and local send accounting; start with a high temperature.
- Add affinity credit for the last-served region and decay it with idle time.
- Add spillover with two thresholds, ramped weights, receiver headroom checks and a cap.
- Run a game day: saturate one region and confirm that spill, residency and failback behave as designed.