"LLM routing" names two different decisions that happen to share a word. The first is which model should answer a request: a small, cheap model or a large, expensive one. The second is which replica of the chosen model should run it: which of the GPUs serving that model gets the prompt. The first decision is about cost and quality; the second is about latency and GPU efficiency. They need different signals, fail in different ways and usually live in different components, but both are called routers, and teams regularly solve one while believing they have solved the other.
This article covers both, from first principles. It starts with model-tier routing (rules, classifiers, routers trained on preference data, and cascades), then explains why ordinary load balancing performs badly in front of LLM replicas, and builds a cache-aware replica router with a load guard. A worked example puts numbers on the prefill work that routing saves, and the article ends with how the Kubernetes ecosystem packages these ideas, the failure modes, and a checklist.
Two routers in one request path
Two decisions, two layers
Keep the layers separate in your head and in your architecture.
| Tier routing | Replica routing | |
|---|---|---|
| Decides | Which model (or provider) answers | Which GPU replica of that model runs it |
| Optimises | Cost per acceptable answer | Latency, throughput, cache hits |
| Inputs | Prompt content, task type, tenant, budget | Replica load, KV cache contents, session |
| Changes output? | Yes: a different model writes the answer | No: same weights, same answer distribution |
| Typical home | Application or AI gateway | Inference gateway in front of the serving fleet |
The last row explains a common confusion. An AI gateway is a policy plane (keys, budgets, rate limits, provider failover) and may host tier routing, but it usually sees a remote provider endpoint, not individual GPUs. Replica routing needs to see inside the fleet: which replica is busy, which holds which prefix in its KV cache. That is why it lives in an inference-aware gateway or a router process next to the model servers.
Tier routing: which model answers
The economics are simple. If a large model costs ten times a small one per token and half your traffic can be answered well by the small one, routing that half cuts spend by 45%. The difficulty is knowing in advance which half. There are four common approaches, in rising order of sophistication.
- Static rules. Route by endpoint, task type or tenant tier: classification and extraction go to the small model, open-ended reasoning to the large one. Cheap, predictable and often enough, because product surfaces already separate easy from hard work.
- A trained classifier. A small model, or a linear head on embeddings, predicts from the prompt whether the cheap model's answer would be acceptable. It adds a few milliseconds and needs labelled data for your traffic.
- Preference-trained routers. RouteLLM (Ong and colleagues, LMSYS and UC Berkeley, 2024) trained routers on Chatbot Arena preference data, including a matrix-factorisation router that scores how much a strong model would beat a weak one on a given query. The authors report cost reductions of over 85% on MT Bench, 45% on MMLU and 35% on GSM8K compared with sending everything to GPT-4, while keeping 95% of GPT-4's performance. The spread across benchmarks is the lesson: savings depend on how much of your traffic is genuinely easy.
- Cascades. Send every request to the small model first, check the answer (with a verifier model, the model's own confidence, a schema validation or a test run), and escalate to the large model only on failure. Cascades need no prediction, but escalated requests pay for both models and twice the latency, so they suit offline and batch work better than interactive chat.
Calibrating the threshold
Whatever produces the score, you still need a threshold, and the threshold should come from data rather than intuition. Collect a few thousand real prompts, run both models, have graders or a rubric mark whether the small model's answer was acceptable, and choose the threshold that meets a quality target at minimum cost.
def calibrate(scores, small_ok, cost_small, cost_large, min_quality):
"""scores: router score per prompt (higher = needs large model).
small_ok: whether the small model's answer was acceptable.
Returns the threshold with the lowest cost meeting min_quality."""
best = None
for t in sorted(set(scores)):
n = len(scores)
to_large = [s >= t for s in scores]
quality = sum(1 if big else ok for big, ok in zip(to_large, small_ok)) / n
cost = sum(cost_large if big else cost_small for big in to_large) / n
if quality >= min_quality and (best is None or cost < best[1]):
best = (t, cost, quality)
return best # (threshold, mean cost, quality)
def route(prompt, router, threshold):
return "large" if router.score(prompt) >= threshold else "small"This assumes the large model's answer is always acceptable; if not, measure it too. Recalibrate whenever either model or the traffic changes, and track the escalation rate in production: if it moves, one of them did.
Why ordinary load balancing fails for LLMs
Now the second layer. A classic HTTP load balancer assumes requests are roughly equal in cost and that servers are stateless. Both assumptions fail for LLM serving.
Request cost varies by orders of magnitude. A 200-token chat turn and a 30,000-token document summary are both one request; the second can occupy a replica's prefill capacity and a large share of its KV cache for a long time. Round robin spreads request counts evenly while the work stays uneven. Least-connections is better, but outstanding-request counts still treat a short and a long request alike.
Replicas are not stateless. Every engine that implements prefix caching keeps the KV blocks of recent prompts in GPU memory, and a new request whose prompt starts with a cached prefix skips the prefill computation for those tokens. That cache is per replica. Send turn two of a conversation to a different replica from turn one and the whole history is prefilled again from scratch. With N replicas and random placement, the chance of landing on the replica that holds your prefix is 1/N, so most of the cache's value is thrown away by the router before the engine ever sees the request.
Replica routing strategies
Replica routing strategies, from simplest to most capable:
- Least load. Pick the replica with the shortest queue, or the fewest running requests, or the lowest KV-cache utilisation, as reported by the server's metrics. This fixes unequal request cost but ignores cache contents.
- Session affinity. Hash a session or user id to a replica. Multi-turn chats hit a warm cache, but popular sessions overload their replicas, and shared prefixes across users (a common system prompt, a shared document) are not exploited.
- Prefix-hash affinity. Hash the first K tokens or characters of the prompt, so requests with the same system prompt or document land together. It captures cross-user sharing, but a single dominant prefix sends all traffic to one replica.
- Cache-aware with a load guard. Keep an approximate index of which prefixes each replica has recently processed, send each request to the replica with the longest match if the match is good enough and that replica is not overloaded, and otherwise fall back to least load. This is the design most modern routers use.
- Adapter affinity. When replicas serve many
LoRAadapters, prefer replicas that already have the request's adapter loaded, since loading one costs time and memory.
SGLang's router is a well-documented example of the fourth strategy. It keeps an approximate radix tree per worker built from the request history, storing raw text rather than token ids to avoid tokenising in the router. If the best worker's prefix match ratio exceeds cache_threshold (default 0.5) it routes there; otherwise it routes to the worker with the smallest tree, the one likely to have the most free cache. When the load gap between the busiest and idlest worker exceeds both an absolute threshold (balance_abs_threshold, default 32) and a relative one, it abandons cache affinity and balances by load until the gap closes.
A cache-aware router in code
A compact version of that logic, simplified to a character-prefix index:
class Replica:
def __init__(self, name):
self.name = name
self.running = 0 # refreshed from server metrics
self.recent = [] # recent prompt prefixes, bounded
def match_len(a, b):
n = min(len(a), len(b))
i = 0
while i < n and a[i] == b[i]:
i += 1
return i
def pick(replicas, prompt, cache_threshold=0.5, abs_gap=32, rel_gap=1.5):
loads = [r.running for r in replicas]
hi, lo = max(loads), min(loads)
if hi - lo > abs_gap and hi > rel_gap * max(lo, 1):
return min(replicas, key=lambda r: r.running) # imbalanced: load wins
best, best_len = None, 0
for r in replicas:
m = max((match_len(prompt, p) for p in r.recent), default=0)
if m > best_len:
best, best_len = r, m
if best is not None and best_len / max(len(prompt), 1) > cache_threshold:
return best # warm cache wins
return min(replicas, key=lambda r: (r.running, len(r.recent)))
def record(replica, prompt, keep=512):
replica.recent.append(prompt[:4096])
del replica.recent[:-keep]Production routers replace the list with a radix tree and evict entries as the engine evicts blocks, but the shape is the same. Note the order of checks: the load guard runs first, so cache affinity can never pile traffic onto a replica that is already far behind.
Worked example: turn five of a support chat
Take a support assistant on eight replicas of one model. Every conversation starts with a 2,000-token system prompt and tool schema; by turn five the history adds another 4,000 tokens, and each new user message is about 150 tokens. Look at turn five.
With round robin, the chance that turn five lands on the replica that served turn four is 1/8. The shared system prompt is probably cached everywhere, because every replica sees it constantly, but the 4,000 tokens of history are cached only where the conversation last ran. Expected prefill for turn five is therefore about 150 + (7/8 x 4,000) = 3,650 tokens.
With cache-aware routing, turn five returns to its replica unless the load guard intervenes. The match ratio is about 6,000 / 6,150, well above 0.5, so prefill is roughly the 150 new tokens. That is about 24 times less prefill work for this turn. Since prefill is the compute-heavy, latency-dominating phase for long prompts, time to first token drops and the freed GPU time goes to other requests. The gain disappears if the KV cache is too small to hold the conversation between turns, which is why routing and KV cache sizing have to be planned together.
How Kubernetes packages it
On Kubernetes, the Gateway API Inference Extension (a Kubernetes SIG project) standardises this layer. An InferencePool resource groups the model-server pods for a workload and references an Endpoint Picker (EPP) service; the gateway's Envoy proxy calls the EPP through its external processing (ext_proc) filter for each request, and the EPP chooses the pod using signals such as queue depth, KV cache use and prefix-cache affinity. Several gateways implement it, and the llm-d project builds its scheduling on the same extension. SGLang ships its router as a standalone process. Whichever you use, check which metrics it scrapes, how often, and whether its prefix index follows the engine's real evictions or only approximates them.
Failure modes
- Herding. One very popular prefix pulls traffic onto one replica until it saturates. The load guard is the fix; tune its thresholds against your traffic, not defaults.
- Stale signals. Load scraped every few seconds lets bursts land on a replica that has already filled up. Count requests the router has itself dispatched since the last scrape.
- Index drift. The router believes a prefix is cached after the engine evicted it. Affinity then buys nothing and still costs balance.
- Retries that double cost. A retried long prompt prefills twice, and cascades pay for both models; cap retries and escalations per request.
- Mid-stream failover. A replica that dies during streaming cannot hand over its KV cache; the retry restarts generation, and clients must tolerate it.
- Silent tier drift. A model upgrade changes which prompts the small model handles well, and the router keeps the old threshold.
Trade-offs
Cache affinity and load balance pull in opposite directions; the right threshold depends on prompt length distribution and how much history is shared, so measure it. Tier routing trades quality risk for cost and should be gated on a quality metric you trust. Cascades cut cost without a predictor but add latency. And every router is now a tier-0 dependency in the request path, so it needs redundancy, a fallback to plain least-load when its state is lost, and its own latency budget, which the full serving stack leaves little room for.
What to do next
- Write down which of the two routing problems you are solving, and where each decision lives today.
- Measure your prefix-cache hit rate per replica; if it is near 1/N of what one replica alone would achieve, your replica router is discarding the cache.
- Replace round robin with least-load on queue depth or KV utilisation, then add cache-aware routing with a load guard.
- Count dispatched-but-unreported requests in the router so bursts do not overrun stale metrics.
- For tier routing, label a few thousand real prompts, calibrate the threshold against a quality target, and alert on escalation-rate changes.
- Recalibrate after every model upgrade, and test the router's fallback when it loses its index.
- Read about continuous batching to see how admitted requests share a replica's step.