"LLM routing" names two different decisions that happen to share a word. The first is which model should answer a request: a small, cheap model or a large, expensive one. The second is which replica of the chosen model should run it: which of the GPUs serving that model gets the prompt. The first decision is about cost and quality; the second is about latency and GPU efficiency. They need different signals, fail in different ways and usually live in different components, but both are called routers, and teams regularly solve one while believing they have solved the other.

This article covers both, from first principles. It starts with model-tier routing (rules, classifiers, routers trained on preference data, and cascades), then explains why ordinary load balancing performs badly in front of LLM replicas, and builds a cache-aware replica router with a load guard. A worked example puts numbers on the prefill work that routing saves, and the article ends with how the Kubernetes ecosystem packages these ideas, the failure modes, and a checklist.

Two routers in one request path

Client requestprompt, session, tenantTier routerwhich model?rules / classifier / cascadehardeasyLarge model poolreplica router + N replicasSmall model poolreplica router + M replicasReplica router (per pool)prefix match + load guardReplica 1KV cache: prefix AReplica 2KV cache: prefix BReplica 3idleSignalsqueue depth, running requests,KV utilisation, prefix index, adapter loadedTier routing trades quality for cost; replica routing trades cache locality for load balance.
A tier router picks the model; a per-pool replica router picks the GPU replica, preferring the one whose KV cache already holds the prompt's prefix unless load is badly skewed.

Two decisions, two layers

Keep the layers separate in your head and in your architecture.

Tier routingReplica routing
DecidesWhich model (or provider) answersWhich GPU replica of that model runs it
OptimisesCost per acceptable answerLatency, throughput, cache hits
InputsPrompt content, task type, tenant, budgetReplica load, KV cache contents, session
Changes output?Yes: a different model writes the answerNo: same weights, same answer distribution
Typical homeApplication or AI gatewayInference gateway in front of the serving fleet

The last row explains a common confusion. An AI gateway is a policy plane (keys, budgets, rate limits, provider failover) and may host tier routing, but it usually sees a remote provider endpoint, not individual GPUs. Replica routing needs to see inside the fleet: which replica is busy, which holds which prefix in its KV cache. That is why it lives in an inference-aware gateway or a router process next to the model servers.

Tier routing: which model answers

The economics are simple. If a large model costs ten times a small one per token and half your traffic can be answered well by the small one, routing that half cuts spend by 45%. The difficulty is knowing in advance which half. There are four common approaches, in rising order of sophistication.

  • Static rules. Route by endpoint, task type or tenant tier: classification and extraction go to the small model, open-ended reasoning to the large one. Cheap, predictable and often enough, because product surfaces already separate easy from hard work.
  • A trained classifier. A small model, or a linear head on embeddings, predicts from the prompt whether the cheap model's answer would be acceptable. It adds a few milliseconds and needs labelled data for your traffic.
  • Preference-trained routers. RouteLLM (Ong and colleagues, LMSYS and UC Berkeley, 2024) trained routers on Chatbot Arena preference data, including a matrix-factorisation router that scores how much a strong model would beat a weak one on a given query. The authors report cost reductions of over 85% on MT Bench, 45% on MMLU and 35% on GSM8K compared with sending everything to GPT-4, while keeping 95% of GPT-4's performance. The spread across benchmarks is the lesson: savings depend on how much of your traffic is genuinely easy.
  • Cascades. Send every request to the small model first, check the answer (with a verifier model, the model's own confidence, a schema validation or a test run), and escalate to the large model only on failure. Cascades need no prediction, but escalated requests pay for both models and twice the latency, so they suit offline and batch work better than interactive chat.

Calibrating the threshold

Whatever produces the score, you still need a threshold, and the threshold should come from data rather than intuition. Collect a few thousand real prompts, run both models, have graders or a rubric mark whether the small model's answer was acceptable, and choose the threshold that meets a quality target at minimum cost.

def calibrate(scores, small_ok, cost_small, cost_large, min_quality):
    """scores: router score per prompt (higher = needs large model).
    small_ok: whether the small model's answer was acceptable.
    Returns the threshold with the lowest cost meeting min_quality."""
    best = None
    for t in sorted(set(scores)):
        n = len(scores)
        to_large = [s >= t for s in scores]
        quality = sum(1 if big else ok for big, ok in zip(to_large, small_ok)) / n
        cost = sum(cost_large if big else cost_small for big in to_large) / n
        if quality >= min_quality and (best is None or cost < best[1]):
            best = (t, cost, quality)
    return best   # (threshold, mean cost, quality)

def route(prompt, router, threshold):
    return "large" if router.score(prompt) >= threshold else "small"

This assumes the large model's answer is always acceptable; if not, measure it too. Recalibrate whenever either model or the traffic changes, and track the escalation rate in production: if it moves, one of them did.

Why ordinary load balancing fails for LLMs

Now the second layer. A classic HTTP load balancer assumes requests are roughly equal in cost and that servers are stateless. Both assumptions fail for LLM serving.

Request cost varies by orders of magnitude. A 200-token chat turn and a 30,000-token document summary are both one request; the second can occupy a replica's prefill capacity and a large share of its KV cache for a long time. Round robin spreads request counts evenly while the work stays uneven. Least-connections is better, but outstanding-request counts still treat a short and a long request alike.

Replicas are not stateless. Every engine that implements prefix caching keeps the KV blocks of recent prompts in GPU memory, and a new request whose prompt starts with a cached prefix skips the prefill computation for those tokens. That cache is per replica. Send turn two of a conversation to a different replica from turn one and the whole history is prefilled again from scratch. With N replicas and random placement, the chance of landing on the replica that holds your prefix is 1/N, so most of the cache's value is thrown away by the router before the engine ever sees the request.

Replica routing strategies

Replica routing strategies, from simplest to most capable:

  • Least load. Pick the replica with the shortest queue, or the fewest running requests, or the lowest KV-cache utilisation, as reported by the server's metrics. This fixes unequal request cost but ignores cache contents.
  • Session affinity. Hash a session or user id to a replica. Multi-turn chats hit a warm cache, but popular sessions overload their replicas, and shared prefixes across users (a common system prompt, a shared document) are not exploited.
  • Prefix-hash affinity. Hash the first K tokens or characters of the prompt, so requests with the same system prompt or document land together. It captures cross-user sharing, but a single dominant prefix sends all traffic to one replica.
  • Cache-aware with a load guard. Keep an approximate index of which prefixes each replica has recently processed, send each request to the replica with the longest match if the match is good enough and that replica is not overloaded, and otherwise fall back to least load. This is the design most modern routers use.
  • Adapter affinity. When replicas serve many LoRA adapters, prefer replicas that already have the request's adapter loaded, since loading one costs time and memory.

SGLang's router is a well-documented example of the fourth strategy. It keeps an approximate radix tree per worker built from the request history, storing raw text rather than token ids to avoid tokenising in the router. If the best worker's prefix match ratio exceeds cache_threshold (default 0.5) it routes there; otherwise it routes to the worker with the smallest tree, the one likely to have the most free cache. When the load gap between the busiest and idlest worker exceeds both an absolute threshold (balance_abs_threshold, default 32) and a relative one, it abandons cache affinity and balances by load until the gap closes.

A cache-aware router in code

A compact version of that logic, simplified to a character-prefix index:

class Replica:
    def __init__(self, name):
        self.name = name
        self.running = 0          # refreshed from server metrics
        self.recent = []          # recent prompt prefixes, bounded

def match_len(a, b):
    n = min(len(a), len(b))
    i = 0
    while i < n and a[i] == b[i]:
        i += 1
    return i

def pick(replicas, prompt, cache_threshold=0.5, abs_gap=32, rel_gap=1.5):
    loads = [r.running for r in replicas]
    hi, lo = max(loads), min(loads)
    if hi - lo > abs_gap and hi > rel_gap * max(lo, 1):
        return min(replicas, key=lambda r: r.running)       # imbalanced: load wins
    best, best_len = None, 0
    for r in replicas:
        m = max((match_len(prompt, p) for p in r.recent), default=0)
        if m > best_len:
            best, best_len = r, m
    if best is not None and best_len / max(len(prompt), 1) > cache_threshold:
        return best                                          # warm cache wins
    return min(replicas, key=lambda r: (r.running, len(r.recent)))

def record(replica, prompt, keep=512):
    replica.recent.append(prompt[:4096])
    del replica.recent[:-keep]

Production routers replace the list with a radix tree and evict entries as the engine evicts blocks, but the shape is the same. Note the order of checks: the load guard runs first, so cache affinity can never pile traffic onto a replica that is already far behind.

Worked example: turn five of a support chat

Take a support assistant on eight replicas of one model. Every conversation starts with a 2,000-token system prompt and tool schema; by turn five the history adds another 4,000 tokens, and each new user message is about 150 tokens. Look at turn five.

With round robin, the chance that turn five lands on the replica that served turn four is 1/8. The shared system prompt is probably cached everywhere, because every replica sees it constantly, but the 4,000 tokens of history are cached only where the conversation last ran. Expected prefill for turn five is therefore about 150 + (7/8 x 4,000) = 3,650 tokens.

With cache-aware routing, turn five returns to its replica unless the load guard intervenes. The match ratio is about 6,000 / 6,150, well above 0.5, so prefill is roughly the 150 new tokens. That is about 24 times less prefill work for this turn. Since prefill is the compute-heavy, latency-dominating phase for long prompts, time to first token drops and the freed GPU time goes to other requests. The gain disappears if the KV cache is too small to hold the conversation between turns, which is why routing and KV cache sizing have to be planned together.

How Kubernetes packages it

On Kubernetes, the Gateway API Inference Extension (a Kubernetes SIG project) standardises this layer. An InferencePool resource groups the model-server pods for a workload and references an Endpoint Picker (EPP) service; the gateway's Envoy proxy calls the EPP through its external processing (ext_proc) filter for each request, and the EPP chooses the pod using signals such as queue depth, KV cache use and prefix-cache affinity. Several gateways implement it, and the llm-d project builds its scheduling on the same extension. SGLang ships its router as a standalone process. Whichever you use, check which metrics it scrapes, how often, and whether its prefix index follows the engine's real evictions or only approximates them.

Failure modes

  • Herding. One very popular prefix pulls traffic onto one replica until it saturates. The load guard is the fix; tune its thresholds against your traffic, not defaults.
  • Stale signals. Load scraped every few seconds lets bursts land on a replica that has already filled up. Count requests the router has itself dispatched since the last scrape.
  • Index drift. The router believes a prefix is cached after the engine evicted it. Affinity then buys nothing and still costs balance.
  • Retries that double cost. A retried long prompt prefills twice, and cascades pay for both models; cap retries and escalations per request.
  • Mid-stream failover. A replica that dies during streaming cannot hand over its KV cache; the retry restarts generation, and clients must tolerate it.
  • Silent tier drift. A model upgrade changes which prompts the small model handles well, and the router keeps the old threshold.

Trade-offs

Cache affinity and load balance pull in opposite directions; the right threshold depends on prompt length distribution and how much history is shared, so measure it. Tier routing trades quality risk for cost and should be gated on a quality metric you trust. Cascades cut cost without a predictor but add latency. And every router is now a tier-0 dependency in the request path, so it needs redundancy, a fallback to plain least-load when its state is lost, and its own latency budget, which the full serving stack leaves little room for.

What to do next

  1. Write down which of the two routing problems you are solving, and where each decision lives today.
  2. Measure your prefix-cache hit rate per replica; if it is near 1/N of what one replica alone would achieve, your replica router is discarding the cache.
  3. Replace round robin with least-load on queue depth or KV utilisation, then add cache-aware routing with a load guard.
  4. Count dispatched-but-unreported requests in the router so bursts do not overrun stale metrics.
  5. For tier routing, label a few thousand real prompts, calibrate the threshold against a quality target, and alert on escalation-rate changes.
  6. Recalibrate after every model upgrade, and test the router's fallback when it loses its index.
  7. Read about continuous batching to see how admitted requests share a replica's step.
Key takeaway: LLM routing is two decisions: which model answers, which trades quality against cost, and which GPU replica runs the request, which trades KV-cache locality against load balance. Choose models with rules, a classifier, a preference-trained router or a cascade, and set the threshold from labelled traffic against a quality target, recalibrating whenever a model or the traffic changes. Choose replicas with cache-aware routing guarded by load: send a request where its prefix is already cached unless that replica is far busier than the rest. Measure cache hit rate, escalation rate and per-replica load, because a router that ignores them silently wastes GPU time.