Every LLM request passes through several caches before a GPU does fresh work. The serving engine may already hold the attention keys and values for your system prompt. A router may have sent you to the replica that holds them. A hosted API may bill most of your prompt at a tenth of the normal input price. And a gateway in front of everything may answer without calling a model at all. Each of these is a caching layer, each has its own hit condition, and they interact in ways that are easy to miss when you tune them one at a time.

This article walks the stack from GPU memory up to the gateway. It covers how prefix caches identify reusable work, how much memory that work occupies, how routing decides whether a cache can hit at all, how provider prompt caching is priced and invalidated, how hit rates combine across layers, and how to stop caches leaking between tenants. Key design, semantic caching and invalidation of response caches are covered in LLM caching strategies; this page is about the machinery underneath.

The stack, named by what each layer reuses

Name the layers by what they reuse, because the hit condition follows directly from that.

LayerReusesHit conditionSaves
L4 gateway response cacheA finished answerSame normalised request (or, for semantic caches, a close one)The whole call
L3 provider prompt cachePrefill of a prompt prefixByte-identical prefix within the cache lifetimeMost input cost and time to first token
L2 cache-aware routingNothing by itselfRequest lands where its prefix is cachedMakes L1 hits possible
L1 engine prefix cacheKV blocks in GPU memoryIdentical token blocks from the start of the sequencePrefill compute
L0 KV offload tierKV blocks in DRAM or SSDSame as L1, after eviction from HBMPrefill compute, at the cost of a copy
Clientapp, agentGatewayL4 response cachemissHosted APIL3 prompt cacheor self-hosted:missPrefix-aware routerL2 placementReplica AL1 HBM blocksReplica BL1 HBM blocksReplica CL1 HBM blocksL0 offload tier: host DRAM, then local SSDevicted KV blocks parked instead of recomputedHit condition at each layer:L4 same requestL3 same prefix, in TTLL2 same replicaL1/L0 same block hashes
Caching layers on the request path. If you call a hosted API, L3 replaces L0 to L2; if you self-host, you own L0 to L2.

L1: the engine prefix cache

A transformer's prefill computes a key and a value vector per token, per layer. Those depend only on the tokens up to that point, so two requests that share a prefix share the same KV entries for it. Engines such as vLLM store KV in fixed-size blocks and give each full block a hash that chains in its parent's hash. Equal hashes mean equal prefixes, so lookup is a dictionary probe per block. According to the vLLM design notes, the hash covers the parent block's hash, the block's tokens and extra keys that change the computation, such as LoRA adapter IDs, multimodal input hashes and an optional cache salt.

import hashlib

def block_hashes(tokens, block_size, extra=(), cache_salt=None):
    # one hash per FULL block; a partial tail block is never shared
    hashes, parent = [], None
    for i in range(0, len(tokens) - block_size + 1, block_size):
        blk = tuple(tokens[i:i + block_size])
        keys = tuple(extra) + ((cache_salt,) if (i == 0 and cache_salt) else ())
        h = hashlib.sha256(repr((parent, blk, keys)).encode()).hexdigest()
        hashes.append(h)
        parent = h
    return hashes

def cached_prefix_len(tokens, block_size, table, **kw):
    n = 0
    for h in block_hashes(tokens, block_size, **kw):
        if h not in table:          # first miss ends the reusable prefix
            break
        n += block_size
    return n

Three consequences follow. Reuse is prefix-only: one changed token early in the prompt invalidates every block after it, so a timestamp at the top of a system prompt defeats the cache entirely. Reuse is block-granular, so the last partial block is always recomputed. And eviction is LRU over blocks no running request is using; vLLM appends freed blocks to its free queue in reverse order, so the tail of a sequence, which is least likely to be shared, is evicted first. The fix for low hit rates is nearly always prompt layout: stable content first, per-request content last.

L0: what a cached prefix weighs, and offload tiers

GPU memory is the scarce resource, so it helps to know what a cached prefix weighs. KV bytes per token are 2 (keys and values) times layers times KV heads times head dimension times bytes per element. For a model shaped like Llama 3 8B, with 32 layers, 8 KV heads under grouped-query attention, head dimension 128 and 16-bit values, that is 2 x 32 x 8 x 128 x 2 = 131,072 bytes, or 128 KiB per token. A 6,000-token system prompt and tool catalogue therefore holds about 750 MiB of KV. Ten such tenants fill several gigabytes of HBM that the scheduler would rather spend on batch size.

Offload tiers park evicted blocks in host DRAM or on local NVMe instead of dropping them, and copy them back on a hit. Whether that wins is a bandwidth question. Moving 750 MiB over a PCIe link that sustains around 25 GB/s takes about 30 ms. Recomputing 6,000 tokens of prefill on an 8B model costs roughly 2 x 8e9 x 6,000, about 1e14 floating-point operations, which on a current data-centre GPU is on the order of a hundred milliseconds or more once real utilisation is counted. For long shared prefixes, restoring from DRAM wins clearly. For short prefixes, or SSD tiers under heavy contention, recompute can be cheaper, so measure on your hardware. KV-cache offloading covers the tier design in detail.

L2: routing decides whether a cache can hit

A prefix cache on replica A does nothing for a request that the load balancer sends to replica B. With round-robin over N replicas, a prefix shared by all traffic is computed N times and occupies HBM N times. Cache-aware routing fixes that by hashing the stable prefix, or the tenant, to a preferred replica. Pure affinity creates hot spots, though, so production routers bound the load: prefer the replica that owns the prefix unless it is much busier than average, then fall back to the next choice.

import hashlib

def score(key, replica):
    return int(hashlib.blake2b(f"{key}|{replica}".encode(), digest_size=8).hexdigest(), 16)

def route(prefix_key, replicas, inflight, slack=1.25):
    # rendezvous hashing gives every key a stable preference order over replicas
    avg = sum(inflight.values()) / len(replicas)
    for r in sorted(replicas, key=lambda r: score(prefix_key, r), reverse=True):
        if inflight[r] + 1 <= slack * avg + 1:   # bounded load: skip overloaded owners
            return r
    return min(replicas, key=inflight.get)

# prefix_key: hash of the first K tokens of the rendered prompt, or the tenant id

Rendezvous hashing has a useful property here: adding or removing a replica moves only the keys that ranked it first, so a scale-out event does not cold-start every cache. Choose the routing key carefully. Hashing the first few thousand tokens captures shared system prompts across tenants. Hashing the conversation ID keeps multi-turn chats on one replica, where each turn extends the previous turn's cached prefix. Many systems combine both. Prefix caching at scale discusses fleet-level placement further.

L3: provider prompt caching

If you call a hosted model, layers L0 to L2 belong to the provider, and what you control is prompt layout and, on some APIs, explicit breakpoints. Mechanics differ by provider and change often, so treat the following as a snapshot and check current documentation before relying on any number.

On Anthropic's API you mark a breakpoint by adding "cache_control": {"type": "ephemeral"} to a content block. The cached prefix is assembled in a fixed order, tools, then system, then messages, and a change at one level invalidates that level and everything after it. The default lifetime is 5 minutes, refreshed on each hit, and a 1-hour lifetime can be requested with "ttl": "1h". Cache writes are billed at 1.25 times the base input price for the 5-minute lifetime and 2 times for one hour. Reads are typically 0.1 times base, lower on some newer models. A request can carry up to four breakpoints. On a read, the system looks back up to 20 blocks from each breakpoint for an earlier cache entry, which is why a growing conversation keeps hitting if each turn adds fewer than 20 blocks. Prompts below a model-specific minimum length are simply not cached, with no error.

{
  "model": "<your-model-id>",
  "max_tokens": 1024,
  "tools": [ ...stable tool definitions... ],
  "system": [
    {"type": "text", "text": "<6,000 tokens of policy and product docs>",
     "cache_control": {"type": "ephemeral"}}
  ],
  "messages": [ ...conversation so far..., {"role": "user", "content": "<new turn>"} ]
}
// response usage: cache_creation_input_tokens, cache_read_input_tokens, input_tokens

OpenAI caches eligible prompts automatically, with no markup, and reports reuse in usage.input_tokens_details.cached_tokens. An optional prompt_cache_key groups requests that share a prefix. Minimum lengths, retention and discounts vary by model. On both platforms the rule is the same as in your own engine: stable content first, volatile content last, and log the cache counters on every response.

How hit rates compose

Layers multiply rather than add, and an upper layer changes the traffic a lower one sees. Take a support assistant with a 6,000-token stable prefix and a 500-token per-request suffix, on a hosted API with a 5-minute cache. Let h4 be the gateway hit rate and h3 the prefix hit rate among requests that reach the API. Input cost relative to no caching is:

def relative_input_cost(h4, h3, prefix=6000, suffix=500, write=1.25, read=0.10):
    per_call = prefix * (h3 * read + (1 - h3) * write) + suffix
    return (1 - h4) * per_call / (prefix + suffix)

print(round(relative_input_cost(h4=0.0, h3=0.0), 2))   # 1.23 - every call writes, nothing reads
print(round(relative_input_cost(h4=0.0, h3=0.9), 2))   # 0.28
print(round(relative_input_cost(h4=0.1, h3=0.9), 2))   # 0.25

Two things stand out. A prompt cache that never hits is worse than none, because writes cost 25 percent more than plain input; caching rarely repeated prefixes loses money. And a 10 percent gateway hit rate removes only about 3 points of cost, because the calls it absorbs had already been discounted by the prompt cache. There is a subtler effect too. Prompt-cache lifetimes are refreshed by traffic, so a low-volume tenant whose requests arrive every four minutes keeps its prefix warm. Put a response cache in front that absorbs a third of those requests, and the gaps stretch past five minutes; h3 collapses. When you add an upper layer, re-measure the layers below it.

Tenant isolation and timing side channels

Shared caches are a side channel. A cached prefix returns its first token faster, so an attacker who can send prompts to the same cache can time responses to test guesses about another tenant's prompt. Researchers auditing hosted APIs have reported cache sharing across users at some providers. Treat cache scope as a security decision.

  • Salt per trust boundary. vLLM accepts a per-request cache_salt that is mixed into the first block's hash, so only requests with the same salt can share blocks. Set it from the authenticated tenant, never from client input.
  • Share only public prefixes across tenants. A common system prompt and tool catalogue are safe to share; retrieved documents and conversation history are not.
  • Scope response caches the same way. The gateway key must include tenant, model version and every permission that shaped the answer.
  • Expire on deletion. When a user deletes data, their salted prefixes must age out or be purged with it.

Operating the layers

Instrument each layer with its own hit rate and its own cost, and alert on changes rather than absolute values.

  • Engine: prefix-cache hit rate by tokens, not requests; HBM blocks in use; evictions per second; offload restores and their latency.
  • Router: fraction of requests sent to their first-choice replica; per-replica in-flight spread.
  • Provider: the sum of cache-read tokens over total input tokens, per prompt template; cache-write tokens, which are pure overhead if reads do not follow.
  • Gateway: hits, staleness complaints, and hit rate per tenant.

The most common regression is invisible in latency dashboards: someone adds a request ID or the current date to a system prompt and the provider hit rate drops from 90 percent to zero. Add a CI check that renders each template twice with different inputs and asserts the leading N tokens are identical. Engine internals are covered in vLLM continuous batching and KV cache.

Failure modes

  • Volatile prefix. Timestamps, user names or shuffled tool lists near the start of the prompt. Hit rate near zero, write charges on every call.
  • Affinity hot spot. One giant tenant pinned to one replica without a load bound; its queue grows while others idle.
  • Scale-out cold start. Modulo hashing remaps most keys when the replica count changes; every cache goes cold at once, exactly during a traffic peak.
  • Tokeniser or template drift. A library upgrade changes how the chat template renders whitespace; prefixes stop matching across the fleet.
  • Cross-tenant timing leak. Unsalted shared cache on a multi-tenant endpoint.
  • Offload thrash. An SSD tier sized below the working set; restores queue behind writes and are slower than recompute.

Trade-offs

DecisionGainCost
Offload to DRAM or SSDLong prefixes survive evictionCopy latency; SSD can lose to recompute
Prefix affinity routingL1 hits across the fleetHot spots without a load bound
Cross-tenant sharingHigher hit rate on shared promptsTiming side channel
1-hour TTL instead of 5 minutesSurvives quiet periodsWrites cost 2x instead of 1.25x
Gateway response cacheSkips whole callsCan starve the prompt cache of warming traffic

What to do next

  1. Log cache counters from every provider response and every engine, per prompt template, starting today.
  2. Reorder prompts: tools and system first, then retrieved context, then history, then the new turn. Remove timestamps from the prefix.
  3. Add a CI test that asserts each template's leading tokens are stable across inputs.
  4. For self-hosted fleets, replace round-robin with bounded-load rendezvous routing on a prefix or conversation key.
  5. Compute KV bytes per token for your model and decide whether a DRAM offload tier pays at your prefix lengths.
  6. Set a cache salt per tenant and review which prefixes may be shared.
  7. Only after the lower layers are healthy, add a gateway response cache, then re-measure prompt-cache hit rates. The prompt caching guide covers the provider layer in more depth.
Key takeaway: LLM caches stack: an engine prefix cache, offload tiers, cache-aware routing, provider prompt caching and gateway response caches. Each hits only on its own condition, and they interact. Keep prefixes stable, route by prefix with a load bound, salt caches per tenant, and re-measure the lower layers whenever you add an upper one.