Every LLM request passes through several caches before a GPU does fresh work. The serving engine may already hold the attention keys and values for your system prompt. A router may have sent you to the replica that holds them. A hosted API may bill most of your prompt at a tenth of the normal input price. And a gateway in front of everything may answer without calling a model at all. Each of these is a caching layer, each has its own hit condition, and they interact in ways that are easy to miss when you tune them one at a time.
This article walks the stack from GPU memory up to the gateway. It covers how prefix caches identify reusable work, how much memory that work occupies, how routing decides whether a cache can hit at all, how provider prompt caching is priced and invalidated, how hit rates combine across layers, and how to stop caches leaking between tenants. Key design, semantic caching and invalidation of response caches are covered in LLM caching strategies; this page is about the machinery underneath.
The stack, named by what each layer reuses
Name the layers by what they reuse, because the hit condition follows directly from that.
| Layer | Reuses | Hit condition | Saves |
|---|---|---|---|
| L4 gateway response cache | A finished answer | Same normalised request (or, for semantic caches, a close one) | The whole call |
| L3 provider prompt cache | Prefill of a prompt prefix | Byte-identical prefix within the cache lifetime | Most input cost and time to first token |
| L2 cache-aware routing | Nothing by itself | Request lands where its prefix is cached | Makes L1 hits possible |
| L1 engine prefix cache | KV blocks in GPU memory | Identical token blocks from the start of the sequence | Prefill compute |
| L0 KV offload tier | KV blocks in DRAM or SSD | Same as L1, after eviction from HBM | Prefill compute, at the cost of a copy |
L1: the engine prefix cache
A transformer's prefill computes a key and a value vector per token, per layer. Those depend only on the tokens up to that point, so two requests that share a prefix share the same KV entries for it. Engines such as vLLM store KV in fixed-size blocks and give each full block a hash that chains in its parent's hash. Equal hashes mean equal prefixes, so lookup is a dictionary probe per block. According to the vLLM design notes, the hash covers the parent block's hash, the block's tokens and extra keys that change the computation, such as LoRA adapter IDs, multimodal input hashes and an optional cache salt.
import hashlib
def block_hashes(tokens, block_size, extra=(), cache_salt=None):
# one hash per FULL block; a partial tail block is never shared
hashes, parent = [], None
for i in range(0, len(tokens) - block_size + 1, block_size):
blk = tuple(tokens[i:i + block_size])
keys = tuple(extra) + ((cache_salt,) if (i == 0 and cache_salt) else ())
h = hashlib.sha256(repr((parent, blk, keys)).encode()).hexdigest()
hashes.append(h)
parent = h
return hashes
def cached_prefix_len(tokens, block_size, table, **kw):
n = 0
for h in block_hashes(tokens, block_size, **kw):
if h not in table: # first miss ends the reusable prefix
break
n += block_size
return nThree consequences follow. Reuse is prefix-only: one changed token early in the prompt invalidates every block after it, so a timestamp at the top of a system prompt defeats the cache entirely. Reuse is block-granular, so the last partial block is always recomputed. And eviction is LRU over blocks no running request is using; vLLM appends freed blocks to its free queue in reverse order, so the tail of a sequence, which is least likely to be shared, is evicted first. The fix for low hit rates is nearly always prompt layout: stable content first, per-request content last.
L0: what a cached prefix weighs, and offload tiers
GPU memory is the scarce resource, so it helps to know what a cached prefix weighs. KV bytes per token are 2 (keys and values) times layers times KV heads times head dimension times bytes per element. For a model shaped like Llama 3 8B, with 32 layers, 8 KV heads under grouped-query attention, head dimension 128 and 16-bit values, that is 2 x 32 x 8 x 128 x 2 = 131,072 bytes, or 128 KiB per token. A 6,000-token system prompt and tool catalogue therefore holds about 750 MiB of KV. Ten such tenants fill several gigabytes of HBM that the scheduler would rather spend on batch size.
Offload tiers park evicted blocks in host DRAM or on local NVMe instead of dropping them, and copy them back on a hit. Whether that wins is a bandwidth question. Moving 750 MiB over a PCIe link that sustains around 25 GB/s takes about 30 ms. Recomputing 6,000 tokens of prefill on an 8B model costs roughly 2 x 8e9 x 6,000, about 1e14 floating-point operations, which on a current data-centre GPU is on the order of a hundred milliseconds or more once real utilisation is counted. For long shared prefixes, restoring from DRAM wins clearly. For short prefixes, or SSD tiers under heavy contention, recompute can be cheaper, so measure on your hardware. KV-cache offloading covers the tier design in detail.
L2: routing decides whether a cache can hit
A prefix cache on replica A does nothing for a request that the load balancer sends to replica B. With round-robin over N replicas, a prefix shared by all traffic is computed N times and occupies HBM N times. Cache-aware routing fixes that by hashing the stable prefix, or the tenant, to a preferred replica. Pure affinity creates hot spots, though, so production routers bound the load: prefer the replica that owns the prefix unless it is much busier than average, then fall back to the next choice.
import hashlib
def score(key, replica):
return int(hashlib.blake2b(f"{key}|{replica}".encode(), digest_size=8).hexdigest(), 16)
def route(prefix_key, replicas, inflight, slack=1.25):
# rendezvous hashing gives every key a stable preference order over replicas
avg = sum(inflight.values()) / len(replicas)
for r in sorted(replicas, key=lambda r: score(prefix_key, r), reverse=True):
if inflight[r] + 1 <= slack * avg + 1: # bounded load: skip overloaded owners
return r
return min(replicas, key=inflight.get)
# prefix_key: hash of the first K tokens of the rendered prompt, or the tenant idRendezvous hashing has a useful property here: adding or removing a replica moves only the keys that ranked it first, so a scale-out event does not cold-start every cache. Choose the routing key carefully. Hashing the first few thousand tokens captures shared system prompts across tenants. Hashing the conversation ID keeps multi-turn chats on one replica, where each turn extends the previous turn's cached prefix. Many systems combine both. Prefix caching at scale discusses fleet-level placement further.
L3: provider prompt caching
If you call a hosted model, layers L0 to L2 belong to the provider, and what you control is prompt layout and, on some APIs, explicit breakpoints. Mechanics differ by provider and change often, so treat the following as a snapshot and check current documentation before relying on any number.
On Anthropic's API you mark a breakpoint by adding "cache_control": {"type": "ephemeral"} to a content block. The cached prefix is assembled in a fixed order, tools, then system, then messages, and a change at one level invalidates that level and everything after it. The default lifetime is 5 minutes, refreshed on each hit, and a 1-hour lifetime can be requested with "ttl": "1h". Cache writes are billed at 1.25 times the base input price for the 5-minute lifetime and 2 times for one hour. Reads are typically 0.1 times base, lower on some newer models. A request can carry up to four breakpoints. On a read, the system looks back up to 20 blocks from each breakpoint for an earlier cache entry, which is why a growing conversation keeps hitting if each turn adds fewer than 20 blocks. Prompts below a model-specific minimum length are simply not cached, with no error.
{
"model": "<your-model-id>",
"max_tokens": 1024,
"tools": [ ...stable tool definitions... ],
"system": [
{"type": "text", "text": "<6,000 tokens of policy and product docs>",
"cache_control": {"type": "ephemeral"}}
],
"messages": [ ...conversation so far..., {"role": "user", "content": "<new turn>"} ]
}
// response usage: cache_creation_input_tokens, cache_read_input_tokens, input_tokensOpenAI caches eligible prompts automatically, with no markup, and reports reuse in usage.input_tokens_details.cached_tokens. An optional prompt_cache_key groups requests that share a prefix. Minimum lengths, retention and discounts vary by model. On both platforms the rule is the same as in your own engine: stable content first, volatile content last, and log the cache counters on every response.
How hit rates compose
Layers multiply rather than add, and an upper layer changes the traffic a lower one sees. Take a support assistant with a 6,000-token stable prefix and a 500-token per-request suffix, on a hosted API with a 5-minute cache. Let h4 be the gateway hit rate and h3 the prefix hit rate among requests that reach the API. Input cost relative to no caching is:
def relative_input_cost(h4, h3, prefix=6000, suffix=500, write=1.25, read=0.10):
per_call = prefix * (h3 * read + (1 - h3) * write) + suffix
return (1 - h4) * per_call / (prefix + suffix)
print(round(relative_input_cost(h4=0.0, h3=0.0), 2)) # 1.23 - every call writes, nothing reads
print(round(relative_input_cost(h4=0.0, h3=0.9), 2)) # 0.28
print(round(relative_input_cost(h4=0.1, h3=0.9), 2)) # 0.25Two things stand out. A prompt cache that never hits is worse than none, because writes cost 25 percent more than plain input; caching rarely repeated prefixes loses money. And a 10 percent gateway hit rate removes only about 3 points of cost, because the calls it absorbs had already been discounted by the prompt cache. There is a subtler effect too. Prompt-cache lifetimes are refreshed by traffic, so a low-volume tenant whose requests arrive every four minutes keeps its prefix warm. Put a response cache in front that absorbs a third of those requests, and the gaps stretch past five minutes; h3 collapses. When you add an upper layer, re-measure the layers below it.
Tenant isolation and timing side channels
Shared caches are a side channel. A cached prefix returns its first token faster, so an attacker who can send prompts to the same cache can time responses to test guesses about another tenant's prompt. Researchers auditing hosted APIs have reported cache sharing across users at some providers. Treat cache scope as a security decision.
- Salt per trust boundary. vLLM accepts a per-request
cache_saltthat is mixed into the first block's hash, so only requests with the same salt can share blocks. Set it from the authenticated tenant, never from client input. - Share only public prefixes across tenants. A common system prompt and tool catalogue are safe to share; retrieved documents and conversation history are not.
- Scope response caches the same way. The gateway key must include tenant, model version and every permission that shaped the answer.
- Expire on deletion. When a user deletes data, their salted prefixes must age out or be purged with it.
Operating the layers
Instrument each layer with its own hit rate and its own cost, and alert on changes rather than absolute values.
- Engine: prefix-cache hit rate by tokens, not requests; HBM blocks in use; evictions per second; offload restores and their latency.
- Router: fraction of requests sent to their first-choice replica; per-replica in-flight spread.
- Provider: the sum of cache-read tokens over total input tokens, per prompt template; cache-write tokens, which are pure overhead if reads do not follow.
- Gateway: hits, staleness complaints, and hit rate per tenant.
The most common regression is invisible in latency dashboards: someone adds a request ID or the current date to a system prompt and the provider hit rate drops from 90 percent to zero. Add a CI check that renders each template twice with different inputs and asserts the leading N tokens are identical. Engine internals are covered in vLLM continuous batching and KV cache.
Failure modes
- Volatile prefix. Timestamps, user names or shuffled tool lists near the start of the prompt. Hit rate near zero, write charges on every call.
- Affinity hot spot. One giant tenant pinned to one replica without a load bound; its queue grows while others idle.
- Scale-out cold start. Modulo hashing remaps most keys when the replica count changes; every cache goes cold at once, exactly during a traffic peak.
- Tokeniser or template drift. A library upgrade changes how the chat template renders whitespace; prefixes stop matching across the fleet.
- Cross-tenant timing leak. Unsalted shared cache on a multi-tenant endpoint.
- Offload thrash. An SSD tier sized below the working set; restores queue behind writes and are slower than recompute.
Trade-offs
| Decision | Gain | Cost |
|---|---|---|
| Offload to DRAM or SSD | Long prefixes survive eviction | Copy latency; SSD can lose to recompute |
| Prefix affinity routing | L1 hits across the fleet | Hot spots without a load bound |
| Cross-tenant sharing | Higher hit rate on shared prompts | Timing side channel |
| 1-hour TTL instead of 5 minutes | Survives quiet periods | Writes cost 2x instead of 1.25x |
| Gateway response cache | Skips whole calls | Can starve the prompt cache of warming traffic |
What to do next
- Log cache counters from every provider response and every engine, per prompt template, starting today.
- Reorder prompts: tools and system first, then retrieved context, then history, then the new turn. Remove timestamps from the prefix.
- Add a CI test that asserts each template's leading tokens are stable across inputs.
- For self-hosted fleets, replace round-robin with bounded-load rendezvous routing on a prefix or conversation key.
- Compute KV bytes per token for your model and decide whether a DRAM offload tier pays at your prefix lengths.
- Set a cache salt per tenant and review which prefixes may be shared.
- Only after the lower layers are healthy, add a gateway response cache, then re-measure prompt-cache hit rates. The prompt caching guide covers the provider layer in more depth.