Teams usually meet LLM caching through one mechanism: a provider's prompt caching, or a semantic cache bolted onto a chatbot. Each mechanism has its own deep dive on this site. This article is about the strategy that sits above them: which layers to use, what goes into each key, when reusing an answer is actually correct, how to invalidate, and how to tell whether a cache pays for itself.
The core idea is that an LLM application is a pipeline. Embedding, retrieval, tool calls, prompt assembly and generation each have different costs, different staleness tolerances and different correctness risks. A good caching strategy caches each stage on its own terms rather than putting one big cache in front of the model.
Six layers worth caching
There are six places worth caching, and they fall into two groups. The first group avoids or shortens a model call. The second makes the pipeline that feeds the model cheaper.
| Layer | Key | What a hit saves | Main risk |
|---|---|---|---|
| Exact response | Hash of the full canonical request | The whole call | Serving a stale or wrongly shared answer |
| Semantic response | Embedding nearest neighbour above a threshold | The whole call | Answering a different question |
| Prompt prefix / KV | Token prefix, handled by provider or server | Prefill compute and input cost | Little; only latency if layout is wrong |
| Embedding | Hash of model id plus exact text | Embedding call | Mixing vectors from two embedding models |
| Retrieval | Normalised query plus index generation | Vector search and rerank | Stale documents after reindexing |
| Tool result | Tool name, version, canonical arguments | External API latency and quota | Stale facts, side effects |
Prefix caching is the one layer that is always safe, because it reuses computation rather than answers: the model still generates fresh output. It is covered in LLM prompt caching and, for self-hosted serving, in prefix caching at scale. The strategic point is layout: put stable content (system prompt, tool definitions, long reference documents) first and variable content last, or every request has a different prefix and nothing is reused. Providers differ on whether caching is automatic or needs explicit markers, on minimum lengths, lifetimes and discounts, so read your provider's current documentation rather than assuming numbers.
Designing the cache key
Most caching bugs in LLM systems are key bugs. A response depends on far more than the user's text, and anything that changes the output but is missing from the key causes a wrong hit. The key must include the model identifier and version, every sampling parameter, the full system prompt, the tool definitions, any retrieved context and the scope the answer may be shared within.
import hashlib, json, unicodedata
def canonical_text(s: str) -> str:
s = unicodedata.normalize("NFC", s)
return " ".join(s.split()) # collapse whitespace only; never lowercase code
def response_key(req: dict, versions: dict, scope: str) -> str:
material = {
"model": req["model"], # exact id, never an alias like "latest"
"params": {k: req.get(k) for k in ("temperature", "top_p", "max_tokens", "stop")},
"system": canonical_text(req["system"]),
"tools": hashlib.sha256(json.dumps(req.get("tools", []), sort_keys=True).encode()).hexdigest(),
"messages": [(m["role"], canonical_text(m["content"])) for m in req["messages"]],
"context_ids": sorted(req.get("context_doc_ids", [])),
"versions": versions, # prompt template, index generation, policy
"scope": scope, # "global", "tenant:42" or "user:9001"
}
blob = json.dumps(material, sort_keys=True, separators=(",", ":"))
return "resp:v3:" + hashlib.sha256(blob.encode()).hexdigest()Three rules follow from the code. Normalise only what cannot change meaning: whitespace and Unicode form are safe, while case is not, because code, identifiers and names are case-sensitive. Never key on a model alias that a provider can repoint. And make scope explicit: an answer computed with one tenant's documents in context must not be visible to another tenant, even if the question text matches exactly.
When reusing an answer is correct
A cache hit returns the same output twice. Whether that is acceptable depends on what the output is for, not on how it was generated.
| Request type | Reuse? | Why |
|---|---|---|
| Classification, extraction, routing at temperature 0 | Yes, long-lived | You want the same answer; caching also makes behaviour reproducible |
| FAQ-style answers over a fixed corpus | Yes, versioned by corpus | Correct until the documents change |
| Creative generation at high temperature | Usually no | Users regenerate to get a different answer; a hit defeats them |
| Answers that depend on time, prices or stock | Short TTL or no | Correctness decays with the world, not the prompt |
| Personalised answers using user data | Only within user scope | Sharing leaks data across users |
| Agent steps that call tools with side effects | Never cache the action | Replaying a send or a payment is a bug |
Temperature 0 does not guarantee identical outputs from every provider or server, because batching and floating-point order can change results. That argues for caching, not against it: when you need a stable answer, the cache is what makes it stable. Record which cached output you served, so a user report can be traced to it.
Where a semantic cache belongs
A semantic cache embeds the question and returns a stored answer when a previous question is close enough. It lifts hit rates on paraphrased traffic, and it is the riskiest layer, because similarity in embedding space is not equivalence of meaning: "How do I cancel my order?" and "How do I cancel my subscription?" can sit very close together. The threshold, the index and evaluation are covered in the semantic cache deep dive.
Strategically, put it behind the exact cache, restrict it to request types from the reuse table that tolerate approximation, include the same scope and versions in its partition key, and measure wrong-hit rate on labelled pairs before switching it on. A semantic cache that is right 97 percent of the time is wrong on three answers in every hundred, and users remember those.
Caching the pipeline: embeddings, retrieval, tools
The pipeline layers are where caching is safest and often most valuable, because their outputs are facts about inputs rather than generated prose.
- Embeddings are deterministic for a given model and text. Key on the embedding model id and a hash of the exact text, and cache indefinitely. When you change embedding models, the model id in the key separates the two spaces; never compare vectors across them.
- Retrieval results map a query to document ids and scores. Key on the normalised query, filters and the index generation number. When you reindex, bump the generation; old entries become unreachable without a scan.
- Tool calls need per-tool policy. A currency rate can be cached for minutes; a weather lookup for an hour; a database read only if the data is versioned. Cache reads only; tools with side effects must be excluded by an allowlist, not a denylist.
Caching document ids rather than document text keeps retrieval entries small and lets a document update show through without invalidating every query that retrieved it.
Invalidation by versioning
Deleting entries is the hard way to invalidate. The easy way is to put versions in keys, as the key function does, and change the version. A small registry holds the current prompt-template version, the index generation and a policy version, and every key builder reads it. When a prompt changes, its version changes, new requests compute new keys, and old entries expire on their own TTL. Nothing is ever served from an old version, and no scan of the cache is needed.
Two cases still need explicit deletion. Content that must disappear for legal or safety reasons should be removable by tag, so store a reverse index from document id to cache keys for retrieval and response layers. And a model output later judged harmful or wrong should be purgeable by key, which is another reason to log the key with every served response.
Stampedes and request coalescing
Popular prompts create a stampede: when an entry expires, every concurrent request misses and calls the model at once, multiplying cost at the worst moment. The fix is request coalescing, also called single-flight: the first miss computes, later identical misses wait for its result.
import asyncio
class SingleFlight:
def __init__(self):
self._inflight: dict[str, asyncio.Future] = {}
async def get(self, key, cache, compute, ttl):
hit = await cache.get(key)
if hit is not None:
return hit
fut = self._inflight.get(key)
if fut: # someone is already computing it
return await asyncio.shield(fut)
fut = asyncio.get_running_loop().create_future()
self._inflight[key] = fut
try:
value = await compute()
await cache.set(key, value, ttl=ttl)
fut.set_result(value)
return value
except Exception as e:
fut.set_exception(e) # waiters fail too; do not cache errors
raise
finally:
del self._inflight[key]This version coalesces within one process. Across a fleet, use a short-lived lock in the shared cache, or refresh popular entries before they expire. Add jitter to TTLs so entries written together do not expire together, and never cache errors, refusals caused by transient failures, or truncated outputs.
Does the cache pay for itself?
A cache is worth running when expected savings exceed its own costs. For a response layer, savings per request equal hit rate times the cost of a model call. Costs are the lookup on every request, storage, the embedding call for semantic lookups, and the price of wrong hits.
Work an example with round, illustrative numbers. A support assistant serves one million requests a day, and a model call costs 0.4 cents on average. Logs show 18 percent of requests are exact repeats after normalisation within the scope rules. The exact cache saves 0.18 x 1,000,000 x 0.4 cents, which is 720 dollars a day, for a key-value store costing a small fraction of that. A semantic layer behind it adds 9 percent more hits but needs an embedding for every remaining miss and has a measured 1.5 percent wrong-hit rate. It saves about 360 dollars a day and produces roughly 1,350 wrong answers a day. Whether that is acceptable is a product decision, and putting the number in front of the product owner is the point of the exercise.
Latency matters as much as cost: a hit returns in milliseconds instead of seconds, which changes how an interface feels. Measure hit rate per request type, not just overall, because a high global rate can hide a layer that never hits.
Failure modes
- Cross-tenant leakage: a missing scope in the key serves one customer's answer to another. Test it with two tenants and identical prompts.
- Alias drift: keys built on a moving model alias mix outputs from two models.
- Prefix churn: a timestamp or request id at the top of the system prompt makes every prefix unique and silently disables prefix caching.
- Cached failures: an empty or truncated output stored with a long TTL turns a transient error into a persistent one.
- Unbounded growth: high-cardinality prompts fill the store with entries that never hit; cap memory and evict least recently used entries.
What to do next
- Log a canonical key for every request for a week, without caching, and measure the repeat rate per request type.
- Reorder prompts so stable content comes first, and confirm in provider or server metrics that cached input tokens rise.
- Add embedding and retrieval caches keyed on model id and index generation; these are low-risk wins.
- Write the reuse policy per request type, then enable the exact response cache only for types marked reusable.
- Build the version registry and include scope and versions in every key before going live.
- Add single-flight and TTL jitter, and refuse to cache errors and truncated outputs.
- Pilot a semantic layer only after measuring its wrong-hit rate on labelled pairs, and report savings and wrong hits together.