Teams usually meet LLM caching through one mechanism: a provider's prompt caching, or a semantic cache bolted onto a chatbot. Each mechanism has its own deep dive on this site. This article is about the strategy that sits above them: which layers to use, what goes into each key, when reusing an answer is actually correct, how to invalidate, and how to tell whether a cache pays for itself.

The core idea is that an LLM application is a pipeline. Embedding, retrieval, tool calls, prompt assembly and generation each have different costs, different staleness tolerances and different correctness risks. A good caching strategy caches each stage on its own terms rather than putting one big cache in front of the model.

Advertisement

Six layers worth caching

There are six places worth caching, and they fall into two groups. The first group avoids or shortens a model call. The second makes the pipeline that feeds the model cheaper.

Requestuser, tenant, promptExact responsecanonical keySemantic cacheembedding, thresholdModel callprefix / KV reuseEmbedding cachetext hash to vectorRetrieval cachequery to doc idsTool cacheargs to result, TTLVersion registrymodel, prompt, index genmissmissEvery key includes the versions it depends on; bumping a version invalidates without deleting
Cache layers on an LLM request path. The top row avoids or shortens generation; the middle row speeds the pieces a pipeline assembles; the registry feeds versions into every key.
LayerKeyWhat a hit savesMain risk
Exact responseHash of the full canonical requestThe whole callServing a stale or wrongly shared answer
Semantic responseEmbedding nearest neighbour above a thresholdThe whole callAnswering a different question
Prompt prefix / KVToken prefix, handled by provider or serverPrefill compute and input costLittle; only latency if layout is wrong
EmbeddingHash of model id plus exact textEmbedding callMixing vectors from two embedding models
RetrievalNormalised query plus index generationVector search and rerankStale documents after reindexing
Tool resultTool name, version, canonical argumentsExternal API latency and quotaStale facts, side effects

Prefix caching is the one layer that is always safe, because it reuses computation rather than answers: the model still generates fresh output. It is covered in LLM prompt caching and, for self-hosted serving, in prefix caching at scale. The strategic point is layout: put stable content (system prompt, tool definitions, long reference documents) first and variable content last, or every request has a different prefix and nothing is reused. Providers differ on whether caching is automatic or needs explicit markers, on minimum lengths, lifetimes and discounts, so read your provider's current documentation rather than assuming numbers.

Designing the cache key

Most caching bugs in LLM systems are key bugs. A response depends on far more than the user's text, and anything that changes the output but is missing from the key causes a wrong hit. The key must include the model identifier and version, every sampling parameter, the full system prompt, the tool definitions, any retrieved context and the scope the answer may be shared within.

import hashlib, json, unicodedata

def canonical_text(s: str) -> str:
    s = unicodedata.normalize("NFC", s)
    return " ".join(s.split())          # collapse whitespace only; never lowercase code

def response_key(req: dict, versions: dict, scope: str) -> str:
    material = {
        "model": req["model"],                      # exact id, never an alias like "latest"
        "params": {k: req.get(k) for k in ("temperature", "top_p", "max_tokens", "stop")},
        "system": canonical_text(req["system"]),
        "tools": hashlib.sha256(json.dumps(req.get("tools", []), sort_keys=True).encode()).hexdigest(),
        "messages": [(m["role"], canonical_text(m["content"])) for m in req["messages"]],
        "context_ids": sorted(req.get("context_doc_ids", [])),
        "versions": versions,                       # prompt template, index generation, policy
        "scope": scope,                             # "global", "tenant:42" or "user:9001"
    }
    blob = json.dumps(material, sort_keys=True, separators=(",", ":"))
    return "resp:v3:" + hashlib.sha256(blob.encode()).hexdigest()

Three rules follow from the code. Normalise only what cannot change meaning: whitespace and Unicode form are safe, while case is not, because code, identifiers and names are case-sensitive. Never key on a model alias that a provider can repoint. And make scope explicit: an answer computed with one tenant's documents in context must not be visible to another tenant, even if the question text matches exactly.

Advertisement

When reusing an answer is correct

A cache hit returns the same output twice. Whether that is acceptable depends on what the output is for, not on how it was generated.

Request typeReuse?Why
Classification, extraction, routing at temperature 0Yes, long-livedYou want the same answer; caching also makes behaviour reproducible
FAQ-style answers over a fixed corpusYes, versioned by corpusCorrect until the documents change
Creative generation at high temperatureUsually noUsers regenerate to get a different answer; a hit defeats them
Answers that depend on time, prices or stockShort TTL or noCorrectness decays with the world, not the prompt
Personalised answers using user dataOnly within user scopeSharing leaks data across users
Agent steps that call tools with side effectsNever cache the actionReplaying a send or a payment is a bug

Temperature 0 does not guarantee identical outputs from every provider or server, because batching and floating-point order can change results. That argues for caching, not against it: when you need a stable answer, the cache is what makes it stable. Record which cached output you served, so a user report can be traced to it.

Where a semantic cache belongs

A semantic cache embeds the question and returns a stored answer when a previous question is close enough. It lifts hit rates on paraphrased traffic, and it is the riskiest layer, because similarity in embedding space is not equivalence of meaning: "How do I cancel my order?" and "How do I cancel my subscription?" can sit very close together. The threshold, the index and evaluation are covered in the semantic cache deep dive.

Strategically, put it behind the exact cache, restrict it to request types from the reuse table that tolerate approximation, include the same scope and versions in its partition key, and measure wrong-hit rate on labelled pairs before switching it on. A semantic cache that is right 97 percent of the time is wrong on three answers in every hundred, and users remember those.

Caching the pipeline: embeddings, retrieval, tools

The pipeline layers are where caching is safest and often most valuable, because their outputs are facts about inputs rather than generated prose.

  • Embeddings are deterministic for a given model and text. Key on the embedding model id and a hash of the exact text, and cache indefinitely. When you change embedding models, the model id in the key separates the two spaces; never compare vectors across them.
  • Retrieval results map a query to document ids and scores. Key on the normalised query, filters and the index generation number. When you reindex, bump the generation; old entries become unreachable without a scan.
  • Tool calls need per-tool policy. A currency rate can be cached for minutes; a weather lookup for an hour; a database read only if the data is versioned. Cache reads only; tools with side effects must be excluded by an allowlist, not a denylist.

Caching document ids rather than document text keeps retrieval entries small and lets a document update show through without invalidating every query that retrieved it.

Invalidation by versioning

Deleting entries is the hard way to invalidate. The easy way is to put versions in keys, as the key function does, and change the version. A small registry holds the current prompt-template version, the index generation and a policy version, and every key builder reads it. When a prompt changes, its version changes, new requests compute new keys, and old entries expire on their own TTL. Nothing is ever served from an old version, and no scan of the cache is needed.

Two cases still need explicit deletion. Content that must disappear for legal or safety reasons should be removable by tag, so store a reverse index from document id to cache keys for retrieval and response layers. And a model output later judged harmful or wrong should be purgeable by key, which is another reason to log the key with every served response.

Stampedes and request coalescing

Popular prompts create a stampede: when an entry expires, every concurrent request misses and calls the model at once, multiplying cost at the worst moment. The fix is request coalescing, also called single-flight: the first miss computes, later identical misses wait for its result.

import asyncio

class SingleFlight:
    def __init__(self):
        self._inflight: dict[str, asyncio.Future] = {}

    async def get(self, key, cache, compute, ttl):
        hit = await cache.get(key)
        if hit is not None:
            return hit
        fut = self._inflight.get(key)
        if fut:                                   # someone is already computing it
            return await asyncio.shield(fut)
        fut = asyncio.get_running_loop().create_future()
        self._inflight[key] = fut
        try:
            value = await compute()
            await cache.set(key, value, ttl=ttl)
            fut.set_result(value)
            return value
        except Exception as e:
            fut.set_exception(e)                  # waiters fail too; do not cache errors
            raise
        finally:
            del self._inflight[key]

This version coalesces within one process. Across a fleet, use a short-lived lock in the shared cache, or refresh popular entries before they expire. Add jitter to TTLs so entries written together do not expire together, and never cache errors, refusals caused by transient failures, or truncated outputs.

Does the cache pay for itself?

A cache is worth running when expected savings exceed its own costs. For a response layer, savings per request equal hit rate times the cost of a model call. Costs are the lookup on every request, storage, the embedding call for semantic lookups, and the price of wrong hits.

Work an example with round, illustrative numbers. A support assistant serves one million requests a day, and a model call costs 0.4 cents on average. Logs show 18 percent of requests are exact repeats after normalisation within the scope rules. The exact cache saves 0.18 x 1,000,000 x 0.4 cents, which is 720 dollars a day, for a key-value store costing a small fraction of that. A semantic layer behind it adds 9 percent more hits but needs an embedding for every remaining miss and has a measured 1.5 percent wrong-hit rate. It saves about 360 dollars a day and produces roughly 1,350 wrong answers a day. Whether that is acceptable is a product decision, and putting the number in front of the product owner is the point of the exercise.

Latency matters as much as cost: a hit returns in milliseconds instead of seconds, which changes how an interface feels. Measure hit rate per request type, not just overall, because a high global rate can hide a layer that never hits.

Failure modes

  • Cross-tenant leakage: a missing scope in the key serves one customer's answer to another. Test it with two tenants and identical prompts.
  • Alias drift: keys built on a moving model alias mix outputs from two models.
  • Prefix churn: a timestamp or request id at the top of the system prompt makes every prefix unique and silently disables prefix caching.
  • Cached failures: an empty or truncated output stored with a long TTL turns a transient error into a persistent one.
  • Unbounded growth: high-cardinality prompts fill the store with entries that never hit; cap memory and evict least recently used entries.

What to do next

  1. Log a canonical key for every request for a week, without caching, and measure the repeat rate per request type.
  2. Reorder prompts so stable content comes first, and confirm in provider or server metrics that cached input tokens rise.
  3. Add embedding and retrieval caches keyed on model id and index generation; these are low-risk wins.
  4. Write the reuse policy per request type, then enable the exact response cache only for types marked reusable.
  5. Build the version registry and include scope and versions in every key before going live.
  6. Add single-flight and TTL jitter, and refuse to cache errors and truncated outputs.
  7. Pilot a semantic layer only after measuring its wrong-hit rate on labelled pairs, and report savings and wrong hits together.
Key takeaway: Treat an LLM application as a pipeline and cache each stage on its own terms. Prefix caching is always safe if stable content comes first. Embedding, retrieval and read-only tool caches are low-risk wins keyed on model and index versions. Response caches need keys that capture model, parameters, tools, context and sharing scope, a per-request-type reuse policy, version-based invalidation and coalescing, and a cost model that counts wrong hits as well as savings.