The cheapest LLM call is the one you never make. A response cache sits in your application in front of the model API and returns a stored answer when an identical request comes in again, trading a few milliseconds of cache lookup for seconds of generation and the whole token bill. On workloads with repetition — support FAQs, autocomplete, classification, templated extraction — this is the single highest-leverage cost and latency optimisation available, and it needs no GPU at all. It is also the cache most likely to serve one user's answer to another if you get the key wrong, so it rewards being built carefully.
This is the application-side, exact-match cache, and it is a different layer from its neighbours. Server-side prefix caching reuses computed attention state inside the serving engine; semantic caching reuses answers for questions that mean the same thing; provider prompt caching discounts repeated prompt prefixes at the API. This article is about the plain response cache you own: byte-identical request in, stored response out.
Where this cache sits
Place it in your AI gateway or service layer, where you already see the full request and can enforce tenancy and policy. The flow is the familiar read-through cache: build a key from the request, look it up, return on a hit, otherwise call the API, store the response under the key, and return it. Everything hard about this cache is in two places — what goes into the key, and when a stored answer stops being correct.
The cache key: exact match done right
A response is only reusable if every input that could change it is in the key. Miss one and you serve a wrong answer; include something irrelevant and your hit rate collapses. For a chat-style request the key material is the model id, the full message list, the system prompt, any tool or function definitions, and every sampling parameter that affects output — temperature, top-p, max tokens, stop sequences, response format. Crucially it must also include the tenant or user scope, because two users must never share a cache entry even when their requests are byte-identical, and a cache version so you can invalidate everything at once.
import hashlib, json
CACHE_VERSION = "v3" # bump to invalidate the whole namespace
def cache_key(req: dict) -> str:
material = {
"cv": CACHE_VERSION,
"model": req["model"], # pin the exact model id
"messages": req["messages"],
"system": req.get("system", ""),
"tools": req.get("tools", []), # tool defs change behaviour
"temperature": req["temperature"],
"top_p": req.get("top_p"),
"max_tokens": req["max_tokens"],
"stop": req.get("stop"),
"response_format": req.get("response_format"),
"tenant": req["tenant"], # NEVER share across tenants
}
blob = json.dumps(material, sort_keys=True, separators=(",", ":"))
return "llmresp:" + hashlib.sha256(blob.encode("utf-8")).hexdigest()Hash the canonical JSON, not the raw request string, so that key order and whitespace do not fragment your cache. The sort_keys=True and compact separators are what make two logically identical requests collapse to one key.
Normalisation: raise the hit rate without lying
Exact-match caching is brittle unless you first normalise the request into a canonical form — but normalise only what is genuinely semantics-preserving, or you will serve answers to requests that were not actually the same. Safe normalisations include trimming trailing whitespace, collapsing runs of spaces, and lower-casing a classification label that you know the model treats case-insensitively. Unsafe “normalisations” include stripping punctuation from a free-text prompt or rounding a temperature, because those change what the user asked. When in doubt, leave it in the key: a lower hit rate is a cost problem, a wrong hit is a correctness problem.
Normalisation is also where exact-match caching ends and semantic caching begins. The moment you find yourself wanting “what time do you open?” and “when are you open?” to share an answer, you have left exact matching behind — that is a similarity judgement, and forcing it into a hash key by aggressive normalisation will eventually collapse two genuinely different questions onto one entry. Keep the exact cache exact and layer meaning-based reuse separately if you need it; the two caches have different correctness guarantees and should not be blended into one key function.
The determinism problem
Here is the subtlety that catches people. An LLM call at temperature > 0 is a sample from a distribution, not a function. Cache it and you freeze one sample forever: every future user with that request gets the identical answer, which may be fine for a classifier but feels broken for a creative assistant. So decide deliberately. Cache aggressively when the output should be deterministic — temperature = 0, or a fixed seed if your provider supports reproducible sampling — and be cautious about caching sampled output unless frozen reuse is actually the product behaviour you want.
The honest rule: only cache what you are willing to serve identically every time. If you want variety, either do not cache that route, or cache a small set of pre-sampled responses and rotate. Record the sampling settings in the key regardless, so a change in temperature is a different cache entry rather than a silently stale one.
Caching streaming responses
Most LLM endpoints stream tokens as server-sent events, and the cache has to cope. The rule is simple: only cache a completed stream. Accumulate the tokens as they arrive, and write the full concatenated response to the cache only after the stream ends successfully. Never cache a partial stream, because a client that disconnected mid-generation would otherwise poison the entry with a truncated answer.
async def stream_with_cache(cache, key, call_api, ttl=3600):
cached = cache.get(key)
if cached is not None:
for chunk in chunk_text(cached): # replay as a stream
yield chunk
return
buf = []
async for chunk in call_api(): # live generation
buf.append(chunk)
yield chunk
cache.set(key, "".join(buf), ttl=ttl) # store only on clean finishOn a hit you replay the stored text as a stream so the client code path is identical whether the answer was cached or generated. You can replay instantly or pace the chunks to mimic live generation; instant replay is usually what users want once the cache is warm. One more subtlety: decide what to do with the token-usage metadata the API returns on a miss. If a downstream system bills or rate-limits on reported tokens, a cached hit that reports zero usage will quietly understate load, so store the original usage figures alongside the response and replay them too, clearly flagged as cached, rather than letting a hit look like free work that never happened.
TTL and invalidation
A cached response is a snapshot that can go stale in three ways, and each has its own control:
- The model changes. A new model version answers differently. Pin the exact model id in the key so a model upgrade naturally misses the old entries.
- The prompt template changes. If you edit the system prompt or tool definitions, old answers no longer reflect current behaviour. The system prompt and tools are already in the key, and
CACHE_VERSIONis the blunt instrument when you want to drop everything at once. - The underlying facts change. An answer about today's prices or inventory is wrong tomorrow. Set a TTL short enough for the content's volatility, and keep time-sensitive data (dates, “now”, live figures) out of cacheable prompts entirely, or those entries will confidently serve yesterday's world.
TTL is your safety net, not your plan. A long TTL maximises hit rate and risk; a short one is safer and cheaper to be wrong about. Choose per route, not globally: a static policy FAQ can live for a day, a stock-level answer for seconds or not at all.
Stampede control: one miss, not a thousand
When a popular entry expires, or a brand-new popular request arrives, every concurrent caller misses at once and they all stampede the API — the thundering-herd problem. Without protection a single cold key can trigger hundreds of identical, expensive calls. The fix is single-flight: let exactly one caller compute the value while the rest wait for it and then read the freshly cached result.
def get_or_compute(cache, lock, key, compute, ttl=3600, lock_ttl=30):
hit = cache.get(key)
if hit is not None:
return hit, "hit"
if lock.acquire(key, ttl=lock_ttl): # I am the one caller
try:
val = compute() # the single API call
cache.set(key, val, ttl=ttl)
return val, "miss"
finally:
lock.release(key)
val = cache.wait_and_get(key, timeout=lock_ttl) # others wait, then read
return (val, "coalesced") if val is not None else (compute(), "fallback")A distributed lock (for example a Redis SET key val NX PX) elects the single computing caller; the losers block briefly and then read the value the winner stored. The lock_ttl must exceed a normal generation time so the lock does not expire mid-call, and the fallback path guarantees a request still completes if the winner crashes.
The economics: when it pays
The arithmetic is simple and almost always favourable. If a fraction h of requests hit the cache, you remove h of the API token cost and replace seconds of latency with milliseconds on those requests. A support assistant where 40% of questions are repeats sheds roughly 40% of its generation cost and serves those answers effectively instantly. Because a cache read costs a tiny fraction of a cent and a generated response costs orders of magnitude more — see the per-request figures in LLM cost analysis — the break-even hit rate is low enough that any genuine repetition pays. The real cost of this cache is never the infrastructure; it is the correctness and privacy risk of a bad entry, which is why the failure modes below matter more than the savings.
Failure modes and trade-offs
| Failure | Cause | Fix |
|---|---|---|
| Cross-tenant leak | Tenant/user scope missing from the key | Put tenant in the key; test with two users and one prompt |
| Stale after a prompt or model change | Template or model id not reflected in the key | Pin model id and system prompt; bump CACHE_VERSION |
| Product feels “stuck” | Caching sampled output at temperature > 0 | Cache only deterministic routes, or rotate pre-sampled answers |
| Wrong answer about current facts | Volatile data baked into a cached prompt | Keep dates/live figures out of cacheable prompts; short TTL |
| Thundering herd on cold keys | No single-flight on miss | Serialise misses with a per-key lock |
| Low hit rate, high memory | Over-specific keys or unbounded cardinality | Normalise safely; cap and evict with an LRU bound |
| Sensitive data at rest | Caching responses that contain PII | Encrypt, scope retention, and exclude PII-bearing routes |
The through-line is that a response cache is a correctness-sensitive component wearing a performance component's clothes. Build the key defensively, invalidate deliberately, and treat a wrong hit as a more serious bug than a miss.
What to do next
- List your routes and mark which ones are deterministic and repetitive enough to cache at all — do not cache everything by default.
- Build the cache key from the full request including model id, system prompt, tools, sampling params and tenant scope, and hash the canonical JSON.
- Set
temperature = 0(or a fixed seed) on cached routes, and leave volatile data out of cacheable prompts. - Accumulate streaming responses and store only on a clean finish; replay cached answers as a stream.
- Add single-flight with a per-key distributed lock so a cold popular key triggers one API call, not a herd.
- Choose a per-route TTL, pin the model id, and keep a
CACHE_VERSIONyou can bump to flush everything after a prompt change. - Before launch, test cross-tenant isolation with two users and one identical prompt, and confirm no entry is shared.