A cache entry is a copy of a fact taken at some moment. Invalidation is everything you do to stop serving that copy after the fact changes. It is hard not because deleting a key is hard but because the change happens in one place, the copies live in many, the messages between them can be delayed, reordered or lost, and every reader can race every writer. Most stale-data incidents are not a missing delete; they are a delete that arrived before the database replica caught up, or a pub/sub message nobody was subscribed to receive.
This article treats invalidation as a delivery system with a staleness budget. It sets out the strategies and when each fits, builds a change-data-capture invalidation pipeline with code, covers generation counters and tag-based purges for groups of keys, works a product-catalog example across database, Redis and CDN, and lists the failure modes and the metrics that catch them. The basics of delete-versus-update and the fill race are summarised here and covered at survey depth in caching in system design.
Start with a staleness budget
Start by deciding how stale each kind of data may be, because the answer chooses the strategy. A product description can be a minute old. A price shown in a cart should be seconds old and must be rechecked at checkout. An account's permissions after revocation must be stale for as close to zero as you can manage, and the authorisation check probably should not be cached at all. Write these as a staleness budget per data class: the maximum time between a committed change and the moment no reader can see the old value.
Then answer three questions for each class. Which keys depend on this row? The direct key is easy; list pages, search results and aggregates that embed the row are where invalidation leaks. Who tells the cache: the writing application, a database log tailer, or nobody, leaving it to expiry? When: synchronously before the write returns, asynchronously within a bound, or at expiry? Every strategy below is a set of answers to these three questions.
The strategies and what each guarantees
| Strategy | Who tells the cache | Staleness bound | Main risk |
|---|---|---|---|
| TTL only | Nobody; entries expire | The TTL | Stale for the whole TTL; short TTLs cost hit rate |
| Delete after commit (app) | The writing service | Milliseconds when it works | Missed by any writer that forgets, crashes or bypasses the service |
| CDC-driven delete | A log tailer on the database | Pipeline lag, usually under a second | Lag spikes; row-to-key mapping drift |
| Generation counter | Writer bumps a version | One version read | Extra lookup per read; orphaned entries use memory |
| Tag or surrogate-key purge | Writer or CDC purges a tag | Purge propagation time | Purge rate limits; over-broad tags empty the cache |
| Write-through update | Writer sets the new value | Immediate on that path | Concurrent writers can leave the older value last |
Two rules apply to all of them. Delete rather than update, because two concurrent updaters can write the cache in the opposite order from the database, while two deletes cannot disagree. And keep a TTL on every entry even when you have active invalidation, because the TTL is the bound on how long any lost message, bug or forgotten writer can keep a value wrong.
Invalidating from the database log
Application-driven deletes fail in a predictable way: someone writes to the database without going through the service that knows the cache keys. A migration script, an admin tool, a second service, a manual fix at 2 a.m. The cure is to invalidate from the database's own change log, which sees every committed write regardless of who made it. Facebook's 2013 paper on scaling memcache describes exactly this: daemons named mcsqueal tail the MySQL commit log, extract the keys to delete, and broadcast deletes to the cache fleet. Today the same shape is commonly built with a CDC connector such as Debezium reading the PostgreSQL write-ahead log or MySQL binlog into Kafka, and a small invalidator consuming it.
Because the log contains only committed changes, deletes never race ahead of a rolled-back transaction, the problem that makes deleting before commit unsafe. Partitioning the topic by primary key keeps changes to one row in order. The invalidator must be idempotent, since deleting twice is harmless, and must commit its consumer offset only after deletes succeed, so a crash replays rather than drops:
def keys_for(table: str, row: dict) -> list[str]:
# The ONE shared definition of which cache keys a row feeds.
# The read path imports the same module to build keys.
if table == "products":
return [f"product:{row['id']}", f"product_card:{row['id']}",
f"tag:category:{row['category_id']}"]
if table == "prices":
return [f"product:{row['product_id']}", f"price:{row['product_id']}"]
return []
for batch in consumer.poll_batches(): # Kafka, partitioned by primary key
keys = set()
for ev in batch:
for img in (ev.before, ev.after): # old AND new image: category may change
if img:
keys.update(keys_for(ev.table, img))
if keys: # DEL with no keys is an error
redis.delete(*keys) # idempotent; retried on error
consumer.commit(batch) # only after the deletes succeed
metrics.observe("invalidation_lag_s", now() - batch.max_commit_ts)Two details in that code are where pipelines usually go wrong. Using both the before and after images matters when a change moves a row between groups, such as a product changing category: both categories' lists are now stale. And keeping the key derivation in one module shared by readers and the invalidator stops the classic drift where someone adds a new cache key on the read path and nothing ever deletes it. The same log can feed other consumers; the outbox variant of this idea is in the outbox pattern.
Invalidating groups: generation counters and tags
Some changes invalidate a group whose members you cannot cheaply list: every page of a user's feed, every cached query touching a tenant's data, every rendered page using a template. Two techniques avoid enumerating keys.
Generation counters. Store a version per group, and build member keys from it: tenant:42:v17:report:monthly. To invalidate the whole group, increment the version. New reads build keys with v18, miss, and refill; the v17 entries are never read again and age out by TTL or eviction. Invalidation becomes a single atomic increment. The cost is an extra read of the version on every request, usually removed by caching the version in process for a second or two, which then becomes part of your staleness bound, and memory spent on orphaned entries until they expire.
Tags and surrogate keys. HTTP caches and CDNs support attaching labels to a cached response and purging everything with a label. Fastly calls the header Surrogate-Key, Cloudflare Cache-Tag, and Varnish offers bans and the xkey module. The origin labels a product page with product-981 category-12, and a price change purges product-981 without knowing which URLs, languages or query strings exist. Check your provider's limits on purge rate and on tags per response, and how long a purge takes to propagate; those numbers, not your code, set the staleness bound at the edge. CDN layering itself is covered in CDN design.
Worked example: a price change across regions
An online store keeps products and prices in PostgreSQL with a read replica per region, caches product objects in Redis per region, and serves product pages through a CDN. The staleness budget: prices within 2 seconds in Redis and 60 seconds at the edge, descriptions within 5 minutes, and checkout always reads the price from the primary.
A merchant changes product 981's price at 10:00:00.000. The transaction commits on the primary. Debezium emits the change at about 10:00:00.150 and the invalidator in region A deletes product:981 and price:981 at 10:00:00.200. The next region A reader misses, reads the primary or an up-to-date replica, and caches the new price with a 10-minute TTL. The invalidator also purges CDN tag product-981; pages are cached at the edge with s-maxage=60 so even a lost purge is bounded at 60 seconds.
Region B is where it gets subtle. Its invalidator receives the same event at 10:00:00.250, but region B's replica is 400 milliseconds behind. A reader that misses at 10:00:00.300 reads the replica, gets the old price, and caches it for 10 minutes. The delete happened; it just happened too early. The fix is to make region B's invalidator wait until the local replica has replayed past the change's log position, checking the replica's replay position against the event's, before deleting. Facebook's paper describes a related device, remote markers, which send reads for recently-written keys to the primary region until the replica catches up. Either way, invalidation in a replicated system must be ordered after the data the refill will read, not after the primary commit.
Failure modes
These are the incidents that recur, with the signal that catches each.
- Lost invalidations. Redis pub/sub delivers only to subscribers connected at that moment, so a restarting invalidator misses messages for good. Use a durable log with committed offsets, and keep TTLs as the backstop. Signal: staleness probes failing while lag looks normal.
- Refill from a lagging replica. As in the example, a delete races ahead of replication and the cache is refilled with old data. Signal: stale reads clustered in non-primary regions.
- The fill race. A reader loads the old value, a writer commits and deletes, then the reader sets the old value. Leases, where a miss hands out a token that a later delete invalidates so the stale set is rejected, close it; memcache leases were described in the same Facebook paper.
- Invalidation storms. A bulk update or backfill touches 5 million rows and the invalidator deletes 5 million keys in seconds, turning into a stampede on the database. Rate-limit the invalidator, collapse duplicate keys per batch, and run bulk jobs off-peak with request coalescing on the read path.
- Key drift. A new read path caches a key the invalidator does not know. Signal: entries that only ever expire, never get deleted. The shared key module prevents it.
- Over-broad tags. Tagging every page with
sitemeans one purge empties the CDN. Tag by entity, not by section.
Measuring and operating invalidation
Measure staleness directly rather than inferring it. Run a canary that writes a timestamped value to a dedicated row every few seconds and reads it back through each cache layer in each region; the difference is your real staleness, and its 99th percentile is the number to hold against the budget. Track invalidation lag from commit timestamp to delete, deletes per second, consumer lag in the log, the ratio of keys deleted to keys expired, and the hit rate around bulk jobs. Give the invalidator a kill switch that falls back to short TTLs, and a replay tool that re-deletes keys for a time window after an incident. The same delivery-and-ordering reasoning underlies read models in CQRS with event sourcing.
What to do next
- Write a staleness budget per data class, including which reads must bypass the cache entirely.
- Put cache key derivation in one module used by both readers and the invalidator.
- Move invalidation from application code to a CDC consumer on the database log, partitioned by primary key, committing offsets only after deletes succeed.
- Delete keys for both the before and after images of each change.
- In replicated regions, delay each region's deletes until the local replica has replayed past the change.
- Keep a TTL on every entry, use generation counters or tags for groups, and rate-limit bulk invalidation.
- Run a staleness canary per layer and region and alert on its 99th percentile against the budget.