A cache entry is a copy of a fact taken at some moment. Invalidation is everything you do to stop serving that copy after the fact changes. It is hard not because deleting a key is hard but because the change happens in one place, the copies live in many, the messages between them can be delayed, reordered or lost, and every reader can race every writer. Most stale-data incidents are not a missing delete; they are a delete that arrived before the database replica caught up, or a pub/sub message nobody was subscribed to receive.

This article treats invalidation as a delivery system with a staleness budget. It sets out the strategies and when each fits, builds a change-data-capture invalidation pipeline with code, covers generation counters and tag-based purges for groups of keys, works a product-catalog example across database, Redis and CDN, and lists the failure modes and the metrics that catch them. The basics of delete-versus-update and the fill race are summarised here and covered at survey depth in caching in system design.

Start with a staleness budget

Start by deciding how stale each kind of data may be, because the answer chooses the strategy. A product description can be a minute old. A price shown in a cart should be seconds old and must be rechecked at checkout. An account's permissions after revocation must be stale for as close to zero as you can manage, and the authorisation check probably should not be cached at all. Write these as a staleness budget per data class: the maximum time between a committed change and the moment no reader can see the old value.

Then answer three questions for each class. Which keys depend on this row? The direct key is easy; list pages, search results and aggregates that embed the row are where invalidation leaks. Who tells the cache: the writing application, a database log tailer, or nobody, leaving it to expiry? When: synchronously before the write returns, asynchronously within a bound, or at expiry? Every strategy below is a set of answers to these three questions.

The strategies and what each guarantees

StrategyWho tells the cacheStaleness boundMain risk
TTL onlyNobody; entries expireThe TTLStale for the whole TTL; short TTLs cost hit rate
Delete after commit (app)The writing serviceMilliseconds when it worksMissed by any writer that forgets, crashes or bypasses the service
CDC-driven deleteA log tailer on the databasePipeline lag, usually under a secondLag spikes; row-to-key mapping drift
Generation counterWriter bumps a versionOne version readExtra lookup per read; orphaned entries use memory
Tag or surrogate-key purgeWriter or CDC purges a tagPurge propagation timePurge rate limits; over-broad tags empty the cache
Write-through updateWriter sets the new valueImmediate on that pathConcurrent writers can leave the older value last

Two rules apply to all of them. Delete rather than update, because two concurrent updaters can write the cache in the opposite order from the database, while two deletes cannot disagree. And keep a TTL on every entry even when you have active invalidation, because the TTL is the bound on how long any lost message, bug or forgotten writer can keep a value wrong.

Invalidating from the database log

App writecommit to primaryPrimary DBWAL / binloglog tailCDC connectorrow change eventsLog (Kafka)partitioned by keyInvalidatorrow to keys, idempotentDEL keysRedis, region ARedis, region Bafter replica catches uppurge tagCDNsurrogate key / cache tagReadersmiss, read DB or replica, set with TTLTTL on every entry bounds what a lost message can cost.
A CDC pipeline: committed changes flow from the database log through Kafka to an invalidator that deletes Redis keys per region and purges CDN tags; TTLs bound any loss.

Application-driven deletes fail in a predictable way: someone writes to the database without going through the service that knows the cache keys. A migration script, an admin tool, a second service, a manual fix at 2 a.m. The cure is to invalidate from the database's own change log, which sees every committed write regardless of who made it. Facebook's 2013 paper on scaling memcache describes exactly this: daemons named mcsqueal tail the MySQL commit log, extract the keys to delete, and broadcast deletes to the cache fleet. Today the same shape is commonly built with a CDC connector such as Debezium reading the PostgreSQL write-ahead log or MySQL binlog into Kafka, and a small invalidator consuming it.

Because the log contains only committed changes, deletes never race ahead of a rolled-back transaction, the problem that makes deleting before commit unsafe. Partitioning the topic by primary key keeps changes to one row in order. The invalidator must be idempotent, since deleting twice is harmless, and must commit its consumer offset only after deletes succeed, so a crash replays rather than drops:

def keys_for(table: str, row: dict) -> list[str]:
    # The ONE shared definition of which cache keys a row feeds.
    # The read path imports the same module to build keys.
    if table == "products":
        return [f"product:{row['id']}", f"product_card:{row['id']}",
                f"tag:category:{row['category_id']}"]
    if table == "prices":
        return [f"product:{row['product_id']}", f"price:{row['product_id']}"]
    return []

for batch in consumer.poll_batches():          # Kafka, partitioned by primary key
    keys = set()
    for ev in batch:
        for img in (ev.before, ev.after):      # old AND new image: category may change
            if img:
                keys.update(keys_for(ev.table, img))
    if keys:                                   # DEL with no keys is an error
        redis.delete(*keys)                    # idempotent; retried on error
    consumer.commit(batch)                     # only after the deletes succeed
    metrics.observe("invalidation_lag_s", now() - batch.max_commit_ts)

Two details in that code are where pipelines usually go wrong. Using both the before and after images matters when a change moves a row between groups, such as a product changing category: both categories' lists are now stale. And keeping the key derivation in one module shared by readers and the invalidator stops the classic drift where someone adds a new cache key on the read path and nothing ever deletes it. The same log can feed other consumers; the outbox variant of this idea is in the outbox pattern.

Invalidating groups: generation counters and tags

Some changes invalidate a group whose members you cannot cheaply list: every page of a user's feed, every cached query touching a tenant's data, every rendered page using a template. Two techniques avoid enumerating keys.

Generation counters. Store a version per group, and build member keys from it: tenant:42:v17:report:monthly. To invalidate the whole group, increment the version. New reads build keys with v18, miss, and refill; the v17 entries are never read again and age out by TTL or eviction. Invalidation becomes a single atomic increment. The cost is an extra read of the version on every request, usually removed by caching the version in process for a second or two, which then becomes part of your staleness bound, and memory spent on orphaned entries until they expire.

Tags and surrogate keys. HTTP caches and CDNs support attaching labels to a cached response and purging everything with a label. Fastly calls the header Surrogate-Key, Cloudflare Cache-Tag, and Varnish offers bans and the xkey module. The origin labels a product page with product-981 category-12, and a price change purges product-981 without knowing which URLs, languages or query strings exist. Check your provider's limits on purge rate and on tags per response, and how long a purge takes to propagate; those numbers, not your code, set the staleness bound at the edge. CDN layering itself is covered in CDN design.

Worked example: a price change across regions

An online store keeps products and prices in PostgreSQL with a read replica per region, caches product objects in Redis per region, and serves product pages through a CDN. The staleness budget: prices within 2 seconds in Redis and 60 seconds at the edge, descriptions within 5 minutes, and checkout always reads the price from the primary.

A merchant changes product 981's price at 10:00:00.000. The transaction commits on the primary. Debezium emits the change at about 10:00:00.150 and the invalidator in region A deletes product:981 and price:981 at 10:00:00.200. The next region A reader misses, reads the primary or an up-to-date replica, and caches the new price with a 10-minute TTL. The invalidator also purges CDN tag product-981; pages are cached at the edge with s-maxage=60 so even a lost purge is bounded at 60 seconds.

Region B is where it gets subtle. Its invalidator receives the same event at 10:00:00.250, but region B's replica is 400 milliseconds behind. A reader that misses at 10:00:00.300 reads the replica, gets the old price, and caches it for 10 minutes. The delete happened; it just happened too early. The fix is to make region B's invalidator wait until the local replica has replayed past the change's log position, checking the replica's replay position against the event's, before deleting. Facebook's paper describes a related device, remote markers, which send reads for recently-written keys to the primary region until the replica catches up. Either way, invalidation in a replicated system must be ordered after the data the refill will read, not after the primary commit.

Failure modes

These are the incidents that recur, with the signal that catches each.

  • Lost invalidations. Redis pub/sub delivers only to subscribers connected at that moment, so a restarting invalidator misses messages for good. Use a durable log with committed offsets, and keep TTLs as the backstop. Signal: staleness probes failing while lag looks normal.
  • Refill from a lagging replica. As in the example, a delete races ahead of replication and the cache is refilled with old data. Signal: stale reads clustered in non-primary regions.
  • The fill race. A reader loads the old value, a writer commits and deletes, then the reader sets the old value. Leases, where a miss hands out a token that a later delete invalidates so the stale set is rejected, close it; memcache leases were described in the same Facebook paper.
  • Invalidation storms. A bulk update or backfill touches 5 million rows and the invalidator deletes 5 million keys in seconds, turning into a stampede on the database. Rate-limit the invalidator, collapse duplicate keys per batch, and run bulk jobs off-peak with request coalescing on the read path.
  • Key drift. A new read path caches a key the invalidator does not know. Signal: entries that only ever expire, never get deleted. The shared key module prevents it.
  • Over-broad tags. Tagging every page with site means one purge empties the CDN. Tag by entity, not by section.

Measuring and operating invalidation

Measure staleness directly rather than inferring it. Run a canary that writes a timestamped value to a dedicated row every few seconds and reads it back through each cache layer in each region; the difference is your real staleness, and its 99th percentile is the number to hold against the budget. Track invalidation lag from commit timestamp to delete, deletes per second, consumer lag in the log, the ratio of keys deleted to keys expired, and the hit rate around bulk jobs. Give the invalidator a kill switch that falls back to short TTLs, and a replay tool that re-deletes keys for a time window after an incident. The same delivery-and-ordering reasoning underlies read models in CQRS with event sourcing.

What to do next

  1. Write a staleness budget per data class, including which reads must bypass the cache entirely.
  2. Put cache key derivation in one module used by both readers and the invalidator.
  3. Move invalidation from application code to a CDC consumer on the database log, partitioned by primary key, committing offsets only after deletes succeed.
  4. Delete keys for both the before and after images of each change.
  5. In replicated regions, delay each region's deletes until the local replica has replayed past the change.
  6. Keep a TTL on every entry, use generation counters or tags for groups, and rate-limit bulk invalidation.
  7. Run a staleness canary per layer and region and alert on its 99th percentile against the budget.
Key takeaway: Treat invalidation as message delivery with a staleness budget per data class. Delete rather than update, keep a TTL on every entry as the bound on lost messages, and drive deletes from the database change log so every writer is covered. Share one key-derivation module between readers and the invalidator, delete for both old and new row images, and use generation counters or CDN tags for groups. In replicated systems, delete only after the local replica has caught up, rate-limit bulk invalidations, and measure real staleness with a canary in every layer and region.