Redis is usually introduced through its data types, and those are covered in depth elsewhere on this site. This article is about the server underneath them: how one process serves hundreds of thousands of operations a second, what really happens when a key expires or memory runs out, which operations are atomic and which only look atomic, how replicas stay in sync and how failover works, how Cluster splits the keyspace, and how to find the cause when latency suddenly jumps.
Almost every Redis incident traces back to one of these internals: a blocked main thread, an eviction policy that removed keys assumed permanent, an unsafe lock, a resync loop. Each section explains the mechanism and the operational rule that follows. A note on naming: in March 2024 Redis moved from the BSD licence to source-available licences, which led to the Valkey fork; Redis 8 later added AGPLv3 as an option. The mechanisms below are shared by both lines unless stated otherwise, but always check behaviour against the version you run.
One command thread, and why that is the whole design
A Redis server runs an event loop over all client sockets using the operating system's readiness API, such as epoll on Linux. When a socket is readable, Redis reads the request, parses the RESP protocol and executes the command against in-memory data structures, then queues the reply. Commands execute strictly one at a time on one thread. That is the source of Redis's simplest guarantee: every individual command is atomic, because nothing else runs while it executes.
It is also why throughput is high and latency fragile. Most commands finish in microseconds with no locking, but any command that takes 50 milliseconds stalls every other client for 50 milliseconds. KEYS * on a large database, SMEMBERS on a huge set, DEL on a key with millions of elements, or a long Lua script all block the world.
The io-threads setting lets extra threads handle socket reads, parsing and reply writes, which helps when network work dominates; execution stays on the main thread. Redis 8 reworked this, so benchmark your version rather than assuming a speed-up. Background threads free large values and run AOF fsync, and a forked child writes snapshots and AOF rewrites.
The keyspace and incremental rehashing
Each database is essentially a hash table from key to value object, plus a second table mapping keys that have a time to live to their expiry time. Hash tables have to grow as keys are added, and copying millions of entries at once would block the server. Redis therefore rehashes incrementally: during a resize it keeps the old and new tables side by side, every operation on the dictionary migrates a little of the old table, and a periodic timer task migrates more when the server is idle. Lookups check both tables until the move completes.
Memory briefly holds both tables during a resize, which can push an instance near maxmemory over the edge. Iteration must survive rehashes, which is why SCAN uses a cursor and guarantees only that elements present for the whole scan are returned at least once; callers must tolerate duplicates. Never use KEYS in production.
Expiry: lazy plus active
Expired keys are removed in two ways. Lazy expiry happens on access: any command that touches a key first checks its deadline, and if it has passed the key is deleted and the command behaves as if it never existed. Active expiry is a background cycle that repeatedly samples keys with a TTL and deletes those that have expired, continuing while the sample suggests many expired keys remain and stopping when a time budget is spent. The sampling constants have changed between versions, and recent versions expose an active-expire-effort setting to trade CPU for faster reclamation.
Expired keys can occupy memory for a while if nothing touches them. On replicas, expiry is driven by the primary, which sends explicit deletes into the replication stream; a replica reports a logically expired key as missing but does not delete it independently. And giving millions of keys created together the same TTL creates an expiry cliff; add random jitter to cache TTLs.
maxmemory and eviction
When maxmemory is set and a write would exceed it, Redis applies the configured maxmemory-policy. With noeviction the write fails with an out-of-memory error, which is right for data you cannot lose. The allkeys- policies evict any key; the volatile- policies evict only keys that have a TTL. Common variants are LRU, LFU, random and, for volatile keys, shortest remaining TTL; check the documentation for your version for the full list.
Eviction is approximate by design. Rather than maintaining an exact global LRU list, which would cost memory and time on every access, Redis samples a handful of candidate keys (maxmemory-samples, 5 by default), keeps a small pool of good candidates across rounds, and evicts the best. Larger samples approach true LRU at more CPU cost. LFU uses a small logarithmic counter per key that decays over time, which the linked LFU article explains in detail.
The operational rules: use volatile- policies only if every evictable key has a TTL, or the server behaves like noeviction and returns errors; never mix cache data and must-keep data on an instance with an allkeys- policy; and leave headroom below the instance's real memory for fragmentation, client output buffers, replication buffers and copy-on-write pages during a fork.
Atomicity: transactions, WATCH, scripts and Functions
Single commands are atomic, but most application logic needs several. MULTI and EXEC queue commands and run them back to back without interleaving. This is isolation, not rollback: if one queued command fails at runtime, for example a type error, the others still apply. And the commands cannot depend on each other's results, because nothing executes until EXEC.
For read-modify-write logic, WATCH adds optimistic concurrency. You watch a key, read it, decide, then run MULTI/EXEC; if any watched key changed in between, EXEC returns null and you retry:
import redis
r = redis.Redis(host="localhost", port=6379, decode_responses=True)
class InsufficientFunds(Exception):
pass
def debit(account_key: str, amount: int) -> int:
with r.pipeline() as pipe:
while True:
try:
pipe.watch(account_key) # optimistic lock starts
balance = int(pipe.get(account_key) or 0)
if balance < amount:
pipe.unwatch()
raise InsufficientFunds(account_key)
pipe.multi() # queue commands from here
pipe.decrby(account_key, amount)
new_balance, = pipe.execute() # aborts if key changed since WATCH
return new_balance
except redis.WatchError:
continue # someone else wrote; retryUnder contention the retry loop wastes work. A Lua script sent with EVAL instead runs atomically on the server and can branch on values it reads. Redis 7 added Functions, which are named libraries of server-side code loaded with FUNCTION LOAD and called with FCALL; they persist and replicate as part of the dataset, unlike the script cache. The classic correct pattern is compare-and-delete for lock release:
# Release a lock only if we still own it: compare and delete in one atomic step.
RELEASE = r.register_script("""
if redis.call('GET', KEYS[1]) == ARGV[1] then
return redis.call('DEL', KEYS[1])
end
return 0
""")
token = "worker-7:3f9c"
if r.set("lock:{invoice:991}", token, nx=True, px=30000):
try:
... # do the work
finally:
RELEASE(keys=["lock:{invoice:991}"], args=[token])Scripts block the server while they run, so keep them short. After the busy-script threshold (5 seconds by default) other clients get a BUSY error, and SCRIPT KILL works only if the script has not written. Declare every key in KEYS so Cluster can route. This lock is not safe across failover: asynchronous replication can lose the lock key when a replica is promoted, so use fencing tokens where mutual exclusion matters.
Replication: IDs, offsets and the backlog
Replication is asynchronous. The primary streams every write command to its replicas and does not wait for them before replying to the client. Each primary has a replication ID, and every byte of the stream advances a replication offset. A replica remembers the ID and offset it has reached. When it reconnects it sends PSYNC with both; if the primary still holds the missing bytes in its replication backlog, a fixed-size ring buffer, it sends only those, which is a partial resync. Otherwise it performs a full resync: it produces an RDB snapshot, either on disk or streamed directly to the socket when diskless sync is enabled, transfers it, then streams the writes that happened meanwhile.
Full resyncs are expensive. If replicas keep doing them after short network blips, the backlog is too small: the default repl-backlog-size is 1 MB, so size it as write bytes per second times the longest disconnection you want to survive. If the replica output buffer overflows during a full sync, the sync restarts in a loop; raise the replica class of client-output-buffer-limit.
WAIT lets a client block until a number of replicas acknowledge its writes, and min-replicas-to-write makes the primary refuse writes when too few replicas are connected. Both narrow the window for losing acknowledged writes on failover. Neither turns Redis into a strongly consistent store.
Failover: Sentinel and Cluster
Sentinel adds automatic failover. Three or more Sentinel processes on separate hosts monitor the primary; when a quorum agree it has been unreachable for down-after-milliseconds, a Sentinel elected by majority promotes the best replica and reconfigures the rest. Clients must ask Sentinel for the current primary (SENTINEL get-master-addr-by-name) instead of hard-coding it.
Cluster shards the keyspace as well as providing failover. Every key maps to one of 16,384 hash slots by CRC16 of the key modulo 16,384, and each primary owns a range of slots. A client that sends a command to the wrong node gets MOVED with the right address, and should update its slot map. During resharding, a slot being migrated can produce ASK, a one-time redirect that the client follows by sending ASKING before the command, without updating its map. Multi-key commands, transactions and scripts work only when all keys are in one slot, otherwise you get a CROSSSLOT error. Hash tags control this: only the part of the key inside the first braces is hashed, so cart:{user:42} and orders:{user:42} land in the same slot. A popular tag concentrates load on one shard, so use them with care.
Worked example: a latency spike every few minutes
A session store serving about 60,000 operations a second reports p99 latency jumping from under 1 millisecond to over 100 milliseconds for a few seconds roughly every five minutes. The numbers are illustrative; the method is general.
# Is the server itself slow, or the network and client?
redis-cli --latency -h cache-1 # round-trip samples from this host
redis-cli -h cache-1 LATENCY LATEST # latest spike per event class
redis-cli -h cache-1 LATENCY DOCTOR # human-readable analysis
redis-cli -h cache-1 SLOWLOG GET 10 # commands slower than slowlog-log-slower-than
redis-cli -h cache-1 INFO stats | grep latest_fork_usec
redis-cli -h cache-1 INFO memory | grep -E "used_memory_human|used_memory_rss_human|mem_fragmentation_ratio"
redis-cli -h cache-1 --bigkeys # largest key per type, via SCAN
redis-cli -h cache-1 CONFIG SET latency-monitor-threshold 50 # record events over 50 msredis-cli --latency from an application host shows the same spikes, so the client is not the cause. LATENCY LATEST shows command and fork events. The slow log holds 90 millisecond DEL calls from a cleanup job deleting sets with millions of members. Switching to UNLINK, which unlinks the key immediately and frees memory on a background thread, removes those spikes; lazyfree-lazy-user-del does the same for DEL calls the team does not control.
The fork events line up with RDB snapshots, and latest_fork_usec reports about 300 milliseconds on a 25 GB instance. Fork time grows with page-table size, and Transparent Huge Pages make the copy-on-write that follows far more expensive; Redis warns about THP in its startup log. The team disables THP and moves snapshots to a replica, and p99 stays under 2 milliseconds across the snapshot cycle.
Failure modes and trade-offs
- Blocking commands. KEYS, large DEL, unbounded range reads and long scripts stall everyone. Rename or disable dangerous commands via ACLs and alert on the slow log.
- Silent eviction. An allkeys policy on an instance also holding queues or locks will eventually evict them. Separate cache and state onto different instances.
- Resync loops. A backlog or replica output buffer that is too small turns brief disconnects into repeated full syncs. Monitor sync counters and size both.
- Lost writes on failover. Asynchronous replication loses the tail of acknowledged writes when a primary dies. Use WAIT and min-replicas settings where it matters, and never treat Redis as the only copy of critical data.
- Hot keys and hot slots. One key or hash tag saturates one shard. Split the key or read from replicas where staleness is acceptable.
What to do next
- Run
INFO,SLOWLOG GETandLATENCY DOCTORon each production instance and note the worst offenders. - Set
latency-monitor-thresholdand alert on slow-log growth andlatest_fork_usec. - Confirm each instance's
maxmemory-policymatches what it stores, and split cache from must-keep data. - Replace DEL of large keys with UNLINK, and KEYS with SCAN, across your codebase.
- Size
repl-backlog-sizefrom your write rate and check for repeated full syncs. - Review every lock and multi-key operation for correctness under failover and Cluster slot rules.
- Rehearse a failover in staging and measure how long clients take to follow the new primary.
Related reading: Redis data structures for types and encodings, Redis persistence modes for RDB, AOF and fsync, Redis Streams architecture for consumer groups, LFU caches and Redis's approximated LFU for eviction internals, and ElastiCache in depth for running it as a managed service.