Amazon ElastiCache is AWS's managed in-memory data store. It runs three open-source engines, Valkey, Redis OSS and Memcached, on infrastructure AWS operates. AWS handles provisioning, patching, replication, failover, backups and, in the serverless option, scaling. You get sub-millisecond reads and writes for caching, sessions, rate limits, leaderboards and queues, without running servers yourself.
Managed does not mean decision-free. You still choose the engine, deployment shape, sharding, client settings, eviction and caching pattern, and those choices decide whether the cache absorbs a spike or turns it into an outage. This article works through each, with limits checked against the AWS documentation. The engine's data structures are covered in Redis data structures in depth.
Engines: Valkey, Redis OSS and Memcached
Valkey is a Linux Foundation fork of Redis, created in 2024 after Redis changed its licence. It is wire-compatible, so Redis clients and commands work against it. AWS supports Valkey from 7.2 and prices it below Redis OSS: AWS states about 20% lower per node and 33% lower on serverless. Recent Valkey versions on ElastiCache also gain features first; for example, node-based Valkey 9.0 and later can optionally persist writes to a Multi-AZ transactional log. Redis OSS remains supported for existing workloads. Memcached is a simpler engine: a multithreaded key-value cache with no replication, no persistence and no data structures beyond strings.
| Need | Valkey / Redis OSS | Memcached |
|---|---|---|
| Data structures (hash, sorted set, stream) | Yes | No, opaque values only |
| Replicas and automatic failover | Yes, up to 5 replicas per shard | No; a lost node loses its keys |
| Backups and restore | Yes | Serverless only |
| Pub/sub, Lua, transactions | Yes | No |
| Multithreaded per node | I/O threads | Yes |
| Typical use | Cache, sessions, leaderboards, rate limits, queues | Pure, rebuildable cache |
For new workloads, default to Valkey. Choose Memcached only if you have an existing Memcached client fleet and a cache you can afford to lose node by node.
Deployment shapes: serverless versus node-based
ElastiCache Serverless asks only for a name. You connect to a single endpoint, and a proxy layer behind a Network Load Balancer routes commands to cache nodes. AWS scales those nodes up and out as memory, CPU and network use grow, replicates data across Availability Zones and states a 99.99% availability SLA. Serverless always encrypts in transit and at rest, runs the engine in cluster mode, and works only with clients that support TLS. You pay for data stored in GB-hours and for compute in ECPUs. A simple read or write costs 1 ECPU per kilobyte transferred, so a 3.2 KB GET costs 3.2 ECPUs, while commands that use more CPU are charged by whichever is higher, CPU time or data. The pricing page lists a minimum metered storage of 100 MB per cache for Valkey and 1 GB for Redis OSS and Memcached.
Node-based clusters let you choose the node type, the number of shards and the number of replicas, the AZ placement and the engine parameters. You pay per node-hour whether or not the memory is used. With Valkey or Redis OSS, a cluster runs in one of two modes. Cluster mode disabled has exactly one shard: one primary plus up to five replicas, a primary endpoint for writes and a reader endpoint that spreads reads across replicas. Cluster mode enabled splits the keyspace into 16,384 hash slots spread over many shards. AWS documents a shard as one to six nodes, and a default limit of 90 nodes per cluster, from 90 shards with no replicas to 15 shards with five replicas each. The limit can be raised to 500 nodes on Valkey 7.2+ or Redis OSS 5.0.6 to 7.1.
Choose serverless for spiky or unknown traffic and many small caches; choose node-based for large, steady traffic where node-hours beat ECPUs, or when you need specific parameters, data tiering or placement control.
Creating a cluster and connecting to it
A production node-based cluster in the CLI looks like this. Every flag is a decision: TLS, encryption at rest, Multi-AZ with automatic failover, private subnets, a dedicated security group and a custom parameter group.
# Node-based Valkey, cluster mode enabled: 3 shards x (1 primary + 1 replica)
aws elasticache create-replication-group \
--replication-group-id orders-cache \
--replication-group-description "orders read cache" \
--engine valkey \
--cache-node-type cache.r7g.large \
--num-node-groups 3 --replicas-per-node-group 1 \
--automatic-failover-enabled --multi-az-enabled \
--transit-encryption-enabled --at-rest-encryption-enabled \
--cache-subnet-group-name private-cache-subnets \
--security-group-ids sg-0abc123 \
--cache-parameter-group-name orders-valkey-params \
--snapshot-retention-limit 3
# Serverless alternative: a name and limits, no shards to plan
aws elasticache create-serverless-cache \
--serverless-cache-name sessions --engine valkeyClients must match the mode. With cluster mode enabled, connect to the configuration endpoint using a cluster-aware client. The client downloads the slot map, hashes each key with CRC16 modulo 16384, and sends the command straight to the right shard. When resharding moves a slot, the server replies MOVED and the client refreshes its map. Multi-key commands such as MGET or a Lua script only work when all keys hash to the same slot. A hash tag, the part of a key inside braces, controls the hash, so cart:{u42}:items and cart:{u42}:total always share a slot.
import json, random
from redis.cluster import RedisCluster # redis-py 4.1+; also works against Valkey
cache = RedisCluster(
host="clustercfg.orders-cache.xxxxxx.use1.cache.amazonaws.com", # configuration endpoint
port=6379, ssl=True,
socket_timeout=0.05, socket_connect_timeout=0.2, # a cache must fail fast
read_from_replicas=True, # accept slightly stale reads
)
TTL = 300
def get_product(pid):
key = f"product:{{{pid}}}" # hash tag: one slot per product
try:
hit = cache.get(key)
if hit is not None:
return json.loads(hit)
except Exception:
metrics.incr("cache.error") # degrade to the database, never fail
row = db.fetch_product(pid)
try:
cache.set(key, json.dumps(row), ex=TTL + random.randint(0, 60)) # jitter
except Exception:
pass
return row
def update_product(pid, fields):
db.update_product(pid, fields) # source of truth first
cache.delete(f"product:{{{pid}}}") # then invalidate, don't overwriteNote the millisecond timeouts (a slow cache is worse than a missing one), the database fallback around every cache call, and replica reads only because product data tolerates a few milliseconds of lag.
Caching patterns and a worked example
The client above implements cache-aside: read the cache; on a miss, read the database and populate the cache with a TTL. On writes, update the database and delete the key, rather than writing the new value into the cache. Deleting avoids a race in which two concurrent writers leave the older value cached. The random jitter on the TTL prevents thousands of keys created together from expiring together. The patterns, including write-through and write-behind, are compared in caching architecture. If your source of truth is DynamoDB, DAX provides a write-through cache with no application code.
When a very hot key expires, every request misses at once and stampedes the database. A short lock lets one caller rebuild while the others wait briefly:
def get_with_lock(key, load, ttl=300):
val = cache.get(key)
if val is not None:
return val
# only one caller rebuilds; others wait briefly or serve a fallback
if cache.set(f"lock:{key}", "1", nx=True, ex=10):
try:
val = load()
cache.set(key, val, ex=ttl)
finally:
cache.delete(f"lock:{key}")
return val
time.sleep(0.05)
return cache.get(key) or load()Worked sizing example. A catalogue serves 20,000 reads per second at a 95% hit rate, so the database sees about 1,000. The working set is 2 million products at about 1.5 KB, roughly 3 GB, or about 4 GB after keys, overhead and fragmentation. Node-based, that means nodes whose usable memory (after the 25% reserve) exceeds 4 GB with headroom, plus a replica per shard. On serverless, each 1.5 KB read costs about 1.5 ECPUs, so about 30,000 ECPUs per second before writes. Price both with current Regional rates; steady high rates usually favour nodes.
Memory: reserved memory, eviction and big keys
Each node type has an advertised memory size, and ElastiCache sets maxmemory below it. The ElastiCache-specific reserved-memory-percent parameter defaults to 25%. That reserve covers replica output buffers, fragmentation and the copy-on-write pages created when the engine forks for a snapshot or a full sync to a replica. AWS's guidance is not to lower it, and to raise it on small instances. If free memory is too low for a forked save, ElastiCache switches to a forkless save, which uses less memory but can add client latency. Redis persistence modes explains why forks need that headroom.
When used memory reaches maxmemory, the maxmemory-policy parameter decides what happens. For a pure cache, allkeys-lru or allkeys-lfu evicts any key. The volatile-* policies only evict keys that have a TTL, so a cache full of keys without TTLs can reject writes with out-of-memory errors. noeviction suits stores where losing a key is wrong. Set the policy explicitly in a custom parameter group (you cannot modify the default groups), rather than relying on whatever default your engine family has.
Watch for big keys. A 50 MB hash or a list with millions of entries is moved, serialised and deleted as a unit, so one DEL or HGETALL can block a single-threaded command loop for many milliseconds. Use UNLINK for asynchronous deletes, cap collection sizes and split hot aggregates across keys. For large, mostly cold datasets, node-based data tiering on R6gd nodes moves less-used values to local SSD, which AWS says can cut cost by over 60% at full utilisation, at the price of higher latency for values read from SSD.
Replication, failover and what you can lose
In Valkey and Redis OSS, replication from primary to replica is asynchronous. With Multi-AZ and automatic failover enabled, ElastiCache detects a failed primary, promotes the replica with the least lag and updates the endpoints. Any writes the old primary acknowledged but had not yet shipped are lost. For a cache this is acceptable. For a session store or a rate limiter it means a few users may be logged out or briefly over quota. If you need more, WAIT blocks until a number of replicas acknowledge a write, which narrows the window but does not provide durable consensus. Node-based Valkey 9.0+ can also commit writes to the Multi-AZ transactional log mentioned earlier.
Clients see failover as connection errors followed by a topology change, so they must reconnect and refresh. Run the test-failover API on staging and measure how long the application takes to recover. Global Datastore adds read-only secondary Regions for disaster recovery and local reads, replicated asynchronously, and promotable when needed.
Scaling node-based clusters
- Vertical: change the node type. ElastiCache performs it online by adding new nodes and syncing, but plan it outside peak hours.
- Read scaling: add replicas, up to five per shard, and send tolerant reads to them.
- Write and memory scaling: add shards with online resharding. Slots migrate while the cluster serves traffic, and clients follow
MOVEDreplies. - Hot keys: resharding does not help a single hot key, because a key lives in exactly one slot. Use a local in-process cache with a short TTL, or replicate the value under several keys.
- IP addresses: large clusters need subnets with enough free addresses. AWS notes that undersized CIDR ranges are a common reason node additions fail.
Security and monitoring
ElastiCache nodes live in your VPC, in a subnet group of private subnets. Allow access only from application security groups on the cache port; see the VPC architecture article for the network design. Enable encryption in transit and at rest on node-based clusters (serverless always has both). For authentication, use role-based access control users and user groups, or IAM authentication on supported Valkey and Redis OSS versions, rather than a single shared AUTH token.
In CloudWatch, alarm on EngineCPUUtilization (the engine's main thread, which saturates before host CPU does), DatabaseMemoryUsagePercentage, Evictions, CurrConnections and sudden jumps in NewConnections (a sign of connection churn), ReplicationLag, and the cache hit rate. For serverless, also track ElastiCacheProcessingUnits, because it is your bill. A falling hit rate with rising evictions means the working set no longer fits. Rising connections with flat traffic usually means clients are not pooling.
Failure modes and trade-offs
| Symptom | Likely cause | Fix |
|---|---|---|
| Database overload after a deploy or failover | Cold cache or stampede | Jittered TTLs, rebuild locks, warm critical keys |
| OOM errors on writes | volatile policy with keys lacking TTLs, or noeviction | Set TTLs; choose allkeys-lru or allkeys-lfu |
| Latency spikes every few minutes | Big-key commands or fork for snapshots | Split keys, UNLINK, keep the 25% reserve |
| CROSSSLOT errors | Multi-key command across slots | Hash tags or per-key commands |
| App errors during failover | Long timeouts, no reconnect | Short timeouts, cluster client, fallback to DB |
| Stale reads | Reading from replicas | Read the primary for read-your-writes paths |
A cache buys latency and offloads the database at the price of staleness and a second system; ElastiCache removes the server work, not the design work.
What to do next
- Pick the engine (Valkey unless you have a reason not to) and the shape: serverless for spiky or small workloads, node-based for large steady ones. Price both with your real request size and rate.
- Create a custom parameter group, set
maxmemory-policyexplicitly and leavereserved-memory-percentat 25% or higher. - Enable TLS, encryption at rest, Multi-AZ with automatic failover and RBAC or IAM authentication; restrict the security group to application tiers.
- Use a cluster-aware client with millisecond timeouts, connection pooling and a database fallback on every cache call.
- Implement cache-aside with delete-on-write, TTL jitter and a rebuild lock for hot keys; design hash tags for any multi-key operations.
- Add CloudWatch alarms on engine CPU, memory usage, evictions, connections, replication lag and hit rate.
- Run
test-failoverin staging and record how long the application takes to recover.