Amazon ElastiCache is AWS's managed in-memory data store. It runs three open-source engines, Valkey, Redis OSS and Memcached, on infrastructure AWS operates. AWS handles provisioning, patching, replication, failover, backups and, in the serverless option, scaling. You get sub-millisecond reads and writes for caching, sessions, rate limits, leaderboards and queues, without running servers yourself.

Managed does not mean decision-free. You still choose the engine, deployment shape, sharding, client settings, eviction and caching pattern, and those choices decide whether the cache absorbs a spike or turns it into an outage. This article works through each, with limits checked against the AWS documentation. The engine's data structures are covered in Redis data structures in depth.

Advertisement

Engines: Valkey, Redis OSS and Memcached

Valkey is a Linux Foundation fork of Redis, created in 2024 after Redis changed its licence. It is wire-compatible, so Redis clients and commands work against it. AWS supports Valkey from 7.2 and prices it below Redis OSS: AWS states about 20% lower per node and 33% lower on serverless. Recent Valkey versions on ElastiCache also gain features first; for example, node-based Valkey 9.0 and later can optionally persist writes to a Multi-AZ transactional log. Redis OSS remains supported for existing workloads. Memcached is a simpler engine: a multithreaded key-value cache with no replication, no persistence and no data structures beyond strings.

NeedValkey / Redis OSSMemcached
Data structures (hash, sorted set, stream)YesNo, opaque values only
Replicas and automatic failoverYes, up to 5 replicas per shardNo; a lost node loses its keys
Backups and restoreYesServerless only
Pub/sub, Lua, transactionsYesNo
Multithreaded per nodeI/O threadsYes
Typical useCache, sessions, leaderboards, rate limits, queuesPure, rebuildable cache

For new workloads, default to Valkey. Choose Memcached only if you have an existing Memcached client fleet and a cache you can afford to lose node by node.

Deployment shapes: serverless versus node-based

Node-based, cluster mode enabledApp (cluster client)TLS, slot mapconfiguration endpointShard 1 primaryslots 0-5460Shard 2 primaryslots 5461-10922Shard 3 primaryslots 10923-16383replica (AZ b)replica (AZ c)replica (AZ a)asyncasyncasyncclient routes by CRC16(key) mod 16384; MOVED on reshardServerlessApp (TLS client)one endpointProxy layer behind NLBhides topologyCache nodes across AZsscale up and outbilled per GB-hour stored plus ECPUsBoth: inside your VPC, security groups, CloudWatch metricsNode-based gives control of shards, replicas and parameters; serverless trades that control for automatic scaling
Left: a node-based cluster with three shards, each a primary plus a replica in another AZ, addressed by a cluster-aware client. Right: a serverless cache, where a proxy layer hides the nodes behind one endpoint.

ElastiCache Serverless asks only for a name. You connect to a single endpoint, and a proxy layer behind a Network Load Balancer routes commands to cache nodes. AWS scales those nodes up and out as memory, CPU and network use grow, replicates data across Availability Zones and states a 99.99% availability SLA. Serverless always encrypts in transit and at rest, runs the engine in cluster mode, and works only with clients that support TLS. You pay for data stored in GB-hours and for compute in ECPUs. A simple read or write costs 1 ECPU per kilobyte transferred, so a 3.2 KB GET costs 3.2 ECPUs, while commands that use more CPU are charged by whichever is higher, CPU time or data. The pricing page lists a minimum metered storage of 100 MB per cache for Valkey and 1 GB for Redis OSS and Memcached.

Node-based clusters let you choose the node type, the number of shards and the number of replicas, the AZ placement and the engine parameters. You pay per node-hour whether or not the memory is used. With Valkey or Redis OSS, a cluster runs in one of two modes. Cluster mode disabled has exactly one shard: one primary plus up to five replicas, a primary endpoint for writes and a reader endpoint that spreads reads across replicas. Cluster mode enabled splits the keyspace into 16,384 hash slots spread over many shards. AWS documents a shard as one to six nodes, and a default limit of 90 nodes per cluster, from 90 shards with no replicas to 15 shards with five replicas each. The limit can be raised to 500 nodes on Valkey 7.2+ or Redis OSS 5.0.6 to 7.1.

Choose serverless for spiky or unknown traffic and many small caches; choose node-based for large, steady traffic where node-hours beat ECPUs, or when you need specific parameters, data tiering or placement control.

Advertisement

Creating a cluster and connecting to it

A production node-based cluster in the CLI looks like this. Every flag is a decision: TLS, encryption at rest, Multi-AZ with automatic failover, private subnets, a dedicated security group and a custom parameter group.

# Node-based Valkey, cluster mode enabled: 3 shards x (1 primary + 1 replica)
aws elasticache create-replication-group \
  --replication-group-id orders-cache \
  --replication-group-description "orders read cache" \
  --engine valkey \
  --cache-node-type cache.r7g.large \
  --num-node-groups 3 --replicas-per-node-group 1 \
  --automatic-failover-enabled --multi-az-enabled \
  --transit-encryption-enabled --at-rest-encryption-enabled \
  --cache-subnet-group-name private-cache-subnets \
  --security-group-ids sg-0abc123 \
  --cache-parameter-group-name orders-valkey-params \
  --snapshot-retention-limit 3

# Serverless alternative: a name and limits, no shards to plan
aws elasticache create-serverless-cache \
  --serverless-cache-name sessions --engine valkey

Clients must match the mode. With cluster mode enabled, connect to the configuration endpoint using a cluster-aware client. The client downloads the slot map, hashes each key with CRC16 modulo 16384, and sends the command straight to the right shard. When resharding moves a slot, the server replies MOVED and the client refreshes its map. Multi-key commands such as MGET or a Lua script only work when all keys hash to the same slot. A hash tag, the part of a key inside braces, controls the hash, so cart:{u42}:items and cart:{u42}:total always share a slot.

import json, random
from redis.cluster import RedisCluster   # redis-py 4.1+; also works against Valkey

cache = RedisCluster(
    host="clustercfg.orders-cache.xxxxxx.use1.cache.amazonaws.com",  # configuration endpoint
    port=6379, ssl=True,
    socket_timeout=0.05, socket_connect_timeout=0.2,   # a cache must fail fast
    read_from_replicas=True,                           # accept slightly stale reads
)

TTL = 300

def get_product(pid):
    key = f"product:{{{pid}}}"                         # hash tag: one slot per product
    try:
        hit = cache.get(key)
        if hit is not None:
            return json.loads(hit)
    except Exception:
        metrics.incr("cache.error")                    # degrade to the database, never fail
    row = db.fetch_product(pid)
    try:
        cache.set(key, json.dumps(row), ex=TTL + random.randint(0, 60))  # jitter
    except Exception:
        pass
    return row

def update_product(pid, fields):
    db.update_product(pid, fields)                     # source of truth first
    cache.delete(f"product:{{{pid}}}")                 # then invalidate, don't overwrite

Note the millisecond timeouts (a slow cache is worse than a missing one), the database fallback around every cache call, and replica reads only because product data tolerates a few milliseconds of lag.

Caching patterns and a worked example

The client above implements cache-aside: read the cache; on a miss, read the database and populate the cache with a TTL. On writes, update the database and delete the key, rather than writing the new value into the cache. Deleting avoids a race in which two concurrent writers leave the older value cached. The random jitter on the TTL prevents thousands of keys created together from expiring together. The patterns, including write-through and write-behind, are compared in caching architecture. If your source of truth is DynamoDB, DAX provides a write-through cache with no application code.

When a very hot key expires, every request misses at once and stampedes the database. A short lock lets one caller rebuild while the others wait briefly:

def get_with_lock(key, load, ttl=300):
    val = cache.get(key)
    if val is not None:
        return val
    # only one caller rebuilds; others wait briefly or serve a fallback
    if cache.set(f"lock:{key}", "1", nx=True, ex=10):
        try:
            val = load()
            cache.set(key, val, ex=ttl)
        finally:
            cache.delete(f"lock:{key}")
        return val
    time.sleep(0.05)
    return cache.get(key) or load()

Worked sizing example. A catalogue serves 20,000 reads per second at a 95% hit rate, so the database sees about 1,000. The working set is 2 million products at about 1.5 KB, roughly 3 GB, or about 4 GB after keys, overhead and fragmentation. Node-based, that means nodes whose usable memory (after the 25% reserve) exceeds 4 GB with headroom, plus a replica per shard. On serverless, each 1.5 KB read costs about 1.5 ECPUs, so about 30,000 ECPUs per second before writes. Price both with current Regional rates; steady high rates usually favour nodes.

Memory: reserved memory, eviction and big keys

Each node type has an advertised memory size, and ElastiCache sets maxmemory below it. The ElastiCache-specific reserved-memory-percent parameter defaults to 25%. That reserve covers replica output buffers, fragmentation and the copy-on-write pages created when the engine forks for a snapshot or a full sync to a replica. AWS's guidance is not to lower it, and to raise it on small instances. If free memory is too low for a forked save, ElastiCache switches to a forkless save, which uses less memory but can add client latency. Redis persistence modes explains why forks need that headroom.

When used memory reaches maxmemory, the maxmemory-policy parameter decides what happens. For a pure cache, allkeys-lru or allkeys-lfu evicts any key. The volatile-* policies only evict keys that have a TTL, so a cache full of keys without TTLs can reject writes with out-of-memory errors. noeviction suits stores where losing a key is wrong. Set the policy explicitly in a custom parameter group (you cannot modify the default groups), rather than relying on whatever default your engine family has.

Watch for big keys. A 50 MB hash or a list with millions of entries is moved, serialised and deleted as a unit, so one DEL or HGETALL can block a single-threaded command loop for many milliseconds. Use UNLINK for asynchronous deletes, cap collection sizes and split hot aggregates across keys. For large, mostly cold datasets, node-based data tiering on R6gd nodes moves less-used values to local SSD, which AWS says can cut cost by over 60% at full utilisation, at the price of higher latency for values read from SSD.

Replication, failover and what you can lose

In Valkey and Redis OSS, replication from primary to replica is asynchronous. With Multi-AZ and automatic failover enabled, ElastiCache detects a failed primary, promotes the replica with the least lag and updates the endpoints. Any writes the old primary acknowledged but had not yet shipped are lost. For a cache this is acceptable. For a session store or a rate limiter it means a few users may be logged out or briefly over quota. If you need more, WAIT blocks until a number of replicas acknowledge a write, which narrows the window but does not provide durable consensus. Node-based Valkey 9.0+ can also commit writes to the Multi-AZ transactional log mentioned earlier.

Clients see failover as connection errors followed by a topology change, so they must reconnect and refresh. Run the test-failover API on staging and measure how long the application takes to recover. Global Datastore adds read-only secondary Regions for disaster recovery and local reads, replicated asynchronously, and promotable when needed.

Scaling node-based clusters

  • Vertical: change the node type. ElastiCache performs it online by adding new nodes and syncing, but plan it outside peak hours.
  • Read scaling: add replicas, up to five per shard, and send tolerant reads to them.
  • Write and memory scaling: add shards with online resharding. Slots migrate while the cluster serves traffic, and clients follow MOVED replies.
  • Hot keys: resharding does not help a single hot key, because a key lives in exactly one slot. Use a local in-process cache with a short TTL, or replicate the value under several keys.
  • IP addresses: large clusters need subnets with enough free addresses. AWS notes that undersized CIDR ranges are a common reason node additions fail.

Security and monitoring

ElastiCache nodes live in your VPC, in a subnet group of private subnets. Allow access only from application security groups on the cache port; see the VPC architecture article for the network design. Enable encryption in transit and at rest on node-based clusters (serverless always has both). For authentication, use role-based access control users and user groups, or IAM authentication on supported Valkey and Redis OSS versions, rather than a single shared AUTH token.

In CloudWatch, alarm on EngineCPUUtilization (the engine's main thread, which saturates before host CPU does), DatabaseMemoryUsagePercentage, Evictions, CurrConnections and sudden jumps in NewConnections (a sign of connection churn), ReplicationLag, and the cache hit rate. For serverless, also track ElastiCacheProcessingUnits, because it is your bill. A falling hit rate with rising evictions means the working set no longer fits. Rising connections with flat traffic usually means clients are not pooling.

Failure modes and trade-offs

SymptomLikely causeFix
Database overload after a deploy or failoverCold cache or stampedeJittered TTLs, rebuild locks, warm critical keys
OOM errors on writesvolatile policy with keys lacking TTLs, or noevictionSet TTLs; choose allkeys-lru or allkeys-lfu
Latency spikes every few minutesBig-key commands or fork for snapshotsSplit keys, UNLINK, keep the 25% reserve
CROSSSLOT errorsMulti-key command across slotsHash tags or per-key commands
App errors during failoverLong timeouts, no reconnectShort timeouts, cluster client, fallback to DB
Stale readsReading from replicasRead the primary for read-your-writes paths

A cache buys latency and offloads the database at the price of staleness and a second system; ElastiCache removes the server work, not the design work.

What to do next

  1. Pick the engine (Valkey unless you have a reason not to) and the shape: serverless for spiky or small workloads, node-based for large steady ones. Price both with your real request size and rate.
  2. Create a custom parameter group, set maxmemory-policy explicitly and leave reserved-memory-percent at 25% or higher.
  3. Enable TLS, encryption at rest, Multi-AZ with automatic failover and RBAC or IAM authentication; restrict the security group to application tiers.
  4. Use a cluster-aware client with millisecond timeouts, connection pooling and a database fallback on every cache call.
  5. Implement cache-aside with delete-on-write, TTL jitter and a rebuild lock for hot keys; design hash tags for any multi-key operations.
  6. Add CloudWatch alarms on engine CPU, memory usage, evictions, connections, replication lag and hit rate.
  7. Run test-failover in staging and record how long the application takes to recover.
Key takeaway: ElastiCache runs Valkey, Redis OSS or Memcached for you, but the design decisions remain yours. Default to Valkey. Choose serverless (one TLS endpoint behind a proxy, billed per GB-hour and ECPU) for spiky workloads, and node-based clusters (shards of one to six nodes, 16,384 hash slots, up to five replicas each) for steady scale. Keep the 25% memory reserve, set the eviction policy explicitly, and use cache-aside with delete-on-write, jittered TTLs and stampede locks. Expect asynchronous replication to lose a few writes on failover, and alarm on engine CPU, memory, evictions and hit rate.