Heroku's managed Redis add-on is the default place a Heroku app puts its cache, its sessions and its job queue. It also surprises teams in production: a full cache that throws write errors, a Sidekiq queue that loses jobs after one setting changes, clients that break after maintenance, and TLS errors on the first deploy. None of these are bugs; they are defaults the provisioning command does not explain.

This article covers what you provision, how failover changes the connection string under you, how to connect and budget clients, how to pick an eviction policy for each job the store does, and how to observe it. It assumes you know what Redis is; for the data structures and persistence internals see Redis in depth and Redis persistence modes. Facts were checked against the Heroku Dev Center in October 2026; read plan sizes and connection limits from the pricing page.

What you are provisioning

The first thing to know is the name. Heroku renamed the product to Heroku Key-Value Store, and it now runs Valkey, the open-source fork of Redis, rather than Redis itself. Valkey speaks the same protocol, so every Redis client library works unchanged. The add-on slug and CLI did not change: you still provision heroku-redis:<plan>, the default config var is still REDIS_URL, and the management commands are still heroku redis:*. When the documentation was checked, Valkey 9.1 was the default version, 8.1 was also available, and 7.2 was being deprecated, with an end date of 19 December 2026.

There are four plan families, and the difference between them is not just size:

Plan familyPersistenceHigh availabilityUse it for
MiniNone. A reboot or failure loses all dataNo standbyDevelopment, review apps, disposable caches
PremiumAOF, written to disk every secondHA standby with automatic failoverProduction caches, sessions, queues
PrivateAOF every secondHA standbyApps in Private Spaces
ShieldAOF every secondHA standbyCompliance workloads in Shield Private Spaces

Two of these facts drive most design decisions. Mini does not persist anything, so never put job queues or anything you cannot rebuild on it. And because production plans write the append-only file once a second, a crash can lose up to about one second of acknowledged writes. That is fine for a cache and usually acceptable for a queue, but it is not the durability of Postgres. Anything that must not be lost belongs in Heroku Postgres, with Redis holding a copy or a pointer.

Architecture and failover

On a production plan you get a primary instance, a standby that receives replication from it, and a control plane that watches both. Your dynos see none of this. They see one URL in a config var.

A production Heroku Key-Value Store: the app sees one config var, the platform owns everything behind itweb dynoscache reads, sessionsworker dynosSidekiq, RQ, BullMQREDIS_URLrediss://, TLS, self-signedPrimaryValkey, AOF every secondHA standbyPremium, Private, ShieldcommandsreplicationControl planehealth checks, failoverOn failover or maintenancestandby promoted, URL may change, new app releaseconfig var rewrittenMini has no standby and no persistence: a reboot or failure loses every key.Clients must read REDIS_URL at boot and reconnect on errors; the dyno restart on release does the rest.
Figure 1. The app connects only through REDIS_URL. Failover and maintenance can point that variable at a new host and trigger a release.

When the control plane sees a problem, it first checks for two minutes from several network locations to confirm that the primary is really unavailable. Then it promotes the standby. The Dev Center says that after a failover the connection string "can change and point to a different host". Heroku updates the config var and creates a new release, which restarts your dynos with the new value. On the single-tenant plans (premium-7 and later) failover is transparent: the server changes but the hostname does not.

So in the worst case a primary failure costs about two minutes of detection, then the promotion, then a dyno restart. Your code must handle that window: commands fail, then the process restarts with a new URL. That leads to the most important rule for this add-on: read the URL from the environment at boot, every time. Never copy it into another app's config, a CI secret or a local file. The Dev Center is explicit that the values can change at any time.

Connecting correctly

Every plan requires TLS. The URL scheme is rediss:// (two s's), and the server presents a self-signed certificate. A client that verifies certificates by default will reject the connection, so Heroku's own examples turn verification off. In redis-py that is ssl_cert_reqs=None, in Ruby verify_mode: OpenSSL::SSL::VERIFY_NONE, and in Node rejectUnauthorized: false.

Be clear about what this gives up. Traffic is still encrypted, so a passive listener on the network learns nothing. But the client no longer proves it is talking to the real server, so an attacker who can redirect traffic could pose as it. If your threat model does not, use Private or Shield plans in a Private Space, where the network itself is isolated. A production-grade redis-py client also needs reconnect behaviour and a cap on its pool:

import os
import redis
from redis.backoff import ExponentialBackoff
from redis.retry import Retry

def make_client(pool_size: int) -> redis.Redis:
    return redis.from_url(
        os.environ["REDIS_URL"],          # read at boot, never cached elsewhere
        ssl_cert_reqs=None,               # self-signed certificate (see trade-off above)
        max_connections=pool_size,        # hard cap per process
        health_check_interval=30,         # PING idle connections before reuse
        socket_keepalive=True,
        socket_timeout=2.0,               # fail fast instead of hanging a web request
        retry=Retry(ExponentialBackoff(cap=2.0, base=0.1), 3),
        retry_on_error=[redis.ConnectionError, redis.TimeoutError],
    )

The equivalents for the two other common stacks look like this. Sidekiq needs the TLS setting on both the server side (the worker process) and the client side (the web process that enqueues jobs):

# config/initializers/sidekiq.rb
redis_opts = { url: ENV["REDIS_URL"], ssl_params: { verify_mode: OpenSSL::SSL::VERIFY_NONE } }
Sidekiq.configure_server { |config| config.redis = redis_opts }
Sidekiq.configure_client { |config| config.redis = redis_opts }

// Node, ioredis
const Redis = require("ioredis");
const client = new Redis(process.env.REDIS_URL, { tls: { rejectUnauthorized: false } });

The server also closes idle connections. The default timeout is 300 seconds, which kills pooled connections in apps with quiet periods. The health_check_interval above handles that on the client. You can also change the server value with heroku redis:timeout <resource> --seconds <n>. Setting it to 0 disables the timeout, but then leaked connections never get reclaimed.

Budgeting connections

Each plan has a maximum number of client connections, and going over it fails new connections with errors. The budget is simple multiplication, and it is easy to get wrong because every layer multiplies:

connections = sum over process types of
              dynos x processes per dyno x (pool size per process + fixed extras)

# fixed extras: a Sidekiq process opens its own pool plus a few for heartbeats;
# a one-off "heroku run" console, release-phase tasks and redis:cli sessions also count.

Leave at least 20 percent of the plan limit free. Deploys temporarily run old and new dynos side by side when preboot is on, so connections can briefly double. Look up your plan's limit on the add-on pricing page, then compare it against connected_clients in INFO clients at your daily peak.

Memory limits and eviction policy

The default maxmemory-policy is noeviction. When the instance reaches its memory limit, every write that needs more memory fails with an out-of-memory error, and reads keep working. For a job queue that is exactly right: an error the app can see and alert on is much better than quietly dropping jobs. For a pure cache it is wrong: the cache fills up and then every cache write throws an error in the request path.

Change it per instance with heroku redis:maxmemory <resource> --policy <policy>. The policies that matter:

PolicyWhat gets evictedRight for
noevictionNothing; writes fail at the limitQueues, locks, rate-limit counters, anything authoritative
allkeys-lruLeast recently used key, any keyA dedicated cache
allkeys-lfuLeast frequently used keyCaches with a stable hot set and scan-like noise
volatile-lruLRU among keys with a TTL onlyMixed use where only TTL keys are disposable

The dangerous combination is a shared instance with an allkeys-* policy: under memory pressure the store will evict a Sidekiq queue list just as happily as a cached page fragment. volatile-lru relies on every cache key having a TTL; one code path that forgets it brings back the write errors. The robust answer is two instances.

Worked example: cache and queue for a Rails app

Take a Rails app with 6 Standard-2X web dynos running 2 Puma processes of 5 threads each, plus 3 worker dynos running 1 Sidekiq process with concurrency 10. It uses Redis for page caching, sessions and Sidekiq.

Step one is to split by job. Provision two instances: the existing one stays on REDIS_URL for Sidekiq and sessions with noeviction, and a second one is attached as the cache:

heroku addons:create heroku-redis:premium-0 --as CACHE -a example-app
heroku redis:maxmemory <cache-resource-name> --policy allkeys-lru -a example-app   # name from redis:info
heroku redis:info -a example-app          # confirm both instances, versions and policies

With --as CACHE the cache instance's URL is exposed as CACHE_URL. Point config.cache_store at ENV["CACHE_URL"] with the same VERIFY_NONE SSL setting, and Sidekiq keeps REDIS_URL. Step two is the connection budget for the queue instance. The web processes each hold a Sidekiq client pool; with a pool of 5 that is 6 x 2 x 5 = 60. The workers need concurrency plus a few extra connections for Sidekiq's internal threads; assuming 5 extras, that is 3 x 1 x (10 + 5) = 45. Add 5 for consoles and release tasks and the steady state is 110, which doubles to about 220 during a preboot deploy. If the plan you chose allows fewer than roughly 275 (so that 220 is at most 80 percent of the limit), either shrink the web-side pools (enqueueing rarely needs 5 connections per process) or move up a plan. Step three is memory. Sessions and queues are small and bounded; the cache will always grow to fill whatever memory it has, which is why it gets its own instance and its own LRU policy.

Maintenance, credentials and migration

Heroku patches every instance at least once every 90 days. On production plans maintenance runs as a failover to the standby, so expect a few minutes of errors and a new release. heroku data:maintenances:info shows the current window and any scheduled run. Mini plans cannot choose a window or run maintenance on demand. On the others, move the window to your quietest hour and, when a maintenance is scheduled, consider running it yourself at a time you are watching:

heroku data:maintenances:info REDIS_URL -a example-app
heroku data:maintenances:window:update REDIS_URL Sunday 14:30 -a example-app
heroku data:maintenances:run REDIS_URL -a example-app        # fails over to the standby now
heroku redis:credentials REDIS_URL --reset -a example-app    # rotate the password on your schedule

Credential rotation and maintenance both end with the same thing your code already handles if you followed the connection rules: a new config value and a restart. To migrate data, create a production plan as a fork of an existing instance by passing -- --fork <url> to addons:create.

Observing the store

Three tools cover most investigations. heroku redis:info shows each instance's plan, version, maintenance status and settings. heroku redis:cli opens a session against the instance, where INFO memory, INFO stats and INFO clients give the numbers that matter: used_memory against maxmemory, evicted_keys, keyspace_hits and keyspace_misses for the cache hit rate, rejected_connections and connected_clients. SLOWLOG GET 10 finds commands blocking the single command thread, which is usually a KEYS * or a large SMEMBERS in application code.

Use heroku redis:stats-reset before a load test so the counters cover only the test. Keyspace notifications are off by default; enable them with heroku redis:keyspace-notifications <resource> --config <flags> only if a consumer really needs expiry or change events, because they add CPU cost to every write. Alert on memory above 80 percent of the limit, on any rejected_connections, and on a falling hit rate.

Failure modes and trade-offs

SymptomCauseFix
SSL certificate verify failed on first connectClient verifies the self-signed certificateSet the client's verify option as shown above
OOM command not allowed on cache writesCache on the default noevictionDedicated cache instance with allkeys-lru
Jobs vanish under loadQueue shares an instance with an allkeys-* policySeparate instances; noeviction for queues
Errors for minutes after a maintenanceURL cached outside the app, or no reconnectRead REDIS_URL at boot; enable retries
max number of clients reached during deploysPools sized for steady state onlyBudget for preboot doubling; shrink web pools
Connection reset after quiet periods300-second idle timeoutHealth checks or keepalive on the client
All data gone after a restartMini plan has no persistenceUse a production plan for anything you cannot rebuild

Two broader trade-offs remain. First, since February 2026 Heroku has been in a sustaining engineering model, as described in the dynos article: the service is maintained, but you should not plan around new features. Second, the add-on gives you no control over replicas, clustering or placement. If you need read replicas, a sharded cluster or more memory than the largest plan, compare it with ElastiCache or a self-managed Valkey. Because the protocol is standard, moving is mostly a matter of creating the new store and changing a URL.

What to do next

  1. Run heroku redis:info and write down the plan, version and maxmemory policy of every instance. Any Valkey 7.2 instance needs an upgrade before 19 December 2026.
  2. Make sure no queue or session store is on a Mini plan.
  3. Split cache and queue into separate instances. Set the cache to allkeys-lru and keep the queue on noeviction.
  4. Grep the code base and CI for hard-coded rediss:// strings and replace them with the config var.
  5. Add retries, socket timeouts and health checks to every client, then rehearse with heroku data:maintenances:run on staging under load.
  6. Write down the connection budget for each process type, including preboot doubling, and compare it with the plan limit.
  7. Set the maintenance window to your quietest hour and schedule credential rotation with redis:credentials --reset.
  8. Alert on memory above 80 percent, rejected connections, evictions on the queue instance and cache hit rate.
Key takeaway: Heroku Redis is now Heroku Key-Value Store running Valkey, provisioned and managed exactly as before. Mini keeps nothing across a restart; production plans write the append-only file every second and fail over to a standby, which can change the URL and restart your dynos. Read REDIS_URL at boot, accept the self-signed certificate knowingly, add retries and health checks, give caches their own instance with an LRU policy, keep queues on noeviction, and budget connections for deploys as well as steady state.