Server-Sent Events look stateless: a client makes one GET request and the server writes text/event-stream lines into a response that never ends. One server handles this easily. The trouble starts with the second server, because SSE's best feature, automatic reconnect with Last-Event-ID, makes promises the cluster has to keep: the client expects to resume exactly where it left off, on whichever node it reaches next.

This guide is about those stateful parts. It assumes you know the wire format, which Server-Sent Events, an introduction covers, and builds a horizontally scaled design: stateless edge nodes, a shared replayable log, and the operational rules for rebalancing, shedding load, proxies and deploys.

Advertisement

Why SSE scale-out is a state problem

Each open stream holds a socket, a buffer and a position: the ID of the last event the client received. When the connection drops, the browser's EventSource waits the reconnection time and reconnects, sending the last ID it saw in the Last-Event-ID header. A correct server sends everything after that ID and then continues live.

With one node, "everything after that ID" can live in memory. With many nodes behind a load balancer, the reconnect may land on a node that never saw the earlier events. There are two ways out: sticky sessions, which pin a client to one node and lose its history when that node dies, or a shared log every node can read from any position. Only the second survives node failure and deploys, so it is the design this page builds.

SSE scale-out: stateless edge nodes, one shared replayable logBrowsersEventSourceMobile / CLISSE clientsLoad balancerleast-conn, no bufferingSSE node A40k streamsSSE node B40k streamsSSE node C (draining)retry: jitterShared logRedis Streams / KafkaXREADProducersappend, get IDXADDEvent id = log ID, so any node can resumea reconnecting client from Last-Event-ID.No sticky sessions: a reconnect may land on any node and still resumes exactly where it left off.
Figure 1. Producers append to a shared log and get back an ID; SSE nodes read from the log and use that ID as the SSE id field. Any node can serve any reconnect.

Event IDs must come from the log

The id: field is an opaque string to the browser, but it must mean the same thing on every node. Per-node counters fail immediately: node A's event 1,041 is not node B's. Timestamps fail under concurrency and clock skew. The robust choice is the position the log assigns when the event is appended: a Redis Streams entry ID, or a Kafka partition and offset encoded together.

Redis Streams fits SSE closely. XADD appends an entry and returns an ID that increases within the stream; XREAD returns entries after a given ID and can block until new ones arrive. A producer looks like this:

async def publish(room, payload):
    # MAXLEN ~ caps memory; the cap is also your resume window.
    return await r.xadd(f"room:{room}", {"data": json.dumps(payload)},
                        maxlen=100_000, approximate=True)

The MAXLEN cap is also your resume window. A client that reconnects with an ID older than the oldest retained entry cannot be resumed exactly; detect that and send an explicit reset event so the client refetches state over a normal API call, rather than silently skipping events.

Advertisement

A resumable SSE node

The node below serves one stream per room. It resumes from Last-Event-ID if present, falls back to a query parameter for clients that start fresh with a known position, and otherwise starts at the live tail. It pins the tail to a concrete ID once instead of passing Redis's $ to every read, because $ re-resolves on each call and would skip entries appended between two reads.

# aiohttp + redis-py asyncio: one SSE endpoint that resumes from a shared Redis stream.
import asyncio, json, random
from aiohttp import web
import redis.asyncio as redis

r = redis.Redis(host="redis.internal")
HEARTBEAT_S = 15
MAX_AGE_S = 30 * 60          # recycle long connections so load rebalances

async def events(request):
    room = request.match_info["room"]
    key = f"room:{room}"
    last = request.headers.get("Last-Event-ID") or request.query.get("since")
    if not last:                             # live tail: pin the current newest ID once, never re-read "$"
        newest = await r.xrevrange(key, count=1)
        last = newest[0][0].decode() if newest else "0-0"
    resp = web.StreamResponse(headers={
        "Content-Type": "text/event-stream",
        "Cache-Control": "no-cache",
        "X-Accel-Buffering": "no",          # tell nginx not to buffer this response
    })
    await resp.prepare(request)
    await resp.write(f"retry: {random.randint(2000, 8000)}\n\n".encode())
    deadline = asyncio.get_running_loop().time() + MAX_AGE_S
    while asyncio.get_running_loop().time() < deadline and not request.app["draining"]:
        batch = await r.xread({key: last}, block=HEARTBEAT_S * 1000, count=100)
        if not batch:
            await resp.write(b": hb\n\n")    # comment line keeps proxies from idling out
            continue
        for _stream, entries in batch:
            for entry_id, fields in entries:
                last = entry_id.decode()
                data = fields[b"data"].decode()
                await resp.write(f"id: {last}\ndata: {data}\n\n".encode())
    return resp                              # ending the response makes the browser reconnect

Four details matter. Each event's id: is the stream entry ID, so resume works on any node. A comment line every 15 seconds keeps idle connections alive through proxies and lets the server notice dead sockets when a write fails. The first thing written is a retry: field with a random value, which the specification defines as the new reconnection time in milliseconds; randomising it spreads out reconnects. And the loop ends after a maximum age, which is how load rebalances, as discussed below.

The sample issues one blocking XREAD per client only to stay readable. At scale that means one Redis connection per viewer, and a single node with tens of thousands of streams would exceed Redis's default maxclients of 10,000 on its own. Production nodes run one reader task per active room and copy each entry into local per-client queues. A per-connection read is used only for the catch-up replay from Last-Event-ID, after which the client joins the room's live queue.

Why plain pub/sub is not enough

Many first designs use Redis Pub/Sub or a similar broadcast channel: producers publish, every node subscribes and forwards to its clients. That handles fan-out, but pub/sub delivers only to subscribers connected at that moment and keeps no history. A client that was reconnecting during a publish misses the event, and there is nothing to replay from Last-Event-ID.

You can combine the two, using pub/sub for live delivery and a separate store for replay, but then you must merge two sources without gaps or duplicates during the switch-over. A log that serves both replay and live tail, as Streams or Kafka do, removes that race. Fan-out cost still matters: if every node reads every room, each event is read once per node. Fan-out architecture covers interest-based subscription, where a node reads only the rooms its clients are in.

Load balancing and rebalancing

Because any node can serve any client, no stickiness is needed. Use least-connections balancing: round-robin counts requests, but SSE load is open connections, and a node that received the last thousand reconnects is far busier than its request count suggests.

Long-lived connections create a rebalancing problem. When you add nodes during a traffic rise, existing connections stay where they are; the new nodes receive only new connections, and the hot nodes stay hot for hours. The fix is a maximum connection age. The server ends each stream after, say, 30 minutes plus jitter, the browser reconnects after its reconnection time, resumes from its last ID, and the balancer places it on the least-loaded node. Clients lose nothing, and load converges within one maximum age. Load balancing WebSockets discusses the same problem for WebSockets, where the client must implement reconnect itself.

Reconnect storms and shedding load safely

When a node dies, every client on it reconnects within roughly one reconnection time. Forty thousand clients arriving within a few seconds is a storm, and each reconnect also triggers a replay read from the log. Randomised retry: values spread the arrivals; replay reads with a COUNT limit keep each catch-up bounded.

Shedding needs care, because the specification is strict: if the response status is not 200, or the content type is not text/event-stream, the browser fails the connection and does not attempt to reconnect. A 503 meant as "try later" therefore disconnects the client for good, until page code creates a new EventSource. A 204 is the documented way to tell a client to stop reconnecting. To ask a client to come back later, answer 200 with a long, jittered retry: and end the response:

async def events_guarded(request):
    app = request.app
    if app["open_streams"] >= app["max_streams"]:
        # NOT 503: a non-200 makes EventSource fail permanently.
        resp = web.StreamResponse(headers={"Content-Type": "text/event-stream"})
        await resp.prepare(request)
        await resp.write(f"retry: {random.randint(5000, 30000)}\n\n".encode())
        return resp                          # client waits the jittered delay, then retries
    app["open_streams"] += 1
    try:
        return await events(request)
    finally:
        app["open_streams"] -= 1

Native clients and libraries outside the browser may implement different retry rules, so test each client type you support. Load shedding for real-time connections covers choosing the threshold.

Proxies, buffering and timeouts

Most SSE outages in production are caused by something between client and server. Reverse proxies buffer responses by default, so events sit in a buffer until it fills; idle timeouts close streams that have been quiet; and some corporate proxies buffer entire responses. A typical nginx block:

location /events/ {
    proxy_pass http://sse_nodes;
    proxy_http_version 1.1;
    proxy_set_header Connection "";
    proxy_buffering off;          # or rely on X-Accel-Buffering: no from the app
    proxy_cache off;
    proxy_read_timeout 1h;        # must exceed the heartbeat interval by a wide margin
}

Every hop needs the same treatment: CDN, cloud load balancer, ingress controller and service mesh sidecar each have their own idle timeout, and the shortest one wins. The heartbeat interval must be comfortably below the smallest of them. Over HTTP/1.1, browsers also limit concurrent connections per origin, so several tabs each holding a stream can starve the page's other requests; serving SSE over HTTP/2, where streams share one connection, removes the problem.

Deploys and draining

A deploy restarts every node, which is a planned reconnect storm. Drain instead. Mark the node as draining so the load balancer stops sending new connections, then end existing streams in batches over a minute or two, after writing a short jittered retry: so clients return promptly but not together. Because resume is log-based, clients reconnect to other nodes without losing events. Set the platform's termination grace period longer than the drain, or the platform kills the process mid-drain.

Worked example: sizing for 200,000 viewers

A live sports page expects 200,000 concurrent viewers across 300 match rooms, with around 20 events per second per room at peak. Suppose a load test, run as in WebSocket scaling patterns, shows one node holding 40,000 idle streams at modest CPU, with memory per stream in the tens of kilobytes. These are your measurements, not universal constants; kernel buffers, TLS and the framework change them a lot.

Five nodes cover the steady state with no headroom. Losing one node sends 40,000 clients to the remaining four, pushing each to 50,000, so plan for seven nodes and keep each under about 30,000 at peak. Outbound volume is the real cost: 300 rooms times 20 events per second times an average of 667 viewers per room is about 4 million event writes per second across the cluster, which is why event size, compression at the edge and interest-based fan-out matter more than connection count. The log sees only 6,000 appends per second, plus one read stream per node per room.

Failure modes

  • Per-node IDs. Resume on a different node skips or repeats events. Use log IDs.
  • Pub/sub only. Events published during a reconnect are lost with no way to replay.
  • 503 for overload. Browsers stop reconnecting for good. Use 200 with a long jittered retry.
  • Buffering proxy. Events arrive in bursts or not at all. Disable buffering on every hop and test through the real path.
  • Timeout shorter than heartbeat. Streams drop at a regular interval, and each drop causes a replay.
  • No maximum age. New nodes stay empty after scale-out while old ones overload.
  • Retention shorter than outages. Clients return with IDs that have been trimmed; send an explicit reset and refetch.

Trade-offs

A shared log adds a dependency and a hop of latency, and its retention bounds how long a client can be away. In return any node can serve any client, deploys lose nothing, and rebalancing is just ending a connection. Sticky sessions avoid the log but tie each client's history to one process. If clients also need to send frequent messages upstream, compare WebSockets; if they only receive, SSE over HTTP/2 with a log behind it is usually simpler to run. Heartbeat tuning is covered in Heartbeats for real-time connections.

What to do next

  1. Make every SSE event's id the position assigned by a shared log, not a per-node counter or timestamp.
  2. Implement resume from Last-Event-ID, and send an explicit reset event when the ID is older than retention.
  3. Send a randomised retry value at the start of each stream and before every server-initiated close.
  4. Replace any non-200 overload response on the SSE endpoint with a 200 plus long jittered retry.
  5. Add a maximum connection age with jitter and switch the balancer to least-connections.
  6. Audit every proxy hop for buffering and idle timeouts, and set the heartbeat well below the smallest timeout.
  7. Load-test one node to find its real stream capacity, then size the cluster to survive losing a node at peak.
Key takeaway: Scaling SSE means keeping the reconnect promise across nodes. Take event IDs from a shared, replayable log such as Redis Streams or Kafka so any node can resume any client from Last-Event-ID, keep nodes stateless and balance by open connections, recycle connections on a jittered maximum age so load rebalances, shed load with 200 plus a long retry because a non-200 permanently stops EventSource, disable buffering and outlast idle timeouts on every hop, and drain on deploy. Size for losing a node at peak, and measure capacity rather than trusting anyone's numbers.