Load balancing HTTP is forgiving. Every request is a fresh decision, so if one server gets too much traffic for a minute, the next minute's requests go elsewhere and the imbalance heals itself. A WebSocket breaks that assumption. The balancer makes one decision when the client opens the connection, and that decision stands for as long as the socket lives, which may be hours. Whatever the balancer got wrong at handshake time, it stays wrong.

This article works through what that changes. It covers where the balancing decision happens, what a proxy must do to pass the handshake, which algorithms suit long-lived connections, the idle timeouts that silently kill sockets, how to drain a server for a deploy, how to rebalance after adding capacity, and how to keep a mass reconnect from becoming an outage. The overall WebSocket server architecture is covered in WebSocket architecture, in depth; this page is about the tier in front of it.

Advertisement

The one-decision problem

A WebSocket starts life as an HTTP/1.1 request carrying Upgrade: websocket and Connection: Upgrade headers, defined in RFC 6455. The server answers 101 Switching Protocols, and from then on the same TCP connection carries WebSocket frames in both directions. Every frame on that connection reaches the server the balancer picked for the handshake.

Three consequences follow. Load is measured in connections and the messages they carry, not requests per second, so a server can be overloaded while its request rate looks idle. Imbalance is sticky: a server that came up late holds few connections until clients happen to reconnect. And every operational action that touches a server, a deploy, a scale-in, a crash, disconnects real users who must then reconnect somewhere else, all at once.

WebSocket load balancing: the decision is made once, at the handshake, and lives for hoursClientsbrowsers, appsL4 balancerTCP, no HTTP viewL7 proxy poolTLS, Upgrade, routingWS server A12,400 connectionsWS server B12,100 connectionsWS server C (new)300 connectionsWS server D (draining)close 1001, no newTCPleast-connBackplanepub/sub between serversPer-request balancingHTTP: every request re-decides, imbalance heals itselfPer-connection balancingWS: imbalance persists until clients reconnectConnection counts are illustrative; they show the post-deploy skew a new server sees
The balancer picks a server once per connection. After a scale-out, the new server C holds a fraction of the connections of A and B until something moves clients, and a draining server D must close its connections so they reconnect elsewhere.

Layer 4 or layer 7

A layer 4 balancer forwards TCP connections without reading them. A WebSocket passes through untouched because the balancer never sees the Upgrade. It is cheap and protocol-agnostic, but it cannot route on path, cookie or header, cannot terminate TLS for you in most configurations, and its health checks can only prove a port is open.

A layer 7 proxy terminates the client connection, reads the HTTP handshake and opens its own connection to a backend. It can route /ws/chat and /ws/market to different pools, authenticate before a socket reaches the application, and pick a backend by user or room. The price is that the proxy holds two sockets and some buffer memory for every live connection, and it must correctly handle the Upgrade, which is where most first deployments fail.

Many production systems use both: an L4 balancer spreads TCP across a pool of L7 proxies, and the proxies make the routing decision. The L4 tier is easy to scale and rarely changes; the L7 tier holds the intelligence.

Advertisement

Passing the handshake through a proxy

Upgrade and Connection are hop-by-hop headers, so a standards-following HTTP proxy strips them before forwarding. A proxy that does that turns a WebSocket handshake into a plain GET and the backend never upgrades. In nginx the handshake must be forwarded explicitly, over HTTP/1.1:

map $http_upgrade $connection_upgrade {
    default upgrade;
    ''      close;
}

upstream ws_backend {
    zone ws_backend 64k;      # share connection counts across worker processes
    least_conn;
    server 10.0.1.11:8080 max_fails=3 fail_timeout=10s;
    server 10.0.1.12:8080 max_fails=3 fail_timeout=10s;
    server 10.0.1.13:8080 max_fails=3 fail_timeout=10s;
}

server {
    listen 443 ssl;
    location /ws/ {
        proxy_pass http://ws_backend;
        proxy_http_version 1.1;
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection $connection_upgrade;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_read_timeout 120s;    # default is 60s; must exceed the heartbeat interval
        proxy_send_timeout 120s;
    }
}

The map sends Connection: upgrade only when the client asked for one, so the same location can serve ordinary requests. The equivalent in HAProxy is to set timeout tunnel, which governs a connection once it has been upgraded, separately from the HTTP-phase timeouts; leave it unset and the shorter client or server timeout applies to the tunnel.

Choosing an algorithm for long-lived connections

AlgorithmHow it behaves for WebSocketsUse when
Round-robinEqual handshakes per server, regardless of how many connections each still holdsServers are identical and were all started together; rarely true for long
Least connectionsSends each new handshake to the server with the fewest open connections, so new servers fill firstThe default choice for a stateless WebSocket tier
Consistent hash on user or roomSame key lands on the same server; minimal remapping when the pool changesServers hold per-room state in memory and a backplane would be too costly
Cookie or session affinityReturning client reconnects to the server it used beforeA reconnect must find buffered state on its old server

Least connections is the right default because it is the only common algorithm that looks at the quantity that actually matters. With round-robin, a server added during peak gets one handshake in N, even though it holds zero connections, so it stays nearly idle for hours. Least connections sends it nearly all new handshakes until it catches up.

Least connections has one trap. Counts are local: in nginx they are per worker process unless the upstream has a shared zone, and across a pool each proxy instance counts only its own connections. That is usually close enough. It also counts connections, not work: a server holding 10,000 idle dashboard sockets and one holding 10,000 busy trading sockets look equal. If message rates vary that much, route the different traffic classes to separate pools rather than trying to weight one pool.

Hashing and affinity trade balance for locality. Use them only when the application keeps state that would be expensive to rebuild. The more common design keeps servers stateless with respect to routing and publishes messages through a backplane, which is covered in real-time fanout.

Idle timeouts: the silent killer

Every hop between client and server may close a connection it thinks is idle: the client's corporate proxy, a NAT gateway, the cloud balancer, your L7 proxy. On AWS the Application Load Balancer's idle timeout defaults to 60 seconds; nginx's proxy_read_timeout also defaults to 60 seconds. A chat socket that sits quiet for a minute dies, and the client sees an abnormal closure with no explanation.

There are two fixes and you want both. Raise the timeouts on the hops you control to comfortably more than your heartbeat interval. Then send application or protocol-level pings often enough that no hop you do not control ever sees a quiet connection. A heartbeat every 25 to 30 seconds is a common choice because it stays under the many 60-second defaults with margin. The design of the heartbeat itself, and how to detect dead peers with it, is in heartbeats and keepalive.

Draining a server without dropping users

To deploy, scale in or replace a server you must move its connections elsewhere. Dropping them all at once forces every client to reconnect in the same second. Draining spreads that out:

  1. Mark the server as draining so the balancer stops sending it new handshakes. Most balancers support this through a deregistration or weight-zero state; in your own proxy it is a health check that starts failing on purpose.
  2. Keep existing connections open while you close them gradually from the server side, a few percent per second, with close code 1001, going away, which RFC 6455 defines for a server that is going down.
  3. Clients reconnect, the balancer places them on healthy servers, and the draining server's count falls toward zero.
  4. After a deadline, close whatever remains and stop the process.
import asyncio, random

async def drain(server, duration_s=120, close_code=1001):
    server.accepting = False                 # health check now reports draining
    conns = list(server.connections)
    random.shuffle(conns)
    if not conns:
        return
    interval = duration_s / len(conns)
    for ws in conns:
        if not ws.closed:
            await ws.close(code=close_code, reason="server draining")
        await asyncio.sleep(interval)

The balancer's own drain window must be at least as long as yours, or it will cut the connections before you finish closing them politely. Watch for one specific nginx behaviour: a configuration reload starts new worker processes, but the old workers keep running until their open connections close. With long-lived WebSockets that can take days, so set worker_shutdown_timeout to bound how long old workers live after a reload.

Rebalancing after scale-out

Draining handles removal. Addition is the subtler problem. Suppose three servers hold 12,000 connections each and you add a fourth. Least connections will route every new handshake to it, but if only 2 percent of connections turn over per minute, it takes most of an hour to approach parity, and in the meantime the old servers are the ones that triggered the scale-out.

The fix is deliberate shedding. Each server compares its own connection count with the pool average, published through the backplane or a shared store, and if it is more than some margin above average it closes a small number of connections per second until it is not. Clients reconnect, least connections sends them to the new server, and the pool converges in minutes instead of an hour. Pick the margin, say 10 percent, and the rate, say 0.5 percent of connections per second, so that the reconnect rate stays well inside what the authentication and session tiers can absorb. Shedding to protect a server under acute overload is a related but separate mechanism, covered in load shedding on bidirectional streams.

Reconnect storms

When a balancer, a zone or a whole server pool fails, every client disconnects at the same moment and every client's reconnect logic fires at the same moment. The handshake path, which is the most expensive part of a connection because it includes TLS and authentication, receives a spike equal to your entire connection count. If it fails under that load, clients retry, and the retries are synchronised too.

Clients must reconnect with exponential backoff and full jitter, and the first retry must be delayed by a random amount, not attempted immediately:

function reconnectDelayMs(attempt, baseMs = 500, capMs = 30000) {
  const ceiling = Math.min(capMs, baseMs * 2 ** attempt);
  return Math.random() * ceiling;      // full jitter: anywhere in [0, ceiling)
}

On the server side, rate-limit handshakes per proxy instance and return a fast rejection rather than queueing, so that clients back off instead of timing out. A server under a reconnect storm that accepts every connection slowly is worse than one that rejects half of them quickly.

WebSockets over HTTP/2

RFC 8441 defines how to bootstrap a WebSocket over an HTTP/2 stream using an extended CONNECT method, so several WebSockets can share one TCP connection with each other and with ordinary requests. For load balancing this cuts both ways. Fewer TCP connections means less handshake cost and less per-connection memory in the proxy. But a balancer that picks a backend per TCP connection will now place every WebSocket a browser opens to your origin on the same backend. Check whether your proxy balances per stream or per connection before you enable it, and check that every hop supports the extended CONNECT; where one does not, browsers fall back to HTTP/1.1.

Failure modes

  • Stripped Upgrade headers. The handshake reaches the backend as a plain GET; the client sees a 400 or a 200 instead of 101.
  • Sixty-second deaths. Quiet connections close at exactly the default idle timeout of some hop; the regularity is the giveaway.
  • Round-robin skew. New servers sit nearly empty for hours while old ones run hot.
  • Thundering reconnect. A deploy closes all connections at once and the handshake tier falls over.
  • Immortal old workers. Proxy reloads accumulate old processes holding long-lived sockets and their memory.
  • Health checks that lie. An L4 port check passes while the application rejects every upgrade.

What to do next

  1. Confirm every hop forwards the Upgrade and that a test client gets 101 through the full production path, not only against a backend directly.
  2. List the idle timeout of every hop, raise the ones you control, and set the heartbeat interval below the smallest.
  3. Switch the WebSocket pool to least connections unless you have a measured need for hashing or affinity.
  4. Implement gradual server-side draining with close code 1001 and align the balancer's drain window and the proxy's worker shutdown timeout with it.
  5. Add above-average shedding so scale-out converges in minutes, and alert on the ratio of the busiest server to the pool average.
  6. Ship jittered exponential backoff in every client, and load-test a full reconnect of your peak connection count before you need it.
Key takeaway: A WebSocket balancer decides once per connection and lives with the decision for hours, so the work moves from the balancing algorithm to the lifecycle around it. Forward the Upgrade, keep every hop's idle timeout above your heartbeat, prefer least connections, drain gradually with close code 1001, shed deliberately after scale-out so new servers fill, and make every client reconnect with full jitter so a failure does not become a storm.