WebSocket gives a browser and a server one long-lived, full-duplex connection over which either side can send at any time. That makes chat, collaborative editing, live dashboards, multiplayer state and agent streaming straightforward on the client. On the server it changes almost everything: connections last hours instead of milliseconds, each one holds state, load balancers and proxies must be told to leave them alone, and a deploy becomes a mass disconnection event.

This article is about the architecture around the socket rather than any single mechanism. It covers what the protocol provides, the path from client to server, how to split the system into tiers, the per-connection loop, authentication, fan-out, slow consumers, deploys and a worked sizing example. Heartbeats, backpressure, compression and fan-out each have their own deep dive linked below.

Advertisement

What the protocol gives you

RFC 6455 defines WebSocket in two parts. The opening handshake is an HTTP/1.1 GET with Upgrade: websocket, Connection: Upgrade and a random Sec-WebSocket-Key; the server answers 101 Switching Protocols with a Sec-WebSocket-Accept derived from the key. After that the TCP connection carries frames, not HTTP. RFC 8441 later defined bootstrapping WebSockets over HTTP/2 with extended CONNECT, and RFC 9220 did the same for HTTP/3, but HTTP/1.1 upgrade remains the path every proxy and client supports.

Frames are small and simple. Each has an opcode (text, binary, continuation, close, ping, pong), a FIN bit for fragmentation and a length. Frames from client to server must be masked; frames from server to client must not be. Control frames carry at most 125 bytes of payload and are never fragmented. A close frame carries a status code: 1000 normal, 1001 going away, 1008 policy violation, 1009 message too big, 1011 internal error, and from the IANA registry 1012 service restart and 1013 try again later. Codes 1005 and 1006 are reserved for reporting locally that no code was received or that the connection dropped abnormally; they must never be sent.

What the protocol does not give you matters as much: no message acknowledgements, no ordering or delivery guarantees beyond those of a single TCP connection, no resume after reconnect, no request and response correlation, and no backpressure signal in the browser API beyond the bufferedAmount counter. All of that is your application protocol's job.

The reference architecture

Split the system into a stateful edge tier and a stateless application tier. Edge nodes terminate WebSocket connections, authenticate them, run heartbeats, enforce per-connection limits and translate frames into calls on application services. Application services hold the business logic and know nothing about sockets. When a service needs to push something to a user, it publishes an event to a topic on a backplane such as Redis pub/sub, NATS or Kafka; each edge node subscribes to the topics its connected users need and delivers events on the right sockets.

WebSocket at scale: a stateful edge tier holds sockets, a stateless app tier holds logic, a backplane joins themClientsbrowser, mobileLoad balancerL4 or L7, idle timeoutEdge node Asockets, auth, heartbeatEdge node Bsockets, auth, heartbeatBackplanepub/sub by topicApp servicesstateless, HTTP/gRPCSession registryuser to edge nodewsspublishcommandsInbound frames become calls to app services; outbound events are published to topics and delivered by whichever edge node holds the socket.
Reference architecture. Only the edge tier holds sockets; everything behind it can be deployed and scaled like any HTTP service.

This split lets you deploy business logic many times a day without dropping a single connection, scale the edge tier on connection count and the application tier on request rate, and keep the hard, long-lived state in one small, rarely changed component. A session registry that maps each user to the edge node and connection ID holding their socket is optional; you need it only for targeted delivery that cannot be expressed as a topic, and it must tolerate being stale.

Advertisement

Getting the upgrade through the path

Every hop between client and edge node has to allow the upgrade and then leave the connection open. Upgrade and Connection are hop-by-hop headers, so a reverse proxy strips them unless told to forward them. The proxy must speak HTTP/1.1 to the upstream. And every hop has an idle timeout that will cut a quiet connection: nginx's proxy_read_timeout defaults to 60 seconds, and cloud load balancers have their own idle timeouts, often also 60 seconds by default.

# nginx in front of the edge tier. Upgrade is hop-by-hop, so the proxy must forward it explicitly.
map $http_upgrade $connection_upgrade {
    default upgrade;
    ''      close;
}

server {
    listen 443 ssl;
    location /ws {
        proxy_pass http://edge_nodes;
        proxy_http_version 1.1;                     # upgrade needs HTTP/1.1 upstream
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection $connection_upgrade;
        proxy_set_header Host $host;
        proxy_read_timeout 120s;                    # default 60s; must exceed heartbeat interval
        proxy_send_timeout 120s;
    }
}

The rule that follows is simple: the heartbeat interval must be shorter than the smallest idle timeout anywhere on the path, including corporate proxies you do not control. 20 to 30 seconds is a common choice. Load balancing also works differently. A WebSocket balancer distributes connections, not messages, so a node that restarts sheds its connections to the survivors and a new node receives almost nothing until clients reconnect. Balance on least connections rather than round robin, and do not rely on sticky sessions for correctness: after a reconnect the client may land anywhere.

The per-connection loop

Each connection needs three concurrent activities: a reader that receives frames, validates them and dispatches them; a writer that drains a bounded outbound queue onto the socket; and a heartbeat that detects dead peers. A half-open TCP connection, where the client vanished without a FIN, looks perfectly healthy to a reader that only waits for data, so heartbeats are not optional. When any of the three ends, cancel the others and clean up subscriptions and registry entries.

# Illustrative per-connection handler (Python asyncio, library-agnostic).
SEND_QUEUE_MAX = 256           # messages, not bytes; tune per product
HEARTBEAT_EVERY = 25           # seconds, below every idle timeout on the path

async def handle(ws, principal):
    outq = asyncio.Queue(maxsize=SEND_QUEUE_MAX)
    sub = await backplane.subscribe(topics_for(principal), sink=lambda m: offer(outq, m, ws))
    await sessions.register(principal.user_id, NODE_ID, ws.id)

    async def reader():
        async for frame in ws:                       # library answers pings with pongs
            msg = parse_and_validate(frame)          # size limit, schema, rate limit
            await app.dispatch(principal, msg)       # call a stateless service

    async def writer():
        while True:
            msg = await outq.get()
            await ws.send(encode(msg))               # awaits when the socket buffer is full

    async def heartbeat():
        while True:
            await asyncio.sleep(HEARTBEAT_EVERY)
            if not await ws.ping_and_wait(timeout=10):
                await ws.close(1011, "heartbeat timeout")   # or just drop the TCP connection

    try:
        await first_completed(reader(), writer(), heartbeat())
    finally:
        await sub.cancel()
        await sessions.unregister(principal.user_id, ws.id)

def offer(outq, msg, ws):
    try:
        outq.put_nowait(msg)
    except asyncio.QueueFull:                        # slow consumer: do not buffer forever
        ws.close_nowait(1013, "slow consumer; reconnect and resync")

Validate everything the reader receives: a maximum message size (close with 1009 when exceeded), a schema, and a per-connection rate limit. A single abusive client should cost you one connection, not one node. The heartbeat mechanics, including why browsers cannot send ping frames themselves and need an application-level heartbeat if the client must detect a dead server, are covered in heartbeat design.

Authentication and authorisation

The browser WebSocket constructor cannot set custom headers, so the usual bearer token in Authorization is not available for the handshake. The practical options are: a session cookie, which the browser sends automatically, combined with a strict check of the Origin header to stop cross-site WebSocket hijacking; a short-lived, single-use ticket fetched over an authenticated HTTP call and passed in the query string, which keeps long-lived tokens out of URLs and logs; or an unauthenticated connection whose first message must be an auth message within a few seconds. Non-browser clients can use headers normally.

Authentication at connect time is not enough for a connection that lasts all day. Tokens expire and permissions change. Record the expiry of the credential used, and either require the client to send a fresh token before it lapses or close the connection with 1008 when it does. Check authorisation per message or per subscription, not once per connection: being connected does not mean being allowed to join every topic.

Fan-out, ordering and resume

The backplane decides your delivery semantics. Plain pub/sub is at most once: an edge node that is restarting or a client that is reconnecting misses messages. If the product needs gap-free delivery, give each stream a monotonically increasing sequence number, keep a short replay buffer per topic in a log or stream store, and have the client send the last sequence it saw when it reconnects. The edge node replays from there, or tells the client to fetch a full snapshot if the gap is older than the buffer. Detailed fan-out topologies are in fan-out architecture.

// Browser reconnect with full jitter and resume-from-sequence.
let lastSeq = 0, attempt = 0;

async function connect() {
  const ws = new WebSocket(`wss://rt.example.com/ws?ticket=${await getTicket()}`);
  ws.onopen = () => { attempt = 0; ws.send(JSON.stringify({type: "resume", after: lastSeq})); };
  ws.onmessage = (e) => { const m = JSON.parse(e.data); lastSeq = m.seq; render(m); };
  ws.onclose = (e) => {
    if (e.code === 1008) return showSignIn();        // policy violation: do not loop
    const cap = Math.min(30000, 500 * 2 ** attempt++);
    setTimeout(connect, Math.random() * cap);         // full jitter spreads the herd
  };
}

Ordering holds within one connection because TCP is ordered, but only if the whole path preserves it: one publisher per topic partition, one subscriber path per edge node, and a single writer per socket.

Slow consumers and backpressure

A client on a poor mobile link reads more slowly than your backplane publishes. If the edge node buffers without limit, memory grows until the node dies and takes every healthy connection with it. The outbound queue in the loop above is bounded for that reason. When it fills, choose a policy per stream: drop and let the client resync from a sequence number, coalesce updates so only the latest state per key is kept, or close with 1013 and let the client reconnect and resume. What you must never do is block the backplane subscriber, which stalls every other connection on the node. See backpressure in bidirectional streams for the detailed policies. Compression with permessage-deflate reduces bandwidth but costs memory per connection; the trade-offs are in permessage-deflate.

Deploys and draining

Restarting an edge node drops every connection on it. Plan for that. Take the node out of the load balancer so it receives no new connections, then close existing connections gradually over a drain window, with code 1001 or 1012, rather than all at once. Clients reconnect with exponential backoff and full jitter, as in the client code above, so a node with 50,000 connections produces a spread of reconnects over tens of seconds rather than a spike that knocks over its neighbours. Deploy the edge tier rarely and in small batches; this is the main reason business logic lives elsewhere. Moving live sessions without a visible interruption is covered in connection migration.

Worked example: sizing 200,000 concurrent connections

Suppose a collaboration product expects 200,000 concurrent connections at peak, each receiving on average one message every 5 seconds of about 400 bytes, and sending one every 30 seconds, with a 25 second heartbeat.

Outbound traffic is 40,000 messages per second and about 16 MB per second before framing and TLS. Inbound is about 6,700 messages per second plus 8,000 pings per second. Memory is usually the binding limit: with TLS state, kernel socket buffers, the runtime's per-connection objects and a small outbound queue, a budget of 30 to 60 KB per idle connection is realistic for many stacks; measure yours. At 50 KB, 200,000 connections need about 10 GB across the tier. With a target of 25,000 connections per node, that is 8 nodes, and running 10 or 11 keeps you under target when one fails and its connections land on the others.

Plan for the reconnect storm too. If one node fails, 25,000 clients reconnect; with jitter spread over 30 seconds that is roughly 830 handshakes per second, each with a TLS negotiation and an auth call. Size the auth service and TLS termination for that burst, not for the steady state.

Failure modes

SymptomCauseFix
Connections drop every 60 secondsIdle timeout on a proxy or balancer below the heartbeat intervalHeartbeat below every idle timeout; raise timeouts
Handshake returns 200 or 400 instead of 101Proxy strips Upgrade and Connection, or uses HTTP/1.0 upstreamForward the headers; HTTP/1.1 upstream
Node memory climbs until it diesUnbounded outbound buffers for slow clientsBounded queues with drop, coalesce or close
Ghost users shown onlineHalf-open connections never detectedServer-side heartbeat with timeout
Neighbours fall over after one node restartsAll clients reconnect at onceDrain gradually; client backoff with full jitter
Users see events from rooms they leftAuthorisation checked only at connectAuthorise per subscription; handle token expiry
Messages missing after reconnectAt-most-once backplane, no resumeSequence numbers, replay buffer, snapshot fallback

What to do next

  1. Draw your path from client to edge node and list every idle timeout on it; set the heartbeat below the smallest.
  2. Move business logic out of the socket-holding process into stateless services behind a backplane.
  3. Give every connection a bounded outbound queue and an explicit slow-consumer policy.
  4. Pick one handshake authentication method, check Origin, and handle credential expiry on open connections.
  5. Add sequence numbers and a resume message if users must not miss events.
  6. Implement client reconnect with exponential backoff and full jitter, and stop retrying on 1008.
  7. Rehearse a drain and a node failure under load, and size auth and TLS termination for the reconnect burst.
Key takeaway: WebSocket moves state into long-lived connections, so the architecture exists to contain that state. Get the upgrade through every proxy and keep the heartbeat under every idle timeout. Hold sockets in a small, rarely deployed edge tier and keep logic in stateless services joined by a backplane. Bound every outbound queue, authorise per subscription and handle expiring credentials, add sequence numbers when delivery matters, and plan deploys and failures around jittered reconnects.