A WebSocket connection will drop. Phones switch from Wi-Fi to cellular, laptops sleep, NAT devices forget idle mappings, load balancers enforce idle timeouts, and you deploy new servers. The browser's WebSocket object does nothing about any of this. When the connection ends it fires close, and that is the last you hear from it. Reconnection is entirely your code, and it is where most real-time products quietly lose messages or overload their own servers.

This article builds reconnection from first principles: how to tell that a connection is dead, how long to wait before trying again, how to pick up where you left off without losing or duplicating messages, and how to keep a fleet of clients from knocking over your servers when they all reconnect at once. It includes a complete client state machine in TypeScript, a server-side resume buffer, and a worked example of a reconnect storm with numbers.

Advertisement

Detecting that the connection is gone

There are two kinds of death. A clean close arrives as a close frame with a status code, and the browser fires close promptly. A silent death happens when the network path disappears: no frame arrives, TCP may not notice for many minutes, and the socket sits in OPEN while nothing gets through. This half-open state is the common case on mobile networks.

The browser API exposes no ping frames, so a browser client cannot rely on protocol-level pings it can observe. Use an application heartbeat instead: the server sends a small message every N seconds, and the client treats a silence of roughly two to three intervals as death. The design of those intervals is covered in heartbeat and keepalive strategies. Two browser signals help too: the online event says the network may be back, and visibilitychange tells you the user returned to a tab whose timers may have been throttled.

Close codeMeaningReconnect policy
1000Normal closureDo not reconnect unless the app expects a permanent connection
1001Going away (page unload, server shutdown)Reconnect with normal backoff
1006Abnormal closure; reserved, never sent on the wireNetwork failure or failed handshake; back off, consider an auth check
1008Policy violationDo not retry blindly; surface to the user or re-authenticate
1011Server internal errorReconnect with backoff
1012Service restart (IANA-registered)Reconnect after a randomised delay of several seconds
1013Try again later, overload (IANA-registered)Back off longer than usual
4000-4999Private use, defined by your applicationYour own semantics, such as token expired or kicked by another session

Note that 1006 is something your client observes, never something a server sends. Also note that the browser's close() method accepts only 1000 or codes from 3000 to 4999, so client-initiated closes must use those.

How long to wait: backoff with full jitter

Retrying immediately wastes battery when the server is down, and after a restart every client retrying at once creates a burst that may knock it down again. Exponential backoff solves the first problem; randomness, called jitter, solves the second.

The variant to default to is full jitter, described in Marc Brooker's 2015 AWS Architecture Blog post on exponential backoff: on attempt n, wait a uniformly random time between zero and min(cap, base * 2^n). Plain exponential backoff without jitter keeps clients synchronised: everyone who disconnected together retries together, at 1 s, 2 s, 4 s and so on. Full jitter spreads them across the whole window.

function backoffMs(attempt: number, baseMs = 500, capMs = 30_000): number {
  const ceiling = Math.min(capMs, baseMs * 2 ** attempt);
  return Math.random() * ceiling;          // full jitter: uniform in [0, ceiling)
}
// attempt 0: up to 0.5 s, 1: up to 1 s, 2: up to 2 s ... 6 and later: up to 30 s

Two adjustments matter in practice. Reset the attempt counter only after the server confirms the session, not when the socket opens, or a server that accepts and immediately drops connections causes a tight loop. And add a floor for codes that carry a hint: a 1012 restart deserves a random delay of several seconds even on the first attempt, and 1013 deserves a longer one.

Advertisement

The client state machine

Client reconnection state machineCONNECTINGfresh token, dialRESUMINGsend session + last seqOPENdeliver, ack, heartbeatBACKOFFfull-jitter timerAUTH CHECKHTTP probe, refreshCLOSEDuser stoppedonopenwelcomeclose / idle timeouterrorN failures, never openedtoken oktimer firesstop()Shortcuts into CONNECTINGonline event, tab visible, user retry: reset attemptBackoff resets only after the server confirms the session, not on onopen,so a server that accepts and immediately drops connections cannot cause a tight loop.
The client moves from CONNECTING through RESUMING to OPEN. Every failure goes through BACKOFF; repeated failures before ever opening route through an HTTP auth check; network and visibility events short-circuit the timer.

The implementation below follows that diagram. It fetches a fresh URL (with a short-lived token) on every dial, resumes from the last sequence number it processed, keeps an outbox of unacknowledged sends, and ignores events from superseded sockets so two connections can never be live at once.

type Opts = {
  url: () => Promise<string>;          // returns wss://... with a fresh short-lived token
  probe: () => Promise<boolean>;       // HTTP check: false means "log in again"
  baseMs?: number; capMs?: number; idleMs?: number;
};

export class ReconnectingSocket {
  private ws?: WebSocket;
  private attempt = 0;
  private failuresBeforeOpen = 0;
  private lastSeq = 0;
  private session?: string;
  private outbox: { id: string; body: unknown }[] = [];
  private idle?: ReturnType<typeof setTimeout>;
  private retry?: ReturnType<typeof setTimeout>;
  private stopped = false;

  constructor(private o: Opts, private deliver: (m: any) => void) {
    addEventListener("online", () => this.kick());
    document.addEventListener("visibilitychange", () => {
      if (document.visibilityState === "visible") this.kick();
    });
  }

  start() { this.stopped = false; this.dial(); }
  stop() { this.stopped = true; clearTimeout(this.retry); this.ws?.close(1000); }

  send(body: unknown) {
    const msg = { id: crypto.randomUUID(), body };   // id lets the server drop duplicates
    this.outbox.push(msg);
    if (this.ws?.readyState === WebSocket.OPEN) this.ws.send(JSON.stringify({ t: "msg", ...msg }));
  }

  private kick() {                      // network or tab came back: retry now
    if (this.stopped || this.ws) return;
    clearTimeout(this.retry);
    this.attempt = 0;
    this.dial();
  }

  private async dial() {
    let url: string;
    try { url = await this.o.url(); } catch { return this.schedule(0); }
    const ws = new WebSocket(url);
    this.ws = ws;
    ws.onopen = () => {
      ws.send(JSON.stringify({ t: "resume", session: this.session, after: this.lastSeq }));
      this.armIdle(ws);
    };
    ws.onmessage = (ev) => {
      if (ws !== this.ws) return;       // a superseded socket must not deliver
      this.armIdle(ws);
      const m = JSON.parse(ev.data);
      if (m.t === "welcome") {
        this.attempt = 0; this.failuresBeforeOpen = 0; this.session = m.session;
        if (m.resync) this.lastSeq = m.seq;          // server will send a snapshot
        for (const x of this.outbox) ws.send(JSON.stringify({ t: "msg", ...x }));
      } else if (m.t === "ack") {
        this.outbox = this.outbox.filter((x) => x.id !== m.id);
      } else if (m.t === "snapshot") {
        this.deliver(m);
      } else if (m.t === "event" && m.seq > this.lastSeq) {
        this.lastSeq = m.seq;                        // duplicates after resume are skipped
        this.deliver(m);
      }
    };
    ws.onclose = (ev) => {
      if (ws !== this.ws) return;
      this.ws = undefined;
      clearTimeout(this.idle);
      if (!this.stopped) this.schedule(ev.code);
    };
  }

  private armIdle(ws: WebSocket) {      // silence means a half-open connection
    clearTimeout(this.idle);
    this.idle = setTimeout(() => {
      if (ws !== this.ws) return;
      this.ws = undefined;
      ws.close(4000, "idle");          // may never complete on a dead path
      this.schedule(4000);
    }, this.o.idleMs ?? 45_000);
  }

  private async schedule(code: number) {
    if (code === 1006 && ++this.failuresBeforeOpen >= 3 && !(await this.o.probe())) {
      this.stopped = true;              // hand control to the login flow
      return;
    }
    const floor = code === 1012 ? 5_000 : code === 1013 ? 15_000 : 0;
    const ceiling = Math.min(this.o.capMs ?? 30_000, (this.o.baseMs ?? 500) * 2 ** this.attempt);
    this.attempt = Math.min(this.attempt + 1, 16);
    this.retry = setTimeout(() => this.dial(), floor + Math.random() * ceiling);
  }
}

Resuming without loss or duplicates

Reconnecting the socket is half the job. The other half is the gap: events published while the client was away. The server numbers every event in a session with a sequence number and keeps the most recent ones in a bounded buffer. On reconnect the client says which session it had and the last sequence it processed. If the buffer still covers that point, the server replays the missing events. If not, it starts a fresh session and sends a snapshot of current state. Ordering guarantees across this boundary are discussed in message ordering in bidirectional streams.

import collections, uuid

class Session:
    def __init__(self, user, maxlen=2000):
        self.sid, self.user, self.seq = uuid.uuid4().hex, user, 0
        self.buf = collections.deque(maxlen=maxlen)      # (seq, payload)

    def publish(self, payload):
        self.seq += 1
        self.buf.append((self.seq, payload))
        return self.seq

async def on_resume(conn, sessions, sid, after):
    s = sessions.get(sid)
    covered = s is not None and s.user == conn.user and (
        not s.buf or after >= s.buf[0][0] - 1)
    if not covered:
        s = sessions.create(conn.user)
        await conn.send({"t": "welcome", "session": s.sid, "seq": s.seq, "resync": True})
        await conn.send({"t": "snapshot", "state": await load_state(conn.user)})
        return s
    await conn.send({"t": "welcome", "session": s.sid, "seq": s.seq, "resync": False})
    for seq, payload in list(s.buf):
        if seq > after:
            await conn.send({"t": "event", "seq": seq, **payload})
    return s

Three details make this correct. The session check includes the user, so a stolen session id cannot replay someone else's stream. The client skips any event at or below its last sequence, so replays are harmless. And client sends carry an id that the server records, so an outbox replay after reconnect does not double-post a chat message or double-submit an order. With several servers, the buffer cannot live in one process: keep it in a shared store such as a Redis stream keyed by session, or route a session back to its node, which load balancing WebSockets covers.

When auth failure looks like a network error

If the handshake is rejected with an HTTP 401 or 403, the browser does not tell your code. You get an error event with no detail and a close with 1006, exactly as if the network had failed. A client that only backs off will retry an expired token forever, every 30 seconds, from every open tab.

The fix has two parts. Fetch a fresh token before every dial, as the url() callback above does, so expiry during a long disconnect is handled by default. And after a few consecutive failures that never reached open, make a plain HTTP request to an authenticated endpoint. HTTP gives you a status code. If it says 401, stop reconnecting and send the user through login. For tokens that expire while connected, have the server close with an application code in the 4000 range so the client knows to refresh rather than back off.

Server side: preventing reconnect storms

Clients that behave well are necessary but not sufficient. The server controls the biggest storm trigger, which is its own deploys. Drain a node before stopping it: stop accepting new connections, then close existing ones with 1012 spread over a window rather than all at once, and only then exit. Pair this with admission control: cap the rate of new handshakes per node and answer the excess with a fast rejection, so overload is cheap instead of cascading. WebSocket scaling patterns covers the rest of the fleet design.

Worked example. Suppose 200,000 clients spread over 10 nodes, 20,000 each, and assume one node can complete about 2,000 TLS handshakes plus authentication per second. These are planning assumptions; measure your own.

  • Rolling restart, no jitter. One node stops and its 20,000 clients retry after a fixed 1 s. The other nine nodes receive 20,000 handshakes within a fraction of a second: roughly 2,200 each, all queued at once. Clients with a 5 s connect timeout start timing out and retrying, adding load.
  • Rolling restart, close with 1012, 5 to 30 s random delay. 20,000 reconnects over 25 s is 800 per second, about 90 per surviving node. Nothing notices.
  • Whole-region network blip, full jitter, 30 s cap. 200,000 clients spread over up to 30 s is about 6,700 per second, 670 per node, well under capacity. Without jitter the first retry wave would be 200,000 handshakes at once, 100 seconds of work for the fleet.

The arithmetic is simple enough to put in a design review: connections to move, divided by the spread window, compared with handshake capacity per node.

Failure modes

  • Two live sockets. A new connection opens before the old one's close fires, and both deliver. Tag handlers with the socket they belong to and ignore stale ones, as the client does.
  • Retrying permanent errors. Policy violations, banned users and removed endpoints are not transient. Stop and surface them.
  • Many tabs, many sockets. Ten open tabs mean ten connections and ten reconnect loops. Share one connection through a SharedWorker or BroadcastChannel when that matters.
  • Throttled timers. Background tabs delay timers, so heartbeats and backoff stretch. Re-check the socket on visibilitychange.
  • Lost subscriptions. A reconnected socket starts with no server-side subscriptions. Make the resume message re-establish them, or the client is connected and silently receiving nothing.
  • Unbounded outbox. A client offline for an hour can queue thousands of messages. Cap it and tell the user what was not sent.

What to measure

Track reconnects per minute per node, time from disconnect to confirmed session at p50 and p99, the share of resumes served from the buffer against those needing a snapshot, close codes by count, and handshake rejections from admission control. Raise the resume buffer size if snapshot fallbacks are frequent. A spike in reconnects after a deploy with no matching spike in errors means draining works; a spike in both means it does not.

What to do next

  1. Read the reconnection code you ship today and check whether it resets backoff on open or on server confirmation.
  2. Replace fixed or plain exponential delays with full jitter and a cap of 30 to 60 seconds.
  3. Add an application heartbeat and an idle timeout of two to three heartbeat intervals.
  4. Number server events per session and add a bounded resume buffer with a snapshot fallback.
  5. Add message ids to client sends and deduplicate them on the server.
  6. Make deploys drain with 1012 spread over a window, and add a per-node handshake rate limit.
  7. Add the HTTP auth probe so expired tokens lead to login instead of an endless loop.
Key takeaway: The browser gives you a close event and nothing else, so reconnection is application code. Detect silent deaths with application heartbeats, honour close codes, and wait using exponential backoff with full jitter, resetting only once the server confirms the session. Resume through per-session sequence numbers, a bounded replay buffer and a snapshot fallback, and deduplicate client sends with ids. Probe auth over HTTP because handshake rejections look like network failures. On the server, drain deploys with spread 1012 closes and rate-limit handshakes so a fleet of reconnecting clients never becomes an outage.