A WebSocket connection will drop. Phones switch from Wi-Fi to cellular, laptops sleep, NAT devices forget idle mappings, load balancers enforce idle timeouts, and you deploy new servers. The browser's WebSocket object does nothing about any of this. When the connection ends it fires close, and that is the last you hear from it. Reconnection is entirely your code, and it is where most real-time products quietly lose messages or overload their own servers.
This article builds reconnection from first principles: how to tell that a connection is dead, how long to wait before trying again, how to pick up where you left off without losing or duplicating messages, and how to keep a fleet of clients from knocking over your servers when they all reconnect at once. It includes a complete client state machine in TypeScript, a server-side resume buffer, and a worked example of a reconnect storm with numbers.
Detecting that the connection is gone
There are two kinds of death. A clean close arrives as a close frame with a status code, and the browser fires close promptly. A silent death happens when the network path disappears: no frame arrives, TCP may not notice for many minutes, and the socket sits in OPEN while nothing gets through. This half-open state is the common case on mobile networks.
The browser API exposes no ping frames, so a browser client cannot rely on protocol-level pings it can observe. Use an application heartbeat instead: the server sends a small message every N seconds, and the client treats a silence of roughly two to three intervals as death. The design of those intervals is covered in heartbeat and keepalive strategies. Two browser signals help too: the online event says the network may be back, and visibilitychange tells you the user returned to a tab whose timers may have been throttled.
| Close code | Meaning | Reconnect policy |
|---|---|---|
| 1000 | Normal closure | Do not reconnect unless the app expects a permanent connection |
| 1001 | Going away (page unload, server shutdown) | Reconnect with normal backoff |
| 1006 | Abnormal closure; reserved, never sent on the wire | Network failure or failed handshake; back off, consider an auth check |
| 1008 | Policy violation | Do not retry blindly; surface to the user or re-authenticate |
| 1011 | Server internal error | Reconnect with backoff |
| 1012 | Service restart (IANA-registered) | Reconnect after a randomised delay of several seconds |
| 1013 | Try again later, overload (IANA-registered) | Back off longer than usual |
| 4000-4999 | Private use, defined by your application | Your own semantics, such as token expired or kicked by another session |
Note that 1006 is something your client observes, never something a server sends. Also note that the browser's close() method accepts only 1000 or codes from 3000 to 4999, so client-initiated closes must use those.
How long to wait: backoff with full jitter
Retrying immediately wastes battery when the server is down, and after a restart every client retrying at once creates a burst that may knock it down again. Exponential backoff solves the first problem; randomness, called jitter, solves the second.
The variant to default to is full jitter, described in Marc Brooker's 2015 AWS Architecture Blog post on exponential backoff: on attempt n, wait a uniformly random time between zero and min(cap, base * 2^n). Plain exponential backoff without jitter keeps clients synchronised: everyone who disconnected together retries together, at 1 s, 2 s, 4 s and so on. Full jitter spreads them across the whole window.
function backoffMs(attempt: number, baseMs = 500, capMs = 30_000): number {
const ceiling = Math.min(capMs, baseMs * 2 ** attempt);
return Math.random() * ceiling; // full jitter: uniform in [0, ceiling)
}
// attempt 0: up to 0.5 s, 1: up to 1 s, 2: up to 2 s ... 6 and later: up to 30 sTwo adjustments matter in practice. Reset the attempt counter only after the server confirms the session, not when the socket opens, or a server that accepts and immediately drops connections causes a tight loop. And add a floor for codes that carry a hint: a 1012 restart deserves a random delay of several seconds even on the first attempt, and 1013 deserves a longer one.
The client state machine
The implementation below follows that diagram. It fetches a fresh URL (with a short-lived token) on every dial, resumes from the last sequence number it processed, keeps an outbox of unacknowledged sends, and ignores events from superseded sockets so two connections can never be live at once.
type Opts = {
url: () => Promise<string>; // returns wss://... with a fresh short-lived token
probe: () => Promise<boolean>; // HTTP check: false means "log in again"
baseMs?: number; capMs?: number; idleMs?: number;
};
export class ReconnectingSocket {
private ws?: WebSocket;
private attempt = 0;
private failuresBeforeOpen = 0;
private lastSeq = 0;
private session?: string;
private outbox: { id: string; body: unknown }[] = [];
private idle?: ReturnType<typeof setTimeout>;
private retry?: ReturnType<typeof setTimeout>;
private stopped = false;
constructor(private o: Opts, private deliver: (m: any) => void) {
addEventListener("online", () => this.kick());
document.addEventListener("visibilitychange", () => {
if (document.visibilityState === "visible") this.kick();
});
}
start() { this.stopped = false; this.dial(); }
stop() { this.stopped = true; clearTimeout(this.retry); this.ws?.close(1000); }
send(body: unknown) {
const msg = { id: crypto.randomUUID(), body }; // id lets the server drop duplicates
this.outbox.push(msg);
if (this.ws?.readyState === WebSocket.OPEN) this.ws.send(JSON.stringify({ t: "msg", ...msg }));
}
private kick() { // network or tab came back: retry now
if (this.stopped || this.ws) return;
clearTimeout(this.retry);
this.attempt = 0;
this.dial();
}
private async dial() {
let url: string;
try { url = await this.o.url(); } catch { return this.schedule(0); }
const ws = new WebSocket(url);
this.ws = ws;
ws.onopen = () => {
ws.send(JSON.stringify({ t: "resume", session: this.session, after: this.lastSeq }));
this.armIdle(ws);
};
ws.onmessage = (ev) => {
if (ws !== this.ws) return; // a superseded socket must not deliver
this.armIdle(ws);
const m = JSON.parse(ev.data);
if (m.t === "welcome") {
this.attempt = 0; this.failuresBeforeOpen = 0; this.session = m.session;
if (m.resync) this.lastSeq = m.seq; // server will send a snapshot
for (const x of this.outbox) ws.send(JSON.stringify({ t: "msg", ...x }));
} else if (m.t === "ack") {
this.outbox = this.outbox.filter((x) => x.id !== m.id);
} else if (m.t === "snapshot") {
this.deliver(m);
} else if (m.t === "event" && m.seq > this.lastSeq) {
this.lastSeq = m.seq; // duplicates after resume are skipped
this.deliver(m);
}
};
ws.onclose = (ev) => {
if (ws !== this.ws) return;
this.ws = undefined;
clearTimeout(this.idle);
if (!this.stopped) this.schedule(ev.code);
};
}
private armIdle(ws: WebSocket) { // silence means a half-open connection
clearTimeout(this.idle);
this.idle = setTimeout(() => {
if (ws !== this.ws) return;
this.ws = undefined;
ws.close(4000, "idle"); // may never complete on a dead path
this.schedule(4000);
}, this.o.idleMs ?? 45_000);
}
private async schedule(code: number) {
if (code === 1006 && ++this.failuresBeforeOpen >= 3 && !(await this.o.probe())) {
this.stopped = true; // hand control to the login flow
return;
}
const floor = code === 1012 ? 5_000 : code === 1013 ? 15_000 : 0;
const ceiling = Math.min(this.o.capMs ?? 30_000, (this.o.baseMs ?? 500) * 2 ** this.attempt);
this.attempt = Math.min(this.attempt + 1, 16);
this.retry = setTimeout(() => this.dial(), floor + Math.random() * ceiling);
}
}
Resuming without loss or duplicates
Reconnecting the socket is half the job. The other half is the gap: events published while the client was away. The server numbers every event in a session with a sequence number and keeps the most recent ones in a bounded buffer. On reconnect the client says which session it had and the last sequence it processed. If the buffer still covers that point, the server replays the missing events. If not, it starts a fresh session and sends a snapshot of current state. Ordering guarantees across this boundary are discussed in message ordering in bidirectional streams.
import collections, uuid
class Session:
def __init__(self, user, maxlen=2000):
self.sid, self.user, self.seq = uuid.uuid4().hex, user, 0
self.buf = collections.deque(maxlen=maxlen) # (seq, payload)
def publish(self, payload):
self.seq += 1
self.buf.append((self.seq, payload))
return self.seq
async def on_resume(conn, sessions, sid, after):
s = sessions.get(sid)
covered = s is not None and s.user == conn.user and (
not s.buf or after >= s.buf[0][0] - 1)
if not covered:
s = sessions.create(conn.user)
await conn.send({"t": "welcome", "session": s.sid, "seq": s.seq, "resync": True})
await conn.send({"t": "snapshot", "state": await load_state(conn.user)})
return s
await conn.send({"t": "welcome", "session": s.sid, "seq": s.seq, "resync": False})
for seq, payload in list(s.buf):
if seq > after:
await conn.send({"t": "event", "seq": seq, **payload})
return sThree details make this correct. The session check includes the user, so a stolen session id cannot replay someone else's stream. The client skips any event at or below its last sequence, so replays are harmless. And client sends carry an id that the server records, so an outbox replay after reconnect does not double-post a chat message or double-submit an order. With several servers, the buffer cannot live in one process: keep it in a shared store such as a Redis stream keyed by session, or route a session back to its node, which load balancing WebSockets covers.
When auth failure looks like a network error
If the handshake is rejected with an HTTP 401 or 403, the browser does not tell your code. You get an error event with no detail and a close with 1006, exactly as if the network had failed. A client that only backs off will retry an expired token forever, every 30 seconds, from every open tab.
The fix has two parts. Fetch a fresh token before every dial, as the url() callback above does, so expiry during a long disconnect is handled by default. And after a few consecutive failures that never reached open, make a plain HTTP request to an authenticated endpoint. HTTP gives you a status code. If it says 401, stop reconnecting and send the user through login. For tokens that expire while connected, have the server close with an application code in the 4000 range so the client knows to refresh rather than back off.
Server side: preventing reconnect storms
Clients that behave well are necessary but not sufficient. The server controls the biggest storm trigger, which is its own deploys. Drain a node before stopping it: stop accepting new connections, then close existing ones with 1012 spread over a window rather than all at once, and only then exit. Pair this with admission control: cap the rate of new handshakes per node and answer the excess with a fast rejection, so overload is cheap instead of cascading. WebSocket scaling patterns covers the rest of the fleet design.
Worked example. Suppose 200,000 clients spread over 10 nodes, 20,000 each, and assume one node can complete about 2,000 TLS handshakes plus authentication per second. These are planning assumptions; measure your own.
- Rolling restart, no jitter. One node stops and its 20,000 clients retry after a fixed 1 s. The other nine nodes receive 20,000 handshakes within a fraction of a second: roughly 2,200 each, all queued at once. Clients with a 5 s connect timeout start timing out and retrying, adding load.
- Rolling restart, close with 1012, 5 to 30 s random delay. 20,000 reconnects over 25 s is 800 per second, about 90 per surviving node. Nothing notices.
- Whole-region network blip, full jitter, 30 s cap. 200,000 clients spread over up to 30 s is about 6,700 per second, 670 per node, well under capacity. Without jitter the first retry wave would be 200,000 handshakes at once, 100 seconds of work for the fleet.
The arithmetic is simple enough to put in a design review: connections to move, divided by the spread window, compared with handshake capacity per node.
Failure modes
- Two live sockets. A new connection opens before the old one's
closefires, and both deliver. Tag handlers with the socket they belong to and ignore stale ones, as the client does. - Retrying permanent errors. Policy violations, banned users and removed endpoints are not transient. Stop and surface them.
- Many tabs, many sockets. Ten open tabs mean ten connections and ten reconnect loops. Share one connection through a SharedWorker or BroadcastChannel when that matters.
- Throttled timers. Background tabs delay timers, so heartbeats and backoff stretch. Re-check the socket on
visibilitychange. - Lost subscriptions. A reconnected socket starts with no server-side subscriptions. Make the resume message re-establish them, or the client is connected and silently receiving nothing.
- Unbounded outbox. A client offline for an hour can queue thousands of messages. Cap it and tell the user what was not sent.
What to measure
Track reconnects per minute per node, time from disconnect to confirmed session at p50 and p99, the share of resumes served from the buffer against those needing a snapshot, close codes by count, and handshake rejections from admission control. Raise the resume buffer size if snapshot fallbacks are frequent. A spike in reconnects after a deploy with no matching spike in errors means draining works; a spike in both means it does not.
What to do next
- Read the reconnection code you ship today and check whether it resets backoff on open or on server confirmation.
- Replace fixed or plain exponential delays with full jitter and a cap of 30 to 60 seconds.
- Add an application heartbeat and an idle timeout of two to three heartbeat intervals.
- Number server events per session and add a bounded resume buffer with a snapshot fallback.
- Add message ids to client sends and deduplicate them on the server.
- Make deploys drain with 1012 spread over a window, and add a per-node handshake rate limit.
- Add the HTTP auth probe so expired tokens lead to login instead of an endless loop.