WebRTC defines how two endpoints exchange media once they know about each other, and deliberately says nothing about how they find each other. That job is signaling: carrying session descriptions, network candidates and room membership between clients that cannot yet talk directly. Every real-time product builds its own, and most first versions are a WebSocket server that relays JSON blobs between sockets.

That version works in a demo and fails in production in predictable ways: messages lost across a deploy, a reconnecting client whose old socket keeps sending stale offers, a duplicated join that leaves a ghost participant, a ten-thousand-person webinar that knocks the servers over at the start time. This article treats signaling as what it is, a small distributed system with its own consistency requirements, and walks through the message envelope, ordering and idempotency, routing between workers, reconnect and resync, scaling for join storms, and the HTTP-only alternatives WHIP and WHEP.

Signaling as a distributed system, separate from the media planeClient AlicePeerConnectionClient BobPeerConnectionRoom serviceHTTP: token, TURN credsjoin tokenEdge worker 1WebSocket, statelessEdge worker 2WebSocket, statelessenvelopePub/subroom and member channelsSession storerooms, epochs, outboxesSFU controlsubscribe, layersSFU mediaRTP, SRTPmedia path (dashed) never touches the signaling workersSignaling carries offers, answers, candidates and room state through stateless workers.Epochs, sequence numbers and outboxes make reconnects and duplicates harmless.
Stateless workers route signaling through pub/sub and a session store; once connected, media flows on its own path and outlives signaling outages.

What signaling has to carry

Signaling traffic is small and bursty, but every message matters. Losing a media packet costs a few milliseconds of concealment; losing an answer costs the call.

MessagePurposeRequirement
join, leaveRoom membershipIdempotent, authenticated
offer, answerSDP session descriptionsOrdered per peer pair, exactly-once effect
candidateTrickled ICE candidatesOrdered after the description they belong to
roster deltaWho is present and publishingReplayable after reconnect
subscribeWhich streams and layers to receive from an SFUAuthorised, auditable
mute, kick, endModeration and controlAuthorised, acknowledged

Note what is not on the list: media. Once ICE and DTLS succeed, the media path is peer to peer or peer to SFU and owes the signaling plane nothing. A signaling outage should prevent new calls and renegotiation, not end the calls already running. That independence is the first architectural goal.

What signaling has to carry

Signaling traffic is small and bursty, but every message matters. Losing a media packet costs a few milliseconds of concealment; losing an answer costs the call.

MessagePurposeRequirement
join, leaveRoom membershipIdempotent, authenticated
offer, answerSDP session descriptionsOrdered per peer pair, exactly-once effect
candidateTrickled ICE candidatesOrdered after the description they belong to
roster deltaWho is present and publishingReplayable after reconnect
subscribeWhich streams and layers to receive from an SFUAuthorised, auditable
mute, kick, endModeration and controlAuthorised, acknowledged

Note what is not on the list: media. Once ICE and DTLS succeed, the media path is peer to peer or peer to SFU and owes the signaling plane nothing. A signaling outage should prevent new calls and renegotiation, not end the calls already running. That independence is the first architectural goal.

The architecture

The production shape has five parts. A room service over plain HTTP issues short-lived join tokens and TURN credentials; it is cacheable, rate-limited and easy to scale. Edge workers terminate WebSockets and hold no authoritative state, so any client can reconnect to any worker. A session store such as Redis holds rooms, members, per-member epochs and short message outboxes, with TTLs so crashed clients age out. A pub/sub layer routes messages between workers, because two peers in one room usually land on different workers. Where an SFU is used, an SFU control service turns subscription intent into forwarding decisions.

Keep signaling and media processes separate. Media servers run hot continuously and drain slowly; signaling workers are mostly idle and should deploy daily. Coupling them means media load starves control traffic and a signaling deploy drops media. How to spread the WebSocket tier itself is covered in horizontal WebSocket scaling.

The message envelope

Every message, in both directions, travels in the same envelope. The fields are what make the rest of the design possible:

{
  "v": 1,
  "id": "01J9Z6Q4K8F2...",        // unique per message: idempotency key
  "type": "offer",
  "room": "r-8812",
  "from": "alice",
  "to": "bob",                     // omitted for room broadcasts
  "epoch": 3,                      // sender's session epoch, bumped on every reconnect
  "seq": 41,                       // per-sender, per-epoch sequence number
  "ack": 57,                       // highest seq received from the server on this session
  "body": { "sdp": "v=0..." }
}

The id lets the server drop duplicates from client retries. epoch identifies which incarnation of a client sent the message. seq restores order and reveals gaps. ack tells the server how much of its outbound stream the client has seen, so it can trim the outbox and know what to replay. Version the envelope from day one; clients in the field will lag the server by months.

Epochs, ordering and idempotency

Epochs solve the zombie problem. When Bob's network flaps, his client opens a new socket while the old one may still be half-alive on another worker, flushing buffered messages. The session store holds Bob's current epoch; the new connection increments it; any message arriving with an older epoch is dropped. Without this, a late offer from the dead socket can overwrite the fresh negotiation and leave both sides with mismatched descriptions.

Sequence numbers handle order within an epoch, and the id handles exactly-once effects. A worker's inbound handler applies all three checks before doing anything else:

async function onClientMessage(conn, msg) {
  const key = `member:${msg.room}:${msg.from}`;
  const current = Number(await redis.hget(key, "epoch"));
  if (msg.epoch < current) return;                          // zombie connection: drop
  const last = Number(await redis.hget(key, "inSeq") ?? 0);  // reset to 0 when the epoch is bumped
  if (msg.seq <= last || await redis.exists(`seen:${msg.id}`)) {
    return conn.send({ type: "ack", id: msg.id });          // duplicate: re-ack, do not re-apply
  }
  if (msg.seq > last + 1) {
    return conn.send({ type: "resync", from: last + 1 });   // gap: ask client to resend
  }
  if (!authorised(conn.claims, msg)) return conn.send({ type: "error", id: msg.id, code: 403 });
  await route(msg);                                         // publish to member or room channel
  await redis.hset(key, "inSeq", msg.seq);                  // record only after routing
  await redis.set(`seen:${msg.id}`, 1, "EX", 300);
  conn.send({ type: "ack", id: msg.id });
}

Joins, leaves and subscription changes are written as set operations on the store, so applying one twice has the same effect as applying it once. Ordering within a single conversation is the same problem as in any messaging system; message ordering goes deeper.

Routing between workers

Routing chooses between room channels, which every worker with a member in the room subscribes to, and member channels, addressed to one client's current worker. Directed messages such as offers go to member channels; roster changes go to room channels. Either way the delivery semantics of the backbone leak into the protocol. Redis pub/sub is at-most-once: a worker that is restarting or briefly disconnected simply misses messages. NATS core behaves the same way. Durable options such as Redis Streams, NATS JetStream or Kafka keep messages for replay, at the cost of latency and storage.

A pragmatic design uses fast at-most-once pub/sub for delivery and a short per-member outbox in the store for recovery: every message routed to Bob is also appended to his outbox with its server sequence number, capped by length and time. Normal delivery is fast; anything lost is recovered from the outbox when Bob's client reports a gap or reconnects.

Reconnect and resync

Reconnect is the normal case, not the exception: phones change networks, laptops sleep, workers deploy. On reconnect the client presents a resume token, its room, and the highest server sequence it acknowledged. The worker bumps the epoch, resets the inbound sequence counter, re-subscribes the member channel, and replays from the outbox. If the outbox no longer reaches back that far, it sends a full snapshot (roster, publications, pending negotiation state) instead of deltas. Client-side backoff and jitter are covered in reconnection strategies.

Signaling recovery and media recovery are separate decisions. If only the WebSocket dropped, media is still flowing and nothing else needs to happen. If the network path itself changed, the client calls pc.restartIce() once signaling is back, which triggers a new offer with fresh ICE credentials and a new candidate exchange without tearing down tracks.

ws.onopen = () => ws.send(envelope("resume", { token, lastAck: state.serverSeq }));
ws.onmessage = ({ data }) => {
  const m = JSON.parse(data);
  if (m.type === "ack" || m.type === "error") return settle(m);  // control replies carry no seq
  if (m.type === "resync") return resendFrom(m.from);
  if (m.type === "snapshot") return applySnapshot(m.body);      // outbox too short
  if (m.seq !== state.serverSeq + 1) return ws.send(envelope("resync", { from: state.serverSeq + 1 }));
  state.serverSeq = m.seq;
  handle(m);
};
pc.oniceconnectionstatechange = () => {
  if (pc.iceConnectionState === "failed" && ws.readyState === WebSocket.OPEN) pc.restartIce();
};

Glare and perfect negotiation

Offers can cross: both peers add a track at the same moment and each sends an offer. The established answer is the perfect negotiation pattern, where one peer is polite and rolls back its own offer on a collision while the impolite peer ignores the incoming one. It needs nothing from the server except that the two roles are assigned deterministically, for example by comparing member ids, and that candidates are delivered in order after their description. The client-side code is walked through in WebRTC architecture.

Join storms and scale

Signaling load is spiky. A scheduled webinar with ten thousand attendees produces ten thousand joins in the first minute, and if each join broadcasts a full roster, the traffic grows with the square of the audience. Defences, in order of effect:

  • Send roster deltas, not full rosters, and for large rooms send only publishers; viewers fetch counts.
  • Let clients fetch join tokens and TURN credentials ahead of the start time, over cacheable HTTP.
  • Spread joins with a small random delay, and admit through a token bucket per room.
  • Shard rooms across store partitions by room id so one hot room cannot saturate the whole store.
  • Drain workers on deploy: stop accepting, tell clients to reconnect over a spread of seconds, then exit.

HTTP-only signaling with WHIP and WHEP

Not every use case needs a persistent socket. For broadcast ingest, WHIP (published as RFC 9725 in 2025) reduces signaling to HTTP: the publisher POSTs an SDP offer to an endpoint, receives the answer in the response and a resource URL in the Location header, may PATCH trickled candidates or an ICE restart to that resource, and DELETEs it to stop. WHEP applies the same pattern to playback; as of mid-2026 it is still an IETF draft, so expect details to change. For one-way streaming these remove the WebSocket tier entirely. For interactive rooms with presence, moderation and renegotiation, a bidirectional channel is still needed.

Worked example: a deploy in the middle of a call

Alice and Bob join room r-8812 on workers 1 and 2. Alice sends an offer, epoch 1, seq 4. Worker 1 checks the epoch, records the id, publishes to Bob's member channel and appends it to his outbox as server seq 17. Worker 2 delivers it. Bob answers, and candidates trickle both ways; the call connects.

Worker 2 is then redeployed. Bob's socket closes; his media keeps flowing. His client reconnects to worker 1 with lastAck 19. Worker 1 bumps Bob's epoch to 2 and replays seq 20 and 21 from the outbox, a roster delta and a candidate published while he was away. Meanwhile his old connection's final buffered message, a duplicate candidate tagged epoch 1, arrives through a lingering proxy and is dropped. Nothing was lost, nothing applied twice, and the users noticed nothing.

Failure modes

  • Ghost participants. Members without TTL heartbeats stay in the roster after a crash; see presence tracking.
  • Lost answers across deploys. At-most-once pub/sub with no outbox means a stuck spinner.
  • Zombie offers. No epochs, so a dead socket's messages corrupt a fresh negotiation.
  • Candidates before descriptions. Out-of-order delivery causes addIceCandidate errors; queue them.
  • Store as a single point of failure. Run it replicated and decide what workers do when it is down.
  • Expired tokens on reconnect. Accept a separate resume token so a long call can recover.

Trade-offs

ChoiceGainCost
At-most-once pub/sub plus outboxLow latency, bounded recoveryTwo mechanisms to maintain
Durable log backboneReplay is built inHigher latency and operations cost
Stateless workersBoring deploys and failoverEvery message touches the store
WHIP or WHEP over HTTPNo socket tier for broadcastNo presence or server push
Signaling inside the SFUOne less serviceMedia load and deploys hit control traffic

What to do next

  1. Write down your envelope and add id, epoch, seq and ack if any are missing.
  2. Make join, leave and subscribe idempotent set operations with TTLs.
  3. Decide your backbone semantics and add a per-member outbox if delivery is at-most-once.
  4. Implement resume with replay and a snapshot fallback, and test it by killing workers under load.
  5. Separate signaling and media processes, and drain workers on deploy.
  6. Load test a scheduled join of your largest expected room with roster deltas.
  7. Use WHIP for any broadcast ingest path instead of a custom socket protocol.
Key takeaway: Signaling is a small distributed system whose job is to make sure no offer, answer or candidate is lost, duplicated or applied out of order, across reconnects and deploys. Stateless workers, a session store, an envelope with ids, epochs and sequence numbers, and an outbox for replay are what turn a demo relay into something that survives production.