A video call looks like one thing to the user and is at least four systems to the engineer. There is a control plane that decides who is in the room, a media plane that moves encrypted packets, a set of media services such as recording, captions and phone dial-in, and the clients, which do most of the encoding, decoding and adapting. The selective forwarding unit gets most of the attention, and it deserves it, but most outages and most of the cost come from the pieces around it: where a room is placed, how rooms span regions, what each receiver actually subscribes to, and what happens when a server dies mid-meeting.
This article is about that surrounding system. Forwarding versus mixing is covered in the SFU deep dive, and ICE, congestion control and simulcast in the WebRTC pipeline article. Here we assume those and build the rest: the planes, the join path, placement and capacity, cascading, subscriptions, end-to-end encryption, participants that are really services, failure recovery and the metrics that tell you whether calls are good.
The planes and why they are separate
The control plane is a signalling service, usually a WebSocket per client, backed by a room service with state in a shared store. It authenticates users, holds the roster, carries SDP offers and answers, relays mute and hand-raise events, and decides which media server each participant uses.
The media plane is a fleet of SFUs. Each receives SRTP from publishers and forwards selected packets to subscribers. SFUs are deliberately ignorant of the product: they know tracks, layers and bandwidth estimates, not meetings, permissions or recordings. Keeping product state out of them means losing one does not lose the meeting.
Media services are everything that consumes or produces media on the server side: composite recording, live captions, phone and SIP gateways, streaming out to a CDN. The design that scales is to make each of these an ordinary participant.
The join path, step by step
Time to first frame is the metric users feel most, and it is the sum of a chain of round trips. A typical join: the client fetches a short-lived token from your application backend, opens the signalling socket, and sends a join. The room service checks the token, locks the room, picks or reuses an SFU for the client's region and returns the SFU address, TURN credentials and the roster. The client then creates a PeerConnection to that SFU, runs ICE and the DTLS handshake, publishes its microphone and camera, and subscribes to others.
To cut seconds: pre-fetch the token, return TURN credentials in the join response, start ICE gathering while signalling is in flight, publish audio first, and request a keyframe for each new video subscription so receivers do not sit on a black tile.
# Room service: join handler (simplified). State lives in a shared store, not in the SFU.
def join(room_id, user, client_region):
token = verify_token(user) # who you are, which rooms you may enter
room = rooms.get_or_create(room_id)
with room.lock(): # serialise joins per room
sfu = room.sfu_for(client_region)
if sfu is None: # first participant from this region
sfu = allocator.pick(client_region, expected_egress_mbps=room.estimate_egress())
room.attach(client_region, sfu)
if len(room.sfus) > 1:
cascade.link_all(room.sfus) # one relay path per SFU pair
seat = room.add_participant(user, sfu)
return {
"sfu_url": sfu.public_url, # client opens its PeerConnection here
"ice_servers": turn_credentials(user, client_region),
"participants": room.roster(), # tracks the client may subscribe to
"seat": seat.id,
}The per-room lock matters. Without it, two participants from a new region who join in the same instant can each allocate an SFU, and the room ends up split across two servers in one region with nobody relaying between them.
Placing rooms and sizing the fleet
The capacity unit of an SFU is not participants, it is forwarded bitrate and packets per second. Forwarding is cheap per packet but not free: each packet is decrypted from the sender's SRTP context and re-encrypted for each receiver, headers are rewritten, and RTCP is generated. Load therefore scales with the number of subscriber-track pairs, which grows roughly with the square of room size when everyone watches everyone.
A worked example makes this concrete. Take a ten-person meeting where each publisher sends three simulcast layers of about 1.5 Mbps, 500 kbps and 150 kbps, plus 40 kbps of audio. Ingress per room is about 10 x 2.19 = 21.9 Mbps. Each receiver watches one active speaker at the top layer and eight thumbnails at the bottom layer: 1.5 + 8 x 0.15 = 2.7 Mbps of video, plus three audio streams at 0.12 Mbps, so about 2.8 Mbps. Egress is 10 x 2.8 = 28 Mbps. On a server with a 10 Gbps interface planned at 60 percent, egress allows about 6,000 / 28, roughly 210 such rooms. In practice CPU often runs out first, so measure packets per second per core under a realistic room mix and set the admission threshold from that.
The allocator should place by estimated load, not by room count, and should keep headroom for rooms that grow after placement. Admission control is the last line: a server above threshold must refuse new rooms.
Cascading SFUs across regions
Pinning a whole room to one SFU is simple, but it means participants far from that server pay the long-haul round trip on every packet and every retransmission. Cascading connects each participant to a nearby SFU and links the SFUs. A track published in one region crosses the backbone once, to each other SFU that has a subscriber for it, and each SFU fans it out locally.
Remote participants get lower latency and faster loss recovery, since retransmission happens over a short local hop, and large rooms spread across machines. The cost is an extra hop for cross-region pairs and a relay protocol that must carry layer selection and bandwidth estimates as well as packets. The receiving SFU should ask the sending SFU only for the layers its local subscribers need, or the relay link carries top layers nobody watches.
Most teams start with single-SFU rooms, but let the room data model hold several SFUs per room from day one; retrofitting it is harder.
What each receiver gets: subscriptions, layers and last-N
The biggest lever on both quality and cost is sending receivers only what they render. The client knows its layout: one large speaker tile, a strip of thumbnails, a shared screen, or a gallery page. It should tell the server, and update the server whenever the layout or window size changes. The SFU then picks the smallest layer that fills each tile, within that receiver's bandwidth estimate, and pauses tracks that are off-screen.
// Client -> room service: what this client will actually render (sent on every layout change)
{
"type": "subscription",
"video": [
{"track": "alice/cam", "max_height": 720, "priority": "speaker"},
{"track": "bob/cam", "max_height": 180},
{"track": "carol/cam", "max_height": 180},
{"track": "dave/screen", "max_height": 1080, "priority": "content"}
],
"audio": "last_n",
"last_n": 3
}# SFU side: choose one simulcast layer per subscribed track within the receiver's estimate.
def choose_layers(subscription, layers_by_track, receiver_bps):
budget = receiver_bps * 0.9 # leave headroom for audio and RTCP
chosen = {t["track"]: None for t in subscription}
ordered = sorted(subscription, key=lambda t: PRIORITY[t.get("priority", "normal")])
for t in ordered: # pass 1: lowest layer for everyone
low = layers_by_track[t["track"]][0]
if low.bps <= budget:
chosen[t["track"]], budget = low, budget - low.bps
for t in ordered: # pass 2: upgrade in priority order
for layer in layers_by_track[t["track"]][1:]:
cur = chosen[t["track"]]
extra = layer.bps - (cur.bps if cur else 0)
if layer.height <= t["max_height"] and extra <= budget:
chosen[t["track"]], budget = layer, budget - extra
return chosen # None means paused: send nothingAudio uses a different rule. Decoding and mixing dozens of audio streams on a phone wastes battery and adds noise, so SFUs usually forward only the loudest few. Senders attach the audio level header extension (RFC 6464) to each packet, so the SFU can rank speakers without decoding anything. Hysteresis stops the selection flapping. Screen share should win bandwidth over thumbnails, at low frame rate and high resolution.
End-to-end encryption with SFrame
WebRTC media is always encrypted on the wire with SRTP, but the SFU terminates SRTP, so the server operator can see the media. End-to-end encryption adds a second layer that the SFU cannot remove. SFrame (RFC 9605) encrypts each encoded frame with a key the SFU never has, and the SFU keeps forwarding because it needs only RTP headers and header extensions to select layers.
In browsers the hook is the encoded transform API: RTCRtpScriptTransform runs a worker that sees each encoded frame after encoding and before packetisation, or after depacketisation and before decoding. It is available across the major engines as of 2025; older Chromium-only code used createEncodedStreams.
// main thread: attach an encoded-frame transform to every sender and receiver
const worker = new Worker("e2ee-worker.js");
sender.transform = new RTCRtpScriptTransform(worker, { side: "send" });
receiver.transform = new RTCRtpScriptTransform(worker, { side: "recv" });
// e2ee-worker.js: runs once per transform; seal/open implement SFrame with WebCrypto
onrtctransform = (event) => {
const { readable, writable, options } = event.transformer;
readable
.pipeThrough(new TransformStream({
async transform(frame, controller) {
frame.data = options.side === "send"
? await seal(frame.data, currentKeyId()) // AEAD over the payload only
: await open(frame.data); // key id read from the SFrame header
controller.enqueue(frame);
},
}))
.pipeTo(writable);
};The hard part is keys, not ciphers. Every participant needs the current key, keys must change when someone leaves so they cannot decrypt what follows, and the change must be coordinated so receivers are not stuck with frames they cannot open. Messaging Layer Security (RFC 9420) is the standard group key agreement for this, and some products use a simpler scheme where each sender distributes its own key over the signalling channel, encrypted to each member. Either way, plan for the consequences: server-side recording, captions and phone gateways cannot read end-to-end encrypted media unless they are key-holding participants, which is a product decision users must see.
Services as participants: recording, captions and phones
A composite recorder is typically a headless client that joins the room, subscribes to the tracks shown in a recording layout, renders them and encodes one file or stream. As a participant it gets adaptation, reconnection and access control for free, and it appears in the roster. Per-track recording by the SFU is cheaper but must be composed and synchronised afterwards.
Live captions work the same way: an audio-only participant forwards speech to a recognition service and publishes text over the data channel or signalling. The captions article covers caption formats and timing. A SIP gateway terminates a phone call and joins as an audio participant, usually receiving a mix.
Failure modes and how the system recovers
- SFU process or host dies. Every client on it loses media at once. Clients detect this through ICE disconnection or a signalling notice and ask the room service to rejoin; the room service allocates a new SFU and the clients republish. Keep room state outside the SFU so this is a reconnect, not a lost meeting.
- Signalling disconnect while media flows. Media keeps flowing, so do not tear down the PeerConnection; reconnect the socket and resynchronise the roster.
- Network change on the client. Wi-Fi to cellular changes the client's address. An ICE restart re-establishes the path without a full rejoin; fall back to rejoin if it fails.
- Keyframe storms. When many receivers join or recover at once, each asks the publisher for a keyframe. The SFU should coalesce these requests and rate-limit them per publisher, or the sender's bitrate spikes and everyone's quality drops.
- Join storms. Meetings start on the hour, so load test that minute, not the average.
Operating it: metrics and deployments
Measure quality from the client, because the SFU cannot see freezes or concealment. The standard getStats() counters are enough to start: for video, freezeCount and totalFreezesDuration on inbound RTP; for audio, concealedSamples divided by totalSamplesReceived; for latency, jitterBufferDelay divided by jitterBufferEmittedCount. Aggregate by region, network and client version, and alert on the share of participant-minutes with freezes or heavy concealment rather than on averages.
Alongside these, track join success rate, time to first audio and first video, reconnect rate, SFU CPU and packets per second, and relay bandwidth between regions. For low-latency one-to-many delivery, where some viewers only watch, compare against the design in the low-latency streaming article.
Deploy SFUs by draining. Mark a server as not accepting new rooms, wait for its rooms to end or migrate them with a forced reconnect at a quiet moment, then replace it.
Trade-offs at a glance
| Decision | Simpler choice | Scalable choice | What you pay |
|---|---|---|---|
| Room placement | One SFU per room | Cascade per region | Relay protocol, more failure paths |
| Video to receivers | Everything, top layer | Layout-driven layers | Client must report layout |
| Audio to receivers | All streams | Last-N by audio level | Quiet speakers are dropped |
| Encryption | SRTP hop by hop | SFrame end to end | Key management; no server recording |
| Recording | Per-track on SFU | Composite participant | Rendering compute per room |
What to do next
- Draw your planes: which service owns room state, and confirm that losing an SFU loses only media, not the meeting.
- Add the per-room lock to joins and test two simultaneous first joiners from one region.
- Work out your capacity unit with the worked example, then load test packets per second per core and set allocator thresholds from the result.
- Send layout-driven subscriptions from clients, pause off-screen tracks and forward only the loudest few audio streams.
- Collect freeze, concealment and jitter buffer counters from clients and alert on the share of bad participant-minutes.
- Kill an SFU during a staged busy hour and measure how long clients take to recover.
- If you need end-to-end encryption, decide the key scheme and what it means for recording before writing the transform.