Zoom became the default meeting tool largely because it worked on bad networks and cheap laptops when competitors stuttered. That was an architectural result, not luck. The company made a few early choices that are worth studying because they apply to any real-time media system: encode each camera once into a layered stream, never decode or re-encode video in the middle of the network, push intelligence to the edges, and treat the transport as something to negotiate rather than assume.
This article reconstructs that architecture from what Zoom has published (security white papers, firewall guidance, engineering talks) and from reported descriptions of those talks; design inferences are labelled as such. For the generic conferencing building blocks such as signalling, rooms and cascading forwarding units, read video conferencing architecture first; this page focuses on what is distinctive about the Zoom approach and on the numbers that make it work.
The core idea: one layered stream per sender
A meeting with N cameras has an N-to-N delivery problem. There are three classic answers. A mesh sends every stream directly to every peer, so each client uploads N-1 copies; it collapses beyond a handful of people. An MCU (multipoint control unit) decodes every stream on a server, composes one picture and re-encodes it per receiver; clients are simple, but the server burns CPU on every pixel and adds encode latency. A selective forwarding unit receives each stream once and forwards packets without decoding them; servers are cheap per stream, but each receiver needs a version of each stream that fits its bandwidth and screen.
Zoom's answer to that last problem is scalable video coding (SVC). The sender encodes a single bitstream organised in layers: a base layer at low resolution and frame rate, plus enhancement layers that add frame rate (temporal layers) and resolution (spatial layers). Each packet is tagged with its layer. A receiver that gets only the base layer can still decode a valid, smaller picture; one that gets more layers gets a better picture. The server Zoom calls a multimedia router never touches the pixels. It reads layer identifiers in packet headers and drops the layers a given receiver should not get.
The alternative most WebRTC systems use is simulcast: the sender encodes two or three independent streams and the forwarding unit picks one per receiver. SVC encodes once and upper layers reuse the lower ones, so the uplink is smaller and the sender's CPU does less.
The architecture end to end
Clients. The native clients carry most of the intelligence: capture, noise suppression, the SVC encoder, bandwidth estimation, forward error correction and jitter buffering, in Zoom's own media stack. The web client inherits the browser's media capabilities (Zoom has described WebAssembly and WebRTC components there), so it is the constrained path.
Meeting zones and routers. Media servers run in Zoom's co-located data centres, with public cloud capacity for bursts, as during the 2020 surge. A meeting is hosted in a zone whose routers forward per-receiver subsets. For participants on several continents, regional routers cascade: each receives a stream once from its peer and fans it out locally.
Connectors and control plane. SIP and H.323 room systems and telephone callers join through gateways that translate protocols, one of the few places where decoding and re-encoding happens. Scheduling, authentication and meeting metadata run as ordinary web services; a client asks the control plane which zone hosts the meeting, then opens media connections there, so a slow API never freezes a running call.
The transport ladder
Real-time media wants UDP. A lost UDP packet is simply gone, and the receiver conceals or repairs the gap; with TCP, one lost segment stalls every later byte until it is retransmitted, which turns a 1 percent loss into visible freezes. But many enterprise and hotel networks block outbound UDP or allow only web traffic through a proxy. Zoom's clients are described as trying transports in order: UDP first, then TCP, then TLS on port 443, and finally an HTTPS proxy tunnel. Zoom's firewall guidance for the meeting client lists TCP 443, 8801 and 8802 and UDP 3478, 3479 and 8801 to 8810, which matches that ladder; check the current support article before writing firewall rules, because ranges and domains change.
A sketch of the client logic shows the important details: rank zones by measured round-trip time, try cheap transports first with short timeouts, record which rung each user lands on, and keep probing for an upgrade because networks change mid-call.
# Client-side connection strategy: try the cheapest transport first, keep the others warm.
CANDIDATES = [
("udp", 8801), # preferred: no head-of-line blocking, lowest latency
("tcp", 8801), # used when UDP is blocked outbound
("tls", 443), # looks like ordinary HTTPS; survives most corporate firewalls
("https-proxy", 443), # CONNECT through the configured web proxy
]
def connect(zone_hosts, deadline_ms=3000):
for host in rank_by_rtt(zone_hosts): # nearest zone first
for kind, port in CANDIDATES:
sock = try_open(kind, host, port, timeout_ms=deadline_ms // len(CANDIDATES))
if sock and handshake(sock):
report_metric("transport", kind) # fleet-wide view of how many users are degraded
if kind != "udp":
schedule_upgrade_probe(host) # networks change; retry UDP later
return sock
raise NoConnectivity("all zones and transports failed")The metric matters: if the share of users on TCP or TLS jumps in one country, a network or firewall vendor changed something. For background on why UDP needs this treatment and how NAT and firewalls interfere, see NAT traversal; QUIC-based media is a newer alternative that keeps UDP while adding congestion control and encryption, covered in QUIC.
What the router decides for each receiver
A router is a packet forwarder with a policy engine attached. For each receiver it maintains a bandwidth estimate (from receiver feedback about arrival times and losses), the layout the receiver is showing (active speaker large, gallery of tiles, or screen share), and the set of layers each sender is producing. On every feedback report it recomputes which layers of which senders to forward.
# Per receiver, per sender: pick the SVC layers to forward. Runs on every feedback report.
def select_layers(sender_stream, receiver):
budget = receiver.estimated_bandwidth * 0.85 # headroom for audio, FEC and retransmits
tile = receiver.tile_size_for(sender_stream.sender) # thumbnail, gallery tile or active speaker
wanted = []
for layer in sender_stream.layers: # ordered base layer first
if layer.height > tile.height * 1.25:
break # no point sending pixels the tile cannot show
if layer.cumulative_bitrate > budget:
break
wanted.append(layer)
if not wanted:
wanted = [sender_stream.layers[0]] # always keep the base layer if anything at all
return wanted # forward packets whose layer id is in this setTwo rules in that sketch carry most of the value. First, never forward pixels the receiver's tile cannot display: a 160-pixel-tall gallery thumbnail gets the base layer even if the receiver has plenty of bandwidth. Second, degrade by dropping enhancement layers, not by stalling: the base layer is protected, so a struggling receiver sees a softer or choppier picture rather than a frozen one. Because the layers are in one bitstream, switching is instantaneous at the next frame that permits it; with simulcast, switching to a different stream usually needs a keyframe request to the sender.
The router also prunes the uplink: if nobody views a sender's high layers, it tells the sender to stop encoding them, saving laptop CPU and upload bandwidth.
Audio is handled differently. A common approach, consistent with Zoom's behaviour, is to forward only the few loudest speakers, chosen from audio-level indicators on packets, and let clients mix them locally.
Worked example: a 25-person all-hands
Assume the following. These are round illustrative numbers, not Zoom's published bitrates. A sender's SVC stream has a base layer at 180p, 15 frames per second, about 150 kbit/s; a middle layer that brings it to 360p and 30 frames per second, cumulative 500 kbit/s; and a top layer at 720p, cumulative 1.2 Mbit/s. Audio is 40 kbit/s per speaker. Twenty-five people join, one presents slides as a screen share at about 600 kbit/s, and the receivers show a 5-by-5 gallery.
| Design | Uplink per sender | Downlink per receiver | Server work |
|---|---|---|---|
| Mesh | 24 copies: about 28.8 Mbit/s at 720p | 24 streams | None; collapses on clients |
| MCU | 1.2 Mbit/s | One composed 720p stream, about 1.2 Mbit/s | 25 decodes plus 25 encodes per meeting |
| Simulcast SFU | 3 encodes: about 1.85 Mbit/s | 24 low streams: about 3.6 Mbit/s | Forward only |
| SVC router | Base layer only while in gallery: 150 kbit/s | 24 base layers: about 3.6 Mbit/s plus share and audio | Forward only |
The interesting line is the uplink: in gallery view nobody needs more than 180p, so each laptop uploads about 150 kbit/s. Downlink per receiver is 24 base layers (3.6 Mbit/s), the screen share (0.6 Mbit/s) and three loudest speakers (0.12 Mbit/s): about 4.3 Mbit/s. A receiver on a 2 Mbit/s hotel connection cannot take that, so the router responds per receiver, not per meeting: it reduces that one person's gallery to the base layer at a lower frame rate for off-screen tiles, or paginates the gallery, keeps the screen share and audio intact, and leaves the other 24 receivers untouched.
When receivers switch to speaker view, the router asks only the speaker for the top layer; one uplink grows to 1.2 Mbit/s and server cost stays at packet forwarding.
Resilience inside the media path
Packet loss is the normal state of consumer networks, so the media path assumes it. Forward error correction sends redundant parity packets so a receiver can rebuild a lost packet without waiting a round trip; the sender raises the FEC ratio when receivers report loss and lowers it when the network is clean, because parity costs bandwidth. Retransmission suits receivers close to the router, and an adaptive jitter buffer trades delay for smoothness.
SVC helps here too: losses in an enhancement layer only damage that layer's frames, while the base layer, given extra FEC, keeps decoding.
At the server level, a router failure drops the streams it carried. Clients detect the loss within seconds, return to the control plane or zone selection, and reconnect to another router in the same zone; the meeting's state (who is in it, who is presenting, roles) lives outside the router so reconnection is a rejoin, not a restart. Zone-level failure is handled the same way at a wider radius, which is why meeting state must never live only in a single media server's memory.
Encryption: what the servers can and cannot see
By default, Zoom meeting media is encrypted with AES-256-GCM between clients and Zoom's servers, and the keys are managed by Zoom's infrastructure. That allows the servers to provide features that need content, such as cloud recording and live transcription, while still forwarding encrypted packets.
Zoom's optional end-to-end encryption changes who holds the key. According to Zoom's own description, the meeting host's client generates the meeting key and distributes it to participants using public-key cryptography, so the routers become relays that forward packets they cannot decrypt. Forwarding still works because the router only needs the unencrypted routing information in packet headers, such as layer identifiers, not the media payload. The trade-off is that any feature that needs the server to see content stops working. At launch in 2020, enabling E2EE disabled cloud recording, live streaming, live transcription, breakout rooms, polling and several other features; that list has changed since, so check Zoom's current documentation before promising a feature to an E2EE user.
Operating a system like this
Measure per participant, not per server: join success and time, transport rung, received bitrate, loss before and after FEC, freezes, audio concealment and round-trip time. Slice by ISP, country and client version, because most incidents are one ISP or one release.
Capacity planning is about forwarding throughput and egress, not CPU. Meetings start on the hour, so keep headroom for top-of-hour spikes, and size transcoding gateways separately.
Stage client rollouts by percentage and keep server-side switches for new codecs and features, so a bad release can be disabled without shipping a new client.
Failure modes and trade-offs
| Problem | Cause | Mitigation |
|---|---|---|
| Everyone in one office freezes | UDP blocked, users on TLS over a proxy | Firewall rules from current guidance; transport metrics by network |
| One user degrades the meeting | Server forwards a single quality to all | Per-receiver layer selection; never per-meeting downgrade |
| Laptops overheat in large meetings | Senders encoding layers nobody views | Router-driven uplink pruning |
| Cross-continent meetings lag | All users on one distant zone | Regional routers cascaded over a private backbone |
| Recording missing in E2EE meetings | Server cannot decrypt | Explain trade-off; local recording where allowed |
| Mass reconnect after a router fails | State held in the media server | Keep meeting state in the control plane; fast rejoin |
SVC versus simulcast is the main trade-off. SVC wins on uplink, sender CPU and instant layer switching, but needs layered-codec support in every client, which owning the clients made practical for Zoom. Browser-based systems usually choose simulcast, described in selective forwarding units. Owning native clients also means owning their media engineering and security on every platform.
What to do next
- Decide whether you control the clients. If you do not, plan on WebRTC simulcast; if you do, evaluate SVC for uplink savings.
- Write the transport ladder explicitly, with timeouts, and emit a metric for the rung each session lands on.
- Implement per-receiver layer selection that respects tile size as well as bandwidth, and protect the base layer.
- Add uplink pruning so senders stop encoding layers nobody is viewing.
- Keep meeting state out of the media servers so a router failure is a fast rejoin.
- Choose an encryption model per product tier, and list which server-side features each model disables before you ship it.
- Build quality dashboards by ISP, network and client version, and roll out client releases in stages.