Zoom became the default meeting tool largely because it worked on bad networks and cheap laptops when competitors stuttered. That was an architectural result, not luck. The company made a few early choices that are worth studying because they apply to any real-time media system: encode each camera once into a layered stream, never decode or re-encode video in the middle of the network, push intelligence to the edges, and treat the transport as something to negotiate rather than assume.

This article reconstructs that architecture from what Zoom has published (security white papers, firewall guidance, engineering talks) and from reported descriptions of those talks; design inferences are labelled as such. For the generic conferencing building blocks such as signalling, rooms and cascading forwarding units, read video conferencing architecture first; this page focuses on what is distinctive about the Zoom approach and on the numbers that make it work.

Advertisement

The core idea: one layered stream per sender

A meeting with N cameras has an N-to-N delivery problem. There are three classic answers. A mesh sends every stream directly to every peer, so each client uploads N-1 copies; it collapses beyond a handful of people. An MCU (multipoint control unit) decodes every stream on a server, composes one picture and re-encodes it per receiver; clients are simple, but the server burns CPU on every pixel and adds encode latency. A selective forwarding unit receives each stream once and forwards packets without decoding them; servers are cheap per stream, but each receiver needs a version of each stream that fits its bandwidth and screen.

Zoom's answer to that last problem is scalable video coding (SVC). The sender encodes a single bitstream organised in layers: a base layer at low resolution and frame rate, plus enhancement layers that add frame rate (temporal layers) and resolution (spatial layers). Each packet is tagged with its layer. A receiver that gets only the base layer can still decode a valid, smaller picture; one that gets more layers gets a better picture. The server Zoom calls a multimedia router never touches the pixels. It reads layer identifiers in packet headers and drops the layers a given receiver should not get.

The alternative most WebRTC systems use is simulcast: the sender encodes two or three independent streams and the forwarding unit picks one per receiver. SVC encodes once and upper layers reuse the lower ones, so the uplink is smaller and the sender's CPU does less.

The architecture end to end

Native clientSVC encoder, FECWeb clientbrowser media stackRoom / phoneSIP, H.323, PSTNTransport ladderUDP, then TCP, then TLSConnectorsgateways for SIP / PSTNMeeting zonemultimedia routerZone in region Bcascaded routerControl planeweb, scheduling, authReceiverschosen layersRecording / captionsserver participantsmediaprivate backbonejoin, tokensEvery sender uploads one layered stream; routers forward the subset each receiver can use.
Clients reach a nearby meeting zone over the best transport they can get. Multimedia routers forward layer subsets, cascade to other regions over a private backbone, and treat recording and captions as participants.

Clients. The native clients carry most of the intelligence: capture, noise suppression, the SVC encoder, bandwidth estimation, forward error correction and jitter buffering, in Zoom's own media stack. The web client inherits the browser's media capabilities (Zoom has described WebAssembly and WebRTC components there), so it is the constrained path.

Meeting zones and routers. Media servers run in Zoom's co-located data centres, with public cloud capacity for bursts, as during the 2020 surge. A meeting is hosted in a zone whose routers forward per-receiver subsets. For participants on several continents, regional routers cascade: each receives a stream once from its peer and fans it out locally.

Connectors and control plane. SIP and H.323 room systems and telephone callers join through gateways that translate protocols, one of the few places where decoding and re-encoding happens. Scheduling, authentication and meeting metadata run as ordinary web services; a client asks the control plane which zone hosts the meeting, then opens media connections there, so a slow API never freezes a running call.

Advertisement

The transport ladder

Real-time media wants UDP. A lost UDP packet is simply gone, and the receiver conceals or repairs the gap; with TCP, one lost segment stalls every later byte until it is retransmitted, which turns a 1 percent loss into visible freezes. But many enterprise and hotel networks block outbound UDP or allow only web traffic through a proxy. Zoom's clients are described as trying transports in order: UDP first, then TCP, then TLS on port 443, and finally an HTTPS proxy tunnel. Zoom's firewall guidance for the meeting client lists TCP 443, 8801 and 8802 and UDP 3478, 3479 and 8801 to 8810, which matches that ladder; check the current support article before writing firewall rules, because ranges and domains change.

A sketch of the client logic shows the important details: rank zones by measured round-trip time, try cheap transports first with short timeouts, record which rung each user lands on, and keep probing for an upgrade because networks change mid-call.

# Client-side connection strategy: try the cheapest transport first, keep the others warm.
CANDIDATES = [
    ("udp", 8801),    # preferred: no head-of-line blocking, lowest latency
    ("tcp", 8801),    # used when UDP is blocked outbound
    ("tls", 443),     # looks like ordinary HTTPS; survives most corporate firewalls
    ("https-proxy", 443),  # CONNECT through the configured web proxy
]

def connect(zone_hosts, deadline_ms=3000):
    for host in rank_by_rtt(zone_hosts):            # nearest zone first
        for kind, port in CANDIDATES:
            sock = try_open(kind, host, port, timeout_ms=deadline_ms // len(CANDIDATES))
            if sock and handshake(sock):
                report_metric("transport", kind)     # fleet-wide view of how many users are degraded
                if kind != "udp":
                    schedule_upgrade_probe(host)     # networks change; retry UDP later
                return sock
    raise NoConnectivity("all zones and transports failed")

The metric matters: if the share of users on TCP or TLS jumps in one country, a network or firewall vendor changed something. For background on why UDP needs this treatment and how NAT and firewalls interfere, see NAT traversal; QUIC-based media is a newer alternative that keeps UDP while adding congestion control and encryption, covered in QUIC.

What the router decides for each receiver

A router is a packet forwarder with a policy engine attached. For each receiver it maintains a bandwidth estimate (from receiver feedback about arrival times and losses), the layout the receiver is showing (active speaker large, gallery of tiles, or screen share), and the set of layers each sender is producing. On every feedback report it recomputes which layers of which senders to forward.

# Per receiver, per sender: pick the SVC layers to forward. Runs on every feedback report.
def select_layers(sender_stream, receiver):
    budget = receiver.estimated_bandwidth * 0.85       # headroom for audio, FEC and retransmits
    tile = receiver.tile_size_for(sender_stream.sender)  # thumbnail, gallery tile or active speaker
    wanted = []
    for layer in sender_stream.layers:                  # ordered base layer first
        if layer.height > tile.height * 1.25:
            break                                       # no point sending pixels the tile cannot show
        if layer.cumulative_bitrate > budget:
            break
        wanted.append(layer)
    if not wanted:
        wanted = [sender_stream.layers[0]]              # always keep the base layer if anything at all
    return wanted                                       # forward packets whose layer id is in this set

Two rules in that sketch carry most of the value. First, never forward pixels the receiver's tile cannot display: a 160-pixel-tall gallery thumbnail gets the base layer even if the receiver has plenty of bandwidth. Second, degrade by dropping enhancement layers, not by stalling: the base layer is protected, so a struggling receiver sees a softer or choppier picture rather than a frozen one. Because the layers are in one bitstream, switching is instantaneous at the next frame that permits it; with simulcast, switching to a different stream usually needs a keyframe request to the sender.

The router also prunes the uplink: if nobody views a sender's high layers, it tells the sender to stop encoding them, saving laptop CPU and upload bandwidth.

Audio is handled differently. A common approach, consistent with Zoom's behaviour, is to forward only the few loudest speakers, chosen from audio-level indicators on packets, and let clients mix them locally.

Worked example: a 25-person all-hands

Assume the following. These are round illustrative numbers, not Zoom's published bitrates. A sender's SVC stream has a base layer at 180p, 15 frames per second, about 150 kbit/s; a middle layer that brings it to 360p and 30 frames per second, cumulative 500 kbit/s; and a top layer at 720p, cumulative 1.2 Mbit/s. Audio is 40 kbit/s per speaker. Twenty-five people join, one presents slides as a screen share at about 600 kbit/s, and the receivers show a 5-by-5 gallery.

DesignUplink per senderDownlink per receiverServer work
Mesh24 copies: about 28.8 Mbit/s at 720p24 streamsNone; collapses on clients
MCU1.2 Mbit/sOne composed 720p stream, about 1.2 Mbit/s25 decodes plus 25 encodes per meeting
Simulcast SFU3 encodes: about 1.85 Mbit/s24 low streams: about 3.6 Mbit/sForward only
SVC routerBase layer only while in gallery: 150 kbit/s24 base layers: about 3.6 Mbit/s plus share and audioForward only

The interesting line is the uplink: in gallery view nobody needs more than 180p, so each laptop uploads about 150 kbit/s. Downlink per receiver is 24 base layers (3.6 Mbit/s), the screen share (0.6 Mbit/s) and three loudest speakers (0.12 Mbit/s): about 4.3 Mbit/s. A receiver on a 2 Mbit/s hotel connection cannot take that, so the router responds per receiver, not per meeting: it reduces that one person's gallery to the base layer at a lower frame rate for off-screen tiles, or paginates the gallery, keeps the screen share and audio intact, and leaves the other 24 receivers untouched.

When receivers switch to speaker view, the router asks only the speaker for the top layer; one uplink grows to 1.2 Mbit/s and server cost stays at packet forwarding.

Resilience inside the media path

Packet loss is the normal state of consumer networks, so the media path assumes it. Forward error correction sends redundant parity packets so a receiver can rebuild a lost packet without waiting a round trip; the sender raises the FEC ratio when receivers report loss and lowers it when the network is clean, because parity costs bandwidth. Retransmission suits receivers close to the router, and an adaptive jitter buffer trades delay for smoothness.

SVC helps here too: losses in an enhancement layer only damage that layer's frames, while the base layer, given extra FEC, keeps decoding.

At the server level, a router failure drops the streams it carried. Clients detect the loss within seconds, return to the control plane or zone selection, and reconnect to another router in the same zone; the meeting's state (who is in it, who is presenting, roles) lives outside the router so reconnection is a rejoin, not a restart. Zone-level failure is handled the same way at a wider radius, which is why meeting state must never live only in a single media server's memory.

Encryption: what the servers can and cannot see

By default, Zoom meeting media is encrypted with AES-256-GCM between clients and Zoom's servers, and the keys are managed by Zoom's infrastructure. That allows the servers to provide features that need content, such as cloud recording and live transcription, while still forwarding encrypted packets.

Zoom's optional end-to-end encryption changes who holds the key. According to Zoom's own description, the meeting host's client generates the meeting key and distributes it to participants using public-key cryptography, so the routers become relays that forward packets they cannot decrypt. Forwarding still works because the router only needs the unencrypted routing information in packet headers, such as layer identifiers, not the media payload. The trade-off is that any feature that needs the server to see content stops working. At launch in 2020, enabling E2EE disabled cloud recording, live streaming, live transcription, breakout rooms, polling and several other features; that list has changed since, so check Zoom's current documentation before promising a feature to an E2EE user.

Operating a system like this

Measure per participant, not per server: join success and time, transport rung, received bitrate, loss before and after FEC, freezes, audio concealment and round-trip time. Slice by ISP, country and client version, because most incidents are one ISP or one release.

Capacity planning is about forwarding throughput and egress, not CPU. Meetings start on the hour, so keep headroom for top-of-hour spikes, and size transcoding gateways separately.

Stage client rollouts by percentage and keep server-side switches for new codecs and features, so a bad release can be disabled without shipping a new client.

Failure modes and trade-offs

ProblemCauseMitigation
Everyone in one office freezesUDP blocked, users on TLS over a proxyFirewall rules from current guidance; transport metrics by network
One user degrades the meetingServer forwards a single quality to allPer-receiver layer selection; never per-meeting downgrade
Laptops overheat in large meetingsSenders encoding layers nobody viewsRouter-driven uplink pruning
Cross-continent meetings lagAll users on one distant zoneRegional routers cascaded over a private backbone
Recording missing in E2EE meetingsServer cannot decryptExplain trade-off; local recording where allowed
Mass reconnect after a router failsState held in the media serverKeep meeting state in the control plane; fast rejoin

SVC versus simulcast is the main trade-off. SVC wins on uplink, sender CPU and instant layer switching, but needs layered-codec support in every client, which owning the clients made practical for Zoom. Browser-based systems usually choose simulcast, described in selective forwarding units. Owning native clients also means owning their media engineering and security on every platform.

What to do next

  1. Decide whether you control the clients. If you do not, plan on WebRTC simulcast; if you do, evaluate SVC for uplink savings.
  2. Write the transport ladder explicitly, with timeouts, and emit a metric for the rung each session lands on.
  3. Implement per-receiver layer selection that respects tile size as well as bandwidth, and protect the base layer.
  4. Add uplink pruning so senders stop encoding layers nobody is viewing.
  5. Keep meeting state out of the media servers so a router failure is a fast rejoin.
  6. Choose an encryption model per product tier, and list which server-side features each model disables before you ship it.
  7. Build quality dashboards by ISP, network and client version, and roll out client releases in stages.
Key takeaway: Zoom's architecture rests on encoding each camera once as layered SVC video and letting routers forward only the layers each receiver can use, so servers never transcode and every participant gets a picture sized to their screen and network. Around that core sit a transport ladder that keeps calls alive behind hostile firewalls, regional cascading, a separate control plane, and an end-to-end encryption mode that trades server-side features for privacy. Copy the principles: encode once, forward selectively, adapt per receiver, and measure quality per participant.