WebRTC is the browser's built-in stack for sending audio, video and arbitrary data directly between peers with sub-second latency. A production WebRTC system is really two systems. The first is a signaling plane that you build yourself and that WebRTC deliberately does not specify. The second is a media plane that the browser runs for you: NAT traversal, encryption, packetisation, congestion control and jitter buffering.
This article explains that media plane from first principles, then the servers you still have to run around it: STUN, TURN and usually a selective forwarding unit. It ends with the failure modes that show up once real users on real networks join. The signaling transport itself, meaning how offers travel between peers, is covered in the signaling architecture article. Here we care about what those messages contain and what happens after they arrive.
The stack, layer by layer
WebRTC is not one protocol. It is a profile that stitches existing IETF protocols together, each chosen for a specific job.
| Layer | Protocol | Job |
|---|---|---|
| Session description | SDP, driven by JSEP (RFC 9429) | Declare codecs, media sections, ICE credentials and DTLS fingerprints |
| Connectivity | ICE (RFC 8445), STUN (RFC 8489), TURN (RFC 8656) | Find a working network path through NATs and firewalls |
| Multiplexing | BUNDLE (RFC 9143), rtcp-mux | Carry every track and the data channel on one 5-tuple |
| Security | DTLS, DTLS-SRTP (RFC 5764) | Authenticate the peer and derive SRTP keys |
| Media | RTP and RTCP (RFC 3550), SRTP | Packetise, sequence, timestamp and report on media |
| Data | SCTP over DTLS (RFC 8831, 8832) | Reliable or unreliable, ordered or unordered messages |
Two design decisions shape everything else. Encryption is mandatory, so there is no unencrypted WebRTC media on the wire. Media is carried over UDP whenever possible, because a retransmitted video packet that arrives 300 ms late is worse than useless. TCP is only a fallback, and when it is used, the transport's head-of-line blocking shows up directly as frozen video.
Offer, answer and perfect negotiation
A connection starts when one side creates an offer: an SDP blob listing each media section (m=audio, m=video, m=application for data), the codecs it supports, its ICE username fragment and password, and a a=fingerprint line holding the hash of its DTLS certificate. The other side applies the offer with setRemoteDescription and produces an answer that selects from what was offered.
The classic bug is glare. Both peers renegotiate at the same moment, for example when both add a screen share, and each receives an offer while holding its own. The robust answer is the "perfect negotiation" pattern described in the W3C spec and on MDN. Both peers run identical code, one is designated polite, the polite peer yields on collision (the browser performs an implicit rollback), and the impolite peer ignores the incoming offer.
// "Perfect negotiation": both peers run the same code; one is polite.
const pc = new RTCPeerConnection({
iceServers: [
{ urls: "stun:stun.example.com:3478" },
{ urls: ["turn:turn.example.com:3478?transport=udp",
"turns:turn.example.com:443?transport=tcp"],
username: creds.username, credential: creds.password },
],
});
let makingOffer = false, ignoreOffer = false;
pc.onnegotiationneeded = async () => {
try {
makingOffer = true;
await pc.setLocalDescription(); // creates offer implicitly
signal.send({ description: pc.localDescription });
} finally { makingOffer = false; }
};
pc.onicecandidate = ({ candidate }) => signal.send({ candidate }); // trickle ICE
signal.onmessage = async ({ description, candidate }) => {
if (description) {
const collision = description.type === "offer" &&
(makingOffer || pc.signalingState !== "stable");
ignoreOffer = !polite && collision;
if (ignoreOffer) return; // impolite peer wins glare
await pc.setRemoteDescription(description); // polite peer rolls back implicitly
if (description.type === "offer") {
await pc.setLocalDescription(); // creates answer
signal.send({ description: pc.localDescription });
}
} else if (candidate) {
try { await pc.addIceCandidate(candidate); }
catch (e) { if (!ignoreOffer) throw e; }
}
};
pc.oniceconnectionstatechange = () => {
if (pc.iceConnectionState === "failed") pc.restartIce();
};Note three details. setLocalDescription() with no argument creates the right description for the current state. Candidates are sent as they are discovered (trickle ICE, RFC 8838) rather than after gathering completes, which saves seconds of setup. And an ICE failure triggers restartIce(), which renegotiates fresh ICE credentials without tearing down the tracks.
ICE: finding a path through NAT
Most devices sit behind at least one NAT, and many sit behind two (home router plus carrier-grade NAT). A peer cannot simply announce its IP address. ICE solves this by gathering candidate addresses of several types and then testing pairs of them.
- host: the local interface address. Works on the same LAN. Browsers usually hide it behind an mDNS name.
- srflx (server reflexive): the public address and port a STUN server observed. Works when the NATs on both sides allow it.
- relay: an address allocated on a TURN server. Always works if the TURN server is reachable, at the cost of server bandwidth and an extra hop.
- prflx (peer reflexive): an address learned during connectivity checks themselves, typically when a NAT maps differently for the peer than for the STUN server.
Each side pairs its candidates with the remote side's, sorts pairs by priority (host above srflx above relay), and sends STUN binding requests on each pair, authenticated with the ICE credentials from SDP. The controlling agent nominates the best pair that succeeded. Consent freshness checks continue every few seconds for the life of the call, so if the path dies the state moves to disconnected and then failed.
Worked example. Alice is on home broadband and Bob is on a corporate network that only allows outbound TCP 443 through a proxy. Alice's host and srflx candidates are fine, but Bob's UDP is blocked outright, so every UDP pair fails. The only pair that succeeds is Bob's relay candidate allocated over TLS on port 443 against Alice's srflx. The call works, but every packet now crosses your TURN server over TCP. Without a turns: URL on port 443, this call would simply fail after about 30 seconds of checks.
TURN is infrastructure, not an afterthought
Teams routinely find that a meaningful share of sessions, often cited in the region of 10 to 20 percent and higher on enterprise and mobile networks, need a relay. Measure your own share rather than trusting a figure. TURN servers are stateful packet relays: each allocation holds a public port, and every media byte of a relayed call crosses the server twice (in and out). Size them by bandwidth, not connection count, and deploy them close to users.
# coturn (turnserver.conf): shared-secret credentials, UDP plus TLS on 443
listening-port=3478
tls-listening-port=443
realm=turn.example.com
use-auth-secret
static-auth-secret=CHANGE_ME_FROM_A_SECRET_STORE
external-ip=203.0.113.10
min-port=49152
max-port=65535
no-multicast-peers
denied-peer-ip=10.0.0.0-10.255.255.255 # stop the relay reaching your private network
denied-peer-ip=172.16.0.0-172.31.255.255
denied-peer-ip=192.168.0.0-192.168.255.255
cert=/etc/ssl/turn.pem
pkey=/etc/ssl/turn.keyNever ship static TURN passwords in client code, because anyone can harvest them and use your relay as a free proxy. The common pattern, implemented by coturn's use-auth-secret mode and described in an IETF draft usually called the TURN REST API, has your backend mint short-lived credentials from a shared secret:
import base64, hashlib, hmac, time
def turn_credentials(user_id: str, secret: bytes, ttl_s: int = 3600) -> dict:
"""Time-limited TURN credentials (the draft 'TURN REST API' convention coturn implements)."""
username = f"{int(time.time()) + ttl_s}:{user_id}" # expiry first, then user
digest = hmac.new(secret, username.encode(), hashlib.sha1).digest()
return {"username": username, "password": base64.b64encode(digest).decode()}The denied-peer-ip lines matter as much as authentication. Without them, an authenticated client can ask the relay to send packets to addresses inside your own network.
Security: DTLS-SRTP and the fingerprint
Once ICE selects a path, the peers run a DTLS handshake over it. Each side presents a certificate, usually self-signed and generated per connection, and checks that its hash matches the a=fingerprint received in SDP. DTLS-SRTP then exports keying material from the handshake to key SRTP for media, while SCTP runs inside the DTLS session for data.
The consequence is that WebRTC's end-to-end security is exactly as strong as the integrity of your signaling channel. If an attacker can rewrite SDP in transit or on your signaling server, they can substitute their own fingerprint and sit in the middle. Signaling must run over TLS with authenticated users. Also, a server that terminates media, such as an SFU or MCU, necessarily decrypts it. End-to-end encryption through an SFU needs extra per-frame encryption in the browser.
The media plane: RTP, feedback and congestion control
Encoded frames are split into RTP packets carrying a sequence number, timestamp and SSRC (stream identifier). RTCP runs alongside and carries the feedback that makes real-time media adapt: receiver reports with loss and jitter, NACKs requesting retransmission of specific packets, PLI requests asking for a new keyframe, and transport-wide congestion control feedback reporting the arrival time of every packet.
Congestion control is the heart of perceived quality. Browsers based on libwebrtc use a delay-based estimator (Google Congestion Control) that watches the growth of one-way queuing delay from that per-packet feedback, and backs off before loss occurs. Its output is a target bitrate that the encoder follows by changing resolution, frame rate and quantisation. The receiver side runs an adaptive jitter buffer that holds packets just long enough to reorder them and smooth arrival variance. That buffer's size is what you are trading when you choose between latency and smoothness.
Data channels
Data channels ride SCTP inside the same DTLS session. Each channel is configured at creation: ordered (default true) and either maxRetransmits or maxPacketLifeTime for partial reliability. Game state updates want {ordered: false, maxRetransmits: 0}; file transfer wants the reliable default. The bottleneck is usually bufferedAmount: if you push faster than the path drains, memory grows without limit. Use bufferedAmountLowThreshold and the bufferedamountlow event to apply backpressure, the same discipline described in the backpressure article.
Topologies: mesh, SFU and MCU
Peer-to-peer is only the right topology for two, perhaps three, participants. Beyond that, the cost moves to the server. Consider a six-person call where each camera sends 720p at about 1.5 Mbps.
| Topology | Each client uploads | Each client downloads | Server work |
|---|---|---|---|
| Mesh (every peer to every peer) | 5 x 1.5 = 7.5 Mbps, five encodes | 7.5 Mbps | None beyond TURN |
| SFU (forwards packets) | 1.5 Mbps, or about 2.2 with three simulcast layers | Chosen layers for 5 senders | Routing and bandwidth only, no decoding |
| MCU (mixes) | 1.5 Mbps | 1.5 Mbps composite | Decode 6, compose, encode per layout |
Mesh collapses on upload bandwidth and encoder CPU. An MCU is expensive and adds transcoding latency. The selective forwarding unit dominates modern systems. Each client sends once, typically using simulcast (RFC 8853): the same video encoded at, for example, 720p, 360p and 180p. The SFU forwards to each receiver the layer that fits that receiver's estimated bandwidth and the size it is displayed at. SVC codecs achieve the same thing with layers inside one stream.
Scaling beyond one SFU means cascading: rooms spread across SFU nodes that forward to each other, usually one per region, so that a participant's media crosses the ocean once rather than once per remote viewer. See the fan-out article for the general pattern.
Observability with getStats
Without per-call telemetry you cannot tell whether a bad call was Wi-Fi, your TURN region or a CPU-starved encoder. The getStats() API exposes the standard statistics objects. Sample them periodically in the client and ship them to your backend.
async function sample(pc) {
const stats = await pc.getStats();
const out = {};
stats.forEach(r => {
if (r.type === "candidate-pair" && r.nominated && r.state === "succeeded") {
out.rttMs = r.currentRoundTripTime * 1000;
out.availableOutKbps = (r.availableOutgoingBitrate || 0) / 1000;
const local = stats.get(r.localCandidateId);
out.path = local.candidateType; // host | srflx | prflx | relay
}
if (r.type === "inbound-rtp" && r.kind === "video") {
out.videoLost = r.packetsLost;
out.framesDropped = r.framesDropped;
out.jitterMs = r.jitter * 1000;
}
});
return out; // ship every 5-10 s to your metrics pipeline, tagged with call id
}Build dashboards around a few ratios: setup success rate (connected within 10 seconds), relay share by network type, p95 RTT, video freeze time per minute, and the reasons recorded in qualityLimitationReason on outbound video (bandwidth, cpu, none).
Failure modes and trade-offs
- No TURN over TLS 443. Calls from locked-down networks fail silently after ICE times out. Always offer a
turns:URL on 443. - Leaked or static TURN credentials. Your relay becomes an open proxy and a bandwidth bill. Mint short-lived credentials and restrict peer IP ranges.
- Glare and stuck negotiation. Ad hoc renegotiation code deadlocks when both sides act at once. Use perfect negotiation, and log
signalingStatetransitions. - Network changes. Wi-Fi to cellular changes every candidate. Without ICE restart the call dies, and with it the call recovers after a brief freeze.
- Trade-off: latency versus smoothness. A deeper jitter buffer hides network variance at the cost of conversational delay. Real-time conversation wants it shallow, while broadcast-like viewing can afford more, or can move to a different protocol altogether, as discussed in the WebTransport article.
What to do next
- Draw your two planes: which service relays signaling, and over what authenticated TLS channel.
- Implement perfect negotiation and ICE restart on
failedbefore adding features such as screen share. - Deploy TURN with UDP 3478 and TLS 443 in each region you serve, with short-lived credentials and private ranges denied.
- Pick a topology from your largest expected room size: peer-to-peer for two, an SFU with simulcast beyond that.
- Ship
getStats()samples tagged with a call ID and chart setup success, relay share, RTT and freeze time. - Test on hostile networks: UDP blocked, 5 percent loss, 300 ms added RTT and a mid-call Wi-Fi to cellular switch.