Most WebRTC guides stop once the connection is up: ICE found a path, DTLS keyed SRTP, and a video element shows a face. That is where the hard part starts. Video in WebRTC is a real-time control loop: the sender encodes against a bitrate it can only estimate, the network drops and delays packets, and the receiver must choose between waiting, asking for a retransmission and asking for a fresh keyframe, all within a few hundred milliseconds. When a call looks blurry, stutters or freezes, one stage of that loop made a choice you can usually see in the statistics.

This article follows one video stream from camera to screen and explains each decision on the way, so you can configure the sender sensibly and diagnose bad calls from data. Connection setup, NAT traversal and topologies are covered in the WebRTC architecture article and the WebRTC video pipeline overview; here the focus is the media engine.

Advertisement

The loop, stage by stage

WebRTC video is a closed loop: the receiver's feedback steers the sender's encoderCapturecamera / screenAdapterscale, drop framesEncoderrate controlPacketizerRTP payload formatPacer + RTX storesmooths burstsNetwork pathqueues, loss, jitterRTP/SRTPJitter bufferreorder, wait, assembleDecoderneeds a referenceRendererfreezes show hereFeedback over RTCPtransport-cc arrival times, NACK, PLIreceiver to senderBandwidth estimatetarget bitrate to encoder and pacernew targetEvery visible symptom, blur, stutter or freeze, traces back to one stage of this loop.
One sender, one receiver. Media flows right then down; feedback flows back over RTCP and becomes a new target bitrate.

Capture produces raw frames at the camera's resolution and rate. Before the encoder sees them, an adapter may scale them down or drop some, because the encoder cannot hit a low bitrate at full resolution without turning the picture to mush. The encoder compresses each frame against a target bitrate, emitting occasional keyframes, which decode alone, and many delta frames, which only decode given earlier frames. The packetizer splits each encoded frame into RTP packets according to the codec's payload format, and the pacer releases them at a smooth rate instead of in one burst per frame, while keeping recent packets in a store for retransmission.

On the receiving side, the jitter buffer collects packets, reorders them, waits for missing ones for a bounded time and assembles complete frames. The decoder can only decode a frame whose references it already holds, so a single unrecoverable loss can stall every following delta frame until the next keyframe. Meanwhile the receiver sends RTCP feedback: per-packet arrival times through the transport-wide congestion control extension, NACKs naming lost packets, and picture loss indications (PLI) asking for a keyframe. The sender turns arrival times into a bandwidth estimate, and the estimate becomes the encoder's new target.

Bandwidth estimation and what the encoder does with it

The sender cannot measure capacity directly, so it infers it. Google Congestion Control, described in an IETF draft and implemented in libwebrtc, combines two signals. The delay-based part watches the gap between packet send times and arrival times: if packets start arriving later and later relative to when they were sent, a queue is building somewhere, and the estimate drops before loss occurs. The loss-based part reacts to packet loss reported by the receiver; the draft decreases the rate when loss exceeds about 10 percent and increases it when loss is below about 2 percent. Browsers have since evolved their estimators, so treat those numbers as the design idea rather than what any given browser does today.

The estimate is a single number, the target bitrate, and the encoder must live inside it. Its rate control chooses a quantizer per frame. When the target falls, the encoder first raises the quantizer, which blurs detail. If that is not enough, something must give: resolution or frame rate. That choice is degradationPreference, a top-level member of the sender parameters with four values: balanced (the default), maintain-framerate (shrink resolution), maintain-resolution (drop frames) and maintain-framerate-and-resolution. Camera video usually wants frame rate; screen sharing of documents wants resolution.

Set track.contentHint to motion, detail or text to tell the browser what the content is; it influences the default trade-off and some encoder tuning. The encoder can also be limited by CPU rather than bandwidth, and outbound statistics report which one is in charge through qualityLimitationReason: none, cpu, bandwidth or other. That one field separates a network problem from an underpowered laptop.

Advertisement

Configuring the send: simulcast, SVC and codecs

With two participants, one encoding is enough. With a selective forwarding unit and several receivers on different networks, the sender must offer choices, because the SFU cannot re-encode. Simulcast sends several independent encodings of the same track at different resolutions, each with its own RTP stream identifier (rid), and the SFU forwards whichever layer suits each receiver. Scalable video coding (SVC) sends one stream whose layers depend on each other, and the SFU drops layers it does not need. The configuration lives in the transceiver's send encodings:

// Three simulcast layers for an SFU, each capped, with resolution-first degradation.
const pc = new RTCPeerConnection(config);
const [track] = (await navigator.mediaDevices.getUserMedia({
  video: { width: 1280, height: 720, frameRate: 30 }
})).getVideoTracks();
track.contentHint = "motion";            // "detail" or "text" for screen share

const tx = pc.addTransceiver(track, {
  direction: "sendonly",
  sendEncodings: [
    { rid: "q", scaleResolutionDownBy: 4, maxBitrate: 150_000 },
    { rid: "h", scaleResolutionDownBy: 2, maxBitrate: 500_000 },
    { rid: "f", scaleResolutionDownBy: 1, maxBitrate: 1_500_000 },
  ],
});

// Prefer VP8; keep the rest as fallbacks. MDN: start from the receiver's list.
const codecs = RTCRtpReceiver.getCapabilities("video").codecs;
tx.setCodecPreferences([
  ...codecs.filter(c => c.mimeType === "video/VP8"),
  ...codecs.filter(c => c.mimeType !== "video/VP8"),
]);

// After negotiation: keep frame rate, give up resolution under pressure.
const params = tx.sender.getParameters();
params.degradationPreference = "maintain-framerate";
await tx.sender.setParameters(params);

Two details matter. First, the encodings and their caps set the ceiling; the bandwidth estimate decides what actually flows. If the estimate cannot sustain all three, the browser drops the top layer first, which is correct behaviour, not a bug. Second, SVC is configured with scalabilityMode on an encoding, defined in the W3C WebRTC-SVC extension. Names read as layer counts: L1T3 is one spatial layer with three temporal layers, L3T3_KEY is three spatial layers that depend on each other only at keyframes, and S2T1 is two simulcast spatial layers. Support differs by browser and codec; an unsupported mode rejects the call, so feature-detect and fall back to plain simulcast.

OptionHow the SFU adaptsCostUse when
Single encodingIt cannot; one quality for allLowest sender CPUOne-to-one calls
Simulcast (2-3 rids)Switches between independent streams, needs a keyframe to switchMore upload bits and CPUDefault for SFU calls; widest support
Temporal layers (L1T3)Drops frame-rate layers instantlySmall overheadCheap rate adaptation on top of either
Spatial SVC (L3T3_KEY, VP9/AV1)Drops resolution layers without a new keyframeCodec and browser support variesLarge rooms where switching cost hurts

Codec choice is negotiated, but setCodecPreferences orders what you offer. VP8 has the broadest simulcast support, H.264 often gets hardware encoders, and VP9 and AV1 compress better and carry SVC at higher CPU cost; see the codec comparison.

Loss recovery: retransmit, refresh or protect

When a packet is lost, the receiver has three tools, and the right one depends on round-trip time. NACK asks the sender to resend specific packets; the sender replays them from its store, usually on a separate RTX stream so retransmissions do not confuse the original sequence numbers. NACK is cheap and precise, but only helps if the resent packet beats the jitter buffer's deadline, so it struggles on long paths.

A PLI tells the sender the decoder has lost its reference and needs a keyframe. Keyframes are several times larger than delta frames, so a PLI costs a burst that, on a lossy link, can itself be lost and trigger another PLI. Forward error correction sends redundant packets proactively, so some losses can be repaired with no round trip at all, at a constant bitrate overhead whether or not loss happens. WebRTC stacks have used ULPFEC and FlexFEC for video, with uneven enablement across browsers; check what your target clients negotiate rather than assuming FEC is on.

MechanismLatency costBitrate costFails when
NACK + RTXOne round tripOnly lost packetsRTT approaches the jitter buffer wait
PLI keyframeOne round trip plus a large frameBurst of several delta frames' worthLoss is bursty; keyframe itself is lost
FECNone for repairable lossConstant overheadBurst loss exceeds the redundancy
Temporal layersNone; lose only enhancement framesSmallLoss hits base-layer frames

Keyframe interval interacts with all of this. WebRTC encoders do not need frequent periodic keyframes, because receivers ask for them, but an SFU switching simulcast layers needs one each time, and recording or late-joining receivers need one to start. The GOP structure article explains the reference structures that decide how far one loss propagates.

The receiver: jitter buffer, decode and why freezes happen

The jitter buffer trades latency for smoothness. Packets of one frame arrive spread over time; the buffer holds frames long enough to absorb that spread and give NACKs time to work, then releases them to the decoder. libwebrtc sizes it adaptively from observed jitter, so the playout delay grows on a jittery path and shrinks when the path calms down. Statistics expose the average wait as jitterBufferDelay divided by jitterBufferEmittedCount, and the current target as jitterBufferTargetDelay.

A freeze is a gap in rendered frames far longer than the frame interval, and the stats count them as freezeCount with totalFreezesDuration. The usual cause is a missing reference: a frame could not be completed in time, every later delta frame depends on it, and nothing renders until a keyframe arrives. So freeze length is roughly the time to notice the loss, plus a PLI round trip, plus the time to send a large keyframe through a constrained link. Shorten it with nearby media servers, temporal layers and affordable keyframes. A freeze with zero packet loss points elsewhere: a stalled decoder, a CPU-starved receiver dropping frames (framesDropped), or a sender that stopped producing frames.

Worked example: a call that is blurry, then frozen

A user on hotel Wi-Fi reports soft video that freezes for seconds at a time. Sample statistics on both ends every two seconds and diff the counters.

// Diff against the previous snapshot: cumulative totals hide what is happening now.
async function videoHealth(pc, prev = {}) {
  const out = {}, snap = {};
  const report = await pc.getStats();
  report.forEach(s => {
    snap[s.id] = s;
    if (s.type === "inbound-rtp" && s.kind === "video") {
      const d = (k) => (s[k] ?? 0) - (prev[s.id]?.[k] ?? 0);
      const emitted = d("jitterBufferEmittedCount");
      out.recv = {
        fps: s.framesPerSecond, res: `${s.frameWidth}x${s.frameHeight}`,
        lossPkts: d("packetsLost"), nacks: d("nackCount"), plis: d("pliCount"),
        freezes: d("freezeCount"), freezeSec: d("totalFreezesDuration"),
        jbMs: emitted ? 1000 * d("jitterBufferDelay") / emitted : null,
        dropped: d("framesDropped"),
      };
    }
    if (s.type === "outbound-rtp" && s.kind === "video") {
      out[`send_${s.rid ?? "0"}`] = { limit: s.qualityLimitationReason,
        res: `${s.frameWidth}x${s.frameHeight}`, fps: s.framesPerSecond,
        rtx: s.retransmittedPacketsSent, enc: s.encoderImplementation };
    }
    if (s.type === "transport" && s.selectedCandidatePairId) {
      const cp = report.get(s.selectedCandidatePairId);
      out.path = { bweKbps: Math.round((cp?.availableOutgoingBitrate ?? 0) / 1000),
                   rttMs: Math.round((cp?.currentRoundTripTime ?? 0) * 1000) };
    }
  });
  return { health: out, snap };
}

let snap = {};
setInterval(async () => {
  const r = await videoHealth(pc, snap);
  snap = r.snap;
  console.log(r.health);
}, 2000);

Illustrative readings from such a call: the sender's candidate pair shows an available outgoing bitrate around 400 kbps and an RTT near 250 ms; the top simulcast layer reports a small frame size and qualityLimitationReason of bandwidth. That explains the blur: the estimator found a thin uplink and the encoder shrank resolution, as configured. The receiver shows bursts of packet loss, many NACKs, several PLIs per minute and a freeze after each PLI burst, with the jitter buffer delay climbing.

The reading is a lossy, high-latency path where NACKs arrive too late and every recovery needs a keyframe the link can barely carry. Useful responses, in order: route the user to a closer media server to cut RTT so NACK works again; enable temporal layers so the SFU can shed frame rate before losses hit referenced frames; lower the top layer's cap so keyframes fit the link; and, if the client supports it, enable FEC for this call. If the sender instead showed cpu as the limitation, the fix would be a lighter codec or a hardware encoder, not the network.

Failure modes

SymptomLikely causeWhere to lookFix
Persistently low resolutionBandwidth estimate or CPU limitqualityLimitationReason, available outgoing bitrateNetwork path, codec or hardware encoder
Repeated freezesLost references, PLI round tripspliCount, freezeCount, RTTCloser servers, temporal layers, lower caps
Top simulcast layer missingEstimate cannot carry all layersPer-rid outbound statsExpected; tune caps or layer count
Video lags audioLarge jitter buffer on a jittery pathjitterBufferDelay per frameReduce jitter; check Wi-Fi and queues
Black video, audio fineCodec mismatch or no keyframe after switchNegotiated codecs, keyFramesDecodedCodec preferences; request keyframe on switch

Operating it

Export a small per-call summary rather than raw samples: time in each quality limitation, freeze seconds per minute of video, NACK and PLI rates, p95 RTT and the received resolution distribution. Freeze seconds per minute tracks user complaints far better than average bitrate. Segment by network type, region, browser and media server, and alert on step changes after deploys and browser releases. For multi-party systems, the SFU article covers how the server uses these signals to pick layers per receiver.

What to do next

  1. Write a stats sampler like the one above and log per-interval deltas for a few real calls, including at least one on a poor network.
  2. Set contentHint and degradationPreference explicitly for camera and screen tracks instead of relying on defaults.
  3. Configure simulcast with explicit per-layer bitrate caps, and add temporal layers where your browsers and SFU support them.
  4. Measure RTT to your media servers by region; if NACK cannot beat the jitter buffer, add closer servers before tuning anything else.
  5. Track freeze seconds per minute and quality-limitation time per call as your primary video health metrics.
  6. Re-run your poor-network test after every browser major release and SFU deploy.
Key takeaway: WebRTC video is a feedback loop: receiver reports become a bandwidth estimate, the estimate becomes an encoder target, and the encoder trades quantizer, resolution and frame rate to meet it, while NACK, keyframe requests and FEC repair losses on the way. Blur usually means a bandwidth or CPU limit, freezes usually mean a lost reference waiting on a keyframe round trip. Configure the trade-offs explicitly, diff getStats counters to see which stage is failing, and fix that stage.