A WebRTC call feels live when a movement in front of one camera appears on the other screen within about a quarter of a second. That number, glass-to-glass latency, is the sum of roughly ten stages, and each one can quietly double.
This article follows one frame from the camera sensor to the remote display. It explains what each stage does, how the frame keeps its identity through RTP timestamps, why frame size turns into delay, and which browser statistics measure which stage. The system architecture of ICE, SFUs and simulcast is covered in WebRTC video pipeline architecture, and the feedback loop between bandwidth estimation, encoder and loss recovery is covered in WebRTC video, in depth. Here the focus is the path itself and the stopwatch you hold against it.
The stages a frame passes through
On the sender, the camera exposes a frame and the driver hands it to the browser. The frame may be scaled or converted to the encoder's pixel format. The encoder compresses it under a bitrate target, producing either a keyframe that decodes on its own or a delta frame that depends on earlier frames. The packetizer cuts the encoded frame into RTP packets small enough to avoid IP fragmentation, typically around 1,200 bytes of payload. The pacer releases packets smoothly rather than in one burst, and each packet is encrypted with SRTP and sent over UDP.
In the network the packets may queue at a congested link, be lost, arrive out of order or pass through a selective forwarding unit. On the receiver, packets are collected until the frame is complete, the frame waits in the jitter buffer until its playout time, the decoder reconstructs pixels, and the renderer hands them to the compositor, which shows them at the next display refresh.
| Stage | What adds delay | Measured by |
|---|---|---|
| Capture | Exposure time, driver buffering | Filmed clock only |
| Encode | Encoder speed, lookahead | totalEncodeTime / framesEncoded |
| Pace and send | Frame size versus pacing rate | totalPacketSendDelay / packetsSent (per packet) |
| Network | Propagation, queueing, SFU hop | currentRoundTripTime (round trip) |
| Assemble | Waiting for the last packet, retransmits | totalAssemblyTime |
| Jitter buffer | Target delay chosen by the receiver | jitterBufferDelay / jitterBufferEmittedCount |
| Decode | Decoder speed, hardware or software | totalDecodeTime / framesDecoded |
| Render | Compositor, vsync | requestVideoFrameCallback |
How a frame keeps its identity: RTP timestamps
Every RTP packet carries a sequence number, which increases by one per packet, and a timestamp, which identifies the frame. For video the timestamp runs on a 90 kHz clock, and all packets belonging to one frame share the same value. At 30 frames per second consecutive frames differ by 3,000 ticks. The last packet of a frame has the marker bit set, which tells the receiver the frame is complete once every sequence number from the first packet to the marked one has arrived. The 32-bit timestamp wraps after about 13.3 hours at 90 kHz, so receivers compare timestamps with wrap-around arithmetic, never as plain integers.
The RTP timestamp starts at a random offset and says nothing about wall-clock time on its own. RTCP sender reports provide the mapping: each report pairs an RTP timestamp with the sender's NTP wall-clock time. The receiver uses those pairs to align audio and video from the same sender (lip sync) and to estimate when a frame was captured. That estimate is only as good as the receiver's knowledge of the offset between the two machines' clocks, which is why browser capture-time figures for remote video are estimates, not facts.
# Wrap-safe comparison of 32-bit RTP timestamps (the same idea applies to 16-bit sequence numbers)
def ts_newer(a, b):
diff = (a - b) & 0xFFFFFFFF
return diff != 0 and diff < 0x80000000
def ticks_to_ms(ticks, clock_hz=90_000):
return 1000.0 * ticks / clock_hz
# 30 fps: consecutive frames are 3,000 ticks apart, i.e. 33.3 ms
assert ts_newer(3000, 0) and not ts_newer(0, 3000)
assert ts_newer(5, 0xFFFFFF00) # survives the wrap
print(ticks_to_ms(3000)) # 33.33...
Encode and send: where frame size becomes delay
A frame cannot be shown until its last packet arrives, so a large frame is a slow frame. Consider a 2 Mbps stream at 30 frames per second: the average frame is about 8.3 KB, or seven packets. A keyframe is often several times larger than a delta frame because it cannot reference anything. Suppose it is 60 KB, about 50 packets. Even if the pacer sends at twice the target bitrate, 60 KB takes 120 ms to leave the sender, and every frame queued behind it waits too. That is why a call stutters briefly each time a keyframe is produced, and why a receiver that keeps asking for keyframes makes the problem worse.
The encoder's rate controller is the other half. When bandwidth estimation lowers the target, the encoder raises its quantizer (QP) to make frames smaller, then reduces resolution or frame rate if that is not enough. The outbound statistic qualityLimitationReason reports whether the sender is currently limited by bandwidth, cpu, other or none, and qualityLimitationDurations gives the cumulative seconds in each state. A sender limited by CPU produces slow encodes and dropped frames; a sender limited by bandwidth produces small, blurry frames on time.
Packetization follows a codec-specific RTP payload format. H.264 uses RFC 6184, which fragments a large NAL unit across packets; VP8 uses RFC 7741.
The receive side: assembly, jitter buffer, decode, render
Assembly ends when the last packet arrives. If a packet is missing, the receiver sends a NACK and the sender retransmits from a short history; recovery costs at least one round trip. If a packet in a reference frame is never recovered, every later delta frame that depends on it is undecodable, so the receiver sends a picture loss indication (PLI) asking for a new keyframe, and video freezes until it arrives.
The jitter buffer exists because packets do not arrive at even intervals. It holds each complete frame until a playout time chosen so that most frames arrive before they are due. The browser adapts this target delay to the measured arrival jitter and, when retransmission is in use, leaves room for a round trip. Applications can influence it with RTCRtpReceiver.jitterBufferTarget, a value in milliseconds between 0 and 4000; it is a hint that the browser clamps to what it can provide, not a command. Raising it trades latency for fewer freezes, which suits one-way broadcasts; interactive calls usually leave it unset.
Decoding normally takes a few milliseconds per frame with hardware decoders and more with software decoders at high resolutions. Finally the frame waits for the compositor and the next display refresh, which adds up to one refresh interval: 16.7 ms on a 60 Hz screen.
Measuring each stage from the browser
The WebRTC statistics API returns cumulative counters, and this is the most common source of misreading. totalDecodeTime is the total seconds spent decoding since the stream started, not the time per frame. To get a per-frame average, take two samples a second apart and divide the change in the total by the change in its matching counter. Note that jitterBufferDelay is measured from when a frame's first packet reaches the buffer, so it includes assembly time.
// Sample once per second on each peer; per-frame figures come from deltas between samples.
let prev = new Map();
async function stageTimes(pc) {
const report = await pc.getStats();
const out = {};
report.forEach((s) => {
const p = prev.get(s.id);
if (!p) return;
const d = (k) => (s[k] ?? 0) - (p[k] ?? 0);
const perFrameMs = (total, count) => (d(count) > 0 ? (1000 * d(total)) / d(count) : null);
if (s.type === 'outbound-rtp' && s.kind === 'video') { // sender side
out.encodeMs = perFrameMs('totalEncodeTime', 'framesEncoded');
out.pacerMsPerPacket = perFrameMs('totalPacketSendDelay', 'packetsSent');
out.limitedBy = s.qualityLimitationReason;
}
if (s.type === 'inbound-rtp' && s.kind === 'video') { // receiver side
out.assemblyMs = perFrameMs('totalAssemblyTime', 'framesAssembledFromMultiplePackets');
out.jitterBufferMs = perFrameMs('jitterBufferDelay', 'jitterBufferEmittedCount');
out.decodeMs = perFrameMs('totalDecodeTime', 'framesDecoded');
out.dropped = d('framesDropped');
out.freezes = d('freezeCount');
out.nacks = d('nackCount');
out.plis = d('pliCount');
}
if (s.type === 'candidate-pair' && s.nominated && s.state === 'succeeded') {
out.rttMs = 1000 * (s.currentRoundTripTime ?? 0);
}
});
prev = new Map([...report]);
return out;
}The tail of the pipeline is visible through video.requestVideoFrameCallback. Its metadata includes expectedDisplayTime and, for WebRTC sources, receiveTime (when the last packet of the frame arrived), rtpTimestamp and captureTime. For remote video the capture time is derived from the RTP timestamp and the sender's clock mapping, so treat capture-to-display figures as estimates and check them against a filmed clock: point the sending camera at a millisecond timer, put the receiving screen next to the timer, photograph both, and subtract. That is the only measurement that includes camera exposure and display latency.
const video = document.querySelector('video');
function onFrame(now, m) {
if (m.receiveTime !== undefined) record('receive_to_display_ms', m.expectedDisplayTime - m.receiveTime);
if (m.captureTime !== undefined) record('capture_to_display_ms_est', m.expectedDisplayTime - m.captureTime);
video.requestVideoFrameCallback(onFrame);
}
video.requestVideoFrameCallback(onFrame);
Worked example: a call that should be 150 ms and measures 260 ms
A support team reports that remote assistance calls feel sluggish. A filmed clock shows about 260 ms glass to glass. Collecting per-stage numbers from both peers gives the budget below. The figures are illustrative of a typical laptop-to-laptop call through one SFU, not benchmarks.
| Stage | Measured | Healthy for this setup |
|---|---|---|
| Capture and driver | about 25 ms (filmed clock minus the rest) | 20 to 35 ms |
| Encode | 9 ms | 5 to 15 ms |
| Pacer | 6 ms, spikes to 90 ms on keyframes | under 10 ms |
| Network one way, via SFU | about 45 ms (RTT 90 ms) | depends on geography |
| Jitter buffer, including assembly | 150 ms | 40 to 70 ms |
| Decode | 4 ms | under 10 ms |
| Render and display | about 20 ms | one or two refreshes |
The jitter buffer is the outlier. The receiver's statistics also show NACK counts climbing every second and a PLI every few seconds. The viewing laptop is on a congested Wi-Fi network: packets arrive in bursts and some are lost, so the buffer has grown its target to absorb both the arrival jitter and retransmission round trips. The PLIs cause keyframes, which cause the pacer spikes on the sender, which add more burstiness. Moving the viewer to a wired connection brought the jitter buffer to about 55 ms and glass to glass to roughly 160 ms. Where the network cannot change, capping resolution so that keyframes are smaller and enabling forward error correction for the worst links are the usual mitigations.
Adding your own stage: encoded transforms
Sometimes you need to touch encoded frames: end-to-end encryption through an SFU, or attaching a frame identifier for measurement. RTCRtpScriptTransform inserts a stream transform, running in a worker, between the encoder and packetizer on the sender, or between assembly and decoder on the receiver. Any work done here adds directly to the latency of every frame, so keep it allocation-free and measure it.
// main thread: attach right after addTrack so the transform sees the first frame
const worker = new Worker('frame-transform.js');
const sender = pc.addTrack(track, stream);
sender.transform = new RTCRtpScriptTransform(worker, { side: 'send' });
// frame-transform.js (worker)
onrtctransform = (event) => {
const { readable, writable } = event.transformer;
readable
.pipeThrough(new TransformStream({
transform(frame, controller) {
// frame.data is an ArrayBuffer holding the encoded frame; frame.getMetadata() has its RTP details
controller.enqueue(frame);
},
}))
.pipeTo(writable);
};
Failure modes
| Symptom | Stage | Likely cause and fix |
|---|---|---|
| Delay grows during the call, then stays high | Jitter buffer | Bursty or lossy link; fix the network, or cap resolution to shrink keyframes |
| Brief stutter every few seconds | Pacer and keyframes | Frequent PLIs or a short keyframe interval; find the receiver asking for keyframes |
| Long freezes after packet loss | Decode | Lost reference frame; check PLI round trip, consider FEC or SVC layers |
| Low frame rate, sharp image | Encode | CPU-limited sender; check qualityLimitationReason, use hardware encoding |
| Blurry image, smooth motion | Encode | Bandwidth-limited; expected behaviour, check the estimate is not stuck low |
| Fine in stats, laggy on screen | Capture or render | Camera buffering or compositor; use the filmed clock to find it |
Operating it and the trade-offs
Collect the per-stage numbers from both peers into the same time series, tagged with call and participant, so a single dashboard shows the whole path. Alert on the jitter buffer and freeze counts rather than on bitrate, since users notice delay and freezes more than resolution. When you add an SFU, it adds a hop and its own queue, so remember that its egress link to each receiver is a separate bottleneck; the SFU guide covers that layer.
Every knob in this pipeline trades one property for another. A bigger jitter buffer means fewer freezes and more delay. A longer keyframe interval saves bandwidth and slows recovery after loss. Higher resolution makes frames larger and so slower to send. If your use case can tolerate a second or more of delay, chunked HTTP delivery is usually cheaper at scale; low-latency live streaming compares it with WebRTC.
What to do next
- Film a millisecond clock through your own call to get a true glass-to-glass number before changing anything.
- Add the getStats sampler above to both peers and chart per-frame encode, pacer, jitter buffer and decode times.
- Check that you divide cumulative totals by their counters; raw totals look like huge delays.
- Add requestVideoFrameCallback on the receiver to measure receive-to-display.
- Find the largest stage in the budget and fix that one first; it is usually the jitter buffer or the network.
- Watch PLI and keyframe counts; reduce keyframe size by capping resolution if bursts drive delay.
- Leave jitterBufferTarget unset for interactive calls; raise it only for one-way viewing.