A WebRTC voice call looks simple from the API: get a microphone track, add it to a peer connection, play the remote track. Underneath, every 20 milliseconds a frame of sound passes through a capture driver, an echo canceller, a noise suppressor, a gain controller, a codec, an encryption layer, the network, an adaptive jitter buffer, a decoder that may invent audio for lost packets, a mixer and a playout driver, and each stage can add delay or damage.

This page follows that frame end to end, as implemented in browsers built on the open-source libwebrtc stack, and gives you the numbers to reason about it: packet sizes, bandwidth on the wire, a mouth-to-ear latency budget and the getStats fields that show where a bad call went wrong. Video and connection setup are covered elsewhere; here the focus is audio.

Advertisement

The path at a glance

One direction of a WebRTC call, plus the echo reference loopSenderMicrophone + ADM10 ms capture framesAudio processingAEC, NS, AGCOpus encoder20 ms frames, FEC, DTXRTP + SRTPseq, timestamp, SSRCICE / DTLS / UDPjitter, loss, reorderingReceiverSRTP decryptRTCP feedbackNetEqjitter buffer, decode, PLCMixer + playout10 ms render framesSpeaker + ADMReceiver's own APM gets the render signal as echo referenceEach side runs both halves:its capture path and its playoutpath meet inside its own APM.
Sender capture path, network, receiver playout path. The receiver's playout signal is fed back into its own audio processing as the echo reference.

Two loops matter. The obvious one is the media path from one mouth to the other ear. The less obvious one is local: what a device plays through its speaker leaks back into its microphone, so the playout signal must be handed to the local echo canceller as a reference. Most audio quality problems in calls come from one of these loops being broken, delayed or mis-timed.

Capture and the audio device module

The audio device module talks to the operating system's audio API, receives microphone samples in small hardware buffers and repackages them into 10 ms frames, the unit the rest of the pipeline processes. At 48 kHz a 10 ms mono frame is 480 samples. Device buffers add delay that varies a lot by platform and driver: a few milliseconds on a well-configured desktop, tens of milliseconds on some Bluetooth headsets and older Android devices.

Microphone and speaker each run on their own clocks, which never quite agree with each other or with the far end. The pipeline must absorb that drift, and the echo canceller must cope with a capture-to-render delay that is unknown and can change mid-call, for example when a user switches output to a headset.

From the browser, processing is requested through constraints. The defaults suit speech; music, or a studio microphone with headphones, needs them off.

// Voice call: keep the platform's processing on (the defaults in browsers).
const voice = await navigator.mediaDevices.getUserMedia({
  audio: { echoCancellation: true, noiseSuppression: true, autoGainControl: true,
           channelCount: 1 }
});

// Music or a studio mic with headphones: turn processing off so it is not
// gated, pumped or 'denoised' away. Check what the browser actually applied.
const music = await navigator.mediaDevices.getUserMedia({
  audio: { echoCancellation: false, noiseSuppression: false, autoGainControl: false,
           channelCount: 2 }
});
console.log(music.getAudioTracks()[0].getSettings());
Advertisement

Audio processing: echo, noise and gain

libwebrtc's audio processing module runs on each 10 ms capture frame. The current echo canceller is AEC3. It estimates the delay between the render reference and the echo in the microphone signal, models the echo path with adaptive filters, subtracts the estimate, then applies a nonlinear suppressor to remove what remains, being careful during double-talk when both people speak at once. If the delay estimate is wrong or the reference is missing, echo leaks through or near-end speech is chopped. Acoustic echo cancellation explains adaptive filtering and double-talk in detail.

Noise suppression estimates the stationary noise floor per frequency band and attenuates bands dominated by noise, with optional neural suppressors in some products; noise suppression compares approaches. Automatic gain control, AGC2 in current libwebrtc, brings speech to a consistent level and limits peaks. A level estimate and voice activity decision from this stage also feed the encoder's discontinuous transmission and the audio-level header extension that conferencing servers use to pick active speakers.

Order matters: echo cancellation must see the signal before gain changes it, or the echo path appears to change every time the gain moves.

Encoding with Opus

WebRTC requires every endpoint to implement Opus and G.711 (RFC 7874), and browsers prefer Opus, which handles everything from narrowband speech to stereo music. Browsers normally send 20 ms frames: long enough for efficient compression, short enough for latency. Opus always uses a 48 kHz RTP clock regardless of the internal bandwidth, so the RTP timestamp advances by 960 per 20 ms frame. Speech at 24 to 32 kbit/s mono sounds good; music benefits from stereo and 64 kbit/s or more.

Two options trade bandwidth for robustness. In-band forward error correction embeds a coarse copy of the previous frame in the current packet, so a single lost packet can be rebuilt from the next one at reduced quality. Discontinuous transmission stops sending full frames during silence and sends occasional comfort-noise updates, saving bandwidth in calls where most participants are quiet. Both are negotiated through SDP format parameters.

m=audio 9 UDP/TLS/RTP/SAVPF 111
a=rtpmap:111 opus/48000/2
a=fmtp:111 minptime=10;useinbandfec=1;usedtx=1
a=extmap:1 urn:ietf:params:rtp-hdrext:ssrc-audio-level
a=rtcp-fb:111 transport-cc

The Opus for real-time voice page covers bitrate, complexity and FEC settings in depth.

RTP, SRTP and what a packet costs

Each encoded frame becomes one RTP packet. The 12-byte RTP header carries a sequence number for loss and reorder detection, the timestamp for playout timing, and the SSRC identifying the stream. Header extensions add the audio level and transport-wide sequence numbers for congestion control. SRTP encrypts the payload with keys negotiated over DTLS and appends an authentication tag. RTCP reports carry loss, jitter and round-trip time back to the sender.

Worked bandwidth example: Opus at 32 kbit/s in 20 ms frames yields 80 bytes of payload per packet. Add 12 bytes of RTP, 10 bytes of SRTP authentication tag with the common 80-bit HMAC profile (AES-GCM suites use 16), 8 bytes of UDP and 20 bytes of IPv4, ignoring header extensions, and the packet is 130 bytes. At 50 packets per second that is 52 kbit/s on the wire, more than one and a half times the codec bitrate. Moving to 10 ms frames would double the header overhead; moving to 60 ms would cut it but add 40 ms of delay and make each loss more audible. This is why 20 ms is the usual compromise.

Audio shares the transport with video in most calls, using the same ICE candidate pair and congestion controller; the WebRTC video pipeline covers ICE, SFUs and bandwidth estimation.

The receiver: NetEq, jitter and concealment

Packets arrive with variable delay, sometimes out of order, sometimes not at all. libwebrtc's NetEq combines the jitter buffer, the decoder and loss handling. It keeps a target buffer level derived from recent arrival-delay statistics, deep enough to absorb typical jitter, shallow enough to keep delay low, and every 10 ms it decides how to produce the next output block.

Its operations are worth knowing by name because stats report them. Normal decodes a packet as is. Expand generates audio to cover a missing packet, packet loss concealment, by extending the pitch and spectral shape of recent speech and fading toward silence if loss persists. Merge blends concealed audio back into real audio when packets resume. Accelerate removes whole pitch periods to shrink the buffer when it is too full, and preemptive expand inserts pitch periods to grow it when too empty; both change timing without changing pitch. Comfort noise fills DTX silences. Adaptive jitter buffers covers the delay estimation behind the target level.

The output then goes to the mixer, which sums remote streams, and to the device module for playout. That same mixed signal is the render reference passed to the local echo canceller.

A mouth-to-ear latency budget

StageTypical one-way contributionWhat moves it
Capture device buffering5-20 msOS audio API, driver, Bluetooth
Audio processinga few msMostly fixed by frame structure
Opus framing and encode20 ms + a few msFrame size; codec lookahead
Network10-150 msDistance, routing, relays, queueing
Jitter buffer20-100 msNetwork jitter; adapts continuously
Decode, mix, playout buffering10-30 msOS audio API and device
Totalroughly 70-320 msInteractive speech degrades as it climbs past about 150 ms

These ranges are illustrative, not measured guarantees; measure your own. The table shows where effort pays. Network and jitter buffer delay dominate and depend on routing and relay placement more than on code. Device buffering is the next largest and is often the hidden cause of a call that measures fine but feels sluggish, especially over Bluetooth. ITU-T G.114 guidance treats one-way delay below about 150 ms as acceptable for most interactive uses, which leaves little slack once relays and a busy Wi-Fi link are involved.

Observability with getStats

The W3C statistics API exposes the receiver's view in the inbound-rtp report. Average jitter buffer delay is jitterBufferDelay divided by jitterBufferEmittedCount; the concealment ratio is concealedSamples over totalSamplesReceived; insertedSamplesForDeceleration and removedSamplesForAcceleration show how much time-stretching NetEq is doing. The remote-inbound-rtp report gives round-trip time as seen by the far end. Always work with deltas between polls, because the counters are cumulative since the stream started.

async function audioHealth(pc, prev = {}) {
  const out = {};
  (await pc.getStats()).forEach(s => {
    if (s.type === 'inbound-rtp' && s.kind === 'audio') {
      const d = k => s[k] - (prev[k] ?? 0);           // deltas since last poll
      out.lossRate      = d('packetsLost') / Math.max(1, d('packetsLost') + d('packetsReceived'));
      out.jitterMs      = s.jitter * 1000;
      out.jbDelayMs     = 1000 * d('jitterBufferDelay') / Math.max(1, d('jitterBufferEmittedCount'));
      out.jbTargetMs    = 1000 * d('jitterBufferTargetDelay') / Math.max(1, d('jitterBufferEmittedCount'));
      out.concealedPct  = 100 * d('concealedSamples') / Math.max(1, d('totalSamplesReceived'));
      out.stretchPct    = 100 * (d('insertedSamplesForDeceleration') + d('removedSamplesForAcceleration'))
                              / Math.max(1, d('totalSamplesReceived'));
      Object.assign(prev, s);
    }
    if (s.type === 'remote-inbound-rtp' && s.kind === 'audio') out.rttMs = s.roundTripTime * 1000;
  });
  return out;   // poll every few seconds and alert on trends, not single samples
}

A useful triage: high loss with low jitter points to a congested or lossy link; high jitter with a growing buffer target points to Wi-Fi or a bufferbloated router; heavy concealment with low loss points to late packets being discarded; healthy numbers with user complaints point back to the device, echo or processing settings.

Failure modes

  • Echo heard by the far end. The echo canceller is missing or mis-timing its reference: audio played outside the WebRTC playout path, a delay change after a device switch, or a nonlinear loudspeaker. Ask the far end, since you never hear your own echo.
  • Choppy near-end speech. Echo suppression too aggressive during double-talk, or noise suppression treating speech as noise. Test with real rooms, not quiet desks.
  • Pumping volume. Gain control fighting a hardware or OS gain stage. Leave only one AGC in the chain.
  • Music sounds underwater. Speech processing on music. Disable the constraints and raise the bitrate.
  • Robotic or warbling audio. Heavy concealment and time-stretching from bursty loss or late packets. Check concealment and loss deltas. Packet loss concealment explains the artefacts.
  • First syllable clipped. DTX or voice activity detection resuming late after silence.
  • Narrowband sound on a headset. Some Bluetooth headset modes limit the microphone path to 8 or 16 kHz sampling. Check the device's mode.
  • Slowly growing delay. Clock drift or a playout path not draining; check the jitter buffer target against actual delay.

Trade-offs

Every lever trades latency against robustness or bandwidth. Longer frames save header bandwidth and cost delay. FEC and redundancy protect against loss and cost bitrate, and FEC only helps for isolated losses. A deeper jitter buffer means fewer concealment events and more delay. Aggressive noise suppression gives cleaner silence and riskier speech. Turning processing off gives music fidelity and requires headphones. Decide per use case: conversational voice favours low delay with moderate FEC; broadcast-style one-to-many audio can afford a deeper buffer and higher bitrate.

What to do next

  1. Add the getStats poller to your client and log loss, jitter, buffer delay, concealment and RTT per call.
  2. Capture a packet trace of one call and confirm 20 ms framing, the negotiated fmtp parameters and the packet size you expect.
  3. Measure mouth-to-ear delay with a loopback clap test across your main device types, including one Bluetooth headset.
  4. Test echo with a laptop on speakers in a real room, then switch output devices mid-call.
  5. Decide processing constraints per use case: speech defaults, music with processing off.
  6. Simulate 2 percent and 10 percent loss and 50 ms jitter with a network emulator and listen to the concealment.
  7. Alert on concealment and jitter buffer trends per network type, not on single samples.
Key takeaway: A WebRTC voice frame passes through capture buffering, echo cancellation, noise suppression and gain control, Opus encoding, RTP and SRTP, the network, NetEq's adaptive jitter buffer and concealment, and the playout path, whose output loops back as the echo reference. Network and jitter buffer delay dominate the latency budget, device buffering is the hidden third, and header overhead makes the wire rate far higher than the codec rate. Instrument calls with getStats deltas, test real devices and rooms, and tune frame size, FEC, processing and buffering per use case.