Every real-time voice system needs a jitter buffer: packets leave the sender every 20 ms but arrive unevenly, and the listener must hear a steady stream. The theory of how to estimate delay and decide when to stretch or conceal audio is covered in jitter buffer design. This article is about building one: how to structure the code so it can be tested without a network, how to tune it against recorded traffic instead of intuition, and which real traffic patterns break designs that look fine on paper.
The running idea: an adaptive jitter buffer is a small control system between two clocks, the network pushing packets and the audio device pulling samples. Isolate everything between them behind two functions driven by a fake clock, and you can replay recorded traffic through the exact production code and compare designs in seconds.
The component boundary
The network thread calls insert(packet, arrival_time) whenever a packet arrives. The audio thread calls pull(10 ms) whenever the device needs the next block. Store, estimator, controller and playout decision live inside and see time only through arguments; nothing reads a wall clock or sleeps.
Two rules follow. The calls arrive on different threads, so guard shared state with a short lock or a lock-free queue; never let the audio callback wait on the network. And log every insert and pull decision with timestamps into a ring buffer: that log is your replay input when a user reports choppy audio.
The packet store
The store holds packets by media timestamp and answers one question: is the next frame here? The bugs live in counters that wrap. RTP sequence numbers are 16 bits and wrap after 65,536 packets, about 22 minutes at 50 packets a second. Timestamps are 32 bits and wrap after about 24.8 hours at 48 kHz, but they start at a random value, so a wrap can happen in the first minute. Extend both to monotonic integers on insert.
from dataclasses import dataclass
SEQ_MOD, TS_MOD = 1 << 16, 1 << 32
def unwrap(value, last_ext, mod):
# Extend a wrapping counter (16-bit seq, 32-bit timestamp) to a monotonic integer.
if last_ext is None:
return value
base = last_ext - (last_ext % mod)
cand = base + value
if cand - last_ext > mod // 2:
cand -= mod
elif last_ext - cand > mod // 2:
cand += mod
return cand
@dataclass
class Packet:
seq: int # extended
ts: int # extended, in samples
payload: bytes
arrival_ms: float
class PacketStore:
def __init__(self, max_packets=250):
self.pkts = {} # ts -> Packet
self.max = max_packets
self.last_seq = self.last_ts = None
def insert(self, seq, ts, payload, arrival_ms, playout_ts):
self.last_seq = unwrap(seq, self.last_seq, SEQ_MOD)
self.last_ts = unwrap(ts, self.last_ts, TS_MOD)
ts_ext = self.last_ts
if playout_ts is not None and ts_ext < playout_ts:
return "late" # too late to play: count it, drop it
if ts_ext in self.pkts:
return "duplicate" # retransmission or network duplicate
if len(self.pkts) >= self.max:
self.pkts.pop(min(self.pkts)) # overflow: drop oldest, then flush-recover
self.pkts[ts_ext] = Packet(self.last_seq, ts_ext, payload, arrival_ms)
return "stored"
def pop(self, ts):
return self.pkts.pop(ts, None)Count late, duplicate and stored packets separately: late arrivals mean the target was too small, while duplicates point at retransmission or a misbehaving middlebox. Cap the store so a flooding sender cannot grow memory, and treat overflow as a signal to resynchronise.
Estimating delay and setting the target
The estimator measures how late each packet is relative to the fastest one seen. That removes the unknown one-way delay and clock offset, leaving the variable part the buffer must absorb. A histogram of relative delays, with old observations decaying through a forgetting factor, gives a quantile such as the 95th percentile: the delay that 95 percent of packets beat. The target buffer level is that quantile plus one frame.
class DelayEstimator:
# Relative delay histogram with exponential forgetting (10 ms buckets).
def __init__(self, bucket_ms=10, buckets=100, forget=0.998, quantile=0.95):
self.b, self.h = bucket_ms, [0.0] * buckets
self.forget, self.q = forget, quantile
self.ref = None # (arrival_ms, ts_ms) of fastest packet seen
def update(self, arrival_ms, ts_ms):
if self.ref is None or (arrival_ms - ts_ms) < (self.ref[0] - self.ref[1]):
self.ref = (arrival_ms, ts_ms) # new minimum transit: reset baseline
rel = (arrival_ms - self.ref[0]) - (ts_ms - self.ref[1])
self.h = [x * self.forget for x in self.h]
self.h[min(int(rel // self.b), len(self.h) - 1)] += 1 - self.forget
def quantile_ms(self):
total, acc = sum(self.h), 0.0
for i, x in enumerate(self.h):
acc += x
if total and acc / total >= self.q:
return (i + 1) * self.b
return len(self.h) * self.b
class TargetController:
def __init__(self, frame_ms=20, min_ms=20, max_ms=500, down_hold_ms=2000, down_step_ms=10):
self.frame, self.min, self.max = frame_ms, min_ms, max_ms
self.hold, self.step = down_hold_ms, down_step_ms
self.target, self.below_since = min_ms + frame_ms, None
def update(self, wanted_ms, now_ms):
wanted = max(self.min, min(self.max, wanted_ms + self.frame))
if wanted > self.target: # rise at once: late audio is worse than delay
self.target, self.below_since = wanted, None
elif wanted < self.target - self.step: # fall slowly, only after a sustained calm
self.below_since = self.below_since or now_ms
if now_ms - self.below_since >= self.hold:
self.target -= self.step
self.below_since = now_ms
else:
self.below_since = None
return self.targetThe controller raises the target immediately when the network worsens, because each late packet is an audible glitch, and lowers it slowly, one small step after a sustained calm, because a network that just calmed down often worsens again. The forgetting factor sets memory: with 0.998 per packet and 50 packets a second, an observation's weight halves in about 7 seconds.
The playout decision
On each pull the buffer picks one action. With the next frame present and the level near target, decode normally. If the level is well above target, accelerate: remove one pitch period from a low-energy stretch of speech, shortening the audio slightly without changing its pitch. If the level is below target, decelerate by inserting a pitch period. If the frame is missing during speech, ask the decoder to conceal it; if it is missing during a silence the sender declared with discontinuous transmission, play comfort noise instead. When real audio returns after concealment, merge it into the concealment tail so there is no click.
def decide(level_ms, target_ms, have_next, prev_was_concealed, in_talkspurt):
# One decision per 10 ms pull. level_ms = audio buffered ahead of playout.
if not have_next:
return "comfort_noise" if not in_talkspurt else "conceal"
if prev_was_concealed:
return "merge" # blend real audio into the concealment tail
if level_ms > target_ms + 20:
return "accelerate" # remove a pitch period, during low energy
if level_ms < target_ms - 20:
return "decelerate" # insert a pitch period before it runs dry
return "normal"Keep the margins, here 20 ms either side of target, wide enough to avoid constant stretching; a buffer that stretches every second sounds subtly wrong even with zero concealment. Concealment quality itself depends on the codec; Opus for real-time voice and packet loss concealment explain what the decoder can do with a missing frame.
A trace-driven simulator
Because the component only sees time through its arguments, a simulator is a short loop: feed recorded arrivals in order up to the current fake time, pull 10 ms, record the action and buffer level, and advance the clock. Use consented real traces plus synthetic ones for Wi-Fi scans, handovers and congested uplinks.
def simulate(trace, jb, pull_ms=10, duration_ms=60_000):
# trace: list of (arrival_ms, seq, ts, payload), sorted by arrival.
# jb: the production component, driven by a fake clock.
events, i, now = [], 0, 0.0
while now < duration_ms:
while i < len(trace) and trace[i][0] <= now:
a, seq, ts, pl = trace[i]
jb.insert(seq, ts, pl, arrival_ms=a)
i += 1
frame, action = jb.pull(pull_ms, now_ms=now)
events.append((now, action, jb.level_ms(), jb.target_ms()))
now += pull_ms
return score(events)
def score(events):
n = len(events)
conceal = sum(1 for e in events if e[1] == "conceal") / n
stretch = sum(1 for e in events if e[1] in ("accelerate", "decelerate")) / n
mean_delay = sum(e[2] for e in events) / n
return {"conceal_ratio": conceal, "stretch_ratio": stretch, "mean_level_ms": mean_delay}Score every design on three numbers together: concealment ratio, stretch ratio and mean buffer level, which is the delay you added. No single number is enough. A design that reaches zero concealment by holding 400 ms of audio has failed conversation. Plot concealment against delay and choose from the frontier, with separate operating points for conversation and listen-mostly streams.
Worked example: a delay spike
Take 20 ms Opus frames on a steady path, and a target of 40 ms. Packets 1 to 10 arrive on time. Packets 11 to 15 are delayed by a Wi-Fi channel scan: all five arrive together 120 ms after packet 11 was due. Packet 16 onwards are on time again.
With the target fixed at 40 ms, the buffer has 40 ms of audio when packet 11 should play. It plays packets 9 and 10, then finds 11 missing and conceals. It keeps concealing for the 120 ms gap minus the 40 ms it held, so 80 ms, four frames, are synthesised. Packet 11's slot was 40 ms after it was due, so it arrives 80 ms too late, and 12, 13 and 14 miss their slots by 60, 40 and 20 ms. The store reports those four as late and drops them, and the listener hears With the adaptive design, one isolated spike changes little: five delayed packets among the several hundred the histogram effectively remembers are about 1 percent of its mass, too little to move a 95th percentile. That is deliberate; one glitch should not cost delay for the rest of the call. If scans recur every 2 seconds, as on some laptops, delayed packets become about 5 percent of recent traffic, with relative delays spread from 40 to 120 ms because the burst arrives together. The 95th percentile climbs into that range and the controller raises the target at once, so later spikes conceal fewer frames or none, at the cost of added delay. Whether 0.95 or 0.98 suits this trace is a question the simulator answers in seconds. When the scans stop, the target falls 10 ms every 2 seconds, so shedding 100 ms takes about 20 seconds of spread-out, inaudible accelerations. Keep this trace as a permanent regression test. The replay of this trace in the simulator is the regression test you keep forever.
Traffic that breaks naive designs
- Bursty AI voice senders. Text-to-speech services often produce audio faster than real time and send it in bursts. If the estimator resets its baseline on those early packets, later normal packets all look late and the target inflates. Cap early audio in the store and prefer a sender that paces packets at real time.
- Barge-in. When a user interrupts an AI agent, the buffered agent audio must be flushed at once, not played out. Expose a
flush()call and reset the estimator's baseline afterwards. - Discontinuous transmission. A silence gap is not loss; re-anchor on the first packet of the next talkspurt.
- Stream restarts. A new SSRC, or a sequence or timestamp jump, means a new stream. Reset the store and estimator instead of unwrapping across the discontinuity.
- Clock drift. Sender and receiver sample clocks differ by tens of parts per million, so the buffer slowly fills or drains. The controller's accelerate and decelerate actions absorb it if the margins allow; clock drift compensation covers resampling approaches.
- Recovered packets. Packets rebuilt by forward error correction arrive with the next packet, not on their own schedule; feed them to the store but not to the delay estimator. FEC and redundancy explains the timing.
Tuning in a browser
In WebRTC you do not replace the browser's jitter buffer, but you can influence it. The jitterBufferTarget attribute on RTCRtpReceiver takes a preferred hold time in milliseconds, from 0 to 4000; values outside that range throw a RangeError. It influences rather than sets the browser's target, which still adapts within its own limits.
// Browsers: hint the receiver's jitter buffer, per track. Value in ms, 0 to 4000.
// It influences the user agent's target; it does not set it.
for (const r of pc.getReceivers()) {
if (r.track.kind === 'audio') r.jitterBufferTarget = 120; // e.g. a listen-mostly stream
}
// Verify the effect from inbound-rtp stats: jitterBufferTargetDelay / jitterBufferEmittedCountUse it for listen-mostly streams and leave it alone for two-way conversation. Confirm the effect with the inbound-rtp statistics jitterBufferTargetDelay, jitterBufferDelay and jitterBufferEmittedCount, which are cumulative, the two delays in seconds, so divide delay deltas by the emitted-count delta between polls.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Delay grows all call long | Target rises but never falls, or drift uncorrected | Enforce the slow-decrease path; check accelerate margins |
| Robotic or warbling speech | Constant stretching from narrow margins | Widen margins; limit stretch rate |
| Choppy audio after a while | Sequence or timestamp wrap mishandled | Unwrap to extended integers; test with traces that wrap |
| Glitches at the start of every AI reply | Burst sender resetting the baseline | Cap early audio; pace the sender |
| Agent keeps talking after interruption | No flush on barge-in | Flush and reset on interruption |
Trade-offs
| Choice | Gains | Costs |
|---|---|---|
| Higher quantile | Fewer concealed frames | More delay |
| Fast increase, slow decrease | Absorbs recurring spikes | Delay lingers after the network recovers |
| Wide stretch margins | Fewer distortions | Buffer level wanders further from target |
What to do next
- Put your jitter buffer behind insert and pull functions that take time as an argument.
- Unwrap sequence numbers and timestamps, and add traces that wrap to your test set.
- Implement the histogram estimator and a controller that rises fast and falls slowly.
- Log every insert and decision in production, and build the replay simulator from the same code.
- Collect real and synthetic traces, and score designs on concealment, stretch and delay together.
- Handle bursty TTS senders, barge-in flushes, DTX gaps and stream restarts explicitly.
- In browsers, use jitterBufferTarget only for listen-mostly streams and verify with getStats deltas.