Packets in a voice call leave the sender at a perfectly steady rate, one every 20 ms for example, and arrive at the receiver with gaps that wobble: 18 ms, 45 ms, 2 ms, 21 ms. The sound card, meanwhile, wants exactly 10 ms of audio every 10 ms, forever. The jitter buffer sits between the two. It holds packets long enough that the late ones usually arrive before they are needed, and it decides what to play when they do not.

Every millisecond of buffering is a millisecond of mouth-to-ear delay, and every packet that arrives after its turn is a gap the listener hears. This article builds a buffer that manages that trade: the timing model, delay estimation, the per-frame playout decision, and operations.

Advertisement

The timing model

Each RTP packet carries a sequence number, incremented by one per packet, and a timestamp in sample-clock units that says where its audio belongs on the sender's timeline. Opus over RTP uses a 48 kHz clock regardless of the audio bandwidth, so a 20 ms packet advances the timestamp by 960. The receiver adds a third value, its local arrival time.

Sequence numbers reveal loss and reordering; timestamps reveal where audio belongs, including silence gaps; arrival time against timestamp reveals delay variation, which is all the buffer needs. Both counters wrap, 16-bit sequence numbers in about 22 minutes at 50 packets per second and 32-bit timestamps much later, so unwrap them into 64-bit values on receipt.

The architecture

Receiver side: packets arrive on the network's schedule, audio leaves on the device'sRTP receiveseq, timestamp, arrivalPacket bufferkeyed by seq, unwrappedDelay estimatorrelative delay quantileTarget delayfloor, cap, smoothingarrival timesPlayout decisionevery 10 msleveltargetNormal decodeFEC decodeConceal (expand)Accelerate / slowMergeOutput ring to audio devicethe device callback pulls fixed frames on its own clockStatsdelay, target, concealed and stretched samplesThe buffer converts variable network delay into a fixed playout delay. Every decision trades latency against gaps.Packet buffer and decoder state belong to the receive and audio threads; keep the hand-off lock-free.
An adaptive jitter buffer. The network side inserts packets and updates the delay estimate; the device side pulls one frame per callback and the playout decision chooses how to produce it.

Packets are inserted whenever the network delivers them; audio is pulled whenever the device callback fires. The buffer level, milliseconds of contiguous audio queued from the next sequence number, connects the two. The delay estimator proposes a target level, and the playout decision compares level with target every frame and picks an operation. For the device side of this picture, periods, callbacks and ring buffers, see Audio Buffer Management for Low Latency.

Advertisement

Fixed or adaptive

A fixed buffer holds every packet for a constant time, say 60 ms. It is simple and predictable but wrong both ways: on a clean network it wastes most of those 60 ms, and on congested Wi-Fi every packet later than 60 ms is discarded, so loss climbs exactly when the network is worst.

An adaptive buffer moves its target with measured delay variation, staying shallow on good networks and growing on bad ones, at the cost of changing depth without audible artifacts and not chasing every spike. Interactive voice almost always uses adaptive buffers.

Measuring jitter, and why the RFC number is not the answer

RFC 3550 defines an interarrival jitter estimate that every RTP receiver reports back to the sender. For consecutive packets, D is the change in transit time, arrival minus timestamp, and the estimate is a running average of |D| with gain 1/16.

class Rfc3550Jitter:
    """Interarrival jitter, RFC 3550 section 6.4.1, in timestamp units."""
    def __init__(self, clock_rate=48000):
        self.rate = clock_rate
        self.prev_transit = None
        self.j = 0.0

    def on_packet(self, rtp_ts, arrival_s):
        transit = arrival_s * self.rate - rtp_ts      # unknown constant offset cancels in D
        if self.prev_transit is not None:
            d = abs(transit - self.prev_transit)
            self.j += (d - self.j) / 16.0
        self.prev_transit = transit
        return 1000.0 * self.j / self.rate            # milliseconds

This is a smoothed mean of packet-to-packet variation. It is good for reports and trends but poor for sizing, for two reasons. First, buffer depth must cover the tail of the delay distribution, not its mean; a link with mostly smooth delivery and occasional 80 ms stalls has a small RFC jitter and needs a large buffer. Second, it measures differences between neighbours, so a slow drift in delay barely registers even though the buffer must absorb all of it.

A delay estimator that sizes the buffer

What the buffer needs is the distribution of each packet's relative delay: how much later it arrived than the fastest packet seen recently, after accounting for where it belongs on the timeline. The fastest recent packet approximates the minimum path delay, so relative delay is the extra wait that packet experienced. A buffer that delays playout by the 95th percentile of relative delay will, in steady conditions, have about 95 percent of packets on time.

from collections import deque

class DelayEstimator:
    """Relative delay of each packet versus the fastest recent one, summarised as a quantile."""
    def __init__(self, bucket_ms=10, max_ms=1000, forget=0.998, quantile=0.95, window_s=10.0):
        self.bucket, self.forget, self.q = bucket_ms, forget, quantile
        self.hist = [0.0] * (max_ms // bucket_ms + 1)
        self.recent = deque()                     # (arrival_ms, transit_ms)
        self.window_ms = window_s * 1000

    def on_packet(self, media_ms, arrival_ms):
        transit = arrival_ms - media_ms           # includes the unknown clock offset
        self.recent.append((arrival_ms, transit))
        while self.recent[0][0] < arrival_ms - self.window_ms:
            self.recent.popleft()
        rel = transit - min(t for _, t in self.recent)   # use a monotonic deque in production
        i = min(int(rel // self.bucket), len(self.hist) - 1)
        self.hist = [h * self.forget for h in self.hist]
        self.hist[i] += 1.0 - self.forget
        return rel

    def target_ms(self, frame_ms=20, floor_ms=20, cap_ms=400):
        total, acc = sum(self.hist), 0.0
        for i, h in enumerate(self.hist):
            acc += h
            if acc >= self.q * total:
                return max(floor_ms, min(cap_ms, (i + 1) * self.bucket + frame_ms))
        return cap_ms

The histogram forgets old arrivals geometrically: with a forgetting factor of 0.998 per packet and 50 packets per second, the effective memory is about 500 packets, roughly 10 seconds. The quantile, floor, cap and memory are tuning parameters, not constants from a standard. A higher quantile means fewer late packets and more delay; a shorter memory adapts faster and fluctuates more. Many designs also make the target rise quickly and fall slowly, because a late packet is audible immediately while extra delay is paid gradually.

The playout decision

Every time the device asks for a frame, the buffer chooses one operation. The essential logic fits in a few lines.

NORMAL, FEC, CONCEAL, MERGE, ACCELERATE, SLOW = range(6)

def decide(buf, seq, target_ms, last_op, margin_ms=20):
    """Called once per 10 ms output frame. Returns (operation, advance_seq)."""
    level = buf.level_ms(seq)                 # contiguous media buffered from seq onward
    if buf.has(seq):
        if last_op == CONCEAL:
            return MERGE, True                # crossfade synthetic audio into real audio
        if level > target_ms + margin_ms:
            return ACCELERATE, True           # remove a pitch period or a silent stretch
        if level < target_ms - margin_ms:
            return SLOW, True                 # insert a pitch period while audio remains
        return NORMAL, True
    if buf.has(seq + 1) and buf.has_fec(seq + 1):
        return FEC, True                      # e.g. Opus in-band FEC carried in the next packet
    if buf.has_any_after(seq):
        return CONCEAL, True                  # later packets arrived: seq is lost, move on
    return CONCEAL, False                     # nothing at all: likely a delay spike, hold seq
  • Normal. The next packet is here and the level is near target: decode and play.
  • Accelerate and slow. The level is far from target, so play the audio slightly faster or slower with time-scale modification, covered below. This is how the buffer changes depth without dropping or inserting audible chunks.
  • FEC. The packet is missing but the next one carries a redundant low-bitrate copy of it, as Opus in-band FEC can. Decoding that copy is better than synthesising. The mechanics are in The Opus Codec for Real-Time Voice.
  • Conceal. Nothing usable is available, so the decoder's packet loss concealment extrapolates from recent audio. See Packet loss concealment.
  • Merge. Real audio has arrived after concealment; crossfade so the seam is not heard as a click.

Late versus lost

When a packet is missing at its playout time, the buffer cannot know whether it is lost or just late, and it has to make the device's deadline anyway, so it conceals. The real decision is what to do with the sequence pointer.

If later packets are already buffered, the missing one is probably lost or badly reordered: advance past it, and if it arrives afterwards, discard it. If nothing at all is buffered, the network has probably stalled, as in a Wi-Fi roam or a cellular handover: hold the pointer and keep concealing. When the packets arrive, they are played in order and the playout point has effectively slipped later, so the buffer has grown by the length of the stall. Accelerate then drains that extra delay over the following seconds.

Count discarded late packets separately from network loss. Both are gaps to the listener, but a high late-discard count says the target is too low, while network loss says nothing about buffer sizing. And always feed late packets to the delay estimator: they are precisely the samples that should raise the target.

Changing depth without artifacts

Time-scale modification changes duration without changing pitch. For speech the classic approach is overlap-add at pitch-synchronous offsets, as in WSOLA: find a segment that matches the waveform one pitch period later and crossfade across it, removing or repeating one period. A pitch period is between about 2.5 and 20 ms for voices from 400 Hz down to 50 Hz, so each operation moves the level by a few milliseconds and is almost inaudible when spread out.

Silence is cheaper still: during pauses, or discontinuous transmission gaps, the buffer can drop or insert comfort noise freely, so prefer pauses when the correction can wait.

Clock drift

The sender's sample clock and the receiver's device clock are different crystals. A mismatch of 100 parts per million, plausible for consumer hardware, is 0.1 ms per second, or 6 ms per minute. Over a one-hour call that is 360 ms. If the receiver plays slower than the sender produces, the buffer fills forever; if faster, it drains and underflows periodically. An adaptive buffer absorbs drift automatically, because the level drifts away from the target and accelerate or slow corrects it. A fixed buffer needs an explicit resampler. A steady one-directional bias in stretch operations is drift, not jitter.

A worked example

Ten 20 ms packets arrive with relative delays of 0, 5, 3, 40, 2, 0, 60, 55, 50 and 10 ms; packets 7 to 9 hit a short stall. A fixed 40 ms buffer plays packet 4 on time, its delay exactly equals the budget, and discards 7, 8 and 9: 30 percent loss in that window and three concealments in a row, which is clearly audible.

Suppose an adaptive buffer's history puts its 95th percentile of relative delay in the 20-30 ms bucket, since delays like packet 4's are under 5 percent of arrivals, so target_ms() returns 30 + 20 = 50 ms. It plays packets 1 to 6 on time and still misses packet 7. But with nothing later buffered, it holds the pointer, conceals through the stall, then plays 7, 8 and 9 late rather than discarding them. The listener hears one stretch of concealment followed by complete speech, and the delay estimator, fed these large relative delays, raises the target. Over the next seconds accelerate drains the slip if the stall does not recur.

Monitoring it in production

In browsers and other WebRTC stacks, the standard statistics expose what the buffer is doing. The fields below are cumulative counters on the audio inbound-rtp stats object, so compute deltas between polls.

let prev = null;
async function pollJitterStats(pc) {
  const report = await pc.getStats();
  for (const r of report.values()) {
    if (r.type !== 'inbound-rtp' || r.kind !== 'audio') continue;
    const emitted = prev ? r.jitterBufferEmittedCount - prev.jitterBufferEmittedCount : 0;
    const samples = prev ? r.totalSamplesReceived - prev.totalSamplesReceived : 0;
    if (emitted > 0 && samples > 0) {
      console.log({
        avgDelayMs: 1000 * (r.jitterBufferDelay - prev.jitterBufferDelay) / emitted,
        avgTargetMs: 1000 * (r.jitterBufferTargetDelay - prev.jitterBufferTargetDelay) / emitted,
        concealedPct: 100 * (r.concealedSamples - prev.concealedSamples) / samples,
        stretchedPct: 100 * ((r.insertedSamplesForDeceleration - prev.insertedSamplesForDeceleration)
                    + (r.removedSamplesForAcceleration - prev.removedSamplesForAcceleration)) / samples,
      });
    }
    prev = r;
  }
}

Alert on concealment ratio and on average delay together: a buffer that keeps concealment low by growing to hundreds of milliseconds has failed just as surely as one that conceals constantly. For a lighter overview of the same components, see Adaptive jitter buffer architecture.

Failure modes

SymptomLikely causeFix
Delay grows all call longClock drift with no correction, or target that rises but never fallsLet accelerate drain to target; decay the histogram
Burst of concealment after every handoverPointer advanced through a stallHold the pointer when nothing later is buffered
Robotic, warbling speechStretching too aggressively in voiced speechLimit stretch rate; prefer pauses
Long silences treated as lossDTX gaps not recognisedUse timestamps, not sequence gaps alone, to detect silence
Garbage after a stream restartSender changed SSRC or jumped timestampsReset buffer and estimator on SSRC change or large jumps
Everything breaks after about 22 minutes16-bit sequence wraparoundUnwrap counters to 64 bits on receipt

Trade-offs

The quantile trades delay against late loss, the forgetting window responsiveness against stability, stretching smooth depth changes against subtle artifacts, and FEC bitrate against concealments. For conversation, mouth-to-ear delay above about 150 ms starts to impair turn-taking, so the jitter buffer must share that budget with capture, encoding, network and playout.

What to do next

  1. Unwrap sequence numbers and timestamps to 64 bits and log arrival time, timestamp and sequence for every packet on a test call.
  2. Replay those logs offline through the delay estimator and plot relative delay, target and late discards over time.
  3. Implement the playout decision with an explicit hold-or-advance rule for missing packets.
  4. Add time-scale modification for accelerate and slow, and verify by listening at different stretch rates.
  5. Enable Opus in-band FEC on lossy links and confirm the FEC path is used before concealment.
  6. Export average delay, target, concealment and stretch ratios, and alert on delay and concealment together.
Key takeaway: A jitter buffer turns variable network delay into a fixed playout delay, and its whole design is the trade between latency and gaps. Size it from the tail of each packet's relative delay rather than the RFC 3550 mean, let the target rise quickly and fall slowly, decide every frame between normal decode, FEC, concealment, merge and gentle time stretching, hold the pointer through stalls instead of discarding the packets that end them, absorb clock drift through the same mechanism, and monitor delay and concealment together.