Opus is a mandatory audio codec of WebRTC and the default in most modern voice systems, and it is easy to get working: create an encoder, feed it 20 milliseconds of samples, send the bytes. It is just as easy to get subtly wrong. Forward error correction that is switched on but never produced, silence suppression that the receiver mistakes for packet loss, a 60 millisecond frame that quietly adds 40 milliseconds of delay: none of these raise errors, they just make calls worse.
This article is about operating Opus in a real-time voice path, not about its internals. The Opus codec architecture article explains how the SILK linear-prediction layer, the CELT transform layer and the hybrid mode fit together. Here we start from that and work through the decisions an engineer actually makes: frame size, bitrate and audio bandwidth, the libopus encoder and decoder calls, how the codec is negotiated in SDP and carried in RTP, how the receiver chooses between a real packet, an in-band FEC copy and concealment, and what all of this does to latency. A worked configuration and a checklist close it out.
What Opus gives a voice application
Opus is standardised in RFC 6716. A single bitstream format covers narrowband telephone speech at a few kilobits per second up to fullband stereo music, and the encoder can change mode, bandwidth and bitrate on every packet without signalling anything to the receiver. For voice, three properties matter most.
First, frames are short (2.5 to 60 milliseconds) and the algorithmic delay is small. Second, the speech path (SILK) can carry a coarser copy of the previous frame inside the current packet, called LBRR; this is Opus in-band forward error correction. Third, discontinuous transmission makes silence almost free on the wire.
The decoder outputs at a rate you choose (8, 12, 16, 24 or 48 kHz) whatever the encoder used internally, while the RTP clock is always 48 kHz, a convenient decoupling that causes several bugs discussed below.
The first decision: frame size and what it costs
The frame size sets the floor on latency and the ceiling on efficiency. Each packet carries a fixed overhead of IP, UDP and RTP headers, around 40 bytes on IPv4 before SRTP authentication tags. At 20 milliseconds per packet that is 50 packets per second and about 16 kbit/s of header overhead, which is comparable to the voice payload itself. At 10 milliseconds the overhead doubles; at 60 milliseconds it falls to a third, but every packet now represents 60 milliseconds of buffering at the sender and a single loss removes 60 milliseconds of speech.
| Frame | Packets/s | Header overhead (approx. 40 B) | Typical use |
|---|---|---|---|
| 10 ms | 100 | 32 kbit/s | Low-latency interactive audio, music jamming |
| 20 ms | 50 | 16 kbit/s | The default for voice and WebRTC |
| 40 ms | 25 | 8 kbit/s | Constrained uplinks, satellite |
| 60 ms | 16.7 | 5.3 kbit/s | Very low bitrate, latency-tolerant links |
Twenty milliseconds is the default for good reason; move off it only for a measured reason. In-band FEC is a SILK feature, so CELT-only frame sizes of 2.5 and 5 milliseconds carry no LBRR at all.
Bitrate, bandwidth and the application mode
Opus separates bitrate from audio bandwidth. Bandwidth is the range of frequencies coded: narrowband (4 kHz), mediumband (6 kHz), wideband (8 kHz), super-wideband (12 kHz) and fullband (20 kHz). The encoder picks bandwidth from the bitrate unless you cap it, and for speech the useful operating points are roughly: narrowband at 8 to 12 kbit/s, wideband at 16 to 20 kbit/s, and fullband speech at 28 to 40 kbit/s. Above that, extra bits for mono speech buy very little; spend them on FEC instead.
The application mode is set at creation. OPUS_APPLICATION_VOIP biases the encoder towards speech intelligibility and is what a call should use. OPUS_APPLICATION_AUDIO favours fidelity for music and mixed content. OPUS_APPLICATION_RESTRICTED_LOWDELAY disables the speech-optimised modes to shave algorithmic delay, which is useful for musicians playing together and almost never for a phone call. Instead of quoting a delay figure for each mode, query the encoder: OPUS_GET_LOOKAHEAD returns the lookahead in samples for your exact build and configuration, and you should put that number into your latency budget.
Also set OPUS_SET_SIGNAL(OPUS_SIGNAL_VOICE) when the input is known to be speech, so the encoder does not drift towards music-oriented modes.
The encoder in code
The libopus encoder is configured through opus_encoder_ctl requests, which can be changed at any time; that is how bandwidth estimation feeds back into the codec. This is a starting point for a mono wideband call that expects some loss.
#include <opus.h>
/* One mono voice encoder at 48 kHz, 20 ms frames (960 samples). */
OpusEncoder *make_voice_encoder(int expected_loss_pct, opus_int32 bitrate_bps) {
int err;
OpusEncoder *enc = opus_encoder_create(48000, 1, OPUS_APPLICATION_VOIP, &err);
if (err != OPUS_OK) return NULL;
opus_encoder_ctl(enc, OPUS_SET_SIGNAL(OPUS_SIGNAL_VOICE));
opus_encoder_ctl(enc, OPUS_SET_BITRATE(bitrate_bps)); /* e.g. 24000 */
opus_encoder_ctl(enc, OPUS_SET_MAX_BANDWIDTH(OPUS_BANDWIDTH_WIDEBAND));
opus_encoder_ctl(enc, OPUS_SET_VBR(1));
opus_encoder_ctl(enc, OPUS_SET_VBR_CONSTRAINT(1)); /* bounded bursts */
opus_encoder_ctl(enc, OPUS_SET_COMPLEXITY(8));
opus_encoder_ctl(enc, OPUS_SET_INBAND_FEC(1));
opus_encoder_ctl(enc, OPUS_SET_PACKET_LOSS_PERC(expected_loss_pct));
opus_encoder_ctl(enc, OPUS_SET_DTX(1));
return enc;
}
/* Returns bytes to send, or 0 when DTX says this frame need not be sent. */
int encode_frame(OpusEncoder *enc, const opus_int16 *pcm960, unsigned char *out, int cap) {
opus_int32 n = opus_encode(enc, pcm960, 960, out, cap);
if (n < 0) return n; /* OPUS_BAD_ARG, OPUS_BUFFER_TOO_SMALL, ... */
if (n <= 2) return 0; /* DTX: libopus documents that such packets need not be sent */
return n;
}Three details matter. In-band FEC is only produced when OPUS_SET_INBAND_FEC is on and OPUS_SET_PACKET_LOSS_PERC is above zero and the encoder is in a SILK or hybrid mode with enough bitrate to afford it; leaving expected loss at its default of zero is the most common reason FEC is enabled in configuration and absent on the wire. Second, constrained VBR keeps bursts bounded, which helps jitter buffers, and unconstrained VBR packet sizes have been shown to leak information about spoken content, so prefer CBR or constrained VBR for sensitive calls. Third, with DTX on, silent frames need not be sent at all. Alternatives to in-band FEC, such as RED and retransmission, are compared in FEC and redundancy for real-time audio.
Negotiation and transport: RTP and SDP
RFC 7587 defines the RTP payload format. Opus uses a dynamic payload type, one Opus packet per RTP packet, and an RTP timestamp clock of 48 kHz no matter which sample rate either side actually uses. The SDP rtpmap is therefore always opus/48000/2, even for mono at 16 kHz; the two channels in the rtpmap are a formality, and real channel preferences go in the fmtp parameters.
m=audio 9 UDP/TLS/RTP/SAVPF 111
a=rtpmap:111 opus/48000/2
a=fmtp:111 minptime=10;useinbandfec=1;usedtx=1;maxaveragebitrate=32000;stereo=0;sprop-stereo=0
a=ptime:20| Parameter | Meaning | Default |
|---|---|---|
useinbandfec | Receiver can use in-band FEC, so the sender should produce it | 0 |
usedtx | Receiver prefers the sender to use DTX | 0 |
maxaveragebitrate | Cap on the average bitrate the receiver wants, in bit/s | codec-dependent |
stereo / sprop-stereo | Receiver prefers stereo / sender is likely to send stereo | 0 |
maxplaybackrate | Highest sample rate the receiver can usefully play | 48000 |
cbr | Receiver prefers constant bitrate | 0 |
ptime, maxptime, minptime | Preferred, maximum and minimum packet durations | none |
These parameters describe what the receiver wants. Each side reads the other side's fmtp and configures its own encoder to match: if the remote offer says useinbandfec=1, turn FEC on in your encoder. Implementations that copy their own fmtp into their encoder get it backwards and usually nobody notices until a loss test.
The receiver: real packet, FEC copy, or concealment
When the jitter buffer releases a frame for playout there are three possibilities, and they must be tried in order.
/* Called by the jitter buffer when frame `seq` is due for playout.
pkt(seq) returns NULL if that packet never arrived. */
int decode_due_frame(OpusDecoder *dec, JitterBuffer *jb, uint16_t seq, opus_int16 *out) {
const int frame = 960; /* 20 ms at 48 kHz */
Packet *cur = jb_get(jb, seq);
if (cur) /* 1. the real packet */
return opus_decode(dec, cur->data, cur->len, out, frame, 0);
Packet *next = jb_get(jb, (uint16_t)(seq + 1));
if (next) /* 2. LBRR copy of seq lives inside seq+1 */
return opus_decode(dec, next->data, next->len, out, frame, 1);
/* next is decoded again normally at seq+1 */
return opus_decode(dec, NULL, 0, out, frame, 0); /* 3. concealment (PLC) */
}The consequence is easy to miss: the LBRR copy of frame N travels in packet N+1, so to use it the jitter buffer must still be waiting when N+1 arrives. That means at least one frame of extra buffering, 20 milliseconds at the default frame size. A jitter buffer that plays out as early as possible will never find the FEC copy in time and will fall straight through to concealment, and the FEC bits are wasted. When calling the decoder with decode_fec=1, the frame size argument must be exactly the duration of the missing audio.
Concealment is a call to opus_decode with a NULL packet. Classic Opus concealment extrapolates the last pitch period and fades out over consecutive losses. Opus 1.5 (2024) added machine-learning features: Deep PLC (compiled in with --enable-deep-plc and active only at decoder complexity 5 or more), a deep redundancy scheme called DRED, and the LACE and NoLACE low-bitrate speech enhancers. Check what your build and both endpoints support before relying on them. General concealment techniques are covered in packet loss concealment and buffering strategy in jitter buffer design.
DTX and comfort noise
In a two-party call each side is silent for roughly half the time, so DTX can cut bandwidth substantially. With DTX on, the Opus encoder sends a frame only occasionally during silence (libopus sends one about every 400 milliseconds to refresh the noise estimate) and the decoder generates comfort noise in between, so the listener hears a natural background instead of dead air.
The receiver must tell a DTX gap from a loss. A jitter buffer that treats every gap as loss runs concealment through silence and reports inflated loss, so the sender wastes bits on FEC. RTP sequence numbers are only incremented for packets actually sent, so a DTX gap shows up as a timestamp jump with consecutive sequence numbers; use that to tell the two apart. The DTX and comfort noise article covers the VAD side and the perceptual tuning.
Worked example: a mouth-to-ear budget
Suppose a mobile voice app targets the commonly cited 150 millisecond one-way limit for comfortable conversation (ITU-T G.114). Here is a budget for 20 millisecond Opus frames, with each line labelled by where its number comes from.
| Stage | Budget | Source of the number |
|---|---|---|
| Capture buffer | 10 ms | Audio device callback size |
| Echo cancellation and noise suppression | about 10 ms | Processing block of the DSP chain |
| Frame accumulation | 20 ms | Frame size: the encoder needs a full frame |
| Encoder lookahead | measure | OPUS_GET_LOOKAHEAD, divide samples by 48 |
| Network one-way | 40 to 60 ms | Measured RTT / 2 for your users |
| Jitter buffer | 40 ms | Adaptive target; includes one frame for FEC |
| Decode and playout buffer | about 15 ms | Output device buffer |
The fixed rows add up to about 135 to 155 milliseconds before lookahead, so the budget is already full. The network is out of your hands and the jitter buffer is adaptive; frame size is the cheapest dial you control. Moving to 40 millisecond frames would add 20 milliseconds and push most calls over the target, while 10 millisecond frames save 10 milliseconds but double the packet rate, which on a congested uplink can cost more in loss than it saves in delay.
With the budget fixed, the encoder settings follow: wideband cap, 24 kbit/s (wideband speech plus headroom for FEC), in-band FEC with expected loss fed from RTCP receiver reports, DTX and constrained VBR. If bandwidth estimation pushes the bitrate below roughly 12 kbit/s, cap the bandwidth to narrowband yourself.
Failure modes
- FEC on, nothing produced. Expected loss left at zero, or the encoder is in CELT-only mode at high bitrate. Verify by decoding the mode from each packet's TOC byte (RFC 6716) and comparing packet sizes with FEC toggled.
- FEC produced, never used. The jitter buffer plays out too eagerly to receive N+1 before N is due. Instrument how often each of the three decode paths runs.
- DTX read as loss. Concealment runs through silence, loss statistics inflate, and the sender wastes bitrate on redundancy.
- 48 kHz clock confusion. RTP timestamps advance by 960 per 20 ms frame even when the decoder runs at 16 kHz. Code that computes timestamps from its own sample rate breaks lip sync and jitter calculations.
- Reversed fmtp semantics. A sender configures itself from its own SDP instead of the peer's.
- Shared decoder across streams. Opus decoders carry state; one decoder per incoming SSRC, and reset it on a stream change.
- Double resampling. Resample in one place only.
What to do next
- Fix the frame size at 20 ms unless you have measured a reason to change it, and write down your mouth-to-ear budget including the value from OPUS_GET_LOOKAHEAD.
- Create encoders with OPUS_APPLICATION_VOIP and OPUS_SIGNAL_VOICE, cap the bandwidth to match your bitrate, and use constrained VBR.
- Turn on in-band FEC and feed OPUS_SET_PACKET_LOSS_PERC from RTCP loss reports; confirm on the wire that packets grow when loss rises.
- Implement the three-step decode (real packet, FEC from N+1, PLC) and keep at least one frame of jitter buffer headroom; log which path each frame took.
- Configure your encoder from the peer's fmtp parameters, and keep RTP timestamps on the 48 kHz clock.
- Distinguish DTX gaps from loss using sequence numbers versus timestamps before reporting loss.
- Run a loss and jitter test matrix (random and bursty loss, 0 to 20 percent) and compare listening quality before you ship any change.