Opus is one codec with a very wide range: narrowband speech at a few kilobits per second, fullband stereo music at a few hundred, and frame sizes from 2.5 to 120 milliseconds. That range is exposed as a dozen or so encoder controls, and their defaults are tuned for a general case that may not be yours. Set them carelessly and you get muffled music, robotic game chat or files twice as large as needed.
This article goes through the parameters one at a time, says what each one changes inside the encoder, shows how it is spelled in the libopus C API, in ffmpeg and in opusenc, and then puts them together into five worked profiles. For the specifics of real-time voice transport (RTP, SDP, the receive-side FEC path and the latency budget), read the Opus voice guide; for the internal SILK and CELT layers, read the Opus architecture article.
The map: what is fixed and what is not
Three things are chosen when you create a libopus encoder: the input sample rate (8, 12, 16, 24 or 48 kHz), the channel count and the application; treat them as fixed once encoding starts. Everything else is a control request, opus_encoder_ctl, which you can change at any time; each packet's first byte (the table of contents) tells the decoder the mode, bandwidth and frame size, so the receiver never needs to be told. That is what lets a call adapt packet by packet.
Internally the encoder chooses between three modes: SILK (a linear-prediction speech coder) for low-bitrate speech, CELT (a transform coder) for music and low delay, and a hybrid that codes the low band with SILK and the high band with CELT. Most parameters below work by biasing that choice or by constraining the bandwidth and bits it has to work with.
Bitrate and rate control: VBR, constrained VBR, CBR
The bitrate request accepts 500 to 512,000 bits per second in libopus, plus automatic and maximum settings. How the encoder interprets it depends on the rate control mode, and that choice matters as much as the number.
| Mode | libopus | ffmpeg | opusenc | Behaviour |
|---|---|---|---|---|
| VBR | VBR(1), VBR_CONSTRAINT(0) | -vbr on | --vbr | Bitrate is a long-run average; hard passages get more bits. Best quality per byte. |
| Constrained VBR | VBR(1), VBR_CONSTRAINT(1) | -vbr constrained | --cvbr | Varies within a bounded window around the target. Good for streams with buffer limits. |
| Hard CBR | VBR(0) | -vbr off | --hard-cbr | Every frame the same size. Lowest quality; nothing leaks through packet sizes. |
Note the defaults: in the libopus API both VBR and the VBR constraint default to on, so an encoder you create and do not touch is in constrained VBR. ffmpeg and opusenc both default to unconstrained VBR. Encrypted calls deserve care here, because the sizes of unconstrained VBR speech packets have been shown to leak information about what was said; prefer constrained VBR or CBR when content confidentiality matters.
The other classic trap is units. ffmpeg's -b:a is in bits per second, while opusenc's --bitrate is in kilobits per second. -b:a 64000 or -b:a 64k is the correct ffmpeg spelling of 64 kbit/s; -b:a 64 asks for 64 bits per second. opusenc's defaults, for input at 44.1 kHz or above, are 64 kbit/s per mono stream and 96 kbit/s per coupled stereo pair.
Complexity
Complexity runs from 0 to 10 and trades encoder CPU for quality at a given bitrate; higher settings enable more exhaustive searches and analysis. ffmpeg's compression_level and opusenc's --comp both default to 10. The libopus API documentation does not state a default, so query it with OPUS_GET_COMPLEXITY on your build rather than assuming. Decoding cost does not depend on encoder complexity.
For offline encoding, always use 10: the CPU is spent once and the files are served many times. For real-time encoding on constrained devices (phones in calls, game clients sharing a core with the renderer, embedded microphones), measure encode time per frame at several settings on the actual hardware and pick the highest that leaves comfortable headroom.
Bandwidth, signal and application
Audio bandwidth is the range of frequencies coded: narrowband (4 kHz), mediumband (6 kHz), wideband (8 kHz), super-wideband (12 kHz) and fullband (20 kHz). The encoder picks one from the bitrate unless you intervene. OPUS_SET_MAX_BANDWIDTH sets a ceiling and still lets the encoder drop lower when bits are short; OPUS_SET_BANDWIDTH forces one value. ffmpeg exposes a cutoff in hertz (4000, 6000, 8000, 12000 or 20000). Capping bandwidth is useful when the source is band-limited (telephone recordings, cheap headsets), because it stops the encoder from spending bits on noise above the useful band.
The signal hint (OPUS_SIGNAL_VOICE, OPUS_SIGNAL_MUSIC or automatic) biases the mode decision at low bitrates; opusenc spells it --speech and --music. Set it when you know the content; automatic detection can flip modes on mixed audio such as speech over background music.
The application is chosen at creation. VOIP favours intelligibility, AUDIO favours fidelity, and RESTRICTED_LOWDELAY disables the SILK and hybrid modes to remove their extra delay. ffmpeg's application option uses voip, audio (the default) and lowdelay. Do not quote lookahead figures from memory: OPUS_GET_LOOKAHEAD returns the delay in samples for your build and configuration, and it is the number that belongs in a latency budget.
Frame duration
An Opus frame lasts 2.5 to 60 milliseconds. In the C API the duration is the number of samples per channel you pass to opus_encode (960 at 48 kHz for 20 ms, 240 for 5 ms), and libopus also accepts 80, 100 and 120 ms, carried as several frames in one packet. ffmpeg and opusenc both default to 20 ms and offer 2.5 to 60.
Shorter frames lower latency and raise overhead: each frame has fixed costs inside the bitstream, and on a network each packet adds IP, UDP and RTP headers. Longer frames are more efficient at low bitrates, and a lost packet costs more audio. For files, 20 ms remains the sensible default; longer frames gain little at music bitrates. The 2.5 and 5 ms sizes are CELT-only, so they cannot carry SILK's in-band FEC.
Stereo, multichannel and the downmix problem
With two channels, Opus can code a stereo image cheaply using intensity stereo, which at low bitrates represents some bands as one signal plus a direction. By default it may use phase inversion for this, which sounds fine on headphones but partly cancels when a player downmixes to mono, as smart speakers and phone loudspeakers often do. OPUS_SET_PHASE_INVERSION_DISABLED, ffmpeg's apply_phase_inv 0 and opusenc's --no-phase-inv trade a little stereo quality for a clean mono downmix. Disable it for podcasts, audiobooks and anything likely to play on a single speaker.
If the content is really mono, encode it as mono (--downmix-mono or one channel at creation) rather than as identical stereo channels; the encoder will waste fewer bits on a non-existent image. OPUS_SET_FORCE_CHANNELS forces mono or stereo coding regardless of input. Surround audio uses the multistream API and a channel mapping family in the Ogg header: family 0 for mono and stereo, family 1 for the Vorbis channel orders up to 7.1, and family 255 for undefined layouts. For spatial and ambisonic formats, see spatial audio.
Loss controls and other expert knobs
In-band FEC (OPUS_SET_INBAND_FEC) makes SILK and hybrid frames carry a low-bitrate copy of the previous frame, but only when expected loss (OPUS_SET_PACKET_LOSS_PERC) is above zero; leaving expected loss at its default of 0 is the usual reason FEC is configured but never transmitted. DTX (OPUS_SET_DTX) sends almost nothing during silence. libopus 1.5 added Deep Redundancy (DRED), set with OPUS_SET_DRED_DURATION in units of 10 ms frames, which carries a compact neural representation of recent past audio; it must be enabled in the library build and supported by the decoder, so check both ends before relying on it. How these compare with RED and retransmission is covered in FEC and redundancy for real-time audio.
OPUS_SET_PREDICTION_DISABLED makes frames almost independent of each other at a quality cost, useful when a stream must be decodable from any packet. OPUS_SET_LSB_DEPTH (8 to 24, default 24) tells the encoder how many bits of the input are real; it is a hint that helps the encoder identify silence and near-silence, so set 16 for 16-bit sources.
Files add two decoder-side parameters in the Ogg header. Pre-skip is the number of samples at the start that the decoder must discard, derived from the encoder's lookahead; players that ignore it add a short burst of silence or garbage and break gapless playback. Output gain is a header gain the decoder must apply, used for loudness normalisation without re-encoding.
Five worked profiles
| Profile | Application | Bitrate | Rate control | Other |
|---|---|---|---|---|
| Voice call | VOIP | 20-32 kbit/s mono | Constrained VBR | Wideband cap, FEC with real loss %, DTX, complexity to CPU budget |
| Game voice chat | VOIP | 24-40 kbit/s mono | Constrained VBR | Complexity around 5 if CPU bound, FEC, 20 ms frames |
| Music streaming file | AUDIO | 96-160 kbit/s stereo | Unconstrained VBR | Complexity 10, music signal, 20 ms |
| Podcast or audiobook file | AUDIO | 32-64 kbit/s | Unconstrained VBR | Mono if the source is mono, phase inversion off, speech signal |
| Live music over a network | RESTRICTED_LOWDELAY | 96-192 kbit/s | Constrained VBR or CBR | 5 or 10 ms frames, measured lookahead |
These ranges are starting points from common practice, not measured answers for your content. Here is the same set expressed in code and on the command line:
#include <opus.h>
typedef enum { CALL, GAME_CHAT, MUSIC_STREAM, SPOKEN_FILE, LIVE_MUSIC } profile_t;
/* All profiles run at 48 kHz. Values are starting points to measure, not answers. */
OpusEncoder *make_encoder(profile_t p, int channels, int expected_loss_pct) {
int err;
int app = (p == CALL || p == GAME_CHAT) ? OPUS_APPLICATION_VOIP
: (p == LIVE_MUSIC) ? OPUS_APPLICATION_RESTRICTED_LOWDELAY
: OPUS_APPLICATION_AUDIO;
OpusEncoder *e = opus_encoder_create(48000, channels, app, &err);
if (err != OPUS_OK) return NULL;
switch (p) {
case CALL:
opus_encoder_ctl(e, OPUS_SET_SIGNAL(OPUS_SIGNAL_VOICE));
opus_encoder_ctl(e, OPUS_SET_BITRATE(24000));
opus_encoder_ctl(e, OPUS_SET_MAX_BANDWIDTH(OPUS_BANDWIDTH_WIDEBAND));
opus_encoder_ctl(e, OPUS_SET_INBAND_FEC(1));
opus_encoder_ctl(e, OPUS_SET_PACKET_LOSS_PERC(expected_loss_pct));
opus_encoder_ctl(e, OPUS_SET_DTX(1));
break;
case GAME_CHAT:
opus_encoder_ctl(e, OPUS_SET_SIGNAL(OPUS_SIGNAL_VOICE));
opus_encoder_ctl(e, OPUS_SET_BITRATE(32000));
opus_encoder_ctl(e, OPUS_SET_COMPLEXITY(5)); /* share the CPU with the game */
opus_encoder_ctl(e, OPUS_SET_INBAND_FEC(1));
opus_encoder_ctl(e, OPUS_SET_PACKET_LOSS_PERC(expected_loss_pct));
break;
case MUSIC_STREAM:
opus_encoder_ctl(e, OPUS_SET_SIGNAL(OPUS_SIGNAL_MUSIC));
opus_encoder_ctl(e, OPUS_SET_BITRATE(channels == 2 ? 128000 : 64000));
opus_encoder_ctl(e, OPUS_SET_VBR_CONSTRAINT(0)); /* unconstrained VBR for files */
opus_encoder_ctl(e, OPUS_SET_COMPLEXITY(10));
break;
case SPOKEN_FILE:
opus_encoder_ctl(e, OPUS_SET_SIGNAL(OPUS_SIGNAL_VOICE));
opus_encoder_ctl(e, OPUS_SET_BITRATE(channels == 2 ? 48000 : 32000));
opus_encoder_ctl(e, OPUS_SET_VBR_CONSTRAINT(0));
opus_encoder_ctl(e, OPUS_SET_COMPLEXITY(10));
break;
case LIVE_MUSIC:
opus_encoder_ctl(e, OPUS_SET_BITRATE(channels == 2 ? 128000 : 96000));
/* 5 ms frames: pass 240 samples per channel to each opus_encode() call */
break;
}
opus_int32 v;
opus_encoder_ctl(e, OPUS_GET_LOOKAHEAD(&v)); /* log it: it is part of your latency */
return e;
}# ffmpeg: -b:a is in bit/s
ffmpeg -i talk.wav -c:a libopus -b:a 32000 -vbr on -application voip talk.opus
ffmpeg -i album.flac -c:a libopus -b:a 128000 -vbr on -compression_level 10 album.opus
ffmpeg -i mix.wav -c:a libopus -b:a 96000 -vbr constrained -frame_duration 10 mix.webm
# opusenc: --bitrate is in kbit/s
opusenc --bitrate 32 --speech talk.wav talk.opus
opusenc --bitrate 128 --vbr --comp 10 album.wav album.opus
opusenc --bitrate 64 --no-phase-inv lecture.wav lecture.opus
opusenc --bitrate 32 --downmix-mono mono_source.wav mono.opusTo choose a real number, sweep. Encode a few minutes of representative material at a ladder of bitrates and modes, record the actual bitrate, and score each output with an objective metric plus a blind listening panel for the bitrates near your candidate:
import subprocess, os, csv
def encode(src, kbps, mode, frame_ms, out):
subprocess.run(["ffmpeg", "-v", "error", "-y", "-i", src, "-c:a", "libopus",
"-b:a", str(kbps * 1000), "-vbr", mode,
"-frame_duration", str(frame_ms), out], check=True)
return os.path.getsize(out)
def sweep(src, seconds, rows):
for kbps in (16, 24, 32, 48, 64, 96, 128):
for mode in ("on", "constrained", "off"):
out = f"sweep_{kbps}_{mode}.opus"
size = encode(src, kbps, mode, 20, out)
# container overhead included; good enough to compare modes
rows.append(dict(kbps=kbps, mode=mode, actual_kbps=round(size * 8 / seconds / 1000, 1), file=out))
# score each file with your objective metric and a blind listening panel, then plot
with open("sweep.csv", "w", newline="") as f:
w = csv.DictWriter(f, fieldnames=rows[0].keys()); w.writeheader(); w.writerows(rows)The curve flattens: past some bitrate, listeners stop hearing a difference. Pick the knee for your content, then add margin for the hardest material you carry.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Music sounds dull or warbly | Bitrate too low for fullband, or VOIP application on music | AUDIO application, music signal hint, higher bitrate |
| ffmpeg bitrate error or wrong rate | Bitrate given in kbit/s to an option expecting bit/s | Write 64000 or 64k |
| Speech fine on headphones, hollow on one speaker | Phase inversion cancelling in mono downmix | Disable phase inversion or encode mono |
| FEC enabled but no recovery | Expected loss left at 0, or CELT-only frame size | Set loss % from measurements; use 10 ms or longer frames |
| Audio glitches on low-end phones | Complexity too high for the real-time deadline | Measure encode time per frame and lower complexity |
| Click or gap at the start of files | Player ignores pre-skip | Use a conforming decoder or apply pre-skip yourself |
If artefacts appear only under loss, look at concealment rather than the encoder; packet loss concealment explains what the decoder does when a packet is missing.
What to do next
- Write down each Opus use in your product, its delivery path (RTP or file) and its latency limit.
- For each, set application, signal hint and channel count deliberately, and query and log lookahead and complexity.
- Pick rate control by path: unconstrained VBR for files, constrained VBR for live streams, CBR where packet sizes must not leak.
- Check bitrate units in every ffmpeg and opusenc command and test one output by ear.
- Turn off phase inversion for content likely to play on a single speaker, and encode mono sources as mono.
- Run the bitrate sweep on representative material and choose the knee, not a number from a table.