Opus is one codec with a very wide range: narrowband speech at a few kilobits per second, fullband stereo music at a few hundred, and frame sizes from 2.5 to 120 milliseconds. That range is exposed as a dozen or so encoder controls, and their defaults are tuned for a general case that may not be yours. Set them carelessly and you get muffled music, robotic game chat or files twice as large as needed.

This article goes through the parameters one at a time, says what each one changes inside the encoder, shows how it is spelled in the libopus C API, in ffmpeg and in opusenc, and then puts them together into five worked profiles. For the specifics of real-time voice transport (RTP, SDP, the receive-side FEC path and the latency budget), read the Opus voice guide; for the internal SILK and CELT layers, read the Opus architecture article.

Advertisement

The map: what is fixed and what is not

Set at creationrate, channels, applicationPer-packet controlsbitrate, VBR, complexitySignal shapingbandwidth, signal, frameLoss controlsFEC, loss %, DTX, DREDEncoder decisionsSILK / hybrid / CELT, bandwidthRTPcalls, game chatOgg / WebMfiles, streamingpacketsDecoderoutput rate, gain, PLC/FECTreat the creation-time settings as fixed; everything else can change on any packet without telling the decoder.The container carries pre-skip and output gain, which the decoder side must apply.
Opus separates creation-time settings from controls that may change on every packet. The decoder learns everything it needs from each packet's table-of-contents byte, plus pre-skip and output gain from the container.

Three things are chosen when you create a libopus encoder: the input sample rate (8, 12, 16, 24 or 48 kHz), the channel count and the application; treat them as fixed once encoding starts. Everything else is a control request, opus_encoder_ctl, which you can change at any time; each packet's first byte (the table of contents) tells the decoder the mode, bandwidth and frame size, so the receiver never needs to be told. That is what lets a call adapt packet by packet.

Internally the encoder chooses between three modes: SILK (a linear-prediction speech coder) for low-bitrate speech, CELT (a transform coder) for music and low delay, and a hybrid that codes the low band with SILK and the high band with CELT. Most parameters below work by biasing that choice or by constraining the bandwidth and bits it has to work with.

Bitrate and rate control: VBR, constrained VBR, CBR

The bitrate request accepts 500 to 512,000 bits per second in libopus, plus automatic and maximum settings. How the encoder interprets it depends on the rate control mode, and that choice matters as much as the number.

ModelibopusffmpegopusencBehaviour
VBRVBR(1), VBR_CONSTRAINT(0)-vbr on--vbrBitrate is a long-run average; hard passages get more bits. Best quality per byte.
Constrained VBRVBR(1), VBR_CONSTRAINT(1)-vbr constrained--cvbrVaries within a bounded window around the target. Good for streams with buffer limits.
Hard CBRVBR(0)-vbr off--hard-cbrEvery frame the same size. Lowest quality; nothing leaks through packet sizes.

Note the defaults: in the libopus API both VBR and the VBR constraint default to on, so an encoder you create and do not touch is in constrained VBR. ffmpeg and opusenc both default to unconstrained VBR. Encrypted calls deserve care here, because the sizes of unconstrained VBR speech packets have been shown to leak information about what was said; prefer constrained VBR or CBR when content confidentiality matters.

The other classic trap is units. ffmpeg's -b:a is in bits per second, while opusenc's --bitrate is in kilobits per second. -b:a 64000 or -b:a 64k is the correct ffmpeg spelling of 64 kbit/s; -b:a 64 asks for 64 bits per second. opusenc's defaults, for input at 44.1 kHz or above, are 64 kbit/s per mono stream and 96 kbit/s per coupled stereo pair.

Advertisement

Complexity

Complexity runs from 0 to 10 and trades encoder CPU for quality at a given bitrate; higher settings enable more exhaustive searches and analysis. ffmpeg's compression_level and opusenc's --comp both default to 10. The libopus API documentation does not state a default, so query it with OPUS_GET_COMPLEXITY on your build rather than assuming. Decoding cost does not depend on encoder complexity.

For offline encoding, always use 10: the CPU is spent once and the files are served many times. For real-time encoding on constrained devices (phones in calls, game clients sharing a core with the renderer, embedded microphones), measure encode time per frame at several settings on the actual hardware and pick the highest that leaves comfortable headroom.

Bandwidth, signal and application

Audio bandwidth is the range of frequencies coded: narrowband (4 kHz), mediumband (6 kHz), wideband (8 kHz), super-wideband (12 kHz) and fullband (20 kHz). The encoder picks one from the bitrate unless you intervene. OPUS_SET_MAX_BANDWIDTH sets a ceiling and still lets the encoder drop lower when bits are short; OPUS_SET_BANDWIDTH forces one value. ffmpeg exposes a cutoff in hertz (4000, 6000, 8000, 12000 or 20000). Capping bandwidth is useful when the source is band-limited (telephone recordings, cheap headsets), because it stops the encoder from spending bits on noise above the useful band.

The signal hint (OPUS_SIGNAL_VOICE, OPUS_SIGNAL_MUSIC or automatic) biases the mode decision at low bitrates; opusenc spells it --speech and --music. Set it when you know the content; automatic detection can flip modes on mixed audio such as speech over background music.

The application is chosen at creation. VOIP favours intelligibility, AUDIO favours fidelity, and RESTRICTED_LOWDELAY disables the SILK and hybrid modes to remove their extra delay. ffmpeg's application option uses voip, audio (the default) and lowdelay. Do not quote lookahead figures from memory: OPUS_GET_LOOKAHEAD returns the delay in samples for your build and configuration, and it is the number that belongs in a latency budget.

Frame duration

An Opus frame lasts 2.5 to 60 milliseconds. In the C API the duration is the number of samples per channel you pass to opus_encode (960 at 48 kHz for 20 ms, 240 for 5 ms), and libopus also accepts 80, 100 and 120 ms, carried as several frames in one packet. ffmpeg and opusenc both default to 20 ms and offer 2.5 to 60.

Shorter frames lower latency and raise overhead: each frame has fixed costs inside the bitstream, and on a network each packet adds IP, UDP and RTP headers. Longer frames are more efficient at low bitrates, and a lost packet costs more audio. For files, 20 ms remains the sensible default; longer frames gain little at music bitrates. The 2.5 and 5 ms sizes are CELT-only, so they cannot carry SILK's in-band FEC.

Stereo, multichannel and the downmix problem

With two channels, Opus can code a stereo image cheaply using intensity stereo, which at low bitrates represents some bands as one signal plus a direction. By default it may use phase inversion for this, which sounds fine on headphones but partly cancels when a player downmixes to mono, as smart speakers and phone loudspeakers often do. OPUS_SET_PHASE_INVERSION_DISABLED, ffmpeg's apply_phase_inv 0 and opusenc's --no-phase-inv trade a little stereo quality for a clean mono downmix. Disable it for podcasts, audiobooks and anything likely to play on a single speaker.

If the content is really mono, encode it as mono (--downmix-mono or one channel at creation) rather than as identical stereo channels; the encoder will waste fewer bits on a non-existent image. OPUS_SET_FORCE_CHANNELS forces mono or stereo coding regardless of input. Surround audio uses the multistream API and a channel mapping family in the Ogg header: family 0 for mono and stereo, family 1 for the Vorbis channel orders up to 7.1, and family 255 for undefined layouts. For spatial and ambisonic formats, see spatial audio.

Loss controls and other expert knobs

In-band FEC (OPUS_SET_INBAND_FEC) makes SILK and hybrid frames carry a low-bitrate copy of the previous frame, but only when expected loss (OPUS_SET_PACKET_LOSS_PERC) is above zero; leaving expected loss at its default of 0 is the usual reason FEC is configured but never transmitted. DTX (OPUS_SET_DTX) sends almost nothing during silence. libopus 1.5 added Deep Redundancy (DRED), set with OPUS_SET_DRED_DURATION in units of 10 ms frames, which carries a compact neural representation of recent past audio; it must be enabled in the library build and supported by the decoder, so check both ends before relying on it. How these compare with RED and retransmission is covered in FEC and redundancy for real-time audio.

OPUS_SET_PREDICTION_DISABLED makes frames almost independent of each other at a quality cost, useful when a stream must be decodable from any packet. OPUS_SET_LSB_DEPTH (8 to 24, default 24) tells the encoder how many bits of the input are real; it is a hint that helps the encoder identify silence and near-silence, so set 16 for 16-bit sources.

Files add two decoder-side parameters in the Ogg header. Pre-skip is the number of samples at the start that the decoder must discard, derived from the encoder's lookahead; players that ignore it add a short burst of silence or garbage and break gapless playback. Output gain is a header gain the decoder must apply, used for loudness normalisation without re-encoding.

Five worked profiles

ProfileApplicationBitrateRate controlOther
Voice callVOIP20-32 kbit/s monoConstrained VBRWideband cap, FEC with real loss %, DTX, complexity to CPU budget
Game voice chatVOIP24-40 kbit/s monoConstrained VBRComplexity around 5 if CPU bound, FEC, 20 ms frames
Music streaming fileAUDIO96-160 kbit/s stereoUnconstrained VBRComplexity 10, music signal, 20 ms
Podcast or audiobook fileAUDIO32-64 kbit/sUnconstrained VBRMono if the source is mono, phase inversion off, speech signal
Live music over a networkRESTRICTED_LOWDELAY96-192 kbit/sConstrained VBR or CBR5 or 10 ms frames, measured lookahead

These ranges are starting points from common practice, not measured answers for your content. Here is the same set expressed in code and on the command line:

#include <opus.h>

typedef enum { CALL, GAME_CHAT, MUSIC_STREAM, SPOKEN_FILE, LIVE_MUSIC } profile_t;

/* All profiles run at 48 kHz. Values are starting points to measure, not answers. */
OpusEncoder *make_encoder(profile_t p, int channels, int expected_loss_pct) {
    int err;
    int app = (p == CALL || p == GAME_CHAT) ? OPUS_APPLICATION_VOIP
            : (p == LIVE_MUSIC) ? OPUS_APPLICATION_RESTRICTED_LOWDELAY
            : OPUS_APPLICATION_AUDIO;
    OpusEncoder *e = opus_encoder_create(48000, channels, app, &err);
    if (err != OPUS_OK) return NULL;

    switch (p) {
    case CALL:
        opus_encoder_ctl(e, OPUS_SET_SIGNAL(OPUS_SIGNAL_VOICE));
        opus_encoder_ctl(e, OPUS_SET_BITRATE(24000));
        opus_encoder_ctl(e, OPUS_SET_MAX_BANDWIDTH(OPUS_BANDWIDTH_WIDEBAND));
        opus_encoder_ctl(e, OPUS_SET_INBAND_FEC(1));
        opus_encoder_ctl(e, OPUS_SET_PACKET_LOSS_PERC(expected_loss_pct));
        opus_encoder_ctl(e, OPUS_SET_DTX(1));
        break;
    case GAME_CHAT:
        opus_encoder_ctl(e, OPUS_SET_SIGNAL(OPUS_SIGNAL_VOICE));
        opus_encoder_ctl(e, OPUS_SET_BITRATE(32000));
        opus_encoder_ctl(e, OPUS_SET_COMPLEXITY(5));        /* share the CPU with the game */
        opus_encoder_ctl(e, OPUS_SET_INBAND_FEC(1));
        opus_encoder_ctl(e, OPUS_SET_PACKET_LOSS_PERC(expected_loss_pct));
        break;
    case MUSIC_STREAM:
        opus_encoder_ctl(e, OPUS_SET_SIGNAL(OPUS_SIGNAL_MUSIC));
        opus_encoder_ctl(e, OPUS_SET_BITRATE(channels == 2 ? 128000 : 64000));
        opus_encoder_ctl(e, OPUS_SET_VBR_CONSTRAINT(0));    /* unconstrained VBR for files */
        opus_encoder_ctl(e, OPUS_SET_COMPLEXITY(10));
        break;
    case SPOKEN_FILE:
        opus_encoder_ctl(e, OPUS_SET_SIGNAL(OPUS_SIGNAL_VOICE));
        opus_encoder_ctl(e, OPUS_SET_BITRATE(channels == 2 ? 48000 : 32000));
        opus_encoder_ctl(e, OPUS_SET_VBR_CONSTRAINT(0));
        opus_encoder_ctl(e, OPUS_SET_COMPLEXITY(10));
        break;
    case LIVE_MUSIC:
        opus_encoder_ctl(e, OPUS_SET_BITRATE(channels == 2 ? 128000 : 96000));
        /* 5 ms frames: pass 240 samples per channel to each opus_encode() call */
        break;
    }
    opus_int32 v;
    opus_encoder_ctl(e, OPUS_GET_LOOKAHEAD(&v));  /* log it: it is part of your latency */
    return e;
}
# ffmpeg: -b:a is in bit/s
ffmpeg -i talk.wav  -c:a libopus -b:a 32000  -vbr on -application voip  talk.opus
ffmpeg -i album.flac -c:a libopus -b:a 128000 -vbr on -compression_level 10 album.opus
ffmpeg -i mix.wav   -c:a libopus -b:a 96000  -vbr constrained -frame_duration 10 mix.webm

# opusenc: --bitrate is in kbit/s
opusenc --bitrate 32  --speech talk.wav talk.opus
opusenc --bitrate 128 --vbr --comp 10 album.wav album.opus
opusenc --bitrate 64  --no-phase-inv  lecture.wav lecture.opus
opusenc --bitrate 32  --downmix-mono  mono_source.wav mono.opus

To choose a real number, sweep. Encode a few minutes of representative material at a ladder of bitrates and modes, record the actual bitrate, and score each output with an objective metric plus a blind listening panel for the bitrates near your candidate:

import subprocess, os, csv

def encode(src, kbps, mode, frame_ms, out):
    subprocess.run(["ffmpeg", "-v", "error", "-y", "-i", src, "-c:a", "libopus",
                    "-b:a", str(kbps * 1000), "-vbr", mode,
                    "-frame_duration", str(frame_ms), out], check=True)
    return os.path.getsize(out)

def sweep(src, seconds, rows):
    for kbps in (16, 24, 32, 48, 64, 96, 128):
        for mode in ("on", "constrained", "off"):
            out = f"sweep_{kbps}_{mode}.opus"
            size = encode(src, kbps, mode, 20, out)
            # container overhead included; good enough to compare modes
            rows.append(dict(kbps=kbps, mode=mode, actual_kbps=round(size * 8 / seconds / 1000, 1), file=out))
    # score each file with your objective metric and a blind listening panel, then plot
    with open("sweep.csv", "w", newline="") as f:
        w = csv.DictWriter(f, fieldnames=rows[0].keys()); w.writeheader(); w.writerows(rows)

The curve flattens: past some bitrate, listeners stop hearing a difference. Pick the knee for your content, then add margin for the hardest material you carry.

Failure modes

SymptomLikely causeFix
Music sounds dull or warblyBitrate too low for fullband, or VOIP application on musicAUDIO application, music signal hint, higher bitrate
ffmpeg bitrate error or wrong rateBitrate given in kbit/s to an option expecting bit/sWrite 64000 or 64k
Speech fine on headphones, hollow on one speakerPhase inversion cancelling in mono downmixDisable phase inversion or encode mono
FEC enabled but no recoveryExpected loss left at 0, or CELT-only frame sizeSet loss % from measurements; use 10 ms or longer frames
Audio glitches on low-end phonesComplexity too high for the real-time deadlineMeasure encode time per frame and lower complexity
Click or gap at the start of filesPlayer ignores pre-skipUse a conforming decoder or apply pre-skip yourself

If artefacts appear only under loss, look at concealment rather than the encoder; packet loss concealment explains what the decoder does when a packet is missing.

What to do next

  1. Write down each Opus use in your product, its delivery path (RTP or file) and its latency limit.
  2. For each, set application, signal hint and channel count deliberately, and query and log lookahead and complexity.
  3. Pick rate control by path: unconstrained VBR for files, constrained VBR for live streams, CBR where packet sizes must not leak.
  4. Check bitrate units in every ffmpeg and opusenc command and test one output by ear.
  5. Turn off phase inversion for content likely to play on a single speaker, and encode mono sources as mono.
  6. Run the bitrate sweep on representative material and choose the knee, not a number from a table.
Key takeaway: Opus asks for sample rate, channels and application at creation; every other parameter can change per packet. Choose rate control by delivery path, set the signal hint and bandwidth when you know the content, spend complexity where the CPU allows, watch bitrate units and phase inversion, and pick bitrates from a measured sweep on your own audio rather than from defaults.