Uncompressed stereo audio at 48 kHz and 16 bits is 1,536 kbit/s. A good music stream uses around a tenth of that, and a clear voice call can use around one percent. The gap is not closed by clever zip-style compression, which barely shrinks audio at all. It is closed by throwing away information the ear will not miss, and spending the remaining bits where the ear is most sensitive.
This article explains how that works inside, from first principles, and builds a small working transform codec in numpy that you can run and modify. Particular codecs are covered elsewhere on this site: Opus for real-time voice and music, neural codecs for learned compression and audio tokens, and MP3 vs AAC vs Opus for choosing a streaming format. If you need a refresher on samples, bit depth and buffers, start with PCM audio. After this page, those designs should read as variations on one idea.
Four families of codec
Codecs split by what they model. Waveform codecs such as G.711 and ADPCM encode the samples themselves with fewer bits, so they are simple and low-delay but need tens of kbit/s even for narrowband speech. Speech codecs based on linear prediction model how a voice is produced: a source of pulses or noise shaped by a vocal-tract filter. They reach very low rates for speech and handle music poorly. Perceptual transform codecs, including MP3, AAC, the CELT layer of Opus and Bluetooth's LC3, move the signal into the frequency domain and shape quantisation noise to hide below what the ear can detect. They handle any sound well. Neural codecs learn an encoder, a quantiser and a decoder end to end, and reach low rates at the cost of compute.
The perceptual transform codec is the workhorse, and its pipeline is the one to understand in detail.
The filterbank: MDCT and aliasing cancellation
Quantising in the frequency domain lets the codec put noise where the signal will mask it. The transform of choice is the modified discrete cosine transform. Each block holds 2N samples and produces only N real coefficients, and consecutive blocks overlap by half. That looks like information loss, and per block it is: the inverse transform returns a time-aliased version of the block. But the aliasing in the second half of one block is equal and opposite to the aliasing in the first half of the next, so overlap-adding the outputs cancels it exactly. This is time-domain aliasing cancellation, and it is why the MDCT is critically sampled, producing N new coefficients for N new samples, while avoiding the block-edge clicks a non-overlapping transform would cause.
Cancellation requires a window satisfying the Princen-Bradley condition: the squares of the two overlapping halves sum to one. The sine window below does; AAC also offers a Kaiser-Bessel-derived window.
import numpy as np
N = 480 # hop size: 10 ms at 48 kHz; each block is 2N samples
n, k = np.arange(2 * N), np.arange(N)
WIN = np.sin(np.pi * (n + 0.5) / (2 * N)) # Princen-Bradley: w[n]^2 + w[n+N]^2 = 1
BASIS = np.cos(np.pi / N * (n[None, :] + 0.5 + N / 2) * (k[:, None] + 0.5)) # [N, 2N]
def mdct(block): # 2N samples -> N real coefficients
return BASIS @ (WIN * block)
def imdct(coefs): # N coefficients -> 2N samples to overlap-add
return WIN * (BASIS.T @ coefs) * (2 / N)This is the O(N squared) matrix form, fine for learning; real codecs use an FFT-based O(N log N) algorithm. With overlap-add and no quantisation the output matches the input to floating-point precision; checking this page, the maximum error was around 1e-13.
Block switching and pre-echo
Long blocks give fine frequency resolution, which is good for tonal music, but spread quantisation noise across the whole block in time. If a castanet click arrives near the end of a long block, the noise it causes spreads to the quiet part before the click, where nothing masks it. The result is pre-echo, a smeared ghost before each attack. Codecs fight it by detecting transients and switching to short blocks: AAC divides a 1,024-coefficient long frame into eight short transforms of 128, using transition windows so aliasing still cancels. Other tools include temporal noise shaping in AAC and adaptive time-frequency resolution in CELT. When tuning or comparing codecs, castanets, glockenspiel and speech plosives are the test signals that expose this.
Psychoacoustics: deciding where noise can hide
The ear analyses sound in critical bands, roughly constant width at low frequencies and widening with frequency. A loud component raises the hearing threshold in its own band and, less, in neighbouring bands and just after it in time; this is masking. There is also an absolute threshold below which nothing is heard, lowest around 2 to 5 kHz. The encoder's perceptual model estimates, per band and per block, the largest noise that would stay inaudible, and that becomes the target for the quantiser. Noise is not removed; it is placed where it will be masked.
Bands are therefore narrow at low frequencies and wide at high ones, as in the band table below. The perceptual model runs only in the encoder, so encoders keep improving while existing decoders still play the stream.
Scale factors, quantisation and bit allocation
For each band the encoder sends a scale factor describing its energy, then quantises the coefficients with a step size derived from that energy and the allowed noise. A uniform quantiser with step size delta adds noise with power delta squared over twelve, which turns the perceptual model's answer directly into a step size. The toy codec replaces a perceptual model with a per-band SNR target so you can see the mechanism.
# band edges in MDCT bins; at N = 480 and 48 kHz each bin is 50 Hz wide
EDGES = [0, 4, 8, 12, 16, 20, 24, 28, 32, 40, 48, 56, 64, 80, 96, 112, 128,
160, 192, 240, 288, 352, 416, 480]
def step_size(sf, snr_db):
# uniform quantiser noise power is step^2 / 12, so this step gives noise = rms^2 / 10^(snr/10)
return 2 ** (sf / 6) * 10 ** (-snr_db / 20) * np.sqrt(12)
def encode_frame(coefs, snr_db):
frame = []
for b, (lo, hi) in enumerate(zip(EDGES[:-1], EDGES[1:])):
band = coefs[lo:hi]
rms = np.sqrt(np.mean(band ** 2)) + 1e-9
sf = int(np.round(6 * np.log2(rms))) # scale factor in ~1 dB steps
frame.append((sf, np.round(band / step_size(sf, snr_db[b])).astype(np.int32)))
return frame
def decode_frame(frame, snr_db):
coefs = np.zeros(N)
for b, (lo, hi) in enumerate(zip(EDGES[:-1], EDGES[1:])):
sf, q = frame[b]
coefs[lo:hi] = q * step_size(sf, snr_db[b])
return coefs
def roundtrip(x, snr_db):
pad = np.concatenate([np.zeros(N), x, np.zeros(2 * N)])
y = np.zeros(len(pad))
for s in range(0, len(pad) - 2 * N + 1, N):
frame = encode_frame(mdct(pad[s:s + 2 * N]), snr_db)
y[s:s + 2 * N] += imdct(decode_frame(frame, snr_db))
return y[N:N + len(x)]
def payload_bits(x, snr_db):
# zeroth-order entropy of the quantised values plus 6 bits per scale factor
qs, sfs = [], 0
pad = np.concatenate([np.zeros(N), x, np.zeros(2 * N)])
for s in range(0, len(pad) - 2 * N + 1, N):
for sf, q in encode_frame(mdct(pad[s:s + 2 * N]), snr_db):
qs.append(q); sfs += 6
_, counts = np.unique(np.concatenate(qs), return_counts=True)
p = counts / counts.sum()
return -(counts * np.log2(p)).sum() + sfsReal codecs run a rate loop around this. For constant bitrate, the encoder adjusts a global offset until the frame fits its budget, often with a bit reservoir that lets hard frames borrow from easy ones. For variable bitrate, it holds the quality target and lets the size vary. In both, the scarce resource is bits, and the perceptual model decides which bands deserve them.
Entropy coding and the bitstream
Quantised coefficients are mostly small integers and zeros, so lossless entropy coding shrinks them further. MP3 and AAC use Huffman code books chosen per region or section; Opus uses a range coder, and its CELT layer codes the shape of each band with pyramid vector quantisation while transmitting band energy separately, so the energy of each band is preserved even at low rates. The bitstream also carries side information, such as block type, scale factors and code book choices, which can be a large share of the bits at low rates. The toy codec's payload_bits estimates the cost from the entropy of the quantised values, a reasonable proxy for a well-designed coder.
Speech codecs and hybrids
Speech codecs take a different route. Linear prediction fits, every 10 to 20 ms, a short filter that predicts each sample from the previous ones; that filter captures the vocal-tract resonances. What remains is an excitation signal, which CELP codecs pick from code books by analysis-by-synthesis, trying candidates through the filter and keeping the one closest to the input with perceptual weighting. A few kbit/s give intelligible speech, but music and background noise suffer because they do not fit the source-filter model. Opus combines both worlds: SILK, a linear-prediction codec, for speech at low rates, CELT, an MDCT codec, for music and higher rates, and a hybrid mode with SILK for low frequencies and CELT above.
Latency is set by frames and overlap
A codec cannot emit a frame until it has seen the frame, plus any lookahead the transform overlap and analysis need. That algorithmic delay comes before network, jitter-buffer and device buffers. Opus frames range from 2.5 to 60 ms, and libopus at its default 20 ms frames has 26.5 ms of algorithmic delay. LC3, the Bluetooth LE Audio codec, uses 7.5 or 10 ms frames. AAC-LC frames are 1,024 samples, about 21 ms at 48 kHz, plus overlap, which is acceptable for streaming but long for conversation. Shorter frames cut delay but cost efficiency, because side information is sent more often and frequency resolution drops. For interactive audio, budget the full mouth-to-ear path, and treat the codec's share as one line item. Losses add another trade-off: shorter frames lose less audio per lost packet, and concealment, covered in packet loss concealment, fills the gaps.
Worked example: spending bits where they count
Run the toy codec on one second of a 440 Hz tone with added noise, once with the same 20 dB SNR target in every band and once with a tilted target: 30 dB in the lowest eight bands, 20 dB in the middle and 8 dB at the top.
fs = 48000
t = np.arange(fs) / fs
x = 0.5 * np.sin(2 * np.pi * 440 * t) + 0.1 * np.random.default_rng(0).standard_normal(fs)
flat = [20] * 23 # same SNR everywhere
tilted = [30] * 8 + [20] * 8 + [8] * 7 # spend bits low, save them high
for name, target in [("flat", flat), ("tilted", tilted)]:
y = roundtrip(x, target)
snr = 10 * np.log10(np.sum(x ** 2) / np.sum((x - y) ** 2))
kbps = payload_bits(x, target) / 1000 # one second of audio
print(f"{name:7s} SNR {snr:5.1f} dB about {kbps:6.1f} kbit/s")When this page was checked, the flat profile gave 20.5 dB SNR at an estimated 189 kbit/s and the tilted one 20.2 dB at about 148 kbit/s: almost the same SNR for 22 percent fewer bits, because the tone's energy sits in the low bands where the bits went. Your figures will depend on the signal. A perceptual model goes further, placing noise where it is masked even when SNR drops, which is why SNR is the wrong metric for a perceptual codec. Measure with listening tests such as MUSHRA, ITU-R BS.1534, or with objective models built for codecs, such as POLQA, ITU-T P.863, for speech and ViSQOL for speech and audio, and listen to critical items yourself.
Failure modes
- Pre-echo. Smeared attacks before transients; check block switching on castanets and plosives.
- Birdies and band dropouts. At low rates, high bands flicker in and out as their coefficients round to zero, creating watery artefacts. Noise filling and band-energy preservation are the usual remedies.
- Tandem coding. Decoding and re-encoding through several lossy codecs compounds artefacts. Transcode once, from the best source.
- Clipping after decode. Quantisation noise can push peaks past full scale even when the input did not clip. Leave headroom.
- Sample-rate mismatches. Feeding 44.1 kHz audio to an encoder set for 48 kHz shifts pitch or forces a hidden resampler.
- Ignoring priming samples. Encoder delay must be trimmed on decode or files gain silence and gapless playback breaks.
Operational guidance and trade-offs
Pick by use case before quality: interactive voice needs low delay and loss resilience, which points to Opus or LC3 on Bluetooth; distribution needs broad decoder support; archives need lossless formats such as FLAC. Store a lossless master and encode derivatives from it. Use variable bitrate for files and constrained or constant bitrate for real-time links whose capacity is fixed. Test with your own content, because speech-heavy, music and noisy field recordings rank codecs differently. And keep the encoder version in your metadata, since perceptual models change between releases and a re-encode can sound different at the same bitrate.
What to do next
- Run the toy codec and confirm perfect reconstruction with quantisation disabled.
- Change the per-band SNR targets and listen to the output on speech and music, comparing against SNR numbers.
- Add a transient detector that halves N for blocks with sudden energy jumps, and hear pre-echo change on a castanet sample.
- Write down the latency budget for your product and check the codec's frame and lookahead fit inside it.
- Set up an objective quality check, ViSQOL or POLQA, plus a small listening panel for release decisions.
- Keep lossless masters, encode once per target, and record encoder versions and settings.