AAC is the audio inside most MP4 files, HLS streams and podcasts. Engineers meet it as a flag, -c:a aac -b:a 128k, and then again as a bug: a stream that plays dull on one device, a drum hit with a smear of noise before it, a file silent in one player. Each of those bugs is explained by how the codec is built.
This article builds AAC from its parts: the MDCT, the hearing model that decides how much noise each band can hide, bit allocation, what HE-AAC adds, and the headers you will end up parsing. For choosing between codecs, and for containers, priming and gapless playback, read the MP3 vs AAC vs Opus comparison.
The core idea: spend bits only where the ear will notice
CD-quality stereo is 44,100 samples per second times 16 bits times 2 channels: 1,411 kbit/s. AAC typically carries the same music at 128 to 256 kbit/s, which lossless coding cannot reach. AAC is perceptual: it throws information away, choosing what to discard with a model of human hearing.
Two facts about hearing do the work. A loud sound masks quieter sounds near it in frequency, and the ear's frequency resolution is coarse at high frequencies. So an encoder moves the signal into the frequency domain, estimates per band how much noise could be added unheard, and quantises each band just coarsely enough to stay under that threshold. The decoder has no hearing model; it reverses quantisation and transform. All the intelligence, and all quality differences between encoders, sit on the encoding side.
Step 1: the MDCT filterbank
AAC frames carry 1,024 new samples per channel, transformed with the modified discrete cosine transform over a 2,048-sample window, so consecutive windows overlap by half. The MDCT is still critically sampled: as many coefficients as new samples. The aliasing each inverse transform introduces is cancelled by the neighbouring window's when they are overlap-added (time-domain aliasing cancellation), provided the window meets the Princen-Bradley condition, as AAC's sine and Kaiser-Bessel-derived windows do.
The overlap makes quantisation errors fade across windows instead of stepping at block edges. You can verify the property in a few lines.
import numpy as np
N = 1024 # coefficients per long window; the window spans 2N samples
n = np.arange(2 * N)
k = np.arange(N)
basis = np.cos(np.pi / N * (n[:, None] + 0.5 + N / 2) * (k[None, :] + 0.5))
win = np.sin(np.pi * (n + 0.5) / (2 * N)) # sine window: win[i]**2 + win[i+N]**2 == 1
def mdct(block): # 2N samples in, N coefficients out
return (win * block) @ basis
def imdct(coefs): # N coefficients in, 2N aliased samples out
return win * (basis @ coefs) * (2 / N)
x = np.random.default_rng(0).normal(size=6 * N)
y = np.zeros_like(x)
for start in range(0, len(x) - N, N): # hop N: each sample is covered by two windows
y[start:start + 2 * N] += imdct(mdct(x[start:start + 2 * N]))
# The aliasing in neighbouring blocks cancels (TDAC); the first and last N samples lack a partner.
print(np.max(np.abs(x[N:-N] - y[N:-N]))) # ~1e-13: perfect reconstruction before quantisationAt 44.1 kHz the 1,024 coefficients each cover about 21.5 Hz, good frequency resolution for tonal music. The price is time resolution: one long window spans about 46 ms.
Step 2: window switching and the pre-echo problem
Quantisation noise spreads across the whole window. If a window holds silence then a castanet click, the noise budget is set by the click but the noise also lands in the silence before it, where nothing masks it: pre-echo. Backward masking covers only a few milliseconds, far less than 46 ms.
The main defence is block switching. When the psychoacoustic model sees perceptual entropy jump, the frame switches to eight short windows of 256 samples, eight sets of 128 coefficients, each spanning about 5.8 ms. Asymmetric start and stop windows keep aliasing cancellation intact. Short windows cost bits, since frequency resolution drops eightfold, so encoders use them only when needed.
The second defence is temporal noise shaping (TNS): a linear prediction filter run across frequency inside one frame. Prediction across frequency is the dual of prediction across time, so it shapes the noise's time envelope to follow the signal's, which helps speech and transients too mild for short windows.
Step 3: the psychoacoustic model and scalefactor bands
The model runs its own FFT analysis. Per frame it estimates energy in partitions matched to the ear's critical bands, spreads it across neighbours to model masking, judges how tonal or noise-like each region is (noise masks better than tones), and applies the threshold of hearing. The output is a signal-to-mask ratio per band and a perceptual entropy figure estimating how many bits the frame needs.
The coefficients are grouped into scalefactor bands, narrow at low frequencies and wide at high ones. At 44.1 and 48 kHz a long window has 49 bands and a short window has 14. Each band's scalefactor sets the quantiser step for all its coefficients: the knob the encoder turns, trading bits for noise.
Step 4: quantisation and the two loops
AAC quantises with a power law rather than uniformly. Roughly, each coefficient's magnitude is raised to the power 3/4, divided by the band's step and rounded with a small bias, so noise grows with signal level, as the ear tolerates. One scalefactor step changes the step size by 2^(1/4) in amplitude, which is 1.5 dB.
The classic reference encoder uses two nested loops. The inner rate loop adjusts a global gain until the Huffman-coded frame fits the budget. The outer distortion loop raises the precision of any band whose noise exceeds its allowed level, then re-runs the rate loop. Modern encoders use faster heuristics and trellis searches, a large part of why encoders sound different at the same bitrate.
Quantised values are packed with one of eleven Huffman codebooks chosen per section of bands. A bit reservoir lets an easy frame donate bits to a later hard one, within a decoder buffer of 6,144 bits per channel, so even constant-bitrate AAC varies frame to frame.
Worked example: the bit budget at 128 kbit/s
Take 44.1 kHz stereo at 128 kbit/s. A frame is 1,024 samples, or 1,024 / 44,100 = 23.2 ms. The budget is 128,000 times 0.0232, about 2,972 bits per frame, roughly 1,486 per channel: under 1.5 bits per coefficient before side information.
That works only because most coefficients quantise to zero: bands above the encoder's lowpass are not coded at all, and quiet bands under a loud masker get large steps. Two more tools stretch the budget. Mid/side stereo, switchable per band, codes sum and difference when channels are similar, leaving the side channel nearly empty. Perceptual noise substitution sends only the energy of a noise-like band, and the decoder synthesises noise there.
Halve the bitrate and the arithmetic stops working for full-band LC: the encoder must lowpass hard or let noise become audible. That is the gap HE-AAC fills.
HE-AAC: spectral band replication and parametric stereo
Spectral band replication (SBR, object type 5) runs the AAC-LC core at half the output sample rate, coding only the lower half of the spectrum. The decoder rebuilds the high band by transposing the low band upward in a QMF filterbank and shaping it with envelope and tonality data costing a few kbit/s. Upper harmonics resemble lower ones, so this beats lowpassing, though it can sound metallic.
Parametric stereo (PS, object type 29, added in HE-AAC v2) codes a mono downmix plus inter-channel level, phase and correlation parameters. It suits very low bitrates and fails audibly on wide, complex mixes. Typically LC serves music at higher bitrates, HE-AAC lower streaming tiers and v2 the lowest; listen on your own material.
Signalling matters. A stream can declare SBR explicitly or only implicitly, embedding SBR data where older decoders skip it. An LC-only decoder given implicit HE-AAC plays the core alone at half the sample rate: dull audio, and sometimes a wrong reported rate.
The bitstream: AudioSpecificConfig and ADTS
Raw AAC frames do not describe themselves, so the transport does. In MP4 it is the AudioSpecificConfig (ASC) in the esds box: 5 bits of object type (2 for LC), 4 of sample-rate index, 4 of channel configuration, then flag bits. AAC-LC at 44.1 kHz stereo is the two bytes 12 10; at 48 kHz it is 11 90. In streams without a container, such as raw .aac files and MPEG-TS, every frame instead carries a 7-byte ADTS header, 9 with CRC. It repeats the same information, using profile = object type minus one, and adds the frame length, so a parser can resynchronise anywhere.
This parser walks ADTS frames and builds an ASC. To get duration, multiply frames by 1,024 and divide by the header's rate; for HE-AAC that is usually the core rate, which still gives the right answer.
RATES = [96000, 88200, 64000, 48000, 44100, 32000, 24000,
22050, 16000, 12000, 11025, 8000, 7350]
def adts_frames(buf: bytes):
i = 0
while i + 7 <= len(buf):
b = buf[i:i + 7]
if b[0] != 0xFF or (b[1] & 0xF6) != 0xF0: # 12-bit sync + layer == 0
raise ValueError(f"lost sync at byte {i}")
crc = not (b[1] & 0x01) # protection_absent == 0 -> 2-byte CRC
aot = (b[2] >> 6) + 1 # profile field is object type minus 1
rate = RATES[(b[2] >> 2) & 0x0F]
chans = ((b[2] & 0x01) << 2) | (b[3] >> 6)
length = ((b[3] & 0x03) << 11) | (b[4] << 3) | (b[5] >> 5) # includes the header
blocks = (b[6] & 0x03) + 1
yield dict(offset=i, aot=aot, rate=rate, chans=chans, length=length, blocks=blocks, crc=crc)
i += length
def audio_specific_config(aot, rate_index, chans): # 2 bytes for AAC-LC
return ((aot << 11) | (rate_index << 7) | (chans << 3)).to_bytes(2, "big")
assert audio_specific_config(2, 4, 2).hex() == "1210" # LC, 44.1 kHz, stereo
assert audio_specific_config(2, 3, 2).hex() == "1190" # LC, 48 kHz, stereo
Encoders and commands
The encoder matters more than the format. FFmpeg's built-in aac encoder is always available and produces AAC-LC. The Fraunhofer libfdk_aac supports HE-AAC, HE-AAC v2 and -vbr 1 to 5, but its licence is GPL-incompatible, so FFmpeg builds with it need --enable-nonfree and cannot be redistributed. On Apple platforms aac_at uses the system encoder. AAC is patent-licensed through a pool; check your own obligations.
# Built-in FFmpeg encoder: AAC-LC, constant target bitrate.
ffmpeg -i master.wav -c:a aac -b:a 160k -ar 48000 music_lc.m4a
# Fraunhofer FDK encoder (FFmpeg must be built with --enable-nonfree; the binary is not redistributable).
ffmpeg -i master.wav -c:a libfdk_aac -profile:a aac_he -b:a 64k music_he.m4a
ffmpeg -i master.wav -c:a libfdk_aac -profile:a aac_he_v2 -b:a 32k music_hev2.m4a
ffmpeg -i master.wav -c:a libfdk_aac -vbr 4 music_vbr.m4a
# Rewrap a raw ADTS stream into MP4 without re-encoding: the 7-byte headers become one ASC.
ffmpeg -i stream.aac -c copy -bsf:a aac_adtstoasc stream.m4a
# Verify what you actually produced, not what you asked for.
ffprobe -v error -show_entries stream=codec_name,profile,sample_rate,channels,bit_rate \
-of default=nw=1 music_he.m4a
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Noise smear before drums or plucks | Pre-echo: block switching or TNS not triggered, or bitrate too low | Raise bitrate, try another encoder, check transient-heavy test clips |
| Dull audio at half the sample rate | LC-only decoder ignoring implicitly signalled SBR | Signal SBR explicitly, or ship an LC rendition for old devices |
| Silence or wrong pitch after remux | ASC does not match the frames, or ADTS headers left inside MP4 | Use aac_adtstoasc; compare ffprobe output with the source |
| Clipping on playback of loud masters | Decoded peaks exceed the source's because quantisation noise adds to peaks | Leave about 1 dB of true-peak headroom before encoding |
| Collapsed or phasey stereo at low rates | Parametric stereo on wide or out-of-phase mixes | Use HE-AAC v1 or a higher-rate LC tier for that content |
| Audio offset from video | Encoder priming not trimmed by the player | Write edit lists; see the comparison article on priming |
Two adjacent problems surface in AAC pipelines: sample-rate mismatches, covered in audio resampling, and inconsistent loudness, covered in loudness normalisation. Normalise and set headroom before the encoder, never after.
Trade-offs against the alternatives
AAC's strength is compatibility: hardware decoders nearly everywhere and native support in HLS and MP4. Its weaknesses are frame latency (low-delay AAC-LD and AAC-ELD exist but are less supported) and efficiency at low bitrates. For real-time voice, Opus is usually better; for the widest device reach, AAC-LC with an HE-AAC low tier remains the default.
What to do next
- Run the MDCT snippet, then quantise the coefficients crudely and listen to the noise you introduced.
- Run ffprobe on your production outputs and confirm profile, sample rate and channels match what you intended.
- Build a 30-second torture clip of castanets, harpsichord, applause and wide stereo, and compare encoders and bitrates on it by ear.
- Decide your ladder explicitly: LC rates for music, an HE-AAC or HE-AAC v2 tier only where bandwidth demands it.
- Set loudness and about 1 dB of true-peak headroom before encoding.
- Add an ADTS or MP4 sanity check to CI: frame count, duration and ASC versus expectations.