AAC is the audio inside most MP4 files, HLS streams and podcasts. Engineers meet it as a flag, -c:a aac -b:a 128k, and then again as a bug: a stream that plays dull on one device, a drum hit with a smear of noise before it, a file silent in one player. Each of those bugs is explained by how the codec is built.

This article builds AAC from its parts: the MDCT, the hearing model that decides how much noise each band can hide, bit allocation, what HE-AAC adds, and the headers you will end up parsing. For choosing between codecs, and for containers, priming and gapless playback, read the MP3 vs AAC vs Opus comparison.

Advertisement

The core idea: spend bits only where the ear will notice

CD-quality stereo is 44,100 samples per second times 16 bits times 2 channels: 1,411 kbit/s. AAC typically carries the same music at 128 to 256 kbit/s, which lossless coding cannot reach. AAC is perceptual: it throws information away, choosing what to discard with a model of human hearing.

Two facts about hearing do the work. A loud sound masks quieter sounds near it in frequency, and the ear's frequency resolution is coarse at high frequencies. So an encoder moves the signal into the frequency domain, estimates per band how much noise could be added unheard, and quantises each band just coarsely enough to stay under that threshold. The decoder has no hearing model; it reverses quantisation and transform. All the intelligence, and all quality differences between encoders, sit on the encoding side.

An AAC-LC encoder: where the bits are spent and whyPCM in1024 new samplesMDCT filterbank2048 long / 8 x 256 shortTNSnoise shaped in timeM/S, PNSper band choicesPsychoacoustic modelFFT, masking, PEwindow decisionallowed noise per band (SMR)Quantise + scalefactorsrate loop / distortion loopHuffman coding11 spectral codebooksraw_data_blockbit reservoir smooths rateADTS or MP4ASC in esds boxHE-AAC: SBR encodercore at half rate + high-band envelopeHE-AAC v2: parametric stereomono core + stereo image parametersThe decoder runs the same chain backwards and never sees the psychoacoustic model: it only reads the decisions.
The AAC-LC encoder chain. The psychoacoustic model steers two decisions, window length and allowed noise per band; everything after quantisation is lossless packing.

Step 1: the MDCT filterbank

AAC frames carry 1,024 new samples per channel, transformed with the modified discrete cosine transform over a 2,048-sample window, so consecutive windows overlap by half. The MDCT is still critically sampled: as many coefficients as new samples. The aliasing each inverse transform introduces is cancelled by the neighbouring window's when they are overlap-added (time-domain aliasing cancellation), provided the window meets the Princen-Bradley condition, as AAC's sine and Kaiser-Bessel-derived windows do.

The overlap makes quantisation errors fade across windows instead of stepping at block edges. You can verify the property in a few lines.

import numpy as np

N = 1024                                   # coefficients per long window; the window spans 2N samples
n = np.arange(2 * N)
k = np.arange(N)
basis = np.cos(np.pi / N * (n[:, None] + 0.5 + N / 2) * (k[None, :] + 0.5))
win = np.sin(np.pi * (n + 0.5) / (2 * N))  # sine window: win[i]**2 + win[i+N]**2 == 1

def mdct(block):                           # 2N samples in, N coefficients out
    return (win * block) @ basis

def imdct(coefs):                          # N coefficients in, 2N aliased samples out
    return win * (basis @ coefs) * (2 / N)

x = np.random.default_rng(0).normal(size=6 * N)
y = np.zeros_like(x)
for start in range(0, len(x) - N, N):      # hop N: each sample is covered by two windows
    y[start:start + 2 * N] += imdct(mdct(x[start:start + 2 * N]))

# The aliasing in neighbouring blocks cancels (TDAC); the first and last N samples lack a partner.
print(np.max(np.abs(x[N:-N] - y[N:-N])))  # ~1e-13: perfect reconstruction before quantisation

At 44.1 kHz the 1,024 coefficients each cover about 21.5 Hz, good frequency resolution for tonal music. The price is time resolution: one long window spans about 46 ms.

Advertisement

Step 2: window switching and the pre-echo problem

Quantisation noise spreads across the whole window. If a window holds silence then a castanet click, the noise budget is set by the click but the noise also lands in the silence before it, where nothing masks it: pre-echo. Backward masking covers only a few milliseconds, far less than 46 ms.

The main defence is block switching. When the psychoacoustic model sees perceptual entropy jump, the frame switches to eight short windows of 256 samples, eight sets of 128 coefficients, each spanning about 5.8 ms. Asymmetric start and stop windows keep aliasing cancellation intact. Short windows cost bits, since frequency resolution drops eightfold, so encoders use them only when needed.

The second defence is temporal noise shaping (TNS): a linear prediction filter run across frequency inside one frame. Prediction across frequency is the dual of prediction across time, so it shapes the noise's time envelope to follow the signal's, which helps speech and transients too mild for short windows.

Step 3: the psychoacoustic model and scalefactor bands

The model runs its own FFT analysis. Per frame it estimates energy in partitions matched to the ear's critical bands, spreads it across neighbours to model masking, judges how tonal or noise-like each region is (noise masks better than tones), and applies the threshold of hearing. The output is a signal-to-mask ratio per band and a perceptual entropy figure estimating how many bits the frame needs.

The coefficients are grouped into scalefactor bands, narrow at low frequencies and wide at high ones. At 44.1 and 48 kHz a long window has 49 bands and a short window has 14. Each band's scalefactor sets the quantiser step for all its coefficients: the knob the encoder turns, trading bits for noise.

Step 4: quantisation and the two loops

AAC quantises with a power law rather than uniformly. Roughly, each coefficient's magnitude is raised to the power 3/4, divided by the band's step and rounded with a small bias, so noise grows with signal level, as the ear tolerates. One scalefactor step changes the step size by 2^(1/4) in amplitude, which is 1.5 dB.

The classic reference encoder uses two nested loops. The inner rate loop adjusts a global gain until the Huffman-coded frame fits the budget. The outer distortion loop raises the precision of any band whose noise exceeds its allowed level, then re-runs the rate loop. Modern encoders use faster heuristics and trellis searches, a large part of why encoders sound different at the same bitrate.

Quantised values are packed with one of eleven Huffman codebooks chosen per section of bands. A bit reservoir lets an easy frame donate bits to a later hard one, within a decoder buffer of 6,144 bits per channel, so even constant-bitrate AAC varies frame to frame.

Worked example: the bit budget at 128 kbit/s

Take 44.1 kHz stereo at 128 kbit/s. A frame is 1,024 samples, or 1,024 / 44,100 = 23.2 ms. The budget is 128,000 times 0.0232, about 2,972 bits per frame, roughly 1,486 per channel: under 1.5 bits per coefficient before side information.

That works only because most coefficients quantise to zero: bands above the encoder's lowpass are not coded at all, and quiet bands under a loud masker get large steps. Two more tools stretch the budget. Mid/side stereo, switchable per band, codes sum and difference when channels are similar, leaving the side channel nearly empty. Perceptual noise substitution sends only the energy of a noise-like band, and the decoder synthesises noise there.

Halve the bitrate and the arithmetic stops working for full-band LC: the encoder must lowpass hard or let noise become audible. That is the gap HE-AAC fills.

HE-AAC: spectral band replication and parametric stereo

Spectral band replication (SBR, object type 5) runs the AAC-LC core at half the output sample rate, coding only the lower half of the spectrum. The decoder rebuilds the high band by transposing the low band upward in a QMF filterbank and shaping it with envelope and tonality data costing a few kbit/s. Upper harmonics resemble lower ones, so this beats lowpassing, though it can sound metallic.

Parametric stereo (PS, object type 29, added in HE-AAC v2) codes a mono downmix plus inter-channel level, phase and correlation parameters. It suits very low bitrates and fails audibly on wide, complex mixes. Typically LC serves music at higher bitrates, HE-AAC lower streaming tiers and v2 the lowest; listen on your own material.

Signalling matters. A stream can declare SBR explicitly or only implicitly, embedding SBR data where older decoders skip it. An LC-only decoder given implicit HE-AAC plays the core alone at half the sample rate: dull audio, and sometimes a wrong reported rate.

The bitstream: AudioSpecificConfig and ADTS

Raw AAC frames do not describe themselves, so the transport does. In MP4 it is the AudioSpecificConfig (ASC) in the esds box: 5 bits of object type (2 for LC), 4 of sample-rate index, 4 of channel configuration, then flag bits. AAC-LC at 44.1 kHz stereo is the two bytes 12 10; at 48 kHz it is 11 90. In streams without a container, such as raw .aac files and MPEG-TS, every frame instead carries a 7-byte ADTS header, 9 with CRC. It repeats the same information, using profile = object type minus one, and adds the frame length, so a parser can resynchronise anywhere.

This parser walks ADTS frames and builds an ASC. To get duration, multiply frames by 1,024 and divide by the header's rate; for HE-AAC that is usually the core rate, which still gives the right answer.

RATES = [96000, 88200, 64000, 48000, 44100, 32000, 24000,
         22050, 16000, 12000, 11025, 8000, 7350]

def adts_frames(buf: bytes):
    i = 0
    while i + 7 <= len(buf):
        b = buf[i:i + 7]
        if b[0] != 0xFF or (b[1] & 0xF6) != 0xF0:          # 12-bit sync + layer == 0
            raise ValueError(f"lost sync at byte {i}")
        crc = not (b[1] & 0x01)                             # protection_absent == 0 -> 2-byte CRC
        aot = (b[2] >> 6) + 1                               # profile field is object type minus 1
        rate = RATES[(b[2] >> 2) & 0x0F]
        chans = ((b[2] & 0x01) << 2) | (b[3] >> 6)
        length = ((b[3] & 0x03) << 11) | (b[4] << 3) | (b[5] >> 5)   # includes the header
        blocks = (b[6] & 0x03) + 1
        yield dict(offset=i, aot=aot, rate=rate, chans=chans, length=length, blocks=blocks, crc=crc)
        i += length

def audio_specific_config(aot, rate_index, chans):        # 2 bytes for AAC-LC
    return ((aot << 11) | (rate_index << 7) | (chans << 3)).to_bytes(2, "big")

assert audio_specific_config(2, 4, 2).hex() == "1210"     # LC, 44.1 kHz, stereo
assert audio_specific_config(2, 3, 2).hex() == "1190"     # LC, 48 kHz, stereo

Encoders and commands

The encoder matters more than the format. FFmpeg's built-in aac encoder is always available and produces AAC-LC. The Fraunhofer libfdk_aac supports HE-AAC, HE-AAC v2 and -vbr 1 to 5, but its licence is GPL-incompatible, so FFmpeg builds with it need --enable-nonfree and cannot be redistributed. On Apple platforms aac_at uses the system encoder. AAC is patent-licensed through a pool; check your own obligations.

# Built-in FFmpeg encoder: AAC-LC, constant target bitrate.
ffmpeg -i master.wav -c:a aac -b:a 160k -ar 48000 music_lc.m4a

# Fraunhofer FDK encoder (FFmpeg must be built with --enable-nonfree; the binary is not redistributable).
ffmpeg -i master.wav -c:a libfdk_aac -profile:a aac_he    -b:a 64k music_he.m4a
ffmpeg -i master.wav -c:a libfdk_aac -profile:a aac_he_v2 -b:a 32k music_hev2.m4a
ffmpeg -i master.wav -c:a libfdk_aac -vbr 4                         music_vbr.m4a

# Rewrap a raw ADTS stream into MP4 without re-encoding: the 7-byte headers become one ASC.
ffmpeg -i stream.aac -c copy -bsf:a aac_adtstoasc stream.m4a

# Verify what you actually produced, not what you asked for.
ffprobe -v error -show_entries stream=codec_name,profile,sample_rate,channels,bit_rate \
        -of default=nw=1 music_he.m4a

Failure modes

SymptomCauseFix
Noise smear before drums or plucksPre-echo: block switching or TNS not triggered, or bitrate too lowRaise bitrate, try another encoder, check transient-heavy test clips
Dull audio at half the sample rateLC-only decoder ignoring implicitly signalled SBRSignal SBR explicitly, or ship an LC rendition for old devices
Silence or wrong pitch after remuxASC does not match the frames, or ADTS headers left inside MP4Use aac_adtstoasc; compare ffprobe output with the source
Clipping on playback of loud mastersDecoded peaks exceed the source's because quantisation noise adds to peaksLeave about 1 dB of true-peak headroom before encoding
Collapsed or phasey stereo at low ratesParametric stereo on wide or out-of-phase mixesUse HE-AAC v1 or a higher-rate LC tier for that content
Audio offset from videoEncoder priming not trimmed by the playerWrite edit lists; see the comparison article on priming

Two adjacent problems surface in AAC pipelines: sample-rate mismatches, covered in audio resampling, and inconsistent loudness, covered in loudness normalisation. Normalise and set headroom before the encoder, never after.

Trade-offs against the alternatives

AAC's strength is compatibility: hardware decoders nearly everywhere and native support in HLS and MP4. Its weaknesses are frame latency (low-delay AAC-LD and AAC-ELD exist but are less supported) and efficiency at low bitrates. For real-time voice, Opus is usually better; for the widest device reach, AAC-LC with an HE-AAC low tier remains the default.

What to do next

  1. Run the MDCT snippet, then quantise the coefficients crudely and listen to the noise you introduced.
  2. Run ffprobe on your production outputs and confirm profile, sample rate and channels match what you intended.
  3. Build a 30-second torture clip of castanets, harpsichord, applause and wide stereo, and compare encoders and bitrates on it by ear.
  4. Decide your ladder explicitly: LC rates for music, an HE-AAC or HE-AAC v2 tier only where bandwidth demands it.
  5. Set loudness and about 1 dB of true-peak headroom before encoding.
  6. Add an ADTS or MP4 sanity check to CI: frame count, duration and ASC versus expectations.
Key takeaway: AAC transforms each 1,024-sample frame with an overlapping MDCT, uses short windows and TNS to keep noise near transients, and sets a scalefactor per band so quantisation noise stays under what the ear can hear. Huffman coding, M/S stereo, noise substitution and a bit reservoir pack the result; HE-AAC reaches low bitrates by rebuilding the high band with SBR and the stereo image with parametric stereo. Most production bugs are configuration: mismatched ASC, implicit SBR on old decoders, missing headroom, untrimmed priming.