H.264, also published as MPEG-4 Part 10 Advanced Video Coding (AVC), was standardised in 2003 and is still the codec that plays everywhere: every browser, phone, smart TV, set-top box and hardware decoder made in the last fifteen years. Newer codecs compress better, and H.264 vs H.265 vs AV1 covers when to switch. This article is about H.264 itself: what the encoder does to each block of pixels, what the bytes on disk look like, what profiles and levels promise a decoder, and which settings matter when you run x264 or a hardware encoder in production.
If you have not met the general hybrid block-based design yet (prediction, transform, quantization, entropy coding), read the video codecs introduction first. Everything here is the H.264-specific version of that pipeline.
The macroblock pipeline
H.264 divides each frame into macroblocks of 16x16 luma samples plus the matching chroma samples, two 8x8 blocks in the usual 4:2:0 sampling. Each macroblock goes through the same steps: predict it from pixels the decoder already has, subtract the prediction to get a residual, transform the residual, quantize the coefficients, and entropy-code the result. The encoder then runs the decoder's inverse steps on its own output, because the next prediction must be made from the reconstructed picture the decoder will see, not from the pristine source. If the two drifted apart, errors would compound frame after frame.
Prediction: intra and inter
Intra prediction builds a block from already-decoded neighbours in the same frame. Luma can be predicted as sixteen 4x4 blocks with nine modes each (eight directions plus DC), or as one 16x16 block with four modes (vertical, horizontal, DC and plane). The High profile adds 8x8 intra blocks. Smooth areas such as sky favour 16x16; textured areas favour 4x4. Chroma has its own four modes.
Inter prediction copies a block from a previously decoded reference picture, displaced by a motion vector. A macroblock can be split into 16x16, 16x8, 8x16 or 8x8 partitions, and each 8x8 can split again down to 4x4, so a moving edge can get its own vectors without paying for them across the whole block. Luma vectors have quarter-sample precision: half-sample positions are interpolated with a six-tap filter and quarter positions are averaged from neighbouring full and half samples. Up to 16 reference frames can be held in the decoded picture buffer (DPB), and weighted prediction scales and offsets a reference, which helps on fades.
B slices predict from two references and average them. Unlike MPEG-2, H.264 lets a B picture itself be a reference, which enables the hierarchical or pyramid B structures that x264 uses by default. Because B pictures can refer forward in display order, decode order differs from display order; the picture order count (POC) in each slice header restores the display sequence. GOP structure covers how I, P and B pictures are arranged and what that does to seeking and latency.
Transform, quantization and entropy coding
The residual is transformed with a 4x4 integer approximation of the discrete cosine transform. Being exact integer arithmetic, every compliant decoder produces identical output, which removes the inverse-transform drift that troubled earlier standards. High profile adds an 8x8 integer transform that the encoder can choose per macroblock, which helps on smooth, high-resolution content.
Coefficients are then quantized under a quantization parameter, QP, from 0 to 51 for 8-bit video. The quantizer step size doubles every 6 QP, so +6 QP roughly halves the bits for the residual and visibly coarsens detail. Rate control is mostly the art of choosing QP per frame and per macroblock.
Two entropy coders exist. CAVLC (context-adaptive variable-length coding) uses lookup tables chosen from neighbouring coefficient counts; it is cheap and is the only option in Baseline. CABAC (context-adaptive binary arithmetic coding) turns each syntax element into binary decisions and codes them with adaptive probability models; it typically saves a meaningful share of bits, often quoted around 10 percent or more depending on content, at the cost of a strictly serial decode within a slice. Finally an in-loop deblocking filter smooths block edges according to their strength and the local QP, and the filtered picture is what goes into the reference buffer.
The bitstream: NAL units and parameter sets
An H.264 stream is a sequence of NAL (network abstraction layer) units. Each starts with a one-byte header: a forbidden zero bit, a two-bit nal_ref_idc saying whether the content is used for reference, and a five-bit type. The types you meet daily are 1 (slice of a non-IDR picture), 5 (slice of an IDR picture), 6 (SEI, supplemental information such as captions or timing), 7 (SPS, sequence parameter set), 8 (PPS, picture parameter set) and 9 (access unit delimiter).
The SPS carries the profile, level, resolution, reference-frame count and VUI data such as colour description; the PPS carries per-picture choices such as which entropy coder is in use. A decoder cannot decode a slice without the SPS and PPS it references. An IDR picture is a clean restart: no later picture may reference anything before it, so it is the safe place to start decoding or to cut a segment.
There are two framings. Annex B, used in MPEG-TS, raw .h264 files and many RTP pipelines, separates NAL units with start codes (00 00 01 or 00 00 00 01). To stop payload bytes from imitating a start code, the encoder inserts an emulation prevention byte, 03, wherever two zero bytes inside a NAL unit would otherwise be followed by a byte from 00 to 03; a decoder removes every 03 that follows two zeros. AVCC, used in MP4, prefixes each NAL unit with its length instead, and stores the SPS and PPS once in the avcC box of the sample entry. With the avc1 sample entry the parameter sets live in avcC; avc3 allows them in-band as well, which helps when they change mid-stream. Video containers explains the surrounding MP4 and TS structures. This parser splits an Annex B stream and reads each header:
NAL_TYPES = {1: "non-IDR slice", 5: "IDR slice", 6: "SEI",
7: "SPS", 8: "PPS", 9: "AUD"}
def split_annexb(data: bytes):
"""Yield NAL unit payloads from an Annex B byte stream."""
i, n, starts = 0, len(data), []
while i + 3 <= n:
if data[i] == 0 and data[i + 1] == 0 and data[i + 2] == 1:
starts.append(i + 3)
i += 3
else:
i += 1
for k, s in enumerate(starts):
end = starts[k + 1] - 3 if k + 1 < len(starts) else n
nal = data[s:end]
while nal.endswith(b"\x00"): # leading zero of the next 4-byte start code
nal = nal[:-1]
yield nal
def describe(nal: bytes):
b = nal[0]
forbidden = b >> 7 # must be 0
ref_idc = (b >> 5) & 0x3 # 0 = not used as a reference
nal_type = b & 0x1F
return forbidden, ref_idc, nal_type, NAL_TYPES.get(nal_type, "other")
def avc_codec_string(sps: bytes) -> str:
# the three bytes after the SPS header: profile_idc, constraint flags, level_idc
return "avc1.%02X%02X%02X" % (sps[1], sps[2], sps[3])
# 00 00 00 01 | 67 64 00 1f ... -> SPS, ref_idc 3, codec string avc1.64001FReal code must also strip emulation prevention bytes before parsing the SPS bit fields beyond those first bytes, and should handle the four-byte start code that a stream may begin with.
Profiles, levels and codec strings
A profile lists the tools a decoder must support; a level caps how much work it must do. Constrained Baseline (no B slices, no CABAC, no interlace) is the common denominator for real-time and WebRTC. Main adds B slices, CABAC and interlace. High adds the 8x8 transform, 8x8 intra and custom quantization matrices, and is the normal choice for streaming. High 10 and High 4:2:2 extend bit depth and chroma sampling, but hardware and browser decode support for them is far thinner than for 8-bit 4:2:0 High, so treat them as production or contribution formats, not distribution formats.
| Level | Max macroblocks/s | Max frame (macroblocks) | Max bitrate Main / High (kbit/s) |
|---|---|---|---|
| 3 | 40,500 | 1,620 | 10,000 / 12,500 |
| 3.1 | 108,000 | 3,600 | 14,000 / 17,500 |
| 4 | 245,760 | 8,192 | 20,000 / 25,000 |
| 4.1 | 245,760 | 8,192 | 50,000 / 62,500 |
| 4.2 | 522,240 | 8,704 | 50,000 / 62,500 |
| 5.1 | 983,040 | 36,864 | 240,000 / 300,000 |
Worked example. 1280x720 is 80 x 45 = 3,600 macroblocks; at 30 fps that is 108,000 macroblocks per second, exactly level 3.1. 1920x1080 is coded as 1920x1088, which is 120 x 68 = 8,160 macroblocks; at 30 fps that is 244,800 per second, inside level 4's 245,760. At 60 fps it is 489,600, so 1080p60 needs level 4.2. Level also bounds the DPB size, so asking for many reference frames at high resolution can push you up a level.
Players and manifests describe a stream with a codec string built from the SPS: avc1 followed by profile_idc, the constraint flag byte and level_idc in hex. avc1.64001F is High (100 = 0x64) at level 3.1 (31 = 0x1F); avc1.640028 is High at level 4; avc1.42E01E is Constrained Baseline at level 3. Generate these from the actual SPS, as the code above does, rather than typing them: a manifest that claims a lower level than the stream has will be accepted by some devices and rejected by others.
Encoding with x264 and hardware encoders
x264 is the reference-quality software encoder. Its presets, from ultrafast to placebo, trade CPU time for compression by changing motion search range, the number of reference frames, subpartition decisions and rate-distortion optimisation; slow or medium is the usual production compromise. CRF (constant rate factor) targets constant perceptual quality and lets bitrate float, which is right for files. For streaming, cap it: VBV settings (maxrate and bufsize) bound the peak so a player's buffer model holds. For ABR ladders, fix the GOP length and disable scene-cut IDRs so every rendition has keyframes at the same timestamps and segments align.
# VOD mezzanine-to-web, quality-targeted (CRF), 8-bit 4:2:0 High profile
ffmpeg -i master.mov -c:v libx264 -preset slow -crf 20 \
-profile:v high -level:v 4.0 -pix_fmt yuv420p \
-c:a aac -b:a 128k -movflags +faststart web_1080p.mp4
# One ABR ladder rung: capped bitrate, fixed 2 s GOP at 24 fps, no scene-cut IDRs
ffmpeg -i master.mov -c:v libx264 -preset slow -b:v 4500k -maxrate 4800k -bufsize 9000k \
-g 48 -keyint_min 48 -sc_threshold 0 -profile:v high -level:v 4.0 -pix_fmt yuv420p \
-an rung_1080p.mp4
# Inspect what you produced: frame types and keyframes
ffprobe -v error -select_streams v:0 -show_entries frame=pict_type,key_frame \
-of csv rung_1080p.mp4 | head
# MP4 (length-prefixed) to a raw Annex B elementary stream without re-encoding
ffmpeg -i rung_1080p.mp4 -c:v copy -bsf:v h264_mp4toannexb -an -f h264 rung.h264Hardware encoders (NVENC, Quick Sync, VideoToolbox and the encoders in phones) run in real time at low power but usually need more bitrate than x264 at a slow preset for the same quality. Use them for live, real-time and high-volume cases where the bitrate premium costs less than the CPU. Compare on your own content with a perceptual metric, not by eye on one clip.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Green or grey frames at stream start | Decoding began without SPS/PPS or before an IDR | Repeat parameter sets in-band before each IDR; start at an IDR |
| Plays in one player, fails on a TV | Level or profile in the codec string is wrong, or stream exceeds the level | Derive codec strings from the SPS; set -level and check DPB size |
| Raw stream unreadable after extracting from MP4 | Length-prefixed AVCC written as if it were Annex B | Use the h264_mp4toannexb bitstream filter |
| ABR switches stall or glitch | Keyframes not aligned across renditions | Fixed GOP, scene-cut off, same frame rate on every rung |
| Washed-out or shifted colours | Missing or wrong VUI colour flags, or 4:2:2/10-bit sent to an 8-bit decoder | Set colour primaries and range explicitly; deliver yuv420p 8-bit |
| Blocky dark scenes and banding | QP too high in flat regions | Lower CRF, use x264's adaptive quantisation, or a film or grain tune |
| Decode stutters on low-end devices | CABAC at high bitrate or too high a level | Lower the level or bitrate; for very weak decoders, use Main or Baseline |
Trade-offs
H.264's value is reach and cheap decoding. Compared with HEVC and AV1 it needs noticeably more bitrate for the same quality, especially at 1080p and above, because its largest block is 16x16 and its tool set is older. The usual production answer is a mixed ladder: H.264 for every device and as the fallback, a newer codec for the clients that decode it in hardware, with the extra encode and storage cost justified only where that traffic is large. Within H.264, the dials trade the same three things: CPU time (preset), bits (CRF or bitrate) and decoder burden (profile and level).