H.264, also published as MPEG-4 Part 10 Advanced Video Coding (AVC), was standardised in 2003 and is still the codec that plays everywhere: every browser, phone, smart TV, set-top box and hardware decoder made in the last fifteen years. Newer codecs compress better, and H.264 vs H.265 vs AV1 covers when to switch. This article is about H.264 itself: what the encoder does to each block of pixels, what the bytes on disk look like, what profiles and levels promise a decoder, and which settings matter when you run x264 or a hardware encoder in production.

If you have not met the general hybrid block-based design yet (prediction, transform, quantization, entropy coding), read the video codecs introduction first. Everything here is the H.264-specific version of that pipeline.

Advertisement

The macroblock pipeline

H.264 divides each frame into macroblocks of 16x16 luma samples plus the matching chroma samples, two 8x8 blocks in the usual 4:2:0 sampling. Each macroblock goes through the same steps: predict it from pixels the decoder already has, subtract the prediction to get a residual, transform the residual, quantize the coefficients, and entropy-code the result. The encoder then runs the decoder's inverse steps on its own output, because the next prediction must be made from the reconstructed picture the decoder will see, not from the pristine source. If the two drifted apart, errors would compound frame after frame.

One macroblock through the H.264 encoder (the decoder runs the shaded loop too)Macroblock16x16 luma, 2x 8x8 chromaPredictionintra modes or motionResidualsource minus predictionTransform4x4 integer (8x8 High)QuantizeQP 0-51, step x2 per +6Entropy codeCAVLC or CABACNAL unitsslice data + SPS/PPSInverse quant + transformsame integers as decoderReconstructprediction + residualDeblocking filterin-loop, edge adaptiveReference framesup to 16 in the DPBmotion searchThe encoder decodes its own output so its references match the decoder's bit for bit.
The forward path (top and right) produces bits; the dashed loop is the reconstruction the encoder shares with every decoder.

Prediction: intra and inter

Intra prediction builds a block from already-decoded neighbours in the same frame. Luma can be predicted as sixteen 4x4 blocks with nine modes each (eight directions plus DC), or as one 16x16 block with four modes (vertical, horizontal, DC and plane). The High profile adds 8x8 intra blocks. Smooth areas such as sky favour 16x16; textured areas favour 4x4. Chroma has its own four modes.

Inter prediction copies a block from a previously decoded reference picture, displaced by a motion vector. A macroblock can be split into 16x16, 16x8, 8x16 or 8x8 partitions, and each 8x8 can split again down to 4x4, so a moving edge can get its own vectors without paying for them across the whole block. Luma vectors have quarter-sample precision: half-sample positions are interpolated with a six-tap filter and quarter positions are averaged from neighbouring full and half samples. Up to 16 reference frames can be held in the decoded picture buffer (DPB), and weighted prediction scales and offsets a reference, which helps on fades.

B slices predict from two references and average them. Unlike MPEG-2, H.264 lets a B picture itself be a reference, which enables the hierarchical or pyramid B structures that x264 uses by default. Because B pictures can refer forward in display order, decode order differs from display order; the picture order count (POC) in each slice header restores the display sequence. GOP structure covers how I, P and B pictures are arranged and what that does to seeking and latency.

Advertisement

Transform, quantization and entropy coding

The residual is transformed with a 4x4 integer approximation of the discrete cosine transform. Being exact integer arithmetic, every compliant decoder produces identical output, which removes the inverse-transform drift that troubled earlier standards. High profile adds an 8x8 integer transform that the encoder can choose per macroblock, which helps on smooth, high-resolution content.

Coefficients are then quantized under a quantization parameter, QP, from 0 to 51 for 8-bit video. The quantizer step size doubles every 6 QP, so +6 QP roughly halves the bits for the residual and visibly coarsens detail. Rate control is mostly the art of choosing QP per frame and per macroblock.

Two entropy coders exist. CAVLC (context-adaptive variable-length coding) uses lookup tables chosen from neighbouring coefficient counts; it is cheap and is the only option in Baseline. CABAC (context-adaptive binary arithmetic coding) turns each syntax element into binary decisions and codes them with adaptive probability models; it typically saves a meaningful share of bits, often quoted around 10 percent or more depending on content, at the cost of a strictly serial decode within a slice. Finally an in-loop deblocking filter smooths block edges according to their strength and the local QP, and the filtered picture is what goes into the reference buffer.

The bitstream: NAL units and parameter sets

An H.264 stream is a sequence of NAL (network abstraction layer) units. Each starts with a one-byte header: a forbidden zero bit, a two-bit nal_ref_idc saying whether the content is used for reference, and a five-bit type. The types you meet daily are 1 (slice of a non-IDR picture), 5 (slice of an IDR picture), 6 (SEI, supplemental information such as captions or timing), 7 (SPS, sequence parameter set), 8 (PPS, picture parameter set) and 9 (access unit delimiter).

The SPS carries the profile, level, resolution, reference-frame count and VUI data such as colour description; the PPS carries per-picture choices such as which entropy coder is in use. A decoder cannot decode a slice without the SPS and PPS it references. An IDR picture is a clean restart: no later picture may reference anything before it, so it is the safe place to start decoding or to cut a segment.

There are two framings. Annex B, used in MPEG-TS, raw .h264 files and many RTP pipelines, separates NAL units with start codes (00 00 01 or 00 00 00 01). To stop payload bytes from imitating a start code, the encoder inserts an emulation prevention byte, 03, wherever two zero bytes inside a NAL unit would otherwise be followed by a byte from 00 to 03; a decoder removes every 03 that follows two zeros. AVCC, used in MP4, prefixes each NAL unit with its length instead, and stores the SPS and PPS once in the avcC box of the sample entry. With the avc1 sample entry the parameter sets live in avcC; avc3 allows them in-band as well, which helps when they change mid-stream. Video containers explains the surrounding MP4 and TS structures. This parser splits an Annex B stream and reads each header:

NAL_TYPES = {1: "non-IDR slice", 5: "IDR slice", 6: "SEI",
             7: "SPS", 8: "PPS", 9: "AUD"}

def split_annexb(data: bytes):
    """Yield NAL unit payloads from an Annex B byte stream."""
    i, n, starts = 0, len(data), []
    while i + 3 <= n:
        if data[i] == 0 and data[i + 1] == 0 and data[i + 2] == 1:
            starts.append(i + 3)
            i += 3
        else:
            i += 1
    for k, s in enumerate(starts):
        end = starts[k + 1] - 3 if k + 1 < len(starts) else n
        nal = data[s:end]
        while nal.endswith(b"\x00"):      # leading zero of the next 4-byte start code
            nal = nal[:-1]
        yield nal

def describe(nal: bytes):
    b = nal[0]
    forbidden = b >> 7                    # must be 0
    ref_idc = (b >> 5) & 0x3              # 0 = not used as a reference
    nal_type = b & 0x1F
    return forbidden, ref_idc, nal_type, NAL_TYPES.get(nal_type, "other")

def avc_codec_string(sps: bytes) -> str:
    # the three bytes after the SPS header: profile_idc, constraint flags, level_idc
    return "avc1.%02X%02X%02X" % (sps[1], sps[2], sps[3])

# 00 00 00 01 | 67 64 00 1f ...  ->  SPS, ref_idc 3, codec string avc1.64001F

Real code must also strip emulation prevention bytes before parsing the SPS bit fields beyond those first bytes, and should handle the four-byte start code that a stream may begin with.

Profiles, levels and codec strings

A profile lists the tools a decoder must support; a level caps how much work it must do. Constrained Baseline (no B slices, no CABAC, no interlace) is the common denominator for real-time and WebRTC. Main adds B slices, CABAC and interlace. High adds the 8x8 transform, 8x8 intra and custom quantization matrices, and is the normal choice for streaming. High 10 and High 4:2:2 extend bit depth and chroma sampling, but hardware and browser decode support for them is far thinner than for 8-bit 4:2:0 High, so treat them as production or contribution formats, not distribution formats.

LevelMax macroblocks/sMax frame (macroblocks)Max bitrate Main / High (kbit/s)
340,5001,62010,000 / 12,500
3.1108,0003,60014,000 / 17,500
4245,7608,19220,000 / 25,000
4.1245,7608,19250,000 / 62,500
4.2522,2408,70450,000 / 62,500
5.1983,04036,864240,000 / 300,000

Worked example. 1280x720 is 80 x 45 = 3,600 macroblocks; at 30 fps that is 108,000 macroblocks per second, exactly level 3.1. 1920x1080 is coded as 1920x1088, which is 120 x 68 = 8,160 macroblocks; at 30 fps that is 244,800 per second, inside level 4's 245,760. At 60 fps it is 489,600, so 1080p60 needs level 4.2. Level also bounds the DPB size, so asking for many reference frames at high resolution can push you up a level.

Players and manifests describe a stream with a codec string built from the SPS: avc1 followed by profile_idc, the constraint flag byte and level_idc in hex. avc1.64001F is High (100 = 0x64) at level 3.1 (31 = 0x1F); avc1.640028 is High at level 4; avc1.42E01E is Constrained Baseline at level 3. Generate these from the actual SPS, as the code above does, rather than typing them: a manifest that claims a lower level than the stream has will be accepted by some devices and rejected by others.

Encoding with x264 and hardware encoders

x264 is the reference-quality software encoder. Its presets, from ultrafast to placebo, trade CPU time for compression by changing motion search range, the number of reference frames, subpartition decisions and rate-distortion optimisation; slow or medium is the usual production compromise. CRF (constant rate factor) targets constant perceptual quality and lets bitrate float, which is right for files. For streaming, cap it: VBV settings (maxrate and bufsize) bound the peak so a player's buffer model holds. For ABR ladders, fix the GOP length and disable scene-cut IDRs so every rendition has keyframes at the same timestamps and segments align.

# VOD mezzanine-to-web, quality-targeted (CRF), 8-bit 4:2:0 High profile
ffmpeg -i master.mov -c:v libx264 -preset slow -crf 20 \
  -profile:v high -level:v 4.0 -pix_fmt yuv420p \
  -c:a aac -b:a 128k -movflags +faststart web_1080p.mp4

# One ABR ladder rung: capped bitrate, fixed 2 s GOP at 24 fps, no scene-cut IDRs
ffmpeg -i master.mov -c:v libx264 -preset slow -b:v 4500k -maxrate 4800k -bufsize 9000k \
  -g 48 -keyint_min 48 -sc_threshold 0 -profile:v high -level:v 4.0 -pix_fmt yuv420p \
  -an rung_1080p.mp4

# Inspect what you produced: frame types and keyframes
ffprobe -v error -select_streams v:0 -show_entries frame=pict_type,key_frame \
  -of csv rung_1080p.mp4 | head

# MP4 (length-prefixed) to a raw Annex B elementary stream without re-encoding
ffmpeg -i rung_1080p.mp4 -c:v copy -bsf:v h264_mp4toannexb -an -f h264 rung.h264

Hardware encoders (NVENC, Quick Sync, VideoToolbox and the encoders in phones) run in real time at low power but usually need more bitrate than x264 at a slow preset for the same quality. Use them for live, real-time and high-volume cases where the bitrate premium costs less than the CPU. Compare on your own content with a perceptual metric, not by eye on one clip.

Failure modes

SymptomLikely causeFix
Green or grey frames at stream startDecoding began without SPS/PPS or before an IDRRepeat parameter sets in-band before each IDR; start at an IDR
Plays in one player, fails on a TVLevel or profile in the codec string is wrong, or stream exceeds the levelDerive codec strings from the SPS; set -level and check DPB size
Raw stream unreadable after extracting from MP4Length-prefixed AVCC written as if it were Annex BUse the h264_mp4toannexb bitstream filter
ABR switches stall or glitchKeyframes not aligned across renditionsFixed GOP, scene-cut off, same frame rate on every rung
Washed-out or shifted coloursMissing or wrong VUI colour flags, or 4:2:2/10-bit sent to an 8-bit decoderSet colour primaries and range explicitly; deliver yuv420p 8-bit
Blocky dark scenes and bandingQP too high in flat regionsLower CRF, use x264's adaptive quantisation, or a film or grain tune
Decode stutters on low-end devicesCABAC at high bitrate or too high a levelLower the level or bitrate; for very weak decoders, use Main or Baseline

Trade-offs

H.264's value is reach and cheap decoding. Compared with HEVC and AV1 it needs noticeably more bitrate for the same quality, especially at 1080p and above, because its largest block is 16x16 and its tool set is older. The usual production answer is a mixed ladder: H.264 for every device and as the fallback, a newer codec for the clients that decode it in hardware, with the extra encode and storage cost justified only where that traffic is large. Within H.264, the dials trade the same three things: CPU time (preset), bits (CRF or bitrate) and decoder burden (profile and level).

Key takeaway: <p><strong>What to do next.</strong> H.264 is a 16x16 hybrid codec with quarter-sample motion, an exact integer transform, CABAC or CAVLC and an in-loop deblocking filter, wrapped in NAL units whose parameter sets every decoder needs. Most production failures are bitstream and signalling mistakes, not compression ones.</p><ol><li>Pick 8-bit 4:2:0 High profile for distribution; use Constrained Baseline only where a real-time or WebRTC endpoint demands it.</li><li>Compute the level from resolution and frame rate in macroblocks per second, and set it explicitly.</li><li>Generate codec strings from the SPS bytes, never by hand.</li><li>For ABR, fix the GOP, disable scene-cut IDRs and verify keyframe alignment with ffprobe.</li><li>Know which framing each stage expects (Annex B or AVCC) and convert with a bitstream filter, not a re-encode.</li><li>Choose the x264 preset by CPU budget, then tune CRF or bitrate against a perceptual metric on your own content.</li></ol>