A video codec is a pair of programs, an encoder and a decoder, that agree on a bitstream format. The standard defines only the decoder: what each bit means and exactly which pixels must come out. Everything the encoder does to choose those bits is left open, which is why two H.264 encoders can produce files of very different quality at the same size.
This page builds the whole idea from the ground up. It starts with how big raw video really is, walks through the hybrid block-based design that every mainstream codec since the 1990s shares, works an 8x8 transform example in code, and then turns to the settings and failure modes you meet in practice.
Why compression is not optional
Start with arithmetic. A 1080p frame has 1920 x 1080 = 2,073,600 luma samples. With 4:2:0 chroma, explained below, each of the two colour planes is 960 x 540, adding 1,036,800 samples, so a frame holds about 3.11 million samples. At 8 bits each that is 24.9 megabits per frame, and at 30 frames per second about 746 Mbit/s. A 4K stream at 60 frames per second with 10-bit samples needs roughly 7.5 Gbit/s.
A good 1080p30 stream is delivered at a few megabits per second. At 5 Mbit/s the codec is throwing away about 99.3 percent of the raw bits, a ratio near 150:1, while viewers judge the result as close to the original. That is only possible because video is full of redundancy (neighbouring pixels and consecutive frames look alike) and because human vision ignores much of the detail that remains. Codecs exploit both: prediction removes redundancy, and quantization discards what the eye is least likely to miss.
Colour: YCbCr and chroma subsampling
Cameras capture red, green and blue, but codecs almost never compress RGB. They convert to YCbCr: one luma channel Y carrying brightness and two chroma channels carrying colour difference. Human vision resolves brightness detail far better than colour detail, so the chroma planes can be stored at lower resolution. In 4:2:0, the default for distribution, each chroma plane has half the width and half the height of luma, which removes half of all samples before any real compression starts. 4:2:2 halves only the width and is common in production; 4:4:4 keeps everything and is used for screen content, where thin coloured text smears badly at 4:2:0.
The conversion matrix matters. HD content normally uses BT.709 coefficients and UHD HDR content uses BT.2020; older SD content used BT.601. The bitstream carries tags for this, the transfer function and limited (16-235) versus full range. When those tags are missing or wrong, the decoder still produces a picture, just with shifted colours or crushed blacks. Bit depth is the other choice: 10-bit encodes reduce banding in skies and gradients even for SDR content, at little cost in bitrate.
The hybrid block-based architecture
Every mainstream codec, from H.261 through H.264, HEVC, VP9, AV1 and VVC, uses the same skeleton. The frame is split into blocks. For each block the encoder forms a prediction from pixels the decoder already has, subtracts it to get a residual, transforms the residual into frequency coefficients, quantizes those coefficients, and entropy-codes the result together with the side information: block sizes, prediction modes and motion vectors.
The key design rule is in the lower loop of the diagram. The encoder must predict from what the decoder will reconstruct, not from the pristine source, because the decoder never sees the source. So the encoder runs inverse quantization, inverse transform and the loop filters itself, and stores the result in a decoded picture buffer. If encoder and decoder ever disagree by a single sample, errors compound frame after frame, which is called drift. This is why decoders are specified to bit-exact integer arithmetic.
Block sizes grew with each generation as resolutions grew. H.264 works on 16x16 macroblocks with smaller partitions; HEVC uses coding tree units of up to 64x64 split by a quadtree; AV1 uses superblocks of up to 128x128 with richer partition shapes; VVC adds binary and ternary splits on top of the quadtree.
Prediction: intra and inter
Intra prediction uses only the current frame. The decoder already holds the reconstructed row above and column to the left of a block, so the encoder can say, for example, extend those pixels diagonally down-right, or average them (DC mode), or fit a smooth plane. HEVC has 35 intra modes and VVC 67. Only the mode number and the residual are coded.
Inter prediction uses earlier decoded frames. For each block the encoder searches a reference frame for a similar region and codes a motion vector pointing at it. Vectors have sub-pixel precision, a quarter pixel in H.264 and HEVC, with interpolation filters producing the in-between samples. Motion search is the most expensive part of encoding.
Frames are labelled by how they may predict. An I frame uses only intra prediction and can be decoded on its own, so seeking and stream joins start there. P frames reference earlier frames; B frames may reference both directions, which forces the encoder to send some frames out of display order and adds delay. The arrangement of these frame types is the group of pictures; GOP structure covers how its length and shape affect seeking, latency and efficiency.
Transform and quantization, worked
The residual after good prediction is small and smooth. A two-dimensional discrete cosine transform rewrites a block as weights on 64 basis patterns, from flat (the DC coefficient) to fine checkerboards. For natural images almost all the energy lands in a few low-frequency coefficients. Codecs use integer approximations of the DCT so encoder and decoder compute identical results, plus a sine-based variant for some small intra blocks.
Quantization divides each coefficient by a step size and rounds. It is the only stage that throws information away; everything else is reversible. The script below makes that concrete on one smooth 8x8 residual block, using a zigzag scan that orders coefficients from low to high frequency.
import numpy as np
N = 8
k = np.arange(N)
# Orthonormal DCT-II basis; codecs use scaled integer approximations of this.
C = np.sqrt(2 / N) * np.cos(np.pi * (2 * k[None, :] + 1) * k[:, None] / (2 * N))
C[0] /= np.sqrt(2)
def dct2(block): return C @ block @ C.T
def idct2(coef): return C.T @ coef @ C
# A smooth 8x8 luma residual: a gentle ramp plus a little noise.
rng = np.random.default_rng(0)
x = np.add.outer(np.arange(8), np.arange(8)) * 3.0 + rng.normal(0, 1.5, (8, 8))
x -= x.mean()
coef = dct2(x)
for step in (4, 16, 40):
q = np.round(coef / step) # the only lossy step
rec = idct2(q * step)
zz = sorted(((i, j) for i in range(8) for j in range(8)),
key=lambda p: (p[0] + p[1], p[1] if (p[0] + p[1]) % 2 else p[0]))
scan = [int(q[i, j]) for i, j in zz]
last = max((n for n, v in enumerate(scan) if v), default=-1)
mse = np.mean((x - rec) ** 2)
print(f"step={step:3d} nonzero={np.count_nonzero(q):2d} "
f"last_sig={last:2d} PSNR={10*np.log10(255**2/mse):5.1f} dB")Run it and you see the trade directly. Even at a small step only a handful of the 64 coefficients survive. As the step grows, the count and the last significant position collapse toward the DC corner, then PSNR keeps falling gently while the count stays put, because the zeroed coefficients carried little energy. That long run of trailing zeros is what the entropy coder turns into very few bits.
The quantization parameter, QP, is the encoder's knob for step size. In H.264 and HEVC the step roughly doubles every 6 QP, so each +6 QP roughly halves bitrate on typical content while costing visible quality. AV1 uses a different index range, so QP numbers do not transfer between codecs. Real encoders also use rate-distortion optimized quantization, rounding a coefficient down when the bits saved outweigh the distortion added.
Entropy coding and loop filters
After quantization the data is a stream of symbols: modes, motion vector differences, coefficient positions and levels. Entropy coding assigns short codes to likely symbols. H.264 offers CAVLC, a set of variable-length code tables, and CABAC, a context-adaptive binary arithmetic coder that keeps running probability estimates per context and typically saves around 10 percent or more over CAVLC. HEVC and VVC use only CABAC; AV1 uses an adaptive multi-symbol arithmetic coder.
Block coding leaves seams. In-loop filters clean them up before a frame becomes a reference, so later frames predict from a better picture. All of these codecs have a deblocking filter. HEVC adds sample adaptive offset; AV1 adds a directional enhancement filter and loop restoration; VVC adds an adaptive loop filter. Because they sit inside the loop, they are normative: every decoder must apply them identically.
Rate control: turning quality into a bitrate
The encoder must decide QP for every frame and block. Constant QP fixes the step size and is useful only for experiments. Constant rate factor, CRF in x264, x265 and SVT-AV1, targets roughly constant perceived quality and lets bitrate float, spending more on complex scenes; it is the right default for files. Streaming needs a bitrate ceiling, so you combine a quality target with a maximum rate and a buffer size that model the decoder's buffer, the video buffering verifier or VBV.
# Quality-targeted file encode (VOD): constant rate factor, no bitrate target.
ffmpeg -i in.mov -c:v libx264 -preset slow -crf 20 -pix_fmt yuv420p \
-c:a copy out_h264.mp4
# Streaming rung: capped VBR so the client buffer model holds.
ffmpeg -i in.mov -c:v libx264 -preset medium -crf 23 \
-maxrate 4500k -bufsize 9000k -g 120 -keyint_min 120 -sc_threshold 0 \
-pix_fmt yuv420p out_rung.mp4
# Same idea with AV1 via SVT-AV1 (lower preset number = slower, better).
ffmpeg -i in.mov -c:v libsvtav1 -preset 6 -crf 32 -g 120 \
-pix_fmt yuv420p10le out_av1.mkvIn the streaming example, a fixed keyframe interval with scene-cut insertion disabled keeps segment boundaries aligned across bitrate rungs, which adaptive streaming needs.
The codec family at a glance
| Codec | Standardised | Largest block | Where it fits |
|---|---|---|---|
| H.264 / AVC | 2003 | 16x16 macroblock | Universal decode support; the safe fallback |
| H.265 / HEVC | 2013 | 64x64 CTU | Broadcast, Apple devices, HDR; patent pools complicate licensing |
| VP9 | 2013 | 64x64 superblock | Web video, widely decoded in browsers and Android |
| AV1 | 2018 | 128x128 superblock | Royalty-free aim, strong efficiency, slow encodes; growing hardware decode |
| H.266 / VVC | 2020 | 128x128 CTU | Highest efficiency of the group; limited device support so far |
Each generation has targeted roughly half the bitrate of its predecessor at equal quality, but real savings depend on content, resolution and encoder maturity, and encoding cost rises steeply. Decode support is the binding constraint for distribution: a codec only helps for viewers whose devices decode it efficiently, ideally in hardware. H.264 vs H.265 vs AV1 works through measured efficiency, device support and a cost model.
Measuring quality
PSNR measures mean squared error and is cheap but correlates weakly with perception across content. SSIM looks at local structure. VMAF combines several features with a model trained on human ratings and is the usual choice for streaming ladders; VMAF covers its models and pitfalls. To compare encoders or settings fairly, encode at several rates, plot quality against bitrate, and compute the Bjontegaard delta rate, the average bitrate difference at equal quality between the two curves.
Content difference dominates everything else. A cartoon and a grainy film need wildly different bitrates for the same score, which is the idea behind per-title encoding.
Artifacts and failure modes
- Blocking. Visible block edges at low bitrate, worst in dark flat areas. Raise bitrate, lower CRF, or check that deblocking was not disabled for speed.
- Banding. Steps in smooth gradients. Encode at 10-bit even for 8-bit sources, and avoid aggressive denoising that removes the dither that hid the steps.
- Ringing and mosquito noise. Halos around sharp edges and text, from quantized high frequencies. Screen content may need 4:4:4 or a codec's screen-content tools.
- Wrong colours or washed-out blacks. Missing or mismatched colour tags, or full-range samples flagged as limited range. Set and check primaries, transfer, matrix and range explicitly.
- Slow seeking or broken stream joins. Keyframes too far apart, or open GOPs where the player expects closed ones.
- Stalls on capped streams. Bitrate peaks beyond the declared maximum because the buffer size was too large for the network. Size the buffer against the delivery path, not the encoder's convenience.
- Unexpected latency. B-frames and long lookahead add frames of delay. Real-time paths usually disable B-frames and lookahead.
Trade-offs to weigh
Every codec decision trades compression efficiency against encode cost, decode compatibility and latency. Slower presets cost encode time, which matters for live but barely for a file watched a million times. Choose per use: a broad H.264 ladder for reach, a more efficient codec where devices support it and viewing hours justify the encode cost, and tuned low-delay settings for real-time paths. Video encoder architecture explains where the encode time actually goes.
What to do next
- Run the DCT script and change the residual to a sharp edge; watch how many coefficients survive at each step.
- Take one representative clip, encode it with x264 at CRF 18, 23 and 28, and record size, VMAF and encode time.
- Repeat at the same settings with a second codec you can decode on your target devices and compute BD-rate between the two curves.
- Run ffprobe on your outputs and confirm profile, level, pixel format and colour tags match what your players expect.
- Fix the keyframe interval to your segment duration and disable scene-cut insertion for adaptive streaming outputs.
- Encode a gradient-heavy clip at 8-bit and 10-bit and compare banding on a real display.