Every mainstream video codec in use today, from H.264 through HEVC to AV1 and VVC, is a hybrid block-based codec. The standards define only the bitstream and how a decoder must interpret it. The encoder is free to make any decisions that produce a valid stream, and almost all of the difference between a good encoder and a poor one, often 30 to 50 percent in bitrate at the same quality, lives in those decisions.
This article opens up the encoder. It walks the pipeline stage by stage, explains the rate-distortion optimization that drives each choice, shows how lookahead and rate control spend bits across time, covers the threading schemes that make encoding fast and what they cost, and ends with concrete settings for x264, x265, SVT-AV1 and NVENC. How frames are grouped into I, P and B structures is covered in GOP structure; here we look at what happens inside each frame.
The core idea: predict, code the difference, and decode as you go
Video is highly redundant: neighbouring pixels are similar, and consecutive frames are mostly the same scene shifted slightly. A hybrid encoder exploits both. For each block it builds a prediction, either from already-coded pixels in the same frame (intra prediction) or from previously decoded frames shifted by a motion vector (inter prediction). It subtracts the prediction from the source, transforms the residual into frequency coefficients, quantizes them, which is where information is discarded, and entropy-codes the result with the prediction parameters.
The key architectural fact is that the encoder contains a full decoder. After quantizing a block, it immediately inverse-quantizes and inverse-transforms it, adds the prediction back, filters the result and stores it as the reference for future predictions. Predictions must be made from what the decoder will actually have, not from the pristine source; otherwise quantization errors would compound frame after frame into visible drift.
Stage by stage
Partitioning. The frame is cut into fixed coding units, 16 by 16 macroblocks in H.264, coding tree units up to 64 by 64 in HEVC, and superblocks of 64 or 128 in AV1, which are recursively split into smaller blocks. Large blocks are cheap to signal and suit flat areas; small blocks follow edges and motion. Choosing the split is one of the most expensive decisions because each option must be evaluated.
Intra prediction. A block is extrapolated from the reconstructed row above and column to the left, along one of many directions (9 luma modes for 4 by 4 blocks in H.264, 35 in HEVC, 56 directional plus smooth and other modes in AV1) or as a smooth planar or DC fill.
Motion estimation and compensation. For inter blocks the encoder searches reference frames for the best match. Exhaustive search is prohibitively expensive, so encoders use patterned searches such as diamond or hexagon steps around predicted vectors, then refine to half and quarter pixel positions (eighth pixel in AV1) using interpolation filters. Motion estimation is typically the largest share of encoding time, and search range and method are what speed presets mostly trade away.
Transform and quantization. The residual is transformed with integer approximations of the DCT (and in HEVC and AV1, other transform types and sizes). Quantization divides each coefficient by a step size set by the quantization parameter (QP). In H.264 and HEVC the step size doubles every 6 QP, so plus 6 QP roughly halves bitrate at a visible quality cost. Most high-frequency coefficients become zero, which is what makes the result compress.
In-loop filters. A deblocking filter smooths block edges; HEVC adds sample adaptive offset (SAO), and AV1 adds CDEF and loop restoration. Because they run inside the loop, they improve both the output and the references future frames predict from.
Entropy coding. Modes, motion vectors and coefficient levels are coded with context-adaptive arithmetic coding (CABAC in H.264 High profile and HEVC, a multi-symbol arithmetic coder in AV1), which adapts probabilities as it goes. It is inherently serial within a slice or tile, which matters for threading.
Rate-distortion optimization: how every decision is made
At each stage the encoder faces many valid options: split or not, intra direction, which reference and motion vector, which transform. RDO picks the one minimizing a Lagrangian cost, J = D + lambda x R, where D is distortion after reconstruction (typically the sum of squared errors), R is the bits the option costs, and lambda converts bits into distortion units. Lambda is derived from QP; the H.264 reference software used 0.85 x 2^((QP - 12) / 3), so as QP rises bits become relatively more expensive and the encoder favours cheaper modes.
def choose_mode(block, candidates, qp, lam_from_qp):
"""Rate-distortion optimized mode decision for one block (simplified)."""
lam = lam_from_qp(qp) # e.g. H.264 reference: 0.85 * 2 ** ((qp - 12) / 3)
best, best_cost = None, float("inf")
for mode in candidates: # intra directions, inter with each motion vector, skip ...
pred = mode.predict(block) # from RECONSTRUCTED neighbours / reference frames
coeffs = quantize(transform(block - pred), qp)
recon = pred + inverse_transform(dequantize(coeffs, qp))
dist = sse(block, recon) # distortion after quantization
bits = mode.header_bits() + entropy_bits(coeffs)
cost = dist + lam * bits
if cost < best_cost:
best, best_cost = (mode, coeffs, recon), cost
return bestA worked example: at a lambda of 50, suppose a block could be coded as skip (reuse the predicted motion, no residual) at 1 bit with distortion 4,000, or as inter with a new motion vector and residual at 60 bits with distortion 1,200. Their costs are 4,000 + 50 = 4,050 and 1,200 + 3,000 = 4,200, so skip wins despite its higher error. At lambda 30 (a lower QP) the inter mode costs 3,000 and wins. That single mechanism explains why low-bitrate encodes look smeared: the encoder is correctly choosing cheap modes.
Full RDO, which actually transforms, quantizes and entropy-estimates every candidate, is expensive, so encoders prune with cheap estimates such as the sum of absolute transformed differences and run full RDO only on the finalists. How aggressively they prune is the main difference between the ultrafast and placebo ends of a preset scale.
Lookahead and rate control: spending bits across time
Rate control decides the QP for each frame and often each block. The common modes are constant QP (fixed step, useful for testing only), constant rate factor or CRF (a quality target that lets bitrate float with content complexity), average bitrate, and constant bitrate for live links. A VBV model, a leaky bucket defined by a maximum rate and a buffer size, can cap any of them so the stream never exceeds what a player's buffer and the network can absorb.
Lookahead makes rate control smarter. The encoder analyses upcoming frames, often at reduced resolution, to find scene cuts, choose frame types and estimate how much each block is referenced by the future. x264's macroblock tree (mbtree) and x265's cutree lower QP for blocks that many future frames predict from, because bits spent there are reused, and raise it for blocks that are quickly overwritten. That is a large efficiency gain and the reason lookahead depth, 40 frames in x264's medium preset, is worth its latency and memory in VOD. For measuring the outcome, use perceptual metrics such as those in VMAF and quality metrics rather than PSNR alone.
Parallelism: how encoders use many cores
A single frame's decisions depend on its neighbours and its references, so encoders parallelise at several levels, each with a cost:
- Frame threads. Several frames encode at once, each waiting only for the rows of its references that it needs. It scales well but adds frames of latency and slightly weakens rate control, since the controller commits QPs before earlier frames finish.
- Wavefront parallel processing (WPP). In HEVC, rows of CTUs encode in parallel, each lagging the row above by two CTUs, with entropy contexts inherited at row starts. Little compression loss, modest parallelism.
- Slices and tiles. The frame is split into independently decodable regions. They parallelise both encoder and decoder, but prediction and entropy context do not cross boundaries, costing a few percent in efficiency. AV1 uses tiles heavily; x264's zerolatency tune uses sliced threads so one frame uses many cores with no frame-level delay.
- Lookahead and analysis threads. Lookahead runs ahead on its own threads. SVT-AV1 is built around this kind of pipelined parallelism across frames, segments and stages, which is why it scales to high core counts.
Hardware encoders
GPUs and many SoCs include fixed-function encoders such as NVIDIA NVENC, Intel Quick Sync and AMD's encoder blocks. They implement the same pipeline in silicon, with fixed motion search strategies, restricted RDO and limited lookahead. They are one or two orders of magnitude cheaper per stream in power and CPU, add little latency, and run beside CUDA work without consuming shader cores, which is why they dominate live, cloud gaming and real-time AI video pipelines. The price is compression efficiency: at the same quality a good software encoder on a slow preset typically needs noticeably fewer bits, and NVENC's own p1 to p7 presets trade speed against the effort it spends. For VOD at scale the bitrate saving usually pays for software encoding; for live and per-user streams it usually does not.
Real settings
The commands below show the main modes. Encoder defaults and option ranges differ, so check the version you run:
# VOD, quality-targeted: CRF with a VBV cap so peaks fit the delivery ladder rung
ffmpeg -i in.mov -c:v libx264 -preset slow -crf 21 \
-maxrate 6M -bufsize 12M -g 96 -keyint_min 96 -sc_threshold 0 -an out_x264.mp4
# HEVC, same idea; x265's CRF scale differs (its default is 28, x264's is 23)
ffmpeg -i in.mov -c:v libx265 -preset medium -crf 24 \
-x265-params "vbv-maxrate=6000:vbv-bufsize=12000:keyint=96:min-keyint=96:scenecut=0" out_x265.mp4
# AV1 with SVT-AV1: preset -1..13 (higher is faster, default 8), CRF 1..70 (default 35)
ffmpeg -i in.mov -c:v libsvtav1 -preset 6 -crf 32 -g 96 out_av1.mp4
# Live, low latency: no B-frames, no lookahead, sliced threads
ffmpeg -i rtmp_in -c:v libx264 -preset veryfast -tune zerolatency \
-b:v 4M -maxrate 4M -bufsize 2M -g 60 -f flv rtmp_out
# Hardware: NVENC presets p1 (fastest) .. p7 (best), tunes hq / ll / ull
ffmpeg -hwaccel cuda -i in.mov -c:v h264_nvenc -preset p5 -tune hq -rc vbr \
-b:v 5M -maxrate 6M -bufsize 12M out_nvenc.mp4A fixed GOP with scene-cut insertion disabled keeps keyframes aligned across ladder renditions so players can switch cleanly, which is what adaptive streaming needs; see adaptive bitrate streaming. For ladders tuned to each title, pair CRF encodes with the approach in per-title encoding.
Failure modes
- Buffer underflow on playback. CRF without a VBV cap produced bitrate spikes on complex scenes that exceed the rung's bandwidth. Always set maxrate and bufsize for delivery.
- Banding in gradients. Smooth skies and dark scenes quantized into steps. Encode in 10-bit even for 8-bit sources, use adaptive quantization, and lower QP in flat areas.
- Pulsing quality at keyframes. Keyframes get a different QP than surrounding frames; tune QP offsets or lookahead so the transition is smooth.
- Unexpected latency in live pipelines. Frame threading, B-frames and lookahead each add frames of delay. The zerolatency tune removes them at an efficiency cost.
- Misaligned keyframes across renditions. Scene-cut detection placed IDR frames differently per rung, breaking segment alignment.
- Grain destroyed or bloated. Film grain is noise the encoder either smears away or spends enormous bits on. AV1 can denoise and resynthesise grain as parameters (SVT-AV1 exposes a film-grain level), which is often the better choice.
Operational guidance and trade-offs
Pick the encoder by workload: software on slow presets for VOD where each encode serves many viewers, hardware for live and per-session streams, and SVT-AV1 or x265 where devices can decode them and the saving justifies the compute. Evaluate presets on your own content with a quality metric, bitrate and encode time together, because a two-step slower preset can cost double the compute for a few percent of bitrate. Keep settings in version control and re-baseline after encoder upgrades: defaults and even option meanings change between releases.
What to do next
- Pick five representative clips, including a dark scene, fast motion and a grainy source.
- Encode each with two or three presets at a fixed CRF and plot bitrate, VMAF and encode time.
- Choose a preset where the next slower one saves less than your compute cost is worth.
- Add VBV maxrate and bufsize matched to each ladder rung, fix the GOP length and check keyframe alignment.
- For live, measure glass-to-glass latency with zerolatency and hardware encoding before and after changes.
- Re-run the comparison whenever you upgrade the encoder or change content type.