AV1 is the royalty-free video format from the Alliance for Open Media, with its bitstream frozen in 2018. Like every modern video standard, it defines only the decoder: the exact bits a decoder must accept and the pixels it must produce. An AV1 encoder is free to make any choices that produce a valid bitstream, and the difference between a fast encoder and a good one is entirely in how it searches the space of choices the format allows.
AV1 offers far more choices than H.264 or HEVC: bigger blocks with more ways to split them, dozens of intra modes, warped and compound motion, sixteen transform types, three in-loop filters and a synthetic film grain model. This article walks through those tools in the order an encoder uses them, explains what each one costs to search, and turns that into practical decisions: which encoder to run, which preset and settings, and how to avoid the common failures. The generic hybrid-encoder loop of prediction, rate-distortion optimisation and rate control is covered in video encoder architecture; this article focuses on what is specific to AV1.
Bitstream structure and the reference model
An AV1 stream is a sequence of OBUs (open bitstream units): a sequence header with stream-wide settings, temporal delimiters marking each presentation time, frame headers, tile groups carrying the coded data, and metadata such as HDR information. Containers such as MP4 and WebM carry the OBUs, and packaging formats layer on top as usual.
The decoder keeps eight reference frame slots. Each inter frame can predict from up to seven named references (LAST, LAST2, LAST3, GOLDEN, BWDREF, ALTREF2 and ALTREF), each mapped to one of the slots, and its header says which slots to overwrite with the new reconstruction. AV1 also allows frames that are decoded but never shown, typically a filtered future frame (an alternate reference) used as a high-quality anchor, and a cheap header that later displays an already-decoded frame. Encoders build their hierarchical GOPs out of these pieces; the general GOP trade-offs are covered in GOP structure.
Superblocks and partition search
Each frame is divided into superblocks of 128×128 or 64×64 pixels, chosen in the sequence header. Each superblock is split recursively. AV1 has ten partition types at each level: no split, horizontal or vertical halves, a four-way split, four T-shaped variants that split one half further, and horizontal or vertical splits into four strips. Blocks go down to 4×4.
This is the largest single source of encoding cost. A full search of every partition tree, with every prediction mode and transform in every block, is far too expensive, so all practical encoders prune. They skip partition types whose parent costs suggest they cannot win, stop splitting when the block is flat, reuse decisions from a fast first pass, and in libaom and SVT-AV1 use small learned models to predict which splits are worth evaluating. Presets mostly change how aggressive this pruning is. 128×128 superblocks help on high-resolution, smooth content; 64×64 is often better for lower resolutions and for encoder parallelism.
Prediction tools
Intra prediction builds a block from already-decoded pixels in the same frame. AV1 has 56 directional modes (eight base angles, each with seven small offsets), smooth and Paeth predictors for gradients and edges, recursive filter intra for textured blocks, chroma-from-luma which predicts colour planes as a scaled copy of the reconstructed luma, palette mode for blocks with few distinct colours, and intra block copy, which copies a block from elsewhere in the same frame. The last two exist for screen content such as slides, code and games, and intra block copy turns off the in-loop filters for that frame.
Inter prediction copies a block from a reference frame at a motion vector with up to eighth-pixel precision. Beyond single-reference prediction, AV1 adds compound prediction from two references with several blend types (average, distance-weighted, difference-weighted and wedge-shaped masks), overlapped block motion compensation to soften block edges, per-block warped motion fitted from neighbouring motion vectors, and frame-level global motion for camera pans and zooms. Each tool is another option to evaluate, so encoders enable the expensive ones only at slower presets or when cheap analysis suggests they will pay off.
Transforms, quantisation and entropy coding
The residual left after prediction is transformed. AV1 combines four one-dimensional kernels (DCT, ADST, flipped ADST and identity) separately for rows and columns into sixteen 2D transform types, on square and rectangular sizes from 4×4 to 64×64. For 64-point transforms only the lowest 32×32 frequencies are coded. Choosing the transform type and size per block is another search, usually pruned by trying DCT first and testing alternatives only when the residual has directional structure.
Quantisation uses an index from 0 to 255 with per-segment and per-superblock adjustments (delta-q), which is how encoders spend more bits on visually important or frequently referenced regions. Coefficients and modes are coded with a multi-symbol arithmetic coder whose probability tables (CDFs) adapt as symbols are coded and can be inherited from a reference frame, so later frames start from probabilities tuned to similar content.
In-loop filters and film grain
After reconstruction, AV1 runs a chain of filters whose output becomes the reference for later frames, in this order:
- Deblocking smooths block edges with adaptive filter lengths.
- CDEF (constrained directional enhancement filter) detects the main edge direction in each 8×8 block and filters along it to remove ringing without blurring the edge. The encoder picks a small set of filter strengths per frame and signals one per 64×64 area.
- Super-resolution, if enabled, upscales a frame coded at reduced width back to full width.
- Loop restoration applies a Wiener filter or a self-guided filter per restoration unit, with parameters chosen by the encoder to bring the reconstruction closer to the source.
Film grain synthesis sits outside this loop. The encoder denoises the source, codes the clean video, and estimates a grain model (an autoregressive pattern plus intensity-dependent scaling) that is sent in the frame header. The decoder adds synthetic grain to the output picture only; references stay clean. On grainy film content this saves a large share of bits, because real grain is noise and almost impossible to compress. The catch is measurement: PSNR and VMAF compare against the original grain and generally score synthesised grain badly, so judge grain settings by eye or with tests designed for it, as discussed in VMAF and quality metrics.
Rate control, lookahead and parallelism
Good AV1 quality depends heavily on lookahead. Encoders buffer future frames, build a temporally filtered alternate reference, and run a temporal dependency model (TPL in libaom and SVT-AV1) that estimates how much each block will be referenced later. Blocks that many future frames copy from get lower quantisers. Constant-quality modes (CRF or CQ) are standard for on-demand content; capped VBR and CBR with short lookahead are used for live and real-time streams.
Parallelism comes from three places. Tiles split a frame into independently coded rectangles, which helps both encoders and decoders run on many cores but costs some compression because prediction and entropy context do not cross tile edges. Row multithreading processes superblock rows in a wavefront within a tile. Frame and segment parallelism, which SVT-AV1 is built around, keeps several pictures in flight at once. For large on-demand jobs, the most scalable form is outside the encoder entirely: split the title at keyframes into chunks and encode them on separate machines.
Choosing an encoder
| Encoder | Strength | Typical use |
|---|---|---|
| libaom | Reference implementation; best compression at its slowest speeds; many tuning knobs | Research, archival, low-volume premium encodes |
| SVT-AV1 | Designed for multi-core scaling; wide preset range with strong speed/quality trade-off | Most software VOD and many live pipelines |
| rav1e | Written in Rust with a focus on safety | Projects that prefer a Rust codebase |
| Hardware (NVENC, Quick Sync, AMF, VA-API) | Real-time speed at low power; fewer tools searched | Live, cloud gaming, conferencing, high-volume UGC |
Hardware AV1 encoding is available on recent NVIDIA, Intel and AMD GPUs; check your exact model and driver, because support varies by generation. In FFmpeg the encoders are libaom-av1, libsvtav1, librav1e, av1_nvenc, av1_qsv, av1_amf and av1_vaapi, depending on how your build was configured.
# SVT-AV1: lower preset = slower and better. tune=0 targets subjective quality, tune=1 PSNR.
ffmpeg -i in.mov -c:v libsvtav1 -preset 5 -crf 30 -g 240 -pix_fmt yuv420p10le \
-svtav1-params tune=0:film-grain=8 -c:a copy out_svt.mkv
# libaom: constant quality needs -b:v 0; row-mt and tiles add parallelism.
ffmpeg -i in.mov -c:v libaom-av1 -crf 30 -b:v 0 -cpu-used 4 -row-mt 1 \
-tile-columns 1 -tile-rows 1 -g 240 out_aom.mkv
# NVIDIA hardware encoder for live or high-volume work.
ffmpeg -i in.mov -c:v av1_nvenc -preset p5 -cq 32 -g 120 out_nvenc.mp4SVT-AV1 presets run from slow, low numbers to fast, high numbers; the exact range has changed between releases, so check what your version accepts. Its film-grain level runs from 0 to 50. Use 10-bit output even for 8-bit sources: it usually improves compression in banding-prone gradients at little cost.
Worked example: sizing a VOD ladder encode
Suppose you publish a two-hour 24 fps film as a six-rung adaptive bitrate ladder with SVT-AV1. The title has 7,200 × 24 = 172,800 frames. Measure, do not guess: encode a representative 60-second clip at each candidate preset and record frames per second on your instance type and the VMAF-versus-bitrate curve.
Say your measurement at 4K shows 4 fps at preset 4 and 12 fps at preset 6, with preset 4 saving 6% bitrate at equal quality (illustrative numbers; yours will differ). At preset 4 the top rung takes 172,800 / 4 = 43,200 seconds, or 12 hours on one instance. Split into 60 keyframe-aligned chunks across 60 instances, it finishes in about 12 minutes plus overhead. Whether 6% is worth three times the compute depends on how often the title is watched: for a catalogue title streamed millions of times, saving bitrate on every view outweighs a one-off encode cost; for a news clip watched for a day, the faster preset wins. Per-title encoding applies the same reasoning to choosing the ladder itself, and adaptive bitrate streaming covers how players move between the rungs.
Failure modes
- Players that cannot decode it. Older devices without hardware AV1 decoding fall back to software or fail. Keep an H.264 or HEVC ladder and use capability detection.
- Too many tiles. Splitting a 1080p frame into many tiles for encoder speed costs visible efficiency. Use tiles for decoder requirements and rely on row and frame threading for encoder speed.
- Mismatched chunk boundaries. Chunked encodes with different GOP settings, or chunks not starting at keyframes, cause bitrate spikes and quality jumps at seams. Fix the GOP length and align chunks to it.
- Film grain judged by metrics. Teams reject grain synthesis because VMAF drops, or enable it at high strength on clean animation where it adds noise. Decide per content type with viewing tests.
- Live latency surprises. Long lookahead, hidden alternate references and frame parallelism add delay. Use the encoder's low-delay or real-time mode for interactive use.
- Upgrades changing output. Encoder releases change preset behaviour. Pin the version and re-run your quality comparison before upgrading.
Operational guidance and trade-offs
The central trade-off is compute against bits. Slower presets and wider tool searches save bitrate, and that saving is multiplied by every view, so popular content deserves slower encodes. Real-time use cases trade efficiency for latency and should usually use hardware.
Keep the encode settings in version control as a per-content-type profile: animation, film with grain, sports, screen content. Screen content benefits from palette and intra block copy, which most encoders enable through a tune or content setting. Track encode cost, average bitrate per rung and quality scores per release so regressions are visible.
What to do next
- Pick five representative clips covering your main content types, including one grainy and one screen-content clip.
- Encode each with SVT-AV1 at three presets and two CRF values, recording fps, file size and VMAF.
- Plot bitrate against quality per preset and choose the preset where extra compute stops paying for your expected views.
- Test film-grain synthesis on the grainy clip by viewing it on a real display, not only with metrics.
- Check decoder support across your audience's devices and keep a fallback codec ladder.
- Pin the encoder version and settings, and automate chunked encoding at keyframe boundaries.