AV1 is a royalty-free video coding format from the Alliance for Open Media (AOMedia), whose members include browser makers, streaming services and chip vendors. The bitstream was frozen in 2018, building on ideas from Google's VP9 and VP10 work, Mozilla and Xiph's Daala and Cisco's Thor. It now carries a large share of streaming video on devices that can decode it. In mid-2026 AOMedia published version 1.0 of its successor, AV2, but at the time of writing no consumer device decodes AV2, so AV1 remains the open codec you can actually deploy.
This article explains what is inside AV1 and how to work with it. It covers the coding tools that give AV1 its efficiency, the bitstream structure with a small parser, the codec strings players need, and practical encoding with SVT-AV1 and libaom and decoding with dav1d. Device support, licensing and cost comparisons with H.264 and HEVC are covered in H.264 vs H.265 vs AV1, and the general structure of a block-based encoder in the video encoder deep dive.
The decoder is the specification
Like other modern codecs, AV1 standardises only the decoder. The specification says exactly how to turn bits into pixels; any encoder that produces a conforming bitstream is valid, and encoders compete on how cleverly they choose among the options. That is why two AV1 encoders can differ greatly in quality and speed while producing files every AV1 decoder can play.
The decoding pipeline has the usual hybrid-codec shape: parse headers, entropy-decode symbols, form a prediction from already-decoded pixels, add a decoded residual, then filter. AV1's differences are in the details: many more prediction modes, more transform types, three in-loop filters instead of one, and a grain synthesis step that runs after the loop on the displayed picture only.
Superblocks and partitions
Each frame is divided into superblocks of 128x128 or 64x64 luma samples, chosen in the sequence header. Each superblock is split recursively by a partition tree with ten partition types: none, horizontal and vertical halves, a four-way split, four T-shaped three-way splits, and horizontal and vertical four-way strips with a 4:1 aspect ratio. Blocks go down to 4x4.
Large blocks code smooth areas cheaply; small and rectangular blocks follow edges and motion boundaries. The cost is search: an encoder that tried every partition with every mode would be impossibly slow, so encoder presets mostly differ in how aggressively they prune this search.
Intra prediction tools
| Tool | What it does | Where it helps |
|---|---|---|
| Directional modes | 8 nominal angles, each with 7 fine offsets of 3 degrees, so 56 directions | Edges and textures at any angle |
| Smooth, smooth-H, smooth-V | Quadratic interpolation from the block edges | Gradients such as sky |
| Paeth | Picks left, top or top-left per pixel | Sharp structured content |
| Chroma from luma (CfL) | Predicts chroma as a scaled copy of the decoded luma pattern | Colour edges that follow brightness edges |
| Recursive filter intra | Predicts small 4x2 patches from their neighbours in sequence | Detailed textures |
| Palette | Codes a block as a few colours plus an index map | Screen content, graphics, text |
| Intra block copy | Copies an already decoded area of the same frame | Repeated glyphs and UI in screen content |
Palette and intra block copy make AV1 strong for screen sharing and slides, which is one reason real-time communication products adopted it. Intra block copy disables the loop filters for that frame, so encoders enable it only when the content benefits.
Inter prediction tools
A frame can reference up to seven of the eight frames held in the decoder's reference buffer, named LAST, LAST2, LAST3, GOLDEN, BWDREF, ALTREF2 and ALTREF. The backward and alternate references allow hierarchical structures in which future frames are coded first and used to predict the frames in between, as described in GOP structure.
- Motion vectors at up to one-eighth pixel precision, predicted from a dynamically built list of spatial and temporal candidates, so most vectors cost very few bits.
- Compound prediction blends two references, either with fixed or distance-based weights, a wedge-shaped mask, or a mask derived from where the two predictions differ.
- Overlapped block motion compensation blends a block's prediction with predictions using neighbours' vectors near its edges, reducing blocking at motion boundaries.
- Warped and global motion model rotation, zoom and shear with an affine transform, per block or for the whole frame, which helps with camera pans and zooms.
- Switchable interpolation filters, chosen separately for horizontal and vertical directions.
Transforms, quantisation and entropy coding
The residual is transformed with one of four one-dimensional kernels in each direction: DCT, ADST, flipped ADST and identity, giving 16 two-dimensional combinations. Sizes run from 4x4 to 64x64 and include rectangular shapes. For 64-point transforms only the lowest 32x32 coefficients are coded, since high frequencies at that size are rarely worth the bits. The identity transform suits screen content, where sharp edges are poorly represented by smooth basis functions.
Quantisation is controlled by a base index from 0 to 255, with optional per-superblock deltas for adaptive quantisation and optional quantisation matrices that weight frequencies differently. Symbols are coded with an adaptive multi-symbol arithmetic coder that handles alphabets of up to 16 values at once rather than one binary decision at a time, and adapts its probabilities as it decodes.
Loop filters and film grain synthesis
After reconstruction, three filters run in order. The deblocking filter smooths block edges. CDEF, the constrained directional enhancement filter, finds the dominant edge direction in each 8x8 block and filters along it, removing ringing without blurring the edge. Loop restoration then applies either a Wiener filter or a self-guided filter per restoration unit, with parameters chosen by the encoder. Optionally, a frame can be coded at reduced horizontal resolution and upscaled before loop restoration, called super-resolution, which helps at very low bitrates.
Film grain synthesis is AV1's most distinctive tool. Grain is random, so coding it exactly costs huge numbers of bits. Instead, the encoder denoises the source, codes the clean picture, and sends a small set of parameters describing the grain: an autoregressive model of its texture and scaling functions linking its strength to brightness. The decoder generates matching grain and adds it to the displayed frame. Because grain is not in the reference frames, it does not disturb prediction.
Two cautions. Objective metrics compare pixels, and synthesised grain never matches the source pixel for pixel, so PSNR and VMAF can score grain-synthesised encodes worse than they look; compare against the denoised source or use subjective checks, as discussed in VMAF and quality metrics. And test playback on your target devices, since a player that skips the grain step shows a cleaner but flatter picture than intended.
The bitstream: OBUs and codec strings
An AV1 stream is a sequence of Open Bitstream Units. Each OBU starts with a one-byte header holding a forbidden bit, a four-bit type, an extension flag and a size-present flag. An optional extension byte carries temporal and spatial layer IDs for scalable streams, and the payload size follows as a LEB128 variable-length integer. The sequence header carries profile, level, bit depth and other stream-wide settings; frame headers and tile groups carry each frame, and the FRAME type combines both.
Tiles split a frame into independently decodable rectangles, which lets decoders use several threads or hardware cores. Containers wrap OBUs: MP4 uses an av1C configuration box holding the sequence header, while WebM and IVF carry them too. The parser below walks the OBUs in an IVF file and prints each frame's structure, which is a fast way to check what an encoder actually produced.
import struct, sys
OBU_TYPES = {1: "SEQUENCE_HEADER", 2: "TEMPORAL_DELIMITER", 3: "FRAME_HEADER",
4: "TILE_GROUP", 5: "METADATA", 6: "FRAME",
7: "REDUNDANT_FRAME_HEADER", 8: "TILE_LIST", 15: "PADDING"}
def leb128(buf, pos):
value = 0
for i in range(8):
byte = buf[pos + i]
value |= (byte & 0x7F) << (7 * i)
if not byte & 0x80:
return value, pos + i + 1
raise ValueError("leb128 longer than 8 bytes")
def walk_obus(buf):
pos = 0
while pos < len(buf):
header = buf[pos]
if header & 0x80:
raise ValueError("forbidden bit set: not an OBU stream or misaligned")
obu_type, has_ext, has_size = (header >> 3) & 0x0F, (header >> 2) & 1, (header >> 1) & 1
pos += 1
temporal_id = spatial_id = 0
if has_ext:
temporal_id, spatial_id = buf[pos] >> 5, (buf[pos] >> 3) & 0x03
pos += 1
if not has_size:
raise ValueError("OBU without size field (only legal as the last OBU in a container sample)")
size, pos = leb128(buf, pos)
yield OBU_TYPES.get(obu_type, f"RESERVED_{obu_type}"), temporal_id, spatial_id, size
pos += size
def ivf_frames(path):
data = open(path, "rb").read()
assert data[:4] == b"DKIF", "not an IVF file"
pos = struct.unpack_from("<H", data, 6)[0] # header length, normally 32
while pos + 12 <= len(data):
size, pts = struct.unpack_from("<IQ", data, pos)
yield pts, data[pos + 12: pos + 12 + size]
pos += 12 + size
for pts, frame in ivf_frames(sys.argv[1]):
print(pts, [(t, size) for t, _, _, size in walk_obus(frame)])Players and manifests describe AV1 streams with a codec string such as av01.0.08M.10. The fields are profile (0 Main, 1 High, 2 Professional), the level index (08 means level 4.0) with tier (M main, H high), and bit depth. Main profile covers 8-bit and 10-bit 4:2:0, which is what almost all distribution uses. Getting the string wrong makes capability checks fail, so generate it from the sequence header rather than by hand.
Encoding and decoding in practice
Three open encoders dominate. SVT-AV1, now developed under AOMedia, is the usual production choice: higher preset numbers are faster and lower ones compress better, the documented range is roughly 0 to 13 (check your build's help output for its exact range), and it scales across many cores. libaom is the reference encoder, slower but useful for comparison and for features other encoders lack. rav1e, written in Rust, focuses on safety and speed. For decoding, dav1d is the widely deployed software decoder used by major browsers and media players.
# SVT-AV1: the usual production choice. Higher preset = faster; lower CRF = higher quality.
ffmpeg -i master.mov -c:v libsvtav1 -preset 6 -crf 32 -g 240 -pix_fmt yuv420p10le \
-svtav1-params film-grain=8 -c:a copy out_svt.mkv
# libaom: slower reference encoder, useful for comparisons and some special modes.
ffmpeg -i master.mov -c:v libaom-av1 -crf 32 -b:v 0 -cpu-used 4 -row-mt 1 \
-tiles 2x2 -g 240 -c:a copy out_aom.mkv
# Extract the raw stream into IVF for inspection, then decode with dav1d.
ffmpeg -i out_svt.mkv -c:v copy -an -f ivf out.ivf
dav1d -i out.ivf -o decoded.y4m --threads 8Use 10-bit encoding even for 8-bit sources where your target devices support it; the extra internal precision usually improves compression and reduces banding. Choose keyframe intervals from your segment length, and use tiles when hardware or low-end software decoders need parallelism, accepting a small efficiency loss at tile boundaries.
Worked example: adding an AV1 ladder
A service with an H.264 ladder wants to add AV1 for capable devices. The steps, with illustrative rather than measured numbers:
- Pick 20 representative titles, including grainy film, animation and screen recordings.
- Encode each at several CRF values with SVT-AV1 at preset 6 and at preset 4, keeping 2-second keyframe intervals aligned with the existing ladder.
- Measure VMAF against the source and build rate-quality curves per title; for grainy titles, also compare with film grain synthesis on and the grain level tuned by eye.
- Choose rungs where AV1 reaches the existing H.264 quality at clearly lower bitrate, as in per-title encoding. Use the slower preset only for titles watched enough to repay the compute.
- Package as CMAF with correct av01 codec strings, and serve AV1 only to clients whose capability check passes, keeping H.264 as the fallback.
- Watch startup time, rebuffering and battery reports on software-decoding devices before widening the rollout.
Failure modes and trade-offs
| Problem | Cause | Fix |
|---|---|---|
| Choppy playback on older phones | Software decode of high-resolution AV1 | Cap resolution for software decoders or serve H.264 |
| Grain looks wrong or missing | Player or decoder handling of grain parameters | Test on real devices; lower grain strength or disable for that target |
| Metrics say worse, eyes say fine | Synthesised grain penalised by pixel metrics | Compare against denoised source; add subjective review |
| Encode farm cannot keep up | Preset too slow for volume | Faster preset for the long tail, slow preset only for popular titles |
| Capability check rejects valid streams | Wrong profile, level or bit depth in codec string | Derive codec strings from the sequence header |
The central trade-off is encode compute against delivery bytes. AV1 encodes cost more CPU time than H.264 at comparable presets, and that cost is repaid only when enough views share the savings. Popular content and bandwidth-constrained audiences justify slow presets; rarely watched content does not.
What to do next
- Encode three of your own titles with SVT-AV1 at two presets and several CRF values, and plot rate against VMAF.
- Run the OBU parser on the output and confirm profile, bit depth and frame structure are what you intended.
- Try film grain synthesis on your grainiest title and judge it on a real screen, not only by metrics.
- Inventory which of your client devices decode AV1 in hardware and which would fall back to software.
- Generate codec strings from the sequence header and test capability checks in each target player.
- Decide which titles justify slow presets by comparing encode cost with expected delivery savings.