Whether AV1 compresses better than H.264 and HEVC is settled, and the comparison, with coding tools, codec strings, licensing and a cost model, lives in H.264 vs H.265 vs AV1. This article starts after that decision. You have chosen to ship AV1; now you must pick an encoder, choose settings, build a ladder, make sure only devices that can decode it well receive it, keep everyone else working, and prove the change saved money without hurting viewers.

Most AV1 trouble in production comes from operations, not the codec: an encoder preset that takes six times longer than budgeted, film grain smeared into plastic, a phone that technically decodes AV1 in software and drains its battery doing so, keyframes that do not line up across codecs, and dashboards that report one blended number for two very different populations. Each section below takes one of those problems and gives the rule or code that handles it.

Advertisement

The production pipeline in one picture

A production AV1 service is a dual-codec service. The mezzanine, ideally a 10-bit master, goes through analysis, then two encode ladders: AV1 for capable devices and H.264 for everything else. Both are packaged into CMAF segments referenced by one manifest. The player probes decode capability before it picks renditions, and telemetry is split by the codec each session actually played, so you can compare cohorts instead of averages.

Mezzanine10-bit masterAnalysisgrain, complexityAV1 ladderSVT-AV1 or HWH.264 ladderfallback, alwaysPackage CMAFone manifest, two codecsCDNsame URLs per renditionPlayer capability gatedecodingInfo: smooth, efficientmanifestAV1 renditionscapable + efficientH.264 renditionseveryone elseQoE telemetry by codec cohortstartup, rebuffer, dropped frames, delivered quality, egressThe fallback ladder is permanent infrastructure, not a migration step: a large share of devices still cannot decode AV1 efficiently
Figure 1. A dual-codec AV1 pipeline. The capability gate and the cohort-split telemetry are what make AV1 safe to ship; the encode is the easy part.

Treat the H.264 ladder as permanent: older televisions, set-top boxes and budget phones will be on your traffic for years. The goal is to move watch hours to AV1 where decoding is efficient, and nowhere else.

Choosing an encoder

There are four realistic choices, and they trade speed, quality per bit and operational fit differently.

EncoderWhat it isFitsWatch out for
SVT-AV1Open-source CPU encoder from Intel and Netflix, now hosted by the Alliance for Open Media; v3.0 shipped in February 2025Default for VOD and much live work: scales across many cores, wide preset rangePresets were renumbered in 3.x, so copied settings from old blog posts are wrong
libaomThe reference encoderHighest-effort archival or reference encodes, conformanceMuch slower at comparable quality; rarely the right production choice
rav1eRust encoder focused on safetyTeams that want a memory-safe encoderSmaller ecosystem and less tuning documentation
HardwareAV1 encode blocks in NVIDIA RTX 40-series (Ada) and later, AMD RDNA 3 (VCN 4.0) and Intel Arc GPUsLive, real-time and high-volume UGC where cost per minute dominatesLower quality per bit than slow software presets; fewer tuning knobs

A common split is software for the VOD catalogue, where encode cost is amortised over many views, and hardware for live channels and user uploads. A tenfold costlier encode pays only if bits saved times views outweigh the compute. For more on how encoders spend their effort, see how video encoders work.

Advertisement

Settings that matter

Three knobs decide almost everything. The preset trades encode time for compression efficiency; lower is slower and better. The constant rate factor (CRF) sets the quality target; higher means fewer bits. The group-of-pictures length sets keyframe spacing, which must be identical across every rung and both codecs so the player can switch cleanly at segment boundaries. Two more are worth deciding deliberately: encode in 10-bit even from 8-bit sources, because AV1's Main profile supports it and it reduces banding in gradients at little cost, and use the perceptual tune rather than the PSNR tune for content people watch.

# One AV1 rung from a 10-bit mezzanine with SVT-AV1 through FFmpeg.
# Parameter names follow the SVT-AV1 docs; run `SvtAv1EncApp --help` for your
# version, because presets were renumbered in the 3.x series.
ffmpeg -i mezz_2160p_10bit.mov -map 0:v:0 \
  -vf "scale=1920:-2:flags=lanczos" \
  -c:v libsvtav1 -preset 5 -crf 32 \
  -g 96 \
  -pix_fmt yuv420p10le \
  -svtav1-params "tune=0:film-grain=8:film-grain-denoise=1" \
  -an rung_1080p_av1.mp4

# Check what you produced before you trust it:
ffprobe -v error -select_streams v:0 \
  -show_entries stream=codec_name,profile,pix_fmt,width,height \
  -of default=nw=1 rung_1080p_av1.mp4
ffprobe -v error -select_streams v:0 -skip_frame nokey \
  -show_entries frame=pts_time -of csv=p=0 rung_1080p_av1.mp4 > keyframes_av1.txt
# keyframes_av1.txt must match the H.264 rung's keyframe list for aligned switching

Treat preset and CRF as a matrix to measure, not a pair to copy. Encode a representative clip set, about ten minutes covering talking heads, sport, animation, dark scenes and grainy film, at three presets and four CRF values, score each output with VMAF and a banding metric, and record wall-clock time per output minute on your actual instance type. Pick the slowest preset your throughput budget allows, then the CRF that reaches your quality target, and always run the ffprobe checks: most ladder bugs show up in the keyframe list.

Film grain synthesis

Grain is the most expensive thing an encoder can try to preserve, because it is random noise and noise does not predict. Encoders that try to keep it spend enormous numbers of bits; encoders that do not, smooth it away and produce the waxy look viewers notice. AV1 has a third option built into the format: the encoder estimates the grain, removes it before compression, and sends a compact parametric description. The decoder synthesises statistically similar grain on top of the clean picture.

In SVT-AV1 the film-grain parameter (0 to 50) sets the denoising and synthesis strength, and film-grain-denoise controls whether the encoded picture itself is denoised; its default of 0 sends grain parameters without denoising, keeping more source texture but saving fewer bits. Two rules keep it safe. First, enable it per title from analysis, not globally: clean animation and screen content gain nothing and can pick up artificial texture. Second, judge it by eye on reference clips, because VMAF compares against the grainy source and penalises synthesised grain that does not match pixel for pixel, even when it looks right.

Building the AV1 ladder

An AV1 ladder should not copy the H.264 ladder rung for rung. Because each AV1 rung reaches a given quality with fewer bits, you can either keep the same resolutions at lower bitrates, saving egress, or keep the bitrates and raise the resolution at each step, improving quality for constrained viewers. Most services do some of each: lower bitrates at the top and higher resolution at the bottom, where quality is worst.

Worked example. Suppose your own encodes show that AV1 at preset 5 reaches the same VMAF as your H.264 ladder with 30 percent fewer bits on your catalogue; this figure is yours to measure, not a constant. Suppose also that 60 percent of watch hours come from devices that pass the capability gate in the next section. If you spend the whole gain on egress, delivered bytes fall by roughly 0.6 times 0.3, or 18 percent. At 5 PB of monthly egress, that is about 0.9 PB. Against that, the AV1 ladder adds encode cost and doubles storage for every title that carries both codecs. If the encode costs 4 compute-hours per content hour and you publish 2,000 content hours a month, you add 8,000 compute-hours monthly; convert both sides to money at your own prices. The usual result is that AV1 pays off quickly for popular titles and never for the long tail, which leads to a sensible policy: encode AV1 for titles above a view threshold, and add it later to titles that cross the threshold.

Gating by decode capability

The question is not whether a device can decode AV1, but whether it can decode this rung smoothly and efficiently. Recent platforms help: Apple added AV1 hardware decode in the A17 Pro and M3 generations, recent NVIDIA, AMD and Intel GPUs decode it in hardware, and Android has shipped a software AV1 decoder since Android 10, with Google later delivering the faster dav1d decoder to newer versions. Software decode works, but on a phone it costs battery and can drop frames at high resolutions, so supported is not the same as suitable.

// Decide per device, per rung, whether AV1 is worth it.
const AV1 = 'video/mp4; codecs="av01.0.08M.10"';   // profile 0, level 4.0, Main tier, 10-bit
const AVC = 'video/mp4; codecs="avc1.640028"';     // H.264 High, level 4.0

async function probe(contentType, r) {
  if (!('mediaCapabilities' in navigator)) {
    return { supported: MediaSource.isTypeSupported(contentType), smooth: false, powerEfficient: false };
  }
  return navigator.mediaCapabilities.decodingInfo({
    type: 'media-source',
    video: { contentType, width: r.width, height: r.height, bitrate: r.bitrate, framerate: r.fps },
  });
}

async function chooseRenditions(ladder) {          // ladder: [{codec, width, height, bitrate, fps}]
  const top = ladder.filter(r => r.codec === 'av1').sort((a, b) => b.height - a.height)[0];
  const info = top ? await probe(AV1, top) : { supported: false };
  const onBattery = navigator.getBattery ? !(await navigator.getBattery()).charging : false;
  const useAv1 = info.supported && info.smooth && (info.powerEfficient || !onBattery);
  report('codec_gate', { useAv1, ...info, onBattery });   // log the decision, not just the outcome
  return ladder.filter(r => r.codec === (useAv1 ? 'av1' : 'avc'));
}

The decodingInfo call returns three booleans: supported, smooth and powerEfficient. The last is a reasonable proxy for hardware decode. Probe with the top rung's parameters, because a device may handle 720p in software but not 4K. Log every decision with its inputs, so that when a device model starts reporting dropped frames you can see what the gate believed. On platforms without the API, such as some smart-TV runtimes, maintain an explicit allow-list by device model rather than guessing.

Packaging and switching

Package AV1 and H.264 renditions into CMAF fragments with identical segment boundaries, as described in CMAF packaging. In DASH, put each codec in its own adaptation set; in HLS, list the AV1 variants with their codec strings so players that cannot handle them skip them. Never switch codecs mid-session; it usually forces a decoder reset and a visible stall. Decide once at startup, adapt bitrate within the chosen codec, and fall back to H.264 only after a decode error, recording that fallback as an event. If you ship HDR, the AV1 ladder needs the same colour metadata discipline as any HDR ladder; see HDR video.

Real-time AV1

AV1 is also used in real-time video, where its screen-content tools, palette mode and intra block copy, make it strong for shared desktops and slides. The constraint here is encode CPU on the sender, often a laptop already busy with the call. Use temporal scalability so a selective forwarding unit can drop layers per receiver instead of asking the sender to re-encode, and keep fallbacks negotiated so a participant on an older device still joins.

// Real-time AV1 in WebRTC: prefer AV1, keep H.264/VP8 as negotiated fallbacks,
// and send three temporal layers so an SFU can thin the stream per receiver.
const tx = pc.addTransceiver(cameraTrack, {
  direction: 'sendonly',
  sendEncodings: [{ scalabilityMode: 'L1T3', maxBitrate: 1_200_000 }],
});
const codecs = RTCRtpReceiver.getCapabilities('video').codecs;   // preferences must come from receiver capabilities
const rank = c => (c.mimeType === 'video/AV1' ? 0 : c.mimeType === 'video/H264' ? 1 : 2);
tx.setCodecPreferences([...codecs].sort((a, b) => rank(a) - rank(b)));

// After negotiation, confirm what is actually being sent and whether CPU is limiting it.
const stats = await tx.sender.getStats();
for (const s of stats.values()) {
  if (s.type === 'outbound-rtp') console.log(s.encoderImplementation, s.qualityLimitationReason, s.framesPerSecond);
}

Watch qualityLimitationReason in the sender statistics. If it often reads cpu, the software AV1 encoder is too heavy for those machines; use a hardware encoder where encoderImplementation shows one, or demote AV1 on those device classes.

Rolling it out and proving it worked

Roll out by cohort, not by title: AV1 for a random slice of capable sessions, with a matched H.264 control from the same device classes, because AV1-capable devices are newer and would flatter any naive comparison. Compare startup time, rebuffering ratio, dropped frames from video.getVideoPlaybackQuality(), delivered resolution and bitrate, and an estimate of delivered quality. Add device-level guards: if a model's dropped-frame rate under AV1 is meaningfully worse than under H.264, remove it from the allow-list automatically. Measure egress per watch hour in the CDN logs, since that is where the savings appear; video CDN design covers how cache efficiency changes when you add a second codec.

Failure modes

  • Software decode on battery: a phone passes a support check, decodes 1080p AV1 in software, heats up and drops frames. Gate on smooth and power-efficient, not on supported.
  • Misaligned keyframes between codecs or rungs, from scene-cut detection placing extra keyframes differently, causing switch failures or stalls.
  • Encode backlog: a slow preset chosen on a benchmark clip cannot keep up with the publishing rate, and AV1 renditions arrive days late.
  • Plastic faces: grain removed without synthesis, or synthesis set too weak, on film content.
  • Blended dashboards that average AV1 and H.264 sessions together and hide a regression in one cohort.

Trade-offs at a glance

ChoiceGainsCosts
Slow software presetMost bits saved per viewEncode time and compute spend
Hardware encoderReal-time, cheap per minuteLess efficient; fewer tuning controls
Spend gain on egressLower CDN billNo visible quality improvement
AV1 for popular titles onlyEncode and storage stay proportional to valuePolicy and re-encode tooling to maintain
Film grain synthesisLarge savings on grainy contentNeeds per-title decisions and human review

What to do next

  1. Build a ten-minute clip set that represents your catalogue, including grainy and screen content.
  2. Measure a preset by CRF matrix with VMAF, a banding check and encode time on your real instance type.
  3. Pick an encoder per workload: software for popular VOD, hardware for live and user uploads.
  4. Decide film grain per title from analysis and review reference clips by eye.
  5. Design the AV1 ladder separately from H.264 and decide where the gain goes: egress, quality or both.
  6. Implement the decodingInfo gate with smooth and power-efficient checks, plus an allow-list for platforms without the API.
  7. Align keyframes across all rungs and both codecs, and verify with ffprobe in CI.
  8. Roll out by randomised cohort with a matched control, and split every dashboard by codec.
  9. Set a view threshold for AV1 encoding and backfill titles that cross it.
Key takeaway: AV1 in production is a dual-codec operation, not a codec swap. Choose the encoder by workload, measure presets and CRF on your own content, use film grain synthesis deliberately, design the AV1 ladder on its own terms, and only send AV1 to devices that decode it smoothly and efficiently. Keep H.264 as permanent fallback, align keyframes everywhere, encode AV1 where views justify it, and prove the savings with cohort-split telemetry.