Choosing between H.264, H.265 (HEVC) and AV1 is rarely a choice of one. Each newer codec spends more encoder compute to save bits, and each is decodable by a smaller, though growing, share of devices. The practical question is which codecs to encode, for which titles, at which rungs of the ladder, and how players pick among them.

This article explains what actually changed between the three generations, shows how to measure the savings on your own content instead of trusting headline percentages, lays out device support and licensing as they stood when this was written, and works through a cost model that tells you when the extra encode is paid back by egress savings. The encoder internals of AV1 are covered separately in AV1 encoder architecture; this page is about the delivery decision.

Advertisement

Same architecture, better tools

All three are block-based hybrid codecs. Each frame is split into blocks; each block is predicted from neighbouring pixels in the same frame (intra) or from previously decoded frames (inter); the prediction error is transformed, quantised and entropy coded; and in-loop filters clean up the reconstruction before it becomes a reference. The GOP structure decides which frames reference which. What changed across generations is the flexibility of every one of those stages:

StageH.264 / AVCH.265 / HEVCAV1
Largest block16x16 macroblockCoding tree unit up to 64x64, quadtree splitSuperblock 128x128 or 64x64, quadtree plus rectangular splits
Transforms4x4 and 8x8 integer DCT4x4 to 32x32, DST for 4x4 intra luma4x4 to 64x64 including rectangular; DCT, ADST, flipped ADST and identity per direction
Intra prediction9 modes for 4x4 and 8x8 blocks, 4 for 16x1635 modes: planar, DC, 33 angles56 directional plus smooth, Paeth, chroma-from-luma, palette, intra block copy
Inter predictionQuarter-pel, multiple referencesMerge mode, advanced motion vector predictionUp to 7 references per frame, compound modes, warped and global motion, OBMC
Entropy codingCAVLC or CABACCABACAdaptive multi-symbol arithmetic coding
In-loop filtersDeblockingDeblocking and SAODeblocking, CDEF, loop restoration; optional film-grain synthesis
ParallelismSlicesTiles and wavefront parallel processingTiles

Larger blocks matter most at high resolutions, where big flat areas can be coded with very few bits. Richer prediction and transform choices let the encoder fit each region more closely. Film-grain synthesis in AV1 is unusual: the encoder can remove grain, code the clean picture cheaply, and send parameters so the decoder re-adds statistically similar grain. For grainy film content the savings can be large; for clean animation they are nil.

The cost of all this flexibility is search. The encoder must evaluate many more partition, mode and transform combinations, which is why AV1 and HEVC encoders run far slower than H.264 at comparable quality settings. Decode complexity rises too, but much less, and dedicated hardware absorbs it.

Measure efficiency on your own content

Published comparisons commonly report HEVC needing roughly 25 to 50 percent fewer bits than H.264 for the same perceived quality, and AV1 roughly a further 10 to 30 percent below HEVC. Treat those as a starting hypothesis. The real number depends on content, resolution, the encoder implementation, its preset, and the quality metric. Measure it with a BD-rate comparison: encode the same clips at several quality points per codec, score each with VMAF, and compare the curves.

# Encode one clip at four quality points per codec (ffmpeg with libx264, libx265, libsvtav1)
for crf in 19 23 27 31; do
  ffmpeg -i clip.mov -an -c:v libx264 -preset slow -crf $crf -pix_fmt yuv420p avc_$crf.mp4
done
for crf in 22 26 30 34; do
  ffmpeg -i clip.mov -an -c:v libx265 -preset slow -crf $crf -tag:v hvc1 -pix_fmt yuv420p10le hevc_$crf.mp4
done
for crf in 28 34 40 46; do
  ffmpeg -i clip.mov -an -c:v libsvtav1 -preset 6 -crf $crf -g 240 -pix_fmt yuv420p10le av1_$crf.mp4
done
# Score each against the source with VMAF (libvmaf filter), keeping bitrate and score per file
ffmpeg -i av1_34.mp4 -i clip.mov -lavfi libvmaf=log_fmt=json:log_path=av1_34.json -f null -
import numpy as np

def bd_rate(rate_a, q_a, rate_b, q_b):
    """Average bitrate difference of B relative to A at equal quality (negative = B saves bits).
    Fits log-rate as a cubic in quality and integrates over the overlapping quality range."""
    pa = np.polyfit(q_a, np.log(rate_a), 3)
    pb = np.polyfit(q_b, np.log(rate_b), 3)
    lo, hi = max(min(q_a), min(q_b)), min(max(q_a), max(q_b))
    ia, ib = np.polyint(pa), np.polyint(pb)
    avg_a = (np.polyval(ia, hi) - np.polyval(ia, lo)) / (hi - lo)
    avg_b = (np.polyval(ib, hi) - np.polyval(ib, lo)) / (hi - lo)
    return (np.exp(avg_b - avg_a) - 1) * 100

Use a dozen clips that represent your catalogue: talking heads, sport, animation, dark scenes, film grain. Report BD-rate per content type, not one average, because the decision often splits by genre. The VMAF article explains which model to use for phones versus televisions and why you should still look at the frames.

Advertisement

Encode cost

Compute cost scales with the encoder preset more than with the codec name. A fast AV1 preset may be cheaper than a very slow x265 preset and still compress better; a slow AV1 preset can cost an order of magnitude more CPU time than x264 for the same clip. Benchmark encode time per minute of content at the presets you would ship, on the instance types you would use, and record it next to the BD-rate figures.

Hardware encoders change the economics for live and high-volume work. Recent consumer and data-centre GPUs and some CPUs include AV1 and HEVC encode blocks that run in real time at low power. They generally compress less efficiently than slow software presets, so a common split is hardware encoding for live and user-generated content, and slow software encoding for premium on-demand titles that are watched many times.

Chunked parallel encoding reduces wall-clock time but not compute: the transcoding pipeline article covers splitting on GOP boundaries so codec choice does not change your job farm design.

Decode support is the real constraint

A codec only saves egress for viewers who can decode it, and on battery devices it must be decoded in hardware to be acceptable. The landscape when this was written:

  • H.264 decodes essentially everywhere, in hardware. It remains the mandatory fallback.
  • HEVC has hardware decode on most phones, televisions and recent PCs. Safari has supported it for years. Chrome enabled HEVC playback via hardware decode by default from version 107 (108 on Linux), initially for clear content only. Firefox 134 added hardware-accelerated HEVC playback on Windows. Support for HEVC with DRM varies by platform and key system, so test your exact DRM path.
  • AV1 has hardware decode on recent Android flagships, many recent smart TVs, and recent desktop GPUs. Apple ties AV1 to its hardware decoders: A17 Pro and later iPhones and M3 and later Macs, with support from Safari 17. Software decoding (dav1d) is fast on desktops but costs battery on phones.

Do not decide from tables like this one; decide from your own player logs. Ask the device what it can do efficiently and record the answer:

const CANDIDATES = [
  { codec: "av1",  type: 'video/mp4; codecs="av01.0.08M.10"' },
  { codec: "hevc", type: 'video/mp4; codecs="hvc1.2.4.L120.90"' },
  { codec: "avc",  type: 'video/mp4; codecs="avc1.640028"' },
];

async function pickCodec(width, height, bitrate) {
  for (const c of CANDIDATES) {
    const info = await navigator.mediaCapabilities.decodingInfo({
      type: "media-source",
      video: { contentType: c.type, width, height, bitrate, framerate: 30 },
    });
    // powerEfficient is the best available hint that decoding is hardware-backed
    if (info.supported && info.smooth && info.powerEfficient) return c.codec;
  }
  return "avc";
}

For protected content, also check the key system with navigator.requestMediaKeySystemAccess using the same codec strings, because a device may decode a codec in the clear but not inside its DRM pipeline. Log the chosen codec with every session so the cost model below uses real shares, not guesses.

Codec strings and packaging

With CMAF, all three codecs share the same fragmented MP4 segment format, so one packager produces them all and the player switches codec between variants without a different container. The manifest must advertise each variant's codec precisely, because players filter variants by codec string before downloading anything:

#EXTM3U
#EXT-X-STREAM-INF:BANDWIDTH=2400000,RESOLUTION=1280x720,CODECS="av01.0.05M.10,mp4a.40.2"
av1/720p.m3u8
#EXT-X-STREAM-INF:BANDWIDTH=3000000,RESOLUTION=1280x720,CODECS="hvc1.2.4.L93.90,mp4a.40.2"
hevc/720p.m3u8
#EXT-X-STREAM-INF:BANDWIDTH=4200000,RESOLUTION=1280x720,CODECS="avc1.64001f,mp4a.40.2"
avc/720p.m3u8

The strings encode profile and level: avc1.64001f is High profile, level 3.1; hvc1.2.4.L93.90 is Main 10, level 3.1; av01.0.05M.10 is Main profile, level 3.1, Main tier, 10-bit. A wrong level causes a capable device to skip a variant or an incapable one to choke on it, so generate strings from the encoded bitstream rather than typing them. For HEVC in HLS use the hvc1 sample entry, which carries parameter sets in the init segment, rather than hev1; with ffmpeg that is -tag:v hvc1. In DASH, put each codec in its own AdaptationSet so players switch rungs within a codec, not across codecs mid-stream.

Licensing, at a high level

This is an engineering summary, not legal advice. H.264 is licensed through a long-established pool now administered by Via LA, which formed from the merger of MPEG LA and Via Licensing. HEVC licensing has been fragmented across several pools and independent licensors, which slowed its adoption on the web; in 2025 Access Advance announced it would take over administration of Via LA's HEVC and VVC program. AV1 is published by the Alliance for Open Media under a royalty-free patent licence, but third parties have formed pools asserting patents against it, including Access Advance's Video Distribution Patent pool launched in January 2025, which covers streaming with HEVC, VVC, VP9 and AV1. Have counsel review your specific distribution model before committing; the engineering decision should not assume any codec is free of licensing questions.

A worked cost model

Take a catalogue served at an average 3 Mbps with H.264. One viewing hour is 3 Mbps times 3,600 s, which is 10.8 gigabits or 1.35 GB. A title watched 200,000 hours a month therefore costs 270 TB of egress; at an illustrative $0.02 per GB, about $5,400 a month.

Suppose your BD-rate tests show AV1 saving 35 percent against your H.264 ladder at equal VMAF, and player logs show 45 percent of viewing on devices that pick AV1. Savings are 0.35 times 0.45, about 16 percent of egress, or roughly $850 a month for that title. If a full AV1 ladder for the two-hour title costs $150 of compute at your chosen preset, it pays back in under a week. A tail title watched 200 hours a month saves under a dollar a month and never pays back its encode.

def payback_months(hours_per_month, avg_mbps, bd_saving, codec_share, usd_per_gb, encode_usd):
    gb_per_hour = avg_mbps * 3600 / 8 / 1000
    monthly_saving = hours_per_month * gb_per_hour * usd_per_gb * bd_saving * codec_share
    return float("inf") if monthly_saving == 0 else encode_usd / monthly_saving

# Head title: about 0.18 months.  Tail title: about 176 months, so never encode it in AV1.
print(payback_months(200_000, 3.0, 0.35, 0.45, 0.02, 150))
print(payback_months(200, 3.0, 0.35, 0.45, 0.02, 150))

The usual policy that falls out is tiered: H.264 for everything, HEVC where you need 10-bit HDR or have many HEVC-only devices such as older televisions, and AV1 for titles above a viewing threshold, re-evaluated monthly as popularity changes. Combine it with per-title encoding so each codec's ladder is also fitted to the content.

Failure modes

  • Software decode on phones. A device reports AV1 as supported but decodes in software; playback works but drains the battery and drops frames at high resolutions. Require powerEfficient for the high rungs.
  • Codec switching mid-session. Mixing codecs in one adaptation set makes some players reinitialise decoders and stall. Keep each codec in its own variant group.
  • Mismatched levels. Hand-written codec strings that understate the level make a device decode a stream it cannot sustain.
  • HDR metadata lost in transcoding. Colour primaries, transfer function and mastering metadata must be carried through every encoder; verify with a probe on the output, not the input.
  • Quality judged on averages. A codec that wins on mean VMAF can lose badly on dark gradients or grain. Inspect the worst segments.
  • Storage blow-up. Three ladders triple storage and packaging work. Apply the payback rule to storage as well as compute.

What to do next

  1. Pick a dozen representative clips and run the BD-rate comparison for all three codecs at the presets you would actually ship.
  2. Record encode time per minute of content per codec and preset on your production instance types.
  3. Add the MediaCapabilities probe to your player and log the chosen codec, device and DRM path for every session for two weeks.
  4. Plug the measured savings, codec shares and encode costs into the payback model and set a viewing threshold for AV1 and HEVC.
  5. Package all codecs as CMAF with generated codec strings, one variant group or AdaptationSet per codec, and keep H.264 as the universal fallback.
  6. Have the licensing position for your distribution model reviewed before launch.
  7. Re-run the share and payback analysis quarterly; device support for AV1 is still rising.
Key takeaway: H.264, HEVC and AV1 share one block-based design; each generation adds more flexible partitioning, prediction, transforms and filtering, trading encoder compute for fewer bits. The right mix depends on measured BD-rate on your content, encode cost at your presets, and the share of viewers whose devices decode each codec efficiently. Keep H.264 as the universal fallback, add HEVC for HDR and HEVC-only devices, encode AV1 where viewing volume pays back the extra compute, and let the player choose with a capability probe.