H.265, also called High Efficiency Video Coding or HEVC, is the codec behind most 4K and HDR delivery of the last decade: broadcast UHD, Blu-ray UHD, much of streaming's HDR catalogue and the default recording format on many phones. It keeps the same hybrid design as H.264 (predict a block, transform the difference, entropy code the result) but replaces almost every tool with a more flexible one, and it adds structure for parallel decoding.
This article explains HEVC from the bitstream up so you can make informed encoding and packaging decisions and debug the problems that actually happen in production. The comparison with H.264 and AV1, codec strings and licensing are covered in H.264 vs H.265 vs AV1; the generic encoder pipeline is covered in Video Encoder Architecture. Here everything is specific to HEVC.
The block structure
H.264 works in fixed 16 by 16 macroblocks. HEVC replaces them with the coding tree unit, or CTU, which the encoder configures at 16, 32 or 64 luma samples square. Large CTUs are the main reason HEVC does well on high resolutions: a flat sky in a 4K frame can be coded as a few 64 by 64 blocks with very little side information.
Each CTU is split by a quadtree into coding units, CUs, down to 8 by 8. The CU is where the encoder decides intra or inter prediction. Each CU is then divided into prediction units, PUs, which carry the prediction parameters. Inter CUs can be split into two rectangles, including asymmetric splits such as one quarter and three quarters, which fit object edges better than squares. Separately, the residual is split by a second quadtree into transform units, TUs, from 32 by 32 down to 4 by 4. Because the PU and TU trees are independent, a large prediction block can still use small transforms where the residual has detail.
Transforms are integer approximations of the discrete cosine transform at every size, with one exception: 4 by 4 luma blocks predicted intra use an integer discrete sine transform, which matches the residual shape of intra prediction better. Transform skip lets the encoder code small blocks without a transform at all, which helps screen content with sharp edges.
Intra and inter prediction
Intra prediction builds a block from the reconstructed samples above and to the left. HEVC has 35 intra modes: planar, DC and 33 angular directions, compared with nine for 4 by 4 blocks in H.264. The extra angles let a diagonal edge be predicted almost exactly. The encoder signals the mode relative to a short list of most probable modes derived from neighbours, so a common direction costs few bits.
Inter prediction copies a block from a reference picture, displaced by a motion vector at quarter-sample precision for luma. Fractional positions are interpolated with longer filters than H.264 used, seven or eight taps for luma and four for chroma, which gives sharper predictions. Motion vectors are coded in one of two ways. Merge mode copies the complete motion information of one neighbouring or temporally co-located candidate from a list of up to five, signalling only an index; it is how large areas of uniform motion become nearly free. Advanced motion vector prediction, AMVP, picks one of two predictor candidates and codes the difference from it. Skip is merge with no residual.
Entropy coding and in-loop filters
HEVC uses only CABAC, context-adaptive binary arithmetic coding; H.264's simpler CAVLC option is gone. CABAC adapts its probability models as it decodes, which is why it compresses well and also why it is serial: every bin depends on the state left by the previous one. HEVC's parallel tools exist largely to work around that.
Two filters run inside the loop, after reconstruction and before a picture is stored as a reference. The deblocking filter smooths block edges on an 8 by 8 grid, with a strength decided by the prediction modes, motion and quantisation on each side. It is designed so vertical and horizontal edges can each be filtered in parallel across the picture. Sample adaptive offset, SAO, is new in HEVC: per CTU, the encoder may classify samples by local edge shape or by intensity band and send a small offset for each class. It removes ringing and banding that the transform introduced. Because both filters are in the loop, the encoder must run them too, or its references drift from the decoder's.
Parallelism: slices, tiles and wavefronts
Slices are independently decodable groups of CTUs and exist mainly for error resilience and packetisation. Tiles divide the picture into a rectangular grid, and each tile resets prediction and CABAC state at its boundary, so tiles can be encoded and decoded on separate cores at a small cost in compression.
Wavefront parallel processing, WPP, keeps the picture whole. Each CTU row starts its CABAC state from the state after the second CTU of the row above, so row n can begin once row n minus one is two CTUs ahead. Decoders then process many rows in a diagonal wave. WPP usually costs less compression than tiles, and x265 uses it by default. Both are signalled in the picture parameter set, and hardware decoders on low-power devices may depend on them for high-resolution, high-frame-rate streams.
The bitstream: NAL units and parameter sets
An HEVC stream is a sequence of NAL units. Each starts with a two-byte header: one forbidden zero bit, a six-bit NAL unit type, a six-bit layer id and a three-bit temporal id plus one. The type tells you what the unit carries. Types 32, 33 and 34 are the video, sequence and picture parameter sets, VPS, SPS and PPS; types 39 and 40 are prefix and suffix SEI messages, which carry metadata such as HDR mastering display information; types 16 to 21 are intra random access point, IRAP, pictures.
IRAP pictures are where decoding can start. An IDR picture resets everything: nothing after it may reference anything before it. A CRA picture is more efficient for open-GOP encoding, because pictures that follow it in decode order but precede it in display order may still reference earlier pictures. Those leading pictures are labelled RASL, skipped when decoding starts at the CRA, or RADL, which are decodable. This matters for streaming: if a player starts at a CRA, it drops the RASL pictures, and a packager that splits segments at CRA pictures must handle that. Many streaming pipelines use closed GOPs with IDR pictures at segment boundaries to keep this simple.
NAL_NAMES = {19: "IDR_W_RADL", 20: "IDR_N_LP", 21: "CRA", 32: "VPS", 33: "SPS",
34: "PPS", 35: "AUD", 39: "SEI_PREFIX", 40: "SEI_SUFFIX"}
def nal_units(data: bytes):
"""Yield (type, layer_id, temporal_id, size) for an Annex B HEVC elementary stream."""
i, starts = 0, []
while (i := data.find(b"\x00\x00\x01", i)) != -1:
starts.append(i + 3)
i += 3
for n, s in enumerate(starts):
end = starts[n + 1] - 3 if n + 1 < len(starts) else len(data)
while end > s and data[end - 1] == 0: # trailing zero of a 4-byte start code
end -= 1
h0, h1 = data[s], data[s + 1]
nal_type = (h0 >> 1) & 0x3F
layer_id = ((h0 & 1) << 5) | (h1 >> 3)
tid = (h1 & 0x07) - 1
yield nal_type, layer_id, tid, end - s
# ffmpeg -i in.mp4 -c:v copy -bsf:v hevc_mp4toannexb -f hevc in.hevc
with open("in.hevc", "rb") as f:
for t, layer, tid, size in nal_units(f.read()):
if t >= 16: # IRAP and non-VCL units only, to keep the output short
print(NAL_NAMES.get(t, t), layer, tid, size)Running this on a stream answers practical questions quickly: are parameter sets repeated before every IDR, are segment boundaries IDR or CRA, are HDR SEI messages present?
Containers: hvc1 versus hev1
In MP4, HEVC tracks use one of two sample entry types. With hvc1 the parameter sets are stored in the sample entry, the configuration record in the file header. With hev1 they may also travel in-band, inside the samples. Both are valid, but Apple's players and HLS require hvc1, and a file tagged hev1 may play in one browser and show a black screen in Safari. When remuxing with ffmpeg, set the tag explicitly with -tag:v hvc1. For segmented streaming, also make sure each segment starts with an IRAP picture.
Profiles, tiers and levels
A profile limits the tools and formats: Main is 8-bit 4:2:0, Main 10 adds 10-bit samples and is the profile for HDR10 and HLG, and the range extensions add 4:2:2 and 4:4:4 for production. A level limits picture size, sample rate and buffer sizes, and a tier sets the bitrate ceiling within a level: Main tier for distribution, High tier for contribution and very high bitrates. A decoder that claims Main 10 at level 5.1 Main tier must decode any stream within those limits, so these numbers are a contract with the device.
| Level | Max luma picture size | Max luma samples per second | Max bitrate, Main / High tier | Typical use |
|---|---|---|---|---|
| 4.1 | 2,228,224 | 133,693,440 | 20 / 50 Mbit/s | 1080p at 60 fps |
| 5 | 8,912,896 | 267,386,880 | 25 / 100 Mbit/s | 2160p at 30 fps |
| 5.1 | 8,912,896 | 534,773,760 | 40 / 160 Mbit/s | 2160p at 60 fps |
Check the arithmetic for your top rendition: 3840 by 2160 at 60 frames per second is 497,664,000 luma samples per second, inside level 5.1 but well over level 5. Signal the lowest level your stream fits, because a device that supports only level 5 will refuse a stream marked 5.1 even if its content would have decoded.
Encoding in practice
x265 is the reference open-source software encoder. Its rate control and presets work like x264's: constant rate factor for quality-targeted files (the default CRF is 28), and capped CRF or two-pass for streaming, where a peak bitrate matters. Slower presets enable more partition and mode searches and give better compression per bit at higher CPU cost. A typical HDR10 top rung looks like this:
ffmpeg -i master.mov -map 0:v:0 -pix_fmt yuv420p10le -c:v libx265 -preset slow \
-x265-params "crf=20:vbv-maxrate=16000:vbv-bufsize=32000:keyint=96:min-keyint=96:\
open-gop=0:scenecut=0:colorprim=bt2020:transfer=smpte2084:colormatrix=bt2020nc:\
master-display=G(13250,34500)B(7500,3000)R(34000,16000)WP(15635,16450)L(10000000,1):\
max-cll=1000,400:hdr10=1:hdr10-opt=1:repeat-headers=1" \
-tag:v hvc1 -an top_2160p.mp4Each choice has a reason. A fixed keyframe interval with scene-cut detection off and closed GOPs gives IDR pictures exactly at four-second segment boundaries at 24 fps. VBV settings cap the peak so the rendition fits its ladder slot. The colour and mastering metadata come from the source's mastering report; copying example values from the internet produces wrong tone mapping. Repeat-headers puts parameter sets before every keyframe so the elementary stream can be cut anywhere a segment starts.
Hardware encoders, exposed in ffmpeg as hevc_nvenc, hevc_qsv, hevc_videotoolbox and others, run many times faster with modest power, and they are the right choice for live and for real-time transcoding. For video-on-demand, where an asset is encoded once and streamed many times, software x265 at a slow preset usually wins on bits, and per-title ladders, described in per-title encoding, multiply the gain. Measure both on your own content with VMAF at matched bitrates before deciding.
Failure modes
- Black screen on Apple devices. The track is tagged hev1. Remux with hvc1.
- Washed-out or too-dark HDR. Missing or wrong colour signalling or mastering metadata. Inspect the SPS colour fields and the SEI messages with the parser above or ffprobe.
- Device refuses the stream. The signalled level or tier exceeds what the decoder supports, or the stream is 10-bit on an 8-bit-only decoder. Check level arithmetic and probe capabilities on the device.
- Glitches after seeking or at segment starts. Segments start at CRA pictures with RASL pictures, or not at an IRAP at all. Use closed GOPs aligned to segment boundaries.
- Banding in dark gradients. 8-bit encode of 10-bit source, or SAO and deblocking too weak at low bitrates. Encode in Main 10 even for SDR when devices allow.
- Live encoder falls behind. Software preset too slow for the resolution. Use hardware encoding, or fewer partition searches.
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| CTU size | 64: best compression at high resolution | 32 or 16: lower latency, cheaper hardware |
| Parallelism | WPP: small compression loss | tiles: more parallel, larger loss at boundaries |
| GOP | closed, IDR at segments: simple packaging | open, CRA: a little more efficient |
| Encoder | x265 slow: fewest bits for VOD | hardware: real time, more bits |
| Bit depth | Main 10: less banding, HDR | Main: widest legacy decode |
If your audience's devices also decode AV1, compare it on your content; AV1, in depth covers its tools.
What to do next
- Extract an elementary stream from one of your files and run the NAL parser: confirm parameter sets, IRAP types at segment starts, and HDR SEI where expected.
- Check every HEVC MP4 you ship is tagged hvc1.
- Compute the luma sample rate of your top rendition and signal the lowest level that fits.
- Encode three test titles with x265 at two presets and with your hardware encoder, and compare VMAF at matched bitrates.
- Fix keyframe interval, closed GOPs and VBV caps to your segment length and ladder.
- Test playback on your oldest supported devices, especially for Main 10 and level 5.1.