H.265, also called High Efficiency Video Coding or HEVC, is the codec behind most 4K and HDR delivery of the last decade: broadcast UHD, Blu-ray UHD, much of streaming's HDR catalogue and the default recording format on many phones. It keeps the same hybrid design as H.264 (predict a block, transform the difference, entropy code the result) but replaces almost every tool with a more flexible one, and it adds structure for parallel decoding.

This article explains HEVC from the bitstream up so you can make informed encoding and packaging decisions and debug the problems that actually happen in production. The comparison with H.264 and AV1, codec strings and licensing are covered in H.264 vs H.265 vs AV1; the generic encoder pipeline is covered in Video Encoder Architecture. Here everything is specific to HEVC.

Advertisement

The block structure

H.264 works in fixed 16 by 16 macroblocks. HEVC replaces them with the coding tree unit, or CTU, which the encoder configures at 16, 32 or 64 luma samples square. Large CTUs are the main reason HEVC does well on high resolutions: a flat sky in a 4K frame can be coded as a few 64 by 64 blocks with very little side information.

Each CTU is split by a quadtree into coding units, CUs, down to 8 by 8. The CU is where the encoder decides intra or inter prediction. Each CU is then divided into prediction units, PUs, which carry the prediction parameters. Inter CUs can be split into two rectangles, including asymmetric splits such as one quarter and three quarters, which fit object edges better than squares. Separately, the residual is split by a second quadtree into transform units, TUs, from 32 by 32 down to 4 by 4. Because the PU and TU trees are independent, a large prediction block can still use small transforms where the residual has detail.

Transforms are integer approximations of the discrete cosine transform at every size, with one exception: 4 by 4 luma blocks predicted intra use an integer discrete sine transform, which matches the residual shape of intra prediction better. Transform skip lets the encoder code small blocks without a transform at all, which helps screen content with sharp edges.

HEVC block hierarchy and decoding loopCTU16, 32 or 64 lumaCU quadtreedown to 8x8PUprediction shapeTU quadtree4x4 to 32x32NAL unitsVPS, SPS, PPS, slicesCABAC decodemodes, MVs, coefficientsInverse transformDCT-like, DST 4x4Add predictionintra or motion comp.Deblocking filter8x8 edge gridSAOedge or band offsetsDecoded picturebufferReference pictures feed motion compensation for later framesthe encoder runs this same loop internally so its references match the decoder's
Top: CTUs split into coding units, prediction units and transform units. Bottom: the decoding loop, with deblocking and SAO applied before a picture becomes a reference.

Intra and inter prediction

Intra prediction builds a block from the reconstructed samples above and to the left. HEVC has 35 intra modes: planar, DC and 33 angular directions, compared with nine for 4 by 4 blocks in H.264. The extra angles let a diagonal edge be predicted almost exactly. The encoder signals the mode relative to a short list of most probable modes derived from neighbours, so a common direction costs few bits.

Inter prediction copies a block from a reference picture, displaced by a motion vector at quarter-sample precision for luma. Fractional positions are interpolated with longer filters than H.264 used, seven or eight taps for luma and four for chroma, which gives sharper predictions. Motion vectors are coded in one of two ways. Merge mode copies the complete motion information of one neighbouring or temporally co-located candidate from a list of up to five, signalling only an index; it is how large areas of uniform motion become nearly free. Advanced motion vector prediction, AMVP, picks one of two predictor candidates and codes the difference from it. Skip is merge with no residual.

Advertisement

Entropy coding and in-loop filters

HEVC uses only CABAC, context-adaptive binary arithmetic coding; H.264's simpler CAVLC option is gone. CABAC adapts its probability models as it decodes, which is why it compresses well and also why it is serial: every bin depends on the state left by the previous one. HEVC's parallel tools exist largely to work around that.

Two filters run inside the loop, after reconstruction and before a picture is stored as a reference. The deblocking filter smooths block edges on an 8 by 8 grid, with a strength decided by the prediction modes, motion and quantisation on each side. It is designed so vertical and horizontal edges can each be filtered in parallel across the picture. Sample adaptive offset, SAO, is new in HEVC: per CTU, the encoder may classify samples by local edge shape or by intensity band and send a small offset for each class. It removes ringing and banding that the transform introduced. Because both filters are in the loop, the encoder must run them too, or its references drift from the decoder's.

Parallelism: slices, tiles and wavefronts

Slices are independently decodable groups of CTUs and exist mainly for error resilience and packetisation. Tiles divide the picture into a rectangular grid, and each tile resets prediction and CABAC state at its boundary, so tiles can be encoded and decoded on separate cores at a small cost in compression.

Wavefront parallel processing, WPP, keeps the picture whole. Each CTU row starts its CABAC state from the state after the second CTU of the row above, so row n can begin once row n minus one is two CTUs ahead. Decoders then process many rows in a diagonal wave. WPP usually costs less compression than tiles, and x265 uses it by default. Both are signalled in the picture parameter set, and hardware decoders on low-power devices may depend on them for high-resolution, high-frame-rate streams.

The bitstream: NAL units and parameter sets

An HEVC stream is a sequence of NAL units. Each starts with a two-byte header: one forbidden zero bit, a six-bit NAL unit type, a six-bit layer id and a three-bit temporal id plus one. The type tells you what the unit carries. Types 32, 33 and 34 are the video, sequence and picture parameter sets, VPS, SPS and PPS; types 39 and 40 are prefix and suffix SEI messages, which carry metadata such as HDR mastering display information; types 16 to 21 are intra random access point, IRAP, pictures.

IRAP pictures are where decoding can start. An IDR picture resets everything: nothing after it may reference anything before it. A CRA picture is more efficient for open-GOP encoding, because pictures that follow it in decode order but precede it in display order may still reference earlier pictures. Those leading pictures are labelled RASL, skipped when decoding starts at the CRA, or RADL, which are decodable. This matters for streaming: if a player starts at a CRA, it drops the RASL pictures, and a packager that splits segments at CRA pictures must handle that. Many streaming pipelines use closed GOPs with IDR pictures at segment boundaries to keep this simple.

NAL_NAMES = {19: "IDR_W_RADL", 20: "IDR_N_LP", 21: "CRA", 32: "VPS", 33: "SPS",
             34: "PPS", 35: "AUD", 39: "SEI_PREFIX", 40: "SEI_SUFFIX"}

def nal_units(data: bytes):
    """Yield (type, layer_id, temporal_id, size) for an Annex B HEVC elementary stream."""
    i, starts = 0, []
    while (i := data.find(b"\x00\x00\x01", i)) != -1:
        starts.append(i + 3)
        i += 3
    for n, s in enumerate(starts):
        end = starts[n + 1] - 3 if n + 1 < len(starts) else len(data)
        while end > s and data[end - 1] == 0:      # trailing zero of a 4-byte start code
            end -= 1
        h0, h1 = data[s], data[s + 1]
        nal_type = (h0 >> 1) & 0x3F
        layer_id = ((h0 & 1) << 5) | (h1 >> 3)
        tid = (h1 & 0x07) - 1
        yield nal_type, layer_id, tid, end - s

# ffmpeg -i in.mp4 -c:v copy -bsf:v hevc_mp4toannexb -f hevc in.hevc
with open("in.hevc", "rb") as f:
    for t, layer, tid, size in nal_units(f.read()):
        if t >= 16:   # IRAP and non-VCL units only, to keep the output short
            print(NAL_NAMES.get(t, t), layer, tid, size)

Running this on a stream answers practical questions quickly: are parameter sets repeated before every IDR, are segment boundaries IDR or CRA, are HDR SEI messages present?

Containers: hvc1 versus hev1

In MP4, HEVC tracks use one of two sample entry types. With hvc1 the parameter sets are stored in the sample entry, the configuration record in the file header. With hev1 they may also travel in-band, inside the samples. Both are valid, but Apple's players and HLS require hvc1, and a file tagged hev1 may play in one browser and show a black screen in Safari. When remuxing with ffmpeg, set the tag explicitly with -tag:v hvc1. For segmented streaming, also make sure each segment starts with an IRAP picture.

Profiles, tiers and levels

A profile limits the tools and formats: Main is 8-bit 4:2:0, Main 10 adds 10-bit samples and is the profile for HDR10 and HLG, and the range extensions add 4:2:2 and 4:4:4 for production. A level limits picture size, sample rate and buffer sizes, and a tier sets the bitrate ceiling within a level: Main tier for distribution, High tier for contribution and very high bitrates. A decoder that claims Main 10 at level 5.1 Main tier must decode any stream within those limits, so these numbers are a contract with the device.

LevelMax luma picture sizeMax luma samples per secondMax bitrate, Main / High tierTypical use
4.12,228,224133,693,44020 / 50 Mbit/s1080p at 60 fps
58,912,896267,386,88025 / 100 Mbit/s2160p at 30 fps
5.18,912,896534,773,76040 / 160 Mbit/s2160p at 60 fps

Check the arithmetic for your top rendition: 3840 by 2160 at 60 frames per second is 497,664,000 luma samples per second, inside level 5.1 but well over level 5. Signal the lowest level your stream fits, because a device that supports only level 5 will refuse a stream marked 5.1 even if its content would have decoded.

Encoding in practice

x265 is the reference open-source software encoder. Its rate control and presets work like x264's: constant rate factor for quality-targeted files (the default CRF is 28), and capped CRF or two-pass for streaming, where a peak bitrate matters. Slower presets enable more partition and mode searches and give better compression per bit at higher CPU cost. A typical HDR10 top rung looks like this:

ffmpeg -i master.mov -map 0:v:0 -pix_fmt yuv420p10le -c:v libx265 -preset slow \
  -x265-params "crf=20:vbv-maxrate=16000:vbv-bufsize=32000:keyint=96:min-keyint=96:\
open-gop=0:scenecut=0:colorprim=bt2020:transfer=smpte2084:colormatrix=bt2020nc:\
master-display=G(13250,34500)B(7500,3000)R(34000,16000)WP(15635,16450)L(10000000,1):\
max-cll=1000,400:hdr10=1:hdr10-opt=1:repeat-headers=1" \
  -tag:v hvc1 -an top_2160p.mp4

Each choice has a reason. A fixed keyframe interval with scene-cut detection off and closed GOPs gives IDR pictures exactly at four-second segment boundaries at 24 fps. VBV settings cap the peak so the rendition fits its ladder slot. The colour and mastering metadata come from the source's mastering report; copying example values from the internet produces wrong tone mapping. Repeat-headers puts parameter sets before every keyframe so the elementary stream can be cut anywhere a segment starts.

Hardware encoders, exposed in ffmpeg as hevc_nvenc, hevc_qsv, hevc_videotoolbox and others, run many times faster with modest power, and they are the right choice for live and for real-time transcoding. For video-on-demand, where an asset is encoded once and streamed many times, software x265 at a slow preset usually wins on bits, and per-title ladders, described in per-title encoding, multiply the gain. Measure both on your own content with VMAF at matched bitrates before deciding.

Failure modes

  • Black screen on Apple devices. The track is tagged hev1. Remux with hvc1.
  • Washed-out or too-dark HDR. Missing or wrong colour signalling or mastering metadata. Inspect the SPS colour fields and the SEI messages with the parser above or ffprobe.
  • Device refuses the stream. The signalled level or tier exceeds what the decoder supports, or the stream is 10-bit on an 8-bit-only decoder. Check level arithmetic and probe capabilities on the device.
  • Glitches after seeking or at segment starts. Segments start at CRA pictures with RASL pictures, or not at an IRAP at all. Use closed GOPs aligned to segment boundaries.
  • Banding in dark gradients. 8-bit encode of 10-bit source, or SAO and deblocking too weak at low bitrates. Encode in Main 10 even for SDR when devices allow.
  • Live encoder falls behind. Software preset too slow for the resolution. Use hardware encoding, or fewer partition searches.

Trade-offs

DecisionOption AOption B
CTU size64: best compression at high resolution32 or 16: lower latency, cheaper hardware
ParallelismWPP: small compression losstiles: more parallel, larger loss at boundaries
GOPclosed, IDR at segments: simple packagingopen, CRA: a little more efficient
Encoderx265 slow: fewest bits for VODhardware: real time, more bits
Bit depthMain 10: less banding, HDRMain: widest legacy decode

If your audience's devices also decode AV1, compare it on your content; AV1, in depth covers its tools.

What to do next

  1. Extract an elementary stream from one of your files and run the NAL parser: confirm parameter sets, IRAP types at segment starts, and HDR SEI where expected.
  2. Check every HEVC MP4 you ship is tagged hvc1.
  3. Compute the luma sample rate of your top rendition and signal the lowest level that fits.
  4. Encode three test titles with x265 at two presets and with your hardware encoder, and compare VMAF at matched bitrates.
  5. Fix keyframe interval, closed GOPs and VBV caps to your segment length and ladder.
  6. Test playback on your oldest supported devices, especially for Main 10 and level 5.1.
Key takeaway: HEVC improves on H.264 with large, flexible coding tree units, independent prediction and transform trees, 35 intra modes, merge and AMVP motion coding, and the SAO filter, while tiles and wavefronts restore the parallelism that CABAC takes away. In production, most problems are not in the coding tools but at the edges: the hvc1 tag, the level you signal, IRAP types at segment boundaries, and HDR metadata. Inspect the NAL units, signal honestly and measure encoders on your own content.