Most adaptive streams, whether delivered as HLS, DASH or CMAF, reach the decoder as fragmented MP4. When it goes wrong the symptoms are vague: a stall at a segment boundary, audio drifting after twenty minutes, a seek that lands on a grey frame. They nearly always come down to a handful of fields inside a few small boxes.
This article works through the format byte by byte: what a box is, how fragmentation moves the sample tables into each fragment, how a parser finds a sample's bytes and timestamp, and which flags mark a keyframe, with a parser you can run. The packaging-level view is covered in CMAF.
Why fragment an MP4 at all
A classic MP4 has an mdat box of compressed samples and a moov index: per track, a sample table giving each sample's size, duration, offset and whether it is a sync sample (keyframe). Without it a player cannot tell where one frame ends and the next begins.
That design breaks streaming. The index cannot be written until the last sample is known, and a live encoder never reaches the last sample. For a long film the index is also large and must be downloaded before the first frame.
Fragmentation, defined in ISO/IEC 14496-12 (the ISO Base Media File Format), splits the index. The moov box keeps only what never changes: tracks, timescales and codec configuration, with empty sample tables. Per-sample information moves into movie fragments, each a moof box describing a short run of samples followed by an mdat holding them. A writer can emit a fragment as soon as its samples exist, and a reader can start at any fragment once it has the initialisation data.
Boxes from first principles
Everything in the file is a box (the older specifications call it an atom). A box starts with an 8-byte header: a 32-bit big-endian size that counts the whole box including the header, then a four-character type code such as moof. Two size values are special. A size of 1 means a 64-bit largesize follows the type, which makes the header 16 bytes. A size of 0 means the box runs to the end of the file, which is only legal for the last box.
Many boxes are "full boxes": after the header they carry an 8-bit version and 24 bits of flags. The version usually decides whether time fields are 32 or 64 bits wide. The flags decide which optional fields are present. A parser that reads a field the flags say is absent reads every later field from the wrong offset.
Boxes such as moov, moof and traf are pure containers of child boxes; others are leaves with a fixed binary layout. A parser skips any box it does not understand by jumping ahead size bytes, which is why players tolerate vendor uuid boxes.
The layout of a fragmented stream
A fragmented stream has two kinds of segment. The initialisation segment is ftyp followed by moov. The media segments each contain one or more moof/mdat pairs. A segment may be preceded by a styp (segment type) box and a sidx (segment index) box. A single-file fragmented MP4 is just the init segment followed by all the fragments, optionally ending with an mfra random-access index.
HLS references the init segment with #EXT-X-MAP; DASH with the initialization attribute. A browser player must append it to a SourceBuffer before any media segment, and again on every rendition switch.
The init segment: empty tables and trex defaults
Each trak still leads down to stbl, but stts, stsc, stsz and stco have zero entries. What matters is stsd, which holds codec configuration: avc1 with avcC (H.264 SPS and PPS), hvc1 with hvcC, or mp4a with esds for AAC. The mdhd box gives the track's timescale, the number of ticks per second used by every timestamp in that track.
The mvex box marks the file as fragmented. It holds one trex per track with default sample description index, duration, size and flags, and optionally an mehd with the overall duration, which live streams omit.
Inside a movie fragment
A moof starts with mfhd, which holds a sequence number that should increase from one fragment to the next. Then comes one traf per track that has samples in this fragment. Each traf holds three boxes: tfhd for the fragment header, tfdt for the decode time, and one or more trun (track run) boxes listing the samples.
| Box | Flag | Meaning |
|---|---|---|
| tfhd | 0x000001 | base-data-offset present: an explicit 64-bit file offset (avoid it in segmented streams) |
| tfhd | 0x000002 | sample-description-index present |
| tfhd | 0x000008 / 0x000010 / 0x000020 | default sample duration / size / flags present for this fragment |
| tfhd | 0x010000 | duration-is-empty: the fragment has no samples |
| tfhd | 0x020000 | default-base-is-moof: offsets count from the start of this moof |
| trun | 0x000001 | data_offset present (signed 32-bit) |
| trun | 0x000004 | first-sample-flags present: overrides flags for sample 0 only |
| trun | 0x000100 / 0x000200 / 0x000400 | per-sample duration / size / flags present |
| trun | 0x000800 | per-sample composition time offset present |
Values cascade: a per-sample field in trun wins, then the tfhd default, then the trex default. A typical video fragment sets a default duration in tfhd, per-sample sizes and composition offsets in trun, and uses first-sample-flags to mark the opening keyframe.
Finding a sample's bytes
A parser needs two numbers to find sample bytes: the base data offset and the run's data_offset. If tfhd sets default-base-is-moof, the base is the file position of the first byte of the enclosing moof. The trun data_offset is added to that base to find the first sample of the run, and later samples follow contiguously, each starting where the previous one's size ended.
Worked example. A segment starts with a moof of 112 bytes at position 0, so the mdat header occupies bytes 112 to 119 and the payload starts at byte 120. The trun must carry data_offset = 120. Its samples of 100, 50 and 60 bytes occupy bytes 120 to 219, 220 to 269 and 270 to 329. If the packager later adds encryption metadata to moof without recomputing data_offset, the decoder receives garbage starting mid-NAL unit.
Default-base-is-moof makes each segment self-contained. An explicit base data offset refers to positions in the original file, which a segment served on its own no longer has. CMAF requires default-base-is-moof for that reason.
Timing: timescales, tfdt and composition offsets
Every duration and timestamp is an integer number of ticks in the track's timescale. Video at 30000/1001 frames per second (29.97) is commonly given a timescale of 30000, so each frame lasts exactly 1001 ticks. A two-second segment of 60 frames covers 60,060 ticks. A timescale of 1000 cannot represent 1001/30000 of a second exactly, and the rounding builds up into drift over a long stream.
The tfdt box carries baseMediaDecodeTime, the decode timestamp of the fragment's first sample. Use version 1 (64 bits): a 90 kHz clock overflows 32 bits after about 13 hours. A continuous stream satisfies one invariant: each fragment's tfdt equals the previous tfdt plus the previous fragment's total duration. In the example, consecutive values go 0, 60060, 120120.
With B-frames, decode order differs from display order. Presentation time is decode time plus the composition time offset in trun. Version 0 makes the offset unsigned, so packagers add an edit list (elst) to shift the start back. Version 1 allows negative offsets and needs no edit list. Browsers have differed in how they apply edit lists to fragmented files, so version 1 offsets remove a class of cross-platform start-time bugs.
sample_flags and keyframes
Each sample's flags are a 32-bit field. Reading from the top: 4 reserved bits, then is_leading (2 bits), sample_depends_on (2), sample_is_depended_on (2), sample_has_redundancy (2), sample_padding_value (3), sample_is_non_sync_sample (1 bit, at 0x00010000) and a 16-bit degradation_priority. Two values account for most real files. 0x02000000 means the sample depends on no other sample and is a sync sample: a keyframe. 0x01010000 means it depends on others and is not a sync sample.
The non-sync bit decides where segments start and seeks land. An IDR frame flagged non-sync breaks seeking; a non-IDR frame flagged sync makes seeks land on corrupted output until the next real keyframe.
Side boxes: sidx, emsg, prft and encryption
sidx: a segment index of subsegment sizes, durations and stream access points, so on-demand players can issue byte-range requests into a single file.emsg: an in-band event message placed beforemoof, used for ad markers (see SCTE-35 signalling) and other timed metadata.prft: pairs a wall-clock time with a media time, so low-latency players can measure how far behind live they are.senc,saiz,saio: Common Encryption per-sample IVs and subsample ranges. Thesaiooffset must also be recomputed whenevermoofchanges size. The licence side is covered in video DRM.
A working parser
The code below walks the box tree and decodes a trun with the cascade applied. It handles largesize, size 0 and signed composition offsets, and rejects boxes that overrun their parent.
import struct
CONTAINERS = {b"moov", b"trak", b"mdia", b"minf", b"stbl", b"mvex",
b"moof", b"traf", b"mfra", b"edts", b"dinf"}
def walk(buf, start=0, end=None, depth=0):
"""Yield (depth, type, offset, size) for every box, descending into containers."""
end = len(buf) if end is None else end
pos = start
while pos + 8 <= end:
size, btype = struct.unpack_from(">I4s", buf, pos)
header = 8
if size == 1: # 64-bit largesize follows the type
size = struct.unpack_from(">Q", buf, pos + 8)[0]
header = 16
elif size == 0: # runs to the end of the enclosing space
size = end - pos
if size < header or pos + size > end:
raise ValueError(f"bad box {btype!r} at {pos}: size {size}")
yield depth, btype, pos, size
if btype in CONTAINERS:
yield from walk(buf, pos + header, pos + size, depth + 1)
pos += size
def parse_trun(buf, pos, default_duration, default_size, default_flags):
"""Return data_offset and [(duration, size, flags, cto)] for one trun box."""
version = buf[pos + 8]
flags = int.from_bytes(buf[pos + 9:pos + 12], "big")
count = struct.unpack_from(">I", buf, pos + 12)[0]
p = pos + 16
data_offset = first_flags = None
if flags & 0x000001:
data_offset = struct.unpack_from(">i", buf, p)[0]; p += 4
if flags & 0x000004:
first_flags = struct.unpack_from(">I", buf, p)[0]; p += 4
samples = []
for i in range(count):
dur, size, sflags, cto = default_duration, default_size, default_flags, 0
if flags & 0x000100: dur = struct.unpack_from(">I", buf, p)[0]; p += 4
if flags & 0x000200: size = struct.unpack_from(">I", buf, p)[0]; p += 4
if flags & 0x000400: sflags = struct.unpack_from(">I", buf, p)[0]; p += 4
if flags & 0x000800:
cto = struct.unpack_from(">i" if version == 1 else ">I", buf, p)[0]; p += 4
if i == 0 and first_flags is not None:
sflags = first_flags
samples.append((dur, size, sflags, cto))
return data_offset, samples
def is_sync(sample_flags):
return not (sample_flags >> 16) & 1 # sample_is_non_sync_samplePass in defaults from tfhd when its flags set them, otherwise from trex. Never hard-code them: a parser that assumes per-sample durations works on one packager and fails on the next.
Producing and inspecting fragments
FFmpeg's MP4 muxer writes fragmented output through -movflags. The combination frag_keyframe+empty_moov+default_base_moof starts a fragment at each keyframe, writes a sample-free init segment and sets default-base-is-moof; -frag_duration (microseconds) sets a target fragment length. Bento4's mp4dump prints the box tree with field values.
ffmpeg -i input.mp4 -c copy \
-movflags frag_keyframe+empty_moov+default_base_moof \
-f mp4 fragmented.mp4
mp4dump fragmented.mp4 | head -80
Failure modes
| Symptom | Likely cause | Check |
|---|---|---|
| Stall or decode error at every segment boundary | data_offset or saio not recomputed after the moof changed size | Verify that data_offset lands on the first mdat payload byte |
| A/V drift over long sessions | Timescale cannot represent the frame duration exactly; durations rounded | Sum of trun durations against the next tfdt |
| Gap or overlap warnings, buffered-range holes | tfdt not continuous across segments or renditions | tfdt[n+1] == tfdt[n] + duration[n] per track |
| Seek lands on garbage frames | sample_flags mark a non-IDR frame as sync | Cross-check sync flags against the NAL unit types |
| Playback fails after a rendition switch | New init segment not appended, or codec string mismatch | Append init on every switch; compare stsd with the manifest CODECS |
| Start time differs across browsers | Version 0 composition offsets plus an edit list | Use trun version 1 and drop elst |
Trade-offs in fragment design
Shorter fragments cut latency, because a player can start decoding a fragment as soon as it arrives. That is the basis of chunked CMAF in low-latency streaming. The cost is a moof header plus several bytes per sample in every fragment, and more request or chunk boundaries. Only chunks that begin with a keyframe can be joined or switched to.
Per-sample durations are worth their bytes only for variable frame rates; explicit per-sample flags are safest when the GOP is not fixed. If you package on the fly, as in just-in-time packaging, compute every offset after the moof is final.