MPEG-DASH (ISO/IEC 23009-1) is a description language more than a protocol. The server does almost nothing: it publishes a manifest called the Media Presentation Description, or MPD, plus a set of media segments as ordinary files over HTTP. Every decision about what to fetch, when, and at which quality is made by the player. That division of labour is why DASH scales through ordinary CDNs, and it is also why most DASH bugs are timing bugs: the player has to compute, from a few attributes in an XML file and its own clock, which segment exists right now and what its URL is.
This article covers the MPD data model, segment addressing, the live timing model and clock sync, multi-period content, the player pipeline and low-latency DASH, with a live-edge calculation worked by hand and the failures seen in production. If you are choosing between DASH and HLS, start with HLS versus DASH; this page assumes you have chosen DASH, or must support it, and want to know how it really works.
The architecture in one picture
On the server side an encoder produces a bitrate ladder, several encodings of the same content at different resolutions and bitrates, with keyframes aligned across all of them. A packager cuts each encoding into segments, usually fragmented MP4 (see CMAF), writes one initialization segment per rendition carrying the codec configuration, and writes the MPD that describes it all.
On the client side a player fetches the MPD, builds an internal timeline, picks a rendition, generates segment URLs, downloads segments and appends them to the browser's decoder through Media Source Extensions (MSE). It switches rendition at segment boundaries.
The MPD data model
The MPD is an XML tree with four levels that matter. The root <MPD> carries global attributes: whether the presentation is static (on demand) or dynamic (live), the profiles it conforms to, and for live content the timing anchors discussed below. Beneath it, one or more <Period> elements divide the presentation in time. Each Period is a stretch where the set of available tracks is constant; an ad break or a programme boundary is typically a new Period.
Inside a Period, each <AdaptationSet> groups tracks the player may switch between seamlessly: all the video renditions, or all the English audio bitrates. Inside an AdaptationSet, each <Representation> is one encoding with a declared bandwidth in bits per second, a codec string and, for video, a resolution. The bandwidth attribute is what the ABR logic compares with measured throughput, so it must be honest: it is defined as the rate needed to play the Representation continuously given the declared minBufferTime, so in practice it should reflect peaks, not the average.
Attributes inherit downward. A <SegmentTemplate> declared on the AdaptationSet applies to every Representation inside it unless a Representation overrides it, which keeps manifests for large ladders short.
Four ways to address segments
DASH offers four ways to tell the player each segment's URL and position in time; choosing among them is the packager's main design decision.
| Mechanism | How URLs and times are known | Best for | Cost |
|---|---|---|---|
| SegmentBase | One file per Representation, indexed by a sidx box | On demand | Range requests plus an index fetch |
| SegmentList | Every segment URL listed explicitly | Irregular naming | MPD grows with content length |
| SegmentTemplate with $Number$ | URL from a counter; constant nominal duration | Live, fixed segment length | Drift if durations vary |
| SegmentTemplate with SegmentTimeline and $Time$ | Exact start and duration of every segment, run-length coded | Variable durations, ad splicing | MPD refetch to learn new segments |
With $Number$ addressing, the player computes the segment number from the clock, so it never needs a fresh manifest to find the next segment. But it assumes every segment lasts exactly duration / timescale seconds. An encoder that inserts an extra keyframe at a scene cut breaks that, and the player drifts away from the content.
With a <SegmentTimeline>, each run of segments is described by an <S> element with a start time t, a duration d and a repeat count r, all in units of the timescale. URLs use $Time$, the exact media start time, so durations can vary freely. The price is that the player learns about new segments only by refetching the MPD, which is why live timelines go with a short minimumUpdatePeriod.
The live timing model
For a dynamic MPD, everything hangs on availabilityStartTime (AST), a wall-clock instant. A Period's start is an offset from AST, and a segment's media time maps to wall-clock time through the Period start and the timescale. A segment becomes available on the server when its end time has passed, because the packager cannot publish a segment before it has finished encoding it. It stays available while it is inside timeShiftBufferDepth, the DVR window behind the live edge.
Three other attributes control player behaviour. minimumUpdatePeriod says how often the MPD may change and therefore how often to refetch it. suggestedPresentationDelay says how far behind the live edge the player should play, leaving slack for publication, CDN propagation and clock error. publishTime tells the player which version of the manifest it has, so it can ignore a stale copy served by a lagging cache.
None of this works if the player's clock is wrong. A player whose clock is two seconds fast requests segments before they exist. The <UTCTiming> element points the player at a time source, such as an HTTP endpoint returning ISO time; the player measures the offset once and applies it everywhere. Always publish one.
A worked example: finding the live edge
Here is a small live MPD (audio omitted) with a three-rung ladder, two-second segments and $Number$ addressing.
<MPD xmlns="urn:mpeg:dash:schema:mpd:2011" type="dynamic"
profiles="urn:mpeg:dash:profile:isoff-live:2011"
availabilityStartTime="2026-10-01T12:00:00Z"
publishTime="2026-10-01T12:41:06Z"
minimumUpdatePeriod="PT2S"
timeShiftBufferDepth="PT2M"
suggestedPresentationDelay="PT6S"
minBufferTime="PT2S">
<Period id="p0" start="PT0S">
<AdaptationSet contentType="video" mimeType="video/mp4" segmentAlignment="true">
<SegmentTemplate timescale="90000" duration="180000" startNumber="1"
initialization="video/$RepresentationID$/init.mp4"
media="video/$RepresentationID$/$Number$.m4s"/>
<Representation id="v1080" bandwidth="6000000" width="1920" height="1080" codecs="avc1.640028"/>
<Representation id="v720" bandwidth="3000000" width="1280" height="720" codecs="avc1.64001f"/>
<Representation id="v360" bandwidth="800000" width="640" height="360" codecs="avc1.64001e"/>
</AdaptationSet>
</Period>
<UTCTiming schemeIdUri="urn:mpeg:dash:utc:http-iso:2014" value="https://time.example.com/now"/>
</MPD>The video timescale is 90000 ticks per second and the duration is 180000 ticks, so each segment is two seconds. Suppose the player's synchronised clock reads 12:41:07.3. Elapsed time since AST is 2467.3 seconds. Segment index k (counting from zero) covers the interval from 2k to 2k+2 seconds after AST. Dividing 2467.3 by two gives 1233.65, so index 1233 is still being encoded and the newest complete segment is index 1232, which with startNumber=1 is segment number 1233. The code below makes the off-by-one explicit, because this is exactly where real players go wrong.
from datetime import datetime, timedelta, timezone
def live_edge_number(ast, now, seg_dur_s, start_number=1, ato_s=0.0): # ato_s: availabilityTimeOffset
# Segment k (0-based) covers [k*d, (k+1)*d] after AST; available once it ends.
elapsed = (now - ast).total_seconds() + ato_s
k = int(elapsed // seg_dur_s) - 1 # last fully available
return start_number + k
def start_number_for_playback(ast, now, seg_dur_s, spd_s, start_number=1):
target = now - timedelta(seconds=spd_s)
k = int((target - ast).total_seconds() // seg_dur_s)
return start_number + k
ast = datetime(2026, 10, 1, 12, 0, 0, tzinfo=timezone.utc)
now = datetime(2026, 10, 1, 12, 41, 7, 300000, tzinfo=timezone.utc)
print(live_edge_number(ast, now, 2.0)) # 1233
print(start_number_for_playback(ast, now, 2.0, 6.0)) # 1231 -> URL video/v720/1231.m4sRather than joining at the edge, the player honours the six-second suggested delay: it targets 12:41:01.3, which falls in index 1230, number 1231, and requests video/v720/1231.m4s after the init segment. Joining at the edge would leave no buffer. An error of one segment is visible either way: too far ahead gives 404s at startup, too far behind adds latency. Unit-test this against your packager's output.
Multi-period presentations
A new Period starts whenever the track set or the timeline changes: a server-side inserted ad with a different encoding ladder, a switch between programmes, or a re-encode after an encoder failover. Each Period carries its own AdaptationSets, and the player must flush or reconfigure its pipeline across the boundary if codecs differ. Within MSE this usually means a timestampOffset change so the new Period's media times land at the right place on the media element's timeline, and sometimes a changeType() call on the SourceBuffer when the codec changes.
Period boundaries are where players glitch: small gaps stall some decoders and overlaps drop frames. Make Period durations match the media exactly, and keep ad ladders close to the content ladder so the decoder configuration can stay.
The player pipeline over Media Source Extensions
In browsers, DASH is implemented in JavaScript on top of MSE. The player creates a MediaSource, attaches it to a video element, and adds one SourceBuffer per media type. It appends the initialization segment for the chosen Representation first, then media segments in order. On a switch it appends the new Representation's init segment first; aligned keyframes keep playback seamless.
// Minimal DASH-style loop over Media Source Extensions (no ABR, one rendition).
const video = document.querySelector('video');
const ms = new MediaSource();
video.src = URL.createObjectURL(ms);
ms.addEventListener('sourceopen', async () => {
const sb = ms.addSourceBuffer('video/mp4; codecs="avc1.64001f"');
const append = buf => new Promise((ok, fail) => {
sb.addEventListener('updateend', ok, { once: true });
sb.addEventListener('error', fail, { once: true });
sb.appendBuffer(buf);
});
await append(await (await fetch('video/v720/init.mp4')).arrayBuffer());
let n = startNumber; // computed from the MPD as above
while (true) {
const res = await fetch(`video/v720/${n}.m4s`);
if (res.status === 404) { await sleep(500); continue; } // not yet available
await append(await res.arrayBuffer());
// (production code also evicts media behind the playhead with sb.remove)
n += 1;
}
});appendBuffer is asynchronous and a SourceBuffer accepts only one operation at a time, so every operation must wait for updateend. Browsers also impose a quota on buffered data and throw QuotaExceededError when it is exceeded, so the player must evict media behind the playhead. Production players such as dash.js wrap all this plus ABR and recovery; in dash.js the entry point is dashjs.MediaPlayer().create() followed by initialize(videoElement, mpdUrl, autoplay). Settings names change across major versions, so check the docs for yours.
ABR rules (throughput-based, buffer-based or hybrid) are covered in ABR selection algorithms. What DASH contributes is the input: honest bandwidth values and segments whose sizes are predictable from them.
Low-latency DASH
Standard live DASH has a latency floor of a few segment durations, because a segment is available only once complete. Low-latency DASH removes that constraint. The packager writes each segment as a series of small CMAF chunks and the origin serves the segment with HTTP chunked transfer while it is still being produced. The MPD advertises this with availabilityTimeOffset, which tells the player it may request a segment that many seconds before its nominal availability, together with availabilityTimeComplete="false".
This changes the client too. Throughput measured over a chunked download is bounded by the encoder's real-time rate, not by the network, so throughput-based ABR underestimates bandwidth and needs a different estimator. Players also speed playback slightly to catch up when they drift behind target. A <ServiceDescription> element can carry the target latency and allowed playback rates from the operator to the player. The latency budget across encoder, packager, CDN and player is analysed in low-latency streaming.
Failure modes in production
- Clock skew. A player with no UTCTiming, or one that ignores it, requests segments early (404 storms at startup) or late (excess latency). Fix: always publish UTCTiming and alert on 404 rates for segment numbers ahead of the edge.
- Nominal duration drift. With
$Number$addressing, encoders that emit a few frames more or less per segment slowly push media time away from computed time. After hours the player is off by a segment. Fix: enforce fixed GOPs at the encoder or switch to SegmentTimeline. - Stale MPD from cache. A CDN that caches the live MPD longer than
minimumUpdatePeriodhides new segments in a SegmentTimeline stream. Fix: short TTLs on the manifest only, and players that comparepublishTime. - Misaligned keyframes across the ladder. Switching causes a visible jump or a decode error. Fix: verify alignment in packager QA by comparing segment start times across Representations.
- Dishonest bandwidth. A Representation declared at 3 Mbps whose segments peak at 6 Mbps causes stalls just after an up-switch. Fix: derive
bandwidthfrom measured peaks over theminBufferTimewindow.
Operating a DASH service
Validate every manifest against the schema and the DASH Industry Forum conformance tools; most interoperability bugs are manifests one player tolerates and another rejects. Set cache rules by object type: segments are immutable and can be cached for their full DVR lifetime, init segments for the life of the stream, and the live MPD for at most half its update period. Monitor from the client: startup time, rebuffer ratio, average bitrate, switches per minute, live latency, and 404 counts per segment number relative to the live edge. Inject faults in testing: delayed publication, a skewed clock, a dropped segment. Encryption and licence acquisition are their own topic; see video DRM.
Trade-offs
$Number$ templates give the smallest, most cache-friendly manifests and need fixed segment durations. SegmentTimeline handles any timing but needs frequent manifest refreshes. Short segments cut latency but multiply requests and cost compression, since each starts with a keyframe. And DASH alone does not reach every device; most services package CMAF once and emit both an MPD and an HLS playlist.
What to do next
- Pick an addressing mode deliberately:
$Number$only if your encoder guarantees fixed GOPs, SegmentTimeline otherwise. - Add a UTCTiming element pointing at a time endpoint you control, and confirm your player uses it.
- Write unit tests for live-edge and start-number arithmetic against real manifests from your packager.
- Run every manifest through conformance validation in CI, and check keyframe alignment across the ladder.
- Set CDN TTLs per object type and verify the live MPD never outlives half its update period.
- Instrument the player for startup time, rebuffering, latency and 404s by segment number, and test with injected clock skew and publication delay.