Classic HLS was designed for reliability at scale, not for speed. A typical live stream uses 6-second segments, and the player starts at least three target durations behind the live edge, because the specification recommends it and because a segment cannot be listed until it is complete. Add encoding, packaging and CDN propagation and viewers routinely sit 20 to 30 seconds behind the event. For sports, betting, auctions and interactive shows, that is the difference between watching live and hearing your neighbour cheer first.
Low-Latency HLS keeps everything that made HLS scale (plain HTTP, CDN caching, adaptive bitrate, segment playlists) and changes how fresh media is published and discovered. This page explains each mechanism from the playlist up, checked against the current HLS second-edition draft (draft-pantos-hls-rfc8216bis-22, May 2026), then works through a latency budget and the operational traps. For the CMAF container that carries the parts, see CMAF architecture.
Where classic HLS loses its seconds
Latency in segment-based streaming is dominated by granularity. If segments are 6 seconds long, the newest media is invisible until its segment closes, so the first frame of a segment is already 6 seconds old when the segment appears in the playlist. The player then polls the playlist, typically once per target duration, adding up to another segment of delay. Finally it holds back several segments as a buffer against network jitter. Shrinking segments to one second helps, but each segment must start with a keyframe, so one-second segments mean one-second GOPs, worse compression, many more requests and playlists that change every second.
LL-HLS attacks each source of delay separately. Partial segments publish media in small pieces while keeping normal-sized segments and GOPs. Preload hints let the player request the next piece before it exists. Blocking playlist reload replaces polling with a request the server answers the moment new media appears. Delta updates and rendition reports keep frequent playlist reloads and bitrate switches cheap. The draft defines Low-Latency Mode as the combined use of partial segments, blocking reload and preload hinting.
Partial segments
A partial segment, or part, is a short slice of a parent segment: for fMP4 streams, one or more CMAF chunks (a moof plus mdat pair, explained in fragmented MP4 internals). The playlist declares a Part Target Duration with EXT-X-PART-INF:PART-TARGET and lists each part with EXT-X-PART, which requires URI and DURATION and may carry INDEPENDENT=YES when the part starts with an independent frame, BYTERANGE when parts are ranges of one file, and GAP=YES when a part is unavailable.
The draft is strict about part durations: each part MUST be no longer than the Part Target Duration and at least 85% of it, except parts marked INDEPENDENT or GAP. Players rely on this regularity to predict when the next part will exist. Once a parent segment completes, it is listed as an ordinary segment after its parts, so a player that knows nothing about LL-HLS simply ignores the part tags and plays the same stream at classic latency.
#EXTM3U
#EXT-X-TARGETDURATION:4
#EXT-X-VERSION:9
#EXT-X-SERVER-CONTROL:CAN-BLOCK-RELOAD=YES,PART-HOLD-BACK=1.5,CAN-SKIP-UNTIL=24.0
#EXT-X-PART-INF:PART-TARGET=0.5
#EXT-X-MEDIA-SEQUENCE:266
#EXT-X-MAP:URI="init.mp4"
#EXTINF:4.0,
... (older segments; their parts age out)
seg270.m4s
#EXT-X-PART:DURATION=0.5,URI="seg271.0.m4s",INDEPENDENT=YES
#EXT-X-PART:DURATION=0.5,URI="seg271.1.m4s"
#EXT-X-PART:DURATION=0.5,URI="seg271.2.m4s"
#EXT-X-PRELOAD-HINT:TYPE=PART,URI="seg271.3.m4s"
#EXT-X-RENDITION-REPORT:URI="../1500k/live.m3u8",LAST-MSN=271,LAST-PART=2
#EXT-X-RENDITION-REPORT:URI="../800k/live.m3u8",LAST-MSN=271,LAST-PART=2
Hold-back: how close to the edge a player may sit
EXT-X-SERVER-CONTROL carries two distances from the end of the playlist. HOLD-BACK applies to classic playback and MUST be at least three times the Target Duration. PART-HOLD-BACK applies in Low-Latency Mode, is required whenever the playlist uses parts, MUST be at least twice the Part Target Duration and SHOULD be at least three times it. With 0.5-second parts, that means a player starts 1.0 to 1.5 seconds behind the newest part instead of 12 seconds behind with 4-second segments.
The draft is candid about the trade: shorter targets reduce latency but leave less buffer, handicap adaptation and raise request overhead, while longer hold-back reduces stalls at the cost of latency. Treat the part target and PART-HOLD-BACK as one tuning pair, chosen from measured network jitter, not as a race to the smallest number.
Preload hints: asking before it exists
A playlist with parts and no EXT-X-ENDLIST should end with EXT-X-PRELOAD-HINT:TYPE=PART naming the next part, optionally as a byte range. A hinted resource MUST be requestable as soon as the hint appears, and the origin holds that request open until the part is complete, then sends it at once. (A byte-range hint spanning several parts of one file is delivered part by part, each part whole.) That removes a full round trip per part: instead of waiting for the playlist to list part 271.3 and then requesting it, the player already has the request in flight.
One rule protects bitrate adaptation: the server MUST NOT send any bytes of a part until it can send the whole part at full link speed. Without it, a player would measure a 0.5-second part trickling out over 0.5 seconds as a slow network and downshift needlessly. If the packager's segmentation changes, for example an ad break cuts a segment short, the server may never publish a hinted resource, and players must tolerate that.
Blocking playlist reload
Polling a playlist every half second from millions of players is wasteful and still late by up to one poll interval. With CAN-BLOCK-RELOAD=YES advertised, the player asks for the playlist version it wants next, using delivery directives in the query string: _HLS_msn=M for a media sequence number and optionally _HLS_part=N for a part index. The server MUST hold the request until the playlist contains that segment or part, then return the entire playlist. A part index beyond the last part of the parent is treated as part 0 of the next segment.
The error rules matter for operations. If _HLS_msn is more than two beyond the last segment, or _HLS_part is further ahead than the Advance Part Limit (three divided by the part target when it is under one second, otherwise three), the server SHOULD return 400 immediately, because the client is confused rather than early. _HLS_part without _HLS_msn is also a 400. If the server cannot produce the requested version after blocking for more than three Target Durations, it SHOULD return 503, which usually means the encoder or packager has stalled.
# Simplified client loop in Low-Latency Mode
pl = fetch(media_uri) # first load: no directives
seek_to(pl.end - pl.part_hold_back)
while live:
msn, part = pl.next_part_after_last() # wraps to (msn + 1, 0)
if pl.preload_hint:
prefetch(pl.preload_hint.uri) # bytes arrive part-at-a-time
try:
pl = fetch(f"{media_uri}?_HLS_msn={msn}&_HLS_part={part}&_HLS_skip=YES",
timeout=3 * pl.target_duration + margin)
except HttpError as e:
if e.status == 400: pl = fetch(media_uri) # resync from scratch
elif e.status == 503: backoff_and_retry() # origin is stalled
merge_delta(pl) # apply EXT-X-SKIP against cached copy
schedule_new_parts(pl)
maybe_switch_rendition(pl.rendition_reports)
Delta updates and rendition reports
Reloading the playlist several times per second makes its size matter, especially with a long DVR window. A server that advertises CAN-SKIP-UNTIL (the Skip Boundary, which MUST be at least six times the Target Duration) responds to _HLS_skip=YES with a delta update: segments older than the boundary are replaced by one EXT-X-SKIP:SKIPPED-SEGMENTS=n tag, and the client splices the response onto its cached copy. _HLS_skip=v2 also drops old EXT-X-DATERANGE tags and lists recently removed ones in RECENTLY-REMOVED-DATERANGES. Servers ignore the directive if they did not advertise support or if the playlist has ended.
Rendition reports solve a quieter problem. When the ABR algorithm switches bitrate, the player needs the other rendition's playlist at the right point. EXT-X-RENDITION-REPORT lists, for each other rendition, its URI, LAST-MSN and LAST-PART, so the player can issue a blocking request for exactly the next part of the new rendition rather than loading a possibly stale playlist and probing. The draft's Low-Latency Server Configuration Profile (Appendix B.1) requires a report for every other rendition and also requires delivery over HTTP/2 or HTTP/3. How players choose when to switch is covered in ABR selection algorithms.
Worked example: a 3-second latency budget
Target: glass-to-glass latency around 3 seconds for a sports stream, with 4-second segments and 0.5-second parts. The numbers for encoding and delivery below are assumptions to replace with your own measurements.
| Stage | Budget | Notes |
|---|---|---|
| Capture and encode | 0.5 s | Low-latency encoder settings, no lookahead-heavy presets |
| Packaging one part | 0.5 s | A part cannot be published until its last frame is encoded |
| Origin to edge to player | 0.3 s | Blocked request already waiting; part streams on completion |
| PART-HOLD-BACK | 1.5 s | Three times the 0.5 s part target, the SHOULD value |
| Decode and render | 0.2 s | Player and device dependent |
| Total | about 3.0 s | Compared with roughly 20 s for 6 s segments and classic hold-back |
Derived settings: PART-TARGET=0.5, TARGETDURATION=4 with 4-second GOPs aligned to segments, an INDEPENDENT part whenever a keyframe lands, PART-HOLD-BACK=1.5, HOLD-BACK=12 for classic clients, CAN-SKIP-UNTIL=24. The Advance Part Limit is 3 / 0.5 = 6 parts. Parts SHOULD leave the playlist once they are more than three target durations (12 s) from the end, and must stay downloadable for three target durations after that. Every playlist update now happens about twice a second per rendition, so origin capacity planning must count blocked requests, not just bytes.
CDN and origin requirements
- Keep the directives in the cache key. If the CDN strips
_HLS_msn,_HLS_partand_HLS_skipor ignores the query string, every blocking request is answered from a stale cached playlist, players spin or fall back to classic latency, and the bug looks like an origin problem. - Collapse identical requests. Thousands of viewers block on the same
_HLS_msn/_HLS_partpair at the same moment; the edge should forward one to the origin and fan out the answer. Without request collapsing, each new part triggers an origin stampede. - Long upstream timeouts. Edge-to-origin timeouts must exceed three target durations, or the CDN turns a normal blocked request into an error.
- Short TTLs only where needed. Directive-carrying playlists can be cached briefly because their content is fixed once produced; parts and segments are immutable and cache normally. Answer unknown, unhinted resources with 404.
Edge architecture and request collapsing in general are covered in video CDN architecture.
Failure modes
| Symptom | Likely cause | Response |
|---|---|---|
| Latency drifts back to 15-20 s | CDN ignores query strings or player fell out of Low-Latency Mode | Check cache keys; log which mode the player is in |
| Bursts of 503 from origin | Encoder or packager stalled for over three target durations | Alert on part publication gaps; fail over the encoder |
| Bursts of 400 | Client clock or state far ahead after a resume or rendition switch | Reload without directives; use rendition reports when switching |
| Needless downshifts | Server trickles hinted parts instead of sending each part whole | Enforce the full-speed delivery rule at origin and edge |
| Slow start or black frames at join | Player joined on a part without an independent frame | Mark INDEPENDENT parts correctly; start at one |
| Rebuffering on mobile networks | Hold-back too small for real jitter | Raise PART-HOLD-BACK or the part target; measure stalls against latency |
What to do next
- Measure your current glass-to-glass latency with a burned-in clock before changing anything, so improvements are real numbers.
- Configure the packager for parts with a target around 0.3 to 1 second and confirm every part is between 85% and 100% of the target.
- Verify the CDN preserves the _HLS_ query parameters in the cache key, collapses identical blocking requests and allows upstream waits longer than three target durations.
- Validate playlists with a conformance tool and with your real player fleet, including a classic-only player to prove backward compatibility.
- Add alerts for part publication gaps, origin 400 and 503 rates, and per-session latency from the player.
- Tune PART-HOLD-BACK against stall rate on real networks, and compare the result with LL-DASH using HLS vs DASH if you serve both.