Video on demand forgives almost everything: a file can be encoded slowly, packaged at leisure, checked and then cached forever. Live streaming forgives nothing. Every stage must keep up with real time, the manifest the player reads changes every few seconds, and a failure that lasts ten seconds is visible to every viewer at once. The pipeline looks similar to on-demand on a diagram, but its design is dominated by three concerns that on-demand barely has: sustained real-time throughput, end-to-end latency, and redundancy without a visible glitch.

This article walks through a live pipeline stage by stage, from the camera's encoder to the viewer's player: how contribution protocols differ, what a live transcoder must guarantee, how live HLS and DASH manifests work, what the origin and CDN do differently for live, where latency comes from with a worked budget, how to build redundancy, and what to monitor. Low-latency modes and packaging formats have their own deep dives, linked where they fit.

Advertisement

The pipeline at a glance

A live streaming pipeline, glass to glassCamera and encodervenue or studioIngestRTMP, SRT, WHIPLive transcoderABR ladder, aligned IDRsPackagerCMAF, HLS, DASHOriginmanifests, DVR windowCDNshield, edges, collapsePlayerbuffer, ABR, clockBackup pathsecond encoder and regionMonitoringingest, encode, edge, QoEcontributionsegmentsEvery stage adds delay; the player's hold-back is usually the largest single term.
Contribution, transcoding and packaging run in a facility or cloud region; origin and CDN distribute; the player decides the final delay.

A camera or production switcher feeds a contribution encoder, which sends one high-quality stream over the internet or a private link to an ingest endpoint. A live transcoder decodes it and produces an adaptive bitrate ladder, several renditions at different resolutions and bitrates. A packager cuts each rendition into segments and writes manifests. The origin serves those files and keeps a rolling window for rewind. A CDN fans them out to viewers, and the player chooses a rendition, maintains a buffer and decides how far behind the live edge to sit.

Contribution: getting the source in

The first hop carries one stream at high quality over networks you do not control, so its protocol choice is about loss recovery and latency.

ProtocolTransportStrengthsWeaknesses
RTMPTCPUniversal in encoders and platforms; simpleHead-of-line blocking on lossy links; classic RTMP is limited to older codecs, which the Enhanced RTMP specification extends with codecs such as HEVC and AV1
SRTUDP with retransmissionHandles loss and jitter with a configurable latency window; built-in encryptionNeeds the latency window sized to the path round trip; firewall and mode setup
WHIPWebRTC over UDPSub-second ingest; works from a browser; standardised as RFC 9725 in 2025WebRTC encoders favour low delay over quality; fewer broadcast encoders support it

SRT's key parameter is its latency: the receiver holds packets for that long so lost ones can be retransmitted. A common starting point is a few times the path round-trip time, raised for lossy links. Units differ by tool, so read the documentation: SRT configures latency in milliseconds, while FFmpeg's srt:// URL option takes microseconds. Whatever the protocol, the contribution encoder should send a constant frame rate, a fixed keyframe interval and a bitrate with headroom below the link capacity.

Advertisement

The live transcoder and its guarantees

The transcoder must encode every rendition faster than real time, continuously. If the 1080p encode runs at 0.97 times real time, it falls about two seconds behind per minute and the stream eventually stalls. Watch the encoder's speed, not its CPU usage, and leave margin; this is why live encoders use fast presets and why GPU and hardware encoders are common in live even where software encoders win on quality for on-demand.

The second guarantee is keyframe alignment. A player switching from 720p to 1080p mid-stream starts decoding the new rendition at a segment boundary, so every rendition must start each segment with an IDR frame at the same timestamp. That means a fixed GOP, scene-cut keyframes disabled, and segment duration equal to a whole number of GOPs. The GOP structure article covers the encoder side of this in detail.

ffmpeg -i "srt://0.0.0.0:9000?mode=listener" \
  -filter_complex "[0:v]split=3[a][b][c];[a]scale=-2:1080[v0];[b]scale=-2:720[v1];[c]scale=-2:480[v2]" \
  -map "[v0]" -map "[v1]" -map "[v2]" -map 0:a -map 0:a -map 0:a \
  -c:v libx264 -preset veryfast -profile:v high \
  -b:v:0 6000k -maxrate:v:0 6600k -bufsize:v:0 6000k \
  -b:v:1 3000k -maxrate:v:1 3300k -bufsize:v:1 3000k \
  -b:v:2 1200k -maxrate:v:2 1320k -bufsize:v:2 1200k \
  -g 60 -keyint_min 60 -sc_threshold 0 \
  -c:a aac -b:a 128k -ar 48000 \
  -f hls -hls_time 2 -hls_list_size 10 \
  -hls_flags delete_segments+independent_segments+program_date_time \
  -master_pl_name master.m3u8 \
  -hls_segment_filename "stream_%v/seg_%05d.ts" \
  -var_stream_map "v:0,a:0 v:1,a:1 v:2,a:2" "stream_%v/index.m3u8"

This FFmpeg command is a single-machine illustration of the idea. It listens for SRT, builds a three-rung ladder, forces a two-second GOP at 30 frames per second with -g 60 -sc_threshold 0, and writes two-second HLS segments with a ten-segment sliding window and program date-time tags, each rendition in its own directory. Production systems split these roles across services, but the invariants are the same. How to choose the rungs themselves is covered in adaptive bitrate streaming.

Packaging and live manifests

For live, the manifest is a moving window. An HLS media playlist lists the most recent segments, with #EXT-X-MEDIA-SEQUENCE giving the number of the first one and #EXT-X-TARGETDURATION the maximum segment length. The packager appends a segment and drops the oldest, and the player reloads the playlist roughly once per target duration. There is no end tag until the event ends. #EXT-X-PROGRAM-DATE-TIME maps segments to wall-clock time, which you need for measuring latency and syncing data such as scores.

DASH does the same with a dynamic MPD. Instead of listing segments, it usually declares a SegmentTemplate with a number or time pattern and an availabilityStartTime, so the player computes which segment should exist now from the clock. That makes clock agreement essential, which is why live MPDs carry a UTCTiming element pointing to a time source; a player with a skewed clock requests segments that do not exist yet. timeShiftBufferDepth declares how far back viewers may seek.

Most modern pipelines package once into CMAF fragmented MP4 and serve both HLS and DASH manifests over the same media files, which halves storage and CDN cache footprint; see CMAF. Ad markers from the source arrive as SCTE-35 signals, which the packager converts into manifest cues at the right segment boundary, as covered in SCTE-35 in live streams.

Origin, DVR window and CDN

The origin is where live and on-demand caching rules diverge most. Segments are immutable once written, so they can be cached for a long time. Manifests change every segment, so their cache lifetime must be short, typically a fraction of the segment duration, or viewers will sit on a stale playlist and fall behind or stall. A single wrong cache header on the manifest path is one of the most common live incidents.

Live traffic is also extremely bursty: at the moment a new segment appears, every viewer of that rendition requests it within a second or two. The CDN must collapse concurrent requests for the same object into one origin fetch, and an origin shield tier should sit between edges and origin so that hundreds of edge locations do not each hit the origin. The DVR window, how far viewers can rewind, is just how many segments the origin keeps and the manifest lists; long windows mean long manifests, so large DVR windows are often served from a separate time-shift path. Video CDN architecture covers shielding and request collapsing in depth.

Worked example: a latency budget

Glass-to-glass latency is the time from light hitting the camera to the frame appearing on a viewer's screen. For a conventional HLS stream with two-second segments, a rough budget, illustrative rather than measured, looks like this:

StageDelay addedWhy
Capture and contribution encode0.3 to 1 sEncoder lookahead and buffering
Contribution network0.2 to 1 sSRT latency window or TCP buffering
Transcode0.5 to 1 sDecode, scale, encode, lookahead
Segment accumulation2 sA segment cannot be published until it is complete
Packaging, origin, CDN0.2 to 0.5 sUpload and first fetch through the shield
Player hold-backabout 6 sHLS clients should start at least three target durations from the end
Totalroughly 9 to 12 sThe player term dominates

The lesson is where to cut. Shorter segments reduce both the accumulation and the hold-back terms but increase request rates and reduce compression efficiency. Low-Latency HLS goes further with partial segments, preload hints and blocking playlist reloads using the _HLS_msn and _HLS_part query parameters, where the part hold-back must be at least twice the part target and three times is recommended. Low-latency DASH uses chunked CMAF transfer. These reach a few seconds; below that, WebRTC delivery is the usual answer. The details are in low-latency streaming.

Redundancy without a glitch

A live event cannot be re-run, so every single point of failure matters. The standard design runs two complete chains, two contribution encoders fed from the same source, two ingest points in different availability zones or regions, and two transcoder and packager stacks, with players or the CDN able to fail over between them.

  • Deterministic segmentation. If both chains cut segments at the same timestamps with the same sequence numbers, derived from a shared clock rather than from when each process started, their outputs are interchangeable and a failover is invisible. If not, the player sees a discontinuity or a jump.
  • Failover points. HLS multivariant playlists can list redundant variant streams with the same attributes on different hosts, and DASH manifests can list multiple base URLs; players move to the backup when the primary errors. Origins can also fail over internally so the change never reaches the player.
  • Input failover. Transcoders can switch to a backup input on loss of signal, but switching sources causes a visible cut, so detect loss quickly and prefer redundant chains to redundant inputs where budget allows.
  • Rehearsal. Kill each component during a test event. Untested failover is the norm in incident reviews.

What to monitor

Monitor each stage by the property that fails there. At ingest: connection state, received bitrate and SRT retransmission or loss counts. At the transcoder: encode speed per rendition, which must stay above 1.0 times real time, and dropped frames. At the packager and origin: manifest freshness, meaning the newest segment number advances every target duration, and segment availability. At the edge: 404 and 5xx rates on segment paths and cache hit ratio. At the player: startup time, rebuffering ratio, bitrate and measured latency using program date-time.

import re, time, urllib.request

def media_sequence(url):
    text = urllib.request.urlopen(url, timeout=5).read().decode()
    target = int(re.search(r"#EXT-X-TARGETDURATION:(\d+)", text).group(1))
    seq = int(re.search(r"#EXT-X-MEDIA-SEQUENCE:(\d+)", text).group(1))
    segments = text.count("#EXTINF")
    return seq + segments, target        # number of the newest segment + 1

def watch(url):
    last, last_change = None, time.time()
    while True:
        newest, target = media_sequence(url)
        if newest != last:
            last, last_change = newest, time.time()
        elif time.time() - last_change > 2 * target:
            alert(f"{url} stale for {time.time() - last_change:.1f}s")
        time.sleep(target / 2)

The staleness check above is deliberately simple: if the newest segment number has not advanced for two target durations, the chain upstream of the manifest has stalled. Run it against both the origin and a few CDN edges; a stale edge with a fresh origin points at cache headers, while a stale origin points at the transcoder or packager.

Trade-offs and failure modes

  • Latency versus stability. Every second removed from the player buffer is a second less of protection against network jitter; low-latency modes rebuffer more on poor connections, so offer them where interactivity matters.
  • Segment length. Short segments lower latency but raise request rates, CDN cost and encoder keyframe overhead.
  • Cloud versus on-premises encode. Cloud transcoders scale per event but add a contribution hop over the internet; on-premises encoders need capacity for the peak event.
  • Common failures. Misaligned keyframes causing glitches on rendition switches, manifests cached too long, encoder speed dipping under real time during complex scenes, clock skew breaking DASH segment addressing, and failover that works only on paper.

What to do next

  1. Draw your pipeline and write a latency budget per stage, then measure it with program date-time.
  2. Pick a contribution protocol per venue: SRT for lossy public internet, RTMP for compatibility, WHIP when sub-second ingest matters.
  3. Enforce a fixed GOP with scene-cut keyframes disabled, and verify IDR alignment across all renditions.
  4. Set short cache lifetimes on manifests and long ones on segments, and confirm request collapsing and an origin shield.
  5. Alert on encoder speed below real time, manifest staleness and edge 404s on segments.
  6. Build a second chain with deterministic segmentation and rehearse failover under load.
  7. Only then pursue low-latency modes, and measure their rebuffering on real networks.
Key takeaway: A live stream is a chain of real-time stages, contribution, transcoding, packaging, origin, CDN and player, each of which must keep pace with the clock. Choose the contribution protocol for the network you have, keep keyframes aligned across the ladder, treat manifests as short-lived and segments as immutable, and remember that the player's hold-back is the largest latency term. Reliability comes from two deterministic chains that can replace each other invisibly, and from monitoring encode speed and manifest freshness, which are the first signals of nearly every live incident.