Every live stream crosses at least two very different networks. The first mile carries one high-quality stream from an encoder, over whatever connection the venue or studio has, to a media server. The last mile carries the processed stream to anywhere from ten viewers to ten million. Protocols that are good at one of these jobs are usually poor at the other, which is why a typical 2026 pipeline ingests over RTMP or SRT and delivers over HLS or DASH, and why WebRTC sits in both columns only when latency must be under a second.

This page is the map. It explains where live latency comes from, then walks through each protocol you are likely to meet in terms of transport, reliability, latency and scale: RTMP and Enhanced RTMP, SRT, RIST and WHIP for contribution; HTTP segment delivery, WebRTC with WHEP, and the emerging Media over QUIC for distribution. Manifests and low-latency HLS have their own pages; here the goal is choosing correctly.

Advertisement

Two jobs: contribution and distribution

Contribution (ingest, first mile) is one sender and one receiver. The stream is high bitrate, often the only copy, and travels over networks you do not control: venue Wi-Fi, bonded cellular, a long-haul internet path. What matters is surviving loss and jitter with predictable delay, and authenticating the sender. Scale is irrelevant.

Distribution (delivery, last mile) is one source and many receivers on consumer networks. What matters is cost per viewer, cacheability, device reach, adaptive bitrate and DRM. Reliability is handled by buffering and by switching renditions rather than by retransmitting at all costs.

Between them sits a media server that transcodes into a bitrate ladder and packages, adding its own latency.

First mile and last mile use different protocols for different reasonsEncodercamera, OBS, hardwareIngestRTMP, SRT, RIST, WHIPMedia servertranscode + package1 streamOrigin + CDNHLS, DASH, CMAFSFU treeWebRTC, WHEPRelaysMoQ (draft)Viewers1 to millionsContribution: one sender, one receiver,unpredictable network, quality first.Distribution: one source, many receivers,cost per viewer and scale first.Pick the ingest protocol for the encoder's network; pick the delivery protocol for the audience's latency need.
Ingest protocols are chosen for the encoder's network; delivery protocols for the audience's size and latency need.

Where live latency comes from

Glass-to-glass latency is the sum of capture and encode, the contribution hop, transcoding, packaging, CDN transfer, the player's buffer and decode. For HTTP segment delivery two terms dominate: a segment cannot be published until it is complete, and players start several segments behind the live edge so that one slow download does not stall playback. Segment duration is therefore the main lever, and it multiplies.

def glass_to_glass_ms(encode=150, ingest=300, transcode=500, segment_s=6.0,
                      segments_held=3, cdn=100, decode=50):
    """Rough live latency for segment delivery: the player sits several segments behind live."""
    packaging = segment_s * 1000                  # a segment exists only once it is complete
    player = segment_s * 1000 * segments_held     # buffer the player keeps behind the live edge
    return encode + ingest + transcode + packaging + cdn + player + decode

print(glass_to_glass_ms())                                 # classic 6 s segments  -> about 25 s
print(glass_to_glass_ms(segment_s=2.0))                    # 2 s segments          -> about 9 s
print(glass_to_glass_ms(segment_s=0.5, segments_held=4))   # parts of ~0.5 s, LL-style -> about 3.6 s
Latency classTypical glass-to-glassUsual protocolFits
Broadcast-like15 to 30 sHLS or DASH, 4 to 6 s segmentslinear channels, very large events, maximum device reach
Reduced5 to 10 sHLS or DASH, 1 to 2 s segmentsmost live events
Low2 to 5 sLL-HLS, LL-DASH with chunked CMAFsports, betting-adjacent feeds, live chat
Interactiveunder 1 sWebRTC through SFUsauctions, watch parties, two-way shows

These are design-time numbers: a player configured to hold three segments sits three segments behind whatever the protocol allows, and a 4-second keyframe interval cannot be cut into 2-second segments.

Advertisement

RTMP and Enhanced RTMP

RTMP is the protocol most encoders and platforms still accept. It runs over a single TCP connection (port 1935, or RTMPS over TLS, commonly on 443), multiplexes audio, video and metadata as interleaved chunk streams, and carries media in FLV tags. Its strength is ubiquity: a URL and a stream key work with OBS, hardware encoders and every platform.

Its weakness is TCP. On a lossy path a single lost packet blocks everything behind it until retransmission, the congestion window collapses, and the sender's buffer grows until the encoder drops frames or the connection stalls. Classic RTMP also only signalled a few codecs, which pinned ingest to H.264 and AAC. Enhanced RTMP, specified by the Veovera group (now at version 2), adds FourCC-based codec signalling so RTMP can carry HEVC, VP9 and AV1 among others; support depends on both the encoder and the ingest server, so verify both ends before relying on it.

# RTMP over TLS (RTMPS) to a platform ingest point: FLV payload over TCP
ffmpeg -re -i event.mp4 -c:v libx264 -preset veryfast -g 60 -keyint_min 60 -sc_threshold 0 \
       -b:v 6M -c:a aac -b:a 160k -f flv "rtmps://ingest.example.com:443/live/STREAM_KEY"

# SRT caller to a listener: MPEG-TS over UDP, retransmission inside a fixed latency window.
# In ffmpeg's libsrt protocol the latency option is in MICROseconds: 800000 = 800 ms.
ffmpeg -re -i event.mp4 -c:v libx264 -g 60 -b:v 6M -c:a aac -f mpegts \
       "srt://ingest.example.com:9000?mode=caller&latency=800000&passphrase=CHANGE_ME_16chars"

Note the keyframe settings in the RTMP command: a fixed GOP of 60 frames at 30 fps gives a keyframe every 2 seconds, with scene-cut keyframes disabled, so the packager downstream can cut segments on exact 2-second boundaries.

SRT and RIST: reliable UDP with a latency budget

SRT (Secure Reliable Transport, open-sourced by Haivision and documented in an IETF Internet-Draft, not an RFC) runs over UDP and replaces TCP's open-ended reliability with a bounded one. The receiver holds packets in a buffer for a configured latency; lost packets are requested with negative acknowledgements and retransmitted, and any packet that still has not arrived when its play time comes is dropped rather than waited for. You trade a fixed, known delay for loss recovery. SRT also encrypts with AES using a shared passphrase, and supports caller, listener and rendezvous connection modes, which helps with firewalls.

The latency setting is the whole game. It must cover several round trips so that a packet can be lost, reported and resent, possibly more than once. SRT's default is 120 ms, which is only enough on short, clean paths; for internet contribution a starting point of about four times the round-trip time, raised further on lossy links, is common guidance, and long-haul or cellular links often run 500 ms to 2 s. Watch the receiver's dropped-packet and retransmission counters and raise the latency until drops stop.

RIST (Reliable Internet Stream Transport, specified by the Video Services Forum in its TR-06 series) solves the same problem over RTP with similar NACK-based retransmission, defined in Simple, Main and Advanced profiles. It is common where broadcast equipment interoperability matters. Both are contribution protocols; no browser plays them.

WHIP: WebRTC as an ingest protocol

WebRTC has always been able to send media with sub-second latency, but every service invented its own signalling, so encoders could not target it generically. WHIP, published as RFC 9725 in March 2025, standardizes the signalling as one HTTP exchange: the client POSTs an SDP offer, the server answers with 201 Created, the SDP answer and a Location header naming the session resource, and the client DELETEs that resource to stop. ICE, DTLS-SRTP and congestion control then work as in any WebRTC session. OBS Studio added WHIP output in version 30, and FFmpeg 8.0 added a WHIP muxer.

# WHIP (RFC 9725): one HTTP POST carries the SDP offer, the response carries the answer.
POST /whip/endpoint/live1 HTTP/1.1
Host: ingest.example.com
Authorization: Bearer <token>
Content-Type: application/sdp

v=0 ... (SDP offer: one audio and one video m-line, ICE credentials, DTLS fingerprint)

HTTP/1.1 201 Created
Content-Type: application/sdp
Location: https://ingest.example.com/whip/session/9f2c

v=0 ... (SDP answer)

# Media now flows over ICE, DTLS-SRTP. To stop publishing, delete the session resource:
DELETE /whip/session/9f2c HTTP/1.1
Authorization: Bearer <token>

Use WHIP when the publisher is a browser, when the whole pipeline must be sub-second, or when you want WebRTC congestion control adapting the bitrate to the uplink. Its limits come from WebRTC: codec choice is negotiated among what both sides support, quality at high bitrates is usually lower than a tuned SRT feed because the encoder favours latency, and the server must speak ICE and DTLS rather than a plain socket.

HTTP segment delivery: HLS and DASH

For distribution at scale, HTTP wins because it is cacheable. The packager writes each rendition as a series of short media segments plus a manifest, and every viewer fetches the same files from the nearest CDN edge. The player measures throughput and buffer and picks a rendition per segment, so a viewer on a weak network drops quality instead of stalling. DRM works through the browser's Encrypted Media Extensions and native platform stacks.

HLS (RFC 8216, with a revised specification maintained by Apple) and MPEG-DASH (ISO/IEC 23009-1) differ in manifest format and device reach, but with CMAF fragmented-MP4 segments one encoded segment set can serve both; HLS vs DASH, in depth covers the differences. Their low-latency modes publish segments in small parts or chunks as they are encoded, so the player can sit a second or two behind live without shrinking segments; Low-Latency HLS explained walks through partial segments, preload hints and blocking playlist reload, and the CMAF page covers the shared segment format.

WebRTC delivery, WHEP and Media over QUIC

Under a second, delivery moves to WebRTC: each viewer holds a peer connection to a selective forwarding unit, and large audiences need trees of SFUs. Latency is excellent, but nothing is cacheable: every viewer is a stateful session that costs server CPU and egress, and simulcast layers replace the bitrate ladder. Studio DRM through Encrypted Media Extensions is not available on this path. See the WebRTC pipeline architecture for ICE, SFUs and congestion control.

WHEP is WHIP's mirror image for playback: a viewer POSTs an offer and receives a session. As of this writing it is still an Internet-Draft in the IETF WISH working group, so implementations follow a moving text; check which draft revision a vendor implements.

Media over QUIC (MoQ) is an IETF effort to get WebRTC-like latency with CDN-like fan-out: publishers and subscribers exchange media objects over QUIC or WebTransport through relays that can cache and forward, with priorities so that stale media can be dropped instead of blocking new frames. The transport specification is still an Internet-Draft (revision 20 appeared at the end of August 2026). Treat it as something to prototype and follow, not something to build a production service on yet.

Choosing a protocol

SituationIngestDeliveryWhy
Platform live stream from OBSRTMPS (SRT if offered)HLS/DASH, 2 s segmentsuniversal encoder support, cheap scale
Remote production over the internetSRT or RISTdepends on audiencebounded latency with loss recovery
Browser-based creator toolsWHIPLL-HLS or WebRTCno plug-in, sub-second first mile
Sports with a second-screen appSRTLL-HLS or LL-DASH2 to 5 s at CDN cost
Auction or interactive showWHIP or SRTWebRTC via SFUssub-second is the requirement
Premium content with studio DRMSRTHLS/DASH with CMAFDRM and device reach

Worked example: one event, three audiences

A regional football final is produced at the stadium and streamed to three audiences: the public app (up to 400,000 viewers), a betting partner that needs to be close to real time, and a studio commentary team watching remotely who talk back on air.

Contribution. The stadium uplink measures 40 ms RTT with occasional 2% loss bursts. The production sends two SRT feeds, primary and backup, over different ISPs, with latency set to 400 ms (ten round trips, comfortably above four), and the media server switches feeds automatically on loss of signal. First-half counters show retransmissions but no drops.

Public audience. The transcoder produces a six-rung ladder with a 2-second GOP and the packager publishes LL-HLS with 2-second segments in about 0.33-second parts, plus DASH from the same CMAF segments. Measured glass-to-glass is around 4 seconds. Holding back to 2-second classic segments would have been about 9 seconds.

Betting partner and commentators. Both need sub-second video, and both are small audiences. The media server forwards the top rung to an SFU, and the partner and commentators connect over WebRTC; the commentators' return audio is published back with WHIP from a browser. Because these audiences are tiny, the per-viewer cost of WebRTC does not matter; sending all 400,000 public viewers through SFUs would have.

Failure modes

  • RTMP stalls on a lossy uplink. TCP head-of-line blocking and congestion collapse; the encoder drops frames. Move contribution to SRT or RIST, or bond connections.
  • SRT latency set too low. Retransmissions arrive too late and packets are dropped as artefacts, often blamed on the encoder. Size latency from measured RTT and loss and watch drop counters.
  • Keyframes not aligned to segments. Segments vary in length or start without a keyframe, renditions cannot be switched cleanly, and low-latency modes misbehave. Fix the GOP at the encoder and disable scene-cut keyframes for live ladders.
  • Manifests cached too long. A CDN rule meant for segments applied to live playlists freezes players at an old edge. Give manifests short or request-collapsed caching, segments long caching.
  • UDP blocked. Corporate networks block SRT and WebRTC media; WebRTC needs TURN over TCP or TLS as a fallback, and SRT needs an alternate path.
  • Timestamp discontinuities on failover. Switching feeds without continuous timestamps stalls players. Rewrite timestamps on switch and test failover under load.

What to do next

  1. Write down the latency each audience actually needs, then compute the budget with your real segment and GOP settings.
  2. Measure RTT and loss on your contribution paths and choose RTMP only where the path is clean.
  3. If you use SRT, set latency from measured RTT, check the units in your tool, and alert on dropped packets.
  4. Fix the encoder GOP to divide your segment duration exactly and disable scene-cut keyframes for live.
  5. Serve HLS and DASH from one CMAF segment set and give manifests and segments separate CDN cache rules.
  6. Reserve WebRTC delivery for small or interactive audiences, and plan TURN fallback.
  7. Track WHEP and MoQ draft revisions and prototype MoQ off the critical path.
Key takeaway: Contribution and distribution are different problems: ingest needs loss recovery with predictable delay over one uncontrolled path, delivery needs cheap cacheable fan-out with adaptive bitrate and DRM. RTMP remains the universal ingest, SRT and RIST bound latency over lossy internet paths, and WHIP (RFC 9725) standardizes WebRTC ingest. HLS and DASH with CMAF serve large audiences at 2 to 30 seconds depending on segment design, WebRTC through SFUs serves small audiences under a second, and WHEP and MoQ are still drafts. Choose per audience from a written latency budget, and align GOPs, segments and cache rules to it.