Every live stream crosses at least two very different networks. The first mile carries one high-quality stream from an encoder, over whatever connection the venue or studio has, to a media server. The last mile carries the processed stream to anywhere from ten viewers to ten million. Protocols that are good at one of these jobs are usually poor at the other, which is why a typical 2026 pipeline ingests over RTMP or SRT and delivers over HLS or DASH, and why WebRTC sits in both columns only when latency must be under a second.
This page is the map. It explains where live latency comes from, then walks through each protocol you are likely to meet in terms of transport, reliability, latency and scale: RTMP and Enhanced RTMP, SRT, RIST and WHIP for contribution; HTTP segment delivery, WebRTC with WHEP, and the emerging Media over QUIC for distribution. Manifests and low-latency HLS have their own pages; here the goal is choosing correctly.
Two jobs: contribution and distribution
Contribution (ingest, first mile) is one sender and one receiver. The stream is high bitrate, often the only copy, and travels over networks you do not control: venue Wi-Fi, bonded cellular, a long-haul internet path. What matters is surviving loss and jitter with predictable delay, and authenticating the sender. Scale is irrelevant.
Distribution (delivery, last mile) is one source and many receivers on consumer networks. What matters is cost per viewer, cacheability, device reach, adaptive bitrate and DRM. Reliability is handled by buffering and by switching renditions rather than by retransmitting at all costs.
Between them sits a media server that transcodes into a bitrate ladder and packages, adding its own latency.
Where live latency comes from
Glass-to-glass latency is the sum of capture and encode, the contribution hop, transcoding, packaging, CDN transfer, the player's buffer and decode. For HTTP segment delivery two terms dominate: a segment cannot be published until it is complete, and players start several segments behind the live edge so that one slow download does not stall playback. Segment duration is therefore the main lever, and it multiplies.
def glass_to_glass_ms(encode=150, ingest=300, transcode=500, segment_s=6.0,
segments_held=3, cdn=100, decode=50):
"""Rough live latency for segment delivery: the player sits several segments behind live."""
packaging = segment_s * 1000 # a segment exists only once it is complete
player = segment_s * 1000 * segments_held # buffer the player keeps behind the live edge
return encode + ingest + transcode + packaging + cdn + player + decode
print(glass_to_glass_ms()) # classic 6 s segments -> about 25 s
print(glass_to_glass_ms(segment_s=2.0)) # 2 s segments -> about 9 s
print(glass_to_glass_ms(segment_s=0.5, segments_held=4)) # parts of ~0.5 s, LL-style -> about 3.6 s| Latency class | Typical glass-to-glass | Usual protocol | Fits |
|---|---|---|---|
| Broadcast-like | 15 to 30 s | HLS or DASH, 4 to 6 s segments | linear channels, very large events, maximum device reach |
| Reduced | 5 to 10 s | HLS or DASH, 1 to 2 s segments | most live events |
| Low | 2 to 5 s | LL-HLS, LL-DASH with chunked CMAF | sports, betting-adjacent feeds, live chat |
| Interactive | under 1 s | WebRTC through SFUs | auctions, watch parties, two-way shows |
These are design-time numbers: a player configured to hold three segments sits three segments behind whatever the protocol allows, and a 4-second keyframe interval cannot be cut into 2-second segments.
RTMP and Enhanced RTMP
RTMP is the protocol most encoders and platforms still accept. It runs over a single TCP connection (port 1935, or RTMPS over TLS, commonly on 443), multiplexes audio, video and metadata as interleaved chunk streams, and carries media in FLV tags. Its strength is ubiquity: a URL and a stream key work with OBS, hardware encoders and every platform.
Its weakness is TCP. On a lossy path a single lost packet blocks everything behind it until retransmission, the congestion window collapses, and the sender's buffer grows until the encoder drops frames or the connection stalls. Classic RTMP also only signalled a few codecs, which pinned ingest to H.264 and AAC. Enhanced RTMP, specified by the Veovera group (now at version 2), adds FourCC-based codec signalling so RTMP can carry HEVC, VP9 and AV1 among others; support depends on both the encoder and the ingest server, so verify both ends before relying on it.
# RTMP over TLS (RTMPS) to a platform ingest point: FLV payload over TCP
ffmpeg -re -i event.mp4 -c:v libx264 -preset veryfast -g 60 -keyint_min 60 -sc_threshold 0 \
-b:v 6M -c:a aac -b:a 160k -f flv "rtmps://ingest.example.com:443/live/STREAM_KEY"
# SRT caller to a listener: MPEG-TS over UDP, retransmission inside a fixed latency window.
# In ffmpeg's libsrt protocol the latency option is in MICROseconds: 800000 = 800 ms.
ffmpeg -re -i event.mp4 -c:v libx264 -g 60 -b:v 6M -c:a aac -f mpegts \
"srt://ingest.example.com:9000?mode=caller&latency=800000&passphrase=CHANGE_ME_16chars"Note the keyframe settings in the RTMP command: a fixed GOP of 60 frames at 30 fps gives a keyframe every 2 seconds, with scene-cut keyframes disabled, so the packager downstream can cut segments on exact 2-second boundaries.
SRT and RIST: reliable UDP with a latency budget
SRT (Secure Reliable Transport, open-sourced by Haivision and documented in an IETF Internet-Draft, not an RFC) runs over UDP and replaces TCP's open-ended reliability with a bounded one. The receiver holds packets in a buffer for a configured latency; lost packets are requested with negative acknowledgements and retransmitted, and any packet that still has not arrived when its play time comes is dropped rather than waited for. You trade a fixed, known delay for loss recovery. SRT also encrypts with AES using a shared passphrase, and supports caller, listener and rendezvous connection modes, which helps with firewalls.
The latency setting is the whole game. It must cover several round trips so that a packet can be lost, reported and resent, possibly more than once. SRT's default is 120 ms, which is only enough on short, clean paths; for internet contribution a starting point of about four times the round-trip time, raised further on lossy links, is common guidance, and long-haul or cellular links often run 500 ms to 2 s. Watch the receiver's dropped-packet and retransmission counters and raise the latency until drops stop.
RIST (Reliable Internet Stream Transport, specified by the Video Services Forum in its TR-06 series) solves the same problem over RTP with similar NACK-based retransmission, defined in Simple, Main and Advanced profiles. It is common where broadcast equipment interoperability matters. Both are contribution protocols; no browser plays them.
WHIP: WebRTC as an ingest protocol
WebRTC has always been able to send media with sub-second latency, but every service invented its own signalling, so encoders could not target it generically. WHIP, published as RFC 9725 in March 2025, standardizes the signalling as one HTTP exchange: the client POSTs an SDP offer, the server answers with 201 Created, the SDP answer and a Location header naming the session resource, and the client DELETEs that resource to stop. ICE, DTLS-SRTP and congestion control then work as in any WebRTC session. OBS Studio added WHIP output in version 30, and FFmpeg 8.0 added a WHIP muxer.
# WHIP (RFC 9725): one HTTP POST carries the SDP offer, the response carries the answer.
POST /whip/endpoint/live1 HTTP/1.1
Host: ingest.example.com
Authorization: Bearer <token>
Content-Type: application/sdp
v=0 ... (SDP offer: one audio and one video m-line, ICE credentials, DTLS fingerprint)
HTTP/1.1 201 Created
Content-Type: application/sdp
Location: https://ingest.example.com/whip/session/9f2c
v=0 ... (SDP answer)
# Media now flows over ICE, DTLS-SRTP. To stop publishing, delete the session resource:
DELETE /whip/session/9f2c HTTP/1.1
Authorization: Bearer <token>Use WHIP when the publisher is a browser, when the whole pipeline must be sub-second, or when you want WebRTC congestion control adapting the bitrate to the uplink. Its limits come from WebRTC: codec choice is negotiated among what both sides support, quality at high bitrates is usually lower than a tuned SRT feed because the encoder favours latency, and the server must speak ICE and DTLS rather than a plain socket.
HTTP segment delivery: HLS and DASH
For distribution at scale, HTTP wins because it is cacheable. The packager writes each rendition as a series of short media segments plus a manifest, and every viewer fetches the same files from the nearest CDN edge. The player measures throughput and buffer and picks a rendition per segment, so a viewer on a weak network drops quality instead of stalling. DRM works through the browser's Encrypted Media Extensions and native platform stacks.
HLS (RFC 8216, with a revised specification maintained by Apple) and MPEG-DASH (ISO/IEC 23009-1) differ in manifest format and device reach, but with CMAF fragmented-MP4 segments one encoded segment set can serve both; HLS vs DASH, in depth covers the differences. Their low-latency modes publish segments in small parts or chunks as they are encoded, so the player can sit a second or two behind live without shrinking segments; Low-Latency HLS explained walks through partial segments, preload hints and blocking playlist reload, and the CMAF page covers the shared segment format.
WebRTC delivery, WHEP and Media over QUIC
Under a second, delivery moves to WebRTC: each viewer holds a peer connection to a selective forwarding unit, and large audiences need trees of SFUs. Latency is excellent, but nothing is cacheable: every viewer is a stateful session that costs server CPU and egress, and simulcast layers replace the bitrate ladder. Studio DRM through Encrypted Media Extensions is not available on this path. See the WebRTC pipeline architecture for ICE, SFUs and congestion control.
WHEP is WHIP's mirror image for playback: a viewer POSTs an offer and receives a session. As of this writing it is still an Internet-Draft in the IETF WISH working group, so implementations follow a moving text; check which draft revision a vendor implements.
Media over QUIC (MoQ) is an IETF effort to get WebRTC-like latency with CDN-like fan-out: publishers and subscribers exchange media objects over QUIC or WebTransport through relays that can cache and forward, with priorities so that stale media can be dropped instead of blocking new frames. The transport specification is still an Internet-Draft (revision 20 appeared at the end of August 2026). Treat it as something to prototype and follow, not something to build a production service on yet.
Choosing a protocol
| Situation | Ingest | Delivery | Why |
|---|---|---|---|
| Platform live stream from OBS | RTMPS (SRT if offered) | HLS/DASH, 2 s segments | universal encoder support, cheap scale |
| Remote production over the internet | SRT or RIST | depends on audience | bounded latency with loss recovery |
| Browser-based creator tools | WHIP | LL-HLS or WebRTC | no plug-in, sub-second first mile |
| Sports with a second-screen app | SRT | LL-HLS or LL-DASH | 2 to 5 s at CDN cost |
| Auction or interactive show | WHIP or SRT | WebRTC via SFUs | sub-second is the requirement |
| Premium content with studio DRM | SRT | HLS/DASH with CMAF | DRM and device reach |
Worked example: one event, three audiences
A regional football final is produced at the stadium and streamed to three audiences: the public app (up to 400,000 viewers), a betting partner that needs to be close to real time, and a studio commentary team watching remotely who talk back on air.
Contribution. The stadium uplink measures 40 ms RTT with occasional 2% loss bursts. The production sends two SRT feeds, primary and backup, over different ISPs, with latency set to 400 ms (ten round trips, comfortably above four), and the media server switches feeds automatically on loss of signal. First-half counters show retransmissions but no drops.
Public audience. The transcoder produces a six-rung ladder with a 2-second GOP and the packager publishes LL-HLS with 2-second segments in about 0.33-second parts, plus DASH from the same CMAF segments. Measured glass-to-glass is around 4 seconds. Holding back to 2-second classic segments would have been about 9 seconds.
Betting partner and commentators. Both need sub-second video, and both are small audiences. The media server forwards the top rung to an SFU, and the partner and commentators connect over WebRTC; the commentators' return audio is published back with WHIP from a browser. Because these audiences are tiny, the per-viewer cost of WebRTC does not matter; sending all 400,000 public viewers through SFUs would have.
Failure modes
- RTMP stalls on a lossy uplink. TCP head-of-line blocking and congestion collapse; the encoder drops frames. Move contribution to SRT or RIST, or bond connections.
- SRT latency set too low. Retransmissions arrive too late and packets are dropped as artefacts, often blamed on the encoder. Size latency from measured RTT and loss and watch drop counters.
- Keyframes not aligned to segments. Segments vary in length or start without a keyframe, renditions cannot be switched cleanly, and low-latency modes misbehave. Fix the GOP at the encoder and disable scene-cut keyframes for live ladders.
- Manifests cached too long. A CDN rule meant for segments applied to live playlists freezes players at an old edge. Give manifests short or request-collapsed caching, segments long caching.
- UDP blocked. Corporate networks block SRT and WebRTC media; WebRTC needs TURN over TCP or TLS as a fallback, and SRT needs an alternate path.
- Timestamp discontinuities on failover. Switching feeds without continuous timestamps stalls players. Rewrite timestamps on switch and test failover under load.
What to do next
- Write down the latency each audience actually needs, then compute the budget with your real segment and GOP settings.
- Measure RTT and loss on your contribution paths and choose RTMP only where the path is clean.
- If you use SRT, set latency from measured RTT, check the units in your tool, and alert on dropped packets.
- Fix the encoder GOP to divide your segment duration exactly and disable scene-cut keyframes for live.
- Serve HLS and DASH from one CMAF segment set and give manifests and segments separate CDN cache rules.
- Reserve WebRTC delivery for small or interactive audiences, and plan TURN fallback.
- Track WHEP and MoQ draft revisions and prototype MoQ off the critical path.