A video-on-demand service looks like one product but is two systems with almost nothing in common. The control plane, sign-in, catalogue, recommendations, entitlement and the playback API, handles small requests and looks like any other web backend. The data plane moves the video bytes, and it dominates everything: cost, capacity planning, failure modes and the quality viewers actually notice. Netflix is the well-known example, running its control plane in the cloud and its own content delivery network, Open Connect, with caching appliances placed inside internet service providers' networks.
This article designs such a system from first principles: what the requirements imply, how a studio master becomes thousands of small files, how those files reach a television, and how the player decides, every few seconds, which quality to fetch next. It uses a worked capacity example with stated assumptions rather than any company's private numbers, and it ends with what to build first if you are doing this yourself.
Requirements and what they imply
Start with the requirements that shape the design, not the feature list:
- Start fast: playback should begin within about two seconds of pressing play.
- Never stall: rebuffering is the event viewers hate most, more than lower picture quality.
- Best quality the network allows: on a phone on mobile data and a 4K television on fibre, from the same title.
- Every device: browsers, phones, smart TVs and set-top boxes, each with its own codecs and DRM system.
- Protect content: studios require encryption and licence enforcement.
- Cost per hour streamed: bandwidth is the dominant variable cost, so every bit saved by better encoding is saved millions of times.
Together these push you to a specific shape: encode each title in advance into many qualities, cut each quality into short independently decodable segments, serve the segments as ordinary cacheable HTTP objects from as close to the viewer as possible, and let the player choose the quality of each segment based on what it measures. No streaming server keeps per-viewer state; the intelligence lives in the encoder and the player.
Ingest and chunked transcoding
A studio delivers a mezzanine file: a very high bitrate master, often hundreds of gigabytes for a feature film, plus audio tracks in several languages and subtitle files. Encoding that serially into a dozen renditions in several codecs would take days. The standard answer is to split the work:
- Validate the source: frame rate, resolution, colour metadata, audio channel layout, duration consistency. Bad inputs caught here are cheap; caught after encoding they are expensive.
- Chunk the video at scene cuts or closed group-of-pictures boundaries into pieces of a minute or so, so each piece can be encoded without its neighbours.
- Encode each chunk for each rendition in parallel on a pool of workers, which turns a multi-day job into hours and makes retries cheap: a failed worker redoes one chunk, not one film.
- Assemble and verify: concatenate chunks per rendition, then check duration, frame count and objective quality scores before anything is published.
The work is a DAG of idempotent tasks keyed by title, rendition and chunk, which suits a workflow engine with a durable queue. Store intermediate outputs in an object store; S3-style object storage gives the durability and parallel throughput this needs. Uploading the mezzanine itself is a resumable multipart upload problem, covered in designing a file upload service.
The bitrate ladder
A ladder is the set of resolution and bitrate pairs a title is encoded at. A fixed ladder applies the same pairs to every title, which wastes bits on simple content, such as animation with flat colours, and starves complex content, such as grain-heavy film or sport. Netflix publicly described moving to per-title encoding in 2015: encode each title at many candidate points, measure quality, and choose the ladder that sits on that title's quality-versus-bitrate curve. Later work applied the same idea per shot.
An illustrative ladder for one title in H.264 might look like this; the numbers are examples, not a recommendation:
| Rung | Resolution | Video bitrate | Typical client |
|---|---|---|---|
| 1 | 416x234 | 235 kbps | Weak mobile signal |
| 2 | 640x360 | 560 kbps | Phone, congested Wi-Fi |
| 3 | 960x540 | 1,500 kbps | Tablet |
| 4 | 1280x720 | 3,000 kbps | Laptop |
| 5 | 1920x1080 | 5,800 kbps | TV, good broadband |
Every codec multiplies the ladder. H.264 plays almost everywhere; HEVC, VP9 and AV1 need fewer bits for the same quality but are supported on fewer devices and cost far more compute to encode. Since encoding is paid once and delivery is paid per view, popular titles justify expensive codecs and slow encoder presets; a long-tail title watched a hundred times may not.
Packaging, manifests and DRM
Players do not fetch whole files. Each rendition is cut into segments of a few seconds, and a manifest lists the renditions and their segments. Two manifest formats dominate: HLS, from Apple, with .m3u8 playlists, and MPEG-DASH, with XML .mpd files. CMAF, a common fragmented-MP4 segment format, lets one set of encrypted media segments serve both, so you store and cache the media once and generate two small manifests. A simplified HLS multivariant playlist:
#EXTM3U
#EXT-X-VERSION:7
#EXT-X-MEDIA:TYPE=AUDIO,GROUP-ID="aud",NAME="English",LANGUAGE="en",DEFAULT=YES,URI="audio/en/index.m3u8"
#EXT-X-STREAM-INF:BANDWIDTH=700000,RESOLUTION=640x360,CODECS="avc1.4d401e,mp4a.40.2",AUDIO="aud"
video/360p/index.m3u8
#EXT-X-STREAM-INF:BANDWIDTH=3400000,RESOLUTION=1280x720,CODECS="avc1.4d401f,mp4a.40.2",AUDIO="aud"
video/720p/index.m3u8
#EXT-X-STREAM-INF:BANDWIDTH=6400000,RESOLUTION=1920x1080,CODECS="avc1.640028,mp4a.40.2",AUDIO="aud"
video/1080p/index.m3u8Segments are encrypted with a content key. The player's content decryption module obtains that key from a licence server after the playback API confirms the viewer is entitled: Widevine on Android and Chrome, PlayReady on Windows and many TVs, FairPlay on Apple devices. CMAF's common encryption lets the same encrypted segments work with more than one DRM system where the encryption modes match. Segment length is a trade-off: shorter segments let the player react faster and start sooner but mean more requests, more manifest entries and slightly worse compression, because each segment must start with a keyframe.
Delivery: caches, shields and steering
Segments are immutable, named by title, rendition and index, so they cache perfectly. The delivery tier is a hierarchy: edge caches close to viewers, a regional mid-tier or origin shield that absorbs edge misses, and origin storage that only the shield talks to. Without the shield, a new release would send a miss from every edge straight to origin at once. The general mechanics are in CDN design.
Video adds two techniques. Proactive placement: because the catalogue is known in advance and popularity is predictable, a service can fill edge caches during off-peak hours, which is how Netflix describes filling its Open Connect appliances. A new episode is then already at the edge when viewers press play. Steering: the playback API, not DNS alone, returns a ranked list of cache URLs for this viewer, chosen by network location, cache health and which caches hold the title. The player fails over down the list on errors.
The economics depend on edge hit ratio. Every percentage point of misses is traffic paid for twice, once into the shield and once out to the viewer, and it loads the tier with the least capacity.
The playback start sequence
- The player calls the playback API with the title, device capabilities (codecs, DRM system, HDR, maximum resolution) and network hints.
- The API checks entitlement and region rights, picks the renditions this device can play, and returns a manifest or manifest URL plus a ranked list of cache hosts.
- The player requests a licence with a challenge from its decryption module; the licence server returns keys bound to that device and session.
- The player fetches the initialisation segment and the first media segments, usually at a conservative rung, and starts rendering once a small buffer is filled.
- From then on it runs the adaptive bitrate loop and sends telemetry: startup time, rung switches, rebuffer events, errors.
Startup time is the sum of these round trips, so they are where engineering effort goes: prefetching manifests and licences for the title under the cursor, starting at a rung that downloads quickly, and keeping the API and licence servers in every region.
Adaptive bitrate: the player's control loop
After each segment, the player chooses the rung for the next one. There are two classic families, and production players blend them.
Throughput-based: estimate bandwidth from recent downloads and pick the highest rung below a safety fraction of it. It reacts quickly but trusts a noisy estimate. Buffer-based: map the current buffer level to a rung, low buffer meaning low quality, ignoring throughput once the buffer is healthy. The BOLA algorithm is a well-known formalisation. It is stable but slow to climb at start-up, when the buffer is empty.
LADDER_KBPS = [235, 560, 1500, 3000, 5800]
def throughput_rule(samples_kbps, safety=0.8):
# Harmonic mean resists one lucky fast download.
recent = samples_kbps[-5:]
if not recent: # nothing downloaded yet
return 0
est = len(recent) / sum(1.0 / s for s in recent)
ok = [i for i, b in enumerate(LADDER_KBPS) if b <= safety * est]
return max(ok) if ok else 0
def buffer_rule(buffer_s, reservoir=8.0, cushion=30.0):
# Below the reservoir: lowest rung. Above reservoir + cushion: highest.
if buffer_s <= reservoir:
return 0
frac = min(1.0, (buffer_s - reservoir) / cushion)
return int(frac * (len(LADDER_KBPS) - 1))
def choose(samples_kbps, buffer_s, current):
if buffer_s < 12.0: # start-up or trouble: trust throughput
nxt = throughput_rule(samples_kbps)
else: # steady state: trust the buffer
nxt = min(buffer_rule(buffer_s), throughput_rule(samples_kbps, 1.2))
if nxt > current + 1: # climb one rung at a time
nxt = current + 1
return nxtThe hybrid uses throughput while the buffer is small, then lets the buffer drive, capped so it never chooses a rung far above measured bandwidth. Climbing one rung at a time avoids visible quality oscillation. Real players add more: dropping a rung mid-download when a segment is arriving too slowly, ignoring cached segments in the throughput estimate, and limiting resolution to what the screen can show.
Worked example: capacity for one region
Assume 2 million concurrent viewers at evening peak in one country, an average delivered bitrate of 4 Mbps after ABR, and a 95% edge hit ratio. All three numbers are assumptions; measure your own.
- Edge egress at peak: 2,000,000 x 4 Mbps = 8 Tbps. If one cache server serves about 40 Gbps sustained, that is 200 servers at full load. Design for one site or one ISP failing, and for 30% growth, and you need roughly 300 to 350 spread across ISPs and exchange points.
- Shield traffic: 5% misses x 8 Tbps = 400 Gbps into the edges. This is why the hit ratio is the single most important number: at 90% the shield tier must carry twice as much.
- Storage per title: the five example rungs add up to about 11 Mbps, so a two-hour film needs 11 Mbps x 7,200 s = 79,200 megabits, about 10 GB per codec, before audio and subtitles. Three codecs make it about 30 GB.
- Requests: with 4-second segments, each viewer fetches one video segment every 4 seconds, so 2 million viewers generate about 500,000 segment requests per second, plus audio. Caches must be tuned for request rate, not only bytes.
The example shows where the money goes: bytes and requests at the edge. The control plane, which handles a few requests per viewer per session, is small by comparison.
Failure modes
- Thundering herd at release: a popular episode at midnight, not yet at the edge, sends misses to the shield and origin together. Pre-position and use request collapsing at every cache tier.
- Licence or playback API outage: caches are healthy but nobody can start. Run these per region, cache entitlement decisions briefly, and degrade features such as personalised artwork before core playback.
- Bad encode published: a corrupt chunk or audio drift reaches millions. Automated quality checks per rendition must gate publishing, with a fast way to pull a rendition from manifests.
- ABR oscillation: two players on one link fight, or an over-optimistic estimate causes repeated up-down switches. Smooth estimates and rate-limit up-switches.
- Cache host failure: without a ranked fallback list, a player retries a dead host. Steering must return alternatives, and players must fail over on the first error.
- Device fragmentation: a TV model mis-decodes one profile. Device-capability rules in the playback API let you exclude a rendition for one model without re-encoding.
Measuring quality of experience
Server metrics cannot tell you whether viewing was good; the player must report it. The core metrics are startup time, rebuffer ratio (stall time divided by watch time), rebuffer frequency, average delivered bitrate or perceptual quality score, rung switch rate, and playback failures by error code. Slice everything by device, ISP, cache site and app version, because problems are almost always local. Changes to ABR, ladders or steering should ship as A/B tests judged on these metrics; a change that raises average bitrate by 5% but adds rebuffers is usually a loss. Recommendation systems such as the one in short-video recommendation architecture consume the same viewing telemetry, so define events once.
What to do next
- Write down your peak concurrency, average bitrate and target hit ratio, and redo the capacity example with your own numbers.
- Build a chunked encoding pipeline with idempotent tasks and quality checks that gate publishing.
- Package as CMAF with HLS and DASH manifests so media is stored and cached once.
- Put an origin shield in front of storage and add request collapsing before your first big release.
- Ship player telemetry for startup, rebuffers and switches before tuning ABR, then test ABR changes as experiments.
- Rehearse failures: kill a cache site, the licence service and a region's playback API in staging and watch what players do.