A YouTube-style platform is usually drawn as a data pipeline: upload, transcode, store, deliver through a CDN. That pipeline moves the bytes, and it is covered in depth elsewhere on this site. This article is about the other half, the control plane: the services that decide what a video is allowed to be at any moment, who may play it, where each viewer left off, and how a takedown reaches every copy and cache. Data-plane failures make video slow; control-plane failures play a deleted video, show a private one to the wrong person, or lose everyone's resume position.
We will design four pieces from first principles: the video lifecycle state machine, the playback-session API, the watch-history and resume path, and takedown propagation. Each comes with code, sizing for a stated workload, failure modes and the trade-offs that decide the design.
The control plane at a glance
Requirements and sizing
Fix a workload so every number has a source. Assume 100 million daily viewers, 6 plays per viewer per day and 40 minutes watched per viewer per day, with peak traffic at three times the daily average. The requirements that shape the control plane are:
- Correct visibility. A private, blocked or deleted video must stop playing within a bounded time, say five minutes, everywhere.
- Fast start. The playback decision sits on the critical path to the first frame, so it needs a p99 of tens of milliseconds.
- Durable resume. Losing a few seconds of position is fine; losing the whole history is not.
- Auditability. Every state change, especially a legal takedown, needs a record of who, when and why.
Plays per day are 100M x 6 = 600M, or about 6,900 per second on average and about 21,000 per second at peak. That is the playback API load. Heartbeats are the larger number: one every 30 seconds for 40 minutes per viewer is 80 per viewer per day, or 8 billion per day, about 93,000 per second on average and roughly 280,000 at peak. Writing every heartbeat to durable storage would make resume the biggest write workload in the company, which is the first design decision we will avoid.
The video lifecycle as a state machine
A video is not a row with a few boolean flags such as is_public, is_deleted and is_processed. Flags allow combinations that make no sense, for example deleted but public, and every reader has to know the precedence rules. Model it instead as an explicit state machine with a version number, where the state service is the only writer and every change goes through one function that checks the transition is legal.
ALLOWED = {
"CREATED": {"UPLOADING", "DELETED"},
"UPLOADING": {"PROCESSING", "FAILED", "DELETED"},
"PROCESSING": {"READY", "FAILED", "DELETED"},
"READY": {"PUBLISHED", "DELETED"},
"PUBLISHED": {"READY", "BLOCKED", "DELETED"}, # READY = unpublished by owner
"BLOCKED": {"PUBLISHED", "DELETED"}, # restore after appeal
"FAILED": {"PROCESSING", "DELETED"},
"DELETED": set(), # terminal
}
def transition(db, video_id, expected_version, new_state, actor, reason):
with db.transaction() as tx:
row = tx.one("SELECT state, version FROM videos WHERE id = %s FOR UPDATE", video_id)
if row.version != expected_version:
raise Conflict("stale version; re-read and retry")
if new_state not in ALLOWED[row.state]:
raise IllegalTransition(f"{row.state} -> {new_state}")
tx.execute("UPDATE videos SET state = %s, version = version + 1, updated_at = now() "
"WHERE id = %s", new_state, video_id)
tx.execute("INSERT INTO outbox (video_id, version, event, actor, reason) "
"VALUES (%s, %s, %s, %s, %s)",
video_id, row.version + 1, new_state, actor, reason)Three details carry the design. The version check makes concurrent writers, such as an owner publishing while trust and safety blocks, fail loudly instead of silently overwriting each other. The outbox row is written in the same transaction as the state change, so an event is published if and only if the change committed; a relay process reads the outbox and appends to the event log. And visibility (public, unlisted, private) and restrictions (regions, age) are separate columns from the lifecycle state, because they change independently and combine at read time.
The playback-session API and signed URLs
When a viewer presses play, the app calls one endpoint, POST /v1/playback/{video_id}, and that call makes every access decision. In order, it reads the video state and visibility (from a cache keyed by video id that stores the version and has a short TTL), checks the viewer's relationship to the video (owner, allowed account for a private video, anyone for public), applies region and age restrictions from the request's geo and the account, checks entitlement for paid content, and only then returns a manifest URL carrying a signed token and a session id.
The token is what lets the CDN enforce the decision without calling back. It binds a path prefix (the video's renditions), an expiry and the session id, and is signed with a key the edge also holds. CDNs implement token authentication in their own formats, so treat the code below as the principle rather than a drop-in.
import hashlib, hmac, time, base64
def sign(path_prefix: str, session_id: str, ttl_s: int, key: bytes) -> str:
path_prefix = path_prefix.rstrip("/") + "/" # /v/abc/ must not match /v/abcd/
expires = int(time.time()) + ttl_s
msg = f"{path_prefix}|{session_id}|{expires}".encode()
mac = base64.urlsafe_b64encode(hmac.new(key, msg, hashlib.sha256).digest()).decode()
return f"exp={expires}&sid={session_id}&sig={mac}"
def verify(path: str, path_prefix: str, session_id: str, expires: int, sig: str, key: bytes) -> bool:
path_prefix = path_prefix.rstrip("/") + "/"
if not path.startswith(path_prefix) or time.time() > expires:
return False
msg = f"{path_prefix}|{session_id}|{expires}".encode()
good = base64.urlsafe_b64encode(hmac.new(key, msg, hashlib.sha256).digest()).decode()
return hmac.compare_digest(good, sig)The token lifetime is the most important number in the control plane. It bounds how long a revoked video keeps playing for someone who already started it, and it bounds the damage of a shared URL. Long-form playback usually needs a refresh: the player calls the playback API again before expiry, which re-runs every check. With a 10-minute token and refresh, a takedown stops existing sessions within about 10 minutes plus cache TTL. DRM licences are a separate, stronger control for premium content and are not replaced by this token.
Watch history and resume position
Resume looks trivial: store the last position per user and video. The scale is what makes it a design problem, at 280,000 heartbeats per second at peak. The answer is to separate the hot session from the durable history. Heartbeats go to a session service that keeps the current position for each active session in memory or in a replicated cache, keyed by session id. It writes to the durable history store only on pause, stop, completion, app backgrounding, or every two minutes of continuous play.
With 40 watch minutes per viewer and a flush every two minutes, durable writes fall to about 20 per viewer per day plus a few for pauses: roughly 2.5 billion per day, around 29,000 per second on average and 90,000 at peak. That is a better than threefold reduction, paid for with at most two minutes of lost position if a session node dies, which the requirements accept.
The history store is a wide-column table partitioned by user id and clustered by last-watched time descending, so "continue watching" is a single-partition read of the first 20 rows. Each row holds video id, position, duration, a completed flag, the update time and a client sequence number. Writes are last-write-wins on the sequence number, not on server time, so a phone and a television watching the same video do not fight over clock skew; the device with the most recent playback wins.
| Quantity | Assumption | Result |
|---|---|---|
| Rows per heavy user | 6 distinct videos a day, 2 years kept | about 4,400 |
| Bytes per row | ids, position, times, flags | about 100 |
| Total, before replication | 100M users x 4,400 x 100 B | about 44 TB |
| With 3 replicas | about 132 TB |
That is an upper bound, since most users are not heavy users, and a TTL on rows older than two years keeps it from growing without limit.
Takedown and deletion propagation
A takedown is one state transition, PUBLISHED to BLOCKED, and the rest of the system holds copies of the old truth: the playback cache, the search index, recommendation candidate lists, CDN-cached manifests and thumbnails, and tokens already issued. The outbox event drives each projection through its own consumer, and each consumer is idempotent by video id and version, so replaying the log is always safe.
- The playback cache entry for the video is invalidated; new plays are refused immediately.
- The search consumer hides the document; it reappears only on a later PUBLISHED event with a higher version.
- The recommendation consumer adds the id to a deny list checked at serving time, because candidate lists are rebuilt in batch and would otherwise carry the video for hours.
- Manifests and thumbnails are purged from the CDN; segments are not purged, because without a fresh token nobody can fetch them after the existing tokens expire.
- Deletion, a separate transition, schedules blob removal after a retention window for appeals and legal holds, and records the purge in the audit log.
Events can be lost by a consumer bug or a skipped partition, so the design does not rely on them alone. A reconciler scans recently changed videos, compares the state service's version with each projection's recorded version, and repairs any projection that lags. The event path gives speed; the reconciler gives the guarantee.
Worked example: a takedown, second by second
Trace a court-ordered takedown issued at 12:00:00 for a video with 5,000 viewers mid-play. The state service commits BLOCKED with version 8 and an outbox row at 12:00:00. The relay publishes by 12:00:01. The playback cache is invalidated by 12:00:02 and new plays get a 451 response with a reason code. Search hides the document by 12:00:05, and the recommendation deny list updates within seconds. The 5,000 active viewers hold tokens that expire within 10 minutes; their players try to refresh, the playback API refuses, and the last of them stops by 12:10. The reconciler's next pass at 12:05 confirms all projections are at version 8, and the audit log holds the actor, legal reference and timestamps. If the requirement were one minute instead of ten, the lever is the token lifetime, at the cost of ten times more refresh calls.
Failure modes
- Cache serves stale visibility. A playback cache with a long TTL and a missed invalidation keeps a blocked video playable. Keep the TTL short, ignore invalidation events older than the cached version, and let the reconciler catch misses.
- Outbox relay stalls. State changes commit but nothing propagates. Alert on outbox age, not just on errors.
- Token key leak or rotation gap. Rotate signing keys with an overlap window in which the edge accepts both.
- Heartbeat storm after an outage. Millions of clients reconnect at once. Add jitter to heartbeat intervals and let the session service shed load by dropping heartbeats, never playback calls.
- Resume regressions. A late write from an old device moves the position backwards. Sequence numbers fix it; server timestamps do not.
- Resurrection by replay. An old PUBLISHED event replayed after a block re-indexes the video. Version checks in every consumer prevent it.
Trade-offs
Short tokens give fast revocation but cost refresh traffic and make playback depend on the playback API being up; long tokens are the reverse. A single state service with an outbox gives one source of truth at the price of one strongly consistent store on the publish path, which is acceptable because state changes are rare compared with plays. Coalescing heartbeats trades a bounded loss of position for a threefold cut in durable writes. A reconciler costs background reads but turns eventual consistency into a measurable bound.
Related reading
This page covers the control plane. For the data plane see YouTube architecture: upload, processing, metadata, view counts and delivery, designing a file upload service and CDN architecture; for packaging and licences see HLS and DRM.
What to do next
- Write down your lifecycle states and the allowed transitions, and replace any boolean flags with them.
- Put an outbox in the same transaction as every state change and alert on its oldest unrelayed row.
- Decide your revocation bound, then derive the token lifetime and refresh interval from it.
- Size heartbeats and durable resume writes for your traffic, and pick a flush interval you can defend.
- Make every projection consumer idempotent by id and version, and build the reconciler before you need it.
- Run a takedown drill on a test video and measure time-to-dark for new plays, active sessions, search and recommendations.