A music service looks like a smaller video service: compressed media, a CDN, a player. The engineering problem is different. A song is a small file played whole, the catalogue runs to tens of millions of tracks, a single listener plays dozens of them a day, and every play may trigger a royalty payment, so counting plays has to be as careful as billing. Playback must feel instant, gapless and available offline on a phone in a tunnel.
This article designs a Spotify-style streaming system from first principles. It separates what Spotify has published, such as its audio quality tiers, its CDN set-up and its event-delivery migration, from design choices that any team building such a system would face. It covers capacity estimates, ingestion, the playback path, CDN strategy, offline keys, playlists and the play-event pipeline, then failure modes and a checklist.
Requirements that shape the design
Start with the requirements that shape the architecture.
- Start fast. Playback should begin within a few hundred milliseconds of a tap, and the next track should start with no gap.
- Huge catalogue, steep popularity curve. A small share of tracks takes most plays, and a very long tail is played rarely. Caching works well at the head and poorly at the tail.
- Rights are per track, per region, per plan. Whether a given account may play a given recording changes with licensing deals, so entitlement is checked at play time, not at upload time.
- Offline playback. Downloaded tracks must play without a network yet stop working when the subscription lapses.
- Accurate play counts. Plays drive payments to rights holders, so they must be complete, deduplicated and auditable.
- Mutable user data. Libraries and playlists, some edited collaboratively, synced across devices.
Capacity: why audio is small
Spotify's support pages list music quality settings of about 24, 96, 160 and 320 kbit/s, plus lossless up to 24-bit/44.1 kHz FLAC; the web player streams AAC at 128 kbit/s on the free tier and 256 kbit/s on Premium. Those numbers make tracks small. A 3.5-minute track, 210 seconds, weighs:
def track_mb(kbps, seconds=210):
return kbps * 1000 * seconds / 8 / 1e6
for q in (24, 96, 160, 320):
print(q, "kbit/s ->", round(track_mb(q), 2), "MB")
# 24 -> 0.63 MB 96 -> 2.52 MB 160 -> 4.2 MB 320 -> 8.4 MB
pcm_24bit = 44_100 * 24 * 2 / 1000 # 2,116.8 kbit/s raw stereo
print("lossless ~", round(track_mb(pcm_24bit * 0.6), 1), "MB") # assumes FLAC at 60% of raw: ~33 MBThe FLAC ratio is an assumption; real compression depends on the music. The lossy figures show why music differs from video. Even at the top lossy tier, a whole track is smaller than a few seconds of 4K video, so there is no need to cut audio into short segments and switch bitrates mid-stream. The client can fetch the file with a few HTTP range requests, and the CDN can cache whole files. Storage per track across all renditions is tens of megabytes, so a catalogue of tens of millions of tracks needs petabytes at origin, which is large but ordinary object storage.
Concurrency sets bandwidth. Ten million concurrent listeners at an average of 160 kbit/s is 1.6 Tbit/s of egress, all served by CDNs; the origin sees only misses.
Ingest: from master to encrypted files
Labels and distributors deliver masters with metadata: track IDs, ISRC codes, contributors, territories and release dates. Ingest is a batch pipeline with a strict order.
- Validate. Check the audio decodes, durations match metadata, and identifiers are not already in use. Reject early; a bad master is far cheaper to bounce than to take down later.
- Measure loudness. Compute integrated loudness and true peak once, store them as metadata, and let the player apply gain at playback. That keeps one encoded file per quality instead of one per loudness setting. See loudness normalisation for the measurement.
- Transcode the ladder. Encode each lossy tier and a lossless rendition. Encodes are independent jobs, so they parallelise across a worker pool keyed by track. Codec choice trades quality per bit against device support; MP3 vs AAC vs Opus walks through it.
- Encrypt. Encrypt each file with a per-file content key, store the key in a key service, and write only ciphertext to origin. Files can then sit on any CDN without exposing music.
- Publish. Write catalogue metadata and territory rights last, so a track never becomes visible before its files exist.
The playback path
When a listener taps a track, the client and control plane run a short protocol.
def resolve_playback(account, track_id, quality, device):
ent = entitlements.get(account) # plan, country, offline allowance
rights = catalogue.rights(track_id)
if ent.country not in rights.territories or rights.blocked(ent.plan):
return Unplayable(reason="not available") # client may skip or show alternative
q = min(quality, ent.max_quality, device.max_quality)
file_id = catalogue.file_for(track_id, q)
return Playback(
urls=cdn_router.signed_urls(file_id, region=ent.country, ttl_s=3600), # 2+ CDNs
key_ref=keys.wrap_for_device(file_id, device.public_key),
gain_db=catalogue.gain(track_id),
)The client fetches the first range from the first URL, decrypts, decodes and starts playing as soon as a small buffer fills, then fetches the rest in the background. Two client behaviours matter as much as any server. First, prefetch: once the current track is well under way, the client resolves and starts downloading the next item in the queue, so the transition is gapless and a slow network has time to catch up. Second, a local cache: recently played files stay on the device, so repeat listens never touch the network. Together they hide most server latency from the listener.
Signed URLs expire, and the client should treat expiry as normal: on a 403, call the playback service again rather than failing the track.
CDN strategy for whole files
In a February 2020 engineering post, Spotify described serving audio through Akamai and AWS, while other content such as images and client updates moved onto Fastly through a self-service tool its CDN team built, called SquadCDN. The design lesson is general. Audio is the critical, high-volume path and deserves dedicated, carefully tuned CDN capacity. Images and small API assets have different cache rules and change more often.
For whole-file audio, the cache key is the encrypted file ID, never the signed URL. The signature belongs in a token the edge validates and then strips, or every user gets a private cache miss. Popular tracks stay hot at every edge. New releases from big artists cause a predictable spike, so pre-warm the edges at release time by requesting the files from each region before the release goes live. Long-tail tracks miss often, so origin and any regional shield tier must be sized for tail misses, not for the average.
Use more than one CDN and let the playback service hand out URLs for two of them, in preference order. A client that times out on the first fails over to the second without another round trip to the control plane. CDN architecture in depth covers edge, shield and origin layers.
Offline downloads and keys
Downloads reuse the streaming files; what differs is key handling. The key service wraps each content key for the device, and the client stores wrapped keys alongside the files. The client must check in with the service periodically to renew a license. If the subscription has lapsed or the device has been removed, renewal fails and the keys are discarded. Spotify's help pages ask users to go online at least once every 30 days to keep downloads, which is this pattern from the user's side. The ciphertext stays on the device, so a lapsed user who resubscribes can resume without downloading again.
Playlists that survive concurrent edits
Playlists are small documents that change often, are edited from several devices and are sometimes shared with collaborators. Last-writer-wins on the whole list loses edits. A safer model stores each playlist as a revision number plus a list of items with stable item IDs, and accepts edits as operations against a base revision.
def apply(playlist, base_rev, ops):
# ops reference item ids, not positions, so concurrent edits rebase cleanly
if base_rev != playlist.rev:
ops = [rebase(op, playlist.log.since(base_rev)) for op in ops]
ops = [op for op in ops if op is not None] # e.g. move of an item someone removed
for op in ops:
if op.kind == "add":
playlist.insert_after(op.after_item_id, op.item_id, op.track_id)
elif op.kind == "remove":
playlist.remove(op.item_id)
elif op.kind == "move":
playlist.move_after(op.item_id, op.after_item_id)
playlist.rev += 1
playlist.log.append(playlist.rev, ops)
return playlist.revBecause operations name items rather than indexes, two people adding songs at the same time both succeed, and a move of a removed item simply drops. Devices sync by asking for the operation log since their last revision, which is cheap even for long playlists.
Play events and royalty counts
Spotify has written publicly about its event delivery system: when it moved to Google Cloud it redesigned the pipeline around Cloud Pub/Sub, and the old Kafka-based system was switched off in February 2017. Whatever the bus, play events must satisfy three requirements: they survive offline periods, they are delivered at least once, and they are counted exactly once.
The client writes each play to a local queue with a client-generated event ID, the track, the timestamps and the listened duration, and uploads batches when online. Retries are expected, so duplicates arrive. The counting job deduplicates on event ID and applies the business rule for what counts as a qualifying play.
QUALIFY_MS = settings.qualifying_play_ms # a business rule; read it from policy
def count(batch, seen, counts):
for ev in batch:
if seen.add_if_absent(ev.event_id, ttl_days=60) is False:
continue # duplicate upload; already counted
if ev.listened_ms >= QUALIFY_MS and not ev.flags.suspected_fraud:
counts.increment(ev.track_id, ev.country, ev.plan, day=ev.played_at.date())The deduplication window must be longer than the longest offline period you accept, or a device that comes back after weeks will be counted twice. Keep raw events immutable so a counting bug can be fixed by recomputation. The same stream feeds recommendation features, which can tolerate looser guarantees; recommendation system architecture covers that side.
Failure modes
- CDN outage in a region. Clients fail over to the second URL. Without multi-CDN URLs, playback stops for everyone not already cached on the device.
- Key service down. New plays cannot start even though the CDN is healthy. Cache wrapped keys on the client for recently played tracks and keep the service multi-region.
- Release-day thundering herd. Millions request the same new files within minutes. Pre-warm edges and collapse concurrent misses at the shield.
- Rights change mid-session. A track loses a territory. Entitlement checks at play time catch it; downloaded copies need the same check at license renewal.
- Double counting. A dedup store with too short a TTL counts late uploads twice, inflating royalties.
- Playlist divergence. Clients applying ops against stale revisions without rebasing. Symptom: duplicate or missing tracks across devices.
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Delivery unit | Whole file with range requests: simple, cache-friendly | Segmented adaptive streaming: switches quality mid-track, more objects |
| Encryption | Per-file keys: one key leak exposes one file | Shared keys: simpler, wider blast radius |
| Loudness | Gain metadata applied in player: one file per tier | Normalised encodes: no player work, more storage |
| Playlist sync | Operation log with revisions | Last-writer-wins documents: simple, loses edits |
| Play counting | Client event IDs and dedup: exact counts | Server-side counting at stream start: misses offline plays |
Segmented streaming, which video depends on, buys little for files of a few megabytes. The video streaming architecture shows the contrasting design.
What to do next
- Write down the quality ladder and compute per-track and catalogue storage with the script above.
- Estimate peak concurrent listeners and average bitrate to size CDN egress; assume origin sees tail misses only.
- Design the playback-resolve API with entitlement checks, two signed CDN URLs and a wrapped key.
- Implement client prefetch of the next queued track and a local file cache, and measure time to first audio.
- Build ingest as an ordered pipeline: validate, loudness, transcode, encrypt, then publish metadata last.
- Model playlists as revisioned operation logs keyed by item IDs.
- Make play events idempotent with client IDs, and set the dedup window longer than your offline limit.
- Run game days for CDN loss, key-service loss and a release-day spike.