Snapchat made disappearing messages a mainstream product: a photo or a chat that is gone after it has been seen. Building that well is harder than building ordinary messaging, because a normal messaging system is designed never to lose data, with replicas, caches, backups, search indexes and analytics copies everywhere, and an ephemeral system has to make all of those copies go away on a schedule while still delivering every message reliably.

This article is a design for Snapchat-style ephemeral messaging. It is not a description of Snap's internal services, which Snap has not documented in enough detail to reproduce; it uses the product's public behaviour as requirements and standard distributed-systems techniques as the solution. It covers the promise and its limits, the data model, view-state tracking, expiry scheduling, crypto-shredding, saves and legal holds, client behaviour and the ways deletion silently fails. Connection handling and the message store for chat in general are covered in designing a real-time chat system.

Advertisement

What ephemeral can promise, and what it cannot

Start with honesty about the guarantee, because it shapes the architecture. A server can promise that it deletes content on a schedule: from primary storage, replicas, caches and CDNs, and, with the right design, effectively from backups. A server cannot promise that the recipient keeps nothing. Screenshots, a second camera, a rooted phone or a modified client defeat any server design. Screenshot detection, where a platform offers it, is a best-effort notification to the sender, not a control, and a product should never describe it as one.

So the system has two jobs: deliver reliably to the intended recipients, and then forget reliably. Snapchat's support pages describe the forgetting as a set of per-type retention windows: content is deleted from servers after all recipients have viewed it, unopened content is deleted after a longer window, group content has its own window, and Stories expire after about a day. Those exact windows have changed over the years, which is the first design lesson: retention windows are configuration, not code.

Requirements and a rough scale

  • One-to-one and group messages; media (photos, short video) and text chat.
  • Per-recipient view state: delivered, opened, replayed; the sender sees it.
  • Deletion when all recipients have viewed (plus a configurable grace period), or when an unopened window expires, whichever comes first.
  • Saving a message in-app exempts it from automatic deletion; a legal preservation order overrides every deletion path.
  • Multi-device clients that may be offline for days.
  • Illustrative scale for sizing, not a Snap figure: 300 million daily senders, 5 billion messages a day, about 58,000 per second on average and several times that at peak; median media object 500 KB.

The scale numbers say something useful. At 5 billion messages a day, deletions are as frequent as writes. Expiry is not a cleanup cron job; it is a second write path of equal volume that needs its own capacity planning, back-pressure and monitoring.

Advertisement

Architecture: separate metadata, media and keys

The central idea is to keep three things apart: small metadata (who sent what to whom, and who has opened it), large media (encrypted blobs), and the keys that decrypt the media. Deleting a message then becomes a small write against the key store, and every leftover copy of the media, wherever it is, becomes noise.

Ephemeral messaging: metadata, media and keys live apart so deletion can be one small writeSender appencrypts mediaMessaging APIauth, policyMessage storemetadata, view stateKey storeone key per mediaObject storeciphertext onlyPush serviceno content in payloadExpiry schedulerdelete_at indexDelete workersidempotentCDN edgesigned, short-livedRecipient appdecrypts, deletessendwrap keyuploadnotifyon opendueshred keyfetchCopies left behind in caches, CDN, replicas and backupsare ciphertext; once the key is gone they are unreadable.
The sender uploads ciphertext directly to object storage; the message store tracks view state; the expiry scheduler fires when the policy says the message is due; delete workers destroy the media key first and clean up storage afterwards.
  1. Send. The client generates a random key for the media, encrypts locally, and uploads the ciphertext to object storage through a short-lived signed URL obtained from the messaging API. It then posts the message metadata, including the media reference and the key (wrapped for the key store, or for recipients if the system is end-to-end encrypted).
  2. Store. The message store writes one message row and one recipient-state row per recipient, partitioned by conversation so that view-state updates and the deletion decision touch one partition.
  3. Notify. A push notification wakes the recipient. The payload carries an id, never content or a preview, because push providers and device notification histories are copies you do not control. The notification pattern is covered in notification system design.
  4. Fetch and open. The recipient fetches ciphertext via a signed, short-lived URL, possibly through a CDN, obtains the key, decrypts in memory and displays. Opening sends an event that updates view state.
  5. Expire. Each view-state change recomputes the message's deletion time and upserts it into a scheduler. When it fires, a delete worker shreds the key, tombstones the metadata and removes the blob.

Deletion policy as a pure function

Rules like "after everyone has viewed it, or after N days unopened, unless saved, unless on legal hold" accumulate exceptions quickly. Encode them in one pure function of the message's state and a versioned policy, and call it from every place that might delete. Never let a queue entry or a cache be the authority on whether something is due.

from dataclasses import dataclass
from datetime import datetime, timedelta
from typing import Optional

@dataclass(frozen=True)
class Policy:                      # loaded from config, versioned; values are product decisions
    after_all_viewed: timedelta    # e.g. 0 for a one-view snap, longer for a chat setting
    unopened_ttl: timedelta        # cap for content nobody opens
    group_unopened_ttl: timedelta

@dataclass
class RecipientState:
    user_id: str
    opened_at: Optional[datetime] = None

@dataclass
class Message:
    message_id: str
    sent_at: datetime
    is_group: bool
    recipients: list
    saved: bool = False            # someone saved it in-app: no automatic expiry
    legal_hold: bool = False       # preservation order: never delete while set

def delete_at(m: Message, p: Policy) -> Optional[datetime]:
    # Pure function: the moment this message becomes deletable, or None for never (for now).
    if m.legal_hold or m.saved:
        return None
    ttl = p.group_unopened_ttl if m.is_group else p.unopened_ttl
    hard_cap = m.sent_at + ttl
    opened = [r.opened_at for r in m.recipients]
    if all(opened):
        return min(max(opened) + p.after_all_viewed, hard_cap)
    return hard_cap

Purity buys three things. It is trivially unit-testable against every product rule. The scheduler can be rebuilt from the message store at any time, because the function recomputes every deadline. And when the product changes a window, a backfill job can recompute deadlines for in-flight messages instead of leaving them on the old rule. Policy changes should apply only in the direction users expect: shortening a window can apply retroactively, lengthening one generally should not, because users sent those messages under the shorter promise.

View state and the scheduler

The open event is the hot path. It arrives from flaky mobile networks, often more than once, and in groups it arrives from many users in any order. Make it a conditional write so only the first open counts, and recompute the deadline on the same partition:

def on_opened(store, scheduler, message_id, user_id, now):
    # Conditional write: only the first open counts; a retried or replayed event is a no-op.
    changed = store.update_recipient(
        message_id, user_id,
        set={"opened_at": now},
        condition="opened_at IS NULL")
    if not changed:
        return
    m = store.get_message(message_id)               # read after write, same partition
    when = delete_at(m, policy_for(m))
    scheduler.upsert(message_id, when)              # None removes the entry

Retries make this idempotent by construction. The scheduler itself can be a table indexed by deadline and sharded by time bucket and hash, a delayed-message queue, or the store's native per-row TTL. Native TTL is attractive and has a catch: most databases expire rows lazily, during compaction or a background sweep, so a TTL'd row can stay readable, or remain on disk, for hours or days after its deadline. Use native TTL as a backstop that guarantees eventual removal, and an explicit scheduler for the deadline you promise. Size the scheduler for peak write volume and make lag a first-class metric: expiry lag, the gap between deadline and actual deletion, is the number that tells you whether the product promise holds.

Crypto-shredding: deleting copies you cannot find

Media gets copied everywhere: object-store replicas in several regions, CDN edges, transcoding outputs, thumbnails, backups and disaster-recovery snapshots. Chasing each copy at deletion time is fragile, and backups are often immutable by design. Crypto-shredding solves this by making sure every copy is ciphertext under a per-message key that lives in exactly one small, tightly controlled store. Destroy the key and every copy is unreadable at once.

def expire(message_id, store, keys, objects, cdn, audit, now):
    m = store.get_message(message_id)
    if m is None:
        return                                       # already gone: idempotent
    when = delete_at(m, policy_for(m))               # re-evaluate; never trust the queue entry
    if when is None or when > now:
        return                                       # saved, held or opened late: not due
    keys.destroy(m.media_key_id)                     # 1. crypto-shred first: content is now unreadable
    audit.record("shredded", message_id, now)        #    log the fact, never the content
    store.tombstone(message_id, now)                 # 2. metadata and view state
    objects.delete(m.media_object)                   # 3. best effort, retried; ciphertext only
    cdn.purge(m.media_path)                          # 4. belt and braces; URLs were short-lived anyway

The order is deliberate. Shredding first means that a crash after step 1 leaves only unreadable leftovers that the retry cleans up; deleting the blob first and crashing would leave a live key pointing at nothing, which is harmless but means the shred never happened if the retry logic is wrong. The key store needs its own care: it must be replicated for durability, so it has backups too. Keep its backup retention shorter than the shortest promise you make, or wrap message keys under rotating epoch keys and destroy each epoch key once every message under it has expired, so old key backups become useless on a schedule. CDN copies are handled the same way: serve only ciphertext, use signed URLs with expiry shorter than the shortest retention window, and keep CDN TTLs short, as discussed in CDN design.

If the system is end-to-end encrypted, the server never holds the media keys at all, and the shred happens on the clients; the server's job shrinks to deleting ciphertext and metadata. The protocol side of that design is covered in the Signal double ratchet article.

Clients, saves and legal holds

Clients are part of the deletion system. The app must delete its local copy and any decrypted cache file when the message expires, including on devices that were offline at the deadline. Treat the server's state as authoritative: on reconnect, the client syncs expired message ids and purges them before showing anything. Decrypt into memory where the platform allows, exclude media caches from device backups, and never write plaintext to shared storage.

Saving changes the policy rather than copying data: the message's saved flag makes the deadline None, and un-saving recomputes it. A legal preservation request is the one path that can override a user-visible promise, so it needs the strongest controls: a separate, audited service sets legal_hold on specific accounts or messages, delete workers check it on every run because the function does, and hold removal recomputes deadlines. Content already shredded before a hold arrives is gone; that is the correct, documented outcome.

Failure modes: how deletion silently fails

  • Leaky side channels. Content or previews in application logs, error-reporting payloads, push notification bodies, search indexes, analytics events or ML training sets. These are the most common way ephemeral systems break their promise. Log ids, never content, and review every new data consumer.
  • Replica resurrection. A deleted row reappears from a lagging replica or a repaired node because the tombstone expired before every replica saw it. Keep tombstones longer than your maximum replica repair interval.
  • Scheduler lag. A backlog in delete workers quietly turns a one-day promise into a three-day reality. Alert on expiry lag percentiles, not on queue length.
  • Key destroyed too early. A group recipient on a slow network fetches the blob after the key is shredded. Compute deadlines from server-observed state only and include a grace period for in-flight fetches.
  • Clock skew. Deadlines computed from client timestamps can be manipulated or wrong. Use server receive times.
  • Partial delete fan-out. Blob deleted in one region, not another. Because the key is gone, it is unreadable; a periodic sweeper finds orphaned blobs by listing objects without live metadata.

Trade-offs

DecisionOption AOption B
Expiry mechanismExplicit scheduler: precise, extra systemNative TTL: simple, lazy and imprecise
Key custodyServer key store: server can enforce shredding and holdsEnd-to-end: stronger privacy, deletion relies on clients
Group deletionAfter all members view: fair, keeps data longerPer-member windows: earlier deletion, more state
BackupsEncrypted, keys shredded: cheap, effectiveRewrite backups to remove data: slow, error-prone

What to do next

  1. Write the product's retention promises as a versioned policy object and implement deletion as one pure function over message state.
  2. Store media only as ciphertext under per-message keys, and keep those keys in a small, separately controlled store.
  3. Build expiry as a capacity-planned write path with idempotent workers, and alert on expiry-lag percentiles.
  4. Inventory every data consumer, including logs, push payloads, analytics and search, and remove content from each.
  5. Set key-store backup retention or epoch-key rotation so that no backup outlives the longest promise.
  6. Test the uncomfortable cases: offline recipients, late opens in groups, saves after expiry, legal holds arriving mid-deletion and region failover during deletion.
Key takeaway: Ephemeral messaging is ordinary reliable messaging plus an equally reliable forgetting path. Keep metadata, media and keys apart; compute every deadline from one pure policy function over server-observed view state; destroy the per-message key first so every cached, replicated and backed-up copy becomes unreadable; and spend as much engineering on the leaks, such as logs, push previews and analytics, as on the deletion itself, because they are where the promise usually breaks.