File sync looks like upload plus download, but it is a different problem. Upload moves bytes from one client to a server once. Sync has to keep many replicas of a tree of files consistent over months, while every replica can change offline, files are renamed and moved, the network drops at arbitrary points, and the local disk is edited by programs that know nothing about sync. A sync engine that is wrong even once in a million operations deletes someone's thesis, so the design is built around being conservative and recoverable rather than around raw speed.

This article builds a Dropbox-style sync architecture from first principles. Where it cites Dropbox specifically, it uses Dropbox's public API documentation and what its engineers have written publicly; the rest is a reference design you can implement. The upload mechanics, multipart transfer and the security caveat in cross-user deduplication are covered in designing a file upload service and real-time co-editing in Google Drive's real-time collaboration architecture, so here the focus is the sync engine itself.

Advertisement

What a sync system must guarantee

  • Convergence: once changes stop and devices are online, every replica of a folder ends up with the same tree and the same bytes.
  • No silent data loss: when two edits conflict, both versions survive; the system never picks a winner by overwriting.
  • Efficiency: changing one paragraph of a 2 GB file must not transfer 2 GB, and a file already present in the account must not be uploaded again.
  • Resumability: any operation can be interrupted by a crash, sleep or network loss and resumed without corruption.
  • Bounded fan-out: a change in a folder shared with thousands of people must reach them without the server scanning each of their trees.

The architecture in two planes

The core idea is to separate metadata from content. Content is stored as immutable blocks addressed by their hash. Metadata says which path, in which folder, at which revision, consists of which list of block hashes. The metadata plane is small, strongly consistent and ordered; the content plane is huge, eventually replicated and needs no ordering at all, because an immutable block named by its hash can never be stale.

Sync in two planes: metadata (journal, cursors) decides what changed, blocks carry the bytesDesktop clientFile watcherOS change eventsHasher4 MiB blocks, SHA-256Plannerlocal / remote / syncedLocal DBtrees, cursor, cacheMetadata servicecommit, list changesJournalper-namespace logBlock serviceput / get by hashNotificationlong-poll: changes?Block storagecontent-addressedcommit(path, hashes)appendmissing blocksnew entryentries after cursorOther devices hold a long-pollopen and fetch the journalwhen it returns changes=true
The client hashes files into blocks, uploads only missing blocks, then commits metadata. The commit appends to the folder's journal; the notification service wakes other devices, which read journal entries after their cursor and fetch the blocks they lack.

On the server side, a metadata service accepts commits and answers list-changes queries, backed by a per-namespace journal: an append-only log of file events with a monotonically increasing position. A block service stores and serves blocks by hash, on top of object storage or, in Dropbox's case, its own storage system, which its engineers have written about under the name Magic Pocket. A notification service holds long-poll connections from idle clients and tells them when something changed.

Advertisement

Blocks and the content hash

Dropbox's public API defines exactly how it fingerprints a file, and the recipe is a good one to copy. Split the file into blocks of 4 MiB (4,194,304 bytes), the last one possibly shorter. Hash each block with SHA-256. Concatenate the binary digests and hash that string with SHA-256 again; the hex result is the file's content_hash. An empty file has no blocks, so its hash is SHA-256 of the empty string.

import hashlib

BLOCK = 4 * 1024 * 1024

def block_hashes(path):
    """Yield (offset, sha256 digest) for each 4 MiB block."""
    with open(path, "rb") as f:
        offset = 0
        while chunk := f.read(BLOCK):
            yield offset, hashlib.sha256(chunk).digest()
            offset += len(chunk)

def content_hash(path):
    """Dropbox-style content hash: SHA-256 over the concatenated block digests."""
    outer = hashlib.sha256()
    for _, digest in block_hashes(path):
        outer.update(digest)
    return outer.hexdigest()

Fixed-size blocks have a known weakness: inserting one byte at the start of a file shifts every block boundary, so every block hash changes. Content-defined chunking, which cuts where a rolling hash matches a pattern, survives insertions, at the cost of variable block sizes and more complex indexing. Fixed blocks win on simplicity and work well for the common cases: appends, in-place edits of databases and media files, and whole-file copies. A system can also layer a binary delta below the block layer for edited blocks, but that is an optimisation, not a requirement.

The upload path: blocks first, then commit

The ordering rule that makes sync safe is that metadata may only reference blocks that already exist. A client that wants to save a file sends a commit naming the path, the parent revision it was based on, and the list of block hashes. The server checks which hashes it lacks and answers with that list instead of committing. The client uploads exactly those blocks, then retries the commit, which now succeeds and appends a journal entry. If the client dies in between, uploaded blocks are orphans that garbage collection eventually removes; no reader ever sees a file pointing at missing data.

def save(path, parent_rev):
    hashes = [h for _, h in block_hashes(path)]
    while True:
        resp = meta.commit(path=path, parent_rev=parent_rev, blocks=hashes)
        if resp.status == "ok":
            local_db.mark_synced(path, resp.rev, hashes)
            return resp.rev
        if resp.status == "need_blocks":
            for h in resp.missing:
                blocks.put(h, read_block(path, hashes.index(h)))   # idempotent by hash
        elif resp.status == "conflict":
            return handle_conflict(path, resp.server_rev)

Every step here is idempotent: putting a block twice is harmless because its name is its content, and the commit is conditional on the parent revision, the same compare-and-set idea described in idempotency in system design. Dropbox's public API exposes the same shape for large files through upload sessions (upload_session/start, append_v2 and finish).

The journal and the cursor

Each shared folder, and each user's home, is a namespace with its own journal. A client does not ask "what does the tree look like?" on every check; it remembers a cursor, an opaque token encoding its position in the journals it follows, and asks for entries after it. The public API mirrors this: list_folder returns entries and a cursor, list_folder/continue returns changes since a cursor, and list_folder/longpoll blocks until there are changes or a timeout passes.

Long polling is what makes idle devices cheap. Instead of polling every few seconds, the client holds one request open to the notification service. When a commit lands in a namespace, the service completes the waiting requests for that namespace with a signal that changes exist; clients then call the metadata service for the entries. The notification carries no file data, so it can be lossy: a dropped signal only delays sync until the next long-poll returns. The general fan-out patterns are covered in notification systems.

Journals also make recovery precise. A cursor that is too old because the journal was compacted gets a reset response, and the client falls back to a full listing and a three-way comparison, slower but still correct.

On the client: three trees and a planner

The hardest code lives on the desktop. Dropbox has written publicly that its rewritten sync engine, Nucleus, models state as three trees: the remote tree (the latest server state the client knows), the local tree (what is on disk, as last observed) and the synced tree (the last state both sides agreed on). The synced tree is the merge base. Without it, a file present remotely but absent locally is ambiguous: was it deleted here, or created there? With it, the answer is mechanical.

def plan(path, local, remote, synced):
    """Return the action that moves this path toward convergence."""
    l, r, s = local.get(path), remote.get(path), synced.get(path)
    if l == r:
        return None                                   # already agree
    if l == s:                                        # only remote changed
        return ("download", path) if r else ("delete_local", path)
    if r == s:                                        # only local changed
        return ("upload", path) if l else ("delete_remote", path)
    # both changed since the merge base
    if l is None or r is None:
        return ("keep_existing", path)               # edit beats delete: never lose data
    return ("conflict_copy", path)

The planner runs repeatedly, applying one batch of safe actions, observing the result and planning again. Each applied action updates the synced tree only after it completes, so a crash leaves a state the next pass can finish. Local edits are detected by a filesystem watcher, but watchers drop events under load, so the client also rescans periodically and compares size, modification time and inode before re-hashing.

Conflicts, moves and deletes

When both sides edited the same file, a sync engine cannot merge arbitrary binary formats, so it keeps both: the server version stays at the path and the local version is saved alongside with a name that marks it as a conflicted copy, plus device and date. Users dislike conflicted copies, but they dislike lost work far more. Edit-versus-delete resolves toward keeping the edit.

Moves deserve care. Treating a folder move as delete plus create would re-upload nothing, because blocks dedupe, but would generate a storm of journal entries and break sharing links. A good engine detects moves, locally through inode or file identifiers and remotely through journal entries that carry a stable file id, and commits a single move. It must also handle cycles (two devices each moving folder A into B and B into A) by rejecting the second move. Deletes are soft on the server: a deleted file keeps its revision history for a retention period, which is the safety net when a planner bug or ransomware deletes things in bulk.

Worked example: one edit, three devices

A user edits a 1 GiB video project file on a laptop; the edit rewrites 30 MiB in the middle, starting on a block boundary. The file is 256 blocks of 4 MiB. The client re-hashes it and finds that 8 block hashes changed. It sends a commit with 256 hashes; the server answers with the 8 it lacks; the client uploads 32 MiB, commits again and gets revision 41. The journal of the user's namespace gains one entry.

The desktop at home is idle with a long-poll open. It returns, the desktop asks for entries after its cursor, sees revision 41 with its block list, finds 248 of the blocks already in its local cache, downloads 8, writes the new file to a temporary path, verifies the content hash and renames it into place atomically. Its synced tree now records revision 41.

The phone was offline and the user, on a train, renamed the same file. When it reconnects, its planner sees: local changed (path differs), remote changed (content differs), synced is the old state. These touch different attributes, path and content, so the engine commits a move based on revision 41 and both changes survive. Had the phone also edited the content, the result would be a conflicted copy, and both versions would be visible to the user.

Shared folders and namespaces

Sharing is why the journal is per namespace rather than per user. A shared folder is one namespace mounted into many users' trees at possibly different paths. A commit appends once to that namespace's journal, and each member's cursor covers the namespaces they mount. This keeps write cost independent of member count: the fan-out happens on the read side, through notifications to subscribed clients. Permissions are checked against the namespace on every read and commit, and removing a member removes the mount without touching the journal.

Failure modes and how the design absorbs them

  • Crash between block upload and commit: orphaned blocks, collected later; no reader impact.
  • Partial download: files are written to a temporary path and renamed only after the content hash matches, so a torn file is never visible.
  • Missed watcher events: periodic rescans and the three-tree comparison repair state without trusting the event stream.
  • Clock skew: ordering comes from journal positions and revisions, never device clocks; timestamps are shown to humans only.
  • Hot shared folder: one namespace with constant writes can dominate a metadata shard; the fix is splitting namespaces and rate-limiting per namespace, not per user.
  • Planner bug deletes data: server-side retention of deleted revisions, mass-delete detection that pauses sync and asks the user, and staged client rollouts.
  • Case and encoding mismatches: macOS and Windows file systems are usually case-insensitive and normalise Unicode differently from Linux, so two remote names can map to one local path; the client must detect the collision and rename one locally.

Trade-offs

DecisionGainCost
Fixed 4 MiB blocksSimple indexing, cheap hashing, good dedupe for appends and copiesInsertions shift boundaries; small files carry per-block overhead
Commit after blocksReaders never see dangling referencesTwo round trips per save; orphan collection needed
Per-namespace journalWrite cost independent of member count, exact cursorsClients follow many journals; hot namespaces need sharding
Conflicted copiesNo lost workUsers must merge by hand
Long-poll notificationIdle devices cost almost nothingMany open connections to manage at the edge

What to do next

  1. Implement the content-hash function above and test it against a file whose hash you get from the Dropbox API or your own server, including the empty file.
  2. Design your commit as conditional on a parent revision and returning missing blocks; write a test that kills the client between upload and commit.
  3. Store a synced tree alongside local and remote state, and write the planner as a pure function with table-driven tests for every combination of create, edit, delete and move.
  4. Give each shared folder its own journal and cursor position; load test one hot namespace with many members.
  5. Add retention of deleted revisions and a mass-delete circuit breaker before shipping to real users.
  6. Test on case-insensitive file systems and with Unicode names that normalise differently.
Key takeaway: A sync engine is a metadata problem with a content cache attached. Address blocks by hash, commit metadata only after its blocks exist, order changes in a per-namespace journal read through cursors, wake idle clients with long-polls, and plan on the client with three trees so every difference has an unambiguous meaning. When in doubt, keep both versions: convergence can wait, lost data cannot be undone.