Open a document, and a colleague on another continent starts typing into the same paragraph. You see their characters within a fraction of a second, your own keystrokes never wait for the network, nobody loses text, and an hour later you can scroll back through the revision history. Behind that experience sit two very different systems: a file service that knows who owns the document, where it lives in a folder tree and who may open it, and an editing service that keeps many copies of the live text converging while everyone types at once.

Google has described parts of this publicly, most notably that the Docs editor uses operational transformation (OT) with the server as the arbiter of order. It has not published its internal service layout, so this article presents a reference architecture that is consistent with those public descriptions and with how collaborative editors are generally built. Treat component names here as roles, not Google product internals. The goal is that you can build, operate or interview on a system like it. The transform functions themselves are covered in the operational transformation architecture article; this one is about everything around them.

Advertisement

Requirements that shape the design

Four requirements drive almost every decision. Local-first latency: a keystroke must render immediately, so the client edits its own copy and synchronises afterwards. Convergence with intent: when two people edit concurrently, every replica must end with the same text, and that text should reflect what each person meant, for example both insertions surviving in sensible positions. Durability and history: an acknowledged edit must never disappear, and past versions must be recoverable. Access control that can change at any moment: an owner can revoke a share while the other person is typing.

Scale is uneven. Most documents have one editor at a time; a few (a meeting agenda, a shared incident doc) have dozens of simultaneous editors and hundreds of viewers. The system must be cheap for the idle majority and robust for the hot minority, which pushes toward state that is loaded on demand when someone opens a document and dropped when the last session closes.

Two planes: the file and the live document

The file plane is Drive-like metadata: a file id, owner, parent folders, sharing permissions, title, trash state, quota accounting and a pointer to the content. It changes rarely and needs strong consistency and rich queries (list this folder, search my files). The editing plane is a stream of small operations against one document at high frequency, with strict ordering inside a document and no ordering needs across documents.

Keeping them separate lets each use the right storage and scaling model. The editing plane consults the file plane when a session opens (can this user view or edit?) and subscribes to permission changes; the file plane learns about content changes asynchronously (update modified time, re-index for search, refresh thumbnails). The Drive public API reflects the same split: files, permissions and a changes feed are file-plane concepts, while the document body is reached through a separate Docs API.

Two planes: Drive owns the file, the editing service owns the live documentEditor client Alocal model + pending opsEditor client Blocal model + pending opsops, acksops, acksSession gatewayWebSocket / long-pollroute by doc idDocument ownerone ordering point per doctransform, assign revision,append, broadcastOp logrev n, n+1, ... durableSnapshotsevery K revsPresence buscursors, ephemeralDrive metadatafile id, parents, owner, ACLRevision historygrouped named versionsSearch, thumbnails, exportasync consumers of revsACL check on openLatency-critical path: client -> gateway -> owner -> log. Everything below the log is asynchronous.
Reference architecture. Only the path from client through gateway to the document owner and its op log is latency-critical; history, search, thumbnails and exports consume the log asynchronously.
Advertisement

The protocol: one ordering point per document

Peer-to-peer OT is notoriously hard to get right because every pair of concurrent operations must transform correctly in every order. A central server removes most of that difficulty. The server assigns every operation a revision number, so there is a single total order per document, and each client only ever has to reconcile its own unacknowledged work against operations the server has already ordered. This is the Jupiter model, published in 1995 and the lineage cited for Google's editors.

The client keeps three things: the last server revision it has seen, at most one operation in flight, and a buffer of edits typed while waiting for the ack. Allowing only one in-flight operation is what keeps the algorithm simple: the server receives ops from a client in order, each stated against a known base revision.

# Client side of a central-server OT protocol (Jupiter-style), simplified.
class Client:
    def __init__(self, doc, rev):
        self.doc, self.rev = doc, rev   # last revision the server confirmed
        self.inflight = None            # at most one op awaiting ack
        self.buffer = None              # local edits composed while waiting

    def local_edit(self, op):
        self.doc = apply(self.doc, op)                  # instant local echo
        if self.inflight is None:
            self.inflight = op
            send(op, base_rev=self.rev, client_seq=next_seq())
        else:
            self.buffer = op if self.buffer is None else compose(self.buffer, op)

    def on_ack(self, new_rev):
        self.rev = new_rev
        self.inflight, self.buffer = self.buffer, None
        if self.inflight is not None:
            send(self.inflight, base_rev=self.rev, client_seq=next_seq())

    def on_remote(self, op, new_rev):
        # op is already ordered by the server; move it past our unacked work
        if self.inflight is not None:
            op, self.inflight = transform(op, self.inflight)
        if self.buffer is not None:
            op, self.buffer = transform(op, self.buffer)
        self.doc = apply(self.doc, op)
        self.rev = new_rev

The server side is equally small. It transforms the incoming op against everything committed since the client's base revision, makes it durable, assigns the next revision, acknowledges the sender and broadcasts to everyone else. The (client_id, client_seq) pair makes the operation idempotent: if the ack is lost and the client resends after reconnecting, the server returns the original revision instead of applying the edit twice.

# Document owner: the only place where revision numbers are assigned.
def receive(doc_id, client_id, client_seq, op, base_rev):
    d = owned_docs[doc_id]                  # this process holds the lease
    if d.seen(client_id, client_seq):       # retry after a lost ack
        return ack(d.rev_of(client_id, client_seq))
    authorize(client_id, doc_id, "edit")    # ACL may have changed mid-session
    for concurrent in d.log.since(base_rev):   # ops the client had not seen
        op, _ = transform(op, concurrent)
    validate(op, d.current_len)            # reject out-of-range ranges
    rev = d.log.append(op, client_id, client_seq)   # durable before ack
    d.apply_in_memory(op)
    broadcast(doc_id, op, rev, exclude=client_id)
    if rev % SNAPSHOT_EVERY == 0:
        schedule_snapshot(doc_id, rev)
    return ack(rev)

Worked example: two edits to the same word

The document is cat at revision 10. Alice places her cursor at position 0 and types A, producing insert(0, 'A') against base 10. At the same moment Bob deletes the t: delete(2, 1) against base 10. Both see their own edit instantly: Alice sees Acat, Bob sees ca.

Alice's op reaches the owner first. Nothing is committed after revision 10, so it is appended unchanged as revision 11 and broadcast. Bob's op arrives with base 10; the owner transforms it against revision 11. An insertion at position 0 shifts every later position right by one, so delete(2, 1) becomes delete(3, 1), which is appended as revision 12. The owner state is now Aca.

Alice receives revision 12 as delete(3, 1), applies it to Acat and gets Aca. Bob receives revision 11 while his delete is still in flight. He transforms the remote insert(0, 'A') against his pending delete(2, 1): the insert is before the deleted range, so it is unchanged, and his pending op becomes delete(3, 1), matching what the server will commit. He applies the insert to ca and gets Aca. When his ack for revision 12 arrives, he just advances his revision. Three replicas, one result, no locks.

Real documents are rich text, so operations also carry formatting attributes, paragraph styles, tables and embedded objects. The shape stays the same, but the transform table grows, and it is the part of the system that needs the most property-based testing.

Storage: an op log with snapshots

The durable truth for a document is its ordered op log. Appending one small record per revision is cheap and gives revision history, audit and replay for free. Replaying millions of ops on every open would be slow, so the owner periodically writes a snapshot: the full document state at some revision. Opening a document means loading the latest snapshot and the log tail after it.

Two details matter. First, the ack must follow the durable append, or a crash between ack and write loses an edit the user believes is saved. Second, snapshots should be written asynchronously and be idempotent, because they are an optimisation; a missing snapshot only makes the next open slower. Old ops can then be compacted, but revision history usually keeps coarser checkpoints (grouped by time and author) rather than every keystroke, which is also what users want to see.

Everything else hangs off the log. Search indexing, thumbnail rendering, export and the file plane's modified time are consumers that read new revisions at their own pace, so a slow indexer never slows typing.

Routing, sessions and hot documents

Because ordering is per document, the natural unit of ownership is one document on one process at a time, enforced by a lease in a coordination store. Gateways terminate client connections (WebSocket, falling back to long-polling where needed) and route each message by document id to the current owner. On first open, some owner acquires the lease and loads state; when the last session closes, it releases memory after a grace period.

The lease is the correctness hinge. If two processes both believe they own a document, both assign revision 13 and histories fork. Use leases with fencing tokens and make the log append conditional on the token and the expected next revision, so a stale owner's write fails instead of forking.

A hot document concentrates load on one owner. The ordering work per op is small, but fan-out is not: each accepted op goes to every connected session. Common mitigations are batching several ops per broadcast frame, coalescing a viewer's updates, and serving read-only viewers from replicas that tail the log rather than from the owner. Presence (cursors, selections, who is here) travels on a separate ephemeral channel with no durability and aggressive rate limiting; see presence tracking for that pattern.

Permissions, comments and suggestions

Access is checked when a session opens, but a check-once model is wrong: the owner can remove a collaborator or downgrade them to commenter mid-session. The editing service should subscribe to permission changes from the file plane and, on a downgrade, reject further edit ops and close or downgrade the session. The server-side authorize call in the pseudocode can be a cached lookup invalidated by those events.

Comments and suggestions anchor to ranges of text that keep moving. Store the anchor as a range at a specific revision and transform it through subsequent ops exactly as you would transform a cursor. If the anchored text is deleted, the comment becomes orphaned and should be shown as such rather than silently re-anchored somewhere unrelated. Suggestion mode is simply edits tagged as proposals, rendered differently and accepted or rejected later by another op.

Offline editing and reconnect

Offline support falls out of the client protocol. The client persists its document, last revision and pending ops locally. On reconnect it asks the server for all ops after its revision, transforms its pending work against them exactly as it would for live remote ops, and resubmits. The idempotency key prevents duplicates if some ops had in fact been committed before the disconnect.

The limit is divergence size. After days offline against a busy document, the transformed result may be technically convergent but surprising to users. Products typically cap how long offline edits may queue, warn when reconnecting after heavy concurrent change, and preserve the offline version in history so nothing is unrecoverable.

Failure modes

  • Ack before durability. An owner crash loses an acknowledged edit and the client's state now contains text the server never had. Append, then ack.
  • Split-brain owners. Two processes accept edits for one document and histories fork. Fence log appends with the lease token and an expected revision.
  • Transform bugs on rare op pairs. Replicas silently diverge. Send a periodic checksum of the document at a revision and force a resync from the server when it mismatches; fuzz the transform table with random concurrent histories.
  • Reconnect storms. After a gateway deploy, thousands of clients reconnect and request log tails at once. Add jittered backoff and serve tails from snapshot plus log without waking every document's owner.
  • Unbounded broadcast queues. A slow client on a busy document accumulates ops in memory. Bound per-session queues and switch that client to snapshot resync when it falls too far behind.
  • Revoked access still editing. A cached ACL outlives a revocation. Invalidate on permission events and re-check on every op path.

OT with a central server versus CRDTs

AspectCentral-server OTCRDT (for example Yjs or Automerge)
OrderingServer assigns a total order per documentNo required order; merge is commutative
Offline and peer-to-peerNeeds the server to commitWorks offline and peer-to-peer natively
MetadataSmall ops; state is plain textPer-character ids and tombstones; needs compaction
Server roleArbiter; must be available to commitRelay and store; can be dumb
ComplexityTransform functions for every op pairData-type design and garbage collection

If your product is always-online and server-authoritative, as a Docs-style editor is, central OT is simple and compact. If local-first or peer-to-peer matters, CRDTs are the better fit; CRDTs in production and CRDT replication architecture cover that path.

What to do next

  1. Implement the client and server pseudocode above for plain text with insert and delete, and fuzz it with random concurrent histories until every replica converges.
  2. Add the idempotency key and test lost acks by dropping responses and forcing reconnects.
  3. Put a lease with fencing in front of the log append and kill owners mid-write to prove no fork occurs.
  4. Add snapshots every K revisions and measure open latency with and without them.
  5. Wire permission-change events into live sessions and test a downgrade while a user is typing.
  6. Add a periodic document checksum so divergence is detected and repaired, not discovered by users.
Key takeaway: A Docs-style editor is two systems: a file plane that owns identity, folders and sharing, and an editing plane that keeps live copies convergent. The editing plane is simplest with one ordering point per document: clients apply edits locally, keep one op in flight, and transform remote ops against their pending work, while the owner transforms incoming ops against what was committed since their base revision, appends durably, then acks and broadcasts. An op log with snapshots gives history and fast opens; leases with fencing prevent forks; permission events, idempotent resubmission and checksums handle the failures that matter in production.