When you edit a note on an iPhone and it appears on a Mac a few seconds later, several systems have cooperated: a local database, a sync engine running inside the app, Apple's push notification service and CloudKit, the structured storage service behind most iCloud app data. The user sees one experience. An engineer building a sync feature sees a protocol with specific guarantees, specific limits and a few sharp edges.
This article explains that protocol from first principles. It covers how CloudKit organises data, how devices learn what changed without downloading everything, how concurrent edits are detected and merged, what Apple's CKSyncEngine now does for you, what Apple has published about the storage underneath, and where real apps break. It ends with a worked example and a checklist. The design is worth knowing even if you never ship an Apple app, because it is a clean example of offline-first, server-ordered sync; compare it with Dropbox's file sync, which solves a similar problem for files.
What "iCloud sync" actually means
iCloud is not one sync system. iCloud Drive syncs files and documents. A small key-value store syncs preferences. Apple's own apps use various internal services. For third-party app data, the main path is CloudKit, either directly or through NSPersistentCloudKitContainer, which mirrors a Core Data store into CloudKit, or SwiftData, which builds on the same mirroring. This article is about CloudKit, because it is the layer with a documented protocol that developers program against.
CloudKit's central idea is that the server is the ordering authority and devices are replicas with local storage. Each device keeps a full working copy of the data it cares about, edits it locally without waiting for the network, and exchanges changes with the server when it can. The server does not merge data for you. It records changes in order, rejects writes based on stale versions, and tells each device what has changed since it last asked. Merging is the app's job.
The data model: containers, databases, zones and records
A container is an app's namespace, identified by a string such as iCloud.com.example.notes. Several apps from the same team can share one. Each container has three databases per user. The public database holds data readable by every user of the app and counts against the developer's quota. The private database holds one user's data and counts against that user's iCloud storage. The shared database is a per-user view of zones or records other people have shared with them.
Within a database, records live in record zones. Every database has a default zone, but sync-friendly apps create custom zones in the private database, because custom zones keep a change history the server can answer delta queries against, and they support atomic commits of a batch of records. A zone is also the natural unit of sharing: since iOS 15 a whole zone can be shared with a CKShare.
A record (CKRecord) is a typed dictionary of fields: strings, numbers, dates, locations, lists, references to other records and CKAsset values for large binary files, which are uploaded separately. Each record carries system fields, including a server-assigned change tag that changes on every save. That tag is the version used for optimistic concurrency. Record types and fields form a schema that is flexible in the development environment and must be deployed to production, where changes are additive: you can add fields and types, but not remove or retype what production already relies on.
Delta sync with change tokens
The fetch path is built on server change tokens, opaque cursors into the server's change history. A device stores one token per database and one per zone. Fetching is a two-level walk:
- Ask the database for changes since the stored database token. The answer is a list of zones that changed or were deleted, plus a new database token.
- For each changed zone, ask for record changes since that zone's token. The answer is a stream of modified records and deleted record IDs, plus a new zone token, possibly in several pages with a more-coming flag.
- Apply the changes to the local store, then persist the new tokens in the same local transaction.
The last step is the one apps get wrong. If tokens are saved before the data is applied and the app crashes in between, the device skips changes forever. If data is applied but the token is not saved, the device re-fetches and must apply the same changes idempotently. Saving both in one transaction makes the fetch exactly-once from the app's point of view. The pattern is identical to a database replication log consumer, or to the journal cursor in production CRDT systems.
Devices learn that a fetch is worthwhile through a CKDatabaseSubscription, which makes the server send a silent push through APNs when anything in the database changes. Pushes are hints, not data: they can be coalesced, delayed or dropped, so an app must also fetch on launch and on returning to the foreground. Treat push as a latency optimisation layered over polling, the same way notification systems treat delivery generally.
Writes and conflicts: optimistic concurrency
Writes go the other way. The app records that a local object changed, and later sends a batch of saves and deletes. Each saved record carries the change tag the device last saw. With the default save policy, the server accepts the save only if that tag still matches; otherwise it rejects the record with serverRecordChanged and returns the current server record in the error. This is compare-and-set at record granularity, and it means a stale device can never silently overwrite a newer version.
Rejected saves must be merged by the app. The usual approach is field-level: for each field, decide whether the local value or the server value wins, then copy the decisions onto the server record, which carries the current change tag, and save that. Pure last-writer-wins by device clock is tempting but unsafe, because device clocks drift; prefer rules tied to the data's meaning, such as union for sets of tags, maximum for monotonic counters, and a user-visible conflict copy for long free text.
// Field-level merge when a save is rejected with .serverRecordChanged.
func merge(local: NoteRecord, serverRecord: CKRecord) -> CKRecord {
// Start from the server copy: it carries the current change tag.
let merged = serverRecord
// Copy only the fields this device changed; keep the server's other fields.
if local.dirtyFields.contains("title") {
merged["title"] = local.title as CKRecordValue
}
if local.dirtyFields.contains("body") {
merged["body"] = local.body as CKRecordValue
}
// Tags: union, so neither device loses a tag.
let serverTags = Set(serverRecord["tags"] as? [String] ?? [])
merged["tags"] = Array(serverTags.union(local.tags)).sorted() as CKRecordValue
return merged
}When both devices changed the same field, this code lets the later save win. To break such ties deliberately, store a per-field logical counter or a vector clock, or model text as a CRDT in a binary field so concurrent edits merge without a winner.
CKSyncEngine: the scheduler Apple now ships
Until 2023, every app wrote its own loop around the fetch and modify operations: tracking tokens, scheduling retries, creating the subscription and handling account changes. CKSyncEngine, available from iOS 17, macOS 14 and watchOS 10, packages that loop. According to Apple's documentation, the engine schedules syncs when system conditions are good, creates or discovers the database subscription, retries transient errors such as networkFailure, requestRateLimited and zoneBusy after the server's retry-after interval, and sends each batch of up to 250 records as one request. It does not resolve serverRecordChanged for you, and Apple says not to use it for the public database.
The app's responsibilities shrink to four: tell the engine what changed, build record batches on request, apply fetched changes, and persist the engine's opaque state whenever it changes.
final class NotesSync: CKSyncEngineDelegate {
let store: LocalStore
var engine: CKSyncEngine!
init(store: LocalStore) {
self.store = store
let config = CKSyncEngine.Configuration(
database: CKContainer(identifier: "iCloud.com.example.notes").privateCloudDatabase,
stateSerialization: store.loadSyncState(), // nil on first launch
delegate: self)
engine = CKSyncEngine(config)
}
func noteDidChange(_ id: CKRecord.ID) {
engine.state.add(pendingRecordZoneChanges: [.saveRecord(id)])
}
func handleEvent(_ event: CKSyncEngine.Event, syncEngine: CKSyncEngine) async {
switch event {
case .stateUpdate(let update):
store.saveSyncState(update.stateSerialization)
case .fetchedRecordZoneChanges(let changes):
store.transaction {
for m in changes.modifications { store.upsert(m.record) }
for d in changes.deletions { store.delete(d.recordID) }
}
case .sentRecordZoneChanges(let sent):
for failure in sent.failedRecordSaves {
if failure.error.code == .serverRecordChanged,
let server = failure.error.serverRecord {
store.upsert(merge(local: store.note(server.recordID), serverRecord: server))
engine.state.add(pendingRecordZoneChanges: [.saveRecord(server.recordID)])
}
}
case .accountChange:
store.resetForAccountChange()
default:
break
}
}
func nextRecordZoneChangeBatch(_ context: CKSyncEngine.SendChangesContext,
syncEngine: CKSyncEngine) async -> CKSyncEngine.RecordZoneChangeBatch? {
let pending = syncEngine.state.pendingRecordZoneChanges
.filter { context.options.scope.contains($0) }
return await CKSyncEngine.RecordZoneChangeBatch(pendingChanges: pending) { id in
self.store.ckRecord(for: id) // nil drops the change from the batch
}
}
}This is a sketch: LocalStore and NoteRecord are app types, and a real implementation also handles zone creation, deletions and other failure codes.
What sits behind the API
Apple has published two papers that describe CloudKit's server side. The 2018 VLDB paper on CloudKit describes a multi-tenant service in which each user's data is logically separate, with records and zones mapped onto underlying stores and large assets kept in a separate blob store. The 2019 SIGMOD paper on the FoundationDB Record Layer describes CloudKit using that layer on top of FoundationDB, with each user's data held in its own logical record store, so the system manages enormous numbers of small, independent databases rather than one huge one. Exact current internals are not public, and should not be assumed beyond what the papers state.
That shape is consistent with the API: atomic commits are scoped to a zone within one user's data, never across users; per-zone change history is an ordered index that tokens point into; and per-user quotas map onto per-user stores.
Encryption is layered on top. Fields placed in a record's encryptedValues are end-to-end encrypted with keys held on the user's devices, and the Advanced Data Protection account setting extends end-to-end encryption to most iCloud categories. Encrypted fields cannot be indexed or queried on the server, which is a design constraint for anything you want to search server-side.
Worked example: one note, two devices, one flight
A user edits a note's title on an iPhone in flight mode, adding a tag, while the Mac, online, edits the body of the same note and adds a different tag. Both notes started at change tag t7.
- The Mac saves first with tag
t7. The server accepts, assignst8and sends a silent push to the user's other devices. The iPhone is offline and does not receive it. - The iPhone lands. On foregrounding, it sends its pending save carrying
t7. The server rejects it withserverRecordChangedand returns thet8record. - The merge function takes the server record, keeps the Mac's body because the iPhone did not change it, applies the iPhone's newer title, and unions the tags. The iPhone saves the merged record with tag
t8; the server accepts and assignst9. - The Mac receives a push, fetches zone changes since its token, gets the
t9record and applies it. Both devices converge on the same title, body and both tags.
Two details make this work. The iPhone must track which fields it changed, not just that the note is dirty, or it will overwrite the Mac's body. And the merge must be deterministic and idempotent, because a conflict can be retried after a crash.
Failure modes and operational guidance
- Expired change token. The server can reject an old token with
changeTokenExpired. Discard the token, refetch the zone from scratch and reconcile against local data without duplicating records. - User deleted the data. Deleting an app's iCloud data from Settings removes its zones; the next fetch reports the zone deleted, and saves fail with
userDeletedZoneorzoneNotFound. Decide deliberately whether to re-upload or to respect the deletion; usually ask. - Quota exhausted. Private-database writes count against the user's storage.
quotaExceededis not transient; surface it in the UI instead of retrying forever. - Partial failures. A batch can succeed for some records and fail for others; handle errors per record, and never mark a whole batch synced from one result.
- Account switch. When the signed-in account changes, local data belongs to the previous user. Clear it or segregate it, then sync from zero.
- Schema drift. Shipping a build that writes a field missing from the production schema fails at runtime. Deploy schema changes before the app release, and treat production schema as append-only.
Trade-offs against building your own
CloudKit gives you authentication, storage, push and an ordered change log without running servers, and private-database storage is paid for by the user's iCloud plan. In exchange, data is tied to Apple platforms (a web JavaScript API exists, but Android clients are not first-class), queries on the server are limited, cross-user transactions do not exist, and you cannot run your own server-side merge logic. Teams needing collaborative real-time editing usually add an operation-based or CRDT layer, as in Google Drive's real-time collaboration, rather than relying on record-level conflict detection alone.
What to do next
- Model data into custom zones in the private database; decide now which zone will be the unit of sharing.
- If you target iOS 17 or later, start with
CKSyncEngineand persist its state in the same store as your data. - Track dirty fields per record, not dirty records, so merges can be field-level.
- Write the merge function first, with unit tests for concurrent edits, retries and deletions.
- Commit fetched changes and tokens in one local transaction.
- Fetch on launch and foreground; treat silent push as a hint.
- Handle token expiry, deleted zones, quota and account changes explicitly, and test each one.
- Deploy schema to production before every release that adds fields.