SQLite is the most widely deployed database in the world, and for a single server it is hard to beat: no network hop, no connection pool, reads measured in microseconds. The problem starts when you want a second server, for availability or for latency in another region. SQLite has no replication protocol. It is a library that writes pages to a file.

LiteFS, from Fly.io, solves that at the file-system layer. It mounts a FUSE file system under your application, watches SQLite commit transactions, copies out the pages each transaction changed, and streams those page sets to replicas, which apply them in order. Your application keeps using the ordinary SQLite library. This article explains how that works, traces one write through the system, shows the configuration, and then spends most of its time on what you must design around: a single writer, asynchronous replication, failover that can drop acknowledged writes, and backups now that the managed LiteFS Cloud service is gone. LiteFS is still labelled beta; the latest release at the time of writing is v0.5.14, from April 2025. Read this as a guide to a useful but young tool, and pin your version.

The architecture: FUSE, LTX and a lease

There are four moving parts. The FUSE mount (conventionally /litefs) is where your app opens its database. LiteFS is a passthrough: reads and writes go to the real files in the data directory (conventionally /var/lib/litefs), but LiteFS sees every call on the way. The lease decides which node is primary; only the primary accepts writes. The replication stream carries changes from primary to replicas over HTTP, on port 20202 by default.

What gets replicated is not SQL and not a byte-level file diff. It is an LTX file: LiteFS's own format containing the set of database pages one transaction changed, with checksums that let a receiver verify it is applying the right pages to the right base state. How LiteFS finds those pages depends on SQLite's journal mode. In rollback-journal modes (DELETE, PERSIST, TRUNCATE) it tracks which pages were dirtied by write(2) calls and exports them when the journal is invalidated at commit. In WAL mode every change is already sitting in the write-ahead log, so it reads the WAL from the previous offset to the new end.

Every node tracks its replication position as a pair: a transaction ID that increments on each write, and a rolling checksum of the whole database contents. The docs show positions in the form 00000000000003e8/a3b2b72f1147c9bc. The checksum is what makes the system safe to reason about. Two nodes with the same TXID but different checksums do not hold the same database, and LiteFS treats that as divergence rather than guessing.

LiteFS: one primary captures page sets, replicas apply them in TXID orderApp (primary node)SQLite library, normal file I/OLiteFS FUSE mount/litefs: intercepts writesdata dir/var/lib/litefs: db + LTX filesLease (Consul)key held by the primary, TTLApp (replica node)reads locally, writes forwardedLiteFS FUSE mountread-only for SQLitedata dirapplies LTX at next TXIDLitestream / backupoff-host copies, point in timeCOMMITLTX page setacquirestream LTX over HTTP :20202write requests (proxy / .primary)backupPosition = TXID + rolling checksum. Replication is asynchronous: COMMIT returns before any replica has the page set.A replica whose checksum diverges from the new primary's history is re-snapshotted, not merged.
LiteFS architecture: the primary's FUSE mount turns each commit into an LTX page set, which replicas apply in TXID order.

One write, traced end to end

Take a small app with two nodes: lhr-1 is primary and syd-1 is a replica. Both are at position TXID 1000. A user in London posts a comment.

  1. The request reaches lhr-1. The app runs INSERT INTO comments ... inside a transaction against /litefs/app.db.
  2. SQLite writes pages through the FUSE mount. In WAL mode those writes land in app.db-wal; LiteFS sees them pass through.
  3. SQLite commits. LiteFS reads the newly appended WAL frames, builds an LTX file for TXID 1001 with the changed pages and the new rolling checksum, and stores it in the data directory.
  4. The commit returns to the app, which responds 201 to the user. No replica has the change yet.
  5. LiteFS pushes the LTX file to syd-1 over its HTTP stream. The replica checks the LTX against its position (TXID 1000 and matching checksum), applies the pages to its copy and advances to 1001.
  6. Within network latency plus apply time, a read on syd-1 sees the comment.

Step 4 is the whole story of LiteFS's consistency model. Replication is asynchronous: the docs say so directly and list synchronous and time-bound asynchronous replication as future work. Between steps 4 and 5 the write exists on exactly one machine. A replica that applied 1001 has the same bytes as the primary had after 1001, which is a stronger guarantee than many logical replication systems give, but nothing promises it applied it in time.

LTX files are kept for a retention window (data.retention, default 10 minutes) so a replica that briefly disconnects can catch up by fetching the missing TXIDs. A replica that falls further behind, or that diverges, is brought back with a full snapshot instead.

Configuring a cluster

LiteFS runs as the entrypoint of your container: it mounts the file system, joins the cluster, and then starts your app as a subprocess via exec, so the app never touches the database before the mount exists. A Consul-leased configuration on Fly.io looks like this:

# litefs.yml
fuse:
  dir: "/litefs"                 # app opens /litefs/app.db

data:
  dir: "/var/lib/litefs"         # real files; put this on a persistent volume
  compress: true                 # default
  retention: "10m"               # how long LTX files are kept for catch-up

exec:
  - cmd: "myapp -addr :8081"

lease:
  type: "consul"
  advertise-url: "http://${HOSTNAME}.vm.${FLY_APP_NAME}.internal:20202"
  candidate: ${FLY_REGION == PRIMARY_REGION}   # only nodes in one region may lead
  consul:
    url: "${FLY_CONSUL_URL}"
    key: "litefs/${FLY_APP_NAME}"
    ttl: "10s"                   # default
    lock-delay: "1s"             # default

proxy:
  addr: ":8080"                  # public port
  target: "localhost:8081"       # your app
  db: "app.db"
  passthrough: ["*.css", "*.js", "*.png"]

http:
  addr: ":20202"

Three choices in that file matter more than the rest. lease.candidate restricts which nodes may become primary. Keep candidates in one region, ideally the one closest to most writers, because every write from elsewhere pays a cross-region round trip. data.dir must be on a persistent volume, or a restarted node rejoins empty and needs a full snapshot. And the lease type can be static instead of consul: one fixed node is always primary, which removes Consul from the picture and removes automatic failover with it. Static is honest for small deployments where you would rather be down than lose writes.

Routing writes to the primary

A replica's copy of the database is read-only to SQLite. A write attempted there fails: the docs say you get SQLITE_READONLY in rollback-journal mode and a generic disk I/O error in WAL mode. The second message is a trap, because it looks like a hardware fault in your logs. Your application needs a routing rule, and LiteFS gives you two.

The .primary file. On a replica, /litefs/.primary exists and contains the primary's hostname. On the primary it does not exist. An app that knows which requests write can check it and forward:

from pathlib import Path

PRIMARY_FILE = Path("/litefs/.primary")

def primary_host():
    """None when this node is primary; otherwise the primary's hostname."""
    try:
        return PRIMARY_FILE.read_text().strip()
    except FileNotFoundError:
        return None

def handle_write(request):
    host = primary_host()
    if host is not None:
        # forward instead of writing locally; on Fly.io, a fly-replay
        # response header asks the platform proxy to replay it there
        return Response(status=409, headers={"fly-replay": f"instance={host}"})
    with db:                      # sqlite3 connection on /litefs/app.db
        db.execute("INSERT INTO comments(body) VALUES (?)", (request.body,))
    return Response(status=201)

Check the exact fly-replay field syntax against the Fly.io dynamic request routing docs for your setup; the important point is that the decision is made per request from a file LiteFS keeps current.

The built-in proxy. For ordinary web apps, the proxy in the config above does this for you. It treats GET requests as reads and serves them locally; methods such as POST and PUT are forwarded to the primary with the fly-replay header, which makes the proxy specific to Fly.io's platform. After a write, the primary sets a cookie carrying the TXID of that commit. When the same client's next read lands on a replica, the proxy waits until the replica's position reaches that TXID before passing the request through. That gives each user read-your-writes consistency without making every read go to the primary. passthrough lists paths the proxy should not apply this to, such as static assets, and primary-redirect-timeout (default 5s) holds writes while a new primary is being elected.

The proxy's limits come straight from that design. It does not work with WebSockets. Clients must accept cookies. And any GET handler that writes, such as a view counter or session refresh, breaks on replicas. Audit for those before you turn the proxy on.

Failover, and what asynchronous replication can lose

Primary election uses a Consul lease with a TTL (10 seconds by default). The primary renews it; if the primary dies or loses Consul, the lease expires and a candidate acquires it. Now apply the asynchronous replication from the trace.

Worked failure. lhr-1 commits TXID 1001 and 1002 and acknowledges both. lhr-2, the other candidate, has applied only 1001 when lhr-1's host dies. The lease expires, lhr-2 becomes primary at 1001 and accepts a new write as its own 1002. When lhr-1 comes back, its 1002 has a different checksum from the cluster's 1002. LiteFS detects the split and snapshots the new primary's database onto lhr-1. The original 1002, which a user was told had succeeded, is gone.

This is not a bug; it is what asynchronous replication means. Your options:

  • Accept a small loss window and say so in your design doc. For a content site or a cache-like store this is often fine.
  • Use a static lease: no automatic failover, so no divergence; recovery is a manual decision.
  • Make important writes idempotent and replayable from an upstream source (a queue, a webhook provider that retries), so a lost tail can be re-applied.
  • Keep candidates close together so the replication lag between them, and therefore the loss window, stays small.

Other failure modes are more mundane. A full volume stops commits on the primary. A long-running write transaction blocks every other writer, as in plain SQLite, but now the whole cluster has one writer. A replica far behind the retention window needs a snapshot, which for a multi-gigabyte database is real network and disk time. And a Consul outage does not take down a healthy primary immediately, but it does prevent any failover until Consul returns.

Backups, monitoring and trade-offs

Backups. Replicas are not backups: a bad DELETE replicates in milliseconds. Fly.io ran a managed backup service, LiteFS Cloud, and shut it down on 15 October 2024. The project's own recommendation since then is Litestream, which continuously copies SQLite changes to S3-compatible object storage and can restore to a point in time. Run it against the primary's database and test a restore on a schedule, not once.

Monitoring. Watch, per node: replication position (TXID) versus the primary's, which is your lag; whether the node is primary, and how often that changes; LTX retention disk usage; and write latency on the primary. Alert on lag measured in TXIDs and in seconds, and on any primary change you did not plan.

When LiteFS fits, and when it does not.

OptionWritersReplicationGood fit
Single SQLite + Litestream1none; async backupone server, fast restore acceptable
LiteFS1 (lease)async page setsread-heavy apps, global read replicas
Cloudflare D1managedmanaged by the platformWorkers apps
Postgres with streaming replicas1async or syncwrite-heavy, needs sync commit
Distributed SQL (CockroachDB etc.)manyconsensus per rangemulti-region writes

LiteFS shines when reads dominate, the dataset fits on one disk, and you want reads served from the node the user hits. It is the wrong tool when you need synchronous durability across machines, many concurrent writers, or a guarantee that an acknowledged write survives any single failure. See Fly volumes for the storage underneath, Fly regions for placing candidates, Cloudflare D1 for the managed alternative, write-ahead logging for the WAL mechanics LiteFS reads, and leases for why a TTL lease cannot rule out a brief overlap of two primaries.

What to do next

  1. Pin a LiteFS version in your Dockerfile and read its release notes; the project is beta.
  2. Put data.dir on a persistent volume and set lease.candidate so only nodes in your write region can lead.
  3. Decide, in writing, whether you accept the async loss window; if not, use a static lease.
  4. Audit every GET handler for writes before enabling the proxy, or forward writes yourself using /litefs/.primary.
  5. Run Litestream against the primary and schedule a restore test.
  6. Dashboard TXID lag per replica and alert on unplanned primary changes.
  7. Rehearse a failover: kill the primary mid-write and confirm what your app and users see.
Key takeaway: LiteFS replicates SQLite by intercepting commits at the file-system layer and shipping each transaction's changed pages to replicas, which apply them in TXID order with checksums. It gives you local reads everywhere and a single writer chosen by a lease. Replication is asynchronous, so a failover can discard acknowledged writes; decide whether that is acceptable, route writes deliberately, back up with Litestream, and pin the version of a project that is still in beta.