SQLite is the most widely deployed database in the world, and for a single server it is hard to beat: no network hop, no connection pool, reads measured in microseconds. The problem starts when you want a second server, for availability or for latency in another region. SQLite has no replication protocol. It is a library that writes pages to a file.
LiteFS, from Fly.io, solves that at the file-system layer. It mounts a FUSE file system under your application, watches SQLite commit transactions, copies out the pages each transaction changed, and streams those page sets to replicas, which apply them in order. Your application keeps using the ordinary SQLite library. This article explains how that works, traces one write through the system, shows the configuration, and then spends most of its time on what you must design around: a single writer, asynchronous replication, failover that can drop acknowledged writes, and backups now that the managed LiteFS Cloud service is gone. LiteFS is still labelled beta; the latest release at the time of writing is v0.5.14, from April 2025. Read this as a guide to a useful but young tool, and pin your version.
The architecture: FUSE, LTX and a lease
There are four moving parts. The FUSE mount (conventionally /litefs) is where your app opens its database. LiteFS is a passthrough: reads and writes go to the real files in the data directory (conventionally /var/lib/litefs), but LiteFS sees every call on the way. The lease decides which node is primary; only the primary accepts writes. The replication stream carries changes from primary to replicas over HTTP, on port 20202 by default.
What gets replicated is not SQL and not a byte-level file diff. It is an LTX file: LiteFS's own format containing the set of database pages one transaction changed, with checksums that let a receiver verify it is applying the right pages to the right base state. How LiteFS finds those pages depends on SQLite's journal mode. In rollback-journal modes (DELETE, PERSIST, TRUNCATE) it tracks which pages were dirtied by write(2) calls and exports them when the journal is invalidated at commit. In WAL mode every change is already sitting in the write-ahead log, so it reads the WAL from the previous offset to the new end.
Every node tracks its replication position as a pair: a transaction ID that increments on each write, and a rolling checksum of the whole database contents. The docs show positions in the form 00000000000003e8/a3b2b72f1147c9bc. The checksum is what makes the system safe to reason about. Two nodes with the same TXID but different checksums do not hold the same database, and LiteFS treats that as divergence rather than guessing.
One write, traced end to end
Take a small app with two nodes: lhr-1 is primary and syd-1 is a replica. Both are at position TXID 1000. A user in London posts a comment.
- The request reaches
lhr-1. The app runsINSERT INTO comments ...inside a transaction against/litefs/app.db. - SQLite writes pages through the FUSE mount. In WAL mode those writes land in
app.db-wal; LiteFS sees them pass through. - SQLite commits. LiteFS reads the newly appended WAL frames, builds an LTX file for TXID 1001 with the changed pages and the new rolling checksum, and stores it in the data directory.
- The commit returns to the app, which responds 201 to the user. No replica has the change yet.
- LiteFS pushes the LTX file to
syd-1over its HTTP stream. The replica checks the LTX against its position (TXID 1000 and matching checksum), applies the pages to its copy and advances to 1001. - Within network latency plus apply time, a read on
syd-1sees the comment.
Step 4 is the whole story of LiteFS's consistency model. Replication is asynchronous: the docs say so directly and list synchronous and time-bound asynchronous replication as future work. Between steps 4 and 5 the write exists on exactly one machine. A replica that applied 1001 has the same bytes as the primary had after 1001, which is a stronger guarantee than many logical replication systems give, but nothing promises it applied it in time.
LTX files are kept for a retention window (data.retention, default 10 minutes) so a replica that briefly disconnects can catch up by fetching the missing TXIDs. A replica that falls further behind, or that diverges, is brought back with a full snapshot instead.
Configuring a cluster
LiteFS runs as the entrypoint of your container: it mounts the file system, joins the cluster, and then starts your app as a subprocess via exec, so the app never touches the database before the mount exists. A Consul-leased configuration on Fly.io looks like this:
# litefs.yml
fuse:
dir: "/litefs" # app opens /litefs/app.db
data:
dir: "/var/lib/litefs" # real files; put this on a persistent volume
compress: true # default
retention: "10m" # how long LTX files are kept for catch-up
exec:
- cmd: "myapp -addr :8081"
lease:
type: "consul"
advertise-url: "http://${HOSTNAME}.vm.${FLY_APP_NAME}.internal:20202"
candidate: ${FLY_REGION == PRIMARY_REGION} # only nodes in one region may lead
consul:
url: "${FLY_CONSUL_URL}"
key: "litefs/${FLY_APP_NAME}"
ttl: "10s" # default
lock-delay: "1s" # default
proxy:
addr: ":8080" # public port
target: "localhost:8081" # your app
db: "app.db"
passthrough: ["*.css", "*.js", "*.png"]
http:
addr: ":20202"Three choices in that file matter more than the rest. lease.candidate restricts which nodes may become primary. Keep candidates in one region, ideally the one closest to most writers, because every write from elsewhere pays a cross-region round trip. data.dir must be on a persistent volume, or a restarted node rejoins empty and needs a full snapshot. And the lease type can be static instead of consul: one fixed node is always primary, which removes Consul from the picture and removes automatic failover with it. Static is honest for small deployments where you would rather be down than lose writes.
Routing writes to the primary
A replica's copy of the database is read-only to SQLite. A write attempted there fails: the docs say you get SQLITE_READONLY in rollback-journal mode and a generic disk I/O error in WAL mode. The second message is a trap, because it looks like a hardware fault in your logs. Your application needs a routing rule, and LiteFS gives you two.
The .primary file. On a replica, /litefs/.primary exists and contains the primary's hostname. On the primary it does not exist. An app that knows which requests write can check it and forward:
from pathlib import Path
PRIMARY_FILE = Path("/litefs/.primary")
def primary_host():
"""None when this node is primary; otherwise the primary's hostname."""
try:
return PRIMARY_FILE.read_text().strip()
except FileNotFoundError:
return None
def handle_write(request):
host = primary_host()
if host is not None:
# forward instead of writing locally; on Fly.io, a fly-replay
# response header asks the platform proxy to replay it there
return Response(status=409, headers={"fly-replay": f"instance={host}"})
with db: # sqlite3 connection on /litefs/app.db
db.execute("INSERT INTO comments(body) VALUES (?)", (request.body,))
return Response(status=201)Check the exact fly-replay field syntax against the Fly.io dynamic request routing docs for your setup; the important point is that the decision is made per request from a file LiteFS keeps current.
The built-in proxy. For ordinary web apps, the proxy in the config above does this for you. It treats GET requests as reads and serves them locally; methods such as POST and PUT are forwarded to the primary with the fly-replay header, which makes the proxy specific to Fly.io's platform. After a write, the primary sets a cookie carrying the TXID of that commit. When the same client's next read lands on a replica, the proxy waits until the replica's position reaches that TXID before passing the request through. That gives each user read-your-writes consistency without making every read go to the primary. passthrough lists paths the proxy should not apply this to, such as static assets, and primary-redirect-timeout (default 5s) holds writes while a new primary is being elected.
The proxy's limits come straight from that design. It does not work with WebSockets. Clients must accept cookies. And any GET handler that writes, such as a view counter or session refresh, breaks on replicas. Audit for those before you turn the proxy on.
Failover, and what asynchronous replication can lose
Primary election uses a Consul lease with a TTL (10 seconds by default). The primary renews it; if the primary dies or loses Consul, the lease expires and a candidate acquires it. Now apply the asynchronous replication from the trace.
Worked failure. lhr-1 commits TXID 1001 and 1002 and acknowledges both. lhr-2, the other candidate, has applied only 1001 when lhr-1's host dies. The lease expires, lhr-2 becomes primary at 1001 and accepts a new write as its own 1002. When lhr-1 comes back, its 1002 has a different checksum from the cluster's 1002. LiteFS detects the split and snapshots the new primary's database onto lhr-1. The original 1002, which a user was told had succeeded, is gone.
This is not a bug; it is what asynchronous replication means. Your options:
- Accept a small loss window and say so in your design doc. For a content site or a cache-like store this is often fine.
- Use a static lease: no automatic failover, so no divergence; recovery is a manual decision.
- Make important writes idempotent and replayable from an upstream source (a queue, a webhook provider that retries), so a lost tail can be re-applied.
- Keep candidates close together so the replication lag between them, and therefore the loss window, stays small.
Other failure modes are more mundane. A full volume stops commits on the primary. A long-running write transaction blocks every other writer, as in plain SQLite, but now the whole cluster has one writer. A replica far behind the retention window needs a snapshot, which for a multi-gigabyte database is real network and disk time. And a Consul outage does not take down a healthy primary immediately, but it does prevent any failover until Consul returns.
Backups, monitoring and trade-offs
Backups. Replicas are not backups: a bad DELETE replicates in milliseconds. Fly.io ran a managed backup service, LiteFS Cloud, and shut it down on 15 October 2024. The project's own recommendation since then is Litestream, which continuously copies SQLite changes to S3-compatible object storage and can restore to a point in time. Run it against the primary's database and test a restore on a schedule, not once.
Monitoring. Watch, per node: replication position (TXID) versus the primary's, which is your lag; whether the node is primary, and how often that changes; LTX retention disk usage; and write latency on the primary. Alert on lag measured in TXIDs and in seconds, and on any primary change you did not plan.
When LiteFS fits, and when it does not.
| Option | Writers | Replication | Good fit |
|---|---|---|---|
| Single SQLite + Litestream | 1 | none; async backup | one server, fast restore acceptable |
| LiteFS | 1 (lease) | async page sets | read-heavy apps, global read replicas |
| Cloudflare D1 | managed | managed by the platform | Workers apps |
| Postgres with streaming replicas | 1 | async or sync | write-heavy, needs sync commit |
| Distributed SQL (CockroachDB etc.) | many | consensus per range | multi-region writes |
LiteFS shines when reads dominate, the dataset fits on one disk, and you want reads served from the node the user hits. It is the wrong tool when you need synchronous durability across machines, many concurrent writers, or a guarantee that an acknowledged write survives any single failure. See Fly volumes for the storage underneath, Fly regions for placing candidates, Cloudflare D1 for the managed alternative, write-ahead logging for the WAL mechanics LiteFS reads, and leases for why a TTL lease cannot rule out a brief overlap of two primaries.
What to do next
- Pin a LiteFS version in your Dockerfile and read its release notes; the project is beta.
- Put
data.diron a persistent volume and setlease.candidateso only nodes in your write region can lead. - Decide, in writing, whether you accept the async loss window; if not, use a static lease.
- Audit every
GEThandler for writes before enabling the proxy, or forward writes yourself using/litefs/.primary. - Run Litestream against the primary and schedule a restore test.
- Dashboard TXID lag per replica and alert on unplanned primary changes.
- Rehearse a failover: kill the primary mid-write and confirm what your app and users see.