A Fly.io Machine is a fast-booting VM whose root filesystem is rebuilt from your image every time it starts. Anything written there is gone after a restart or a deploy. When an app needs state that survives, such as a SQLite database, an upload directory or a search index, Fly gives you a volume: a piece of persistent disk mounted at a path inside the Machine.

The interface looks like any cloud block device, but the model underneath is very different from AWS EBS or Azure managed disks. A Fly volume is a slice of an NVMe drive in one physical server. It is not network storage, it is not replicated, and it cannot follow a Machine to another host. This page explains that model from first principles, walks through deploying a SQLite-backed app on a volume, and covers snapshots, growth, forks, the ways volumes fail and how to design so those failures do not lose data. Every command and configuration key below was checked against Fly's documentation in October 2026.

The model: a slice of local NVMe on one host

Fly's documentation describes a volume as a slice of an NVMe drive on the same physical server as the Machine that uses it. Three consequences follow, and the rest of the article is mostly about living with them.

  1. Locality. Reads and writes go to a local drive rather than across a storage network, so latency is that of local flash. There is no separate storage service between your process and the disk.
  2. Pinning. A volume exists on one server in one region. The Machine that mounts it must run on that server. If the host is down, the Machine cannot be started somewhere else with the same data.
  3. No replication. Fly states plainly that volumes are independent of one another and that it does not replicate data between them. Two volumes with the same name are two separate disks. Keeping them in sync is the application's job.

The mapping is strictly one to one: a volume attaches to one Machine, and a Machine mounts at most one volume. There is no shared filesystem across Machines, and no multi-attach mode. If ten Machines need persistent state, you create ten volumes, and each starts empty unless you seed it from a snapshot or a fork.

Volumes are encrypted at rest by default; --no-encryption turns that off at creation time, and there is rarely a reason to use it. The default size is 1GB and the maximum is 500GB. A volume can grow but never shrink, which matters for planning: over-provisioning is a permanent cost until you migrate the data to a new, smaller volume.

Placement: regions, hardware zones and Machine size

Region: ordHost A (hardware zone 1)Machine app-1mounts /dataVolume dataNVMe slice, vol_aDaily snapshotkept 5 days by default, 1 to 60Host B (hardware zone 2)Machine app-2mounts /dataVolume dataNVMe slice, vol_bDaily snapshotindependent of vol_aapp-level replication (your code)No automatic sync between volumessame name, separate disksHost A failsapp-1 cannot move; vol_b and snapshots survive
Two volumes with the same name in one region. Each lives on its own host and hardware zone; nothing copies data between them unless the application does.

Because a volume is tied to a host, where it is created decides what a single hardware failure can take out. When you create a volume, flyctl's --require-unique-zone behaviour, on by default, places it in a hardware zone that does not already hold a volume of the same name for the app. That spreads a set of volumes across failure domains inside a region. Setting the flag to false lets two volumes share a zone, which is sometimes needed when capacity in a region is tight, but it means one hardware event can affect both.

Fly also accepts a --vm-size hint at creation time. It tells the scheduler which Machine size will mount the volume, so the volume lands on a host that can actually run that Machine. Without it, you can end up with a volume on a host that has no room for the CPU or memory you later ask for, and the Machine fails to be placed.

Choose the region by where the users and other dependencies are. Once data is on a volume in ord, the only way to move it to ams is to copy it: fork the volume into the new region, or restore a snapshot there, and accept that writes made after the copy started must be reconciled.

Worked example: SQLite on a volume

Take a small web service that keeps its data in SQLite. The goal is one primary Machine with a durable database, daily snapshots kept for two weeks, and automatic growth so a full disk does not take the service down at night.

Start with the configuration. The [mounts] section in fly.toml names the volume, the path where it appears in the Machine, and its lifecycle settings:

app = "notes-api"
primary_region = "ord"

[mounts]
  source = "notes_data"            # volume name
  destination = "/data"            # mount point inside the Machine
  initial_size = "10GB"            # used only when the volume is first created
  snapshot_retention = 14          # days, 1 to 60, default 5
  scheduled_snapshots = true       # daily automatic snapshots, the default
  auto_extend_size_threshold = 80  # percent used that triggers growth
  auto_extend_size_increment = "5GB"
  auto_extend_size_limit = "60GB"

[env]
  DATABASE_PATH = "/data/notes.db"

The three auto_extend_size_* keys only work together: the threshold says when to act, the increment how much to add, and the limit where to stop. initial_size is read once, when the first deploy creates the volume, so changing it later does nothing to an existing volume.

You can also create the volume explicitly before deploying, which is clearer when you want to control placement:

fly volumes create notes_data --region ord --size 10
fly volumes list                         # note the volume ID, e.g. vol_xxx
fly deploy

The application only needs to put its database file under the mount point and to open SQLite in write-ahead-log mode, which makes crash recovery after an abrupt stop reliable:

import os, sqlite3

path = os.environ.get("DATABASE_PATH", "/data/notes.db")
conn = sqlite3.connect(path, timeout=5)
conn.execute("PRAGMA journal_mode=WAL")      # crash-safe, readers do not block the writer
conn.execute("PRAGMA synchronous=NORMAL")    # durable at checkpoint, fast commits
conn.execute("CREATE TABLE IF NOT EXISTS notes (id INTEGER PRIMARY KEY, body TEXT NOT NULL)")
conn.commit()

A common mistake is to write the database path into the image, for example /app/notes.db. The app works perfectly until the first deploy, then starts with an empty database. Check where your files land by listing /data from inside the Machine with fly ssh console before relying on it.

The data flow is now: requests reach the Machine, SQLite writes pages to its WAL on the local NVMe slice, the daily snapshot captures the block device, and those snapshots are kept for 14 days. That is a backup with up to a day of exposure. The next sections show how to narrow that gap and what the snapshot actually protects.

Snapshots and restore

Fly takes an automatic snapshot of each volume once a day. Retention defaults to five days and can be set from 1 to 60 with snapshot_retention or at creation. You can also take one on demand, which is the right move immediately before a risky migration:

fly volumes snapshots create vol_xxx      # on-demand snapshot
fly volumes snapshots list vol_xxx        # find the snapshot ID
fly volumes create notes_data --snapshot-id vs_xxx --size 10 --region ord

Restoring always creates a new volume; it never rolls an existing volume back in place. The new volume must be the same size as the original or larger. You then point a Machine at it, which in practice means destroying the old Machine and letting the deploy attach the restored volume, or creating a Machine with that volume explicitly.

Treat a snapshot as an image taken while the Machine is running, not as a coordinated application backup. A database with a write-ahead log, such as SQLite in WAL mode or Postgres, recovers from that kind of image the same way it recovers from a power cut. Files your application writes in several steps without a journal may be captured half-written. If that matters, quiesce the writer or use an application-level backup alongside the snapshot.

Daily snapshots set a recovery point objective of up to 24 hours. For most production data that is too coarse. The usual fix is continuous replication to object storage: Litestream streams SQLite WAL changes to an S3-compatible bucket within seconds, and Postgres can archive WAL the same way. Snapshots then become the second line of defence rather than the only one.

Growing, shrinking and forking

A volume can be extended with fly volumes extend vol_xxx -s 20, or automatically by the auto_extend_size_* settings. It cannot be shrunk. If you grew a volume to 200GB during a one-off import and now use 30GB, the only way down is to create a smaller volume and copy the data across, which is a migration with a write freeze or a replication cut-over.

Auto-extension is a safety net, not capacity planning. Set the threshold low enough that the increment lands before the disk fills during your peak write rate, and set the limit to the most you are willing to pay for. A volume that hits the limit fills up just like one that never had auto-extension, so alert on usage independently.

A fork, fly volumes fork vol_xxx --region ams, makes a new volume with a copy of the data, optionally in another region. The documentation is explicit that the fork and the source are independent afterwards and do not continue to sync. Forks are good for staging copies, for debugging against production-shaped data, and for seeding a replica before replication takes over. They are not replicas.

Failure modes

These are the failures that actually lose data or availability on Fly volumes, roughly in order of how often teams meet them.

FailureWhat happensMitigation
Host failure or maintenanceThe Machine is pinned to the volume's host and cannot start elsewhere; the app is down until the host returnsRun at least two Machines with their own volumes and replicate at the application level
Data written outside the mountLost on the next restart or deploy because the root filesystem is rebuiltWrite only under the destination path; verify with fly ssh console
Volume destroyedPermanent; there is no undo. The volume ID can be looked up for 24 hours after deletion to find its snapshotsSnapshot before destructive work; restrict who can run destroy
Disk fullWrites fail; databases may refuse to startAuto-extend with a sensible limit, plus a usage alert
Scaling outNew Machines get new, empty volumes, not copiesSeed from a fork or snapshot, then let replication catch up
Placement failureNo host in the region has room for the Machine size next to the volumePass --vm-size when creating volumes; keep a spare volume in another zone

Note the order of operations for removal: a volume attached to a Machine cannot be destroyed until that Machine is stopped and destroyed. That guard is useful, but it is not a backup policy.

Replication is your job

Since Fly does not replicate volumes, high availability is something you build. The standard shape is a primary and one or more replicas, each on its own Machine and volume, with the database's own replication protocol moving data between them. The database replication article covers the synchronous and asynchronous trade-offs that decide how much data a failover can lose.

For SQLite, LiteFS replicates a single-writer database to read replicas and handles primary election, and Litestream gives continuous off-host backup without replicas. For Postgres, streaming replication between two volumes in different hardware zones, with automated failover, is the common pattern. In each case the Fly volume provides fast, local, durable-on-this-host storage, and the replication layer provides durability across hosts.

If an app does not need a filesystem at all, the simpler answer is often not to use a volume. Object storage, including Tigris, which is integrated with Fly, keeps files reachable from every Machine in every region with no pinning. A managed database keeps replication someone else's problem. Volumes are the right tool when you want local-disk latency and are willing to own replication.

Trade-offs

OptionStrengthCost
Single Machine plus volumeSimplest; local NVMe latencyDowntime on host loss; up to 24h of data at risk with snapshots alone
Volume plus Litestream to object storageSeconds of data at risk; cheapRestore takes time; still one writer, one host
Primary and replica volumes in different zonesSurvives host loss with failoverYou run replication, failover and monitoring
Object storage or managed databaseNo pinning; durability handled for youNetwork latency; less control

Compared with network block storage elsewhere, such as the options in EBS volume types or OCI block volumes, a Fly volume gives up the ability to detach and reattach to a new host in exchange for local-disk performance and simplicity. That is a good trade for apps that replicate themselves and a poor one for apps that assume the disk is indestructible.

What to do next

Use this checklist before you trust a Fly volume with data you cannot recreate.

  1. Confirm every persistent path is under the destination mount; restart the Machine and check the data survives.
  2. Set snapshot_retention deliberately rather than living with the 5-day default, and take an on-demand snapshot before every schema change.
  3. Practise a restore: fly volumes create --snapshot-id into a scratch app, and time it.
  4. Add continuous backup (Litestream, WAL archiving) so a host loss costs seconds, not a day.
  5. Configure all three auto_extend_size_* keys and a separate disk-usage alert.
  6. For anything user-facing, run a second Machine with its own volume in another hardware zone and replicate at the application level.
  7. Create volumes with --vm-size matching the Machine you will run, and leave --require-unique-zone at its default.
  8. Restrict who can run fly volumes destroy; there is no undo.
Key takeaway: A Fly volume is fast local NVMe tied to one host, attached to exactly one Machine, encrypted, growable but not shrinkable, and never replicated for you. Daily snapshots, kept 5 days unless you change it, are a coarse safety net with up to a day of exposure. Put all state under the mount, add continuous backup, practise restores, and for anything that must stay up run a second Machine and volume in another hardware zone with replication you control.