Fly.io runs applications as Machines, fast-booting VMs placed in regions close to users. For years, Postgres on Fly meant running a Postgres app on your own Machines with some flyctl conveniences, and the operations were yours. That offering is now legacy: Fly's documentation states that unmanaged Fly Postgres is no longer maintained and that Fly support cannot help with it. The supported path is Fly Managed Postgres, usually shortened to MPG, a service where Fly runs the cluster, handles failover and backups, and gives your app connection strings.
This article explains how MPG is put together, how to create a cluster and attach an app, the difference between pooled and direct connections and why it matters for your driver, sizing by plan and connection count, backups and point-in-time restore, migrating off legacy Fly Postgres, and the failure modes that show up in practice. Everything specific here was checked against Fly's documentation on 2026-10-04; where the docs are silent, such as backup frequency or failover timing, this article says so rather than guessing.
What Managed Postgres is
An MPG cluster is a managed Postgres deployment in one Fly region. According to Fly's docs, every plan includes a primary and a replica, PgBouncer connection poolers, and backups, plus automatic failover, encryption at rest and in transit, metrics and support. You do not see or manage the Machines underneath. Storage can reach 1 TB, and a cluster can be created with up to 500 GB initially. Extensions such as pgvector and PostGIS are available if enabled when the cluster is provisioned, so decide on them at creation time; other third-party extensions are listed as still under development, so check the extensions page before you depend on one.
Compare that with the legacy model, which was a Fly app like any other: you chose Machine sizes, attached Fly Volumes, handled replication and upgrades, and fixed it when it broke. If you run one of those today, plan a migration; there is a section on it below. For a different design that keeps SQLite on each Machine and replicates it, see LiteFS.
Architecture: pooled and direct endpoints
Each cluster exposes two hostnames on Fly's private network. The pooled host, pgbouncer.<cluster>.flympg.net, routes through PgBouncer and is what applications should use. The direct host, direct.<cluster>.flympg.net, bypasses the pooler. Fly recommends the direct URL for migrations, and it is required for advisory locks and LISTEN/NOTIFY when the pooler is in transaction mode. TLS is enabled by default. Because the hosts live on the private network, your app should run in the same organisation and ideally the same region; Fly regions explains placement. Your app runs on Fly Machines, and the database is just another private-network service from their point of view.
Creating a cluster and attaching an app
Creation and attachment are two flyctl commands. fly mpg create takes a name, organisation, plan, region and initial volume size in GB, which defaults to 10. The flag help lists basic, launch and scale as plan values.
# create a cluster next to the app
fly mpg create --name shop-db --org acme --plan launch --region ams --volume-size 40
# list clusters and note the cluster ID
fly mpg list
# attach it to an app: creates the DATABASE_URL secret and triggers a deploy
fly mpg attach <clusterID> -a shop-web
# a psql session, or a local port for GUI tools and scripts
fly mpg connect
fly mpg proxy # e.g. localhost:16380 -> the cluster's private address on 5432Attach writes the connection string into a secret named DATABASE_URL; if your app already has a secret by that name, choose another variable name. You can also set the secret yourself with fly secrets set DATABASE_URL=..., which is how you would add a second URL for the direct host. The fly mpg databases and fly mpg users command groups manage logical databases, roles and extensions, and the same operations exist in the Managed Postgres section of the Machines API for automation.
That API is also your day-2 toolkit. Alongside create, attach, backup and restore, it lists endpoints to show active queries and slow queries, to fork a cluster into a new cluster, and to create users, change their roles and rotate a user's password. Wire the slow-query listing into a weekly review, and script password rotation so that a leaked credential becomes a routine fix rather than an incident. A fork is the cheapest way to give a migration rehearsal or a load test realistic data without touching production.
PgBouncer modes and your driver
PgBouncer multiplexes many client connections onto fewer server connections. It has two modes in MPG. Session mode, the default, assigns a server connection to a client for the whole session, so everything Postgres supports works: named prepared statements, advisory locks, session settings. Transaction mode returns the server connection to the pool after each transaction, which supports far more concurrent clients but breaks anything that assumes the next statement runs on the same backend. Named prepared statements, advisory locks and LISTEN/NOTIFY do not work through it.
If you switch to transaction mode, configure your driver to stop using named prepared statements. Fly's client configuration guide lists the settings:
| Client | Setting for transaction mode |
|---|---|
| Prisma | ?pgbouncer=true on the connection string |
| Ecto | prepare: :unnamed |
| ActiveRecord | prepared_statements: false |
| psycopg 3 | prepare_threshold=None |
| pgx | default_query_exec_mode=exec |
Two client-side pool settings apply in either mode: a maximum connection lifetime of 600 seconds and an idle timeout of 300 seconds. The point is to recycle connections before the proxy closes them, so your app never hands a dead connection to a request. Changing the pool mode restarts the pooler nodes, which causes a brief interruption; the database itself keeps running.
Sizing: plans, memory and connections
Plans set CPU and memory per node. As of the date checked, Fly lists:
| Plan | vCPUs | RAM |
|---|---|---|
| Basic | 2 shared | 1 GB |
| Starter | 2 shared | 2 GB |
| Launch | 2 performance | 8 GB |
| Scale | 4 performance | 32 GB |
| Performance | 8 performance | 64 GB |
Client connection limits run from 200 on Basic and Starter to 1,000 on Scale and Performance; check the docs for the plans in between. Size against your working set first: Postgres is fast when hot tables and indexes fit in memory, so compare the size of your most-read tables and indexes with the plan's RAM. Then do the connection arithmetic. A web app with 6 Machines, each running 2 processes with a pool of 10, opens 120 connections, plus the release command, background workers and any analytics tool. On a 200-connection plan that leaves little room for a deploy, when old and new Machines overlap and briefly double the count. Lower per-process pools, or switch to transaction mode, before you hit the limit rather than after.
You can change plan at any time. Fly notes that nodes restart during the change and there may be a brief period of downtime during the switchover, so schedule it.
Worked example: a Rails app growing up
A Rails app, shop-web, runs in Amsterdam on 4 Machines and uses Sidekiq for background jobs. The team creates a Launch cluster in ams with 40 GB, attaches it, and keeps the default session mode, because Sidekiq and the app together open about 90 connections, well within limits, and session mode needs no driver changes.
They add a second secret, DIRECT_DATABASE_URL, pointing at the direct host, and the release command runs migrations against it, because Rails migrations take an advisory lock and must not go through a transaction-mode pooler if the team ever switches modes. The app's database.yml sets pool size per process and a reaping frequency so idle connections close within the 300-second guidance. Six months later traffic quadruples. Rather than jump straight to Scale, they switch the pooler to transaction mode, set prepared_statements: false, and keep migrations on the direct URL. Connection count stops being the constraint, and the next upgrade is driven by memory: when the hot order tables outgrow 8 GB and cache hit rate drops, they move to Scale in a quiet window.
Backups and point-in-time restore
Backups are included. You can trigger one with fly mpg backup create and list them with fly mpg backup list. Restores never touch the source: fly mpg restore creates a new cluster from a backup or from a point in time.
# restore a specific backup into a new cluster
fly mpg restore <clusterID> --backup-id <backupID> --name shop-db-restored
# or restore to a moment just before a bad deploy (RFC 3339, UTC)
fly mpg restore <clusterID> --pitr-time 2026-10-03T14:05:00Z --name shop-db-pitrThe two flags are mutually exclusive, and a point-in-time restore only works inside the cluster's recovery window. The docs consulted do not state backup frequency or retention, so look them up for your plan and write them in your runbook. Rehearse the whole path: restore into a new cluster, point a staging copy of the app at it, run checks, and time it. Recovery from a bad migration usually means restoring to a new cluster and either switching the app's secret or copying the damaged rows back.
Migrating from legacy Fly Postgres
Fly does not offer an automated migration from legacy Fly Postgres. The documented path for simple databases is a dump piped into the new cluster through the proxy:
# 1. create an MPG cluster at least as large as the source database
fly mpg create --name app-db --plan launch --region ams --volume-size 60
# 2. in one terminal, open a local proxy to the new cluster
fly mpg proxy
# 3. in another, copy schema and data (source must be reachable from here)
pg_dump "$LEGACY_PG_CONNECTION_STRING" | psql "postgres://fly-user:<password>@localhost:16380/fly-db"Put the app into maintenance or read-only mode before the dump so no writes are lost, verify row counts on the main tables, then fly mpg attach the app and deploy. Unsupported extensions block the import, so list the extensions in the source first. For larger databases, Fly points to tools such as pgloader or pgcopydb, which copy in parallel and reduce downtime.
Failure modes
- Prepared statement errors after switching to transaction mode. Messages about prepared statements that do not exist or already exist mean the driver still uses named statements. Apply the setting from the table above.
- Migrations hang or fail on locks. Advisory locks through a transaction-mode pooler do not hold. Run migrations on the direct URL.
- LISTEN/NOTIFY silently stops. Same cause. Listeners need the direct host.
- Dead connections after idle periods. The proxy closed them first. Set the 600-second lifetime and 300-second idle timeout in the client pool.
- Too many connections during deploys. Old and new Machines overlap. Budget for twice the steady-state count, or use transaction mode.
- Cross-region latency. An app in one region and a cluster in another pays a round trip on every query. Keep them together.
- Brief outage on plan change. Nodes restart. Schedule plan changes and warn users.
What to do next
- If you still run legacy Fly Postgres, inventory those apps and schedule migrations now; it is unsupported.
- Create an MPG cluster in your app's region with a plan sized to your hot working set.
- Attach it, then add a second secret for the direct host and run migrations against it.
- Set client pool lifetime to 600 seconds and idle timeout to 300 seconds in every service.
- Add up worst-case connections across Machines, workers and deploy overlap; compare with your plan's limit.
- Decide on session or transaction mode, and if transaction, change the driver setting and test LISTEN/NOTIFY and locks.
- Restore a backup and a point-in-time copy into new clusters, time it, and record backup retention in your runbook.