A Fly Machine is a fast-starting virtual machine that runs one container image on one host in one region. Everything else on Fly.io is built from Machines: fly deploy creates and updates them, the proxy starts and stops them, and Fly's Postgres runs on them. You can also drive them directly through a REST API, which turns Fly into a place where you can start a VM per job, per user or per tenant in a fraction of a second and pay for it only while it runs.
This article covers the Machine itself: what it is underneath, the state machine it moves through, the API you use to control it, how automatic stopping, starting and suspending work, and what fails. Routing between regions and replaying requests is covered in Fly.io regions and persistent disks in Fly.io volumes. The worked example is a job runner that starts one Machine per task, which exercises nearly every part of the API.
What a Machine is
Underneath, a Machine is a Firecracker microVM. Fly takes your OCI image, unpacks it into a root filesystem, and boots a small VM with its own kernel, so tenants are isolated by hardware virtualization rather than only by Linux namespaces. Three properties follow, and they shape how you should build on it.
A Machine is pinned. It lives on a specific host in a specific region. It does not float across a pool like a Kubernetes pod. If that host has a problem, the Machine is unavailable until the host recovers or Fly migrates it, which is why you run at least two Machines for anything that must stay up.
The root filesystem is disposable. When a stopped Machine starts again, its rootfs is reset to the image. Anything written outside a volume is gone. Persistent data goes on a volume, which is tied to the same host, or into an external database or object store.
Configuration is versioned. Updating a Machine (a new image, more memory, a changed environment variable) produces a new version; the old one moves to the terminal replaced state. The API requires the whole configuration on every update, not a patch, so a client that sends only the field it changed wipes the rest.
The state machine
Fly documents five persistent states: created (exists, never started), started (running and reachable), stopped (exited on its own or was stopped), suspended (memory saved to disk, will resume on next start) and failed (could not start). The transient states are creating, starting, stopping, restarting, suspending, destroying, launch_failed, updating and replacing. The terminal states are destroyed, replaced and migrated.
Two operational facts come from the same page. A Machine that sits in starting, stopping, restarting or destroying for more than five minutes, or in updating for more than ten, is probably wedged, so build that timeout into your orchestration. And a stopped Machine costs much less than a running one but not nothing: Fly bills root filesystem storage for stopped Machines under its current pricing, so thousands of idle Machines left behind by a job runner do add up. Check the pricing page rather than assuming zero.
The Machines API
The Machines API is a REST API. From the internet use https://api.machines.dev; from inside your Fly private network use http://_api.internal:4280. Authenticate with Authorization: Bearer <token>, ideally a deploy token scoped to one app (fly tokens deploy). The core endpoints all sit under /v1/apps/{app}/machines:
| Call | Purpose |
|---|---|
POST /machines | Create (and by default start) a Machine from a config |
GET /machines/{id} | Read state, config, region, instance id |
POST /machines/{id} | Update; send the full config |
POST /machines/{id}/start | Start a stopped Machine, or resume a suspended one |
POST /machines/{id}/stop | Stop: process gets a signal, then the VM shuts down |
POST /machines/{id}/suspend | Snapshot memory to disk and pause |
GET /machines/{id}/wait | Block until the Machine reaches a requested state, or time out |
POST /machines/{id}/lease | Take an exclusive lease so two controllers do not fight |
POST /machines/{id}/cordon | Stop the proxy routing new requests to it |
DELETE /machines/{id} | Destroy |
Rate limits are per action and per Machine: about 1 request per second per action, with bursts to 3, and 5 per second (burst 10) for reading a Machine. That makes a polling loop on GET the wrong design; use /wait, which holds the request open until the state changes. And creating a Machine can fail because the region has no capacity for that size at that moment. The docs are explicit that handling this is your job: retry, or fall back to another region.
Worked example: one Machine per job
Here is a job runner that starts one Machine per task, waits for it to finish, and cleans up. Each job runs a container that does its work and exits. With auto_destroy set and the restart policy set to no, the Machine destroys itself when the process exits, so a crashed controller does not leave it behind.
import os, time, httpx
API = "https://api.machines.dev/v1/apps/thumbs-jobs/machines"
H = {"Authorization": f"Bearer {os.environ['FLY_API_TOKEN']}"}
REGIONS = ["ams", "fra", "cdg"] # preferred first; fall back on capacity errors
def run_job(job_id: str, src_url: str) -> str:
cfg = {
"image": "registry.fly.io/thumbs-worker:2026-10-01",
"env": {"JOB_ID": job_id, "SRC_URL": src_url},
"guest": {"cpu_kind": "shared", "cpus": 2, "memory_mb": 1024},
"auto_destroy": True, # destroy when the process exits
"restart": {"policy": "no"}, # a failed job is retried by us, not by Fly
"metadata": {"job_id": job_id}, # lets a sweeper find orphans
}
last_err = None
for region in REGIONS:
r = httpx.post(API, headers=H, json={"region": region, "config": cfg}, timeout=30)
if r.status_code in (200, 201):
m = r.json()
break
last_err = r.text # capacity or validation error; try next region
if r.status_code == 422:
raise ValueError(last_err) # bad config: no point retrying elsewhere
else:
raise RuntimeError(f"no capacity in {REGIONS}: {last_err}")
# Long-poll until the job exits. Each wait call is bounded (default 60 s), so loop
# to a job deadline and confirm the state with a GET rather than trusting the wait.
deadline = time.time() + 15 * 60
while time.time() < deadline:
httpx.get(f"{API}/{m['id']}/wait", headers=H,
params={"state": "destroyed", "timeout": 60}, timeout=75)
g = httpx.get(f"{API}/{m['id']}", headers=H, timeout=30)
if g.status_code == 404 or g.json().get("state") == "destroyed":
return m["id"]
httpx.post(f"{API}/{m['id']}/stop", headers=H, timeout=30) # past deadline: kill it
raise TimeoutError(job_id)Four details in that code carry the design. The region list is the capacity fallback. A 422 means invalid config and is not retried, while other errors move on to the next region. The metadata field lets a periodic sweeper list Machines and destroy any whose job is already marked done, which covers the case where both the job and the controller die. And the job's real result goes to object storage or a database, never to the Machine's disk, because the disk disappears with the Machine. The /wait call accepts state (started, stopped, suspended or destroyed), timeout in seconds (default 60) and instance_id, which the reference says is required when waiting for stopped. The code re-reads the Machine after each wait instead of inferring the outcome from the wait's status code, which is the safer habit.
For a queue with steady traffic, creating a Machine per job is wasteful. The usual pattern is a pool: create N Machines once, leave them stopped, and start one per job with /start. A start reuses the existing config and image on the host, so it is faster and less likely to hit capacity limits than a create. Use a lease when more than one controller might act on the same pool member, so two schedulers do not hand the same Machine two jobs.
Autostop and autostart
For web services you usually let the Fly Proxy manage state instead of calling the API. Three settings in fly.toml control it:
[http_service]
internal_port = 8080
force_https = true
auto_stop_machines = "suspend" # "off", "stop" or "suspend"
auto_start_machines = true
min_machines_running = 1 # kept running in the primary region
[http_service.concurrency]
type = "requests"
soft_limit = 40 # proxy's idea of "busy" for one Machine
hard_limit = 60New apps default to auto_stop_machines = "stop", auto_start_machines = true and min_machines_running = 0. The proxy runs a loop every few minutes. On each pass it checks whether a region has more capacity than its current load needs, judged against each Machine's soft_limit, and if so it stops or suspends at most one Machine in that region. On the way up, a request that arrives when every Machine is busy or asleep makes the proxy start one. So scaling down is slow and gradual by design, and scaling up is driven by requests. The ceiling is the number of Machines you have created: autostart never creates new ones. Create enough with fly scale count to cover your peak.
A worked example: an internal API with 6 Machines in one region, a soft_limit of 40, and overnight load of 30 concurrent requests. One Machine can carry that, so the proxy removes one Machine per pass until it reaches min_machines_running. At a few minutes per pass, 6 down to 1 takes roughly a quarter of an hour. At 08:00 the load climbs, the remaining Machine passes its soft limit, and the proxy starts others as requests arrive. The first request routed to each waking Machine pays the wake-up time.
Suspend: faster wake-ups, with caveats
Suspend saves the Machine's memory to disk instead of shutting the process down, and a later start resumes it from that snapshot. That skips both the boot and your application's own start-up (JIT warm-up, loading a model, filling a cache), which is where most cold-start time goes. The documented requirements are 2 GB of memory or less (larger works but suspends slowly), no swap, no schedule, and a Machine updated after June 20, 2024.
The caveats come from what a snapshot is. After resume, the clock can be a few seconds behind until NTP catches up, which affects token expiry checks and anything comparing timestamps. TCP connections that were open look alive to your process but are dead on the network, so database pools must validate connections on borrow, or reconnect on the first error. Unlike stop, suspend does not reset the rootfs. And snapshots are not guaranteed: a deploy, host migration or hardware fault discards them, and the Machine cold-starts instead. Your application must therefore work correctly on both paths, so test both.
Failure modes
- Capacity errors on create. A popular region can be full for a size. Keep a fallback region list, or keep a pool of pre-created stopped Machines.
- Single-Machine services. One Machine on one host is one outage away from downtime. Run two or more, in different hosts or regions.
- Data on the rootfs. It disappears on the next start. Use a volume, a database or object storage.
- Partial updates. An update with only the changed field drops the rest of the config. Always read, modify and write the full config.
- Orphan Machines. A crashed controller leaves Machines behind. Tag them with metadata, use
auto_destroy, and run a sweeper. - Wedged transitions. Treat more than five minutes in a transient state (ten for updating) as stuck. Alert on it and destroy or replace the Machine.
- Stale connections after resume. Validate pooled connections, and expect clock skew for a few seconds.
Trade-offs
Compared with a function platform such as AWS Lambda, a Machine is a whole VM running your container: any language, long-lived processes, WebSockets, and real control over placement. In return you get less managed scaling, since you decide how many Machines exist, and per-host failure is your concern. Compared with Cloud Run, Fly gives you explicit placement per region and an API for individual VMs. Cloud Run gives you automatic instance creation up to a limit and a regional service that hides hosts. Machines fit best for per-user or per-tenant sandboxes, regional low-latency apps, and job runners that need a full VM. They fit worst for very spiky workloads that need thousands of new instances in seconds with no planning. For the wider trade-off, see serverless architecture.
What to do next
- List your app's Machines with
fly machine listand confirm every production process has at least two Machines on different hosts. - Decide per service between
"stop"and"suspend"; for suspend, test resume with a stale database connection and a skewed clock. - Set
soft_limitfrom a load test, not a guess, and create enough Machines to cover peak, since autostart never creates new ones. - If you call the API, replace polling with
/wait, handle capacity errors with a region fallback, and always send full configs on update. - Tag API-created Machines with metadata and schedule a sweeper that destroys orphans.
- Alert on Machines in a transient state for more than five minutes.