A Fly Machine is a fast-starting virtual machine that runs one container image on one host in one region. Everything else on Fly.io is built from Machines: fly deploy creates and updates them, the proxy starts and stops them, and Fly's Postgres runs on them. You can also drive them directly through a REST API, which turns Fly into a place where you can start a VM per job, per user or per tenant in a fraction of a second and pay for it only while it runs.

This article covers the Machine itself: what it is underneath, the state machine it moves through, the API you use to control it, how automatic stopping, starting and suspending work, and what fails. Routing between regions and replaying requests is covered in Fly.io regions and persistent disks in Fly.io volumes. The worked example is a job runner that starts one Machine per task, which exercises nearly every part of the API.

What a Machine is

Underneath, a Machine is a Firecracker microVM. Fly takes your OCI image, unpacks it into a root filesystem, and boots a small VM with its own kernel, so tenants are isolated by hardware virtualization rather than only by Linux namespaces. Three properties follow, and they shape how you should build on it.

A Machine is pinned. It lives on a specific host in a specific region. It does not float across a pool like a Kubernetes pod. If that host has a problem, the Machine is unavailable until the host recovers or Fly migrates it, which is why you run at least two Machines for anything that must stay up.

The root filesystem is disposable. When a stopped Machine starts again, its rootfs is reset to the image. Anything written outside a volume is gone. Persistent data goes on a volume, which is tied to the same host, or into an external database or object store.

Configuration is versioned. Updating a Machine (a new image, more memory, a changed environment variable) produces a new version; the old one moves to the terminal replaced state. The API requires the whole configuration on every update, not a patch, so a client that sends only the field it changed wipes the rest.

A Fly Machine's life: persistent, transient and terminal statescreatedconfig storedstartedrunning, routablestoppedrootfs reset on startsuspendedmemory snapshotfailedcould not startdestroyedterminalreplacedold versionstartstop / exitstartsuspendstart = resumedestroyupdateFly Proxyanycast edge, TLS, load balancingautostart: wakes a stopped/suspended Machineautostop: every few minutes, stops at mostone Machine per region per pass when there isexcess capacity (judged by soft_limit)Machines APIcreate, update (full config), start, stopsuspend, wait for state, lease, cordonabout 1 req/s per action per Machineapi.machines.dev or _api.internal:4280capacity errors are yours to retryTransient states (starting, stopping, suspending, updating, replacing, destroying) sit on the arrows.
Machine states from the Fly docs, and the two controllers that move Machines between them.

The state machine

Fly documents five persistent states: created (exists, never started), started (running and reachable), stopped (exited on its own or was stopped), suspended (memory saved to disk, will resume on next start) and failed (could not start). The transient states are creating, starting, stopping, restarting, suspending, destroying, launch_failed, updating and replacing. The terminal states are destroyed, replaced and migrated.

Two operational facts come from the same page. A Machine that sits in starting, stopping, restarting or destroying for more than five minutes, or in updating for more than ten, is probably wedged, so build that timeout into your orchestration. And a stopped Machine costs much less than a running one but not nothing: Fly bills root filesystem storage for stopped Machines under its current pricing, so thousands of idle Machines left behind by a job runner do add up. Check the pricing page rather than assuming zero.

The Machines API

The Machines API is a REST API. From the internet use https://api.machines.dev; from inside your Fly private network use http://_api.internal:4280. Authenticate with Authorization: Bearer <token>, ideally a deploy token scoped to one app (fly tokens deploy). The core endpoints all sit under /v1/apps/{app}/machines:

CallPurpose
POST /machinesCreate (and by default start) a Machine from a config
GET /machines/{id}Read state, config, region, instance id
POST /machines/{id}Update; send the full config
POST /machines/{id}/startStart a stopped Machine, or resume a suspended one
POST /machines/{id}/stopStop: process gets a signal, then the VM shuts down
POST /machines/{id}/suspendSnapshot memory to disk and pause
GET /machines/{id}/waitBlock until the Machine reaches a requested state, or time out
POST /machines/{id}/leaseTake an exclusive lease so two controllers do not fight
POST /machines/{id}/cordonStop the proxy routing new requests to it
DELETE /machines/{id}Destroy

Rate limits are per action and per Machine: about 1 request per second per action, with bursts to 3, and 5 per second (burst 10) for reading a Machine. That makes a polling loop on GET the wrong design; use /wait, which holds the request open until the state changes. And creating a Machine can fail because the region has no capacity for that size at that moment. The docs are explicit that handling this is your job: retry, or fall back to another region.

Worked example: one Machine per job

Here is a job runner that starts one Machine per task, waits for it to finish, and cleans up. Each job runs a container that does its work and exits. With auto_destroy set and the restart policy set to no, the Machine destroys itself when the process exits, so a crashed controller does not leave it behind.

import os, time, httpx

API = "https://api.machines.dev/v1/apps/thumbs-jobs/machines"
H = {"Authorization": f"Bearer {os.environ['FLY_API_TOKEN']}"}
REGIONS = ["ams", "fra", "cdg"]          # preferred first; fall back on capacity errors

def run_job(job_id: str, src_url: str) -> str:
    cfg = {
        "image": "registry.fly.io/thumbs-worker:2026-10-01",
        "env": {"JOB_ID": job_id, "SRC_URL": src_url},
        "guest": {"cpu_kind": "shared", "cpus": 2, "memory_mb": 1024},
        "auto_destroy": True,                # destroy when the process exits
        "restart": {"policy": "no"},         # a failed job is retried by us, not by Fly
        "metadata": {"job_id": job_id},      # lets a sweeper find orphans
    }
    last_err = None
    for region in REGIONS:
        r = httpx.post(API, headers=H, json={"region": region, "config": cfg}, timeout=30)
        if r.status_code in (200, 201):
            m = r.json()
            break
        last_err = r.text                    # capacity or validation error; try next region
        if r.status_code == 422:
            raise ValueError(last_err)       # bad config: no point retrying elsewhere
    else:
        raise RuntimeError(f"no capacity in {REGIONS}: {last_err}")

    # Long-poll until the job exits. Each wait call is bounded (default 60 s), so loop
    # to a job deadline and confirm the state with a GET rather than trusting the wait.
    deadline = time.time() + 15 * 60
    while time.time() < deadline:
        httpx.get(f"{API}/{m['id']}/wait", headers=H,
                  params={"state": "destroyed", "timeout": 60}, timeout=75)
        g = httpx.get(f"{API}/{m['id']}", headers=H, timeout=30)
        if g.status_code == 404 or g.json().get("state") == "destroyed":
            return m["id"]
    httpx.post(f"{API}/{m['id']}/stop", headers=H, timeout=30)   # past deadline: kill it
    raise TimeoutError(job_id)

Four details in that code carry the design. The region list is the capacity fallback. A 422 means invalid config and is not retried, while other errors move on to the next region. The metadata field lets a periodic sweeper list Machines and destroy any whose job is already marked done, which covers the case where both the job and the controller die. And the job's real result goes to object storage or a database, never to the Machine's disk, because the disk disappears with the Machine. The /wait call accepts state (started, stopped, suspended or destroyed), timeout in seconds (default 60) and instance_id, which the reference says is required when waiting for stopped. The code re-reads the Machine after each wait instead of inferring the outcome from the wait's status code, which is the safer habit.

For a queue with steady traffic, creating a Machine per job is wasteful. The usual pattern is a pool: create N Machines once, leave them stopped, and start one per job with /start. A start reuses the existing config and image on the host, so it is faster and less likely to hit capacity limits than a create. Use a lease when more than one controller might act on the same pool member, so two schedulers do not hand the same Machine two jobs.

Autostop and autostart

For web services you usually let the Fly Proxy manage state instead of calling the API. Three settings in fly.toml control it:

[http_service]
  internal_port = 8080
  force_https = true
  auto_stop_machines = "suspend"     # "off", "stop" or "suspend"
  auto_start_machines = true
  min_machines_running = 1           # kept running in the primary region

  [http_service.concurrency]
    type = "requests"
    soft_limit = 40                  # proxy's idea of "busy" for one Machine
    hard_limit = 60

New apps default to auto_stop_machines = "stop", auto_start_machines = true and min_machines_running = 0. The proxy runs a loop every few minutes. On each pass it checks whether a region has more capacity than its current load needs, judged against each Machine's soft_limit, and if so it stops or suspends at most one Machine in that region. On the way up, a request that arrives when every Machine is busy or asleep makes the proxy start one. So scaling down is slow and gradual by design, and scaling up is driven by requests. The ceiling is the number of Machines you have created: autostart never creates new ones. Create enough with fly scale count to cover your peak.

A worked example: an internal API with 6 Machines in one region, a soft_limit of 40, and overnight load of 30 concurrent requests. One Machine can carry that, so the proxy removes one Machine per pass until it reaches min_machines_running. At a few minutes per pass, 6 down to 1 takes roughly a quarter of an hour. At 08:00 the load climbs, the remaining Machine passes its soft limit, and the proxy starts others as requests arrive. The first request routed to each waking Machine pays the wake-up time.

Suspend: faster wake-ups, with caveats

Suspend saves the Machine's memory to disk instead of shutting the process down, and a later start resumes it from that snapshot. That skips both the boot and your application's own start-up (JIT warm-up, loading a model, filling a cache), which is where most cold-start time goes. The documented requirements are 2 GB of memory or less (larger works but suspends slowly), no swap, no schedule, and a Machine updated after June 20, 2024.

The caveats come from what a snapshot is. After resume, the clock can be a few seconds behind until NTP catches up, which affects token expiry checks and anything comparing timestamps. TCP connections that were open look alive to your process but are dead on the network, so database pools must validate connections on borrow, or reconnect on the first error. Unlike stop, suspend does not reset the rootfs. And snapshots are not guaranteed: a deploy, host migration or hardware fault discards them, and the Machine cold-starts instead. Your application must therefore work correctly on both paths, so test both.

Failure modes

  • Capacity errors on create. A popular region can be full for a size. Keep a fallback region list, or keep a pool of pre-created stopped Machines.
  • Single-Machine services. One Machine on one host is one outage away from downtime. Run two or more, in different hosts or regions.
  • Data on the rootfs. It disappears on the next start. Use a volume, a database or object storage.
  • Partial updates. An update with only the changed field drops the rest of the config. Always read, modify and write the full config.
  • Orphan Machines. A crashed controller leaves Machines behind. Tag them with metadata, use auto_destroy, and run a sweeper.
  • Wedged transitions. Treat more than five minutes in a transient state (ten for updating) as stuck. Alert on it and destroy or replace the Machine.
  • Stale connections after resume. Validate pooled connections, and expect clock skew for a few seconds.

Trade-offs

Compared with a function platform such as AWS Lambda, a Machine is a whole VM running your container: any language, long-lived processes, WebSockets, and real control over placement. In return you get less managed scaling, since you decide how many Machines exist, and per-host failure is your concern. Compared with Cloud Run, Fly gives you explicit placement per region and an API for individual VMs. Cloud Run gives you automatic instance creation up to a limit and a regional service that hides hosts. Machines fit best for per-user or per-tenant sandboxes, regional low-latency apps, and job runners that need a full VM. They fit worst for very spiky workloads that need thousands of new instances in seconds with no planning. For the wider trade-off, see serverless architecture.

What to do next

  1. List your app's Machines with fly machine list and confirm every production process has at least two Machines on different hosts.
  2. Decide per service between "stop" and "suspend"; for suspend, test resume with a stale database connection and a skewed clock.
  3. Set soft_limit from a load test, not a guess, and create enough Machines to cover peak, since autostart never creates new ones.
  4. If you call the API, replace polling with /wait, handle capacity errors with a region fallback, and always send full configs on update.
  5. Tag API-created Machines with metadata and schedule a sweeper that destroys orphans.
  6. Alert on Machines in a transient state for more than five minutes.
Key takeaway: A Fly Machine is a Firecracker microVM pinned to one host, with a disposable root filesystem and a versioned config that must be replaced whole. Drive it through the Machines API with wait calls, region fallback and orphan cleanup, or let the proxy stop, suspend and start it based on soft_limit. Suspend skips cold starts but brings clock skew and dead connections, so test both the resume path and the cold path.