A dyno is the unit Heroku runs your code in: an isolated Linux container started from your app's compiled slug, with your config vars in its environment, running one command from your Procfile. You never see the machine underneath. You choose how many dynos of each process type to run and how big each one is, and Heroku places, restarts and routes to them.
That abstraction is simple, and it carries rules that decide whether your app is reliable: how long a web process has to boot, how long it has to shut down, how long a request may take, what happens when memory runs out, and how often the platform restarts everything. This article covers those rules with the documented numbers, shows code that respects them, works through a sizing example, and ends with the platform's changed status in 2026 and when to choose something else.
The process model
Heroku follows the twelve-factor model. A build turns your repository into a slug, a release pairs that slug with the current config vars, and dynos run the release. The Procfile at the root of the repository names the process types:
web: gunicorn app:app --config gunicorn.conf.py
worker: python -m jobs.worker
release: python manage.py migrate --noinputThree kinds of dynos exist. Web dynos are the only ones that receive HTTP traffic from the router, and they must listen on the port in the $PORT environment variable. Worker dynos run background processes with no inbound HTTP; they pull work from a queue in an add-on. One-off dynos run an ad hoc command such as heroku run python manage.py shell and stop when it exits. The release entry is special: it runs once per release, before new dynos start, and a failed release command stops the deploy.
Each dyno gets its own ephemeral filesystem with a fresh copy of the deployed code. Files written while it runs are discarded when it stops or restarts, and no other dyno can see them. Uploads, sessions and caches therefore belong in object storage, Postgres or Redis, never on local disk.
Choosing a dyno type
On the Cedar generation, which most apps run on, the Common Runtime offers these types. Eco dynos cost a flat $5 a month for a pool of 1,000 dyno hours, and an Eco web dyno sleeps after 30 minutes without traffic, taking any Eco worker with it; the first request after that waits for the dyno to wake.
| Type | Memory | CPU | Notes |
|---|---|---|---|
| Eco | 0.5 GB | 1x share | Sleeps after 30 min idle; hobby use |
| Basic | 0.5 GB | 1x share | Never sleeps |
| Standard-1X | 0.5 GB | 1x share | Horizontal scaling, preboot |
| Standard-2X | 1 GB | 2x share | Same, more memory |
| Performance-M | 2.5 GB | dedicated | Autoscaling available |
| Performance-L | 14 GB | dedicated | Autoscaling available |
| Performance-L-RAM / XL / 2XL | 30 / 62 / 126 GB | dedicated | Memory-heavy workloads |
Private Spaces and Shield Private Spaces have their own Private-S to Private-2XL and Shield-S to Shield-2XL types, all with dedicated CPU, from 1 GB to 126 GB. The newer Fir generation, available only in Private Spaces, runs dynos as Kubernetes pods. It names sizes by vCPU and memory, such as dyno-2c-8gb, in Classic, General Purpose, Compute and Memory families from 0.5 GB up to 128 GB, and offers both x86 and ARM.
Shared-CPU types suit I/O-bound web apps whose dynos mostly wait on databases and APIs. Choose dedicated CPU when latency varies with neighbour load, or when you run CPU-bound work such as image processing or tokenisation-heavy LLM pre-processing. Choose memory first, though: most dyno trouble is memory, not CPU.
Lifecycle: boot, cycle, crash and stop
Four timers govern every dyno. Memorise them, because each maps to an error code you will see in heroku logs --tail.
- Boot. A web process must bind to
$PORTwithin 60 seconds or the dyno is killed and marked crashed (R10). Meanwhile the router queues requests for up to 75 seconds waiting for a web dyno to come up; after that it returns H20. - Daily cycling. On Cedar, every dyno restarts at least once a day: 24 hours plus up to 216 random minutes after its last restart, so dynos do not all cycle together. Deploys and manual restarts reset the clock. Fir apps can opt out through a Heroku Labs flag.
- Crash backoff. In the Common Runtime, the first crash restarts immediately. Repeated crashes wait through cool-offs of up to 20, 40, 60, 180 and then 320 minutes, and every 320 minutes after that, until a successful start, a new release, a manual restart or a scale to zero and back resets it. Cedar Private Spaces restart continuously with no cool-off; Fir uses Kubernetes' own backoff.
- Shutdown. Stopping a dyno sends SIGTERM to every process in it. Processes have 30 seconds to exit; anything still running then gets SIGKILL and the log shows R12.
The shutdown rule is the one code must handle. A web server should stop accepting connections and finish in-flight requests well inside 30 seconds. A worker should stop taking jobs, finish or return the current one to the queue, and exit:
# gunicorn.conf.py
import os
bind = f"0.0.0.0:{os.environ.get('PORT', '8000')}"
workers = int(os.environ.get("WEB_CONCURRENCY", "2"))
timeout = 28 # kill a stuck worker before the router's 30 s
graceful_timeout = 20 # drain in-flight requests after SIGTERM, inside the 30 s budget
# jobs/worker.py
import signal, sys
stopping = False
def on_term(signum, frame):
global stopping
stopping = True # finish the current job, take no new ones
signal.signal(signal.SIGTERM, on_term)
def run(queue):
while not stopping:
job = queue.reserve(timeout=5) # visibility timeout held by the queue
if job is None:
continue
try:
job.perform() # keep each job well under 30 s, or checkpoint
queue.ack(job)
except Exception:
queue.release(job) # put it back for another worker
raise
sys.exit(0)The queue object stands for whatever client you use; what matters is reserve, acknowledge and release semantics, so a job killed by SIGKILL reappears instead of vanishing. Because daily cycling guarantees a SIGTERM to every dyno every day, a worker that mishandles it loses work regularly, not just during deploys.
The router's 30-second rule
The HTTP router gives an app 30 seconds to send the first byte of a response. After that, each byte sent or received resets a rolling 55-second window, so streamed responses can run longer as long as data keeps flowing. Missing the first-byte limit logs H12; an app that goes silent for 55 seconds mid-response logs H15. Either way the client gets an error page. The dyno is not told: it keeps working on a request whose response nobody will read, holding a worker slot while new requests queue behind it.
That is why the gunicorn timeout above sits just under 30 seconds. More importantly, it is why slow work belongs on workers. Accept the request, enqueue a job, return 202 with a status URL, and let the client poll or receive a webhook. For LLM features, stream tokens so the first byte arrives quickly and the rolling window keeps resetting.
Memory errors
When a dyno needs more memory than its type provides, Heroku logs R14 and the dyno starts using swap, which slows it sharply. If it consumes vastly more than its quota, the platform kills it with SIGKILL and logs R15, a check based on resident memory plus swap. To see memory over time, enable runtime metrics with heroku labs:enable log-runtime-metrics, which adds periodic memory and load samples to the log stream on Cedar.
Most R14s come from concurrency, not leaks. Each gunicorn or Puma worker is a full copy of the app. Set WEB_CONCURRENCY from measured memory per process, leave about a quarter of the quota as headroom, and move up a dyno size before adding processes that do not fit.
Worked example: sizing a web tier
An API serves a peak of 120 requests per second with a mean service time of 150 ms. By Little's law, concurrency is arrival rate times service time: 120 × 0.15 = 18 requests in flight on average. Plan for bursts at roughly twice that, 36 concurrent requests.
Each app process uses 180 MB and handles one request at a time. A Standard-2X has 1 GB; keeping a quarter as headroom leaves about 750 MB, so four processes fit per dyno. Thirty-six slots divided by four is nine Standard-2X dynos. A Performance-M with 2.5 GB fits about ten processes, so four dynos cover the peak with dedicated CPU, at a different price point. Price both from the current Heroku pricing page rather than from this article, then load-test, because a slow database call doubling service time doubles every number here.
heroku ps:scale web=9:standard-2x worker=2:standard-1x
heroku config:set WEB_CONCURRENCY=4
heroku ps # confirm the formation
Scaling and autoscaling
Scaling is changing the formation: the count and type per process type. Horizontal scaling is instant to request and takes a boot time to take effect. Heroku's built-in autoscaling is available for Performance, Private and Shield web dynos, and not yet on Fir. You set a desired p95 response time and a minimum and maximum count. It forecasts from the last hour of one-minute measurements and scales up when latency is degrading and predicted to cross your target. It adds or removes one web dyno per event, at least a minute apart, and scales down more cautiously than up. Workers need a queue-depth autoscaler from the add-on marketplace or your own scheduler.
The one-dyno-per-minute pace matters for spikes: going from three dynos to twelve takes roughly nine minutes, plus boot time. For predictable peaks, scale ahead on a schedule. For general strategy, see cloud autoscaling.
Deploys: release phase and preboot
By default a deploy restarts dynos with the new release, briefly dropping capacity. Preboot, enabled with heroku features:enable preboot on Standard and Performance dynos in the Common Runtime, starts new dynos first and switches traffic about three minutes after the deploy completes, then shuts the old ones down. For that window both versions run, so database migrations must be backward-compatible, and twice the dynos means twice the database connections, so use connection pooling. Disable preboot for migrations that need downtime. Private Spaces use rolling deploys instead. This is blue-green deployment managed for you.
Failure codes and what they mean
| Code | Meaning | Usual cause and fix |
|---|---|---|
| R10 | Web process did not bind $PORT in 60 s | Hard-coded port, slow boot work; bind first, warm caches after |
| H10 | App crashed | Read the lines before it; often a missing config var or R10 |
| H12 | No response within 30 s | Slow query or external call; move work to a worker |
| H13 | Connection closed without response | Process crashed mid-request or server timeout too low |
| H14 | No web dynos running | Web scaled to zero |
| H15 | No data from the dyno for 55 s mid-response | Stalled stream; send data or heartbeats |
| H20 | No web dyno up within 75 s | Slow or crashing boot |
| R12 | Did not exit within 30 s of SIGTERM | Missing signal handling; drain faster |
| R14 / R15 | Memory over quota / vastly over | Too many processes per dyno, or a leak |
Platform status in 2026 and trade-offs
On 6 February 2026, Heroku announced a move to a sustaining engineering model: stability, security, reliability and support continue, but new feature development has stopped. Enterprise contracts are no longer offered to new customers, though existing ones can renew. Fir, introduced in 2025, is generally available only in Private Spaces, and features like autoscaling have not reached it.
What Heroku still offers is a mature, low-operations runtime: git-based deploys, a clear process model, managed Postgres and Redis, and the lifecycle rules above. What you give up is control over networking and placement, GPU workloads, long requests through the router, and now a roadmap. For a stable app it remains a reasonable home. For new systems, compare with serverless platforms and managed Kubernetes, and keep apps portable: twelve-factor config, stateless dynos, and add-ons with standard protocols.
What to do next
- Check your Procfile: bind web to
$PORT, put migrations inrelease, and move slow work to aworker. - Add SIGTERM handling to every process and test it with
heroku ps:restartwhile traffic and jobs run; look for R12. - Set server timeouts just under 30 seconds and convert any endpoint that can exceed them into a job or a stream.
- Enable runtime metrics, measure memory per process, and set
WEB_CONCURRENCYfrom it with headroom. - Size the web tier with Little's law, then load-test at twice the expected peak.
- Enable preboot or confirm rolling deploys, and make migrations backward-compatible.
- Alert on H12, H10, R14 and R15 counts from the log drain.
- Write down an exit plan: which add-ons and platform features you depend on and what replaces each.