Aurora Serverless v2 is often described as an Aurora database that grows and shrinks on its own, and that sentence hides almost everything you need to run it well. It does not scale the cluster. It scales each database instance separately, in fractional capacity units, while the storage underneath stays exactly as it is. Whether that saves you money, keeps your failover safe and keeps your latency flat depends on three settings and one design choice: the minimum capacity, the maximum capacity, the auto-pause timeout, and the promotion tier of each reader.

This article builds the model from first principles: what an Aurora capacity unit is, what the scaler watches, how readers are tied to the writer, what happens to memory and connections as capacity moves, and how scale-to-zero behaves. It then works through a capacity and cost model for a real-looking daily load, gives connection code that survives a resume, and finishes with failure modes and a checklist. For how the shared storage volume works, read Aurora storage architecture; for managed relational databases in general, read AWS RDS.

Advertisement

The unit: what an ACU is, and what does not scale

An Aurora capacity unit (ACU) is roughly 2 GiB of memory plus proportional CPU and networking. You give the cluster a range, for example a minimum of 2 and a maximum of 64 ACUs, and every Serverless v2 instance in the cluster (instance class db.serverless) lives somewhere inside that range. Capacity is a floating-point number, sampled every second, and moves in increments as small as 0.5 ACU. On current engine versions the range runs from 0 to 256 ACUs: Aurora MySQL 3.08.0 and higher and Aurora PostgreSQL 13.15, 14.12, 15.7 and 16.3 and higher support the full 0 to 256 range, while older versions stop at a floor of 0.5 and a ceiling of 128 or 256. AWS also assigns a serverless platform version per cluster; the documentation lists versions 1 to 4, with versions 3 and 4 each claiming up to 30 percent better performance than the one before. You can read yours from the ServerlessV2PlatformVersion field of the cluster.

What does not scale is storage. Aurora separates compute from a shared, replicated cluster volume, so a cluster idling at half an ACU can still hold many terabytes. That separation is why scaling is cheap: no data moves when capacity changes, and in most cases the instance stays on the same host. It also means storage and I/O charges continue regardless of compute. Serverless v2 is a compute-billing model, not a whole-database billing model.

Aurora Serverless v1, which AWS now describes as deprecated, was a different architecture. Everything here is about v2, which uses the same engine and storage as provisioned Aurora.

The architecture: one range, many independent instances

Cluster (writer) endpointreads and writesReader endpointspreads read sessionsCustom endpointe.g. reporting readersWriterdb.serverless, 2-64 ACU nowReader, tier 0 or 1min = writer's capacityReader, tier 2-15scales and pauses alonecapacity follows writerShared cluster volumesix copies across three AZs; compute capacity never limits stored dataredo logpage readspage readsEach instance has its own capacity, measured in ACUs every second, inside the cluster-wide min-max range.Billing = sum over instances of ACU-hours actually used. With min 0, idle instances pause and bill nothing for compute.Storage and I/O are billed separately and do not pause.
A Serverless v2 cluster. The range is set once for the cluster, but each instance scales on its own load. Readers in promotion tiers 0 and 1 never fall below the writer's current capacity; readers in tiers 2 to 15 scale, and pause, independently. All instances share one storage volume.

The scaler tracks the load of each instance: CPU, memory pressure and network, including background work such as purge or vacuum and Aurora's own housekeeping. When any of them is constrained it adds capacity, and when capacity is idle it gives some back. Scaling happens while transactions are open, tables are locked and temporary tables exist; Aurora does not wait for a quiet point and does not disconnect sessions. Increments grow with current capacity, so an instance at 32 ACUs climbs faster in absolute terms than one at 1 ACU. So a very low minimum makes the first seconds of a burst slower.

Readers add a twist through their promotion tier, the number that also decides failover order. A reader in tier 0 or 1 has a floor equal to the writer's current capacity (or, for a provisioned writer, the ACU equivalent of its memory), so it is always big enough to take over. A reader in tiers 2 to 15 scales on its own load inside the cluster range. That suits a reporting reader busy for an hour a day, and is dangerous as a failover target: a promoted tier-2 reader may start at a fraction of the writer's size and must climb under full write load.

Advertisement

Memory and connections: the two consequences people miss

Because an ACU is mostly memory, scaling changes the buffer pool: Aurora grows the page cache as an instance scales up and shrinks it as it scales down. A database that scales down overnight throws away its hot pages, and the first morning queries read from storage until the cache warms. AWS's guidance is to set the minimum high enough that each instance holds the application's working set.

Connections work the other way. Aurora does not let max_connections follow current capacity, because shrinking the limit would drop live sessions; it derives the limit from the maximum ACU setting and keeps it fixed. A cluster with a small maximum therefore has a small connection limit, and raising the maximum is sometimes needed purely for connection headroom.

Scale to zero: auto-pause and what blocks it

Setting the minimum to 0 ACUs turns on auto-pause. An instance with no user-initiated connections for SecondsUntilAutoPause seconds pauses, and its instance charge stops; storage is still billed. The interval runs from 300 seconds, which is also the default, to 86,400 seconds. The instance status still reads available. A connection attempt, even one with a wrong password, wakes it; AWS gives about 15 seconds as a typical resume time, and 30 seconds or more if the instance has been paused for longer than 24 hours and moved into a deeper sleep. A resumed instance starts small and scales up from there, whatever capacity it had before it paused.

The writer and the tier-0 and tier-1 readers pause and resume together, and the writer cannot pause while any reader is active; connecting to any reader wakes the writer too. Several features keep an instance awake whatever you set: an RDS Proxy in front of the cluster (it holds connections open), being the primary or a secondary of a global database, logical replication on PostgreSQL or binlog replication on MySQL, a zero-ETL integration to Redshift, and, for Babelfish, activity on the T-SQL port. Any provisioned instance in the cluster also keeps the serverless writer awake. Maintenance wakes paused instances and then waits at least 20 minutes before pausing them again.

One surprise catches teams in production: scheduled jobs inside the engine, such as pg_cron or the MySQL event scheduler, do not wake a paused instance, and jobs due during a pause are skipped, not queued. Wake the instance from outside, for example with a scheduled Lambda that connects a minute early.

Setting it up

The range lives on the cluster; the instance class and promotion tier live on each instance. This creates a PostgreSQL cluster that can pause after 30 minutes, a writer, and a tier-1 reader in another Availability Zone that tracks it:

aws rds create-db-cluster \
  --db-cluster-identifier orders \
  --engine aurora-postgresql --engine-version 16.3 \
  --master-username app --manage-master-user-password \
  --serverless-v2-scaling-configuration MinCapacity=0,MaxCapacity=32,SecondsUntilAutoPause=1800

aws rds create-db-instance --db-cluster-identifier orders \
  --db-instance-identifier orders-w --engine aurora-postgresql \
  --db-instance-class db.serverless

aws rds create-db-instance --db-cluster-identifier orders \
  --db-instance-identifier orders-r1 --engine aurora-postgresql \
  --db-instance-class db.serverless --promotion-tier 1 \
  --availability-zone us-east-1b

aws rds describe-db-clusters --db-cluster-identifier orders \
  --query 'DBClusters[0].ServerlessV2ScalingConfiguration'

Turning auto-pause off is a modify-db-cluster call with a minimum of 0.5 or more. Changing the range, the parameter group or the engine version wakes paused instances. The --manage-master-user-password flag stores the master password in Secrets Manager; Secrets Manager rotation covers how rotation then works.

Worked example: sizing a range and pricing it

Take an order service with a writer and one tier-1 reader. Measured on a provisioned cluster, its writer needs about 16 ACUs' worth of memory and CPU for eight busy hours, about 4 for four shoulder hours, and almost nothing for the remaining twelve. The working set of hot tables and indexes is about 3 GiB. The question is what range to choose and what it costs compared with a provisioned instance sized for the peak.

Start with the floor. A 3 GiB working set needs about 2 ACUs of buffer pool headroom, so a minimum of 2 keeps mornings warm; a minimum of 0 saves more but accepts a cold cache and a 15-second resume. The ceiling should cover the peak with margin and give enough connections, so 32 rather than 16. The reader is in tier 1, so it costs the same as the writer. This small model turns an hourly profile into ACU-hours:

def acu_hours(profile, min_acu, max_acu, instances):
    # profile: list of (hours, needed_acu); returns ACU-hours per day for the cluster
    total = 0.0
    for hours, need in profile:
        cap = min(max(need, min_acu), max_acu)
        total += hours * cap
    return total * instances

day = [(8, 16), (4, 4), (12, 0.5)]
for floor in (0.5, 2):
    h = acu_hours(day, floor, 32, instances=2)
    print(f"min={floor}: {h:.0f} ACU-hours/day, average {h / 48:.2f} ACU per instance")

# Break-even: serverless is cheaper when
#   ACU-hours/day * price_per_ACU_hour  <  24 * instances * provisioned_hourly_price

With a floor of 0.5 the cluster uses 2 x (128 + 16 + 6) = 300 ACU-hours a day, an average of 6.25 ACUs per instance. With a floor of 2 it uses 2 x (128 + 16 + 24) = 336, an average of 7 ACUs. A provisioned writer sized for the 16-ACU peak needs roughly 32 GiB of memory and runs 24 hours. Serverless wins when 7 times the ACU-hour price is less than the hourly price of that 32 GiB instance; look both prices up for your Region and engine, because the ratio between them is the whole decision. A flat load at 15 ACUs all day would lose to provisioned.

Connecting to a database that might be asleep

Applications need two changes for auto-pause. Connection timeouts must exceed the resume time, so 30 seconds rather than the 5 or 10 many pools default to, and connection attempts must retry on transient errors, because a burst of connections during a resume can exceed limits on the first try. Do not retry authentication failures. In Python with psycopg:

import time, psycopg

TRANSIENT = (psycopg.OperationalError,)

def connect(dsn, attempts=4):
    delay = 1.0
    for i in range(attempts):
        try:
            # connect_timeout above the documented resume time (about 15 s, 30 s+ after a day paused)
            return psycopg.connect(dsn, connect_timeout=35)
        except TRANSIENT as e:
            if "password authentication failed" in str(e) or i == attempts - 1:
                raise
            time.sleep(delay)
            delay *= 2

In Java the equivalent settings in the PostgreSQL JDBC driver are connectTimeout and, with TLS, sslResponseTimeout. Long-lived idle pools defeat auto-pause, since any open user connection keeps an instance awake, so set pool minimum idle to zero and an idle timeout below the pause interval for clusters that should sleep. For Lambda-heavy traffic the RDS Data API is an alternative to connection pools, and a Data API call wakes a paused writer; see AWS serverless architecture for the surrounding patterns.

Monitoring and failure modes

Two CloudWatch metrics carry most of the signal. ServerlessDatabaseCapacity is the current ACUs per instance; its minimum over a period tells you whether the instance paused, its maximum shows burst height. ACUUtilization is capacity as a percentage of the maximum; sustained values near 100 mean the ceiling is the bottleneck. DatabaseConnections and the instance's auto-pause log explain why an instance did not pause.

SymptomLikely causeFix
ACUUtilization pinned at 100, latency risingMaximum too low for the peakRaise the maximum; add tier 2-15 readers for read load
Slow first queries after idle periodsBuffer pool shrank with capacityRaise the minimum to hold the working set
Never pausesRDS Proxy, replication, global database, a pooled idle connectionCheck the auto-pause log and DatabaseConnections
Connection errors after a quiet nightClient timeout shorter than resume timeTimeouts of 30 s or more, plus retry
Slow failoverOnly failover target is a tier 2-15 readerKeep one reader in tier 0 or 1
Nightly pg_cron job silently missingInstance paused at job timeWake it from outside or raise the pause interval

Trade-offs: when to choose it, and when not

Serverless v2 fits spiky or unpredictable loads, many small per-tenant or per-developer databases, test environments that should sleep, and readers that serve occasional heavy queries. It fits poorly for loads that are flat and high all day, where provisioned or reserved capacity is cheaper per unit of work, and for latency-critical services that cannot accept a cold cache or a resume. Mixed clusters are a good middle path: a provisioned writer for steady writes and serverless readers for elastic reads.

What to do next

  1. Export a week of CPU and memory use from your current cluster and turn it into an hourly ACU profile.
  2. Set the minimum from the working set size (about 2 GiB per ACU), not from the lowest number allowed.
  3. Set the maximum from peak need plus connection headroom, then check the resulting max_connections.
  4. Put exactly one reader in tier 0 or 1 for failover; put reporting readers in tiers 2 to 15 behind a custom endpoint.
  5. If you enable auto-pause, raise client connect timeouts above 30 seconds and add retry on transient errors.
  6. Move pg_cron or event-scheduler jobs that must run to an external scheduler that connects first.
  7. Alarm on sustained ACUUtilization near 100 and on unexpected zero periods of ServerlessDatabaseCapacity.
  8. Run the cost model with your Region's ACU-hour and instance prices before migrating production.
Key takeaway: Aurora Serverless v2 scales each instance, not the cluster, in ACUs of about 2 GiB of memory each, inside a range you set once. The minimum decides how warm the cache stays, the maximum decides peak capacity and the connection limit, and the promotion tier decides whether a reader is a safe failover target or an independent, pausable worker. A minimum of zero adds auto-pause, with a resume of about 15 seconds and a list of features that prevent it. Size the range from a measured profile, price it against a provisioned instance, and make clients patient enough to wait for a resume.