Amazon Aurora is a MySQL- and PostgreSQL-compatible relational database that AWS rebuilt around one idea: separate the database engine from a distributed, replicated storage volume that every instance in the cluster shares. Your SQL, drivers and most extensions work as they do on the open-source engines, but replication, durability, backups, failover and scaling behave very differently from a self-managed database or standard RDS.

This is an operator's guide. It summarises the storage design only as far as you need it to reason about behaviour (the internals have their own article, Aurora storage architecture), then covers cluster topology and endpoints, failover and client behaviour, a worked setup with the AWS CLI, Serverless v2, Global Database, the two storage billing configurations, cloning, Backtrack and blue/green deployments, failure modes and a checklist. Numbers were checked against AWS documentation in October 2026; confirm limits for your engine version.

Advertisement

What Aurora is, and is not

An Aurora DB cluster is one storage volume plus up to sixteen DB instances: one writer and up to fifteen Aurora Replicas. Instances hold no durable data of their own; they are compute with a buffer cache. That is the source of most of Aurora's operational advantages: adding a replica does not copy data, failover does not need to catch a replica up from a binlog, and backups are continuous from storage rather than taken from an instance.

Aurora is a different product from standard RDS for MySQL or PostgreSQL, which run the stock engine on block storage with a separate standby (see AWS RDS). It is also distinct from Aurora DSQL, a separate distributed SQL service with its own architecture and limits, and from Aurora PostgreSQL Limitless Database, which shards across many writers. This page is about provisioned and Serverless v2 Aurora clusters.

The storage volume in one section

The cluster volume is split into segments, and each segment is stored six times, two copies in each of three Availability Zones. A write is durable when four of six copies acknowledge it; a read that needs storage quorum uses three of six. Those numbers mean the volume survives the loss of a whole AZ and still accepts writes, and survives an AZ plus one more copy without losing data. The writer sends only redo log records to storage, and storage nodes materialise pages from the log in the background.

Two practical consequences follow. The volume grows automatically and you pay for what is used, up to 128 TiB, or 256 TiB on recent versions (Aurora PostgreSQL 17.5, 16.9 and 15.13 and later, and Aurora MySQL 3.10 and later). And replicas read the same volume, so replica lag is the time to apply log records to their caches, usually far smaller than binlog replication lag. Monitor it rather than assuming it, with the AuroraReplicaLag metric.

Advertisement

Cluster topology and endpoints

Applicationwrites + readscluster endpointreader endpointWriterAZ a, tier 0Reader 1AZ b, tier 1Reader 2AZ c, tier 1Reader 3custom endpoint: analyticsredo logShared cluster volume: 6 copies across 3 AZs, write quorum 4/6, read quorum 3/6AZ a2 storage copiesAZ b2 storage copiesAZ c2 storage copiesstorage-level replicationGlobal Database secondaryanother Region, read-only until promoted
One writer and three readers in three AZs share a six-copy volume. Applications use the cluster endpoint for writes, the reader endpoint for scaled reads, and a custom endpoint to isolate analytics. A Global Database secondary replicates at the storage layer to another Region.

Aurora gives you DNS endpoints rather than instance addresses. The cluster endpoint always resolves to the current writer. The reader endpoint spreads connections across available replicas, at connection time, not per query. Custom endpoints point at a chosen subset of instances, which is how you send heavy reporting queries to a large replica without disturbing the replicas that serve users. Instance endpoints exist for diagnostics and should not be in application configuration.

Because balancing happens when a connection opens, long-lived pools stick to whichever replica they first reached. If you add replicas under load, old pools will not move until connections are recycled; set a maximum connection lifetime in the pool so load redistributes within minutes.

Failover and what clients see

When the writer fails, Aurora promotes a replica, choosing by promotion tier (0 is highest priority, 15 lowest) and then by size. It updates the cluster endpoint's DNS to the new writer. With no replicas, Aurora must create a new instance, which takes much longer, so production clusters should always have at least one replica in another AZ. Failover typically completes well under a minute, but the client experience depends on how quickly your drivers notice.

Clients see broken connections, then possibly stale DNS that still resolves to the old writer, now read-only or gone. Defences: keep DNS caching short in the runtime (some runtimes, the JVM among them, cache lookups, so set the TTL explicitly), use a failover-aware driver such as the AWS Advanced JDBC Wrapper, or put RDS Proxy in front so the application keeps its connections to the proxy while the proxy reconnects. Then make writes retryable:

import contextlib, time, psycopg2

RETRYABLE = {"40001", "57P01", "08006", "08003"}   # serialization failure, admin shutdown, connection lost

def run_tx(dsn, work, attempts=5):
    for i in range(attempts):
        try:
            conn = psycopg2.connect(dsn, connect_timeout=3)         # dsn uses the cluster endpoint
            with contextlib.closing(conn), conn:                   # psycopg2's `with conn` alone does not close
                with conn.cursor() as cur:
                    cur.execute("SELECT pg_is_in_recovery()")       # guard against stale DNS
                    if cur.fetchone()[0]:
                        raise psycopg2.OperationalError("connected to a reader")
                    result = work(cur)                              # must be idempotent
                return result
        except psycopg2.OperationalError:
            time.sleep(min(0.2 * 2 ** i, 5))
        except psycopg2.Error as e:
            if e.pgcode not in RETRYABLE:
                raise
            time.sleep(min(0.2 * 2 ** i, 5))
    raise RuntimeError("transaction failed after retries")

The pg_is_in_recovery() check catches the case where stale DNS sends you to a demoted instance; on Aurora MySQL the equivalent is checking @@innodb_read_only. Retries are only safe if the work is idempotent, so use natural keys or idempotency tokens for inserts.

Worked example: a three-AZ PostgreSQL cluster

An order service needs one writer, two readers for user traffic in other AZs, and I/O-heavy workloads that make I/O-Optimized attractive. Secrets come from AWS Secrets Manager, managed by RDS.

aws rds create-db-cluster \
  --db-cluster-identifier orders \
  --engine aurora-postgresql --engine-version 16.9 \
  --master-username app_admin --manage-master-user-password \
  --db-subnet-group-name orders-private --vpc-security-group-ids sg-0abc123 \
  --storage-type aurora-iopt1 \
  --backup-retention-period 14 --deletion-protection --storage-encrypted

AZS=(us-east-1a us-east-1b us-east-1c)
for i in 1 2 3; do
  aws rds create-db-instance \
    --db-instance-identifier orders-$i --db-cluster-identifier orders \
    --availability-zone ${AZS[$((i-1))]} \
    --engine aurora-postgresql --db-instance-class db.r7g.xlarge \
    --promotion-tier $([ $i = 1 ] && echo 0 || echo 1)
done

The first instance created becomes the writer. The subnet group must span all three AZs, and the explicit --availability-zone puts one instance in each. Give both readers the same tier and the same class as the writer: a failover to a smaller replica is a performance incident on top of an availability incident. Add a custom endpoint containing a fourth, larger replica if analytics arrive later. For credential rotation patterns see secrets rotation.

Serverless v2

Aurora Serverless v2 is an instance class, db.serverless, that scales capacity in Aurora Capacity Units (ACUs) in fine increments while running, without dropping connections. You set a cluster-wide minimum and maximum, for example --serverless-v2-scaling-configuration MinCapacity=0.5,MaxCapacity=16. On newer engine versions the range runs from 0 to 256 ACUs. A minimum of 0 enables automatic pause after an idle period you configure between 300 seconds and one day; while paused you pay only for storage, and the first connection resumes the instance, which AWS says can take up to about 15 seconds.

Serverless and provisioned instances can be mixed in one cluster, which is a good pattern: a provisioned writer for steady load and serverless readers for spiky reads, or the reverse for development. Keep the minimum high enough that the buffer cache survives quiet periods, because capacity also bounds memory and a cold cache can make the first minutes after a scale-up slow. Readers in promotion tiers 0 and 1 scale with the writer so they are ready to take over.

Global Database

A Global Database adds read-only secondary clusters in other Regions, up to ten since May 2025, replicated at the storage layer rather than through the engine. Readers in each Region serve local queries. There are two ways to move the writer. A switchover is planned: Aurora waits for secondaries to catch up, so no data is lost. A failover is for a Regional outage: you promote a secondary, and anything not yet replicated is lost, so your recovery point is the replication lag at the moment of failure. Watch lag continuously and alarm on it.

Write forwarding lets applications in a secondary Region send writes that are executed on the primary, which simplifies code at the cost of cross-Region latency on every forwarded write. Multi-Region design trade-offs more broadly are in multi-Region on AWS.

Standard versus I/O-Optimized

Aurora has two storage configurations. Standard charges for storage and per I/O request; I/O-Optimized has no per-request I/O charge but higher instance and storage prices. The decision is arithmetic: look at a month of bills, and if I/O is a large share of total Aurora spend, model the cluster under I/O-Optimized. AWS's guidance points at roughly a quarter of spend as the crossover. You can switch with modify-db-cluster --storage-type, but after a change you cannot switch back for 30 days, so measure first.

Clones, Backtrack, snapshots and blue/green

  • Fast clones create a new cluster that shares pages with the source copy-on-write. Creating one is quick regardless of size, and you pay only for pages that diverge. Use them for testing migrations against production-sized data.
  • Point-in-time restore creates a new cluster from continuous backups within your retention period. It never rewrites the original cluster.
  • Backtrack (Aurora MySQL only) rewinds the existing cluster in place to a point within a window of up to 72 hours. It must be enabled when the cluster is created or restored, and it rewinds everything, so it fits a bad deploy on a single-application database rather than one tenant's mistake.
  • Blue/green deployments create a synchronised copy for major version upgrades or parameter changes, then switch over with brief downtime. Test the green side with a clone of real traffic before switching.

Failure modes

FailureSymptomDefence
Writer fails with no replicaLong outage while a new instance is createdAt least one replica in another AZ, same class
Stale DNS after failoverWrites fail with read-only errorsShort DNS TTL in runtime, failover-aware driver or RDS Proxy
Pools pinned to one readerOne replica hot, others idleMaximum connection lifetime in the pool
Long transaction on writerUndo/history growth, replica lag, slow purgeStatement and idle-in-transaction timeouts
Serverless minimum too lowCold cache, slow first queries after idleRaise minimum ACUs for latency-sensitive tiers
Unplanned Global failoverRecent writes missing in new primaryAlarm on lag; reconcile from upstream events
Connection stormsMax connections hit during deploysRDS Proxy or pooled connections with limits

Trade-offs and when not to choose Aurora

Aurora buys fast failover, cheap replicas and painless storage growth at a higher price per instance than standard RDS and with some engine features unavailable or behaving differently. A small, steady database that tolerates a minute or two of failover may be cheaper on standard RDS. Write throughput is still bounded by a single writer per cluster, so very high write volumes need sharding, Limitless Database or a different store such as DynamoDB. For a refresher on what transactional guarantees you are buying, see transactional databases.

What to do next

  1. Inventory your clusters: number of replicas, their AZs, instance classes and promotion tiers; fix any cluster with no replica in another AZ.
  2. Check how your application connects: cluster endpoint for writes, reader or custom endpoints for reads, a maximum pool connection lifetime and an explicit DNS TTL.
  3. Run a failover with failover-db-cluster in a staging environment and measure the client-side error window; repeat with RDS Proxy or a failover-aware driver.
  4. Pull three months of bills and compute the I/O share to decide between Standard and I/O-Optimized.
  5. Alarm on AuroraReplicaLag and, for Global Database, cross-Region lag; write down your recovery point objective.
  6. Practise a migration on a fast clone, then rehearse a blue/green switchover before the next major version upgrade.
Key takeaway: Aurora separates compute from a six-copy, three-AZ storage volume, so replicas share data, failover promotes rather than rebuilds, and storage grows on its own. Operate it through its endpoints: the cluster endpoint for writes, reader and custom endpoints for reads, with pools that recycle connections and clients that survive failover. Use promotion tiers deliberately, size Serverless v2 minimums for cache warmth, treat Global Database failover lag as data loss, and choose the storage configuration from billing data.