Product comparison pages for cloud databases go out of date within a year, and they mostly compare feature lists. The decisions that stay with you for a decade are architectural: where durable state lives, how writes scale, what a commit waits for, and what happens when a zone or region disappears. Those follow from a small number of designs, and every product on the market is a variation on one of them.

This article sorts cloud-native databases into families, compares them on five axes, and turns that into a selection procedure you can run against your own workload. Individual products get only as much detail as the comparison needs. The site has deeper pages on Aurora's storage layer, Spanner, AlloyDB and DynamoDB.

Advertisement

What makes a database cloud-native

Running PostgreSQL on a virtual machine with a network disk is a database in the cloud, not a cloud-native one. The difference is that the classic engine assumes a local disk it owns. Replication, backup and failover are bolted on around it: a standby replays the whole write-ahead log, backups copy entire data files, and failover means promoting a replica that may be seconds behind. Managed services such as RDS and Cloud SQL automate this well, but the architecture is unchanged.

Cloud-native designs change the architecture. They take the storage layer apart from compute, or remove the single writer altogether, and they are built around facts particular to the cloud: availability zones fail independently, network bandwidth inside a region is plentiful, and compute can be added in minutes. The result falls into three families, shown in the diagram, plus serverless packaging that can be applied to any of them.

Where the data lives: three storage shapes behind cloud-native databasesPrimary (read-write)one writer, full engineReplicareadsReplicareadsShared log-structuredstorage, spread over 3 AZsredo log onlyA: shared-storageAurora, AlloyDB, NeonSQL gateway (any node)parse, plan, coordinateRange 1RaftRange 2RaftRange 3Rafteach range: 3 or 5 replicasleader per range, spread over zonesB: distributed SQLSpanner, CockroachDB, YugabyteDB, TiDBRequest routerhash(partition key)Part. A3 copiesPart. B3 copiesPart. C3 copiessingle-key ops, no joinssplit by size or throughputC: partitioned key-valueDynamoDB, Cosmos DB, BigtableWrites scale: A no (one writer), B yes (per range), C yes (per partition).SQL surface: A full, B mostly full, C minimal.
Three storage shapes. Shared-storage engines keep one writer and move durability into a replicated storage service. Distributed SQL shards data into ranges, each its own consensus group. Partitioned key-value stores hash keys onto partitions and give up joins.

Family A: shared storage, the log is the database

Aurora introduced the idea at scale. The compute node runs a mostly standard MySQL or PostgreSQL engine with its buffer pool, but it never writes data pages. It ships only redo log records to a storage service that keeps six copies of each 10 GB segment, two in each of three availability zones. A write is durable when four of the six acknowledge it. Normal reads do not need a quorum: the writer tracks which copies are current and reads one of them. The three-of-six read quorum is used during recovery, and because 4 + 3 is more than 6, it always overlaps the latest write quorum. The storage nodes apply log records to pages themselves, in the background.

The consequences are practical. Replicas attach to the same storage and receive only log records to update pages they already cache, so they don't need their own copy of the data. Failover promotes one of them with no data copy. Backups are continuous at the storage layer. But there is still exactly one writer, so write throughput is bounded by one machine. AlloyDB follows the same pattern for PostgreSQL, with a log-processing layer and a columnar cache for analytical queries. Neon splits PostgreSQL into stateless compute, a quorum of safekeepers that make the write-ahead log durable, and pageservers that rebuild pages on demand. That split makes branching a database and scaling compute to zero cheap.

Advertisement

Family B: distributed SQL

Spanner, CockroachDB, YugabyteDB and TiDB split the key space into ranges and make each range an independent consensus group, using Paxos in Spanner and Raft in the others. CockroachDB's ranges split when they pass a default of 512 MiB. Any node can accept SQL, plan the query and send reads and writes to the leaders of the ranges involved. A transaction touching several ranges uses a two-phase commit layered over the consensus groups, as described in the sharding article.

Writes scale with the number of ranges, and a zone or even a region can fail without data loss if replicas are placed for it. The costs are latency and coordination. Every write waits for a consensus round, so commit latency is at least one round trip to a majority of replicas. Transactions must be ordered without a single clock. Spanner uses TrueTime, whose timestamps carry an explicit uncertainty bound, and waits out that uncertainty before acknowledging a commit. CockroachDB uses hybrid logical clocks and an uncertainty interval. A read that finds a value inside the interval restarts at a later timestamp. The default isolation level in CockroachDB is SERIALIZABLE, so applications must retry transactions that fail with SQLSTATE 40001.

Family C: partitioned key-value and document stores

DynamoDB, Cosmos DB and Bigtable hash or range-partition items by key and replicate each partition. They give up joins and, mostly, multi-item transactions in exchange for predictable single-digit-millisecond access at any scale. The limits come from the partition. In DynamoDB, a partition serves up to 3,000 read capacity units and 1,000 write capacity units per second and holds about 10 GB. A key that attracts more than that is a hot partition, and no amount of table capacity fixes it.

Consistency is a per-request choice. DynamoDB offers eventually consistent reads at half the cost of strongly consistent ones. Cosmos DB offers five levels: strong, bounded staleness, session, consistent prefix and eventual. Data modelling runs in reverse compared with SQL. You list the access patterns first, then design keys, often in a single table, so that each pattern is one query on one partition key.

The five axes of comparison

AxisShared storage (A)Distributed SQL (B)Partitioned KV (C)
DurabilityQuorum storage across 3 AZsConsensus per rangeReplicated partitions
Write scalingOne writer; scale up onlyHorizontal, per rangeHorizontal, per partition
ConsistencyStrong on writer; replicas lag slightlySerializable or snapshot, globalPer request, strong to eventual
Multi-region writesSecondary regions read-only (Aurora Global Database)Native; commit pays inter-region RTTMulti-active; last writer wins by default, strong mode available
Query surfaceFull PostgreSQL or MySQLMost of PostgreSQL or its own SQLKey lookups, limited queries

Read the table as physics, not marketing. A shared-storage engine commits in one intra-region quorum round, so its single-row writes are fast, but it cannot absorb more writes than one instance can generate. Distributed SQL adds a consensus round per range touched, and more for cross-range transactions. That is small inside a region and large across regions, where a majority must be reached over tens of milliseconds of network. Partitioned stores are fast at any scale, provided each request touches one key and keys are spread evenly.

Axis detail: multi-region is where the families diverge most

Light in fibre covers about 200 km per millisecond, so a round trip between US East and US West is roughly 60 to 70 ms in practice. No design avoids that. Each family only decides who pays. Aurora Global Database replicates asynchronously to secondary regions, so local commits stay fast and a regional failover can lose the last moments of writes. Spanner and CockroachDB can place replicas so a majority survives a region loss, and then every write pays a cross-region consensus round. Both let you pin a table's or row's leaders to a home region, so writes from that region are fast and writes from elsewhere are slow. DynamoDB global tables by default accept writes in every region and resolve conflicts by last writer wins, which is fine for idempotent upserts and wrong for counters and balances. A multi-Region strong consistency mode, generally available since mid-2025, instead replicates each write synchronously to another region before acknowledging it.

Decide the recovery point objective first. If losing a few seconds of writes during a regional disaster is acceptable, asynchronous replication gives you local latency everywhere. If it is not, you are buying consensus across regions, and the latency comes with it. The broader trade-offs are laid out in the multi-region architecture guide.

Portability code: retries are not optional

Two patterns account for most of the application changes when you move to a distributed database. The first is the transaction retry loop that serializable distributed SQL requires.

import random, time
import psycopg

def run_txn(conn, fn, attempts=5):
    for i in range(attempts):
        try:
            with conn.transaction():
                return fn(conn)
        except psycopg.errors.SerializationFailure:     # SQLSTATE 40001
            time.sleep(min(1.0, 0.05 * 2 ** i) * random.random())
    raise RuntimeError("transaction kept conflicting; look for a hot row")

def transfer(conn, src, dst, amount):
    def body(c):
        cur = c.execute("UPDATE accounts SET balance = balance - %s WHERE id = %s AND balance >= %s",
                        (amount, src, amount))
        if cur.rowcount != 1:
            raise ValueError("insufficient funds or unknown account")   # rolls back
        c.execute("UPDATE accounts SET balance = balance + %s WHERE id = %s", (amount, dst))
    return run_txn(conn, body)

The second is the conditional write that replaces a transaction in a key-value store. It turns read-modify-write into a compare-and-set that the partition evaluates atomically.

import boto3
from botocore.exceptions import ClientError

table = boto3.resource("dynamodb").Table("orders")

def advance(order_id, expected, new_status):
    try:
        table.update_item(
            Key={"pk": f"ORDER#{order_id}", "sk": "META"},
            UpdateExpression="SET #s = :new",
            ConditionExpression="#s = :expected",
            ExpressionAttributeNames={"#s": "status"},
            ExpressionAttributeValues={":new": new_status, ":expected": expected},
        )
        return True
    except ClientError as e:
        if e.response["Error"]["Code"] == "ConditionalCheckFailedException":
            return False          # someone else moved it first; re-read and decide
        raise

Code written without these patterns works in testing and fails under concurrency. Put the retry wrapper in the data-access layer on day one, even before you migrate.

Worked example: choosing for an order platform

An e-commerce team needs a store for orders and payments. The requirements are 4,000 writes per second at peak, growing about 50 percent a year; 40,000 reads per second, mostly by order ID or customer ID; ad-hoc reporting joins for the finance team; customers in North America and Europe; and a regional disaster may lose no committed payment.

Step one: the zero-loss requirement across regions rules out asynchronous cross-region replication for payments, so plain Aurora Global Database alone does not satisfy it. Step two: write volume. 4,000 writes per second fits one large writer today, but at 50 percent annual growth it reaches about 13,500 in three years, which is uncomfortable for a single instance. Step three: queries. Finance needs joins, so a pure key-value store would push reporting into a separate analytical copy, which is acceptable but adds a pipeline.

Two designs survive. The first is distributed SQL with regional leader pinning, European customers' rows homed in Europe and American rows in North America, plus replicas placed so each region's data survives a regional loss. Local writes pay an extra consensus round to a nearby region, around 10 to 30 ms depending on placement, and joins work. The second splits the domain. Payments go to distributed SQL for zero loss, and order browsing goes to DynamoDB global tables, whose last-writer-wins semantics are acceptable for data owned by one customer session, with change data capture feeding the warehouse. The team chose the first because it kept one consistency model, and wrote down the alternative as the fallback if cost per write proved too high.

Cost models and elasticity

Pricing follows architecture. Shared-storage engines charge for instance hours and storage, and some tiers also charge per I/O, which punishes scan-heavy workloads and rewards a large buffer pool. Distributed SQL charges for nodes or for request units, and replication factor multiplies storage. Key-value stores charge per request in on-demand mode or per provisioned capacity unit, and because an eventually consistent read costs half a strong one, the consistency choice appears on the invoice. Serverless variants of all three scale compute with load, sometimes to zero. Cold starts and minimum capacities vary by product and change often, so test them before relying on them.

Failure modes and trade-offs

  • Hot keys. A single counter row or popular item saturates one range or partition in every family. Shard the key or aggregate asynchronously.
  • Retry storms. Serializable retries without jitter or limits amplify contention into an outage.
  • Replica lag surprises. Reading from replicas after a write returns stale data. Route read-your-writes traffic to the writer or use session consistency.
  • Cross-region transactions by accident. One unpinned table forces every transaction that touches it into a cross-region commit.
  • Lock-in by feature. Engine-specific extensions, TTL semantics and change-stream formats are the parts that do not port.

What to do next

  1. Write down peak and three-year write rates, read patterns, join needs, regions and the recovery point objective for each dataset.
  2. Eliminate families with those numbers: one writer, cross-region loss tolerance, need for joins.
  3. Prototype the two survivors with your real hot paths and measure p99 commit latency from every region.
  4. Add a jittered retry wrapper and idempotency keys to the data-access layer now.
  5. Model cost at today's and the three-year load, including replication factor and consistency choice.
  6. Run a zone failure and a region failover drill before production, and record how much data and time each costs.
Key takeaway: Cloud-native databases come in three shapes: shared-storage engines with one writer and quorum storage, distributed SQL built from consensus groups, and partitioned key-value stores. Compare them on durability, write scaling, consistency, multi-region behaviour and cost, and let your write growth, join needs and cross-region recovery point objective eliminate families before you look at products. Whatever you pick, build retries and conditional writes into the data layer from the start.