DynamoDB is rarely a bad database. It is often the wrong one for the job. Teams that pick it for the right reasons get single-digit-millisecond key lookups at almost any scale with no servers to patch. Teams that pick it because it is the AWS default end up with Scans, a second database for reporting, and a key design nobody dares change.
This article is about the decision, not the mechanics. If you need the details of partitions, capacity modes, secondary indexes and Streams, read the DynamoDB deep dive first. Here we cover the one property that decides fit, a method for testing your workload against it, a worked design and sizing, the workloads where DynamoDB strains, the alternatives, and the escape hatches you should plan before day one.
What DynamoDB is, architecturally
DynamoDB is a managed, partitioned key-value and document store. Every item has a partition key, and optionally a sort key. The partition key is hashed to choose a storage partition. Items that share a partition key are stored together, sorted by the sort key. The fast operations follow directly from that layout: get one item by its full key, or Query a contiguous range of sort keys within one partition key. Everything else is either an index you maintained in advance or a Scan that reads the whole table.
That is the core trade. A relational database lets you decide your queries later and pays for that flexibility at query time, with a planner, joins and shared resources. DynamoDB makes you decide your queries up front and pays for them at design time, and in return every request costs about the same no matter how big the table gets. The hard limits are part of the design surface: an item can be at most 400 KB, a Query or Scan returns at most 1 MB per page, and AWS documents a per-partition ceiling of 3,000 read capacity units and 1,000 write capacity units per second. A transaction can group up to 100 actions, up to 4 MB in total.
Billing follows the same model. A write capacity unit covers one write per second of up to 1 KB. A read capacity unit covers one strongly consistent read per second of up to 4 KB, or two eventually consistent reads. Transactional reads and writes cost twice as much. You pay for these either per request (on-demand) or as provisioned throughput with auto scaling. Either way, cost is a function of request count and item size, not of CPU or table size, apart from storage.
The access-pattern inventory
The test for fit is an exercise, not an opinion. Write down every way the application will read and write data, before any schema exists. For each access pattern, record the inputs you have at request time, the shape of the result, the expected rate at peak, the size of each item, and the consistency it needs. Then try to answer each pattern with a single GetItem or one Query against the table or an index.
If every pattern maps to a key lookup or a key-range Query, DynamoDB fits and will stay fast as you grow. If one or two patterns need filtering on non-key attributes, you can usually add a global secondary index (GSI) with a key designed for that pattern. If patterns need arbitrary combinations of filters, sorting by user-chosen columns, joins across entities, or aggregates, you are describing a relational database or a search engine, and forcing it into DynamoDB means Scans or a pile of indexes.
Be honest about how often access patterns change. A payments ledger keeps the same queries for years. An early-stage product changes them every sprint. Adding a GSI is easy. Changing the base table's key is a migration: you copy every item into a new table while live traffic continues.
Worked example: a multi-tenant order system
A B2B SaaS stores orders for many tenants. The inventory produced five patterns: AP1, list a tenant's orders by date, newest first; AP2, fetch one order by id; AP3, fetch an order with its line items; AP4, list a customer's recent orders; AP5, list a tenant's open orders. All have the key they need at request time. None needs a join that cannot be pre-built, and reporting happens elsewhere. That is a fit.
One table design serves all five. The META item is the only copy of an order's header, so a status change is one write and every index follows it; the three indexes are projected from it:
| Where | Partition key | Sort key | Serves |
|---|---|---|---|
| Table, META item | ORDER#o981 | META | AP2: get an order by id (status, total, version live only here) |
| Table, LINE items | ORDER#o981 | LINE#001 | AP3: an order with its lines, one Query |
| GSI1 (from META) | TENANT#t42 | 2026-10-01T09:12:00Z#o981 | AP1: a tenant's orders by date |
| GSI2 (from META) | CUST#c7 | 2026-10-01T09:12:00Z#o981 | AP4: a customer's recent orders |
| GSI3 (sparse, from META) | TENANT#t42#OPEN | 2026-10-01T09:12:00Z#o981 | AP5: open orders per tenant |
AP1 is a Query on GSI1 with a date prefix on the sort key. AP2 is a GetItem. AP3 is a Query on the order partition that returns the META item and its LINE items together, which is the single-table trick that replaces a join. AP4 uses GSI2, keyed on customer. AP5 uses a sparse GSI3: only open orders carry its key attributes, so the index holds only what the query needs, and closing an order removes the attribute and the index entry with it. Keeping one header copy costs you GSI reads that are eventually consistent; duplicating the header into a tenant partition would make AP1 strongly consistent but then every status change must update both copies, ideally in one transaction at double the write cost. The code for the main patterns is short:
import boto3
from boto3.dynamodb.conditions import Key
table = boto3.resource("dynamodb").Table("orders")
# AP1: a tenant's orders for one day, newest first, one page at a time.
resp = table.query(
IndexName="GSI1",
KeyConditionExpression=Key("gsi1pk").eq("TENANT#t42")
& Key("gsi1sk").begins_with("2026-10-01"),
ScanIndexForward=False,
Limit=50,
)
page, cursor = resp["Items"], resp.get("LastEvaluatedKey") # pass back as ExclusiveStartKey
# AP2: create an order only if it does not exist (idempotent on retries).
table.put_item(
Item={"pk": "ORDER#o981", "sk": "META", "status": "OPEN", "total": 4200, "version": 1,
"gsi1pk": "TENANT#t42", "gsi1sk": "2026-10-01T09:12:00Z#o981",
"gsi2pk": "CUST#c7", "gsi2sk": "2026-10-01T09:12:00Z#o981",
"gsi3pk": "TENANT#t42#OPEN", "gsi3sk": "2026-10-01T09:12:00Z#o981"},
ConditionExpression="attribute_not_exists(pk)",
)
# Optimistic concurrency: change status only if nobody else changed it first.
# Removing gsi3pk drops the order out of the sparse open-orders index.
table.update_item(
Key={"pk": "ORDER#o981", "sk": "META"},
UpdateExpression="SET #s = :new, version = version + :one REMOVE gsi3pk",
ConditionExpression="version = :v",
ExpressionAttributeNames={"#s": "status"},
ExpressionAttributeValues={":new": "PAID", ":v": 1, ":one": 1},
)Sizing in request units. Peak is 2,000 new orders per second, each a 2 KB META item plus three LINE items under 1 KB, plus 10,000 reads per second at about 3 KB, eventually consistent. Per order that is 2 + 3 = 5 WCU on the table, so 10,000 WCU, or 20,000 if you write the header and lines atomically with TransactWriteItems. Each META write also writes to the three GSIs; with small projections under 1 KB that is about 1 WCU per index, another 6,000 WCU of index capacity. The 3 KB reads round up to one 4 KB unit, and at half a unit each when eventually consistent that is 5,000 RCU. Multiply by the current price in your region to compare with the alternatives. Do not compare against a server's hourly price alone: the alternative also needs replicas, backups, patching and people.
Hot keys. The table itself spreads well, because every order has its own partition key. The hot spot moves to GSI1: one tenant's partition there takes 1 WCU per order, so a single tenant tops out near 1,000 orders per second, and a throttled GSI pushes back on base-table writes. If one large tenant can exceed that, shard the index key, for example TENANT#t42#3 with a random suffix from 0 to 9, and have AP1 query all shards in parallel. This decision belongs in the design, not in an incident review.
Where DynamoDB fits well
| Workload | Why it fits |
|---|---|
| User profiles, sessions, carts, preferences | Keyed by user id; small items; huge fan-in; TTL clears old sessions |
| Order and entitlement records with known queries | Key-range reads per owner; conditional writes for idempotency |
| Event-driven backends on Lambda | No connection pools to exhaust; scales to zero with on-demand |
| Metadata for objects in S3 | Small pointers to large blobs; lookups by id or prefix |
| Spiky or unpredictable traffic | On-demand mode absorbs bursts without capacity planning |
| Active-active multi-Region data | Global tables replicate across Regions; multi-Region strong consistency became generally available in June 2025 |
For the replication trade-offs, see DynamoDB global tables.
Where DynamoDB strains
- Unknown or fast-changing queries. Every new filter means a new index or a Scan. Early products pay this repeatedly.
- Ad hoc analytics and reporting. Aggregates over many items are Scans, which are expensive and slow. Export the data to an analytical store instead.
- Relational integrity across many entities. There are no foreign keys and no joins. Transactions help, but they are capped at 100 actions and cost double.
- Large items. The 400 KB item limit pushes documents, images and long histories out to S3, so you manage two stores.
- Full-text or fuzzy search. DynamoDB has no text index. You need OpenSearch or similar, fed from Streams.
- Portability. The API and the cost model are AWS-specific. Leaving means rewriting the data layer, not just running a dump and restore.
Alternatives side by side
| Option | Choose it when | Cost of choosing it |
|---|---|---|
| Postgres (RDS, Aurora, self-hosted) | Queries are unknown, relational or analytical; strong integrity needed | You manage scaling, vacuum and connection pools; vertical limits |
| Cassandra, ScyllaDB, Amazon Keyspaces | Wide-row, write-heavy, multi-cloud or self-hosted needs | Similar modelling discipline; you run the cluster unless managed |
| MongoDB / DocumentDB | Flexible documents with secondary queries that change | Ad hoc queries tempt you into designs that do not scale |
| Redis / ElastiCache | Sub-millisecond reads of hot, derived data | Memory cost; durability is not the point |
The When to Pick Postgres article is the natural counterpart. Many systems use both: DynamoDB for the high-rate keyed path, Postgres for the relational and reporting path, with Streams connecting them. The partitioning ideas that make DynamoDB work are covered more generally in database partitioning.
Failure modes in production
- Hot partitions. One key takes more than its partition's share and gets throttled, even though the table has spare capacity. Adaptive capacity helps with uneven load but cannot exceed the per-partition ceiling. Shard hot keys in the design.
- GSI back-pressure. If a GSI cannot keep up with writes, DynamoDB throttles writes to the base table. Give GSIs enough capacity and keep projections small.
- Scan creep. A one-off admin Scan becomes a nightly job and then a user-facing feature. Alarm on Scan usage in production code paths.
- Eventual consistency on GSIs. GSIs do not support strongly consistent reads. Read-your-write flows must read the base table.
- Item growth. Lists appended inside one item creep towards 400 KB and the write cost grows with them. Model growing collections as separate items under one partition key.
- TTL is not a timer. Expired items are deleted in the background, not at the expiry moment, so filter expired items on read.
- Transaction conflicts. Concurrent transactions on the same items get cancelled. Retry with jitter, and keep transactions small.
What saying yes costs, and the escape hatches
Choosing DynamoDB commits you to a second path for analytics and search. Plan it now: enable Streams and feed a search index or a change log, or use the export to S3 feature for periodic analytics in a query engine. AWS also offers managed integrations into some analytics services; check the current list. Turn on point-in-time recovery from day one. Write your access-pattern inventory into the repository next to the table definition, so the next engineer adds an index instead of a Scan.
What to do next
- Write the access-pattern inventory: inputs, result shape, peak rate, item size and consistency for every read and write.
- Map each pattern to a GetItem or Query on the table or a named GSI. Any pattern that needs a Scan means rethinking the design or the choice.
- Size peak load in WCU and RCU, including GSI writes and transactional doubling, and price it against a managed relational alternative with its operating cost.
- Check every partition key's peak rate against 1,000 WCU and 3,000 RCU and shard any key that can exceed them.
- Plan the analytics and search path, through Streams or export, before launch.
- Load-test with production-shaped keys, not random ones, and alarm on throttling, Scan usage and item size.