Amazon DynamoDB is usually explained from the outside: tables, keys and indexes. That view stops being enough the first time a table nowhere near its capacity throws throttling errors, or a strongly consistent read costs twice what you budgeted. Each symptom has a cause inside the service, and it is easier to fix when you can picture where it happens.

This page goes inside. It follows a single request through the components AWS described in its 2022 USENIX ATC paper on DynamoDB, then turns that picture into practice: exact capacity unit arithmetic, a sizing calculator for a realistic table, write code that survives retries, and the throttling reason codes DynamoDB now attaches to every throttled request. The data model, single-table design, streams and TTL are covered in the DynamoDB architecture overview; this page assumes you know what a partition key is.

Advertisement

The request path: routers, metadata and admission

Every data-plane call, a GetItem, a PutItem or a Query, arrives at a fleet of request routers. According to the paper, the request routing service authenticates and authorizes the call and routes it to the right storage node. To do that it needs to know which partition owns the key, so it consults the metadata service, which maps tables and indexes to the replication groups holding each key range, and caches the answer. Control-plane calls such as CreateTable go to a separate autoadmin service.

The router is also where table-level admission happens. The paper describes global admission control, GAC: a service that tracks a table's capacity as tokens, while each router keeps a local token bucket refilled from GAC every few seconds. Capacity can then be spent wherever traffic lands, instead of being carved into fixed per-partition slices, which is how early DynamoDB worked and why early users saw throttling on mostly idle tables.

One PutItem, from SDK to quorumSDK clientsigns, retriesHTTPSRequest routerauthn, authz, admissionMetadata servicekey -> replication groupGACtable-wide token budgettokensroute to leaderReplication group for one partition (replicas spread across Availability Zones)Leader replicaWAL + B-treeStorage replicaWAL + B-treeStorage or log replicaWAL only if loglog recordwrites, strong readseventual reads: any storage replicaAck to the client after a quorum persists the write-ahead log recordThrottling can happen at the router (table budget) or at the partition (key-range limit)
A write passes the router, which checks identity and table-level tokens, is routed by key to the leader of one replication group, and is acknowledged once a quorum of replicas has persisted the log record.

Partitions and replication groups

A table is split into partitions by key range, and each partition is stored as a replication group of replicas in different Availability Zones. The group uses Multi-Paxos to elect a leader. The leader holds a lease that it renews periodically; only the leader serves writes and strongly consistent reads. For a write, the leader creates a write-ahead log record and sends it to its peers, and the write is acknowledged once a quorum has persisted it. A storage replica keeps both the write-ahead log and a B-tree holding the items. Rebuilding a failed storage replica takes minutes, so the leader adds a log replica, which stores only recent log entries and takes seconds to add, to keep the quorum healthy meanwhile.

This structure explains the read pricing you see in the API. An eventually consistent read can be served by any storage replica, so it may miss a write that is still propagating, and it costs half a read unit per 4 KB. A strongly consistent read must go to the leader, costs a full unit, and is unavailable for the short window while a new leader waits for the old lease to expire. A transactional read costs two units. Where a read path can tolerate that, it is cheaper and more available, and caches such as DynamoDB Accelerator build on that same trade.

AWS's guidance gives each partition about 10 GB of data and a ceiling of 3,000 read and 1,000 write units per second. You cannot see partitions, but every throttling decision at this layer comes from those limits.

Advertisement

Capacity units, exactly

One write unit is one write per second of an item up to 1 KB; one read unit is one strongly consistent read per second of an item up to 4 KB, or two eventually consistent reads. Sizes round up per item, so a 1.1 KB write costs 2 units, and a 4.1 KB strong read costs 2. Item size includes attribute names as well as values. Transactional reads and writes cost double. The same rules apply to on-demand request units; only the billing differs.

Global secondary indexes are where budgets go wrong. Every write to the base table that adds, changes or removes an item in an index is also a write to that index, charged at the size of the projected index entry. An update that changes an indexed key attribute costs two index writes, a delete of the old entry and a put of the new one. A Query or Scan returns at most 1 MB per call and is charged on the data read before any filter expression is applied, so a filter saves bandwidth, not capacity. Batch calls group requests, not cost: BatchWriteItem takes up to 25 puts or deletes and BatchGetItem up to 100 items, each charged as if sent alone. Transactions take up to 100 unique items and 4 MB.

How DynamoDB absorbs skew, and where it stops

Real traffic is never uniform, so the service has layers of defence. The paper describes bursting, which retained a partition's unused capacity for up to 300 seconds, and split for consumption: once a partition's consumed throughput crosses a threshold, it is split at a point chosen from the key distribution it has observed, which usually completes in minutes. The paper is explicit about the limits: a partition taking heavy traffic to a single item, or one whose key range is written sequentially, gains nothing from a split, so DynamoDB does not split it. The user-facing documentation calls the combined behaviour adaptive capacity; treat it as how the service tends to behave, not as a guarantee.

On-demand tables add a second dimension. AWS documents that a new on-demand table can sustain up to 4,000 writes and 12,000 reads per second and instantly accommodates up to double its previous peak; exceeding double the previous peak within 30 minutes can throttle. Before a launch or migration, set warm throughput on the table, which AWS calls pre-warming. An optional maximum throughput per table or index is a cost guard that throttles on purpose. Default quotas are 40,000 read and 40,000 write units per table, adjustable on request.

Worked example: sizing an orders table

An order service writes 2,000 orders a second at 2.5 KB each and reads 5,000 orders a second at 6 KB with eventual consistency. The table will hold about 500 GB. It has two indexes: by_status, KEYS_ONLY with roughly 0.2 KB entries, and by_customer, projecting ALL. One tenant produces 40 percent of all orders. The calculator below applies the rules from the previous sections.

import math

PARTITION_RCU, PARTITION_WCU, PARTITION_GB = 3000, 1000, 10


def wcu(item_kb, transactional=False):
    units = math.ceil(item_kb / 1.0)          # 1 WCU per 1 KB, rounded up
    return units * (2 if transactional else 1)


def rcu(item_kb, mode="eventual"):
    units = math.ceil(item_kb / 4.0)          # 1 RCU per 4 KB, rounded up
    return {"eventual": units / 2, "strong": units, "transactional": units * 2}[mode]


def plan(writes, item_kb, reads, read_kb, read_mode, table_gb, gsis, hot_share):
    base_w = writes * wcu(item_kb)
    gsi_w = sum(writes * wcu(kb) for _, kb in gsis)
    base_r = reads * rcu(read_kb, read_mode)
    parts = max(math.ceil(base_w / PARTITION_WCU),
                math.ceil(base_r / PARTITION_RCU),
                math.ceil(table_gb / PARTITION_GB))
    hot_w = writes * hot_share * wcu(item_kb)
    shards = math.ceil(hot_w / (PARTITION_WCU * 0.5))   # keep each shard under half the ceiling
    print(f"base table writes  {base_w:>8,.0f} WCU/s ({wcu(item_kb)} per {item_kb} KB item)")
    for name, kb in gsis:
        print(f"  GSI {name:<14} {writes * wcu(kb):>8,.0f} WCU/s")
    print(f"total write units  {base_w + gsi_w:>8,.0f} WCU/s")
    print(f"reads ({read_mode:<8})  {base_r:>8,.0f} RCU/s")
    print(f"partitions needed  >= {parts} (throughput and size)")
    print(f"hot tenant         {hot_w:>8,.0f} WCU/s on one key -> {shards} write shards")


plan(writes=2000, item_kb=2.5, reads=5000, read_kb=6, read_mode="eventual",
     table_gb=500, gsis=[("by_status", 0.2), ("by_customer", 2.5)], hot_share=0.4)

Running it prints:

base table writes     6,000 WCU/s (3 per 2.5 KB item)
  GSI by_status         2,000 WCU/s
  GSI by_customer       6,000 WCU/s
total write units    14,000 WCU/s
reads (eventual)     5,000 RCU/s
partitions needed  >= 50 (throughput and size)
hot tenant            2,400 WCU/s on one key -> 5 write shards

Three findings matter. First, the indexes more than double the write bill, from 6,000 units to 14,000, and by_customer alone costs as much as the table; projecting only the attributes the customer screen needs cuts that. Second, size sets the partition floor here: 500 GB implies at least 50 partitions, useful context but not something you configure. Third, the hot tenant writes 2,400 units a second to one partition key, well past the 1,000-unit ceiling of any single partition. A split takes minutes and does not help sequential sort keys such as timestamps, so the reliable fix is write sharding: spread that tenant over five keys, each carrying under half a partition's limit, and fan reads out across the five.

Writing to it safely

Make writes idempotent, because the SDK retries failed calls, and ask for consumed capacity, because your arithmetic will drift from the bill. The sketch uses boto3's adaptive retry mode, which adds client-side rate limiting to exponential backoff, a conditional put, and the sharded key from the sizing step.

import random
import boto3
from botocore.config import Config
from botocore.exceptions import ClientError

ddb = boto3.client("dynamodb", config=Config(retries={"max_attempts": 10, "mode": "adaptive"}))
HOT_TENANTS = {"t-0042": 5}            # tenant -> shard count, from the sizing step


def tenant_pk(tenant):
    shards = HOT_TENANTS.get(tenant, 1)
    return tenant if shards == 1 else f"{tenant}#{random.randrange(shards)}"


def put_order(tenant, order_id, body):
    try:
        resp = ddb.put_item(
            TableName="orders",
            Item={"pk": {"S": tenant_pk(tenant)}, "sk": {"S": f"ORDER#{order_id}"}, **body},
            ConditionExpression="attribute_not_exists(sk)",   # idempotent create
            ReturnConsumedCapacity="INDEXES",
        )
        return resp["ConsumedCapacity"]            # log it: table and per-GSI units
    except ClientError as e:
        code = e.response["Error"]["Code"]
        if code == "ConditionalCheckFailedException":
            return None                            # already written by an earlier retry
        reasons = e.response.get("ThrottlingReasons", [])
        print("dynamodb error", code, reasons)     # e.g. TableWriteKeyRangeThroughputExceeded
        raise

The handler reads the throttling reasons with a fallback, because what your SDK version exposes may differ, and logs the error code with them. It does not retry a ConditionalCheckFailedException, which is a correct answer rather than a fault. Stream consumers, such as AWS Lambda functions reading DynamoDB Streams, need the same idempotency, since they can see a record more than once.

Reading throttling: the reason codes

Throttled requests now carry a ThrottlingReasons list. Each reason is a resource type, Table or Index, an operation, Read or Write, and a limit type joined into one string, such as TableWriteKeyRangeThroughputExceeded, with the ARN of the table or index. The limit type tells you which layer refused the request.

Limit typeMeaningUsual fix
KeyRangeThroughputExceededOne partition exceeded its internal limit: a hot key, sequential writes, or traffic above the table's warm throughputPre-warm for the event; shard hot keys; parallelise scans
ProvisionedThroughputExceededA provisioned table or GSI used more than it was givenRaise capacity or auto-scaling limits; check GSI capacity separately
AccountLimitExceededAn on-demand table and its indexes hit the account-level table quotaRequest a quota increase
MaxOnDemandThroughputExceededTraffic exceeded the maximum you configured on an on-demand table or indexRaise or remove your own cap

Each reason maps to a CloudWatch metric of the same shape, for example WriteKeyRangeThroughputThrottleEvents, while ReadThrottleEvents, WriteThrottleEvents and ThrottledRequests aggregate across the table and its indexes. For key-range throttling, enable CloudWatch Contributor Insights in its throttled-keys mode to see which keys are responsible. If no key stands out, suspect sequential access: rate-limit scans under 3,000 read units per key range and use parallel segmented scans, as AWS advises. The dashboards and alarms around these metrics are covered in the CloudWatch deep dive.

Failure modes

  • Throttling with idle capacity: a hot key or sequential key range hits a partition's 1,000 write or 3,000 read limit while the table as a whole is barely used.
  • GSI back-pressure: an under-provisioned global secondary index throttles writes to the base table, because the base write cannot complete until the index can take it.
  • Launch-day cliff: an on-demand table that has never seen more than 5,000 writes a second is asked for 50,000 within minutes.
  • Duplicate writes on retry: unconditional puts or counter increments replayed by the SDK after a timeout whose original request succeeded.
  • Cross-region surprises: transactions are ACID only in the region where they ran; replicas of global tables can observe them partially applied.

Trade-offs

DynamoDB buys predictable latency at any scale by refusing work it cannot do quickly: no joins, no ad hoc queries, a 400 KB item limit, and a per-partition ceiling that only key design removes. On-demand trades a higher unit price for no capacity planning; provisioned capacity with auto-scaling is cheaper for steady traffic but reacts in minutes. Strong consistency costs double, and every index is a second table you pay to write. It is excellent when access patterns are known in advance and poor when they are not.

What to do next

  1. List every access pattern and confirm each is served by a Get, a Query on the table, or a Query on an index, with no Scan in the request path.
  2. Compute write units for the table and every GSI from real item and index entry sizes, and trim projections that double the bill.
  3. Estimate each partition key's share of traffic and shard any key that could exceed about half of 1,000 write or 3,000 read units per second.
  4. Make every write idempotent with condition expressions, and log ConsumedCapacity on a sample of requests.
  5. Log throttling reasons with their ARNs, alarm on the KeyRange and Provisioned throttle metrics, and enable Contributor Insights on hot tables.
  6. Before a launch or backfill, set warm throughput or ramp traffic over more than 30 minutes, and test at the target rate.
Key takeaway: A DynamoDB request is admitted by a router against a table-wide token budget, then served by the leader or a replica of one partition's Paxos replication group. Throttling comes from one of those two places, and the reason code tells you which. Sizes round up per item, strong reads and transactions cost double, and every GSI write is billed again. No partition exceeds about 1,000 write or 3,000 read units per second, so shard hot keys, make writes idempotent, and pre-warm before traffic jumps.