Amazon Keyspaces is a serverless service that speaks the Cassandra Query Language over the Cassandra wire protocol. Your existing drivers and most of your CQL work, but there are no nodes to size, no repair to schedule, no compaction to tune and no JVM to garbage-collect. You pay per request or per provisioned unit of throughput, and per gigabyte stored.
That makes it attractive for teams that chose Cassandra for its data model and dislike operating it, and confusing for teams who expect it to behave exactly like the cluster they run today. This page explains what Keyspaces is underneath, how connections and capacity actually work, what changes in data modelling and CQL, and how to diagnose the errors that show up after a migration. If you are still deciding between data stores, read when to pick Cassandra first. All limits quoted here come from the Keyspaces developer guide as of October 2026.
What Keyspaces is underneath
Keyspaces is compatible with the CQL 3.11 API. It does not expose Cassandra nodes, JVMs or storage files to you in any form, and AWS documents it as serverless with no clusters or hosts to configure. That is why settings for compaction, compression, caching, bloom filters and garbage collection are accepted and ignored, and why concepts such as gc_grace_seconds and repair are listed as not applicable.
Data is stored three times across multiple Availability Zones in a Region, and writes are durably stored at LOCAL_QUORUM before they are acknowledged, so send writes at that level. Reads may use ONE, LOCAL_ONE or LOCAL_QUORUM. QUORUM, ALL, EACH_QUORUM, ANY, TWO, THREE and the serial levels are unsupported and raise exceptions, which is a common first error for tools whose default is QUORUM.
The driver connects to a regional endpoint such as cassandra.us-east-2.amazonaws.com on port 9142, always over TLS. The system.peers table it reads describes nine endpoints, which are load balancers, not storage nodes. Token-aware routing therefore buys nothing; a round-robin policy is the recommended choice.
Connecting: SigV4, TLS and connection math
Authentication uses IAM. You can generate service-specific credentials, a user name and password tied to an IAM user, but the better option is the SigV4 authentication plugin, which signs requests with normal IAM credentials, including temporary credentials from a role. For Python the plugin is the cassandra-sigv4 package:
# pip install cassandra-driver cassandra-sigv4 boto3
from ssl import SSLContext, PROTOCOL_TLSv1_2, CERT_REQUIRED
import boto3
from cassandra.cluster import Cluster, ExecutionProfile, EXEC_PROFILE_DEFAULT
from cassandra.policies import RoundRobinPolicy
from cassandra import ConsistencyLevel
from cassandra_sigv4.auth import SigV4AuthProvider
ssl_context = SSLContext(PROTOCOL_TLSv1_2)
ssl_context.load_verify_locations("keyspaces-bundle.pem") # Amazon root CAs, per the guide
ssl_context.verify_mode = CERT_REQUIRED
auth = SigV4AuthProvider(boto3.Session(region_name="us-east-2"))
profile = ExecutionProfile(load_balancing_policy=RoundRobinPolicy(),
consistency_level=ConsistencyLevel.LOCAL_QUORUM)
cluster = Cluster(["cassandra.us-east-2.amazonaws.com"], port=9142,
ssl_context=ssl_context, auth_provider=auth,
execution_profiles={EXEC_PROFILE_DEFAULT: profile})
session = cluster.connect()The certificate bundle is the set of Amazon root CAs the developer guide tells you to download and concatenate. AWS has been moving Keyspaces from the older Starfield root to Amazon Trust Services roots, so include both until the guide says otherwise.
Throughput per connection is the first limit most migrations hit. Each TCP connection can carry at most 3,000 CQL requests per second. With nine peers and the common driver default of one connection per peer, a client tops out at 27,000 requests per second, and in practice lower, because load is never perfectly even. AWS recommends planning on 500 requests per second per connection. Worked through for a service that needs 18,000 requests per second: 18,000 divided by 500 gives 36 connections, and 36 divided by nine peers gives four connections per peer. With VPC endpoints the driver may see fewer peers, so the per-peer count must rise to keep the total. Exceeding the per-connection rate shows up in CloudWatch as PerConnectionRequestRateExceeded.
Capacity: request units and a worked example
Capacity is measured in request units. A write consumes one unit per 1 KB of row written. A read consumes one unit per 4 KB at LOCAL_QUORUM and half a unit per 4 KB at ONE or LOCAL_ONE. Tables run in one of two modes: on-demand, which bills per read and write request unit with no planning, and provisioned, where you set read and write capacity units per second and can attach Application Auto Scaling.
Work an example. An orders table stores rows of about 2.5 KB. The service writes 400 rows per second and reads 3,000 rows per second by primary key, half at LOCAL_QUORUM and half at LOCAL_ONE.
- Writes: 2.5 KB rounds up to 3 units, so 400 times 3 is 1,200 write units per second.
- Quorum reads: 2.5 KB fits in one 4 KB unit, so 1,500 reads cost 1,500 read units.
- Single-replica reads: 1,500 reads at half a unit each cost 750.
- Total: 1,200 write and 2,250 read units per second at steady state, before headroom.
Two lessons fall out. Row size matters more for writes than for reads, because the write unit is a quarter the size of the read unit; trimming that 2.5 KB row below 2 KB would cut write cost by a third. And LOCAL_ONE halves read cost where slightly stale data is acceptable.
Provisioned capacity can be raised at any time. Decreases are rationed: four at any time in a UTC day, plus one more for each hour with no decrease, for a maximum of 27 a day. Default per-table limits are 40,000 read and 40,000 write units per second, with account-level provisioned limits of 80,000 each, all adjustable through Service Quotas.
Storage partitions and hot keys
Underneath every table are storage partitions, and each storage partition supports up to 3,000 read units and 1,000 write units per second. A logical Cassandra partition can span several storage partitions, so a large partition is not itself a problem, but traffic concentrated on a narrow key range is. When it exceeds the partition's limit, requests are throttled, and CloudWatch reports StoragePartitionThroughputCapacityExceeded.
In the orders example, partitioning by customer_id is fine for normal customers. A wholesale customer placing 600 orders a second, each 2.5 KB, needs 1,800 write units on one hot key, nearly double the partition ceiling. The usual fix is write sharding: add a bucket to the partition key and spread writes across a known number of buckets.
CREATE TABLE shop.orders_by_customer (
customer_id text,
bucket smallint, -- 0..7, chosen as hash(order_id) % 8
order_ts timestamp,
order_id uuid,
total_cents bigint,
PRIMARY KEY ((customer_id, bucket), order_ts, order_id)
) WITH CLUSTERING ORDER BY (order_ts DESC);
-- read the latest orders: query the 8 buckets in parallel, merge client-side
SELECT order_id, total_cents FROM shop.orders_by_customer
WHERE customer_id = 'c-42' AND bucket = 3 LIMIT 50;Size limits shape keys as well. A row may be at most 1 MB excluding static data, static data per logical partition at most 1 MB, a compound partition key at most 2,048 bytes and all clustering columns together at most 850 bytes. Those are smaller than open-source Cassandra will tolerate, so check wide-row and large-blob designs before migrating.
CQL differences that matter
Most day-to-day CQL works unchanged. The differences that break real applications are these:
| Area | Keyspaces behaviour |
|---|---|
| Secondary indexes, materialized views, triggers | CREATE INDEX, CREATE MATERIALIZED VIEW and CREATE TRIGGER are not supported; denormalise into query tables |
| User code in the database | CREATE FUNCTION and CREATE AGGREGATE are not supported |
TRUNCATE | Not supported; drop and recreate the table, or delete by partition |
| Batches | Logged: up to 100 statements and 4 MB; unlogged: up to 30; only INSERT, UPDATE and DELETE |
IN | SELECT only, up to 100 values; each value is a subquery counted against the per-connection rate; results follow the order of the IN list |
| Paging | A page ends after 1 MB read, so filtered queries can return pages with fewer rows than the page size |
| Range delete | Up to 1,000 rows per statement, not isolated, and processed asynchronously since 1 July 2026: success means accepted, not finished |
| Timestamps | USING TIMESTAMP and WRITETIME need client-side timestamps enabled on the table |
| DDL | Asynchronous; a table is usable when system_schema_mcs.tables shows status ACTIVE |
Lightweight transactions are fully supported, and AWS states there is no performance penalty for them as there is with Paxos in open-source Cassandra; the conditional semantics are described in Cassandra lightweight transactions. Failed conditions still consume write capacity based on row size and are counted in ConditionalCheckFailedRequests.
The asynchronous DDL point catches every migration script that creates a table and inserts into it straight away. Poll before writing:
import time
def wait_active(session, ks, table, timeout=300):
q = ("SELECT status FROM system_schema_mcs.tables "
"WHERE keyspace_name = %s AND table_name = %s")
deadline = time.time() + timeout
while time.time() < deadline:
row = session.execute(q, (ks, table)).one()
if row and row.status == "ACTIVE":
return
time.sleep(2)
raise TimeoutError(f"{ks}.{table} not ACTIVE after {timeout}s")
Operating it: metrics, retries, deletes and recovery
Operations shift from nodes to metrics. Watch ReadThrottleEvents and WriteThrottleEvents, which count requests rejected for any capacity reason, and then split by cause: PerConnectionRequestRateExceeded means add connections, StoragePartitionThroughputCapacityExceeded means fix a hot key, and throttling with neither means table or account capacity. SystemErrors count server errors, and UserErrors count invalid requests, often a bad signature or a missing table.
Retries need care. Throttled requests should be retried with exponential backoff and jitter, but only idempotent statements should be retried blindly. A counter increment or a list append retried after a timeout can apply twice, exactly as in open-source Cassandra.
Tombstones and their read cost are an operator's headache in self-managed Cassandra, as the tombstones article explains. Keyspaces removes the gc_grace and compaction tuning, but deletes and TTL expiry are not free: TTL deletions are billed and reported as TTLDeletes, and deletions under client-side timestamps as SystemReconciliationDeletes.
For data protection, point-in-time recovery restores a table to a chosen moment into a new table; default quotas allow four concurrent restores and 5 TB restored per 24 hours. For change streams, Keyspaces CDC streams provide ordered, de-duplicated change records per table, with the row before and after the change by default.
Trade-offs
Keyspaces is a strong fit when the data model is already Cassandra-shaped, traffic is spiky or modest, and the team does not want to run a cluster. It is a weaker fit when the application depends on secondary indexes, materialized views, user-defined functions or TRUNCATE, when partitions are very hot and cannot be sharded, or when steady very high throughput makes per-request pricing more expensive than well-utilised instances.
If the choice is really between Keyspaces and DynamoDB, the capacity model will feel familiar, and DynamoDB internals is useful background. The deciding factors are usually the query language and existing code: CQL applications move to Keyspaces with driver configuration changes, while DynamoDB requires a rewrite of the data access layer.
What to do next
A practical path from a self-managed cluster or a new project:
- Inventory your CQL for unsupported features: indexes, materialized views, triggers, UDFs, UDAs, TRUNCATE,
INin UPDATE or DELETE. - Check row, partition key and clustering key sizes against the 1 MB, 2,048-byte and 850-byte limits.
- Estimate peak requests per second and set connections per peer at 500 requests per second per connection.
- Compute write and read units from real row sizes; decide per query whether
LOCAL_ONEis acceptable. - Find the hottest partition keys and shard any that could exceed 1,000 write units per second.
- Add a wait-for-ACTIVE step to every schema migration.
- Alarm on throttle events,
PerConnectionRequestRateExceededandStoragePartitionThroughputCapacityExceeded. - Enable point-in-time recovery and rehearse a restore.