Tablestore (historically OTS) is Alibaba Cloud's serverless NoSQL store. Its main model is a wide column table in the Bigtable lineage: rows sorted by a composite primary key, any number of attribute columns per row, cell versions, and a table that splits itself into partitions by primary-key range. It also offers a TimeSeries model and a message (Timeline) model, but almost every design decision you make lives in the wide column model, so that is what this page covers.
The goal is that you leave able to design a table, write correct client code, and run it in production. Limits and SDK names below were checked against Alibaba Cloud's Tablestore documentation on 2026-10-04; prices change by region and are deliberately left out.
The wide column model
An instance is the unit you create in a region; it has its own endpoint and holds tables. A table is declared with a primary key of one to four columns, each typed String, Integer or Binary. Attribute columns are not declared at all: every row can carry a different set, typed String, Integer, Double, Boolean or Binary. Rows are stored sorted by the full primary key, compared column by column, which is what makes range reads cheap.
The first primary key column is the partition key. Tablestore splits a table into partitions, each owning a contiguous range of partition-key values, and spreads partitions across its servers; it can split a partition as data or traffic grows. All rows with the same partition key live in one partition, so a single partition-key value is a throughput ceiling you cannot configure away. The remaining primary key columns order rows inside that key.
Each attribute cell can keep several versions, each stamped with a timestamp. Two table settings govern them: max versions (how many to keep) and time to live (how long data survives before it is expired in the background). Reads must say which versions they want: a single-row read that sets neither max versions nor a time range is rejected.
| Limit (wide column model) | Value |
|---|---|
| Primary key columns | 1 to 4 |
| String or Binary primary key value | up to 1 KB |
| String or Binary attribute value | up to 2 MB |
| PutRow or UpdateRow request | up to 4 MB |
| Attribute columns written per request | 1,024 |
| BatchWriteRow | 200 rows and 4 MB |
| BatchGetRow | 100 rows |
| GetRange per call | 5,000 rows or 4 MB, whichever comes first |
How a request finds its partition
Read the figure from left to right. Your code talks to one endpoint per instance; the service decides which partition serves each row from the first primary key column alone. Point operations (GetRow, PutRow, UpdateRow, DeleteRow) touch exactly one partition. A range read walks partitions in key order. Everything below the table is derived data: indexes and change streams that trail the base table by some delay, except local secondary indexes, which are kept in step.
Designing the primary key
Primary key design is where Tablestore tables succeed or fail, because it fixes both how you can query and how load spreads. Work from the access patterns. Suppose a fleet of 200,000 devices each report a reading every ten seconds, and you need two queries: the latest readings for one device, and all readings for one device between two times.
The naive key is (device_id, ts). It answers both queries with one range read, and device IDs are high-cardinality, so load spreads. The trap appears when IDs are sequential, such as dev-000001 upwards: devices provisioned together sit next to each other in key order, land in the same partition, and a batch of new devices becomes one hot partition. The fix is to prefix the partition key with a short hash of the ID:
pk = md5(device_id).hex()[:4] + "|" + device_id # e.g. "9f3a|dev-000123"
ts = epoch_millis # second PK column, Integer
# Query one device: GetRange from (pk, t0) to (pk, t1)
# Latest N readings: GetRange from (pk, INF_MAX) backwards, limit NThe hash prefix destroys ordering across devices, which you did not need, and keeps ordering within a device, which you did.
Tablestore can also generate a value for you: a primary key column declared auto-increment receives a server-assigned, increasing integer when you write PrimaryKeyValue.AUTO_INCREMENT and ask for the key back with ReturnType.RT_PK. It is useful for message sequences inside a conversation. It must be an Integer column other than the partition key, and its values are unique and strictly increasing within a partition key but not consecutive, so do not use them as invoice numbers.
Reads, writes and conditions
The core API is small. Single-row operations are atomic per row. Batch operations are a convenience for round trips: by default each row in a BatchWriteRow succeeds or fails on its own, and the response tells you which.
| Operation | What it does | Use it for |
|---|---|---|
| PutRow | Replaces the whole row (or creates it) | Writing a complete record |
| UpdateRow | Puts or deletes individual columns; can increment | Partial updates, counters |
| DeleteRow | Removes a row | Explicit deletes; prefer TTL for expiry |
| GetRow | Reads one row by full primary key | Point lookups |
| BatchGetRow / BatchWriteRow | Many rows, per-row results | Fan-in reads, bulk loads |
| GetRange | Reads rows in key order between two keys | Time windows, prefixes |
Every write can carry a condition. A row existence expectation is one of IGNORE, EXPECT_EXIST or EXPECT_NOT_EXIST; a column condition can additionally compare attribute values. If the condition fails, nothing is written and the call returns an error you can recognise. Two patterns fall out. Idempotent create: write with EXPECT_NOT_EXIST and a request ID column; if a retry fails the condition, read the row and treat it as success only if it carries your request ID, because another writer may have created it. Optimistic locking: store a version column, read it, and update with a column condition that the version still equals what you read, writing version plus one.
Client code in Java
The Java SDK shows the shapes. The client is thread-safe and should be created once per process and shut down on exit.
import com.alicloud.openservices.tablestore.SyncClient;
import com.alicloud.openservices.tablestore.model.*;
SyncClient client = new SyncClient(endpoint, accessKeyId, accessKeySecret, instanceName);
// shardedId(id) adds the 4-character hash prefix; key(pk, ts) builds the two-column PrimaryKey.
// Idempotent create of one reading.
PrimaryKeyBuilder pkb = PrimaryKeyBuilder.createPrimaryKeyBuilder();
pkb.addPrimaryKeyColumn("pk", PrimaryKeyValue.fromString(shardedId("dev-000123")));
pkb.addPrimaryKeyColumn("ts", PrimaryKeyValue.fromLong(1759546800000L));
PrimaryKey pk = pkb.build();
RowPutChange put = new RowPutChange("telemetry", pk);
put.addColumn("temp_c", ColumnValue.fromDouble(21.4));
put.addColumn("battery", ColumnValue.fromLong(87));
put.setCondition(new Condition(RowExistenceExpectation.EXPECT_NOT_EXIST));
client.putRow(new PutRowRequest(put));
// One device, one hour, paginated until the server says the range is done.
RangeRowQueryCriteria q = new RangeRowQueryCriteria("telemetry");
q.setInclusiveStartPrimaryKey(key(shardedId("dev-000123"), t0));
q.setExclusiveEndPrimaryKey(key(shardedId("dev-000123"), t1));
q.setMaxVersions(1);
while (true) {
GetRangeResponse r = client.getRange(new GetRangeRequest(q));
for (Row row : r.getRows()) handle(row);
PrimaryKey next = r.getNextStartPrimaryKey();
if (next == null) break; // the only correct exit
q.setInclusiveStartPrimaryKey(next);
}Three details matter. The loop exits on a null continuation key, never on a short page: a call stops at 5,000 rows or 4 MB of data, and a filter is applied to rows the server has already read, so a page can come back with fewer rows than you hoped while more remain. Setting setMaxVersions(1) returns only the latest version of each cell. And a full-table scan, bounded by PrimaryKeyValue.INF_MIN and PrimaryKeyValue.INF_MAX, belongs in a batch pipeline.
For bulk loads, BatchWriteRow is the right tool, with one discipline: after the call, check isAllSucceed(); if it is false, walk getFailedRows(), whose entries carry the index of each failed row in the request and its error, and build a new request from only those rows. Resending the whole batch is wrong whenever rows are not idempotent.
Secondary indexes, search indexes and change data
A table answers queries on its primary key. Everything else needs an index, and Tablestore offers two different families that are easy to confuse.
Secondary indexes are index tables: you choose index primary key columns from the base table's primary key columns and predefined (declared) attribute columns, and Tablestore maintains a copy of the rows sorted that way, with the base primary key appended so every index row is unique. A global secondary index can start with any column and is synchronised asynchronously, so it is eventually consistent; the docs describe the lag as typically milliseconds, but your code must tolerate more. A local secondary index must keep the base table's first primary key column, so it lives with the same partition, and it is written synchronously, so a read after a successful write sees it. Tables with secondary indexes must keep max versions at 1.
Search indexes are inverted and columnar structures built asynchronously from the base table. They answer what a sorted index cannot: arbitrary combinations of conditions, full-text match, geo queries, sorting by non-key columns, and aggregations. They cost more to maintain and are eventually consistent.
| You need | Choose |
|---|---|
| Lookup by one alternate key, read-your-writes, within a partition key | Local secondary index |
| Lookup or range by one alternate key across the whole table | Global secondary index |
| Ad hoc filters on many columns, text search, geo, aggregates | Search index |
| React to every change, rebuild caches, feed analytics | Tunnel Service (change data) |
In the telemetry example, an alerts screen that lists devices whose battery is below 20 percent, newest first, across the fleet is a search index query. Looking up a device by its serial number is a global secondary index on a predefined serial column.
Capacity units and cost
In CU mode, every read and write consumes capacity units: one CU per 4 KB of data, rounded up per operation. You can reserve read and write CUs per table and pay a lower rate for that baseline; consumption above the reservation is metered at a higher unit price. The alternative VCU mode buys computing capacity for the instance instead of per-operation units and suits steady, cost-sensitive workloads.
The rounding drives row design. A 700-byte reading costs 1 write CU. A row that also carries a 9 KB JSON blob costs 3 write CUs on every write and 3 read CUs on every full-row read, even when the dashboard only wanted the battery level. Two fixes: keep large, rarely read payloads in Object Storage Service and store the object key (a 600-byte pointer row costs 1 CU), and pass the columns you need to reads so the response stays small.
Failure modes
These are the failures that reach production.
- Hot partition. Sequential or low-cardinality partition keys (dates, tenant IDs where one tenant dominates) concentrate traffic. Symptom: throttling errors on some keys while the table total is far below its reservation. Fix the key with a hash prefix or a bucket suffix; splitting cannot help a single key value.
- Treating BatchWriteRow as atomic. By default each row is processed independently. Half a batch lands and the caller retries everything, double-counting increments. Inspect per-row results and retry only failed rows.
- Read-after-write through a global index. A user creates a record, the page queries the global secondary index or search index, and the record is missing. Read the base table by primary key after a write, use a local secondary index, or tolerate the lag in the UI.
- Pagination that stops early. Loops that exit on an empty or short page silently drop data, mostly when filters are involved. Exit only on a null continuation key.
- Unbounded retries. The SDK retries retriable errors; wrapping it in your own aggressive loop turns throttling into a retry storm. Add jittered backoff and a cap, and alert on throttle rate.
Trade-offs and when to choose it
Tablestore sits next to familiar systems, and the comparison clarifies when to pick it.
| Concern | Tablestore | Comparable choice |
|---|---|---|
| Key model | 1 to 4 typed PK columns, first one partitions | DynamoDB: partition key plus optional sort key |
| Ordering | Sorted by full PK, range reads across the table | Bigtable and HBase: sorted row key |
| Operations burden | Serverless, no nodes to size | HBase: you run the cluster |
| Alternate access paths | Local and global secondary indexes, search index | DynamoDB: LSI and GSI; search needs another service |
| Multi-row atomicity | Per row; batch rows independent by default | Similar for all of these at scale |
Choose Tablestore when you run on Alibaba Cloud and the workload is key-addressed at high volume: telemetry, logs and metadata, user feeds, chat histories, order state. Do not choose it for workloads that need multi-row transactions across arbitrary keys or rich joins; a relational service is the honest answer there.
Related reading
Related reading on this site: Alibaba Object Storage Service for the blobs you should keep out of rows, Alibaba Function Compute for event-driven writers and readers, Cloud Bigtable and DynamoDB in depth for the closest relatives, HBase hotspotting for row-key techniques that transfer directly, and LSM trees for the storage structure behind this family of databases.
What to do next
- Write down every query the application will run, with its expected rate and latency target, before choosing a primary key.
- Design the primary key so the first column has high cardinality and no sequential pattern; add a short hash prefix if IDs are sequential.
- Create the table with max versions 1 and a TTL if data expires; change these only with a reason.
- Use EXPECT_NOT_EXIST for creates and a version-column condition for read-modify-write updates.
- Write every range read as a loop that ends only on a null nextStartPrimaryKey.
- Retry only failed rows of a BatchWriteRow, and cap retries with jittered backoff.
- Pick a local secondary index for read-your-writes, a global one for alternate keys, a search index for ad hoc filters, and Tunnel Service for change data.
- Move payloads above a few kilobytes to OSS, then load-test with production-like key distributions and watch per-partition throttling before launch.