HBase and Cassandra are both wide-column stores built on log-structured merge trees: writes go to a log and an in-memory table, flush to immutable sorted files, and compaction merges those files later. Look only at the storage engine and they seem like the same system. They are not. They made opposite choices about who owns a piece of data, and almost every practical difference follows from that one choice: consistency, failover, multi-datacenter behaviour, what queries are cheap, and what operations staff spend their nights on.
This page compares them at that level. It explains each system's ownership model from first principles, traces a write and a failure through both, models the same telemetry workload in each with real code, compares conditional writes, and ends with a decision guide. Translating row-key designs between HBase and a hashed store is covered in HBase vs DynamoDB; this page does not repeat it.
Two ownership models
HBase descends from Google's Bigtable. The key space is one sorted range of row keys, cut into regions. At any moment exactly one RegionServer serves each region, and every read and write for those rows goes through it. Durability comes from below: the write-ahead log and the data files live in HDFS, which replicates blocks across machines. The RegionServer itself is a single point of service for its regions, and a single point of truth, which is why reads after writes are simply consistent. The HBase RegionServer covers the WAL, MemStore and block cache in detail.
Cassandra takes its storage engine from Bigtable and its distribution from Amazon's Dynamo. A partition key is hashed onto a token ring, each node owns token ranges (16 virtual nodes per node by default since Cassandra 4.0), and each row is stored on replication factor nodes, typically three. There is no owner. Any node can coordinate a request, every replica accepts writes, and the client chooses per request how many replicas must answer: the consistency level.
| Property | HBase | Cassandra |
|---|---|---|
| Who serves a key | One RegionServer per region | Any of RF replicas, via any coordinator |
| Copies made by | HDFS block replication | The database itself, per keyspace RF |
| Key placement | Sorted ranges, split as they grow | Hash of partition key on a token ring |
| Dependencies | HDFS, ZooKeeper, HMaster | None beyond the nodes |
| Query language | Java client API; SQL via Apache Phoenix | CQL |
Consistency
In HBase, a successful put is visible to every subsequent get of that row, because there is one server to ask. Operations on a single row are atomic, including several columns and families at once. There is no tunable knob, and no cross-row transaction in the core.
In Cassandra, consistency is a property of the pair of operations. With replication factor 3, writing at QUORUM (two acks) and reading at QUORUM (two replicas) guarantees the read overlaps the write on at least one replica, because 2 + 2 is greater than 3. Writing at ONE and reading at ONE is faster and can return stale data. Within a partition, a write is atomic and isolated; a logged batch across partitions is eventually applied in full but is not isolated. Conflicts between concurrent writes to the same cell are resolved by timestamp: the last write wins, which makes clock discipline on clients and servers a correctness concern. Cassandra consistency levels has the full matrix.
What happens when a server dies
When an HBase RegionServer dies, its regions are unavailable until three things happen: the master notices (driven by the ZooKeeper session expiring, so tune zookeeper.session.timeout deliberately), the dead server's WAL is split and replayed so unflushed edits are not lost, and the regions are opened on other servers. Reads and writes to those regions fail or wait during that window; everything else carries on. The window is typically seconds to minutes depending on configuration and WAL size, so measure it on your cluster rather than trusting a number. Read availability can be improved with region replicas, which serve timeline-consistent, possibly stale reads from secondary servers; writes still need the primary.
When a Cassandra node dies, nothing is reassigned. Requests that can still reach enough replicas for their consistency level succeed immediately. The coordinator stores hints for the dead node for a bounded window and replays them when it returns. Data that misses hints is fixed by read repair and by anti-entropy repair, which compares replicas with Merkle trees. The cost is moved, not removed: repair must run regularly, and specifically within gc_grace_seconds (864000 seconds, ten days, by default) of a delete, or a replica that missed a tombstone can resurrect deleted data after the tombstone is purged.
| Event | HBase | Cassandra |
|---|---|---|
| One server dies | Its regions unavailable until reassigned | Requests at QUORUM continue with RF 3 |
| Data on dead server | Safe in HDFS; WAL replayed | On other replicas; dead node catches up later |
| Ongoing cost | Low: no repair process | Scheduled repair; tombstone management |
| Network partition | Minority side cannot serve its regions | Both sides may accept writes at low CL |
Worked example: device telemetry in both
Take a fleet of 200,000 devices each sending a reading every ten seconds, with two queries: the last hour for one device, and a nightly job over everything. In HBase the row key must avoid a hotspot. A key that starts with a timestamp sends every write to the region holding the newest range, so one server takes all the load. Putting the device id first spreads writes by device; a one-byte salt computed from the device id spreads them further across pre-split regions without breaking the per-device range scan, because one device always lands in the same bucket. The reversed timestamp makes the newest readings come first.
// Row key: 1-byte salt | deviceId | Long.MAX_VALUE - epochMillis (newest first)
byte[] rowKey(String deviceId, long ts) {
byte salt = (byte) Math.floorMod(deviceId.hashCode(), 16);
return Bytes.add(new byte[] {salt}, Bytes.toBytes(deviceId + "|"),
Bytes.toBytes(Long.MAX_VALUE - ts));
}
try (Table t = conn.getTable(TableName.valueOf("telemetry"))) {
Put put = new Put(rowKey("pump-17", now));
put.addColumn(Bytes.toBytes("m"), Bytes.toBytes("temp"), Bytes.toBytes(71.5));
t.put(put);
// Last hour for one device: a contiguous range in one salt bucket.
byte[] start = rowKey("pump-17", now);
byte[] stop = rowKey("pump-17", now - 3_600_000L);
Scan scan = new Scan().withStartRow(start).withStopRow(stop, true).setCaching(500);
try (ResultScanner rs = t.getScanner(scan)) {
for (Result r : rs) { /* ... */ }
}
}In Cassandra the partition key is hashed, so there are no ordered-key hotspots to design around. The design problem moves to partition size: one partition per device for its whole life grows without bound, so the key buckets by day. Within the partition, rows are sorted by the clustering column, which makes the last-hour query a single sequential slice.
CREATE TABLE telemetry (
device_id text,
day date,
ts timestamp,
temp double,
PRIMARY KEY ((device_id, day), ts)
) WITH CLUSTERING ORDER BY (ts DESC);
-- Last hour for one device: one partition, a slice of clustering order.
SELECT ts, temp FROM telemetry
WHERE device_id = 'pump-17' AND day = '2026-10-03'
AND ts > '2026-10-03 09:00:00+0000';The nightly job shows the other side. HBase can scan the whole table in key order, and integrates with MapReduce and Spark through region-aligned input splits; bulk loading prepared HFiles is a standard ingestion path. Cassandra can also be read in full by Spark using its connector, which splits by token range, but a range query across partition keys in sorted order is not something CQL offers; ordering exists only within a partition. If your workload needs global ordered scans, such as all devices whose id starts with a prefix, HBase fits naturally and Cassandra needs a second table keyed for that query.
Conditional writes and counters
Compare-and-set is where the ownership model shows most clearly. In HBase it is cheap: the single owner holds a row lock, checks, and writes. Older clients used checkAndPut; the CheckAndMutate builder below is the HBase 2.4 and later form. Increments are atomic in the same way.
// HBase 2.4+: atomic compare-and-set on one row, served by its single owner.
CheckAndMutate cam = CheckAndMutate.newBuilder(Bytes.toBytes("order-991"))
.ifEquals(Bytes.toBytes("s"), Bytes.toBytes("state"), Bytes.toBytes("PAID"))
.build(new Put(Bytes.toBytes("order-991"))
.addColumn(Bytes.toBytes("s"), Bytes.toBytes("state"), Bytes.toBytes("SHIPPED")));
boolean applied = table.checkAndMutate(cam).isSuccess();Cassandra has no owner to lock, so a conditional write runs a Paxos consensus round among the partition's replicas, at a serial consistency level. It is correct and several times more expensive than a normal write, with more round trips and contention when many clients race for one partition. Use it for the uncommon operations that need it, such as uniqueness and state transitions, never in a hot loop; when to use Cassandra LWT goes deeper. Cassandra counters are a separate type with a known weakness: a retried increment after a timeout may apply twice.
-- Cassandra lightweight transaction: Paxos among the partition's replicas.
UPDATE orders SET state = 'SHIPPED'
WHERE order_id = 'order-991'
IF state = 'PAID';
-- Result row includes [applied]; mixing LWT and plain writes on the same cells breaks the guarantee.
Multiple datacenters
Cassandra was built for several datacenters. A keyspace using NetworkTopologyStrategy places a set number of replicas in each datacenter, and clients use LOCAL_QUORUM so a request waits only for replicas in its own region. Every datacenter accepts writes; replication between them is part of the normal write path, and conflicts resolve by timestamp. Losing a whole datacenter is a routing change, not a failover.
HBase replication ships WAL edits from one cluster to another and is asynchronous by default: after losing the source cluster, the peer may be missing the latest writes. Two-way replication is possible, and concurrent writes to the same cell in both clusters then resolve by timestamp, with no conflict detection. HBase 2 added synchronous replication (HBASE-19064) for an active and standby pair: the active cluster also writes its WAL to the standby's HDFS, the standby rejects client reads and writes, and failover is an explicit state transition. It works at table level and costs write latency, since every write now crosses to the other site.
Operations
Operating HBase means operating HDFS and ZooKeeper too: NameNode health, DataNode disks, ZooKeeper quorum, plus region sizing, splits, the balancer, and compaction tuning. Teams that already run Hadoop absorb this easily; teams that do not are adopting three systems. In return there is no repair process, and capacity is added by adding RegionServers and DataNodes.
Operating Cassandra means one kind of node, which is simpler to deploy, plus duties HBase does not have: scheduled repair, tombstone monitoring, and compaction strategy choice per table (Cassandra 5.0 added the Unified Compaction Strategy and Storage-Attached Indexes). Both are JVM systems, so garbage-collection pauses show up as tail latency in both, and both punish a schema that ignores the access pattern.
Failure modes
| Failure | System | Cause and fix |
|---|---|---|
| One server at 100% while others idle | HBase | Monotonic row key; salt or lead with a high-cardinality field, pre-split |
| Writes stall during flush pressure | HBase | Too many store files; compaction falling behind; tune flush and compaction |
| Deleted rows come back | Cassandra | Repair not run within gc_grace_seconds |
| Reads time out on one partition | Cassandra | Unbounded partition or tombstone-heavy slice; bucket the key, use TTL with care |
| Stale read after a write | Cassandra | Consistency levels that do not overlap; use QUORUM/QUORUM or LOCAL_QUORUM |
| Peer cluster missing recent data after failover | HBase | Async replication lag; monitor it, or use synchronous replication |
| LWT latency spikes | Cassandra | Contention on one partition; redesign so fewer clients race |
How to choose
| If you need | Lean towards | Why |
|---|---|---|
| Strong per-row reads without thinking about it | HBase | Single owner per region |
| Writes accepted through any single node failure | Cassandra | Leaderless replicas at QUORUM |
| Active-active across regions | Cassandra | Native per-datacenter replication |
| Global ordered scans and prefix queries | HBase | Sorted key space |
| Tight Hadoop, Spark batch and bulk-load integration | HBase | HDFS-native files and input splits |
| Cheap compare-and-set and increments | HBase | Row lock on the owner |
| No HDFS or ZooKeeper to run | Cassandra | Self-contained nodes |
Many teams decide on operations rather than data model: if a Hadoop platform and its people already exist, HBase is cheap to add; if not, Cassandra's single node type usually wins. Then check the decision against the two or three queries that matter most, written out in code as above.
What to do next
- Write down your top three queries and your consistency requirement for each, in plain words.
- Decide whether you need active-active writes across regions; if yes, start from Cassandra.
- Model the hottest table in both systems as shown here, and check for hotspots (HBase) and unbounded partitions (Cassandra).
- Load test a server failure in each: measure region unavailability in HBase and latency at QUORUM in Cassandra.
- Count the operational systems each choice brings, including HDFS and ZooKeeper, and who will run them.
- If you pick Cassandra, schedule repair within gc_grace_seconds before launch; if you pick HBase, monitor replication lag and compaction backlog.