HBase is not dead in 2026. The Apache project shipped 3.0.0 on 5 August 2026 and maintenance releases 2.6.7 and 2.5.16 on 1 October 2026, and large installations still run it for workloads it handles very well. What has changed is where new wide-column and key-value workloads land. Teams leaving on-premises Hadoop no longer want to run HDFS, ZooKeeper and a fleet of JVM region servers to get a sorted key-value store, and managed services now cover most of what HBase did. The slug calls HBase declining; there is no reliable public data on its market share, so this article will not invent any. It deals with the question you can actually answer: for each of your tables, is HBase still the right home, and if not, which system is?

Moving the bytes is covered in HBase migration, in depth and the AWS comparison in HBase vs DynamoDB. This page is the step before both: understand the contract you depend on, classify the workload, map each shape to a destination, and find the semantics that change when you leave.

Advertisement

The contract you are actually depending on

Before comparing products, write down what HBase gives you, because an alternative is only an alternative if it keeps the parts your code relies on. HBase stores rows sorted by a byte-array row key. The keyspace is cut into regions, contiguous key ranges, and each region is served by exactly one region server at a time. That single-owner design is the source of most of its properties.

  • Ordered keyspace. A scan from any start key to any stop key returns rows in key order, across entity boundaries. Prefix scans, reversed-timestamp tricks and time-range reads all depend on this.
  • Strong consistency per row. Because one server owns a row, a read after an acknowledged write sees it. There is no tunable consistency level and no read repair.
  • Atomic single-row operations. Every mutation to one row is atomic across its column families, and the server offers checkAndMutate, increment and append as compare-and-set primitives. There are no multi-row transactions outside one region without Phoenix or application logic.
  • Cell versions and timestamps. Each cell can keep several timestamped versions; reads can ask for a time range or the last N versions.
  • Server-side code. Filters run on the region server, and coprocessors let you run observers and endpoints inside it.
  • Hadoop integration. Bulk loads write HFiles directly; snapshots and MapReduce or Spark jobs read table files from HDFS.

The costs come from the same design. When a region server dies, its regions are unavailable until the write-ahead log is split and the regions reopen elsewhere, typically seconds to minutes depending on WAL volume. The cluster needs HDFS and ZooKeeper, JVM heap and GC tuning, compaction management and a team that understands all of it. A poorly designed row key sends every write to one region, as HBase hotspotting explains, and no alternative fixes that for you: most of them hash or range-partition keys in the same way.

The landscape in 2026

The project is active: 3.0.0 is the new major line and 2.6 and 2.5 still receive releases, so staying is a supported choice. Read the 3.0 release notes and the upgrade section of the reference guide before planning a move to it; this page does not summarise its feature list. The alternatives fall into five groups.

OptionModelClosest to HBase inGives up or changes
Google Cloud BigtableSorted wide-column, managedData model, row keys, HBase client library, ordered scansCoprocessors; operational control; runs only on Google Cloud. Adds GoogleSQL queries and continuous materialized views
Apache Cassandra 5.xPartitioned wide-column, leaderlessWrite throughput, multi-datacentre replication, TTLGlobal key order; per-row strong consistency becomes a per-query consistency level
ScyllaDBCassandra-compatible, C++ shard-per-coreAs Cassandra, with lower tail latency per nodeLicence: 6.2 was the last AGPL release; later versions are source-available with a free tier up to 10 TB and 50 vCPUs per organisation
Amazon DynamoDBManaged key-value and documentPoint access at any scale with no serversRange scans only within one partition key; item size limits; AWS only
TiKVOrdered key-value on RaftOrdered keyspace split into ranges, like regionsNo column families or cell versions in the HBase sense; adds distributed transactions
Distributed SQL or PostgreSQLRelationalReplaces Phoenix workloadsWide sparse rows and very high write rates need care
Iceberg on object storageOpen table formatReplaces scan-heavy and batch tablesNot a low-latency point store

Two licensing and platform facts change some decisions. ScyllaDB's December 2024 change means that new open-source ScyllaDB clusters either stay on 6.2.x or accept the source-available terms; check where your footprint sits against the free limits. Of the options here, Bigtable is the only one that speaks the HBase client API, which makes it the cheapest move in code terms, but it ties you to one cloud.

Advertisement

Classify each table by workload shape

A cluster is not one workload. An HBase installation that has run for years usually holds user profiles, event logs, counters, a Phoenix schema someone added for reporting and a few tables nobody reads. Each has a different best home, so classify per table using evidence: region-server request metrics split into gets, mutations and scans; scan sizes from slow-query logs; and a code search of every client for checkAndMutate, increment, coprocessor endpoints and filter classes.

Classify each HBase table by workload shape, then pick a destination per shapeHBase tableone row key designPoint reads/writesget, put, small scansOrdered range scanstime series, prefixesRow-atomic updatescheckAndMutate, incrementSQL over HBasePhoenix, joins, indexesBulk scans, analyticsMapReduce / Spark jobsCassandra / ScyllaDBor DynamoDBBigtablesame model, HBase clientTiKV / stay on HBaseordered, transactionalDistributed SQLor PostgreSQL if it fitsIceberg + Spark/Trinoobject storageA single cluster usually holds several shapes; the answer is per table, not per cluster.
Workload shapes found in a typical HBase cluster and the destination each one maps to. The arrows are defaults, not rules: a hard latency target or a cloud commitment can override them.

The shape that decides most migrations is the cross-entity range scan. If a query reads rows for many different entities in key order, such as every device whose key starts with one region prefix, you need a store with a globally ordered keyspace: HBase, Bigtable or TiKV. Cassandra, ScyllaDB and DynamoDB keep order only inside a partition, so the same query becomes a fan-out over many partitions.

A decision procedure you can run

The rules above fit in a function. Feed it facts gathered per table and it gives a first answer that a human then reviews. The point is to make the reasoning explicit and repeatable across dozens of tables, not to automate the decision.

def destination(table):
    """Map one HBase table's observed behaviour to a candidate home. Inputs come from
    a week of region-server metrics and a code search of the clients (see the checklist)."""
    if table.needs_cross_row_transactions:
        return "distributed SQL or TiKV (transactional API)"
    if table.used_via_phoenix and table.queries_have_joins:
        return "distributed SQL, or PostgreSQL if the data fits one primary"
    if table.full_scan_share > 0.5 and table.p99_read_slo_ms is None:
        return "Iceberg tables on object storage, queried by Spark or Trino"
    if table.range_scans_cross_entity_prefixes:          # scans over many devices at once
        return "Bigtable (ordered keyspace) or stay on HBase"
    if table.uses_check_and_mutate or table.uses_increments:
        return "Bigtable, or Cassandra/ScyllaDB with LWT and counters after a semantics review"
    if table.on_aws and table.access == "point":
        return "DynamoDB"
    return "Cassandra or ScyllaDB (partition key = entity, clustering key = order)"

The order of the checks matters. Transactions and SQL needs come first because they rule out every wide-column store. Analytics comes next because moving a scan-heavy table to another low-latency store is paying for a property you do not use. Only then do access patterns choose between ordered and partitioned stores.

Translating a row key

Most HBase row keys pack several ideas into bytes: a salt or hash prefix to spread writes, an entity identifier, and often a reversed timestamp so the newest row sorts first. In Bigtable the same key works unchanged, since the model is the same, although Bigtable's guidance also says to avoid sequential leading keys. In Cassandra and ScyllaDB each part becomes an explicit column with a role.

// HBase: one table, row key = salt(1 byte) | device_id | Long.MAX_VALUE - ts
Scan scan = new Scan()
    .withStartRow(Bytes.add(salt(dev), Bytes.toBytes(dev), Bytes.toBytes(Long.MAX_VALUE - to)))
    .withStopRow(Bytes.add(salt(dev), Bytes.toBytes(dev), Bytes.toBytes(Long.MAX_VALUE - from)))
    .addFamily(Bytes.toBytes("m"))
    .setCaching(500);
try (ResultScanner rs = table.getScanner(scan)) {
    for (Result r : rs) handle(r);          // newest first, because of the reversed timestamp
}

-- Cassandra / ScyllaDB: the salt becomes an explicit day bucket in the partition key,
-- the reversed timestamp becomes a clustering order.
CREATE TABLE metrics (
    device_id text,
    day       date,
    ts        timestamp,
    cpu       double,
    mem       double,
    PRIMARY KEY ((device_id, day), ts)
) WITH CLUSTERING ORDER BY (ts DESC)
  AND default_time_to_live = 2592000;      -- 30 days

SELECT ts, cpu, mem FROM metrics
 WHERE device_id = ? AND day = ? AND ts >= ? AND ts < ?;   -- one query per day bucket

Three things happened in the translation. The salt, which existed only to spread writes, became a day bucket that also bounds partition size; a device writing every second produces 86,400 rows per day, which is a reasonable partition. The reversed timestamp became CLUSTERING ORDER BY (ts DESC). The scan across a time range became one query per day bucket, issued in parallel by the client. What did not survive is the ability to scan many devices in key order with one request.

Worked example: splitting one 40-node cluster

Consider a 40-node on-premises HBase cluster that a company wants to retire with its Hadoop estate. Classification of its six tables gives the following plan.

TableObserved shapeDestinationReason
user_profile99% gets and puts by user id, 5 ms p99 targetCassandra or DynamoDBPure point access; no ordered scans
device_metricsWrites per device; reads are time ranges for one deviceCassandra with day bucketsOrder needed only within one device
fleet_indexPrefix scans across all devices in a siteBigtableCross-entity ordered scans
countersincrement on hot rowsBigtable, or Cassandra counters after reviewCounter semantics differ; see below
report_phxPhoenix SQL with joins, read by a BI toolPostgreSQLFits one primary at 300 GB; needs joins
clicks_rawWritten once, read by nightly Spark jobsIceberg on object storageNo point reads at all

The result is not one replacement but four destinations. That sounds worse than it is: clicks_raw and report_phx usually make up most of the bytes and the operational pain, and moving them to systems the data team already runs removes most of the cluster. A sensible order is the analytical tables first, because they have no latency target and are easy to verify by row counts, then point-access tables, and the ordered-scan tables last.

Semantics that change silently

Most migration bugs are not data loss. They are behaviour that differs with no error message.

  • Conditional writes. HBase checkAndMutate is a cheap single-row operation. Cassandra lightweight transactions use Paxos with several round trips, cost far more and should not sit on a hot path; plan capacity separately.
  • Counters. Cassandra counters cannot be mixed with other columns in a table, cannot carry a TTL and are not idempotent on retry, so a timeout followed by a retry can double-count.
  • Versions and timestamps. HBase can keep several versions per cell. Cassandra keeps one value per cell and resolves conflicts by write timestamp, last write wins, so client-supplied timestamps from skewed clocks can lose writes.
  • Deletes and TTL. Both use tombstones, but Cassandra's tombstones interact with gc_grace_seconds and repair; deleting in bulk on Cassandra needs a plan HBase never required.
  • Server-side logic. Coprocessors and custom filters have no equivalent in managed stores. Move observers into a stream processor reading a change feed, and filters into the client.
  • Consistency. A Cassandra read at ONE after a write at ONE can miss the write. Use LOCAL_QUORUM for both if code assumed read-your-writes.

The way to catch these is a shadow phase: dual-write to the new store, read from both, and log every mismatch with its key before any traffic depends on the new system. The LSM-tree mechanics behind both families, which explain why tombstones and compaction behave as they do, are in LSM trees.

When staying is the right answer

Leaving is not automatically progress. Stay on HBase, and plan an upgrade path instead, when several of these hold: you already run Hadoop for other reasons and have people who operate it well; you depend on coprocessors or bulk-loaded HFiles in ways that would need rewrites; your access pattern is cross-entity ordered scans and you cannot use Google Cloud; or your data volume makes the managed alternatives' per-node pricing much higher than your amortised hardware. A cluster on a supported 2.5 or 2.6 release with good row-key design is a perfectly reasonable system in 2026. An unsupported 1.x cluster is not, whatever you decide about the long term.

What to do next

  1. Export a week of per-table request metrics (gets, mutations, scans and scan sizes) from your region servers.
  2. Search every client codebase for checkAndMutate, increment, append, coprocessor endpoints and custom filters, and record which tables they touch.
  3. Fill in the decision function's inputs for each table and review the suggested destinations as a team.
  4. Move scan-only tables to Iceberg first; verify with row counts and checksums.
  5. For each point-access table, write the new schema, then dual-write and compare reads for at least one full business cycle.
  6. If any cluster is on 1.x or an unsupported 2.x line, upgrade it to 2.5 or 2.6 now, whatever the long-term plan.
  7. Record licence terms (ScyllaDB) and cloud commitments (Bigtable, DynamoDB) as explicit constraints in the plan.
Key takeaway: HBase's value is a strongly consistent, globally ordered keyspace with cheap single-row atomics and tight Hadoop integration; its cost is HDFS, ZooKeeper, JVM operations and slow failover. The project is maintained, with 3.0.0 released in August 2026, so the question is per table, not whether HBase survives. Classify each table by shape: point access goes to Cassandra, ScyllaDB or DynamoDB, cross-entity ordered scans to Bigtable or TiKV or stay, SQL to a relational system, and scan-only data to Iceberg. Then shadow-test the semantics that change silently: conditional writes, counters, versions and consistency levels.