HBase and Google Cloud Bigtable look like the same database. Both store sparse, sorted, multi-versioned maps: a row key, a column family, a column qualifier and a timestamp address one cell, rows are kept in lexicographic key order, and the key range is cut into contiguous pieces served by different machines. HBase was written as an open-source implementation of Google's 2006 Bigtable paper, and the Java client Google ships for Bigtable implements the HBase client interfaces, so a lot of HBase application code compiles and runs against Bigtable with a dependency and configuration change.

The resemblance stops at the API. Underneath, the two systems put data in different places, recover from failure in different ways, implement some operations with different semantics, and charge you in different currencies: HBase costs you machines and an operations team, Bigtable costs you nodes, storage and network by the hour. This page compares them as systems, so you can predict what will behave differently before you depend on it. Moving live tables from one to the other is covered separately in HBase migration, in depth, and the general Bigtable product overview lives in Cloud Bigtable.

Same data model, different machines

In HBase, a table is split into regions, each a contiguous key range, and each region is hosted by exactly one RegionServer at a time. The RegionServer writes every mutation to a write-ahead log (WAL) on HDFS, applies it to an in-memory MemStore, and flushes MemStores to immutable HFiles on HDFS. HDFS writes the first replica of each block to the DataNode beside the RegionServer. That is locality: reads are fast while the HFiles are local and slow down after a region moves, until a major compaction rewrites them. The internals of that server are in the RegionServer deep dive.

Bigtable divides a table into tablets, also contiguous key ranges, served by nodes in a cluster. Nodes do not own any storage. Tablet data and the commit log live in Colossus, Google's distributed file system, and every node can read every file. Google's documentation states the consequence directly: rebalancing a tablet from one node to another is fast because the data is not copied, and recovering from a node failure only requires moving metadata to a replacement. A Bigtable node is closer to a cache and request processor for a slice of the keyspace than to a storage server.

HBase: storage attached to serversClient (HBase API)meta lookup via ZooKeeper + hbase:metaRegionServer Aregions r1, r2RegionServer Bregions r3, r4DataNode (local)HFiles + WAL blocksDataNode (local)HFiles + WAL blocksHMaster + ZooKeeperassignment, failover, WAL splitshort-circuit readFailover: detect, split WAL, replay, reassignBigtable: storage shared by all nodesClient (Bigtable or HBase API)gRPC to a regional endpointRouting layerapp profile picks the clusterNode 1serves tablets t1, t2Node 2serves tablets t3, t4Colossus (shared file system)SSTables + shared commit logFailover: move tablet metadata, no data copied
The same logical model on two physical designs. HBase binds regions to servers whose local disks hold the data; Bigtable nodes serve tablets whose data sits in a shared file system.

What a server failure costs

The architecture difference shows up most clearly when a server dies. When an HBase RegionServer stops, its ZooKeeper session must first expire before the HMaster declares it dead; the default session timeout in hbase-default.xml is 90 seconds, and many operators lower it. The master then has to split or replay the dead server's WAL so that edits that were only in its MemStore are not lost, and reassign every region the server hosted. Until each region is open again, reads and writes to its key range fail or wait. With good tuning that takes seconds to tens of seconds; with large WALs or many regions per server it takes minutes, and moved regions lose locality.

When a Bigtable node fails, the commit log is already in shared storage, so tablets are reassigned by updating metadata. You cannot observe or tune that process. The trade is visibility for convenience: an HBase operator sees every recovery step in the master log; a Bigtable user sees a brief latency blip and, at worst, retried requests.

Splitting follows the same pattern: HBase operators pre-split, merge and normalise regions by hand, while Bigtable splits and merges tablets automatically and exposes no region APIs.

API compatibility, feature by feature

The Bigtable HBase client implements the HBase Connection, Table, BufferedMutator and Admin interfaces, so Put, Get, Scan, Delete, checkAndMutate, Increment and Append all work. Google publishes a list of what does not. The important entries, checked against that page:

FeatureHBaseBigtable via the HBase client
Coprocessors (observers, endpoints)SupportedNot supported; no server-side code of any kind
Custom filtersAny class deployed to the serversNot supported; built-in filters only, filter expressions capped at 20 KB
NamespacesSupportedNot used; tables live directly in an instance
Cell tags, cell visibility labels, cell ACLsSupportedNot supported; access control is IAM at instance and table level
Column family block size and compressionConfigured per familyManaged by Bigtable; settings ignored
Scan.setBatch, setCaching, setMaxResultsPerColumnFamily, setColumnFamilyTimeRangeSupportedNot supported or ignored
Delete(row, timestamp), addColumn(family, qualifier) latest-version delete, addFamily(family, ts), addFamilyVersionSupportedNot supported
Region APIs, split and merge, balancer switchesSupportedNot applicable; tablets are managed automatically
Append atomicityReaders can observe partial appendsFully atomic for readers and writers

Coprocessors are the biggest gap in practice. Secondary-index maintainers, server-side aggregations and Apache Phoenix all depend on them, so a Phoenix application cannot move to Bigtable without a rewrite. The usual replacements are client-side index writes, a change stream feeding a separate index, or pushing aggregation into Dataflow or BigQuery. The HBase alternatives guide covers what to do with workloads that cannot give these features up.

One codebase, two backends

Because both sides implement the same interfaces, the cleanest way to keep an application portable is to isolate connection creation and use only the shared API everywhere else. The Bigtable artifacts are published under the com.google.cloud.bigtable group as bigtable-hbase-2.x (standalone), bigtable-hbase-2.x-hadoop (for Hadoop classpaths) and bigtable-hbase-2.x-shaded (when protobuf or Guava versions clash), with 1.x equivalents for HBase 1 clients.

// Only this class knows which backend is in use.
public final class Connections {
  public static Connection open(Config cfg) throws IOException {
    if (cfg.backend().equals("bigtable")) {
      // com.google.cloud.bigtable.hbase.BigtableConfiguration
      return BigtableConfiguration.connect(cfg.projectId(), cfg.instanceId());
    }
    Configuration conf = HBaseConfiguration.create();
    conf.set("hbase.zookeeper.quorum", cfg.zkQuorum());
    return ConnectionFactory.createConnection(conf);
  }
}

// Everything else uses the shared HBase interfaces.
try (Connection conn = Connections.open(cfg);
     Table t = conn.getTable(TableName.valueOf("events"))) {
  Put put = new Put(Bytes.toBytes("device#0042#20261003T2149"));
  put.addColumn(Bytes.toBytes("m"), Bytes.toBytes("temp"), Bytes.toBytes(21.5d));
  t.put(put);

  boolean claimed = t.checkAndMutate(Bytes.toBytes("lock#job-7"), Bytes.toBytes("m"))
      .qualifier(Bytes.toBytes("owner")).ifNotExists()
      .thenPut(new Put(Bytes.toBytes("lock#job-7"))
          .addColumn(Bytes.toBytes("m"), Bytes.toBytes("owner"), Bytes.toBytes("worker-3")));
}

Run the same integration suite against both backends in CI. The gcloud Bigtable emulator covers functional behaviour, not performance, replication or garbage-collection timing.

Semantics that change under the same API

Several behaviours differ even when the code compiles. Each one has caught real migrations.

  • Timestamp units. HBase cell timestamps are milliseconds since the epoch. Bigtable stores microseconds. The HBase client converts for you, multiplying by 1,000 on write and dividing on read, which is why timestamps above Long.MAX_VALUE / 1000 cannot be represented. Google warns that large reversed-timestamp values are capped there and may not convert correctly. Mixing clients is the other trap: a Go or Python service on the native Bigtable client writes microseconds directly, so test how its cells read back through the HBase client. Check any code that stores version numbers or reversed times in the timestamp field first.
  • Deletes do not mask later puts. In HBase, a delete marker at timestamp T hides every version with timestamp at or below T until a major compaction removes both, so a late-arriving put carrying an old timestamp stays invisible. Bigtable applies deletes to the data that exists when the delete runs; a put sent afterwards is visible even if its timestamp is older. Replay and backfill jobs that rely on tombstones to suppress stale data behave differently.
  • Garbage collection is lazy and visible. In HBase, cells past a family's TTL or beyond its max-versions are filtered out at read time, so readers never see them. In Bigtable, garbage-collection rules run in the background and Google documents that eligible data can still be returned until it is actually removed. If correctness depends on expiry, add a timestamp-range or cells-per-column filter to the read.
  • Limits. Bigtable enforces a 4 KB row key, 16 KB column qualifier, 100 MB per cell (10 MB recommended), 256 MB per row (100 MB recommended), 100 column families per table and 100,000 mutations per batch. HBase's defaults are looser and mostly configurable, so check row-size outliers before moving.

Replication and consistency

HBase replication ships WAL edits asynchronously from a source cluster to peer clusters, as described in HBase replication architecture. It is eventually consistent, the operator configures peers, filters and serial ordering, and a failover to the replica is an application-level decision.

Bigtable replication is a property of an instance: an instance can have clusters in up to 8 regions, and Bigtable replicates between them automatically. Applications choose behaviour through app profiles. A profile with single-cluster routing sends all requests to one cluster, which gives read-your-writes consistency for that application, and it is the only routing that allows single-row transactions such as checkAndMutate, Increment and Append. A profile with multi-cluster routing sends each request to the nearest available cluster and fails over automatically, but reads are eventually consistent across clusters, and concurrent writes to the same cell in different clusters are resolved by Bigtable rather than serialised. A common design uses a single-cluster profile for writers that need conditional updates and a multi-cluster profile for serving reads.

Operations and cost model

HBase's operating cost is people and machines. You run ZooKeeper, HDFS NameNodes and DataNodes, HMasters and RegionServers; you tune garbage collection, block cache, MemStore and compactions; you upgrade Hadoop and HBase in lockstep; and you size the cluster for peak because adding RegionServers is slow and moving regions costs locality. In return you get full control and a hardware bill that does not grow with request volume.

Bigtable's cost is a metered bill: nodes per cluster per hour, storage per GB-month (SSD or HDD), and network egress and replication traffic. Storage per node is capped, 5 TB on SSD and 16 TB on HDD without tiered storage, so a large, cold dataset can force you to pay for nodes you do not need for throughput; tiered storage raises the SSD ceiling on higher editions. Autoscaling adjusts node count to a CPU target and to storage, and adding a node takes effect in minutes because no data moves. You lose control of compaction, compression and server code, and gain tools such as Key Visualizer, which draws access heat maps over the keyspace to find hotspots.

ConcernHBaseBigtable
Scaling upAdd servers, rebalance, wait for localityRaise node count or let autoscaling do it
Node failureZooKeeper timeout, WAL split, region reassignMetadata move, no data copy
Server-side logicCoprocessors, custom filtersNone
EcosystemPhoenix, Hive, Spark on HFiles, Kerberos, RangerDataflow, BigQuery federation, IAM, Spark connector
BillHardware and staff, mostly fixedNodes, storage, egress, metered

Worked example: a 40 TB telemetry table

Take a concrete workload: 40 TB of device telemetry, row keys of the form device#<id>#<reverse-timestamp>, around 60,000 writes per second at peak, point reads and short scans for dashboards, a 90-day TTL, and an existing HBase cluster of 24 RegionServers that uses one coprocessor to maintain a daily roll-up table.

  1. Check the hard incompatibilities. The coprocessor has to go. Replace it with a Dataflow or Spark job that reads a change stream or scans the previous day's keys and writes the roll-up. If the team cannot accept that, stop here and stay on HBase.
  2. Check the semantics. Timestamps are real event times in milliseconds, so the unit change is handled by the client. Expiry is by TTL; dashboards must add a time-range filter so lazily collected rows older than 90 days are not drawn.
  3. Size the cluster. 40 TB on SSD at 5 TB per node needs at least 8 nodes for storage alone. Size throughput with a load test on your own row sizes and read/write mix.
  4. Decide on replication. Telemetry ingest can use multi-cluster routing; the roll-up writer uses a single-cluster profile if it needs Increment.
  5. Compare cost honestly. Set the Bigtable estimate against the fully loaded HBase cost, including engineering time.

If either check fails, you now know exactly which feature or cost keeps you on HBase.

Failure modes

  • Hot tablets from monotonic keys. Timestamp-first keys hotspot both systems. Bigtable cannot split a single hot row, and its automatic splitting does not help when every write targets the end of the keyspace. Salt or reverse the key exactly as you would in HBase.
  • Unexpected expired data. A report counts rows past their GC rule because Bigtable had not collected them yet. Filter by time on read.
  • Silent no-ops. Scan.setCaching and setBatch are ignored, so memory-bounded scans that relied on them can return huge batches. Bound scans by row limit or key range instead.
  • Conditional writes failing under multi-cluster routing. Single-row transactions are rejected with that routing; use a single-cluster profile for those writers.
  • Storage-bound bills. Cold data forces node count up through the per-node storage cap. Consider HDD clusters, tiered storage or exporting cold ranges.

What to do next

  1. Grep the codebase for coprocessor classes, custom Filter subclasses, namespace usage and the unsupported Delete and Scan methods listed above.
  2. List every place that writes or interprets cell timestamps, and every job that depends on delete markers masking late writes.
  3. Wrap connection creation in one factory and run your integration tests against both HBase and the Bigtable emulator.
  4. Measure row-key, qualifier, cell and row size distributions against Bigtable's limits.
  5. Size Bigtable nodes from storage and from a load test, and write down which app profile each service will use.
  6. Build a fully loaded cost comparison, then follow the migration runbook if the numbers and features line up.
Key takeaway: HBase and Bigtable share a data model and a client API but not an architecture. HBase binds regions to servers with local HDFS storage, which gives control and server-side extensibility at the price of slow failover and heavy operations. Bigtable serves tablets from shared Colossus storage, so scaling and recovery are fast and invisible, but there are no coprocessors or custom filters, timestamps are microseconds, deletes do not mask later puts, garbage collection is lazy, and conditional writes need single-cluster routing. Inventory those differences in your code before comparing costs.