HBase rarely fails because of a bug. It fails because a reasonable-looking design decision runs into one of its architectural facts. A timestamp row key is natural, and it sends every write to one server. Ten column families mirror the domain model neatly, and they multiply flush and compaction work. A Connection per request looks like ordinary resource hygiene, and it takes the cluster's metadata lookups with it. Each of these anti-patterns is a choice that works in testing, where data is small and traffic is even, and breaks in production.

This article is a symptom-first catalogue. Each entry gives what you will see, why HBase behaves that way, and the fix. It closes with a worked design review of one table and a list of signals for finding these problems in a running cluster. Deep treatments live elsewhere on the site, and this page links to them instead of repeating them. Start with HBase schema design if you are designing a new table rather than diagnosing an old one.

The five facts behind every anti-pattern

Five facts explain almost every anti-pattern on this page.

  1. Rows are sorted by key, and a region is a contiguous key range served by exactly one RegionServer. Load follows key distribution, not row count.
  2. The row is the unit of atomicity. Mutations to one row are atomic. Nothing spanning rows is, unless you use special mechanisms. A row also never splits across regions.
  3. Writes go to the WAL and an in-memory MemStore per column family. MemStores flush to immutable HFiles. Background compaction merges HFiles, and only a major compaction removes deleted and expired cells for good.
  4. Reads merge the MemStore with every relevant HFile. More files and more delete markers make every read more expensive.
  5. Clients cache region locations from hbase:meta. A long-lived client is cheap. A client that starts cold for every request has to look up those locations again each time.
Where each anti-pattern hits the HBase architectureClientConnection, Tablehbase:metakey range to serverlocateRegionServerRegion A[a, f)Region B[f, +inf) hotMemStore per familyWALHFilesPut / Get / ScanCompactionmerge HFiles, drop deletesSplitregion over max sizeAnti-patterns by layer:Client: a Connection per request, single Puts in a loop, giant multi batches, no timeoutsKeys: monotonic keys pile onto the last region (B); one hot counter row serialises on its row lockMemStore: many column families multiply stores, flush work and file countsHFiles: huge cells and giant rows; delete-heavy queues leave markers until major compactionOperations: no pre-split, major compactions at peak, SKIP_WAL, thousands of tiny regions
The write and read path, annotated with the layer each anti-pattern stresses. Region B is hot because monotonic keys always land at the end of the key space.

Key anti-patterns

Monotonically increasing row keys. Symptom: one RegionServer at full CPU and write latency while the others sit idle. The busy region splits, the new top half becomes the hot region, and the problem moves without going away. Cause: timestamps, sequence numbers and time-ordered IDs always sort to the end of the key space. Fix: lead the key with something well distributed, such as a hash prefix of an entity ID, a small salt bucket, or a reversed counter. Accept that time-range scans then fan out across buckets. The diagnosis workflow is in HBase hotspot analysis.

One hot counter row. Symptom: increments on one row time out under load, even though the cluster has spare capacity. Cause: an increment is a read-modify-write under that row's lock, so concurrent increments on one row run one at a time. Fix: shard the counter across N rows, for example counter#07, and sum on read. Or aggregate in the application and write one increment per second.

Keys that do not match the reads. Symptom: the important query is a full scan with a filter. Cause: HBase indexes exactly one thing, the row key. Fix: design the key from the read patterns, and maintain a second table keyed for the second query if you need one. This is the most expensive anti-pattern to fix late, because it means rewriting the data.

Data-shape anti-patterns

Too many column families. Symptom: many small HFiles, frequent flushes, high compaction load. Cause: each family is a separate store with its own MemStore and files in every region. Older HBase versions flushed all families together. Since 1.1 the default policy flushes only families above a lower bound (with a 16 MB floor, hbase.hregion.percolumnfamilyflush.size.lower.bound.min), but every family still adds files to compact, store objects to manage and MemStore to account for. Fix: one or two families, split only by access pattern or by settings such as TTL and compression. Do not split by domain concept. See column family design.

Giant rows. Symptom: RowTooBigException, RegionServer heap pressure, and regions that will not split below some size. Cause: a row is never split, and a whole-row Get materialises it. The server refuses to return rows larger than hbase.table.max.rowsize, which defaults to 1 GB. Fix: move the unbounded dimension, usually time, into the row key. Read wide rows with Scan.setBatch or column pagination. Limits and patterns are in wide row limits.

Huge cells. Symptom: client exceptions on Put, slow flushes and compactions that rewrite the same large values again and again. Cause: by default the client rejects cells larger than hbase.client.keyvalue.maxsize (10 MB). Below that, large values still take block cache space and get rewritten on every compaction. Fix: store objects above a few megabytes in an object store and keep a pointer in HBase. Use the MOB feature for values between roughly 100 KB and 10 MB.

Versions as a history table. Symptom: "we lost history" after a compaction. Cause: a column family keeps only its configured VERSIONS (1 by default in HBase 2), and extra versions are dropped at compaction. Fix: if history is data, put the timestamp in the row key or the qualifier.

Long family and qualifier names. Every cell stores its full coordinates: row, family, qualifier and timestamp. A 30-byte qualifier on a 4-byte value is mostly overhead. Block encoding such as FAST_DIFF and compression reduce this, but short names cost nothing.

Access anti-patterns

Unbounded scans on online paths. Symptom: scanner timeouts (the client lease defaults to 60 seconds, hbase.client.scanner.timeout.period), block cache churn, and latency spikes for every other tenant. Cause: a scan without start and stop rows reads every region. Filters do not help with I/O, because they run on the server after the data has been read. Fix: every online scan gets a key range and a limit. Full-table work goes to MapReduce, Spark or a snapshot export, away from the serving path.

Relational habits. Symptom: client-side joins making hundreds of Gets per request, and secondary indexes that disagree with the base table after a client crash. Cause: HBase has no joins and no cross-row transactions. A dual write to a base table and an index table can half-succeed. Fix: denormalise, so one read returns what one page needs. Make index writes idempotent, and repair them with a periodic job, or use a layer that maintains indexes for you, such as Apache Phoenix.

HBase as a queue. Symptom: consumers that scan for the oldest item get slower every hour. Cause: a consumed item is deleted, and a delete writes a marker. Scans must skip every marker until a major compaction removes them, so the head of the queue becomes a field of tombstones. Fix: use a real log or queue such as Kafka. If you must stay in HBase, write time-bucketed rows and drop whole buckets with TTL instead of deleting items one by one.

Client anti-patterns

A Connection per request. Connection is heavyweight and thread-safe. It owns the ZooKeeper or registry session, the region location cache and thread pools. Table is lightweight and not thread-safe. The anti-pattern is to create a Connection per request. The fix is one Connection per process and a Table per unit of work.

Single Puts in a loop. Each table.put(p) is a round trip. Ingest at scale should use BufferedMutator, which groups mutations by region server. Very large table.batch calls fail the other way: they produce RPCs that hit timeouts and retry as a whole. The trade-offs are in batching operations.

// Anti-pattern: new connection, single puts, no timeouts
try (Connection conn = ConnectionFactory.createConnection(conf);      // per request!
     Table t = conn.getTable(TableName.valueOf("events"))) {
    for (Event e : events) t.put(toPut(e));                           // one RPC each
}

// Better: one Connection for the process, BufferedMutator for ingest, explicit timeouts
Configuration conf = HBaseConfiguration.create();
conf.setInt("hbase.rpc.timeout", 2000);
conf.setInt("hbase.client.operation.timeout", 10000);
Connection conn = ConnectionFactory.createConnection(conf);          // created once, shared

BufferedMutatorParams params = new BufferedMutatorParams(TableName.valueOf("events"))
        .writeBufferSize(4L * 1024 * 1024)
        .listener((ex, m) -> {
            for (int i = 0; i < ex.getNumExceptions(); i++) deadLetter(ex.getRow(i), ex.getCause(i));
        });
try (BufferedMutator mut = conn.getBufferedMutator(params)) {
    for (Event e : events) mut.mutate(toPut(e));                      // flushed in region-grouped batches
}

// Online read: bounded range, bounded result
Scan scan = new Scan().withStartRow(prefix).withStopRow(nextPrefix(prefix)).setLimit(100).setCaching(100);

No timeouts, unbounded retries. The client's default retry schedule can keep a request alive for minutes while a region moves. That is reasonable for a batch job and wrong behind a user request with a 2-second budget. Set hbase.rpc.timeout and hbase.client.operation.timeout for each workload, ideally on separate Connections.

Operational anti-patterns

No pre-split. A new table has one region, so the first hours of a bulk load go to one server until splits catch up. Pre-split from the key distribution you expect, for example one region per salt bucket. Sizing guidance is in region count and sizing.

Periodic major compaction at peak. hbase.hregion.majorcompaction defaults to 7 days, with jitter. On large tables, a time-based major compaction during business hours rewrites terabytes alongside user traffic. Many operators set it to 0 and trigger major compactions off-peak, table by table, from a scheduler.

Skipping the WAL for speed. Durability.SKIP_WAL makes writes faster by making them lost if a RegionServer crashes before flushing. Use it only for data you can regenerate, and write that decision down. ASYNC_WAL is the middle ground when a small loss window is acceptable.

Too many tiny regions, or a few giant ones. Thousands of regions per server mean many MemStores competing for heap and frequent small flushes. Very large regions make splits, moves and recovery slow. The default split ceiling, hbase.hregion.max.filesize, is 10 GB.

Worked example: reviewing a device_events table

A team proposes a device_events table for an IoT fleet of 200,000 devices, each sending a reading every 10 seconds. Their design: row key <epoch_millis>#<device_id>. Families meta, temp, power, net and alerts. Two queries: the last hour for one device, and all devices with an alert today. Writes go through a REST handler that opens a Connection per call. A review finds five anti-patterns:

FindingAnti-patternChange
Key starts with epoch millisMonotonic key: all 20,000 writes per second hit one region<hash4(device)><device_id><reverse_ts>
Five families by domainToo many column familiesOne family d with short qualifiers; alerts move to a second table
"All alerts today" by scan and filterKey does not match the readTable alerts_by_day keyed <yyyymmdd><bucket><device_id>
Connection per REST callClient lifecycleOne shared Connection; BufferedMutator in the ingest path
Single region at creationNo pre-splitPre-split on the 4-hex-digit hash prefix into 64 regions

With the device ID behind a hash prefix and the timestamp reversed, "last hour for device D" becomes a single prefix scan that starts at the newest reading. The 20,000 writes per second spread across all 64 regions. The alerts table is written alongside the main table, idempotently, and is small enough to scan by day. Nothing about the hardware changed. The design was simply aligned with the five facts above.

Finding anti-patterns in a running cluster

Most of these anti-patterns leave a signature you can find without reading any application code.

  • Request skew by region. The RegionServer web UI and the JMX metrics at /jmx expose per-region read and write request counts. One region with ten times the median is a key problem.
  • Store file counts near the blocking limit. Writes to a store stall when it reaches hbase.hstore.blockingStoreFiles (16 by default). Regularly reaching it points at too many families, undersized flushes or a compaction backlog.
  • Exceptions in client logs. RowTooBigException, scanner lease expiries and RetriesExhaustedWithDetailsException each point at one entry in this catalogue.
  • The slow log. RegionServers log slow and oversized RPCs. On recent HBase 2.x releases the shell can query the online slow log with get_slowlog_responses. Full-table scans show up immediately.
  • Connection churn. A steady stream of new ZooKeeper sessions, or meta lookups that track the request rate, means clients are not reusing Connections.

What to do next

  1. Write down every production read pattern and check that each is a Get or a bounded prefix scan on some table's row key.
  2. Check row keys for monotonic prefixes, and look at per-region request counts for skew.
  3. Count column families per table and justify any beyond two by access pattern or settings.
  4. Search client logs for RowTooBigException, scanner timeouts and retries exhausted, and map each to a catalogue entry.
  5. Audit client code for Connection lifecycle, BufferedMutator use and explicit timeouts.
  6. Check whether time-based major compaction is enabled on large tables, and schedule it off-peak instead.
  7. List every use of SKIP_WAL or ASYNC_WAL and confirm the data can be regenerated.
  8. Put this catalogue into your design review template, so the next table is reviewed before it holds data.
Key takeaway: HBase anti-patterns are reasonable-looking choices that collide with its architecture: sorted keys served by one region, the row as the unit of atomicity, per-family stores, merge-on-read files and cached region locations. Lead keys with a distributed prefix, design keys from reads, keep one or two families, bound rows, cells and scans, reuse one Connection and batch writes, pre-split, schedule major compactions, and treat SKIP_WAL as data loss. Detect problems from region skew, store file counts and client exceptions.