HBase stores a table as a sorted map from row key to row, and splits that sorted key space into contiguous ranges called regions, each served by one RegionServer. That design makes range scans fast and makes one kind of key disastrous: a key that increases monotonically, such as a timestamp, a sequence number or an auto-increment ID. Every new row sorts after the last one, so every write lands in the final region, and a 40-node cluster performs like a one-node cluster.

Hashing the key is the most common fix. This article explains what hashing does to the key space, the three ways to apply it, how to pick a hash and a prefix length, how to pre-split the table so the hash pays off on day one, and what you give up. The salting article covers bucket salts, where a small number of buckets are chosen by a hash modulo N; this one covers hash prefixes and fully hashed keys, which behave differently once a table grows, and the code that builds them correctly in more than one language.

Why monotonic keys melt one RegionServer

A region is the unit of load in HBase. One RegionServer hosts it, one memstore buffers its writes per column family, and one write-ahead log on that server records them. If a write pattern sends every row to the same region, adding servers adds nothing until the region splits, and after a split the new rows still go to the upper daughter. You see this as one RegionServer with a high request count and growing flush and compaction queues while the others are idle; hotspot analysis describes how to confirm it from metrics.

Hashing replaces the order of the key with a pseudo-random order. A good hash maps similar inputs to unrelated outputs, so consecutive IDs or timestamps land in different, evenly spread parts of the key space. If the table is split into regions along that hashed space, writes spread across all of them in proportion to their share of the space. The price is that rows which used to be adjacent are no longer adjacent, and HBase's main fast path, the contiguous range scan, no longer covers them.

Sequential keysR1idleR2idleR3idleR4all writests=nowHash-prefixed keys00-3f1/440-7f1/480-bf1/4c0-ff1/4row key = hash(user_id)[0:2] | user_id | Long.MAX - tsprefix spreads writes; id keeps rows unique; reversed time keeps a user's newest rows firstGet(user, ts)recompute prefix: one regionScan(one user)prefix + id as start row: one regionScan(all users, last hour)no longer contiguous: needs every region, or a secondary index
Sequential keys send every write to the last region; a hash prefix spreads writes while keeping per-user Gets and scans in one region.

Three shapes: hash prefix, full hash, bucket salt

There are three distinct shapes, and choosing between them is the main design decision.

ShapeKey layoutPoint GetScan for one entityOriginal key recoverable
Hash prefixhash(id)[0:k] + id + restYes, recompute prefixYes, if id leads the restYes, it is in the key
Full hashhash(id)Yes, recompute hashNoNo, store id in a column
Bucket salt(hash(id) mod N) + id + restYesYesYes

A hash prefix puts the first few bytes of a hash in front of the natural key. Uniqueness comes from the natural key, so the prefix can be short; its only job is distribution. This is the default choice for entity-keyed tables such as users, devices or orders.

A full hash replaces the key with the hash. It gives fixed-length keys and hides the original identifier, which some teams want for privacy, but uniqueness now depends on the hash not colliding, and any query other than an exact lookup becomes impossible. With a full MD5 (128 bits) collisions are not a practical concern. With a hash truncated to 32 bits, by the birthday bound you reach a 50 percent chance of at least one collision at about 77,000 keys, so never truncate a full-hash key.

A bucket salt takes the hash modulo a small N, often 8 to 64, and prefixes that number. It bounds read fan-out to N but fixes the parallelism at N forever. A hash prefix instead gives 65,536 possible prefixes with two bytes, so the table can grow from 16 regions to 1,000 and keep splitting evenly without re-keying. That growth behaviour is the main reason to prefer a prefix for large tables.

Building the key the same way everywhere

The rule that makes hashed keys work is that every writer and every reader computes the key in exactly the same way, byte for byte. Put the key builder in one small library and use it everywhere. Here is a Java version for an events table keyed by user and time:

import java.nio.charset.StandardCharsets;
import java.security.MessageDigest;
import org.apache.hadoop.hbase.util.Bytes;

public final class EventKeys {
    static final int PREFIX_BYTES = 2;          // 65,536 prefixes

    private static byte[] prefix(byte[] id) {
        try {
            MessageDigest md5 = MessageDigest.getInstance("MD5");  // distribution, not security
            byte[] h = md5.digest(id);
            return Bytes.head(h, PREFIX_BYTES);
        } catch (java.security.NoSuchAlgorithmException e) {
            throw new IllegalStateException(e);
        }
    }

    /** prefix | userId | 0x00 | (Long.MAX_VALUE - tsMillis) */
    public static byte[] rowKey(String userId, long tsMillis) {
        byte[] id = userId.getBytes(StandardCharsets.UTF_8);
        return Bytes.add(prefix(id), Bytes.add(id, new byte[] {0}),
                         Bytes.toBytes(Long.MAX_VALUE - tsMillis));
    }

    /** Start and stop rows that cover every event of one user, newest first. */
    public static byte[][] userRange(String userId) {
        byte[] id = userId.getBytes(StandardCharsets.UTF_8);
        byte[] start = Bytes.add(prefix(id), id, new byte[] {0});
        byte[] stop  = Bytes.add(prefix(id), id, new byte[] {1});
        return new byte[][] {start, stop};
    }
}

Three details matter. The hash input is the UTF-8 bytes of the ID, stated explicitly, because a Python or Go service will hash bytes, not Java strings. The zero byte after the ID stops user ab from sharing a scan range with user abc. And the timestamp is subtracted from Long.MAX_VALUE so a scan of one user returns the newest events first. A Python service must produce identical bytes:

import hashlib, struct

LONG_MAX = 2**63 - 1

def row_key(user_id: str, ts_millis: int) -> bytes:
    uid = user_id.encode("utf-8")
    prefix = hashlib.md5(uid).digest()[:2]
    return prefix + uid + b"\x00" + struct.pack(">q", LONG_MAX - ts_millis)

# Never use Python's built-in hash(): string hashing is randomised per process
# (PYTHONHASHSEED), so two workers would write the same user to different keys.

Java's String.hashCode() is a poor choice for the same family of reasons: it is language-specific and its low bits are weak for short similar strings. MD5 is fine here because nobody relies on it for security; Murmur3 is faster if you have a shared implementation in every language. Whatever you choose, record it in the table's documentation, because changing it later means rewriting every row.

Pre-splitting so the hash pays off

A hashed key only spreads load across regions that exist. A new table starts as one region, so without pre-splitting the hash buys nothing until the table has grown through several splits. Create the table with split points that divide the prefix space evenly. For raw byte prefixes, HBase's UniformSplit algorithm divides the byte space evenly; for keys whose prefix is a hexadecimal string, HexStringSplit does the same over hex characters. In the shell:

# raw two-byte MD5 prefix, as in the code above
create 'events', {NAME => 'e', COMPRESSION => 'SNAPPY'},
       {NUMREGIONS => 64, SPLITALGO => 'UniformSplit'}

# if you encode the prefix as hex text instead (e.g. '3fa1' + id)
create 'events_hex', 'e', {NUMREGIONS => 64, SPLITALGO => 'HexStringSplit'}

Do not mix them: HexStringSplit boundaries such as 04000000 are ASCII text, and a table whose keys start with raw bytes would put almost every row into a handful of regions. If you prefer explicit split points, generate them from the same prefix definition so they cannot drift. For a two-byte prefix and 64 regions the boundaries are multiples of 1,024:

splits = [(i * 65536 // 64).to_bytes(2, "big") for i in range(1, 64)]
# b'\x04\x00', b'\x08\x00', ... b'\xfc\x00'  -> pass as SPLITS to create

Because the prefix space is uniform, later automatic splits stay balanced as each region grows. See regions and splits for how split policies and the normaliser interact with a pre-split table.

Worked example: an events table

Worked example. A product analytics service writes 60,000 events per second for 20 million users into a 24-node cluster. The original key was tsMillis + userId, chosen so the team could scan the last hour. In production one RegionServer carried nearly the whole write rate; its memstores flushed constantly, compaction queues grew, and clients saw rising write latency and eventually RegionTooBusyException retries.

The access patterns, written down before redesigning, were: (1) show one user's recent events, many times per second; (2) fetch one event by user and time; (3) a nightly job reading the previous day for all users. Pattern 1 dominates, and it is an entity scan, so the key above fits: hash prefix, user ID, reversed time. The team created the table with 96 regions, four per RegionServer, using UniformSplit. Each region now receives about 1/96 of the writes, roughly 625 events per second, and RegionServer request counts are within a few percent of each other.

Pattern 3 was the cost. The last day is no longer a contiguous range, so the nightly job became a full-table scan with a time-range filter, run as a parallel job with one task per region. Because each region's data for one user is sorted newest first, and HFiles carry time-range metadata, many files are skipped, but the job still reads more than before. The team accepted that because it runs once a day off-peak. Had the global time query been interactive, the right answer would have been a second table keyed by time bucket, written alongside the first.

Failure modes

Hashing fixes one problem and introduces several of its own.

  • Writers disagree on the key. One service hashes UTF-16, another UTF-8; one trims whitespace, another does not. Rows silently split into two keys, and reads miss data. Ship one key library and add a cross-language test that compares bytes for a fixed list of IDs.
  • A single hot entity is still hot. Hashing spreads many keys; it cannot split one. If one device sends half the traffic, all its rows share a prefix and a region. Handle that with a per-entity sub-bucket or by aggregating before writing.
  • No pre-split. The table starts with one region and the hash only helps after weeks of splits.
  • Wrong split algorithm. Hex split points on raw byte prefixes leave most regions empty.
  • Range scans that used to work now need every region. A scan with only a time range must touch the whole table. Check the scan behaviour of each query before migrating.
  • Changing the hash or prefix length later. Every key changes, so it is a full rewrite into a new table with dual writes during the move. Choose two bytes unless you have a reason not to.
  • Truncated full hashes. Keys that are only a short hash collide and overwrite each other without any error.

Trade-offs

Choose a hash prefix when your hot path looks up or scans by entity and your write rate is high enough that a single region cannot absorb it. Choose a full hash only when you never scan and want fixed-width, opaque keys. Choose a bucket salt when you need bounded read fan-out over a time dimension. Choose none of them if writes are modest; a natural, well-distributed key such as a random UUID already spreads load, and adding a hash only removes useful order.

Alternatives worth weighing are key reversal, which spreads keys whose low digits vary fastest but keeps per-entity locality only for reversible IDs, and Apache Phoenix, which manages salting for you with a table option. The broader trade-offs of key design are in HBase schema design.

What to do next

  1. List every read and write pattern with its rate before choosing a key shape.
  2. If the hot path is per-entity, use a two-byte hash prefix followed by the natural ID.
  3. Write one key-builder library, define the byte encoding, and test it across languages.
  4. Never use language-built-in hashes such as Python's hash() or Java's String.hashCode() for keys.
  5. Pre-split with UniformSplit for byte prefixes or HexStringSplit for hex text, a few regions per server.
  6. Add a terminator byte after variable-length IDs so scan ranges do not overlap.
  7. Plan how global time-range queries will run: full parallel scan, secondary table or index.
  8. After launch, compare per-RegionServer request counts to confirm writes are balanced.
Key takeaway: Monotonic row keys send every write to one region, so one RegionServer does all the work. Putting a short hash of the entity ID in front of the natural key spreads writes evenly while keeping exact lookups and per-entity scans to a single region; a full hash goes further but gives up scans and needs the untruncated hash for uniqueness. The scheme only works if every writer builds identical bytes, so use one shared key library and never a language's built-in hash, and it only helps from day one if the table is pre-split with a split algorithm that matches the prefix encoding. The cost is that ranges across entities, such as all events in the last hour, now touch every region.