Every HBase schema discussion reaches the same fork: put many values for one entity in one row as separate columns (a wide table), or give each value its own row and push the item identifier into the row key (a tall table). The usual reasons offered for either choice, that wide rows save space and tall tables are slow to read, do not survive a look at how HBase stores a cell.

This article works from the storage format up. It shows what a row really is on disk, names the three things the wide-versus-tall decision changes, sizes both designs for a concrete time-series workload, covers the middle-ground layouts that production systems usually end up with, and lists the failure modes each design invites. The end-to-end key design process, including the short version of this rule, lives in HBase schema design; this page is the deep dive underneath it.

Advertisement

A row is a sorted run of cells, not a record

HBase does not store rows. It stores cells, each one a KeyValue whose key is the tuple (row, family, qualifier, timestamp, type). Cells are sorted by that key in the MemStore and in every HFile, so all cells of a row in one column family happen to sit next to each other. A "row" is simply the set of cells that share a row key. The layout of a single cell, before any block encoding, looks like this:

KeyValue serialization (one per cell, before block encoding)
  key length        4 bytes
  value length      4 bytes
  row length        2 bytes
  row               variable   <- repeated in EVERY cell of the row
  family length     1 byte
  family            variable
  qualifier         variable   <- in a wide row, this is where the item id lives
  timestamp         8 bytes
  type              1 byte     (Put, Delete, DeleteColumn, DeleteFamily ...)
  value             variable
Fixed overhead: 20+ bytes per cell plus row + family + qualifier
(HFile v3 can add a tags length and a sequence id on top).

Two consequences follow immediately. First, the row key is written into every cell. A wide row with a million columns does not store its key once; it stores it a million times, exactly as a tall table with a million rows does. The intuition that wide rows avoid repeating the key is false at the format level. What actually differs is where the item identifier lives: in the tall design it is a suffix of the row key, in the wide design it is the qualifier. The bytes per cell are close to identical.

Second, repetition is what data block encoding is for. Setting DATA_BLOCK_ENCODING => 'FAST_DIFF' (or PREFIX, DIFF, ROW_INDEX_V1) on the family lets HBase store each key as a delta against the previous one, so long shared prefixes cost little in either layout. If you chose wide rows to save disk, enable block encoding and measure instead.

The same readings for one device, stored two waysWide: one row per devicerow dev42d:t1 | d:t2 | d:t3 | ... | d:t3,153,600row dev43d:t1 | d:t2 | ...one region holds the whole row, foreverTall: one row per readingdev42#9998 | d:vdev42#9997 | d:vdev42#9996 | d:v... millions of short rowsregion boundaries can fall between any twoAtomic unitthe rowDistribution unitrows, split between keysRead unitGet = row, Scan = row rangeEvery cell carries its full row key in both layouts.
Both layouts store the same cells. The wide design keeps the item id in the qualifier, the tall design in the row key, and that single move changes what is atomic, what can split and what one read returns.

The three things the choice really changes

Since storage cost is roughly a wash, the decision is about three properties that HBase attaches to the row, and only to the row.

The unit of atomicity. A Put, Delete, Increment, Append or checkAndMutate on one row is atomic across all its families and columns. Nothing larger is atomic by default. If ten values must change together, a wide row gives you that for free; ten tall rows give you ten independent writes that can half-succeed.

The unit of distribution. A table is divided into regions by row-key ranges, and a region boundary can only fall between rows. A single row can never be split, however large it grows. With the default hbase.hregion.max.filesize of 10 GB, a store normally splits long before trouble, but a row that is itself several gigabytes forces its region to stay at least that big, on one RegionServer, taking all of that row's traffic.

The unit of reading. A Get returns one row, or the selected families and columns of one row; a Scan returns a range of rows. Fetching "everything about entity X" is a single Get in a wide design and a prefix scan in a tall one. The real difference is the ceiling: the server refuses to materialise a row larger than hbase.table.max.rowsize (1 GB by default) in a Get or a non-batched Scan, and throws RowTooBigException. Long before that, giant Results strain client memory and RPC handlers.

Advertisement

Worked example: sizing a sensor time series

Take a fleet of 50,000 devices, each reporting one 8-byte reading every 10 seconds: 8,640 readings per device per day, about 3.15 million per device per year. Device ids are 8 characters, the family is d, and we keep a year of data.

Estimate the per-cell size before encoding. Tall, with a key of dev00042# plus an 8-byte reversed timestamp (17 bytes), a 1-byte qualifier and an 8-byte value: 20 + 17 + 1 + 1 + 8 = about 47 bytes. Wide, one row per device, with the bare 8-byte device id as key and an 8-byte timestamp qualifier: 20 + 8 + 1 + 8 + 8 = about 45 bytes. The same, as promised. (The bucketed variant below, keyed dev00042#2026-10-01 with a 4-byte seconds-of-day qualifier, comes to about 52 bytes.) The interesting numbers are elsewhere:

PropertyWide, row per deviceBucketed wide, row per device-dayTall, row per reading
Cells per row after a year3.15 million8,6401
Raw row size after a yearabout 140 MBabout 450 KBabout 47 bytes
Can a heavy device spread out?No, one row, one regionYes, by dayYes, anywhere
Last hourGet with column range filterGet with column range filterScan, limit 360
Atomic multi-reading writeAny readings, everReadings in one dayOne reading
Rows in the table50,00018 million158 billion

The unbucketed wide row is the trap. It looks tidy at 50,000 rows, but each row grows without bound, and a busy device's region can never shed it. The tall design has the opposite profile: tiny rows, perfect splitting, and a row count large enough that you must think about hotspotting on the key prefix and about scan efficiency. The bucketed wide design caps the row at a known size and still lets a single Get return a whole day.

Reading each layout without hurting yourself

The tall layout reads with ordinary range scans. Putting the reversed timestamp in the key makes "latest first" the natural order, so the most common query, the newest N readings, is a scan that stops after N rows:

// Tall: row key = deviceId + reversed timestamp, one cell per reading.
static byte[] tallKey(String deviceId, long epochMillis) {
    return Bytes.add(Bytes.toBytes(deviceId), Bytes.toBytes("#"),
                     Bytes.toBytes(Long.MAX_VALUE - epochMillis));   // newest first
}

table.put(new Put(tallKey("dev42", ts)).addColumn(D, V, Bytes.toBytes(reading)));

// Latest 100 readings: a bounded scan that touches one region in the common case.
byte[] prefix = Bytes.toBytes("dev42#");
Scan latest = new Scan()
        .setRowPrefixFilter(prefix)      // sets start and stop rows from the prefix
        .setLimit(100)
        .setCaching(100);
try (ResultScanner rs = table.getScanner(latest)) {
    for (Result r : rs) { handle(r); }
}

Two scanner settings matter here. setCaching is the number of rows per RPC, and should match the limit for small bounded reads. hbase.client.scanner.max.result.size caps the bytes per RPC (2 MB by default in HBase 2), which protects both sides when rows are bigger than you expected. See the scans article for caching, timeouts and leases.

The wide layout reads with Gets restricted to columns. Never issue an unrestricted Get against a row whose size you do not control. Restrict with a column list, a ColumnRangeFilter, a ColumnPrefixFilter or a ColumnPaginationFilter, and when you must walk an entire large row, use a single-row Scan with setBatch so the server returns it in slices:

// Bucketed wide: row key = deviceId + day, qualifier = seconds-of-day, at most 8,640 cells per row.
static byte[] wideKey(String deviceId, LocalDate day) {
    return Bytes.add(Bytes.toBytes(deviceId), Bytes.toBytes("#"), Bytes.toBytes(day.toString()));
}
table.put(new Put(wideKey("dev42", day)).addColumn(D, Bytes.toBytes(secondOfDay), Bytes.toBytes(reading)));

// One hour of one day: a Get restricted to a qualifier range, never the whole row.
Get hour = new Get(wideKey("dev42", day))
        .setFilter(new ColumnRangeFilter(Bytes.toBytes(3600), true, Bytes.toBytes(7200), false));
Result r = table.get(hour);

// Walking a row that might be huge: stream it in slices instead of materialising it.
Scan oneRow = new Scan()
        .withStartRow(rowKey).withStopRow(rowKey, true)
        .setBatch(1000)                  // at most 1,000 cells per Result
        .setAllowPartialResults(true);   // let the server cut a row at the size limit too

Note what setBatch gives up: a Result no longer equals a row, so code that assumes "one Result, one entity" breaks silently. Callers must accumulate Results until the row key changes. Column filters also have a cost the filters article explains: the server still reads the blocks and skips cells, so a filter reduces network traffic far more than it reduces disk I/O.

The middle ground most systems end up using

Pure wide and pure tall are the endpoints. Two intermediate designs deserve to be the default for growing data.

Bucketed wide rows. Put a coarse bucket into the row key (device plus day, thread plus page) and keep fine-grained items as qualifiers. You choose the bucket so that the worst-case row stays under a few megabytes. You keep single-row atomicity within a bucket and single-Get reads of a bucket, and the table still splits along bucket boundaries. Cross-bucket reads become multi-row scans, and the bucket size is hard to change later, so size it against the worst entity.

Tall rows held together by a split policy. If you need atomic writes across several tall rows of one entity, HBase can give it to you, provided those rows live in one region. MultiRowMutationEndpoint applies a batch of mutations to multiple rows atomically when they share a region, and the prefix split policies guarantee they do by refusing to split inside a key prefix:

// Keep every row of one tenant in one region so MultiRowMutationEndpoint can apply
// several tall rows atomically. Row keys look like  tenant0042|order|...
TableDescriptor td = TableDescriptorBuilder.newBuilder(TableName.valueOf("orders"))
    .setColumnFamily(ColumnFamilyDescriptorBuilder.of("d"))
    .setRegionSplitPolicyClassName(DelimitedKeyPrefixRegionSplitPolicy.class.getName())
    .setValue("DelimitedKeyPrefixRegionSplitPolicy.delimiter", "|")
    .setCoprocessor(MultiRowMutationEndpoint.class.getName())
    .build();

KeyPrefixRegionSplitPolicy does the same with a fixed prefix length. This keeps the tall layout's small rows and natural paging while restoring a scoped form of atomicity. The price is the same floor that wide rows have, at the prefix level: one tenant can never spread over more than one region, so this suits bounded groups such as an order and its lines, not a tenant with unbounded history. How regions split and how split policies are chosen is covered in regions and splits.

Failure modes

  • The unsplittable hot region. A wide row that keeps growing pins its region to one server. Symptoms: one RegionServer with outsized request counts, a region far above the split size that never splits, and the balancer unable to help. Only re-keying fixes it.
  • RowTooBigException or client OOM. A Get on a row that crossed hbase.table.max.rowsize, or a smaller row that still exceeds the client heap. Raising the limit only hides the problem.
  • Tombstone build-up inside wide rows. Deleting old columns from a long-lived row writes delete markers that every subsequent read of that row must skip until a major compaction removes them. A queue modelled as a wide row with columns added and deleted continually gets slower with every message. TTLs on the family are cheaper than explicit deletes for aging data out.
  • Monotonic tall keys. Tall rows keyed by timestamp alone send every write to the last region. Lead with an entity id or a salt.
  • Half-applied multi-row writes. A tall design that needs several rows to change together will eventually observe one changed and one not. Either keep them in one region with the pattern above, or write in an order that leaves recoverable states and make every write idempotent.

Choosing, and changing your mind

ChooseWhen
WideThe column set is bounded and known (a profile, a few dozen counters), the values are read together, and they must change atomically.
Bucketed wideItems grow without bound but are read in natural groups (a day, a page), and you want one Get per group.
TallItems grow without bound, are read by range or individually, and per-item atomicity is enough.
Tall with a prefix split policyItems of a bounded group must change atomically together, and the group will never outgrow a region.

Migrating between layouts is a full rewrite, because the row key changes. The safe sequence is the usual one for any HBase re-key: create the new table with its own pre-splits; dual-write from the application with the new key alongside the old; backfill history with a MapReduce or Spark job that reads a snapshot of the old table and bulk-loads HFiles into the new one; compare per-entity counts; switch reads; stop old writes. Sample the worst entities, not random ones.

What to do next

  1. List every read pattern for the table and mark which ones need several values atomically; that list, not storage cost, decides the layout.
  2. For each candidate row, write down the worst-case cell count after the retention period, using your heaviest entity, not your average one.
  3. If any row can exceed a few megabytes, bucket it, and record the bucket size in the schema document.
  4. Enable FAST_DIFF or another data block encoding on the family and measure real bytes per cell before arguing about key repetition.
  5. Audit client code for unrestricted Gets and for any assumption that one Result is one entity; add column restriction and setBatch where rows are unbounded.
  6. Add an alert on regions far above hbase.hregion.max.filesize that have not split; that is the signature of a wide row out of control.
Key takeaway: Wide and tall tables store nearly the same bytes, because every HBase cell carries its full row key either way. What the choice changes is the row's three roles: the boundary of atomicity, the unit that can never split, and what one Get returns. Use wide rows for bounded sets that change together, tall rows for anything that grows, bucketed wide rows or prefix-pinned tall rows in between, and size every row against your heaviest entity.