HBase lets a row hold any number of columns, and many schemas take advantage of it: one row per user with a column per message, one row per device with a column per reading. That freedom has no single hard ceiling, which is exactly why it causes outages. A wide row grows quietly for months, passing a series of soft limits, each with its own symptom, until a Get throws RowTooBigException, a RegionServer runs out of heap, or one region refuses to split and takes a whole server's traffic.

This article lists those limits in the order a growing row meets them, with the configuration property and default for each, explains why each exists, and shows how to read a wide row without tripping them, how to find rows that are approaching them before they fail, and how to re-key a table so they never come back. The choice between wide and tall designs is covered in HBase wide vs tall tables; this article assumes you already have wide rows, possibly too wide.

Where a wide row meets each limit

Client Putcell <= keyvalue.maxsize 10 MBRow lock + memstoreflush 128 MB, block at 4xHFiles in one regionrow never split across regionsRegion sizemax.filesize 10 GB; row stays wholeGet / unbatched Scanrow > max.rowsize 1 GB: RowTooBigScanner RPC sizingclient 2 MB, server 100 MBWhole row returnedunless partial results allowedClient heapResult holds every cell in memoryFixessetBatch, partial results, bucketingEach box is a separate ceiling with its own symptom; a growing row meets them in turn.
A growing row passes cell-size, memstore, read-size and region-size limits in turn; each has its own property and symptom, and none of them splits the row for you.

Why rows have limits at all

The row is special in HBase because four properties attach to it and only to it.

  • Atomicity. A single-row mutation (Put, Delete, Increment, Append, checkAndMutate, mutateRow) is atomic across all its columns and families. To guarantee that, the RegionServer takes a lock on the row for each write.
  • Distribution. Regions are ranges of row keys, and a region boundary can only fall between rows. A row always lives in exactly one region on one server.
  • Reads. A Get returns a row, and by default a Scan returns whole rows, so the server assembles the selected cells of the row before sending them.
  • Storage. Every cell stores the full row key alongside its family, qualifier and timestamp. A long key on a wide row is repeated once per cell.

Every wide-row limit is one of these properties meeting a resource bound: memory, RPC size, region size or lock throughput.

The limits, with defaults

The defaults below are from hbase-default.xml in the HBase source tree. Check your version and your site overrides; distributions sometimes change them.

PropertyDefaultGuardsWhat you see when a wide row hits it
hbase.client.keyvalue.maxsize10485760 (10 MB)Single cell size at the clientPut rejected before it is sent
hbase.server.keyvalue.maxsize10485760 (10 MB)Single cell size at the serverPut rejected by the RegionServer
hbase.hregion.memstore.flush.size134217728 (128 MB)Memstore size before flushFrequent flushes on the region holding the hot row
hbase.hregion.memstore.block.multiplier4Memstore ceiling (4 x flush size)Writes to the region blocked until flush catches up
hbase.hregion.max.filesize10737418240 (10 GB)Region split thresholdRegion grows past it and cannot split inside the row
hbase.table.max.rowsize1073741824 (1 GB)Row materialised by a Get or unbatched ScanRowTooBigException at the client
hbase.client.scanner.max.result.size2097152 (2 MB)Bytes per scanner RPC requested by the clientMany RPCs per wide row
hbase.server.scanner.max.result.size104857600 (100 MB)Bytes per scanner call at the serverIgnored for a single row unless partial results are allowed
hbase.client.scanner.timeout.period60000 msScanner lease and RPC waitTimeouts while the server assembles a huge row

The note in hbase-default.xml on the server scanner limit deserves emphasis: when a single row is larger than the limit, the row is still returned completely. The size limits protect you from many rows, not from one big one. Only hbase.table.max.rowsize or partial results stop a single giant row, and the first does it by failing the read.

Write side: cells, locks and memstores

Writes meet the limits first, but quietly. A cell over 10 MB is rejected outright, which is a value-size problem rather than a width problem; large values belong in MOB columns or an object store with a pointer in HBase. The more important write-side limit is throughput. All writes to one row take that row's lock in turn, and all writes to the row's region share one memstore. A wide row that is also a hot row (a global counter, a popular user's inbox, a busy device) serialises its writers on the lock, and if it carries enough of the region's traffic it drives frequent flushes. When the memstore reaches four times the flush size, the server blocks updates to the whole region, so neighbouring rows suffer too.

Increments and checkAndMutate are the worst case because each is a read-modify-write under the row lock. A design where every event increments a column in one row per tenant caps that tenant's write rate at whatever one lock can sustain, regardless of cluster size.

Regions cannot split a row

Regions split by size, and the split point is chosen between rows. A table of reasonably sized rows splits smoothly as it grows. A region dominated by one enormous row cannot be split usefully: whatever the policy computes, the row stays whole on one side. The region keeps growing past hbase.hregion.max.filesize, compactions on it rewrite ever larger files, and every read and write for that row goes to one RegionServer. Adding servers does not help, and the balancer can only move the problem, not divide it.

In practice the read limits below usually bite before a row approaches 10 GB. But a table with many rows of several hundred megabytes already makes splits coarse and uneven, so region sizing and hotspot analysis both get harder long before any exception appears.

Read side: RowTooBigException and how to page a row

A Get, or a Scan without batching, makes the RegionServer collect the requested cells of the row before returning them. The server checks the accumulated size against hbase.table.max.rowsize and throws RowTooBigException once it is crossed. That limit is a guard for the server, and 1 GB is far too much for any sane client anyway: the whole Result is held in client heap, often several times over while it is decoded.

The fixes all stop treating the row as one unit. Ask only for what you need, by family and column, and page through the rest:

// 1. Scan one row in slices: at most 1,000 cells per Result.
Scan scan = new Scan()
    .withStartRow(rowKey)
    .withStopRow(rowKey, true)            // inclusive: exactly this row
    .addFamily(Bytes.toBytes("m"))
    .setBatch(1000)                       // cells per Result
    .setAllowPartialResults(true)         // let size limits cut the row too
    .setMaxResultSize(4L * 1024 * 1024);  // bytes per RPC
try (ResultScanner rs = table.getScanner(scan)) {
    for (Result slice : rs) {
        handle(slice);                    // a slice, not the whole row
        // slice.mayHaveMoreCellsInRow() is true until the last slice
    }
}

// 2. Read one page of columns with a Get: 100 columns from a cursor.
//    The qualifier offset is inclusive: pass the next qualifier, or drop the first cell.
Get get = new Get(rowKey)
    .addFamily(Bytes.toBytes("m"))
    .setFilter(new ColumnPaginationFilter(100, pageStartQualifier));
Result page = table.get(get);

// 3. Read a column range (qualifiers are sorted bytes).
Get recent = new Get(rowKey).setFilter(
    new ColumnRangeFilter(fromQualifier, true, toQualifier, false));

With setBatch or partial results, one Result is no longer one row, so code that assumes "one Result per entity" must be changed to group slices by row key. ColumnPaginationFilter with a qualifier cursor seeks directly to the offset; the integer-offset form has to skip every earlier column on each call, which gets slower as pages go deeper. Design qualifiers so that the reads you need are ranges, for example a reversed timestamp so that "newest 100" is the first 100 columns.

Worked example: a messaging inbox

A messaging service stores one row per user in family m, one column per message. Row key: an 8-byte user id. Qualifier: an 8-byte reversed timestamp. Values average 600 bytes. Each cell in the KeyValue format carries a 4-byte key length and 4-byte value length, then a 2-byte row length, the 8-byte row, a 1-byte family length, the 1-byte family, the 8-byte qualifier, an 8-byte timestamp and a 1-byte type, then the value: 37 bytes of overhead plus 600, about 637 bytes per message before compression.

A heavy user (an automated account) receives 2,000 messages a day, about 1.27 MB per day. The row passes 100 MB in about 80 days, at which point an unbatched read of it already holds 100 MB in one Result. It reaches the 1 GB hbase.table.max.rowsize at 1,073,741,824 / 1,274,000, about 843 days, roughly 2.3 years. The 10 GB region size is decades away, so this table fails on reads, not on splits. Meanwhile a scan with default settings returns about 2,097,152 / 637, roughly 3,300 messages per RPC, so reading the row page by page is fine, while reading it whole is not.

Re-key by month: userId + yyyymm. A row then holds at most 31 x 2,000 = 62,000 messages, about 39.5 MB, for the heaviest user. "Latest messages" becomes a Get on the current month's row; history is a prefix scan over the user's month rows. The heavy user also stops being a single-row hotspot for older data, since past months are read-only. Combine with a column-family TTL if old messages expire, rather than deleting cells and leaving tombstones for every read of the row to skip.

Finding wide rows before they fail

Find wide rows before they find you. Four sources, from cheapest to most thorough:

  • Exceptions and slow logs. Alert on RowTooBigException in client logs and on large responses in the RegionServer's slow log; both name the table and are early warning if you catch the first one.
  • Region metrics. A region far above the others in size or request count, or one that stays above the split threshold, often contains a giant row.
  • HFile statistics. The HFile pretty-printer reports per-file statistics, including the key of the biggest row, with hbase hfile -s -f <path-to-hfile>. Run it against the largest store files of suspicious regions; it reads the file directly, so run it off-peak.
  • A full row census. A MapReduce or Spark job over a table snapshot that emits row key, cell count and byte total per row, then the top thousand. Reading a snapshot keeps the load off the RegionServers.
# sketch: per-row size census from a snapshot export (PySpark, pseudocode for the reader)
cells = read_hbase_snapshot("messages_snap")       # rows of (row, family, qualifier, value)
census = (cells
    .groupBy("row")
    .agg(F.count("*").alias("cells"),
         F.sum(F.length("value") + F.length("row") + 37).alias("bytes"))
    .orderBy(F.desc("bytes")))
census.limit(1000).write.csv("hdfs:///reports/widest_rows")

The reader call is a placeholder for whichever snapshot input format your stack uses; the aggregation is the point. Track the top row's size over time; the slope tells you how long you have.

Failure modes

  • Raising max.rowsize. It makes the exception go away and moves the failure into client heap or GC pauses. Fix the read or the key instead.
  • Batching that breaks grouping. Adding setBatch to an existing job turns one Result per row into many; aggregations that assumed one Result per entity double-count or truncate.
  • Tombstone build-up. Queues and inboxes that delete old columns from a long-lived row leave delete markers every read must skip until major compaction.
  • Unbounded versions. A family with many versions retained multiplies width; set VERSIONS to what you read.
  • Long row keys. Repeated in every cell; data block encoding such as FAST_DIFF reduces the cost on disk, but keys still dominate memory for small values.
  • Hot-row lock contention. Increments on one row per tenant cap throughput no matter how many servers you add.

Trade-offs

Bucketing trades single-Get convenience for bounded rows: a time bucket keeps recent reads to one Get, a hash bucket spreads writes but makes "read everything" a multi-row scan. Bucket size is the dial: small buckets mean more rows and more Gets, large buckets bring back the limits. The atomicity of one row is lost across buckets, so anything that must change together has to share a bucket. Re-keying is a full rewrite (new table, dual writes, snapshot backfill, verification, cutover), so choose a bucket that leaves a factor of ten in headroom over the heaviest entity you can foresee.

What to do next

  1. Look up the effective values of the limits in the table above on your cluster, including site overrides.
  2. Alert on RowTooBigException and on regions whose size stays above the split threshold.
  3. Run hbase hfile -s on the largest HFiles of your biggest regions, or a snapshot census job, and record the widest rows and their growth rate.
  4. Change every read of a potentially wide row to request specific families and columns and to page with setBatch, ColumnPaginationFilter or column ranges.
  5. Compute the time until your heaviest entity crosses 100 MB and 1 GB, as in the worked example.
  6. If it is under two years, re-key with a time or hash bucket and plan the migration.
  7. Further reading: HBase scans, in depth, schema design, hotspotting and regions and splits.
Key takeaway: A wide HBase row has no single hard limit; it passes a series of soft ones. Cells over 10 MB are rejected, writes to one row serialise on its lock and a hot row can block its whole region's memstore, a region can never split inside a row, and a Get or unbatched Scan fails with RowTooBigException once a row passes hbase.table.max.rowsize, 1 GB by default, while the server's scanner size limit never cuts a single row unless partial results are allowed. Read wide rows by column and in pages with setBatch, partial results or ColumnPaginationFilter, find the widest rows with HFile statistics or a snapshot census, and re-key with a time or hash bucket before the heaviest entity runs out of headroom.