In HBase the column family is the unit of physical storage. Rows and qualifiers are free-form, you can add a new qualifier on any write, but families are declared in the table schema and each one gets its own files on disk, its own memory buffer and its own tuning. That makes the family decision one of the few places in HBase schema design where you choose, up front, how data is laid out, compressed, cached and expired.

It is also the decision most often made badly, usually by treating families like tables in a relational database and creating one per entity or per column group. This article explains what a family is physically, what families share and what they isolate, the specific reasons that justify a second family, how to configure each family's attributes, and a worked schema with the shell and Java code to build and query it. Row key design, the other half of schema design, is out of scope here.

Advertisement

What a column family is on disk

A table is split by row key into regions, and each region is served by one RegionServer. Inside a region, every column family is a separate store. Each store has its own MemStore, the in-memory sorted buffer that receives writes, and its own set of HFiles, the immutable sorted files written when the MemStore is flushed. Compactions run per store and rewrite only that family's files.

Every cell is stored as a key-value whose key contains the row, the family, the qualifier, the timestamp and a type. The family name is written into every single cell in every HFile, which is why the standard advice is to use one-character family names such as d or m. With billions of small cells, a family called metadata instead of m is seven wasted bytes per cell before compression and a real cost in block cache capacity.

Table 'docs' with two column families: one region, two independent storesRegion [row a .. row m)shared key range and splitStore 'm' (metadata)small cells, read by list scansStore 'b' (body)large cells, read on openMemStore mHFiles m: FAST_DIFF, 16 KB blocksMemStore bHFiles b: compressed, 128 KB blocksPer family: own MemStore, own HFiles, own compactions, own VERSIONS / TTL / encoding / bloom / block sizeShared by families: row key, region boundaries, splits, the WAL, and (for some triggers) flushesA scan with addFamily('m') opens only store m's files; the body bytes are never read
A two-family table. Each family is a separate store within every region, with its own MemStore, HFiles and settings; the families share the region boundaries, the WAL and some flush triggers.

What families isolate and what they share

Families isolate storage and configuration. A read that names only one family opens only that family's store. Compression, block encoding, bloom filters, block size, caching, version count and time to live are all set per family. Compactions of one family do not rewrite the other's files.

Families share almost everything structural. They share the row key and the region boundaries, so when a region splits, every family in it is split at the same row. The split point is chosen from the largest store, so a small family is carried along by the growth of a large one: if one family holds 99 percent of the bytes, the small family ends up scattered thinly across many regions, and a scan over it touches far more regions and files than its size would suggest. Families also share the region's write-ahead log, so a single WAL entry covers a row mutation across families.

Flushing is the classic coupling. Historically a region flushed all its stores together, so a lightly written family produced a stream of tiny HFiles every time its busy neighbour filled up. HBase added a flush policy, FlushAllLargeStoresPolicy, that flushes only stores above a lower bound: the region's flush size divided by the number of families, but never less than hbase.hregion.percolumnfamilyflush.size.lower.bound.min, 16 MB by default. If no store is above the bound, all are flushed. Some triggers, such as too many WAL files, still force whole-region flushes. MemStore flush mechanics covers each trigger; the design consequence is that families with very different write rates remain a problem that configuration only softens.

Advertisement

The default answer is one family

Because of those couplings, the HBase reference guide's long-standing advice is to keep the number of families small, one where possible and rarely more than two or three. One family gives you one MemStore per region, one set of HFiles, one compaction queue and no cross-family flush or split distortion. Most tables need nothing more: qualifiers already give you as many columns as you want, and filters or column selection already restrict what a read returns.

The question to ask is not whether data is logically different, but whether it needs to be physically different. If two groups of columns are read together, written together and expire together, they belong in one family.

Four reasons that justify a second family

  1. Different read paths. A frequent query needs a small subset of a row's bytes, and the rest is large. Listing documents needs titles and timestamps, not bodies. With one family, a scan reads every body block; with two, it opens only the small store.
  2. Different retention. Part of the row must expire after a period or keep several versions while the rest keeps one. VERSIONS and TTL are per family, so different retention requires different families. See TTL and versions for how expiry is actually reclaimed.
  3. Different storage format. Small, repetitive cells benefit from prefix-style block encoding and small blocks for point reads; large, opaque blobs benefit from strong compression and large blocks and gain nothing from encoding.
  4. Different caching. Hot metadata should stay in the block cache, while large, rarely re-read bodies should not evict it. Cache participation is set per family.

Each of these is about physical behaviour. Notice also what is not on the list: different write rates. A family that is written far more often than another is a reason for a separate table, not a separate family, because the flush and split couplings above are exactly what goes wrong.

Per-family settings and how to choose them

AttributeWhat it controlsGuidance
VERSIONS / MIN_VERSIONSHow many timestamped versions of a cell are kept1 unless you read history; bound it with TTL
TTLAge in seconds after which cells stop being returnedSet for any data with a retention policy
COMPRESSIONCodec applied to HFile blocksA fast codec by default; a stronger codec for cold, large values
DATA_BLOCK_ENCODINGEncoding of keys within a blockFAST_DIFF or PREFIX for many small cells with long shared keys; skip for blobs
BLOOMFILTERPer-file filter to skip HFiles on point readsROW for Gets by row; ROWCOL only if Gets name specific columns in wide rows
BLOCKSIZETarget size of an HFile data blockSmaller for random point reads, larger for scans and big values
BLOCKCACHE / IN_MEMORYWhether blocks are cached, and with what priorityDisable caching for large, rarely re-read families
IS_MOBStore large values outside normal HFilesConsider for values from roughly 100 KB into the megabytes

Bloom filter trade-offs are covered in HBase bloom filters, and cache priorities in the block cache article. For large values, HBase MOB often beats a separate family, because it keeps big values out of the compaction path entirely.

Worked example: a document store

A service stores documents keyed by a hashed document id. Users list documents in a folder, which needs the title, owner, size and modified time of perhaps a hundred rows; and they open one document, which needs the metadata plus the rendered body, averaging 40 KB. Bodies keep three revisions for 90 days; metadata keeps only the latest. Both are written in the same mutation when a document is saved, so the families see the same number of writes, though very different byte volumes.

Every justification applies: a different read path (listing), different retention (versions and TTL), different formats (small repetitive cells against large compressible HTML) and different caching. So the schema uses two families, m and b:

create 'docs',
  {NAME => 'm', VERSIONS => 1, BLOOMFILTER => 'ROW',
   DATA_BLOCK_ENCODING => 'FAST_DIFF', COMPRESSION => 'SNAPPY', BLOCKSIZE => 16384},
  {NAME => 'b', VERSIONS => 3, TTL => 7776000,
   COMPRESSION => 'ZSTD', BLOCKSIZE => 131072, BLOCKCACHE => false},
  {SPLITS => ['2', '4', '6', '8', 'a', 'c', 'e']}

describe 'docs'

The m family uses 16 KB blocks and FAST_DIFF encoding because listing reads many small, adjacent cells, and ROW bloom filters because opening a document is a Get by row. The b family uses 128 KB blocks because a body read wants the whole value, stronger compression because HTML compresses well, and no block cache because a body is rarely re-read soon enough to be worth evicting metadata. Codec availability depends on your build, so check that ZSTD is installed on every RegionServer before creating the table. The same definition in Java, with the two read paths:

TableDescriptor docs = TableDescriptorBuilder.newBuilder(TableName.valueOf("docs"))
    .setColumnFamily(ColumnFamilyDescriptorBuilder.newBuilder(Bytes.toBytes("m"))
        .setMaxVersions(1)
        .setBloomFilterType(BloomType.ROW)
        .setDataBlockEncoding(DataBlockEncoding.FAST_DIFF)
        .setCompressionType(Compression.Algorithm.SNAPPY)
        .setBlocksize(16 * 1024)
        .build())
    .setColumnFamily(ColumnFamilyDescriptorBuilder.newBuilder(Bytes.toBytes("b"))
        .setMaxVersions(3)
        .setTimeToLive(90 * 24 * 3600)          // seconds
        .setCompressionType(Compression.Algorithm.ZSTD)
        .setBlocksize(128 * 1024)
        .setBlockCacheEnabled(false)
        .build())
    .build();

// Listing documents: touch only the metadata store.
Scan listing = new Scan()
    .setRowPrefixFilter(prefix)
    .addFamily(Bytes.toBytes("m"))
    .setCaching(500);

// Opening one document: read both families for one row.
Get open = new Get(rowKey)
    .addFamily(Bytes.toBytes("m"))
    .addColumn(Bytes.toBytes("b"), Bytes.toBytes("html"));

Now the arithmetic. A listing of 100 rows reads about 100 small metadata cell groups, a few blocks of m. With a single family, the same listing would read through about 4 MB of bodies in the same blocks. Separation makes listing roughly two orders of magnitude cheaper in I/O. The price is the byte imbalance: b reaches the flush size far sooner than m, so under the per-family flush policy m stays in memory until it passes its lower bound or a whole-region trigger fires, and region splits are driven by the size of b. If bodies grew to hundreds of kilobytes, the better design would be one family with bodies in MOB, or bodies in a separate table.

Changing families on a live table

Family attributes can be changed online with alter. Because HFiles are immutable, a change such as a new block size or codec applies to newly written files and reaches existing data only when compaction rewrites it. Adding a family is cheap; removing one deletes its data. Families cannot be renamed, so renaming means creating a new family, copying data and dropping the old one.

# Change a family attribute online. New HFiles use it at once;
# existing files are rewritten by the next major compaction.
alter 'docs', {NAME => 'm', BLOCKSIZE => 8192}
major_compact 'docs', 'm'

# Add a family online (empty until written).
alter 'docs', {NAME => 't', VERSIONS => 1, TTL => 86400}

# Drop a family: its data is deleted. There is no rename.
alter 'docs', 'delete' => 't'

Schedule the major compaction off-peak and one table at a time, because it rewrites every file of the family and competes with foreground I/O. HBase compaction explains how to throttle it.

Failure modes

  • Families as entity types. A family per logical entity multiplies stores, flushes and small files with no physical benefit. Use qualifiers or separate tables.
  • Hot and cold families together. An audit-trail family written on every event next to a profile family written monthly gives endless small flushes of the cold one and distorted splits. Move the hot data to its own table.
  • Long family names. Repeated in every cell; multi-character names cost storage and cache space at scale.
  • Unbounded versions. A high VERSIONS without TTL keeps history forever and inflates Gets that read it.
  • Reading every family by default. A Get or Scan with no family selection reads all stores, which silently cancels the benefit of separating them. Always name the family.
  • Expecting alter to act immediately. Compression or encoding changes take effect on disk only after compaction.

What to do next

  1. List every read path and retention rule for the table before creating any family.
  2. Start with one family; add a second only for a different read path, retention, format or caching need, never for a different write rate.
  3. Use one-character family names.
  4. Set VERSIONS, TTL, encoding, compression, bloom filter and block size per family from the table above, and check codec availability on every RegionServer.
  5. Make every Get and Scan name the families it needs, and check with a request-count metric that listing scans stay on the small store.
  6. After any attribute change, run a major compaction off-peak and verify with describe and store file sizes that it took effect.
Key takeaway: A column family is a physical store inside every region, with its own MemStore, HFiles, compactions and settings, and it shares row keys, region boundaries, splits and the WAL with every other family in the table. Default to one family with a one-character name. Add a second only when data needs a different read path, retention, format or caching, and put data with a very different write rate in a separate table instead. Configure each family deliberately, name the family in every read, and remember that attribute changes reach disk only through compaction.