A region split looks like one event in the master UI: a region disappears and two take its place. Underneath it is a pipeline that spans two processes and several minutes to days. The RegionServer notices the region is too big and asks for a split. The master runs a durable procedure that closes the parent, writes small pointer files for the daughters, rewrites hbase:meta and opens the daughters. The daughters serve reads through those pointers until compactions rewrite the data. Only then does a background chore delete the parent. Each hand-off has its own failure behaviour, and most split problems in production are one of those hand-offs stalling.
Which policy triggers a split and how the split key is chosen are covered in the regions and splits article. This article follows the machinery from trigger to cleanup, with the configuration keys and defaults read from the Apache HBase source on the master branch, so it can be checked against the version you run. Where a behaviour is recent, it says so.
The lifecycle at a glance
Five phases, owned by three components. The RegionServer evaluates the split policy and sends a request. The master runs SplitTableRegionProcedure, a Procedure v2 state machine persisted in the procedure store, so it survives a master crash. The daughter regions, usually on the same RegionServer at first, open with reference files and compact them away. The master's CatalogJanitor removes the parent once nothing refers to it. A split is cheap at the moment it happens because no data moves; the cost is paid later, in the daughters' compactions.
Phase 1: the RegionServer decides
After a flush or compaction finishes, the RegionServer's CompactSplit component asks whether the region should split. The first check is a server-wide cap: hbase.regionserver.regionSplitLimit, default 1000 online regions. Above 90 percent of it the server logs a warning, and at the limit it stops requesting splits at all, so regions keep growing silently. The second check is the split policy. The default is SteppingSplitPolicy: if exactly one region of this table lives on this server, the threshold is twice the memstore flush size, 256 MB with the 128 MB default; otherwise it is hbase.hregion.max.filesize, default 10 GB, perturbed by a random jitter of up to plus or minus 12.5 percent (the 0.25 jitter setting is the width of that band) so that regions filled at the same rate do not all split at once. On the master branch hbase.hregion.split.overallfiles defaults to true, which compares the sum of all stores against the threshold rather than the largest single store.
A region cannot split if any store still holds reference files from a previous split, and it cannot split if it is a meta region. If the policy says yes, the server computes a split point from the largest store and hands a SplitRequest to a small thread pool (hbase.regionserver.thread.split, default 1). The request goes to the master, which re-validates everything, because the master, not the RegionServer, owns region state.
Phase 2: the master's procedure and its point of no return
The master first runs its preparation checks. It refuses to split while the table is taking a snapshot or while a table-modification procedure is running, and it re-checks both split switches, the cluster-wide one and the table's own SPLIT_ENABLED attribute, so a switch turned off after submission still cancels the split. The parent must be OPEN or CLOSED. It asks the hosting RegionServer whether the region is splittable and, if no key was given, for the best split row, then rejects a row equal to the region's start key or outside its range. Then the states run in this order:
| Order | State | What happens | On failure |
|---|---|---|---|
| 1 | PREPARE | checks above; parent marked SPLITTING | rolled back |
| 2 | PRE_OPERATION | coprocessor pre-split hooks | rolled back |
| 3 | CLOSE_PARENT_REGION | parent closed and flushed | rolled back |
| 4 | CHECK_CLOSED_REGIONS | confirm the parent is really closed | rolled back |
| 5 | CREATE_DAUGHTER_REGIONS | reference files and daughter directories written | rolled back |
| 6 | WRITE_MAX_SEQUENCE_ID_FILE | daughters get a sequence id floor | rolled back |
| 7 | PRE_OPERATION_BEFORE_META | coprocessor hooks | rolled back |
| 8 | UPDATE_META | parent marked split, daughters added in hbase:meta | no rollback: rolls forward |
| 9 | PRE_OPERATION_AFTER_META | coprocessor hooks | rolls forward |
| 10 | OPEN_CHILD_REGIONS | daughters assigned and opened | rolls forward |
| 11 | POST_OPERATION | post-split hooks | rolls forward |
The line between states 7 and 8 is the split's point of no return. Before the meta update, any failure rolls the procedure back and the parent comes back online as if nothing happened. From the meta update onwards, rollback is unsupported: the daughters exist in the catalogue, so a crashed master's successor replays the procedure forward until they open. This is why a stuck split never leaves data half-owned, and why the fix for a split wedged after state 8 is to get the daughters open, not to resurrect the parent. The Procedure v2 article explains the executor that makes this replay possible, and the hbase:meta article the rows being rewritten.
Phase 3: reference files and links
Creating daughters writes pointers, not data. For every store file in the parent, the master decides per daughter. A file whose keys lie entirely outside a daughter's range gets nothing in that daughter. A file that straddles the split row gets a reference file: a tiny file holding the split row and a half, bottom for keys below the split and top for keys at or above it, named after the original HFile with the parent's encoded region name appended. On the current master branch, a file that lies wholly inside one daughter gets an HFileLink instead, a pointer to the whole file with no half-file filtering.
The work runs in a thread pool sized to the smaller of hbase.regionserver.region.split.threads.max (falling back to the blocking store file count) and the number of files, and must finish within hbase.master.fileSplitTimeout, which falls back to hbase.regionserver.fileSplitTimeout, default 600,000 ms. It is idempotent: if a reference file already exists after a crash, the code logs that it is assuming a recovery and reuses it. The parent's HFiles are not touched; they stay in the parent's directory, and both daughters read them.
Phase 4: daughters serve, then compact
A daughter opens with its references and serves immediately. Reads through a reference use a half-file reader that filters keys against the split row, so a scan in a daughter touches blocks of the parent's file, including blocks that belong to the other daughter. It works, but it is slower than reading a file of your own, and bloom filters and block indexes still describe the whole parent file.
Two rules follow. First, a store with references cannot split, so a daughter that grows quickly cannot split again until its references are gone. Second, the daughter's compactions select reference files and rewrite them into new, real HFiles containing only the daughter's keys. That rewrite is the true cost of a split: every byte of the parent is read and written once more, across the two daughters, on top of the normal compaction schedule described in the compaction article. A cluster that splits thousands of regions at once feels it as a compaction storm and a jump in disk and network I/O.
Phase 5: CatalogJanitor retires the parent
The parent stays in hbase:meta, flagged as split and offline, because its files are still in use. The master's CatalogJanitor chore runs every hbase.catalogjanitor.interval, default 300,000 ms. For each split parent it checks both daughters; only when neither holds references does it submit a GCRegionProcedure that archives the parent's files and removes its meta row. If either daughter still has references it logs that removal is deferred and tries again next time. The parent's meta row records its daughters in the splitA and splitB columns, which is how the janitor finds them, and a parent that is itself the product of a merge waits until its merge parents are cleaned first. The practical consequence is that disk usage does not fall when a split finishes: the parent's files remain until both daughters have compacted, so a burst of splits temporarily holds up to two copies of the affected data in HDFS. Archived files can still be held by snapshots, which is expected and handled by the archive cleaner, not by the janitor.
Worked example: a new table under load
A table is created with one region and receives writes at 50 MB/s spread evenly across its key range. At about 256 MB, roughly five seconds of data, the single region splits. Now two regions of the table are on that server, so the threshold jumps to 10 GB. As the balancer spreads daughters around, a server that again holds exactly one region of the table applies the 256 MB threshold to it once more, which is how a new table spreads across the cluster quickly.
In steady state a region reaches about 10 GB, splits into two daughters of about 5 GB, and each daughter grows another 5 GB before it splits in turn. At 50 MB/s the table gains about 4.3 TB per day, so it performs roughly one split per 5 GB ingested, about 860 splits a day, and each makes its daughters rewrite about 10 GB of parent data. Split-driven rewrites therefore run at roughly twice the ingest rate, before counting normal compaction. Pre-splitting the table at creation, with split keys matched to the row key distribution, avoids the early cascade entirely; raising the max file size reduces how often it happens later but makes each split and each compaction bigger.
Operating splits
The shell and the Admin API give you manual control:
# pre-split at creation, keys matching the row key distribution
create 'events', 'd', SPLITS => ['1', '2', '3', '4', '5', '6', '7', '8', '9']
# split every region of a table at its computed midpoint, a table at a key, or one region
split 'events'
split 'events', '5a'
split 'REGION_NAME'
# stop all splits during maintenance, then check and restore
splitormerge_switch 'SPLIT', false
splitormerge_enabled 'SPLIT'
splitormerge_switch 'SPLIT', true
# disable splitting for a single table
alter 'events', SPLIT_ENABLED => falsetry (Admin admin = conn.getAdmin()) {
admin.split(TableName.valueOf("events"), Bytes.toBytes("5a"));
admin.splitRegionAsync(regionName).get(10, TimeUnit.MINUTES);
boolean previous = admin.splitSwitch(false, true); // synchronous
}Watch four signals. On RegionServers, splitQueueLength and splitRequestCount show pending and requested splits. On the master, the assignment manager's split operation metrics count submitted and failed split procedures and their time, and ritCount with ritOldestAge show regions stuck in transition. In the logs, the master's line reporting how many store files it is splitting, and the janitor's deferral messages, tell you which phase a slow split is in.
Failure modes
- Silent split cap. A server at regionSplitLimit stops requesting splits and regions grow far past the max file size; watch the 90 percent warning.
- Switch left off. A maintenance splitormerge_switch never restored; check it in every runbook's exit step.
- Split refused during snapshots. Frequent snapshot schedules can make splits fail repeatedly; they retry on the next trigger, but a big table can outgrow its threshold meanwhile.
- References that never compact. Compactions disabled or starved leave daughters unable to split and parents never cleaned; region directories and meta rows accumulate.
- Unsplittable hot row. A single row or tight key range cannot be divided; splitting moves nothing and the hotspot remains a row key design problem.
- Stuck after the meta update. The procedure must roll forward; fix what prevents the daughters opening rather than trying to restore the parent.
What to do next
- Confirm your split policy, max file size, flush size and whether overallfiles applies in your version.
- Pre-split new large tables with keys matched to the row key distribution.
- Alert on splitQueueLength, failed split procedures, ritOldestAge and servers near regionSplitLimit.
- Keep compaction healthy so daughters shed references and parents are collected; watch janitor deferrals.
- Schedule snapshots and table alterations away from peak split activity.
- Add restoring the split switch to every maintenance runbook, and verify it with splitormerge_enabled.
- Reproduce one split on a test table and follow it in the master log from the store-file count line to the janitor's GCRegionProcedure, so you recognise each phase when one stalls in production.