HBase clusters tend to live for years, and the upgrade is the operation teams put off longest. The fear is justified. An HBase cluster is a distributed state machine whose state lives in several places at once: HFiles and write-ahead logs in HDFS or object storage, region assignments in the hbase:meta table, in-flight operations in the master's procedure store, and coordination data in ZooKeeper. An upgrade changes the code that reads and writes all of it, often while the cluster keeps serving traffic.
This article treats the upgrade as an engineering procedure. It explains what changes underneath, which version jumps have special documented requirements, how to run a rolling upgrade with health gates between steps, and why rollback must be designed in advance rather than improvised. The version-specific facts come from the Apache HBase reference guide's upgrade paths.
What an upgrade actually changes
It helps to list the state an upgrade touches, because each item has its own compatibility story and its own rollback risk.
| Layer | What can change | Why it matters |
|---|---|---|
| Binaries and dependencies | HBase jars, bundled Hadoop client, Netty, protobuf shading | Coprocessors and custom filters compiled against old APIs may fail to load |
| Master procedure store | Format and procedure types for assign, split, merge, DDL | A new master may not understand procedures an old one left in flight |
| hbase:meta and system tables | Schema or contents of catalogue tables | Once rewritten, older masters may not read them |
| On-disk data | HFile and WAL formats, usually backward readable | New writers can produce files old readers cannot parse |
| RPC and client protocol | New calls, removed calls, changed defaults | Old clients against new servers, or the reverse, can break |
| Configuration | Renamed, deprecated or new-default properties | A silently ignored property changes behaviour |
The reason ordering and rollback get so much attention is the middle three rows. Binaries can be swapped back in minutes. Rewritten metadata and new-format files cannot be un-written, so the point of no return is the first time a new component persists something an old one cannot read. Find that point for your jump before you start. The procedure framework article explains the procedure store that most version-specific steps revolve around.
Compatibility rules and the documented special paths
HBase's general policy is that patch releases and minor releases within a major line are rolling-upgradeable: you can restart servers one at a time and a mixed-version cluster keeps working during the transition. Major version jumps may require a full shutdown. Several jumps inside the 2.x line also have explicit prerequisites. The ones below come from the reference guide and are the most common sources of trouble.
| From and to | Documented requirement | Practical meaning |
|---|---|---|
| 1.x to 2.x | Full shutdown is the safe path. A rolling upgrade is described but marked experimental: start from the latest 1.4.x, set hbase.assignment.usezk to false, confirm hbck reports no inconsistencies, upgrade region servers first and masters last | Plan a maintenance window unless you have rehearsed the experimental path on a copy of production |
| 2.0 or 2.1 to 2.2+ | 2.2 uses new assign and unassign procedure types and cannot run the old ones. Stop all masters, set hbase.procedure.upgrade-to-2-2 to true, start one old-version master and let it drain the procedure store and exit, then start 2.2+ masters and roll region servers | The drain step is mandatory; skipping it leaves the new master unable to process leftover procedures |
| 2.2 to 2.3+ | The procedure store moves from the MasterProcWALs directory into a local HBase region; the new master migrates automatically on startup and deletes the old directory | Upgrade backup masters first, then the active one. After migration an old master cannot read the store |
| 2.4+ to 3.0 | hbase:namespace is removed and folded into hbase:meta; RegionServer grouping is reimplemented. The guide lists no special steps from 2.4.x | Treat this as a major upgrade: rehearse, and check rsgroup and namespace tooling that reads system tables directly |
HBase 3.0.0 is listed as released in August 2026, so many teams are now deciding between the latest 2.x line and 3.0. There is no reason to rush: 2.x lines continue to receive releases, and a major jump deserves its own rehearsal. Anything in your estate that reads hbase:namespace or the old group tables directly needs changing before a 3.0 move; the hbase:meta article and the RegionServer groups article cover what those structures do.
Pre-flight checks
Most failed upgrades were already unhealthy clusters. An upgrade restarts every server and reassigns every region, which exposes problems that a stable cluster was quietly tolerating: a region stuck in transition, a table half-disabled, a procedure waiting on a lock since last month. Clear them first.
- No regions in transition and no stuck procedures. Check the master UI's procedures and locks pages; anything older than minutes is a problem to resolve on the old version, where it was created.
- Consistency. On 2.x, run HBCK2's read-only reports, such as
reportMissingRegionsInMetaandextraRegionsInMetawithout--fix, and the master's own consistency report. Do not run the 1.x hbck against a 2.x cluster. - Snapshots of critical tables, exported off-cluster if the data is irreplaceable. Snapshots are cheap to take because they reference existing HFiles; see HBase snapshots.
- Coprocessors and custom code rebuilt against the target version and tested on a staging cluster. A coprocessor that fails to load can abort a region server when
hbase.coprocessor.abortonerroris true, which is the default. - Configuration diff. Compare your hbase-site.xml against the target's defaults and deprecation list; note properties that no longer exist.
- Client compatibility. List every client version in use, including Spark, Phoenix and replication peers, and confirm each against the target server version.
- Hadoop and Java. Confirm the target release supports your Hadoop and JDK versions from its documentation before touching anything.
The rolling procedure
For a patch or minor upgrade inside a major line, the usual order is masters first, then region servers. Upgrade and restart the backup masters, then stop the active master so an upgraded backup takes over, then upgrade the old active and let it rejoin as a backup. Region servers keep serving their existing regions while masters restart. The 1.x to 2.x path is the documented exception, with region servers first; when release notes specify an order, they win.
Before touching region servers, turn the balancer off so it does not shuffle regions onto servers you are about to restart, and pause scheduled major compactions and large bulk loads. Then upgrade one canary region server and let it soak under real traffic for long enough to see compactions, flushes and a few hours of load. Only then roll the rest.
#!/usr/bin/env bash
# Rolling region server upgrade with a health gate between hosts.
set -euo pipefail
HOSTS_FILE=${1:?usage: roll.sh hosts.txt}
MAX_P99_MS=${MAX_P99_MS:-50}
MASTER=${MASTER:?set MASTER to the active master host}
echo "balance_switch false" | hbase shell -n
gate() {
# Wait until no regions are in transition and no servers are dead.
for i in $(seq 1 60); do
rit=$(curl -s "http://$MASTER:16010/jmx?qry=Hadoop:service=HBase,name=Master,sub=AssignmentManager" | jq '.beans[0].ritCount')
dead=$(curl -s "http://$MASTER:16010/jmx?qry=Hadoop:service=HBase,name=Master,sub=Server" | jq '.beans[0].numDeadRegionServers')
if [[ "$rit" == "0" && "$dead" == "0" ]]; then return 0; fi
sleep 10
done
return 1
}
while read -r host; do
echo "== $host"
# Unload regions, stop, then restart and reload them (script ships in HBase's bin/).
ssh -n "$host" 'sudo /opt/hbase/deploy_new_version.sh' # swap binaries and config only
bin/graceful_stop.sh --restart --reload "$host" < /dev/null
gate || { echo "health gate failed after $host; stopping"; exit 1; }
./check_latency.sh "$MAX_P99_MS" || { echo "latency gate failed after $host"; exit 1; }
done < "$HOSTS_FILE"
echo "balance_switch true" | hbase shell -nJMX bean and attribute names vary between versions, so read them from the master's /jmx page on both versions and fix the query before the upgrade. graceful_stop.sh moves a server's regions away before stopping it, and --reload brings them back afterwards to preserve locality. Keep the deploy step idempotent and let configuration management place binaries; the script only restarts.
What to watch during the roll
The health gate catches hard failures. Softer regressions need dashboards compared against a baseline captured the day before: p99 read and write latency per table, regions in transition, request rates per server, flush and compaction queue lengths, block cache hit ratio, GC pause time and WAL sync latency. A new version sometimes changes a default, for example a compaction or cache setting, and the effect shows up as a slow drift over hours rather than an error.
Watch locality too. Each restart moves regions, and even with reload some regions come back to a different server or lose HDFS block locality until the next major compaction. A read latency increase right after the roll that fades over days is usually locality, not the new version.
Rollback is a design decision, not a button
Rolling back binaries is easy: redeploy the old version and restart. Rolling back state is often impossible. After a 2.3+ master migrates the procedure store, the old directory is gone. After a 3.0 master folds namespaces into meta, an older master will not find what it expects. Files written by new-version region servers may use features old readers cannot parse. The practical rule is that rollback is only possible until the first component persists new-format state, so identify that step in advance and put your final go or no-go check immediately before it.
For minor upgrades inside a line, rolling back region servers is usually fine because data formats rarely change; test it anyway on staging by upgrading, writing data, flushing and compacting, then downgrading. For anything that migrates metadata, the real rollback is a restore: snapshots taken before the upgrade, or a second cluster that never upgraded.
Blue-green: upgrading by building a second cluster
For major jumps, or when downtime and rollback risk are unacceptable, build a new cluster on the target version and move traffic to it. The pattern is: create tables on the new cluster, seed them with ExportSnapshot from snapshots of the old cluster, set up replication from old to new to carry writes made since the snapshot, verify row counts and sampled reads, then switch clients and stop writes to the old cluster. Rollback is switching clients back, provided you also replicate new-to-old after the cut or accept losing writes made in between.
The costs are real: double hardware for the transition, replication compatibility between versions to verify, and client configuration changes to coordinate. The benefit is that the old cluster is untouched until you decide it is safe to retire. The HMaster article explains the assignment and DDL paths a new cluster exercises heavily during seeding.
Failure modes and recovery
- Regions stuck in transition after the new master starts. Usually leftover procedures or a region that failed to open. Find the cause in the region server log first; HBCK2's
assignscan reschedule an assignment, andbypasscan release a stuck procedure, but both are last resorts that can leave state inconsistent if used blindly. - Region server aborts on startup. Most often a coprocessor or custom filter compiled against the old version. Remove it from the table descriptor or configuration, or ship the rebuilt jar, before continuing.
- Tables left DISABLING or ENABLING. An interrupted DDL across the restart. HBCK2's
setTableStatecan correct the recorded state after you have confirmed what the regions are actually doing. - Meta holes or overlaps. Report with
reportMissingRegionsInMetaand repair withfixMetaonly after taking a snapshot of the affected tables. - Clients failing after the roll. An old client using a removed call, or a shaded dependency clash in the application. Upgrade clients in their own change, tested against the new servers in staging.
Worked example: a 40-node cluster, 2.4 to 2.5
A 40-region-server cluster with two masters, about 12,000 regions and a 99th-percentile read latency objective of 20 ms is moving from 2.4.x to 2.5.x, a minor upgrade. The team rebuilds its two coprocessors, runs integration tests against a staging cluster on 2.5, and diffs configuration. Pre-flight on production shows three regions in transition from a failed split last week; they are fixed on 2.4 before the window.
On the day: snapshots of the six critical tables, balancer off, backup master upgraded and restarted, active master stopped so the upgraded backup takes over, old active upgraded. One canary region server is rolled and soaked for four hours while dashboards are compared with the previous day. Then the script rolls the remaining 39 servers. With about 300 regions per server, each graceful stop and reload takes six to eight minutes, so the roll takes roughly five hours with gates. The latency objective holds, apart from a short read-latency bump from locality that recovers after the weekend's major compactions. The balancer is re-enabled at the end.
What to do next
- Write down your exact source and target versions and read the upgrade section and release notes for every line in between.
- Identify the point of no return for your jump, and decide whether rollback means redeploying binaries or restoring from snapshots or a second cluster.
- Clear regions in transition and stuck procedures, run HBCK2's read-only reports and snapshot critical tables.
- Rebuild coprocessors and custom filters against the target version and test them on staging with production-like load.
- Script the roll with a health gate between hosts, balancer off, and a canary region server soak before the rest.
- For major versions, rehearse on a copy of production, or plan a blue-green move with snapshots and replication.