The HBase shell is the first tool most people touch and the last one they fully understand. It looks like a SQL console, but it is a JRuby interactive Ruby session in which every command, such as put, scan or major_compact, is a Ruby method that calls the ordinary HBase Java client. That one fact explains its strengths (you can loop, branch and call Java classes), its surprises (values are raw bytes, and a typo can be valid Ruby) and its dangers (a single unbounded scan or count is a real workload on a production cluster).
This article treats the shell as an operator's instrument. We build a small table, read and write it with the correct byte encodings, scan it safely, change its schema, run cluster operations, and finally script it so a deployment pipeline can use it without misreading failures. If you want the Java API for application code, the shell is not the place; it is for inspection, administration and small, deliberate changes.
What the shell actually is
Running hbase shell starts a JVM, loads JRuby and the HBase client jars, and opens a connection using the hbase-site.xml on the classpath. Like any client, it first asks ZooKeeper where the hbase:meta table lives, reads meta to find the region that holds a given row key, and then sends RPCs directly to that RegionServer. Administrative commands such as create, alter and balancer go through the Admin interface to the active HMaster instead.
Because the shell is a client, a command against a region in transition retries with the client's settings, which is why it can hang for minutes before failing. Kerberos credentials come from your ticket cache, so on a secure cluster you kinit first. And because the session is Ruby, anything that is not a known command is evaluated as Ruby, so Scan 't1' (capital S) gives a Ruby error rather than an HBase one.
Starting it against the right cluster
The most expensive shell mistake is running a correct command against the wrong cluster. The shell uses whichever configuration is on its classpath, so make the target explicit. You can override any client property at start-up with -D, for example hbase shell -Dhbase.zookeeper.quorum=zk1.prod.example.org. Pass JVM options such as heap size or GC flags through the HBASE_SHELL_OPTS environment variable, and add -d when you need DEBUG logging to see what the client is doing.
Run status as the first command of every session so the output names the cluster, and keep history across sessions with a .irbrc; the reference guide's example is require 'irb/ext/save-history' followed by IRB.conf[:SAVE_HISTORY] = 100.
Worked example: writing and reading bytes correctly
HBase stores every row key, qualifier and value as an uninterpreted byte array. The shell converts Ruby strings to bytes, which is convenient for text and a trap for numbers. The session below creates a namespaced table with one column family, three versions, a 30-day TTL and three split points so it starts with four regions instead of one. It then writes a text value, a 4-byte big-endian integer, and a counter.
hbase> create_namespace 'ops'
hbase> create 'ops:user_events', {NAME => 'e', VERSIONS => 3, TTL => 2592000},
hbase* {SPLITS => ['4', '8', 'c']}
hbase> put 'ops:user_events', 'a1f3|u42', 'e:type', 'login'
hbase> put 'ops:user_events', 'a1f3|u42', 'e:score', "\x00\x00\x00\x2A"
hbase> incr 'ops:user_events', 'a1f3|u42', 'e:clicks', 1
hbase> get 'ops:user_events', 'a1f3|u42'
hbase> get 'ops:user_events', 'a1f3|u42', {COLUMN => 'e:type', VERSIONS => 3}
hbase> get_counter 'ops:user_events', 'a1f3|u42', 'e:clicks'
hbase> scan 'ops:user_events', {COLUMNS => ['e:score:toInt', 'e:clicks:toLong', 'e:type'],
hbase* ROWPREFIXFILTER => 'a1f3', LIMIT => 20}Notice three things. The row key a1f3|u42 starts with a short hash prefix so writes spread across the pre-split regions, a pattern covered in depth in HBase hotspotting. The score is written with a double-quoted Ruby string containing escape sequences, which produces exactly four bytes; writing '42' would have stored two ASCII bytes, and any Java reader calling Bytes.toInt on it would fail or return garbage. And incr stores an 8-byte long, so get_counter and the :toLong formatter read it back, while a plain get shows escaped bytes such as \x00\x00\x00\x00\x00\x00\x00\x01.
The column list in the final scan uses per-column formatters: e:score:toInt asks the shell to render that column with the toInt method of org.apache.hadoop.hbase.util.Bytes. The scan help also allows a custom class with the form cf:q:c(MyFormatterClass).format, and the FORMATTER and FORMATTER_CLASS options change the default for all columns. Formatters change only what you see; the stored bytes are untouched.
Scanning without hurting production
A scan with no bounds reads every region of the table, pulls blocks through the RegionServers' block caches and streams results back to one client. On a large table that is a full-table read issued from a laptop. The rules are simple: always bound the key range, name the columns you need, set a LIMIT, and switch off block caching for one-off reads so you do not evict hot data that applications depend on.
# Bounded range, only the columns you need, don't pollute the block cache
scan 'ops:user_events', {STARTROW => 'a1f3', STOPROW => 'a1f4',
COLUMNS => ['e:type'], LIMIT => 100, CACHE_BLOCKS => false}
# Filter language string, evaluated on the RegionServer
scan 'ops:user_events', {ROWPREFIXFILTER => 'a1f3',
FILTER => "SingleColumnValueFilter('e', 'type', =, 'binary:login')"}
# Forensics: see delete markers and old versions that a normal scan hides
scan 'ops:user_events', {ROWPREFIXFILTER => 'a1f3|u42', RAW => true, VERSIONS => 10}
# Time-bounded (epoch millis, end exclusive)
scan 'ops:user_events', {TIMERANGE => [1790812800000, 1790899200000], LIMIT => 50}
# Counting: fine for thousands of rows, wrong tool for billions
count 'ops:user_events', INTERVAL => 100000, CACHE => 1000The FILTER string uses the filter language and runs on the RegionServer, which saves network but not disk: a filter still reads every row in the range to decide what to drop. A filter is no substitute for a good start and stop row. See HBase filters for which filters can skip data and which only discard it, and HBase scans for caching and batching behaviour.
RAW => true is the forensic switch: it returns delete markers and cells that are deleted but not yet removed by compaction. When someone asks "why did my data disappear", a raw scan of the row often shows a delete marker with a timestamp that names the culprit's clock.
Treat count with respect. Its help text warns that it may take a long time; it is a client-side scan that prints progress every INTERVAL rows (1,000 by default) and does not cache blocks by default. Raise CACHE so each RPC fetches more rows, and for large tables use the distributed counter instead:
# Distributed count as a MapReduce job instead of one client-side scan
hbase org.apache.hadoop.hbase.mapreduce.RowCounter ops:user_events
DDL, schema changes and destructive commands
Schema changes in the shell are online for most attributes: alter updates a family's versions, TTL, compression or block size and the change is rolled through the regions. Removing a family deletes its data, so it deserves the same care as a drop.
describe 'ops:user_events'
alter 'ops:user_events', {NAME => 'e', VERSIONS => 5, COMPRESSION => 'ZSTD'}
alter 'ops:user_events', {NAME => 'tmp', METHOD => 'delete'}
# Destructive changes: snapshot first, then disable, then act
snapshot 'ops:user_events', 'user_events_20261001_pre_drop'
disable 'ops:user_events'
drop 'ops:user_events'
# Recover: clone to a new name (restore_snapshot needs the table disabled)
clone_snapshot 'user_events_20261001_pre_drop', 'ops:user_events_restored'
# Emptying a table: truncate drops and recreates it with ONE region;
# truncate_preserve keeps the existing region boundaries
truncate_preserve 'ops:user_events'The ordering matters. drop refuses to run on an enabled table, which is a deliberate speed bump, not an obstacle to script around. A snapshot is a metadata operation that references existing HFiles, so it is cheap and should precede every destructive change; HBase snapshots covers retention and export.
The truncate versus truncate_preserve distinction catches many teams. truncate disables, drops and recreates the table, and the new table has a single region. If the table was pre-split for write distribution, the next bulk load hammers one RegionServer until splits catch up. truncate_preserve keeps the region boundaries. Be equally careful with the regex commands such as disable_all and drop_all: they list the matching tables and ask for confirmation, and a loose pattern matches more than you intended.
Operational commands
The shell is also the quickest cluster console. These are the commands operators reach for during incidents and maintenance:
status 'detailed' # per-server load, region counts, requests
list_regions 'ops:user_events' # regions, their servers and sizes
flush 'ops:user_events' # memstores to HFiles
major_compact 'ops:user_events' # rewrites every HFile: I/O heavy, schedule it
compaction_state 'ops:user_events'
balance_switch false # returns the PREVIOUS state; note it
move 'ENCODED_REGION_NAME', 'host,16020,STARTCODE'
balance_switch true
balancer
scan 'hbase:meta', {ROWPREFIXFILTER => 'ops:user_events,', COLUMNS => ['info:server']}status 'detailed' is the fastest way to spot a RegionServer serving far more requests than its peers. major_compact rewrites all of a region's files, so trigger it table by table off-peak. Moving regions by hand only sticks if the balancer is paused; balance_switch prints the previous setting, so record and restore it. For regions stuck in transition, the shell is the wrong tool: use HBCK2, and HBase troubleshooting walks through that path. Rate limits for noisy tenants are set with set_quota, explained in HBase quotas and throttling.
Scripting: table references, command files and exit codes
Because the session is Ruby, you can hold a table reference and call methods on it: t = get_table 'ops:user_events' and then t.put, t.scan or t.count. create also returns a reference. Combine that with Ruby iteration for small backfills:
# backfill.rb - run with: hbase shell -n backfill.rb
t = get_table 'ops:user_events'
%w[a1f3|u42 a1f3|u43 b7c0|u99].each do |row|
t.put row, 'e:schema', 'v2'
end
puts t.get('a1f3|u42', 'e:schema')
exitEach put is a separate RPC, so a shell loop handles hundreds or a few thousand rows; for more, write a client program or bulk load. A file of commands runs with hbase shell ./commands.txt. The commands are not echoed, which makes the output hard to line up with the input, so print markers with puts. End every script with exit so the process terminates cleanly instead of waiting at a prompt.
For automation, use non-interactive mode, -n or --non-interactive. In interactive mode the shell returns its own exit status, which is nearly always 0 even when a command failed. In non-interactive mode it passes the command's status back. The reference guide adds an important caveat: a non-zero status does not necessarily mean failure; it means the outcome is unknown. The client may have lost its connection after the server applied the change. A script must therefore re-check state before retrying or rolling back. This wrapper creates a table only if it is missing and handles the unknown case explicitly (it matches the phrase the exists command prints, so confirm that wording on your version):
#!/usr/bin/env bash
# ensure_table.sh - create the table only if it is missing; safe to re-run
set -u
check() { echo "exists 'ops:user_events'" | hbase shell -n 2>&1 | grep -q 'does exist'; }
if check; then echo "table present, nothing to do"; exit 0; fi
echo "create 'ops:user_events', {NAME => 'e', VERSIONS => 3}, {SPLITS => ['4','8','c']}" \
| hbase shell -n
status=$?
if [ $status -ne 0 ]; then
# Non-zero means "unknown", not "failed": the create may have been applied
if check; then echo "create landed despite exit $status"; exit 0; fi
echo "create failed (exit $status) and table is absent" >&2
exit 1
fiWrite scripts to be idempotent. Check exists before create, compare describe output before alter, and never chain a destructive command to a non-zero exit without a verification step in between.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Java reader gets garbage from a value written in the shell | Number stored as ASCII text | Write bytes explicitly (escaped string) or write from code; read with :toInt/:toLong |
| Command hangs for minutes, then fails | Client retries against a region in transition | Check list_regions and the master UI; fix assignment with HBCK2 |
| Latency spike during an ad-hoc query | Unbounded scan or count on a hot table | Bound the range, LIMIT, CACHE_BLOCKS => false, RowCounter for counts |
| Bulk load hot-spots one server after a reset | truncate recreated a single region | Use truncate_preserve, or recreate with SPLITS |
| Script reports success but nothing changed | Interactive mode returned the shell's own status | Run with -n and check state afterwards |
| Script retried a create and failed | Non-zero exit read as failure; first attempt had succeeded | Treat non-zero as unknown; check exists first |
| Changes landed on the wrong cluster | Wrong hbase-site.xml on the classpath | -D overrides, run status first, separate configs per environment |
Trade-offs: shell, client code or HBCK2
The shell wins for inspection, one-off administration and small, reviewed changes. It loses for volume (each call is a synchronous RPC) and for anything that must be tested and rolled back, where a program with the Java client is safer. For repair of meta and assignment state, HBCK2 is the supported tool; do not hand-edit hbase:meta from the shell. A good rule: if a shell command will run more than once, put it in a reviewed script with -n, idempotency checks and a snapshot step.
What to do next
- Add a
.irbrcwith saved history on every bastion host where operators run the shell. - Make
statusthe first command of every session and every script, and log its output. - Practise writing an integer with an escaped string and reading it back with
:toInt. - Adopt a scan rule: start row, stop row, columns,
LIMITandCACHE_BLOCKS => falsefor ad-hoc reads. - Replace large
countcalls with the RowCounter job. - Put a snapshot step before every
drop, family delete or truncate, and prefertruncate_preserve. - Convert repeated shell tasks into
-nscripts that treat non-zero exit as unknown and re-check state.