An HBase client is more than a network wrapper. It finds the cluster, finds which RegionServer holds each row, caches that map, sends RPCs, notices when regions have moved, and retries with a backoff policy that, left at its defaults, can hold a request for many minutes. Most production incidents blamed on HBase itself, such as threads piling up, requests hanging for minutes or a slow meta table after a server dies, are really client configuration problems.
This article covers what every client does, the Java options (blocking and asynchronous), the shaded artifact, how clients bootstrap, the retry and timeout settings with a worked worst case, buffered writes, and the non-JVM choices. Scanner internals are covered in the scans deep dive and gateway internals in the Thrift gateway article; this page is about choosing and operating the client.
What every client does
Every HBase request goes through three steps. First, bootstrap: the client asks a connection registry for cluster-level facts, chiefly the cluster ID and the location of the hbase:meta region. Second, locate: for a row key, the client finds the region whose key range contains it by looking in its region location cache and, on a miss, reading hbase:meta. The meta table article explains that lookup in detail. Third, call: the client sends the RPC to the RegionServer hosting the region.
Failures loop back. If the region has moved, split or its server has died, the server replies with an exception such as NotServingRegionException, or the connection fails. The client invalidates the cached location, sleeps, looks the region up again and retries. That loop gives HBase clients resilience across region moves, and it is also why a badly configured client can wait far longer than the application expected.
The Java landscape
| Client | Model | Use it for |
|---|---|---|
| Connection + Table | Blocking calls from a thread | Most services and batch jobs; simplest to reason about |
| AsyncConnection + AsyncTable | CompletableFuture-based, non-blocking | High-concurrency services where threads are the bottleneck |
| BufferedMutator | Client-side write buffer, flushed in batches | High-throughput ingest where losing the unflushed buffer is acceptable or handled |
| hbase-shaded-client | Same APIs, dependencies relocated | Applications with their own Guava, protobuf or Netty versions |
| Phoenix JDBC | SQL over HBase | When you want SQL and secondary indexes; see Apache Phoenix |
Use the shaded artifact by default in applications. The unshaded client pulls in Hadoop, Guava, protobuf and Netty at particular versions, and the classic failure is a NoSuchMethodError at runtime because the application's own Guava won a classpath conflict. hbase-shaded-client relocates those dependencies under an HBase-specific package so they cannot collide. Use the unshaded client only when you run inside a Hadoop environment that already provides matching versions, such as a MapReduce job built against the cluster's own jars.
Connection lifecycle and thread safety
The blocking API has a rule that decides whether a service is healthy: Connection is heavyweight and thread-safe, while Table, Admin and RegionLocator are lightweight and not thread-safe. A Connection owns the registry client, the region location cache, the RPC client and its sockets, and thread pools. Creating one per request repeats bootstrap and throws the cache away, so every request pays for meta lookups, and connection churn can exhaust sockets or overload the registry. Create one Connection per process (per cluster), share it, and get a Table per unit of work.
import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.hbase.HBaseConfiguration;
import org.apache.hadoop.hbase.TableName;
import org.apache.hadoop.hbase.client.*;
import org.apache.hadoop.hbase.util.Bytes;
public final class ProfileStore implements AutoCloseable {
private static final TableName TABLE = TableName.valueOf("app:profiles");
private static final byte[] CF = Bytes.toBytes("p");
private final Connection conn; // one per process, thread-safe
public ProfileStore() throws java.io.IOException {
Configuration conf = HBaseConfiguration.create(); // reads hbase-site.xml
conf.setInt("hbase.rpc.timeout", 2000); // per attempt
conf.setInt("hbase.client.operation.timeout", 5000); // whole call, all retries
conf.setInt("hbase.client.retries.number", 5);
this.conn = ConnectionFactory.createConnection(conf);
}
public byte[] email(String userId) throws java.io.IOException {
try (Table t = conn.getTable(TABLE)) { // cheap, per call, not shared
Result r = t.get(new Get(Bytes.toBytes(userId)).addColumn(CF, Bytes.toBytes("email")));
return r.getValue(CF, Bytes.toBytes("email"));
}
}
@Override public void close() throws java.io.IOException { conn.close(); }
}
The asynchronous client
Since HBase 2.0 there is a separate asynchronous API. ConnectionFactory.createAsyncConnection(conf) returns a CompletableFuture<AsyncConnection>, and tables come in two flavours. getTable(name, executor) returns an AsyncTable<ScanResultConsumer> whose callbacks run on your executor, so it is safe to do slow work in them. getTable(name) returns an AsyncTable<AdvancedScanResultConsumer> whose callbacks run on the RPC framework's own threads. It is faster, but a callback that blocks stalls every other request sharing that thread.
AsyncConnection aconn = ConnectionFactory.createAsyncConnection(conf).get();
AsyncTable<ScanResultConsumer> t = aconn.getTable(TABLE, callbackPool);
List<CompletableFuture<Result>> futures = t.get(userIds.stream()
.map(id -> new Get(Bytes.toBytes(id)).addColumn(CF, EMAIL))
.collect(Collectors.toList()));
CompletableFuture.allOf(futures.toArray(new CompletableFuture[0]))
.orTimeout(200, TimeUnit.MILLISECONDS) // application deadline
.whenComplete((v, err) -> { /* respond, count errors per future */ });The async client removes the thread-per-request ceiling, but not the need for backpressure. A service that turns every incoming request into unbounded futures will queue work faster than RegionServers can serve it, and the queued work eventually fails with timeouts. Bound in-flight requests with a semaphore, and choose between the async client and a blocking client on a bounded pool by where the latency budget is spent, not by fashion. Batch jobs are usually simplest with the blocking API.
How clients bootstrap: ZooKeeper and RPC registries
Historically clients bootstrapped through ZooKeeper, reading the meta location and cluster ID from znodes. That meant every client needed network access to the ZooKeeper quorum and added load to it. A master-based registry came next; in 2.5.0 it was deprecated in favour of RpcConnectionRegistry, which is the default from HBase 3.0.0. With it, clients bootstrap over RPC from a configured set of nodes, and RegionServers as well as masters can serve the request, so the list can be spread over the cluster, refreshed and cleaned of dead nodes.
<property>
<name>hbase.client.registry.impl</name>
<value>org.apache.hadoop.hbase.client.RpcConnectionRegistry</value>
</property>
<property>
<name>hbase.client.bootstrap.servers</name>
<value>rs1.example.com:16020,rs2.example.com:16020,rs3.example.com:16020</value>
</property>If the bootstrap servers are not configured, the client falls back to the master addresses. The practical benefit is network design: application subnets need to reach RegionServers anyway, and no longer need to reach ZooKeeper. Check which registry your client and server versions support before switching, and roll the change out per application.
Retries and timeouts: a worked worst case
Four settings govern how long a call can take. hbase.client.pause (default 100 ms) is the base sleep between retries. hbase.client.retries.number (default 15) caps attempts. The sleep before each retry is the pause times a multiplier from a fixed table in the client, {1, 2, 3, 5, 10, 20, 40, 100, 100, 100, 100, 200, 200}, staying at 200 once the table runs out, plus a small random jitter. hbase.rpc.timeout (default 60,000 ms) bounds a single attempt, and hbase.client.operation.timeout (default 1,200,000 ms, twenty minutes) bounds the whole operation across all retries.
Now work the worst case with defaults. Summing the multipliers for 15 retries gives 881 for the first 13 entries plus 200 and 200 for the last two, 1,281 in total, so the client can spend about 128 seconds just sleeping. Each attempt can also wait up to 60 seconds for the RPC, so the theoretical total is far beyond anything a user will wait for, and the twenty-minute operation timeout is the only backstop. In an online service this shows up as a thread pool that fills with calls stuck on one dead RegionServer until the whole service stops responding.
The fix is to size the settings from the latency budget backwards. For a service with a 500 ms deadline, a per-attempt RPC timeout of 200 to 300 ms, three or four retries, a 50 ms pause and an operation timeout equal to the deadline will fail fast and let the caller degrade. Batch jobs can keep longer timeouts, because waiting out a region move is cheaper than failing a task. Scanners have their own lease, hbase.client.scanner.timeout.period (default 60,000 ms), which must be longer than the time your code spends processing between next() calls. For reads that must survive a server failure without waiting for reassignment, region replicas offer timeline-consistent reads from a secondary.
Writes: put, batch and BufferedMutator
A single Table.put is one RPC. put(List<Put>) and batch group mutations by RegionServer and send them in parallel, which is how you load thousands of rows without thousands of round trips. BufferedMutator goes further: it accumulates mutations in a client-side buffer (hbase.client.write.buffer, default 2,097,152 bytes) and flushes in the background when the buffer fills.
BufferedMutatorParams params = new BufferedMutatorParams(TABLE)
.writeBufferSize(8L * 1024 * 1024)
.listener((e, mutator) -> {
for (int i = 0; i < e.getNumExceptions(); i++) {
log.error("failed row {}", Bytes.toStringBinary(e.getRow(i).getRow()), e.getCause(i));
deadLetter.add(e.getRow(i)); // do not silently drop
}
});
try (BufferedMutator m = conn.getBufferedMutator(params)) {
for (Event ev : events) m.mutate(toPut(ev));
m.flush(); // before acknowledging the source
}The trade-off is durability at the edge. Mutations sitting in the buffer are only in the client's memory; if the process dies they are gone, and errors surface asynchronously through the exception listener rather than from mutate. Flush before you acknowledge upstream, such as committing a Kafka offset, and always install a listener that records failed rows.
Non-JVM clients
Outside the JVM there are two families. Gateway clients call a Thrift or REST server that runs a Java client on your behalf: HappyBase for Python speaks the original Thrift interface, and any language with a Thrift compiler or HTTP library can use the gateways. Their cost is an extra hop and a second set of timeouts, pools and scanner state in the gateway, all covered in the gateway deep dive. Native clients implement the HBase RPC protocol themselves, like gohbase for Go, and avoid the hop, but they lag the Java client in features and must track protocol and security changes.
import happybase
pool = happybase.ConnectionPool(size=8, host="thrift.example.com", port=9090, timeout=2000)
with pool.connection() as conn:
t = conn.table("app:profiles")
row = t.row(b"user:48213", columns=[b"p:email"])
with t.batch(batch_size=500) as b: # client-side batching via Thrift
b.put(b"user:48214", {b"p:email": b"x@example.com"})Choose the Java client when you can, a gateway when the language is fixed and throughput per client is modest, and a native client when the gateway hop is the bottleneck and you can accept feature gaps.
Failure modes
- Connection per request. Latency dominated by bootstrap and meta lookups; registry and meta load spikes. Share one Connection.
- Shared Table across threads. Intermittent corrupted batches or exceptions under load. Get a Table per unit of work.
- Default timeouts in online paths. Threads stuck for minutes when a RegionServer dies. Set per-attempt and operation timeouts from the deadline.
- Meta stampede. After a server failure, many clients invalidate locations at once and hammer
hbase:meta. Fewer, longer-lived connections and sane retry pauses reduce it. - Classpath conflicts.
NoSuchMethodErrorinvolving Guava or protobuf. Switch to the shaded client. - Blocking in async callbacks. Throughput collapses without errors. Use the executor-backed table or move work off the callback.
- Kerberos expiry. Long-running clients start failing authentication; ensure credentials are renewed from a keytab rather than a short-lived ticket cache.
What to do next
- Audit every service for how many Connections it creates; make it one per process per cluster.
- Switch applications to
hbase-shaded-clientunless they run inside the cluster's own Hadoop classpath. - Set
hbase.rpc.timeout,hbase.client.operation.timeout, retries and pause from each service's latency budget, and test by killing a RegionServer under load. - Decide between blocking and async clients by measured thread usage, and bound in-flight async requests.
- Give every BufferedMutator an exception listener and flush before acknowledging upstream.
- Plan the move to
RpcConnectionRegistryso clients no longer need ZooKeeper access. - Enable client metrics and alert on retries, meta lookups and operation timeouts.