The HBase Thrift gateway lets programs in Python, Go, C++, PHP and other languages read and write HBase without the Java client. It is a separate process that speaks Apache Thrift RPC to clients and the ordinary HBase client protocol to the cluster. That makes it a translation layer, a connection pool and, under load, a bottleneck, and most production problems with it come from not knowing which of those roles is failing.
The broader choice between the REST gateway, Thrift, the native client and SQL, together with the basics of scanner state and whose identity HBase sees, is covered in Thrift and REST gateways in depth. This article goes one level down, into the gateway itself: which server to run, how its thread pool and queue behave, the transport settings that must match on both ends, the frame-size limit, how to scan without pinning clients to one gateway, and how to operate a fleet of gateways. Configuration keys and defaults below were checked against the HBase source.
Two servers, two interfaces
HBase ships two Thrift servers, started by different commands and speaking different interface definitions (IDLs). hbase thrift runs the original server (the thrift1 IDL), and hbase thrift2 runs the newer one. A client generated from one IDL cannot talk to the other server, and the commonest first-day error is exactly that mismatch.
| thrift (thrift1) | thrift2 | |
|---|---|---|
| Start command | hbase thrift start | hbase thrift2 start |
| Data model | Column names as family:qualifier strings, Mutation lists | Structs mirroring the Java API: TGet, TPut, TScan, TDelete |
| Batch reads and writes | Per-row and row-list calls | getMultiple, putMultiple, deleteMultiple |
| Atomic operations | Increments, checkAndPut | checkAndPut, checkAndDelete, checkAndMutate, mutateRow, increment, append |
| Stateless scan | No | getScannerResults |
| Typical client | happybase (Python) | Code generated from the thrift2 hbase.thrift |
For new code, generate a client from the thrift2 IDL: it maps closely to the Java client, supports multi-row batches and conditional mutations, and has a stateless scan call that matters for load balancing. Keep thrift1 for existing happybase applications. Both servers accept the same server-model and transport options, and the thrift2 interface includes getThriftServerType, so a client can confirm which one it reached.
Inside one gateway
A request travels through four stages. The server implementation accepts the connection and reads a Thrift message. A worker thread runs the handler, which converts Thrift structs into HBase Get, Put or Scan objects. The handler obtains an HBase connection from the connection cache, keyed by effective user, and that connection locates the right region through hbase:meta (cached after the first lookup; see the meta table) and calls the RegionServer directly. The result is converted back into Thrift structs and written to the client.
Two consequences follow. The gateway adds a network hop and a serialisation step to every call, so batching matters more than with the Java client. And the gateway holds no data: it can be restarted or replaced freely, except for open scanners, which live in its memory.
Region movement is handled for you. When a region splits, moves during balancing or reopens after a RegionServer failure, the HBase client inside the gateway gets a not-serving or moved error, refreshes its cached location from hbase:meta and retries within its own retry budget. The Thrift client sees only a slower call, or, if the client-side retries are exhausted, an error. Budget the timeouts in one direction: bound the gateway's own HBase client operation timeouts, set in the gateway's hbase-site.xml, below the Thrift client's socket timeout. Then the gateway gives up first and returns an error the caller can retry, instead of the caller timing out while the gateway is still retrying, and resending a request the cluster may yet complete, which doubles load exactly when the cluster is recovering.
Server implementations and transports
The server model is set with hbase.regionserver.thrift.server.type or a command-line flag. The source defines four, and three of them only work with framed transport.
| Option | Thrift class | Framed required | Behaviour |
|---|---|---|---|
| threadpool (default) | TBoundedThreadPoolServer | No | One worker thread per connection, bounded pool and queue |
| hsha | THsHaServer | Yes | One selector thread for I/O, a worker pool for calls |
| nonblocking | TNonblockingServer | Yes | Single-threaded selector; calls run on the I/O thread |
| threadedselector | TThreadedSelectorServer | Yes | Several selector threads plus a worker pool |
Transport and protocol must match on both ends. hbase.regionserver.thrift.framed and hbase.regionserver.thrift.compact both default to false, so a default server expects buffered transport and the binary protocol. A client that sends framed or compact messages to it, or the reverse, does not get a clean error: calls hang until a timeout or fail with protocol exceptions that look like corruption. Decide once, write it into the client library wrapper, and alert on protocol errors in gateway logs.
The threadpool server ties a thread to each open client connection for as long as it stays open. With many idle, pooled client connections, workers are consumed by connections doing nothing. Selector-based servers (hsha, threadedselector) multiplex idle connections cheaply and hand only active calls to workers, which suits large fleets of clients with long-lived connections. HTTP mode, described below, is a separate path.
Sizing the worker pool and queue
The threadpool server's pool is controlled by hbase.thrift.minWorkerThreads (default 16), hbase.thrift.maxWorkerThreads (default 1000), hbase.thrift.maxQueuedRequests (default 1000) and hbase.thrift.threadKeepAliveTimeSec (default 60). When all threads are busy and the queue is full, the server logs Queue is full, closing connection and drops the client's connection. Clients see that as a reset socket, not a busy signal, so it must be alerted on explicitly. The hsha and threadedselector servers bound their call queue with the same hbase.thrift.maxQueuedRequests setting; the nonblocking server has no worker pool at all and runs every call on its single I/O thread, so one slow call stalls every client.
Size with Little's law. Suppose a service issues 3,000 calls per second through the gateways at 8 ms average latency, including the hop to the RegionServer. Calls in flight: 3,000 x 0.008 = 24. With the threadpool server, however, threads track open connections, not calls: if 40 application instances each hold a pool of 10 connections, that is 400 threads spread across the gateways even when only 24 calls are active. Either cap client pool sizes, or switch to a selector server whose worker count follows active calls. Across three gateways behind a load balancer, plan so that two can carry the full load during a restart. For the threadpool server the binding figure is connections: with one gateway down, 400 connections become 200 threads on each survivor, still under the 1,000 maximum but a long way above the 16 minimum. For hsha or threadedselector the binding figure is calls: 24 in flight becomes 12 per surviving gateway, which a modest worker pool handles easily.
Latency is set mostly downstream. When a RegionServer slows because of compaction, garbage collection or a hot region, in-flight calls at the gateway rise at the same arrival rate, the pool fills and connections start dropping. Gateway saturation is usually a symptom; check RegionServer health before adding gateway threads.
The thrift2 API in practice
Generate Python bindings from the thrift2 IDL with thrift --gen py hbase.thrift; its Python namespace is hbase. The example writes a batch and reads it back with a stateless scan, using framed transport to match a gateway started with the framed option.
from thrift.transport import TSocket, TTransport
from thrift.protocol import TBinaryProtocol
from hbase import THBaseService
from hbase.ttypes import TPut, TColumnValue, TScan, TColumn, TIOError
sock = TSocket.TSocket("hbase-thrift.internal", 9090)
sock.setTimeout(10_000) # ms; keep above the gateway's HBase operation timeout
transport = TTransport.TFramedTransport(sock) # must match server's framed setting
client = THBaseService.Client(TBinaryProtocol.TBinaryProtocol(transport))
transport.open()
table = b"orders"
puts = [TPut(row=f"cust42#{i:06d}".encode(),
columnValues=[TColumnValue(b"o", b"total", str(i * 10).encode())])
for i in range(500)]
for start in range(0, len(puts), 100): # keep each frame small
client.putMultiple(table, puts[start:start + 100])
scan = TScan(startRow=b"cust42#", stopRow=b"cust42$",
columns=[TColumn(b"o", b"total")])
rows = client.getScannerResults(table, scan, 200) # open, read 200, close
for r in rows:
print(r.row, [(c.qualifier, c.value) for c in r.columnValues])
transport.close()Push filtering to the cluster rather than the client. TScan and TGet accept a filterString written in the HBase filter language, which the gateway parses into a real filter that runs on the RegionServers, so only matching cells cross the network twice. Custom filter classes can be registered for that parser with hbase.thrift.filters. Fetching whole rows and filtering in Python multiplies gateway traffic, worker time and frame sizes for no benefit. The same applies to column lists: name the columns you need in columns instead of reading entire families.
Two details matter. The batch is split into chunks of 100 rows, because framed transport has a maximum frame size: hbase.regionserver.thrift.framed.max_frame_size_in_mb defaults to 2 MB, and one large putMultiple or wide result silently exceeds it. And TIOError carries a canRetry field, so retry logic can distinguish transient failures from ones that will fail again.
Scanners and load balancers
openScanner creates an HBase scanner inside the gateway and returns an integer ID; getScannerRows pulls the next batch and closeScanner releases it. The ID is meaningful only to the gateway that created it. Put gateways behind a round-robin load balancer and a scan whose second call lands on another gateway fails. There are three workable designs.
- Stateless scans.
getScannerResultsopens a scanner, reads up to numRows and closes it in a single call; the source confirms the scanner is closed before returning. Page through a range by setting the next startRow just after the last row returned. Any gateway can serve any page, so this is the design to prefer. - Connection affinity. Keep each scan on one TCP connection so a layer-4 balancer keeps it on one gateway. It works until the connection breaks mid-scan, when the client must restart from its last row.
- Client-side gateway choice. The client picks a gateway per scan from a list. More code, but no dependence on balancer behaviour.
Always close scanners in a finally block. An abandoned scanner holds gateway memory and RegionServer resources until its lease expires, and a client that leaks scanners in a loop can exhaust both. See HBase scans for caching, batching and lease behaviour on the server side.
HTTP mode, security and the connection cache
Setting hbase.regionserver.thrift.http to true runs Thrift over HTTP in an embedded web server, with its own thread pool (hbase.thrift.http_threads.min and .max, defaults 2 and 100). HTTP mode is what enables impersonation: hbase.thrift.support.proxyuser only takes effect with HTTP on, and the server authorises each impersonated user against the Hadoop proxy-user rules. Kerberos uses hbase.thrift.keytab.file and hbase.thrift.kerberos.principal for the gateway's own identity, plus the SPNEGO keytab and principal for HTTP clients.
The connection cache holds one HBase connection per effective user and closes connections idle longer than hbase.thrift.connection.max-idletime (10 minutes), checked every hbase.thrift.connection.cleanup-interval (10 seconds). With impersonation and many distinct end users, that is many connections, each with its own region-location cache.
Operating a gateway fleet
- Run at least two gateways per application tier, sized so the rest carry full load during a rolling restart.
- Scrape the info server on port 9095 (
hbase.thrift.info.port) for JVM and call metrics, and alert on the queue-full log line and on protocol errors. - Set
hbase.thrift.readonlyto true on gateways serving read-only consumers, so a bug cannot write. - Apply per-user limits on the cluster with quotas; the gateway itself does no fair sharing between clients.
- Remember that the server drops a client connection idle for longer than
hbase.thrift.server.socket.read.timeout(60 seconds by default); client connection pools must validate or reconnect rather than assume a pooled socket is still open.
Failure modes
- IDL mismatch. A thrift1 client against a thrift2 server fails with unknown-method errors.
- Transport mismatch. Framed against buffered, or compact against binary, hangs or produces garbage exceptions.
- Frame too large. Batches or wide rows over 2 MB fail under framed transport; chunk writes and limit columns.
- Queue full. Connections are dropped silently under load; usually a slow RegionServer upstream.
- Scanner ID on the wrong gateway. Stateful scans break behind round-robin balancing; use getScannerResults paging.
- Leaked scanners. Memory and lease pressure grows until the gateway or RegionServers slow down.
What to do next
- Decide thrift or thrift2 per application, and record the server type, framed and compact settings in one client wrapper.
- Pick threadpool for a few busy connections, or hsha or threadedselector for many mostly idle ones.
- Size worker threads and client pools with Little's law, keeping capacity for one gateway down.
- Chunk batch writes and cap result width to stay under the frame limit.
- Replace stateful scans with getScannerResults paging wherever gateways sit behind a load balancer.
- Alert on queue-full closures, protocol errors and scanner counts, and set per-user quotas on the cluster.