HBase speaks its own protobuf RPC protocol, and the full-featured client that speaks it is the Java one. It locates regions through hbase:meta, caches region locations, retries on region moves and talks directly to each RegionServer. A Python analytics job, a Go service or a shell script does not want to reimplement any of that. The Thrift and REST gateways solve this by running the Java client inside a server process and exposing HBase operations over a simpler protocol: Thrift RPC or HTTP.

The idea is simple, and that simplicity hides three facts that decide whether a gateway deployment works in production. The gateway is not stateless: scanners live in its memory. HBase does not see the calling user unless you configure impersonation; it sees the gateway. And every request pays an extra hop and serialisation step, which matters for scans far more than for single gets. This article explains all three, with commands and code for both gateways, and ends with a checklist.

Advertisement

What a gateway actually does

A gateway is an ordinary HBase client process. It reads hbase-site.xml, connects to ZooKeeper and hbase:meta, and holds a long-lived connection to the cluster. For each incoming Thrift call or HTTP request it builds the equivalent Get, Put, Delete or Scan, executes it with the Java client against the owning RegionServers, and serialises the result back. Region routing, retries and location caching all happen inside the gateway, so the external client only needs the gateway's address.

Two gateways ship with HBase. The REST gateway maps tables, rows, columns and scanners to URLs, and returns XML, JSON or protobuf. The Thrift gateway exposes an interface definition from which Thrift generates clients for many languages. There are two Thrift IDLs: the original one, started with bin/hbase-daemon.sh start thrift and used by the popular Python library happybase, and a second, thrift2, whose API mirrors the Java client's Get, Put and Scan objects more closely and is started with the thrift2 command. A client must use the IDL the server runs.

HBase gateways: translation, identity and scanner statePython servicehappybase (Thrift)Web / curl clientHTTP + JSON / protobufLoad balancersticky for scanners90908080Thrift gatewayJava client + scanner mapREST gatewayJava client + scanner maphbase:metaregion lookupRegionServersACL checkRPCWithout doAsHBase sees only the gateway principalone identity for ACLs, quotas and auditWith doAs (proxy user)gateway authenticates, impersonates callerper-user ACLs, quotas and auditStateful scannersa scanner id is valid only on the gateway that created it
Clients reach gateways through a load balancer; each gateway runs the Java client, which routes to RegionServers. Scanner ids are held in the gateway that created them, and HBase authorises the gateway's principal unless impersonation is on.

The REST gateway: resources and encodings

The REST gateway listens on port 8080 by default. Its URLs follow the data model: /table/schema for the table descriptor, /table/row for a row, /table/row/family:qualifier for a column, and /table/scanner for scanners. With Accept: application/json, row keys, column names and values are all base64-encoded inside a Row, Cell structure, because HBase stores arbitrary bytes and JSON strings cannot carry them safely. Writes use PUT or POST with the same structure. Encode row keys that contain reserved characters such as | or / in the URL path.

# Start a REST gateway (default port 8080) on a host with the cluster's hbase-site.xml
bin/hbase rest start -p 8080

# Cluster version and a table's schema
curl -s -H "Accept: application/json" http://gw:8080/version/cluster
curl -s -H "Accept: application/json" http://gw:8080/orders/schema

# Read one row; row keys, columns and values come back base64-encoded
curl -s -H "Accept: application/json" http://gw:8080/orders/u123%7C20260930
# {"Row":[{"key":"dTEyM3wyMDI2MDkzMA==",
#          "Cell":[{"column":"ZDpzdGF0dXM=","timestamp":1790000000000,"$":"U0hJUFBFRA=="}]}]}

# Stateless range read: no scanner is created on the gateway
curl -s -H "Accept: application/json" \
  "http://gw:8080/orders/*?startrow=u123%7C20260901&endrow=u123%7C20261001&limit=100"

# Stateful scanner: create, page, delete
curl -si -X PUT -H "Content-Type: text/xml" \
  -d '<Scanner batch="500" startRow="dTEyMw==" endRow="dTEyNA=="/>' \
  http://gw:8080/orders/scanner
# HTTP/1.1 201 Created
# Location: http://gw:8080/orders/scanner/SCANNER_ID
curl -s -H "Accept: application/json" http://gw:8080/orders/scanner/SCANNER_ID
curl -s -X DELETE http://gw:8080/orders/scanner/SCANNER_ID

A small decoding helper is worth writing once rather than scattering base64 calls through application code. Note that values stay as bytes: HBase has no types, so the application must know that d:total holds a decimal string and not an eight-byte long.

import base64, requests

def get_row(gw, table, key):
    r = requests.get(f"{gw}/{table}/{requests.utils.quote(key, safe='')}",
                     headers={"Accept": "application/json"}, timeout=5)
    if r.status_code == 404:
        return None
    r.raise_for_status()
    b64 = base64.b64decode
    row = r.json()["Row"][0]
    return {b64(c["column"]).decode(): b64(c["$"]) for c in row["Cell"]}
Advertisement

Scanners hold state in the gateway

A common claim, repeated in older notes on this topic, is that the gateways are stateless enough for any gateway to serve any request. For gets and puts that is true. For scanners it is false, in both gateways. Creating a REST scanner with PUT /table/scanner opens a real HBase scanner inside that gateway process and returns its id in the Location header; each GET on that id returns the next batch; DELETE closes it. The Thrift gateway works the same way: opening a scanner returns an integer id that indexes a scanner map in the gateway's memory, and the client then calls get-next with that id.

Put two gateways behind a round-robin load balancer and the second page of a scan lands on a gateway that has never heard of the id, and fails. The fixes are to make the load balancer sticky by client for the duration of a scan, to give long scans a direct connection to one gateway, or, for bounded reads, to use the REST gateway's stateless form, /table/*?startrow=...&endrow=...&limit=..., which carries the whole scan in one request and leaves nothing behind. Scanners that clients abandon also hold RegionServer resources until they time out, so always DELETE REST scanners and close Thrift scanners in a finally block. A gateway restart kills every open scan; clients must be able to resume from the last row key they processed.

The Thrift gateway with happybase

The Thrift gateway listens on port 9090 by default and is the better choice for high request rates from a single language, because Thrift's binary encoding is compact and a persistent connection avoids HTTP overhead per call. Client and server must agree on transport and protocol: the server's hbase.regionserver.thrift.framed and hbase.regionserver.thrift.compact settings have to match the client's framed or buffered transport and compact or binary protocol. A mismatch tends to show up as hangs or garbled-frame errors rather than a clear message, so set a client timeout.

# Thrift gateway: bin/hbase-daemon.sh start thrift   (default port 9090)
import happybase   # speaks the original ("thrift1") HBase IDL

conn = happybase.Connection("thrift-gw.internal", port=9090,
                            timeout=10000,              # ms; fail fast instead of hanging
                            transport="buffered",       # must match the server's framed setting
                            protocol="binary")          # must match the server's compact setting
orders = conn.table("orders")

# Batched writes: one Thrift call per flush, not per row
with orders.batch(batch_size=500) as b:
    for o in new_orders:
        b.put(f"{o.user}|{o.day}".encode(), {b"d:status": o.status.encode(),
                                              b"d:total": str(o.total).encode()})

# Scans are gateway-side scanners; batch_size rows per round trip
for key, data in orders.scan(row_prefix=b"u123|", columns=[b"d:status"], batch_size=1000):
    handle(key, data)
conn.close()

The two habits in that example are the ones that matter for performance. Batch mutations, because a put per row pays a full client-to-gateway-to-RegionServer round trip for each row. And set the scan batch size, because a scan that fetches one row per call turns a million-row read into a million Thrift calls. Push filtering to the server where you can; the gateways accept HBase filters, and a prefix or column filter evaluated on the RegionServer is far cheaper than shipping rows through the gateway to discard them.

Security: whose identity does HBase see?

On a Kerberos-secured cluster, the gateway authenticates to HBase with its own keytab and principal. The HBase reference guide is explicit about the consequence: without impersonation, all client access through the Thrift gateway uses the gateway's credentials and has its privileges. Authorisation is still enforced by HBase itself, at the RegionServers, but against the gateway principal. Every user of the gateway can do whatever the gateway may do, and audit logs and request quotas see a single user.

Impersonation, called doAs or proxy user, fixes this. The client authenticates to the gateway (SPNEGO for REST, Kerberos over SASL or HTTP for Thrift), and the gateway performs each operation as that user, so HBase applies that user's ACLs and quotas. For Thrift, doAs requires the HTTP transport mode. The gateway principal must also be allowed to impersonate users in the Hadoop proxy-user configuration, scoped to the gateway hosts and the groups it may act for, and it must itself be granted whatever permissions the reference setup requires. hbase.thrift.security.qop selects authentication, integrity or privacy for the Thrift SASL channel; use privacy on untrusted networks, and TLS for REST.

<!-- Thrift gateway with Kerberos and per-user impersonation (hbase-site.xml) -->
<property><name>hbase.thrift.keytab.file</name><value>/etc/security/keytabs/hbase-thrift.keytab</value></property>
<property><name>hbase.thrift.kerberos.principal</name><value>thrift/_HOST@EXAMPLE.COM</value></property>
<property><name>hbase.thrift.security.qop</name><value>privacy</value></property>
<property><name>hbase.regionserver.thrift.http</name><value>true</value></property>
<property><name>hbase.thrift.support.proxyuser</name><value>true</value></property>

<!-- REST gateway with SPNEGO for clients and impersonation -->
<property><name>hbase.rest.keytab.file</name><value>/etc/security/keytabs/hbase-rest.keytab</value></property>
<property><name>hbase.rest.kerberos.principal</name><value>rest/_HOST@EXAMPLE.COM</value></property>
<property><name>hbase.rest.authentication.type</name><value>kerberos</value></property>
<property><name>hbase.rest.authentication.kerberos.principal</name><value>HTTP/_HOST@EXAMPLE.COM</value></property>
<property><name>hbase.rest.authentication.kerberos.keytab</name><value>/etc/security/keytabs/spnego.keytab</value></property>
<property><name>hbase.rest.support.proxyuser</name><value>true</value></property>

Deployment and sizing

Run gateways as a separate tier, not on RegionServers, so a gateway traffic spike cannot starve region serving of CPU or heap. Two or more instances behind a load balancer give availability; size by concurrent scanners and request rate rather than data volume, since the gateway holds no data beyond in-flight results. Each open scanner's batch sits in gateway heap, so large batches multiplied by many concurrent scans is the usual source of gateway garbage-collection pauses. Watch request latency at the gateway and at the RegionServers separately: if the gateway's latency is much higher, the gateway is the bottleneck, not HBase. The REST gateway also has a read-only mode, hbase.rest.readonly, which is a cheap safety measure for gateways that serve dashboards.

Worked example: a Python service reading order history

An order-history API written in Python serves about 2,000 requests per second, each reading one user's orders for a month: a prefix scan returning on average 40 rows. Row keys are user|yyyymmdd. The team first deploys one REST gateway, uses stateful scanners with a batch of 10, and does not delete them.

Three problems appear. Each request makes one call to create the scanner, five to page through 40 rows and none to close it, so the gateway handles about 12,000 HTTP requests per second and accumulates abandoned scanners until they expire. Adding a second gateway behind round-robin makes about half of the page requests fail with unknown scanner ids. And the security team notices that audit logs attribute every read to the gateway principal.

The fixed design switches to the stateless form with limit=100, making one HTTP request per API call and leaving no scanner behind, so round-robin becomes safe. Heavy batch exports move to happybase against a Thrift pool with batch_size=1000 and sticky connections. SPNEGO and proxy-user are enabled so HBase sees the calling service account, per-service quotas apply, and the gateway principal's own grants are reduced to what impersonation needs. Gateway request rate drops by roughly a factor of six, and failures from scanner routing disappear.

Failure modes

  • Round-robin with stateful scanners: intermittent unknown-scanner errors that vanish when you test against one gateway.
  • Leaked scanners: gateway heap growth and RegionServer lease pressure from clients that never close scans.
  • Single-identity gateways: ACLs, quotas and audits collapse to one principal; a compromised client inherits everything the gateway may do.
  • Transport mismatch: framed versus buffered or compact versus binary produces hangs, not clear errors.
  • Unbatched access: per-row puts and one-row scan batches make the extra hop dominate latency.
  • Gateway co-located with RegionServers: gateway load competes with region serving for heap and CPU.

Choosing between REST, Thrift, native and SQL

OptionStrengthsCostsUse when
REST gatewayAny HTTP client, curl-debuggable, stateless range readsBase64 JSON overhead, HTTP per callLow to moderate rates, tooling, dashboards
Thrift gatewayCompact binary, persistent connections, many languagesStateful scanners, IDL and transport matchingHigh-rate services in Python, Go and others
Native Java clientNo extra hop, full API, direct region routingJVM requiredLatency-critical or high-throughput paths
SQL layerSchema, types, secondary indexes, JDBCAnother system to operateRelational access patterns; see Apache Phoenix

What to do next

  1. Inventory which applications reach HBase through gateways, and whether each uses stateful scanners.
  2. Make the load balancer sticky for scanner traffic or move bounded reads to the stateless REST form, then test with two gateways.
  3. Audit client code for unclosed scanners, per-row puts and missing scan batch sizes.
  4. Check whether HBase sees end users or the gateway principal; if the gateway, enable doAs with scoped proxy-user settings and reduce the gateway's own grants.
  5. Set hbase.thrift.security.qop to privacy or use TLS on any network you do not fully trust.
  6. Move gateways off RegionServer hosts and alert on gateway latency separately from RegionServer latency.
Key takeaway: The Thrift and REST gateways run the HBase Java client on behalf of non-Java callers. Gets and puts can go to any gateway, but scanners live in the gateway that created them, so route scans stickily or use stateless range reads. Without doAs, HBase authorises and audits only the gateway principal. Batch writes, size scan batches, close scanners, and keep gateways on their own hosts. See the <a href="hbase_overview.html">HBase overview</a> for how RegionServers and hbase:meta fit together.