HBase's native RPC is built around the Java client, and native clients for other languages are few. The Thrift gateway fills the gap: it exposes an interface described in a Thrift IDL file, and any language with a Thrift code generator gets a typed client. Python, Go, C++ and Ruby services use it to read and write HBase without a JVM.

The gateway's operation, sizing and load balancing are covered in the Thrift gateway architecture article, and the choice between REST, Thrift and native clients in Thrift and REST gateways. This article is about the API itself, read as a contract: what each group of calls promises, which operations are atomic, how data is encoded, what happens on partial failure, and how mismatched IDL versions fail without an error. Names are taken from the hbase.thrift file for the thrift2 interface in the Apache HBase source.

Advertisement

Two IDLs, one gateway process

hbase.thrift (thrift2 IDL)pinned to the cluster versionthrift --genGenerated stubspy, go, cpp, rb, ...Your serviceTHBaseService.ClientTPut, TGet, ...Thrift gatewayTHBaseService handlerHBase Java clientTable and Admin callsRegionServers / Masteratomicity is per rowContract surfacereads, writes, check-and-mutate, scanners, region locations,DDL and namespaces, slow log, grant and revokeErrors back to the clientTIOError(message, canRetry), TIllegalArgument, transport exceptions
From IDL to RegionServer. The client is generated from hbase.thrift; the gateway translates each call into Java Table or Admin operations, so HBase semantics such as per-row atomicity pass through unchanged.

HBase ships two Thrift interfaces. The original, usually called thrift1, is the one the Python library happybase speaks: tables and columns are addressed with family:qualifier byte strings and writes are lists of Mutation structs. The newer thrift2 interface, service THBaseService, mirrors the Java client's object model with TGet, TPut, TDelete, TScan, TIncrement and TAppend, and in recent releases also covers table, column family and namespace administration. A gateway runs one handler or the other, and getThriftServerType() tells a client which one it reached. New code should use thrift2; the rest of this article does.

Generate bindings from the IDL file that matches your cluster's HBase version, for example thrift --gen py hbase.thrift. The file declares its namespaces, so the Python package is hbase, Java classes land in org.apache.hadoop.hbase.thrift2.generated, and so on. Commit the IDL file you generated from, next to the generated code, so the contract your service was built against is on record.

The method surface, grouped by purpose

GroupMethodsWhat the contract promises
Point readsexists, existsAll, get, getMultipleRead one or many rows; existsAll returns one boolean per TGet
Writesput, putMultiple, deleteSingle, deleteMultiple, increment, appendEach row's mutation is atomic; a batch as a whole is not
Conditional and row-atomiccheckAndPut, checkAndDelete, checkAndMutate, mutateRowCompare-and-set on one cell, or several puts and deletes applied atomically to one row
ScanningopenScanner, getScannerRows, closeScanner, getScannerResultsStateful scanner ids held by one gateway, or a one-shot open, read and close
TopologygetRegionLocation, getAllRegionLocationsRegion boundaries and hosting servers for a table
AdministrationcreateTable, modifyTable, deleteTable, enableTable, ... , createNamespace, ...DDL through the gateway's Admin connection
Operations and securitygetSlowLogResponses, clearSlowLogResponses, getClusterId, grant, revokeDiagnostics and access-control changes

The last two groups matter for security as much as for features: any client that can reach a thrift2 gateway can call deleteTable or grant, subject only to what the gateway's own HBase identity may do. Treat gateway reachability as an administrative privilege unless you have configured authentication and impersonation, as described in the gateway articles.

Advertisement

The data model: bytes all the way down

Every row key, family, qualifier and value in the data methods is Thrift binary. The API does not know about strings, integers or JSON; it moves bytes, and the encoding is your contract with every other reader of the table. Three conventions matter in practice.

Numbers. The Java client's Bytes.toBytes(long) writes a long as 8 big-endian bytes, and HBase counters use exactly that layout. increment fails on a cell that is not 8 bytes long, so a Python writer that stores b"42" as text breaks every later increment. Read and write counters with struct.pack(">q", n) and struct.unpack(">q", b).

Timestamps. Cell timestamps are i64 milliseconds since the epoch. Leave them unset and the RegionServer assigns them; set them explicitly only when you need idempotent rewrites or deliberate versioning.

Results. A TResult is a row key plus a flat list of TColumnValue cells, each with family, qualifier, value and timestamp. There is no nested map; group cells yourself. TColumn with only a family selects the whole family, and with a qualifier selects one column.

Worked example: an order state machine

import struct
from thrift.transport import TSocket, TTransport
from thrift.protocol import TBinaryProtocol
from hbase import THBaseService
from hbase.ttypes import (TPut, TGet, TColumn, TColumnValue, TIncrement, TColumnIncrement,
                          TDelete, TMutation, TRowMutations, TCompareOperator, TIOError)

sock = TSocket.TSocket("hbase-thrift.internal", 9090)
sock.setTimeout(10_000)
transport = TTransport.TFramedTransport(sock)          # must match the gateway setting
client = THBaseService.Client(TBinaryProtocol.TBinaryProtocol(transport))
transport.open()

T, F = b"orders", b"o"
row = b"cust42#20260930#A17"

# 1. Create the order only if it does not exist: value omitted means "column absent".
created = client.checkAndPut(T, row, F, b"status", None,
                             TPut(row=row, columnValues=[TColumnValue(F, b"status", b"PENDING"),
                                                         TColumnValue(F, b"total", b"129.90")]))

# 2. PENDING -> PAID, atomically with removing the payment lock column.
paid = client.checkAndMutate(T, row, F, b"status", TCompareOperator.EQUAL, b"PENDING",
                             TRowMutations(row=row, mutations=[
                                 TMutation(put=TPut(row=row, columnValues=[TColumnValue(F, b"status", b"PAID")])),
                                 TMutation(deleteSingle=TDelete(row=row, columns=[TColumn(F, b"lock")])),
                             ]))

# 3. Per-customer counter: HBase stores it as an 8-byte big-endian signed long.
res = client.increment(T, TIncrement(row=b"cust42#stats",
                                     columns=[TColumnIncrement(F, b"paid_orders", 1)],
                                     returnResults=True))
count = struct.unpack(">q", res.columnValues[0].value)[0]

# 4. Read back only what you need; cells come back as a flat list.
r = client.get(T, TGet(row=row, columns=[TColumn(F, b"status"), TColumn(F, b"total")]))
cells = {cv.qualifier: cv.value for cv in r.columnValues}
transport.close()

Step 1 relies on a documented detail of checkAndPut: when the expected value is not provided, the check is for the non-existence of the column. That turns it into a create-if-absent primitive, so two retries of the same order creation cannot both succeed. The return value tells you whether your put was applied.

Step 2 uses checkAndMutate with TRowMutations, whose mutations are a TMutation union of either a put or a delete. The comparison and all mutations execute atomically on one row, so no reader ever sees a status of PAID with the lock column still present. Every mutation must target the row named in TRowMutations. mutateRow offers the same atomic multi-mutation without the check. The example sticks to EQUAL; if you use the ordering operators such as LESS or GREATER, test which side of the comparison your value lands on against a real cluster before relying on it.

Step 3 shows returnResults on TIncrement: ask for the new value only when you need it, because returning it costs bytes and a little server work on every call.

Batches are not transactions

putMultiple, getMultiple and deleteMultiple send many rows in one Thrift call, which is how you avoid a network round trip per row. Behind the gateway they become multi-row Java operations spread across regions and servers, and HBase guarantees atomicity only within a row. If a putMultiple raises TIOError, some rows may already be written.

The IDL contains a trap here. deleteMultiple is declared to return a list of TDelete, which reads like "the deletes that failed", but its documentation says it throws TIOError if any delete fails and always returns an empty list, kept only for backward compatibility. Code that checks the returned list for failures will never find any.

The safe pattern is to make batches idempotent and retry the whole batch. Puts and deletes of fixed cells are naturally idempotent. increment and append are not: retrying after a timeout can apply them twice, because the first attempt may have succeeded after the client gave up. For counters that must be exact, record an operation id with checkAndPut before incrementing, or accept approximate counts.

Read and write options worth knowing

  • Server-side filtering. filterString on TGet and TScan takes the HBase filter language, evaluated on the RegionServers; see HBase filters for what each filter costs.
  • Versions and time. maxVersions, timeRange and, on scans, colFamTimeRangeMap for per-family ranges.
  • Existence without data. existence_only on TGet, or exists and existsAll, avoid moving values you will discard.
  • Wide rows. storeLimit and storeOffset page through columns within a row; batchSize on a scan caps cells per result, and TResult.partial flags a row split across results.
  • Replica reads. consistency=TIMELINE, optionally with targetReplicaId, lets reads be served by secondary region replicas, and TResult.stale tells you when one was; region replicas explains the trade-off.
  • Scan shape. reversed, limit, caching and readType (DEFAULT, STREAM, PREAD) mirror the Java Scan.
  • Durability. Mutations accept TDurability: USE_DEFAULT, SKIP_WAL, ASYNC_WAL, SYNC_WAL or FSYNC_WAL. SKIP_WAL loses acknowledged writes if a RegionServer dies before flushing; WAL durability covers what each level guarantees.

Region-parallel scans and DDL

A single scan over a large table is serial: one scanner walks regions in order. getAllRegionLocations returns every region with its start and end keys, server and replica id, which lets a client split the key space on region boundaries and scan the pieces in parallel, one connection per worker. Empty start and end keys represent the table's open ends, and tables with region replicas list each range once per replica, so keep replica id 0.

from concurrent.futures import ThreadPoolExecutor
from hbase.ttypes import TScan

def region_ranges(client, table):
    locs = client.getAllRegionLocations(table)
    return sorted((l.regionInfo.startKey or b"", l.regionInfo.endKey or b"")
                  for l in locs if not l.regionInfo.replicaId)     # primaries only

def scan_range(bounds):
    start, end = bounds
    c, transport = new_client()               # one connection per worker, never shared
    try:
        return c.getScannerResults(b"orders", TScan(startRow=start, stopRow=end or None,
                                                    caching=500), 10_000)
    finally:
        transport.close()

with ThreadPoolExecutor(max_workers=8) as pool:
    parts = list(pool.map(scan_range, region_ranges(client, b"orders")))

Each worker owns its own transport because a Thrift client is not safe to share between threads. getScannerResults returns at most numRows rows, so for ranges larger than that, loop with openScanner and getScannerRows on the same gateway, as the gateway article explains. Keep the parallelism modest: eight workers can hit eight RegionServers at once, which is the point, but also eight gateway worker threads.

For DDL, createTable(desc, splitKeys) creates a table pre-split at the keys you pass, which avoids the hot single region that every new table starts with. Region boundaries from getAllRegionLocations are also a quick way to check whether a table has split as expected.

When the IDL versions drift

Thrift identifies struct fields by numeric id, and a receiver skips any field id it does not know. That keeps old and new peers wire-compatible, and it also means a newer client talking to an older gateway can have options silently ignored. If your stubs were generated from a recent IDL and you set a newer field such as limit or filterBytes on a scan, a gateway built before that field existed drops it and returns an unlimited or unfiltered result, with no error. The reverse case is benign: an old client simply never sends the new fields.

Defend against this deliberately: generate from the IDL of the oldest gateway version you run, pin that file in your repository, regenerate as part of HBase upgrades, and add an integration test that sets each option you depend on and checks its effect, for example that a scan with a limit of 10 returns 10 rows.

Errors and retries

ErrorMeaningClient action
TIOError with canRetry trueTransient HBase condition, for example a region movingRetry with backoff; the operation should be idempotent
TIOError otherwiseRequest will fail again: missing table, bad counter cell, permissionSurface the message; do not retry
TIllegalArgument from scanner callsUnknown scanner id: expired, closed, or created on another gatewayReopen the scan from the last row key you processed
Transport exceptionConnection reset, timeout, frame or protocol mismatchReconnect; outcome of the last write is unknown

The last row is the subtle one. A timeout on a write does not mean the write failed; it means you do not know. That is the practical reason every write path should be idempotent, and why the Java client, discussed in HBase client libraries, invests so much in retry semantics that a Thrift client must build for itself.

Failure modes

  • Text-encoded counters. A writer stores numbers as strings, and increment fails on those cells. Standardise on 8-byte big-endian longs.
  • Trusting deleteMultiple's return. It is always empty. Treat an exception as "some unknown subset applied" and retry the idempotent batch.
  • Double-counted increments. Retried increments after timeouts inflate counters. Deduplicate with an operation id or accept approximation.
  • Silently ignored options. Stubs newer than the gateway drop fields without error. Pin the IDL and test each option's effect.
  • Cross-row atomicity assumed. Two rows updated in one batch can be seen half-applied. Keep invariants inside one row, or add application-level recovery.
  • Open administrative surface. Any reachable thrift2 gateway exposes DDL and grants. Restrict network access and enable authentication.

What to do next

  1. Find the HBase version of your gateways, copy that release's thrift2 hbase.thrift into your repository and generate stubs from it.
  2. Call getThriftServerType() at startup and fail fast if the gateway is not running thrift2.
  3. Write down the byte encoding for every row key, qualifier and value type, including counters as 8-byte big-endian longs.
  4. Rewrite create-if-absent and state transitions with checkAndPut and checkAndMutate, and add a concurrency test that races two clients.
  5. Make every batch idempotent, retry whole batches only on retryable errors, and remove any code that inspects deleteMultiple's return value.
  6. Add an integration test for each read option you rely on, run it against every gateway version in your fleet, and restrict network access to the gateway's administrative surface.
Key takeaway: The thrift2 API is a thin, typed mirror of the HBase Java client, so HBase's semantics pass straight through it: atomicity is per row, batches can partly apply, and every value is bytes whose encoding you own. Use checkAndPut with an absent value and checkAndMutate for safe state changes, store counters as 8-byte big-endian longs, make batches idempotent, ignore deleteMultiple's always-empty return, parallelise scans on region boundaries, and pin the IDL to your gateway version so options are never silently dropped.