Teams that run HBase for serving and Trino (forked from Presto, which continues as PrestoDB) for analytics eventually want one SQL query that joins the two. The intuitive answer, "add an HBase catalog to Trino", no longer exists in a supported form. The current Trino documentation (release 483 at the time of writing) lists no HBase connector and no Phoenix connector. The Phoenix connector that let Trino read HBase through Apache Phoenix was marked for removal on Trino's published roadmap in 2024, and release 447 had already dropped support for Phoenix 5.1.x and earlier. The PrestoDB connector lists we checked include Accumulo and Kudu but not HBase or Phoenix. Third-party HBase connectors exist on GitHub; we have not verified them against current releases.

So this article is about the real options in 2026: querying live HBase through a pinned engine release, exporting snapshots to an open table format, or streaming changes through replication, plus the decoding work all three share. It explains why SQL engines and HBase fit poorly, works an example, and lists the failure modes that take down serving clusters. If you are weighing whether to keep HBase at all, read HBase and Its Alternatives first.

Why SQL engines and HBase fit poorly

A distributed SQL engine wants four things from a storage system: splits it can read in parallel, typed columns, predicates it can push down, and statistics for planning joins. HBase offers the first and partly the third, and neither of the others.

  • Splits. Regions are natural splits: each covers a contiguous row-key range on one region server. A connector lists regions and gives each worker a scan over one range, ideally on a worker near the region server.
  • Types. HBase stores byte arrays. A long, a string and a Phoenix-encoded integer are all just bytes, and the encoding is a convention of whatever wrote them. Every SQL layer needs an external mapping from column qualifiers to types and encodings.
  • Pushdown. Only predicates on the leading part of the row key become cheap scans with start and stop rows. Predicates on other columns become server-side filters, which still read every row in the range.
  • Statistics. There are none a planner can use, so join order and broadcast decisions are guesses. A large HBase table placed on the wrong side of a join is read in full by every worker that needs it.

The fourth point is the operational one. An analytic query that scans a large table pulls data through the region servers' read path, evicting hot rows from the block cache and adding latency to the serving traffic the cluster exists for. Any design must answer how analytic scans are kept from harming online reads.

Three options

Three ways to put SQL over HBase data; only one touches region servers at query timeHBase tablesregions on region serversA. Phoenix connectorpinned older Trinolive scansB. Snapshot exportHFiles read offlineC. CDC via replicationWAL edits to KafkaIceberg tablestyped, partitioneddailyminutesTrinojoins, BI, ad hocA gives the freshest data and the most risk to the serving cluster; B and C move analytic load off HBase.
Option A reads live data through region servers. Options B and C copy data into an open table format that Trino reads natively.
A. Phoenix connectorB. Snapshot exportC. CDC stream
FreshnessLiveHours to a dayMinutes
Load on region serversHigh during scansNone for readsReplication only
Engine versionPinned to an older TrinoCurrentCurrent
Works for non-Phoenix tablesNoYes, with a decoderYes, with a decoder
Build effortLowMediumHigh

Option A: a pinned Phoenix catalog

If your HBase tables were created through Apache Phoenix, the shortest path is a Trino release that still ships the Phoenix connector, configured against your ZooKeeper quorum. Phoenix supplies what HBase lacks: a typed schema, row-key encodings, secondary indexes and its own pushdown, and the connector maps those into Trino.

# etc/catalog/phoenix.properties on a pinned, older Trino release
connector.name=phoenix5
phoenix.connection-url=jdbc:phoenix:zk1,zk2,zk3:2181:/hbase
phoenix.config.resources=/etc/hbase/conf/hbase-site.xml
-- federated join: live Phoenix table with an Iceberg dimension
SELECT d.region, count(*) AS orders, sum(o.amount) AS revenue
FROM phoenix.default.orders o
JOIN iceberg.dim.customers d ON o.customer_id = d.customer_id
WHERE o.order_ts >= TIMESTAMP '2026-10-03 00:00:00'
GROUP BY d.region;

Treat this as a bridge, not a destination. Pinning a query engine means forgoing its security fixes, and your Phoenix version must be one that release supports: from release 447 the connector requires Phoenix 5.2.0 or later, as the release 464 documentation also states. Run the pinned cluster separately from your main Trino deployment, restrict it to the few queries that genuinely need live data, and plan the move to option B or C. Before choosing it at all, check whether an engine you already run, such as Impala with HBase, covers the need.

Option B: snapshot export to Iceberg

For most analytics, yesterday's data is fine, and the cleanest architecture never lets analytic reads touch a region server. HBase snapshots make that possible. A snapshot is a metadata operation that records which HFiles make up a table at a moment; it copies no data. TableSnapshotInputFormat then reads those HFiles directly from HDFS or object storage, bypassing region servers entirely. A Spark job decodes the rows and writes a typed, partitioned Iceberg table, which Trino reads natively; see the Iceberg and Trino stack.

# hbase shell, once a day
snapshot 'orders', 'orders_20261004'
// Scala: read the snapshot offline and write Iceberg
import org.apache.hadoop.hbase.client.Result
import org.apache.hadoop.hbase.io.ImmutableBytesWritable
import org.apache.hadoop.hbase.mapreduce.TableSnapshotInputFormat
import org.apache.hadoop.hbase.util.Bytes
import org.apache.hadoop.mapreduce.Job
import org.apache.hadoop.fs.Path
import spark.implicits._

val job = Job.getInstance(hbaseConf)
TableSnapshotInputFormat.setInput(job, "orders_20261004", new Path("/tmp/restore/orders_20261004"))

val rows = spark.sparkContext.newAPIHadoopRDD(job.getConfiguration,
    classOf[TableSnapshotInputFormat], classOf[ImmutableBytesWritable], classOf[Result])

val cf = Bytes.toBytes("d")
val orders = rows.map { case (_, r) =>
  val key = r.getRow                                   // salt(1) | customer_id(16) | reverse_ts(8)
  (Bytes.toString(key, 1, 16),
   Long.MaxValue - Bytes.toLong(key, 17),             // undo reverse timestamp
   Bytes.toLong(r.getValue(cf, Bytes.toBytes("amt"))),
   Bytes.toString(r.getValue(cf, Bytes.toBytes("st"))))
}.toDF("customer_id", "order_ts_ms", "amount_cents", "status")

orders.writeTo("lake.sales.orders_snapshot").overwritePartitions()

The restore directory holds references to the snapshot's files and must be cleaned up after the job. Delete old snapshots too: while a snapshot exists, HFiles it references are moved to the archive directory when compactions replace them instead of being deleted, so forgotten snapshots grow storage quietly. Snapshot mechanics are covered in HBase Snapshots.

Option C: change data capture through replication

When analysts need minutes rather than a day, stream changes instead of copying tables. HBase replication ships write-ahead-log edits to peers asynchronously, and the peer does not have to be another HBase cluster: the replication endpoint is pluggable, and the Apache hbase-connectors project includes a Kafka proxy that presents itself as a replication peer and publishes edits to Kafka topics. A streaming job then decodes edits and merges them into an Iceberg table keyed on the row key.

Three semantic gaps need explicit handling. Edits arrive per cell, not per row, so a job must assemble cells into rows and decide what a partial update means. Deletes come in several kinds (cell, column, family, row), and each needs a mapping to a SQL delete or a column set to NULL. And replication is at-least-once with ordering only within a region, so merges must be idempotent and must compare cell timestamps rather than trusting arrival order. Run a periodic snapshot export as a reconciliation baseline; CDC pipelines drift.

Decoding bytes into columns

Every option except Phoenix-through-Phoenix needs a decoder, and the decoder is where most correctness bugs live. Write the mapping down as data, version it, and test it against real rows:

table: orders
row_key:
  - {name: salt,        bytes: 1,  type: uint8,   drop: true}
  - {name: customer_id, bytes: 16, type: string}
  - {name: order_ts,    bytes: 8,  type: int64_be, transform: "Long.MAX_VALUE - x"}
columns:
  d:amt: {name: amount_cents, type: int64_be}
  d:st:  {name: status,       type: utf8}
versions: latest_only

Watch for three traps. HBase's Bytes.toLong is big-endian two's complement, but Phoenix flips the sign bit of signed numbers so they sort correctly as bytes; decoding a Phoenix table with plain HBase utilities gives wrong values for every number. Phoenix also encodes column qualifiers as numbers by default since version 4.10, so the qualifiers you see in raw HBase are not the column names in the DDL. And HBase keeps multiple versions and honours TTLs; decide whether the SQL view shows only the latest version and whether expired cells that a snapshot still contains should be filtered out.

Worked example: orders for finance and operations

An e-commerce team keeps 9 TB of orders in HBase, serving order-history pages at a p99 under 20 ms. Finance wants daily revenue by region joined to a customer dimension in Iceberg; operations wants a dashboard of orders in the last 15 minutes.

Finance gets option B. A nightly snapshot plus a Spark job reading HFiles offline produces orders_snapshot, partitioned by order date. The job reads 9 TB from storage, but no region server serves a single byte, so order-history latency is unaffected. Rewriting the full table each night is the price of simplicity; measure the job's runtime and cost on the first runs, and when that starts to hurt, the table is a candidate for the CDC path.

Operations gets a narrower answer than option C. Their query touches one day of orders and needs only counts. A small service that already sees order events can publish them to Kafka, and a streaming job can write them to an Iceberg table every minute; Trino reads that. Full HBase CDC is justified only when many tables need fresh analytic copies. The lesson generalises: start from the freshness each consumer needs, not from the storage system.

A custom service behind the Thrift connector

One more connector is worth knowing about. Trino's current list includes a Thrift connector, which delegates metadata and reads to an external service implementing Trino's Thrift interface. A team that must expose live HBase data to Trino without pinning an old release can write that service: it lists tables from a mapping file, turns row-key predicates into scans with start and stop rows, and returns typed pages. That is real engineering work, and the service owns all the load-protection decisions a connector would have made, such as concurrency limits per region server and refusing scans with no key predicate. It is justified when live, ad hoc SQL over HBase is a long-term requirement rather than a migration step.

Failure modes

  • Full scans against serving clusters. An analytic query without a row-key predicate reads every region and churns the block cache. If live scans are unavoidable, disable block caching for the scan, use large scanner caching, and run against a read replica cluster.
  • Snapshot sprawl. Unreleased snapshots pin archived HFiles; storage grows with no change in table size.
  • Wrong decoder. Sign-bit, endianness or salt-offset mistakes produce plausible wrong numbers. Compare a sample of decoded rows with application reads on every mapping change.
  • Version and TTL drift. The SQL copy shows rows that HBase considers expired, or loses history the application relies on.
  • Pinned engine forgotten. The old Trino release used for Phoenix stays in production for years without security fixes.
  • CDC ordering. Merges that trust arrival order apply an older cell over a newer one after a replication retry.

Trade-offs

ChoiceGainCost
Pinned Phoenix catalogLive data, little to buildUnpatched engine, scan load on region servers
Snapshot exportZero serving impact, typed Iceberg tablesDay-old data, full rewrite per run
CDC through replicationMinutes of lag, incrementalCell-to-row assembly, ordering, reconciliation
Custom Thrift serviceLive data on a current engineYou build and operate a connector service
Application events to KafkaSimplest fresh feedCovers only what the application emits

What to do next

  1. List every consumer that wants SQL over HBase and write down the freshness each one actually needs.
  2. If any consumer depends on a Phoenix connector, record the pinned engine version and set a retirement date.
  3. Write the row-key and column mapping for each table as versioned data, and test it against real rows, including negative numbers.
  4. Build a nightly snapshot-to-Iceberg export for the most-queried table and point its consumers at Trino's Iceberg catalog.
  5. Add snapshot cleanup and restore-directory cleanup to the same job, and alert on archive directory growth.
  6. Only for consumers needing minutes, evaluate CDC through a replication endpoint, with a periodic snapshot reconciliation.
  7. Read HBase Replication before designing that pipeline.
Key takeaway: Current Trino releases ship no HBase or Phoenix connector, so SQL over HBase is now an architecture choice rather than a catalog file. A pinned Phoenix catalog gives live data at the cost of an unpatched engine and load on serving clusters. Snapshot export to Iceberg gives day-old data with zero region-server load. CDC through replication gives minutes at the highest build cost. Whichever you choose, the row-key and value decoder decides whether the numbers are right.