Apache HBase

Apache HBase

Deep technical articles on this topic.

148Articles
148Topics covered
Articles in this category

All 115 articles, sorted alphabetically

ARTICLE · 001

Cloudera Operational DB (COD)

Managed HBase in Cloudera CDP: architecture, auto-scaling clusters, multi-tenancy, backup and snapshots, performance knobs, licensing, vendor lock-in …

Read article →
ARTICLE · 002

Flink + HBase, in depth: the SQL connector, upsert sinks, lookup joins with caching and checkpoint guarantees

Using Apache HBase with Apache Flink: how the connector maps rows to column families, the upsert sink and its at-least-once guarantee, lookup joins wi…

Read article →
ARTICLE · 003

HBase for Ad Tech, in depth: bid-path profile lookups, frequency caps, sharded budget counters and attribution

Designing HBase for a demand-side platform: hashed device keys, profile reads with region replicas inside the bid deadline, frequency-cap increments a…

Read article →
ARTICLE · 004

HBase Alerting Best Practices

A production guide to alerting on HBase: which metrics to monitor (Get/Put latency, compaction backlog, regions in transition, GC pauses), how to set …

Read article →
ARTICLE · 005

HBase Anti-Patterns to Avoid, in depth: a symptom-first catalogue of key, schema, client and operations mistakes

A symptom-first catalogue of HBase anti-patterns: monotonic keys, hot counters, too many column families, giant rows and cells, unbounded scans, queue…

Read article →
ARTICLE · 006

HBase backup and restore architecture

Deep-dive on HBase backup and restore: an occasional full backup that snapshots and exports a self-contained base image, frequent incremental backups …

Read article →
ARTICLE · 007

HBase Backups, in depth: designing a backup program from snapshots, exported copies, archived WALs, point-in-time replay and restore drills

How to protect HBase data as a program rather than a command: map each failure to the mechanism that survives it, combine snapshots, ExportSnapshot an…

Read article →
ARTICLE · 008

HBase Balancer

A deep dive on the HBase Balancer: how regions are distributed across RegionServers, balancing strategies by region count, request rate, and region si…

Read article →
ARTICLE · 009

HBase batching client operations, in depth: how batches fan out, the client throttles, partial failure, retries and sizing

An operations guide to HBase client batching: how Table.batch and BufferedMutator group mutations into multi RPCs, the four request checkers and their…

Read article →
ARTICLE · 010

HBase block cache architecture

Deep-dive on the HBase block cache: the read-path lookup and populate-on-miss cycle, the three priority levels (single-access, multi-access, in-memory…

Read article →
ARTICLE · 011

HBase BlockCache Tuning, in depth: sizing L1 and L2, eviction shares, bucket sizes, per-table controls and the metrics that prove it

A tuning guide for the HBase BlockCache: the heap and direct-memory budget, sizing L1 for index and bloom blocks and L2 for the hot set, LRU priority …

Read article →
ARTICLE · 012

HBase bloom filters architecture

Deep-dive on HBase bloom filters: how probabilistic per-HFile filters cut read amplification by skipping files that cannot hold a key, ROW vs ROWCOL s…

Read article →
ARTICLE · 013

HBase BucketCache -- caching without GC pain

Deep-dive on HBase BucketCache: the on-heap-cache GC-pressure problem, the block cache, LruBlockCache vs BucketCache, off-heap/SSD storage, fixed-size…

Read article →
ARTICLE · 014

HBase Bulk Load: Generating HFiles and LoadIncrementalHFiles

How HBase bulk load works: write HFiles with MapReduce or Spark, align them to regions, load with LoadIncrementalHFiles, why it skips the WAL and repl…

Read article →
ARTICLE · 015

HBase Cell-Level Security, in depth: cell ACLs, the early_out switch, covering checks on writes and designing per-record access

How HBase cell-level ACLs really work: storage as cell tags, why the default early_out=true makes them inert, read filtering, covering-permission chec…

Read article →
ARTICLE · 016

HBase Cell Tags and Visibility Labels, in depth: per-cell metadata in HFile v3, label expressions, scan label generators, cell ACL ordering and where tags get dropped

How HBase cell tags work and how visibility labels use them: HFile v3 tag storage, VisibilityController configuration and ordering, labels and authori…

Read article →
ARTICLE · 017

HBase Client Libraries, in depth: sync and async Java clients, the shaded artifact, connection registries, retries, buffered writes and non-JVM options

How to choose and operate an HBase client: what every client does (bootstrap, meta lookup, location cache, retries), the blocking Connection/Table API…

Read article →
ARTICLE · 018

HBase Cluster Topology, in depth: placing ZooKeeper, masters, RegionServers and HDFS so one failure stays small

How to lay out an HBase cluster: control-plane versus data-plane roles, co-locating RegionServers with DataNodes and short-circuit reads, HDFS rack aw…

Read article →
ARTICLE · 019

HBase Column Families, in depth: the cell format, versions, TTL, delete markers and how a read touches each family

What an HBase column family means at runtime: the cell format and its byte cost, versions, TTL and MIN_VERSIONS, the four delete marker types, deletes…

Read article →
ARTICLE · 020

HBase Column Family Design, in depth: when to split data into families, and how to configure each one

HBase column family design from first principles: what a family is physically, what families share and what they isolate, when a second family is just…

Read article →
ARTICLE · 021

HBase Compaction: Minor vs Major Compaction Explained

How HBase compaction works: the LSM write path, minor vs major compaction, tombstone and TTL cleanup, write amplification, throttling, and the tuning …

Read article →
ARTICLE · 022

HBase Minor + Major Compaction, in depth

How HBase minor and major compactions differ: what each rewrites and drops, the ExploringCompactionPolicy ratio test, promotion to major, weekly sched…

Read article →
ARTICLE · 023

HBase Compaction Tuning, in depth

HBase compaction tuning in depth: how ExploringCompactionPolicy selects files with the ratio test, every key property with verified 2.x defaults, pres…

Read article →
ARTICLE · 024

HBase coprocessor architecture, in depth: hosts, loading, hook ordering, failure isolation and safe patterns

How HBase 2.x coprocessors are hosted, loaded and ordered, where observer hooks fire relative to row locks and the WAL, how exceptions abort servers, …

Read article →
ARTICLE · 025

HBase Coprocessors

How coprocessors let you run user code inside the RegionServer for triggers, secondary indexes, and custom RPCs. Powerful and dangerous.

Read article →
ARTICLE · 026

HBase Coprocessors from Client, in depth: calling endpoints with Table and AsyncTable, ranges, errors and AggregationClient

The client side of HBase coprocessor endpoints: protobuf services and stubs, Table.coprocessorService for one row or a key range, the per-region fan-o…

Read article →
ARTICLE · 027

HBase Alternatives + Landscape 2026, in depth: what HBase guarantees, which workloads still belong on it, and where the rest should go

A decision guide to HBase alternatives in 2026: the contract HBase actually provides, the state of the project, a workload-shape classification, Bigta…

Read article →
ARTICLE · 028

HBase Denormalization Patterns, in depth: wide rows, tall rows, index tables, counters and keeping copies consistent

How and why to denormalize in Apache HBase: single-row atomicity and sorted rowkeys, five patterns (embedded children, tall child rows, duplicated ind…

Read article →
ARTICLE · 029

HBase Disaster Recovery, in depth: RPO and RTO per table, standby clusters, synchronous replication, failover and failback runbooks

How to design disaster recovery for Apache HBase at the site level: setting RPO and RTO per table, mapping failure classes to tools, the dependencies …

Read article →
ARTICLE · 030

HBase Disaster Recovery Strategy, in depth: choosing replication or snapshot export per table from measured WAL rates, link budgets, copy windows and drain time

How to build an HBase disaster recovery strategy with arithmetic: measure table size and peak WAL rate, assign async replication, snapshot export or s…

Read article →
ARTICLE · 031

HBase on Erasure-Coded HDFS, in depth: which directories can be erasure coded, why the WAL cannot, the ERASURE_CODING_POLICY table setting, migrating by major compaction and the read-path costs

An HBase-specific guide to erasure-coded HDFS: the HBase directory layout and which parts may be erasure coded, the hflush rule that keeps WALs replic…

Read article →
ARTICLE · 032

HBase for Event Tracking, in depth: timeline and index tables, reverse-time keys, TTL against cell timestamps and sizing

Designing an HBase event store from its queries: a salted per-user timeline keyed by reverse event time and event ID, a bucketed per-type index table,…

Read article →
ARTICLE · 033

HBase Filters in Depth: How the RegionServer Evaluates Them, Which Ones Seek, and How to Compose Them Safely

How HBase filters work inside the RegionServer scanner: the Filter hook sequence and ReturnCode values, which built-in filters seek and which read eve…

Read article →
ARTICLE · 034

HBase Client-Side Filters, in depth: building, serializing, testing and shipping filters, and when to post-filter instead

How HBase filters look from the client, based on the HBase 2.6 source: filters are built and serialized on the client but run on RegionServers. Covers…

Read article →
ARTICLE · 035

HBase for Graph Data, in depth: adjacency rows, edge tables, supernodes, batched traversal and JanusGraph

How to store and query graphs on HBase: the vertex-row adjacency layout versus a tall edge table, salted row keys and compact qualifiers, bucketing su…

Read article →
ARTICLE · 036

HBase GC Tuning, in depth: choosing G1, ZGC or Shenandoah for a RegionServer, and proving the choice with an A/B test

How to choose a garbage collector for HBase RegionServers: what the RegionServer asks of a collector, how G1, ZGC and Shenandoah work and fail, which …

Read article →
ARTICLE · 037

HBase JVM GC Tuning in Depth: Heap Layout, MSLAB, G1 Settings, Off-Heap Caching and Reading the Logs

How to tune garbage collection for HBase RegionServers from first principles: why a long pause is treated as death, how MemStore, block cache and RPC …

Read article →
ARTICLE · 038

HBase Hashing Row Keys, in depth: hash prefixes, full hashes and pre-splitting

How hashing HBase row keys fixes write hotspots: hash prefix vs full hash vs bucket salt, a byte-exact key builder in Java and Python, UniformSplit an…

Read article →
ARTICLE · 039

HBase HFile Format

The internal structure of HFile: data blocks, indices, bloom filters, and trailer. How the format enables fast point lookups and range scans over sort…

Read article →
ARTICLE · 040

HBase + Hive Integration, in depth: column mapping, what actually pushes down, snapshot reads, HFile loads and lifecycle traps

How Hive's HBaseStorageHandler works: one split per region, the column mapping syntax including families, prefixes, timestamps and binary encodin…

Read article →
ARTICLE · 041

HBase HMaster

How the HBase HMaster elects itself, initializes, drives assignment through AMv2 procedures, executes DDL, runs its chores, and why it sits off the da…

Read article →
ARTICLE · 042

HBase Hotspot Analysis, in depth: finding the hot server, region and key, and classifying the cause before you fix anything

A diagnostic workflow for HBase hotspots: a taxonomy of write, read, single-row and false hotspots; per-region JMX counters and why you must use delta…

Read article →
ARTICLE · 043

HBase Hotspot Fixes, in depth: traffic-midpoint splits, placement, quotas, read replicas, sharded rows and key migration

Fixing an HBase hotspot on a live table: matching the fix to a single hot row, hot range or monotonic tail; splitting at the traffic midpoint; balance…

Read article →
ARTICLE · 044

HBase Hotspotting: Row Key Design, Salting, Pre-Splitting

Why HBase regions hotspot and how to fix it: monotonic row keys, salting, hashing, field reordering, pre-splitting, and the spread vs range-scan trade…

Read article →
ARTICLE · 045

HBase In-Memory Column Family, in depth: what IN_MEMORY really does to the block cache

What the HBase IN_MEMORY column family flag actually does: block cache priorities, the eviction algorithm and force mode from the HBase 2.6 source, be…

Read article →
ARTICLE · 046

HBase Incremental Update Patterns, in depth

HBase incremental update patterns in depth: upsert, versioned and delta models, partial updates and NULLs, delete masking, source commit time as the c…

Read article →
ARTICLE · 047

HBase for IoT, in depth: a four-table data model for device fleets, event-time timestamps, late uploads, rollups and sizing

How to store device-fleet telemetry in HBase: read patterns first, raw telemetry, latest-state, rollup and registry tables, salted device-first row ke…

Read article →
ARTICLE · 048

HBase Java API, in depth: Put, Get, Scan and Delete semantics, atomic operations, batch results and Admin, with a tested data-access layer

The HBase 2.x Java client data API: how Put, Get, Delete and Scan behave, parsing Result and Cell, Increment, Append, CheckAndMutate and RowMutations,…

Read article →
ARTICLE · 049

HBase Log Aggregation for Ops, in depth: daemon log files, log4j2 layouts, slow-RPC and WAL signals, parsing, the truncation trap and the slow log

How to aggregate HBase Master and RegionServer logs for operations: file naming and rotation, log4j2 configuration, the responseTooSlow, Slow sync cos…

Read article →
ARTICLE · 050

HBase MemStore and Flushes

HBase MemStore internals: the skip list and MSLAB, every flush trigger including global heap pressure and WAL count, and the write-stall cascade.

Read article →
ARTICLE · 051

HBase MemStore Tuning, in depth: global watermarks, flush size and block multiplier, per-family flushes, WAL pressure, MSLAB cost and in-memory compaction

A knob-by-knob guide to tuning HBase MemStores: the global upper and lower watermarks, region flush size and block multiplier, the per-family flush po…

Read article →
ARTICLE · 052

HBase write path architecture

Deep-dive on HBase's write path: WAL append and sync durability semantics, MemStore skiplists and MVCC visibility, flush triggers and WAL retenti…

Read article →
ARTICLE · 053

HBase hbase:meta architecture

How HBase routes a row key to the RegionServer that serves it: the hbase:meta catalog table, its row schema, the ZooKeeper bootstrap pointer, the deat…

Read article →
ARTICLE · 054

HBase Metrics Deep Dive, in depth: counters, gauges and reset-per-snapshot histograms, and the arithmetic that turns them into correct answers

What HBase metrics actually measure and how to compute with them: counters, gauges and histograms in the RegionServer and IPC sources, why HBase perce…

Read article →
ARTICLE · 055

HBase Metrics and Monitoring, in depth: the metrics2 pipeline, JMX beans, the /prometheus endpoint and reading a cluster by layer

How HBase metrics work and how to read them: Hadoop metrics2 sources and JMX beans, /jmx?qry= and the /prometheus servlet, exporter configuration, cou…

Read article →
ARTICLE · 056

HBase Migration, in depth: moving live tables between clusters and platforms with snapshots, replication catch-up, hash-based verification and a reversible cutover

How to migrate HBase tables between clusters, data centres, distributions or to Bigtable without losing writes: inventory and bandwidth planning, a di…

Read article →
ARTICLE · 057

HBase MOB -- storing medium objects efficiently

Deep-dive on HBase MOB (Medium OBject storage): the medium-object write-amplification problem, separate MOB files, decoupled compaction, cell referenc…

Read article →
ARTICLE · 058

HBase Monitoring Metrics, in depth: headroom, measuring each RegionServer metric against the limit that stops it

HBase monitoring by headroom: each key RegionServer and Master metric paired with its configured limit and what HBase does at it, with defaults from t…

Read article →
ARTICLE · 059

HBase Multi-Tenancy, in depth: namespaces, ACLs, quotas and RegionServer groups combined into tenant tiers

How to run many teams on one HBase cluster: what each isolation layer controls (namespaces and ACLs, namespace limits, throttle quotas, space quotas, …

Read article →
ARTICLE · 060

HBase Region Normalizer architecture

Deep-dive on the HBase Region Normalizer: why region shape sets the load ceiling, how the Master chore computes split and merge plans against the tabl…

Read article →
ARTICLE · 061

HBase Off-Heap Memory, in depth: the RegionServer direct-memory budget, off-heap read and write paths, ByteBuffAllocator and sizing HBASE_OFFHEAPSIZE

How Apache HBase uses off-heap (direct) memory: why it exists, the off-heap read path with BucketCache and ByteBuffAllocator, off-heap MSLAB chunks fo…

Read article →
ARTICLE · 062

Operating HBase Off-Heap Memory, in depth: reconciling RegionServer RSS, finding buffer and native leaks, and rolling out safely

Operate HBase off-heap memory in production: the layers of RegionServer RSS, the MaxDirectMemorySize default trap, reconciling with Native Memory Trac…

Read article →
ARTICLE · 063

HBase on Kubernetes, in depth: stable identity, HDFS locality, container memory and drain-before-stop rolling restarts

How to run Apache HBase on Kubernetes: StatefulSets and headless Services for RegionServer identity, client reachability, DataNode co-location and sho…

Read article →
ARTICLE · 064

HBase Overview

The HBase data model (row key, column families, cells, versions), how it complements HDFS's sequential nature with random access, and where HBase…

Read article →
ARTICLE · 065

HBase + Ozone Integration, in depth: running HBase on Apache Ozone

How Apache HBase runs on Apache Ozone: the filesystem guarantees HBase needs (hsync, lease recovery, atomic rename), how Ozone provides them with ofs,…

Read article →
ARTICLE · 066

HBase Read Performance, in depth: a layer-by-layer procedure for finding and fixing p99 read latency

A practical guide to HBase read latency: why p99 matters under fan-out, how to split a slow read across client, RPC queue, handler, block cache, HDFS …

Read article →
ARTICLE · 067

HBase Write Performance, in depth: a layer-by-layer procedure for put latency, ingest throughput and write stalls

How to diagnose and fix HBase write performance: the cost of a Put at the client, RPC, WAL and MemStore layers, how flush and compaction backpressure …

Read article →
ARTICLE · 068

Apache Phoenix

How Phoenix maps SQL onto HBase: composite row keys, coprocessor pushdown, global versus local indexes, skip scans, and statistics-driven parallelism.

Read article →
ARTICLE · 069

HBase Procedure v2 architecture

Deep-dive on HBase Procedure v2 and AssignmentManager v2: the ProcedureExecutor, procedure store and WALs, parent-child trees, locks, TransitRegionSta…

Read article →
ARTICLE · 070

HBase Put, Get, Scan and Delete, in depth: the write path, MVCC read points, delete markers and the scan RPC protocol

What HBase Put, Get, Scan and Delete do inside a RegionServer: row locks, WAL durability, MVCC visibility, timestamps and versions, delete markers and…

Read article →
ARTICLE · 071

HBase quota throttling architecture

Deep-dive on HBase quota throttling: throttle quotas (req/s, bytes/s, size caps) scoped to user, table, and namespace; token-bucket enforcement at the…

Read article →
ARTICLE · 072

HBase read amplification, in depth: measuring and bounding the cost of every Get and Scan

Read amplification in HBase as a quantity you can measure and bound: file, block and cell amplification, a cost model for single-row Gets, how compact…

Read article →
ARTICLE · 073

HBase Read/Write Perf Tuning, in depth: tuning a RegionServer when reads and writes compete for heap, files, disks and handlers

How writes degrade reads in HBase and how to isolate them: the shared heap split, per-region flush size against global MemStore pressure, file counts …

Read article →
ARTICLE · 074

HBase RegionServer Architecture in Depth: how a server hosts, opens, splits, moves and recovers regions

HBase RegionServer architecture through the region: HDFS layout, shared WAL and sequence ids, the assignment state machine, open and close costs, spli…

Read article →
ARTICLE · 075

HBase Region Count Sizing, in depth: heap budgets, write-active regions, split thresholds and a worked capacity plan

How to size HBase region count deliberately: what each region and store costs in memstore, MSLAB, WAL and compaction, the write bound and data bound f…

Read article →
ARTICLE · 076

HBase Region Data Locality, in depth: how HDFS placement makes reads local, how moves and restarts destroy it, how to measure it and the cheapest ways to restore it

What HBase data locality is and why it matters: how HDFS replica placement puts a region's files on its own host, how balancer moves, restarts, c…

Read article →
ARTICLE · 077

HBase region replicas architecture

Deep-dive on HBase region replicas: read-only secondary copies of a region that serve stale, timeline-consistent reads for high availability and low p…

Read article →
ARTICLE · 078

HBase RegionServer

How the RegionServer serves reads and writes, why the WAL is the durability foundation, and how memstore, block cache, and HFiles interact per region.

Read article →
ARTICLE · 079

HBase Region Split: How Regions Split, Pre-Split and Merge

How HBase regions split past a size threshold, how split points and policies are chosen, how pre-splitting avoids hotspots, and when to disable splits…

Read article →
ARTICLE · 080

HBase replication architecture

Deep-dive on HBase replication: WAL edit shipping, filters, serial ordering, throttling, and DR drills.

Read article →
ARTICLE · 081

HBase RegionServer Groups architecture

Deep-dive on HBase RSGroups: the hbase:rsgroup metadata table and ZooKeeper mirror, RSGroupAdminEndpoint, the group-aware stochastic balancer, group-s…

Read article →
ARTICLE · 082

HBase Salting Row Keys, in depth: deterministic salt prefixes, bucket counts, pre-splits and scatter-gather reads

How to salt HBase row keys properly: why the salt must be a deterministic function of the natural key, what to hash, how to choose the bucket count fr…

Read article →
ARTICLE · 083

HBase Scans, in Depth: The Scanner Lifecycle, How Each RPC Is Sized, Leases and Heartbeats, and Making Range Reads Fast

How an HBase scan really runs: region-by-region scanner RPCs, the server-side merge across MemStore and HFiles, caching versus maxResultSize versus ba…

Read article →
ARTICLE · 084

HBase Schema Design Deep Dive, in depth: designing from access patterns, row key byte layout, tall versus wide, atomicity boundaries and index tables, with a worked messaging schema

How to design an HBase schema from its queries: the data model you are really designing against, a worked messaging schema, composite row keys and the…

Read article →
ARTICLE · 085

HBase Security, in depth: Kerberos, ZooKeeper SASL, delegation tokens, wire encryption and Ranger authorization

Securing Apache HBase end to end: following one request through Kerberos authentication, ZooKeeper SASL, SASL or native TLS RPC protection, delegation…

Read article →
ARTICLE · 086

HBase for Session Data, in depth

Build a durable HBase session store: hashed-token row keys, sliding expiry without losing cells, CheckAndMutate concurrency, log-out-everywhere indexi…

Read article →
ARTICLE · 087

HBase Shell, in depth: the JRuby REPL for data, admin and scripted operations

A practical guide to the HBase shell: how it talks to the cluster, reading and writing bytes correctly, scanning without hurting production, DDL and d…

Read article →
ARTICLE · 088

HBase Slow Query Analysis, in depth: slow logs, the ring buffer, hbase:slowlog, scan metrics and classifying slow calls

How to find and explain slow HBase queries: the responseTooSlow and responseTooLarge thresholds, the per-RegionServer ring buffer and the hbase:slowlo…

Read article →
ARTICLE · 089

HBase Snapshot Export/Import, in depth: ExportSnapshot internals, throttling, checksums and restore runbooks

How HBase ExportSnapshot really works: manifest-first staging, the MapReduce copy job, per-mapper bandwidth, checksum failures to object stores, impor…

Read article →
ARTICLE · 090

HBase Snapshots

How HBase snapshots create instant point-in-time views of a table via HFile references, and how to use them for backup, cloning, and restore.

Read article →
ARTICLE · 091

HBase + Spark Connector, in depth: catalogs, pushdown, HBaseContext and bulk load

How the Apache HBase Spark connector works: building it, JSON catalogs and column mappings, how row-key predicates become scan ranges and column predi…

Read article →
ARTICLE · 092

HBase Sparse Data Modeling, in depth: cell cost, null semantics, filters and blooms

How to model sparse data in HBase: why absent columns cost nothing while present cells pay full key overhead, a byte-level worked catalogue example, a…

Read article →
ARTICLE · 093

HBase region splitting architecture, in depth: the split procedure, reference files, daughter compaction and parent cleanup

The machinery behind an HBase region split: how a RegionServer decides and asks the master, SplitTableRegionProcedure's states and its point of n…

Read article →
ARTICLE · 094

HBase Split Management, in depth: split points from real keys, budgeted managed splitting and runbooks for split trouble

How to manage HBase region splits deliberately: what the default SteppingSplitPolicy does in numbers, computing pre-split boundaries from sampled keys…

Read article →
ARTICLE · 095

HBase StochasticLoadBalancer architecture

Deep-dive on the HBase StochasticLoadBalancer: cluster state snapshots, the weighted cost ensemble, candidate generators, the hill-climbing loop under…

Read article →
ARTICLE · 096

HBase Streaming Ingest Patterns, in depth: idempotent puts, flush-then-commit, counters, late events and backpressure

How to stream events from Kafka or Flink into HBase without duplicates or loss: at-least-once delivery with idempotent puts, event-derived row keys an…

Read article →
ARTICLE · 097

HBase Thrift gateway architecture, in depth: thrift versus thrift2, server models, worker pools, frame limits, stateless scans and load balancing

Inside the HBase Thrift gateway: the thrift and thrift2 servers and their IDLs, the four server implementations and which need framed transport, sizin…

Read article →
ARTICLE · 098

The HBase Thrift API, in depth: the thrift2 IDL as a contract, row atomicity with checkAndMutate, counters, partial failures, region locations and IDL version drift

A client developer's guide to the HBase thrift2 API: the THBaseService method surface grouped by purpose, the binary data model and flat cell lis…

Read article →
ARTICLE · 099

HBase Thrift and REST Gateways in Depth: How Requests Are Translated, Where Scanner State Lives, and Whose Identity HBase Sees

The HBase Thrift and REST gateways from the inside: the translation path to the Java client, the REST resource model with base64 JSON, stateful versus…

Read article →
ARTICLE · 100

HBase for Time-Series, in depth: bucketed row keys, packed Gorilla-style blocks, FIFO-compacted raw data, resolution routing and sizing

Building a metrics store on HBase: series-first versus time-first row keys, salting, raw points in a FIFO-compacted family, packing closed hours into …

Read article →
ARTICLE · 101

HBase Troubleshooting, in depth: triage order, RegionServer aborts, stuck regions and HBCK2, write blocking, slow reads and safe fixes

A symptom-first HBase troubleshooting guide: how to narrow the blast radius, the evidence to collect in the first five minutes, RegionServer aborts fr…

Read article →
ARTICLE · 102

HBase TTL and MAX_VERSIONS

How HBase TTL, VERSIONS, MIN_VERSIONS and KEEP_DELETED_CELLS interact: why expired data is still on disk, and what only major compaction can reclaim.

Read article →
ARTICLE · 103

HBase Upgrade, in depth: rolling upgrades, version-specific paths, rollback limits and a runbook that survives production

How to upgrade Apache HBase safely: what an upgrade actually changes (binaries, procedure store, meta, wire protocol, coprocessors), the compatibility…

Read article →
ARTICLE · 104

HBase vs Google Bigtable, in depth

HBase and Cloud Bigtable compared as systems: storage attached to RegionServers versus tablets on shared Colossus storage, failure recovery, an API co…

Read article →
ARTICLE · 105

HBase vs Cassandra, in depth: one owner per region against leaderless replicas, failure and repair, data modelling, conditional writes, multi-datacenter and how to choose

HBase and Cassandra compared at the architecture level: single RegionServer ownership versus leaderless replicas with tunable consistency, what happen…

Read article →
ARTICLE · 106

HBase vs DynamoDB, in depth: key models, ordering, atomicity, capacity and hot keys, translating a row-key design, and when to choose which

A practical comparison of Apache HBase and Amazon DynamoDB: how each turns a key into a storage location, global ordering versus per-partition orderin…

Read article →
ARTICLE · 107

HBase Write-Ahead Log

The HBase WAL as a subsystem: WALKey and WALEdit on disk, the asyncfs, filesystem and multiwal providers, group-commit sync mechanics, log rolling and…

Read article →
ARTICLE · 108

HBase WAL Durability Levels: Architecture Deep-Dive

How HBase's WAL durability levels — SKIP, ASYNC, SYNC, and FSYNC — trade write speed for crash safety per table or per write, with recovery, repl…

Read article →
ARTICLE · 109

HBase WAL Splitting: How RegionServer Crash Recovery Works

How HBase recovers a crashed RegionServer: ZooKeeper failure detection, distributed WAL splitting into recovered.edits, region reassignment and edit r…

Read article →
ARTICLE · 110

HBase Wide Row Limits, in depth: every ceiling a growing row hits, and how to find and fix rows that hit them

The limits a wide HBase row runs into, with properties and defaults: cell size, row locks and memstore blocking, unsplittable regions, hbase.table.max…

Read article →
ARTICLE · 111

HBase Wide vs Tall Tables, in depth: what the choice really changes, sized examples and the middle ground

Wide versus tall HBase table design from the storage format up: why every cell repeats its row key, the three things the choice changes (atomicity, di…

Read article →
ARTICLE · 112

Iceberg and HBase Integration, in depth: snapshot exports, CDC merges, bulk loads back and keeping the copies consistent

How to integrate HBase with Apache Iceberg when no native connector exists: cell-to-column mapping, snapshot exports via TableSnapshotInputFormat, rep…

Read article →
ARTICLE · 113

Impala + HBase Integration, in depth: table mapping, row-key pushdown, reading EXPLAIN for SCAN HBASE, joins with Parquet, writes and when to choose Kudu instead

A practical guide to querying HBase from Apache Impala: how the Hive storage handler maps columns, which predicates become start/stop keys or HBase fi…

Read article →
ARTICLE · 114

OpenTSDB, in depth: UIDs, row keys, salting, compaction and query semantics on HBase

How OpenTSDB stores time series in HBase: stateless TSDs, the UID dictionary, the row-key and qualifier byte layout, the put API, TSD compaction versu…

Read article →
ARTICLE · 115

Presto/Trino + HBase, in depth: why current Trino has no HBase connector, and how to query HBase data with Phoenix, snapshots or CDC

Query HBase data from Trino or Presto in 2026: the missing HBase and Phoenix connectors, why SQL engines fit HBase poorly, a pinned Phoenix catalog, o…

Read article →