All 115 articles, sorted alphabetically
Cloudera Operational DB (COD)
Managed HBase in Cloudera CDP: architecture, auto-scaling clusters, multi-tenancy, backup and snapshots, performance knobs, licensing, vendor lock-in …
Read article →Flink + HBase, in depth: the SQL connector, upsert sinks, lookup joins with caching and checkpoint guarantees
Using Apache HBase with Apache Flink: how the connector maps rows to column families, the upsert sink and its at-least-once guarantee, lookup joins wi…
Read article →HBase for Ad Tech, in depth: bid-path profile lookups, frequency caps, sharded budget counters and attribution
Designing HBase for a demand-side platform: hashed device keys, profile reads with region replicas inside the bid deadline, frequency-cap increments a…
Read article →HBase Alerting Best Practices
A production guide to alerting on HBase: which metrics to monitor (Get/Put latency, compaction backlog, regions in transition, GC pauses), how to set …
Read article →HBase Anti-Patterns to Avoid, in depth: a symptom-first catalogue of key, schema, client and operations mistakes
A symptom-first catalogue of HBase anti-patterns: monotonic keys, hot counters, too many column families, giant rows and cells, unbounded scans, queue…
Read article →HBase backup and restore architecture
Deep-dive on HBase backup and restore: an occasional full backup that snapshots and exports a self-contained base image, frequent incremental backups …
Read article →HBase Backups, in depth: designing a backup program from snapshots, exported copies, archived WALs, point-in-time replay and restore drills
How to protect HBase data as a program rather than a command: map each failure to the mechanism that survives it, combine snapshots, ExportSnapshot an…
Read article →HBase Balancer
A deep dive on the HBase Balancer: how regions are distributed across RegionServers, balancing strategies by region count, request rate, and region si…
Read article →HBase batching client operations, in depth: how batches fan out, the client throttles, partial failure, retries and sizing
An operations guide to HBase client batching: how Table.batch and BufferedMutator group mutations into multi RPCs, the four request checkers and their…
Read article →HBase block cache architecture
Deep-dive on the HBase block cache: the read-path lookup and populate-on-miss cycle, the three priority levels (single-access, multi-access, in-memory…
Read article →HBase BlockCache Tuning, in depth: sizing L1 and L2, eviction shares, bucket sizes, per-table controls and the metrics that prove it
A tuning guide for the HBase BlockCache: the heap and direct-memory budget, sizing L1 for index and bloom blocks and L2 for the hot set, LRU priority …
Read article →HBase bloom filters architecture
Deep-dive on HBase bloom filters: how probabilistic per-HFile filters cut read amplification by skipping files that cannot hold a key, ROW vs ROWCOL s…
Read article →HBase BucketCache -- caching without GC pain
Deep-dive on HBase BucketCache: the on-heap-cache GC-pressure problem, the block cache, LruBlockCache vs BucketCache, off-heap/SSD storage, fixed-size…
Read article →HBase Bulk Load: Generating HFiles and LoadIncrementalHFiles
How HBase bulk load works: write HFiles with MapReduce or Spark, align them to regions, load with LoadIncrementalHFiles, why it skips the WAL and repl…
Read article →HBase Cell-Level Security, in depth: cell ACLs, the early_out switch, covering checks on writes and designing per-record access
How HBase cell-level ACLs really work: storage as cell tags, why the default early_out=true makes them inert, read filtering, covering-permission chec…
Read article →HBase Cell Tags and Visibility Labels, in depth: per-cell metadata in HFile v3, label expressions, scan label generators, cell ACL ordering and where tags get dropped
How HBase cell tags work and how visibility labels use them: HFile v3 tag storage, VisibilityController configuration and ordering, labels and authori…
Read article →HBase Client Libraries, in depth: sync and async Java clients, the shaded artifact, connection registries, retries, buffered writes and non-JVM options
How to choose and operate an HBase client: what every client does (bootstrap, meta lookup, location cache, retries), the blocking Connection/Table API…
Read article →HBase Cluster Topology, in depth: placing ZooKeeper, masters, RegionServers and HDFS so one failure stays small
How to lay out an HBase cluster: control-plane versus data-plane roles, co-locating RegionServers with DataNodes and short-circuit reads, HDFS rack aw…
Read article →HBase Column Families, in depth: the cell format, versions, TTL, delete markers and how a read touches each family
What an HBase column family means at runtime: the cell format and its byte cost, versions, TTL and MIN_VERSIONS, the four delete marker types, deletes…
Read article →HBase Column Family Design, in depth: when to split data into families, and how to configure each one
HBase column family design from first principles: what a family is physically, what families share and what they isolate, when a second family is just…
Read article →HBase Compaction: Minor vs Major Compaction Explained
How HBase compaction works: the LSM write path, minor vs major compaction, tombstone and TTL cleanup, write amplification, throttling, and the tuning …
Read article →HBase Minor + Major Compaction, in depth
How HBase minor and major compactions differ: what each rewrites and drops, the ExploringCompactionPolicy ratio test, promotion to major, weekly sched…
Read article →HBase Compaction Tuning, in depth
HBase compaction tuning in depth: how ExploringCompactionPolicy selects files with the ratio test, every key property with verified 2.x defaults, pres…
Read article →HBase coprocessor architecture, in depth: hosts, loading, hook ordering, failure isolation and safe patterns
How HBase 2.x coprocessors are hosted, loaded and ordered, where observer hooks fire relative to row locks and the WAL, how exceptions abort servers, …
Read article →HBase Coprocessors
How coprocessors let you run user code inside the RegionServer for triggers, secondary indexes, and custom RPCs. Powerful and dangerous.
Read article →HBase Coprocessors from Client, in depth: calling endpoints with Table and AsyncTable, ranges, errors and AggregationClient
The client side of HBase coprocessor endpoints: protobuf services and stubs, Table.coprocessorService for one row or a key range, the per-region fan-o…
Read article →HBase Alternatives + Landscape 2026, in depth: what HBase guarantees, which workloads still belong on it, and where the rest should go
A decision guide to HBase alternatives in 2026: the contract HBase actually provides, the state of the project, a workload-shape classification, Bigta…
Read article →HBase Denormalization Patterns, in depth: wide rows, tall rows, index tables, counters and keeping copies consistent
How and why to denormalize in Apache HBase: single-row atomicity and sorted rowkeys, five patterns (embedded children, tall child rows, duplicated ind…
Read article →HBase Disaster Recovery, in depth: RPO and RTO per table, standby clusters, synchronous replication, failover and failback runbooks
How to design disaster recovery for Apache HBase at the site level: setting RPO and RTO per table, mapping failure classes to tools, the dependencies …
Read article →HBase Disaster Recovery Strategy, in depth: choosing replication or snapshot export per table from measured WAL rates, link budgets, copy windows and drain time
How to build an HBase disaster recovery strategy with arithmetic: measure table size and peak WAL rate, assign async replication, snapshot export or s…
Read article →HBase on Erasure-Coded HDFS, in depth: which directories can be erasure coded, why the WAL cannot, the ERASURE_CODING_POLICY table setting, migrating by major compaction and the read-path costs
An HBase-specific guide to erasure-coded HDFS: the HBase directory layout and which parts may be erasure coded, the hflush rule that keeps WALs replic…
Read article →HBase for Event Tracking, in depth: timeline and index tables, reverse-time keys, TTL against cell timestamps and sizing
Designing an HBase event store from its queries: a salted per-user timeline keyed by reverse event time and event ID, a bucketed per-type index table,…
Read article →HBase Filters in Depth: How the RegionServer Evaluates Them, Which Ones Seek, and How to Compose Them Safely
How HBase filters work inside the RegionServer scanner: the Filter hook sequence and ReturnCode values, which built-in filters seek and which read eve…
Read article →HBase Client-Side Filters, in depth: building, serializing, testing and shipping filters, and when to post-filter instead
How HBase filters look from the client, based on the HBase 2.6 source: filters are built and serialized on the client but run on RegionServers. Covers…
Read article →HBase for Graph Data, in depth: adjacency rows, edge tables, supernodes, batched traversal and JanusGraph
How to store and query graphs on HBase: the vertex-row adjacency layout versus a tall edge table, salted row keys and compact qualifiers, bucketing su…
Read article →HBase GC Tuning, in depth: choosing G1, ZGC or Shenandoah for a RegionServer, and proving the choice with an A/B test
How to choose a garbage collector for HBase RegionServers: what the RegionServer asks of a collector, how G1, ZGC and Shenandoah work and fail, which …
Read article →HBase JVM GC Tuning in Depth: Heap Layout, MSLAB, G1 Settings, Off-Heap Caching and Reading the Logs
How to tune garbage collection for HBase RegionServers from first principles: why a long pause is treated as death, how MemStore, block cache and RPC …
Read article →HBase Hashing Row Keys, in depth: hash prefixes, full hashes and pre-splitting
How hashing HBase row keys fixes write hotspots: hash prefix vs full hash vs bucket salt, a byte-exact key builder in Java and Python, UniformSplit an…
Read article →HBase HFile Format
The internal structure of HFile: data blocks, indices, bloom filters, and trailer. How the format enables fast point lookups and range scans over sort…
Read article →HBase + Hive Integration, in depth: column mapping, what actually pushes down, snapshot reads, HFile loads and lifecycle traps
How Hive's HBaseStorageHandler works: one split per region, the column mapping syntax including families, prefixes, timestamps and binary encodin…
Read article →HBase HMaster
How the HBase HMaster elects itself, initializes, drives assignment through AMv2 procedures, executes DDL, runs its chores, and why it sits off the da…
Read article →HBase Hotspot Analysis, in depth: finding the hot server, region and key, and classifying the cause before you fix anything
A diagnostic workflow for HBase hotspots: a taxonomy of write, read, single-row and false hotspots; per-region JMX counters and why you must use delta…
Read article →HBase Hotspot Fixes, in depth: traffic-midpoint splits, placement, quotas, read replicas, sharded rows and key migration
Fixing an HBase hotspot on a live table: matching the fix to a single hot row, hot range or monotonic tail; splitting at the traffic midpoint; balance…
Read article →HBase Hotspotting: Row Key Design, Salting, Pre-Splitting
Why HBase regions hotspot and how to fix it: monotonic row keys, salting, hashing, field reordering, pre-splitting, and the spread vs range-scan trade…
Read article →HBase In-Memory Column Family, in depth: what IN_MEMORY really does to the block cache
What the HBase IN_MEMORY column family flag actually does: block cache priorities, the eviction algorithm and force mode from the HBase 2.6 source, be…
Read article →HBase Incremental Update Patterns, in depth
HBase incremental update patterns in depth: upsert, versioned and delta models, partial updates and NULLs, delete masking, source commit time as the c…
Read article →HBase for IoT, in depth: a four-table data model for device fleets, event-time timestamps, late uploads, rollups and sizing
How to store device-fleet telemetry in HBase: read patterns first, raw telemetry, latest-state, rollup and registry tables, salted device-first row ke…
Read article →HBase Java API, in depth: Put, Get, Scan and Delete semantics, atomic operations, batch results and Admin, with a tested data-access layer
The HBase 2.x Java client data API: how Put, Get, Delete and Scan behave, parsing Result and Cell, Increment, Append, CheckAndMutate and RowMutations,…
Read article →HBase Log Aggregation for Ops, in depth: daemon log files, log4j2 layouts, slow-RPC and WAL signals, parsing, the truncation trap and the slow log
How to aggregate HBase Master and RegionServer logs for operations: file naming and rotation, log4j2 configuration, the responseTooSlow, Slow sync cos…
Read article →HBase MemStore and Flushes
HBase MemStore internals: the skip list and MSLAB, every flush trigger including global heap pressure and WAL count, and the write-stall cascade.
Read article →HBase MemStore Tuning, in depth: global watermarks, flush size and block multiplier, per-family flushes, WAL pressure, MSLAB cost and in-memory compaction
A knob-by-knob guide to tuning HBase MemStores: the global upper and lower watermarks, region flush size and block multiplier, the per-family flush po…
Read article →HBase write path architecture
Deep-dive on HBase's write path: WAL append and sync durability semantics, MemStore skiplists and MVCC visibility, flush triggers and WAL retenti…
Read article →HBase hbase:meta architecture
How HBase routes a row key to the RegionServer that serves it: the hbase:meta catalog table, its row schema, the ZooKeeper bootstrap pointer, the deat…
Read article →HBase Metrics Deep Dive, in depth: counters, gauges and reset-per-snapshot histograms, and the arithmetic that turns them into correct answers
What HBase metrics actually measure and how to compute with them: counters, gauges and histograms in the RegionServer and IPC sources, why HBase perce…
Read article →HBase Metrics and Monitoring, in depth: the metrics2 pipeline, JMX beans, the /prometheus endpoint and reading a cluster by layer
How HBase metrics work and how to read them: Hadoop metrics2 sources and JMX beans, /jmx?qry= and the /prometheus servlet, exporter configuration, cou…
Read article →HBase Migration, in depth: moving live tables between clusters and platforms with snapshots, replication catch-up, hash-based verification and a reversible cutover
How to migrate HBase tables between clusters, data centres, distributions or to Bigtable without losing writes: inventory and bandwidth planning, a di…
Read article →HBase MOB -- storing medium objects efficiently
Deep-dive on HBase MOB (Medium OBject storage): the medium-object write-amplification problem, separate MOB files, decoupled compaction, cell referenc…
Read article →HBase Monitoring Metrics, in depth: headroom, measuring each RegionServer metric against the limit that stops it
HBase monitoring by headroom: each key RegionServer and Master metric paired with its configured limit and what HBase does at it, with defaults from t…
Read article →HBase Multi-Tenancy, in depth: namespaces, ACLs, quotas and RegionServer groups combined into tenant tiers
How to run many teams on one HBase cluster: what each isolation layer controls (namespaces and ACLs, namespace limits, throttle quotas, space quotas, …
Read article →HBase Region Normalizer architecture
Deep-dive on the HBase Region Normalizer: why region shape sets the load ceiling, how the Master chore computes split and merge plans against the tabl…
Read article →HBase Off-Heap Memory, in depth: the RegionServer direct-memory budget, off-heap read and write paths, ByteBuffAllocator and sizing HBASE_OFFHEAPSIZE
How Apache HBase uses off-heap (direct) memory: why it exists, the off-heap read path with BucketCache and ByteBuffAllocator, off-heap MSLAB chunks fo…
Read article →Operating HBase Off-Heap Memory, in depth: reconciling RegionServer RSS, finding buffer and native leaks, and rolling out safely
Operate HBase off-heap memory in production: the layers of RegionServer RSS, the MaxDirectMemorySize default trap, reconciling with Native Memory Trac…
Read article →HBase on Kubernetes, in depth: stable identity, HDFS locality, container memory and drain-before-stop rolling restarts
How to run Apache HBase on Kubernetes: StatefulSets and headless Services for RegionServer identity, client reachability, DataNode co-location and sho…
Read article →HBase Overview
The HBase data model (row key, column families, cells, versions), how it complements HDFS's sequential nature with random access, and where HBase…
Read article →HBase + Ozone Integration, in depth: running HBase on Apache Ozone
How Apache HBase runs on Apache Ozone: the filesystem guarantees HBase needs (hsync, lease recovery, atomic rename), how Ozone provides them with ofs,…
Read article →HBase Read Performance, in depth: a layer-by-layer procedure for finding and fixing p99 read latency
A practical guide to HBase read latency: why p99 matters under fan-out, how to split a slow read across client, RPC queue, handler, block cache, HDFS …
Read article →HBase Write Performance, in depth: a layer-by-layer procedure for put latency, ingest throughput and write stalls
How to diagnose and fix HBase write performance: the cost of a Put at the client, RPC, WAL and MemStore layers, how flush and compaction backpressure …
Read article →Apache Phoenix
How Phoenix maps SQL onto HBase: composite row keys, coprocessor pushdown, global versus local indexes, skip scans, and statistics-driven parallelism.
Read article →HBase Procedure v2 architecture
Deep-dive on HBase Procedure v2 and AssignmentManager v2: the ProcedureExecutor, procedure store and WALs, parent-child trees, locks, TransitRegionSta…
Read article →HBase Put, Get, Scan and Delete, in depth: the write path, MVCC read points, delete markers and the scan RPC protocol
What HBase Put, Get, Scan and Delete do inside a RegionServer: row locks, WAL durability, MVCC visibility, timestamps and versions, delete markers and…
Read article →HBase quota throttling architecture
Deep-dive on HBase quota throttling: throttle quotas (req/s, bytes/s, size caps) scoped to user, table, and namespace; token-bucket enforcement at the…
Read article →HBase read amplification, in depth: measuring and bounding the cost of every Get and Scan
Read amplification in HBase as a quantity you can measure and bound: file, block and cell amplification, a cost model for single-row Gets, how compact…
Read article →HBase Read/Write Perf Tuning, in depth: tuning a RegionServer when reads and writes compete for heap, files, disks and handlers
How writes degrade reads in HBase and how to isolate them: the shared heap split, per-region flush size against global MemStore pressure, file counts …
Read article →HBase RegionServer Architecture in Depth: how a server hosts, opens, splits, moves and recovers regions
HBase RegionServer architecture through the region: HDFS layout, shared WAL and sequence ids, the assignment state machine, open and close costs, spli…
Read article →HBase Region Count Sizing, in depth: heap budgets, write-active regions, split thresholds and a worked capacity plan
How to size HBase region count deliberately: what each region and store costs in memstore, MSLAB, WAL and compaction, the write bound and data bound f…
Read article →HBase Region Data Locality, in depth: how HDFS placement makes reads local, how moves and restarts destroy it, how to measure it and the cheapest ways to restore it
What HBase data locality is and why it matters: how HDFS replica placement puts a region's files on its own host, how balancer moves, restarts, c…
Read article →HBase region replicas architecture
Deep-dive on HBase region replicas: read-only secondary copies of a region that serve stale, timeline-consistent reads for high availability and low p…
Read article →HBase RegionServer
How the RegionServer serves reads and writes, why the WAL is the durability foundation, and how memstore, block cache, and HFiles interact per region.
Read article →HBase Region Split: How Regions Split, Pre-Split and Merge
How HBase regions split past a size threshold, how split points and policies are chosen, how pre-splitting avoids hotspots, and when to disable splits…
Read article →HBase replication architecture
Deep-dive on HBase replication: WAL edit shipping, filters, serial ordering, throttling, and DR drills.
Read article →HBase RegionServer Groups architecture
Deep-dive on HBase RSGroups: the hbase:rsgroup metadata table and ZooKeeper mirror, RSGroupAdminEndpoint, the group-aware stochastic balancer, group-s…
Read article →HBase Salting Row Keys, in depth: deterministic salt prefixes, bucket counts, pre-splits and scatter-gather reads
How to salt HBase row keys properly: why the salt must be a deterministic function of the natural key, what to hash, how to choose the bucket count fr…
Read article →HBase Scans, in Depth: The Scanner Lifecycle, How Each RPC Is Sized, Leases and Heartbeats, and Making Range Reads Fast
How an HBase scan really runs: region-by-region scanner RPCs, the server-side merge across MemStore and HFiles, caching versus maxResultSize versus ba…
Read article →HBase Schema Design Deep Dive, in depth: designing from access patterns, row key byte layout, tall versus wide, atomicity boundaries and index tables, with a worked messaging schema
How to design an HBase schema from its queries: the data model you are really designing against, a worked messaging schema, composite row keys and the…
Read article →HBase Security, in depth: Kerberos, ZooKeeper SASL, delegation tokens, wire encryption and Ranger authorization
Securing Apache HBase end to end: following one request through Kerberos authentication, ZooKeeper SASL, SASL or native TLS RPC protection, delegation…
Read article →HBase for Session Data, in depth
Build a durable HBase session store: hashed-token row keys, sliding expiry without losing cells, CheckAndMutate concurrency, log-out-everywhere indexi…
Read article →HBase Shell, in depth: the JRuby REPL for data, admin and scripted operations
A practical guide to the HBase shell: how it talks to the cluster, reading and writing bytes correctly, scanning without hurting production, DDL and d…
Read article →HBase Slow Query Analysis, in depth: slow logs, the ring buffer, hbase:slowlog, scan metrics and classifying slow calls
How to find and explain slow HBase queries: the responseTooSlow and responseTooLarge thresholds, the per-RegionServer ring buffer and the hbase:slowlo…
Read article →HBase Snapshot Export/Import, in depth: ExportSnapshot internals, throttling, checksums and restore runbooks
How HBase ExportSnapshot really works: manifest-first staging, the MapReduce copy job, per-mapper bandwidth, checksum failures to object stores, impor…
Read article →HBase Snapshots
How HBase snapshots create instant point-in-time views of a table via HFile references, and how to use them for backup, cloning, and restore.
Read article →HBase + Spark Connector, in depth: catalogs, pushdown, HBaseContext and bulk load
How the Apache HBase Spark connector works: building it, JSON catalogs and column mappings, how row-key predicates become scan ranges and column predi…
Read article →HBase Sparse Data Modeling, in depth: cell cost, null semantics, filters and blooms
How to model sparse data in HBase: why absent columns cost nothing while present cells pay full key overhead, a byte-level worked catalogue example, a…
Read article →HBase region splitting architecture, in depth: the split procedure, reference files, daughter compaction and parent cleanup
The machinery behind an HBase region split: how a RegionServer decides and asks the master, SplitTableRegionProcedure's states and its point of n…
Read article →HBase Split Management, in depth: split points from real keys, budgeted managed splitting and runbooks for split trouble
How to manage HBase region splits deliberately: what the default SteppingSplitPolicy does in numbers, computing pre-split boundaries from sampled keys…
Read article →HBase StochasticLoadBalancer architecture
Deep-dive on the HBase StochasticLoadBalancer: cluster state snapshots, the weighted cost ensemble, candidate generators, the hill-climbing loop under…
Read article →HBase Streaming Ingest Patterns, in depth: idempotent puts, flush-then-commit, counters, late events and backpressure
How to stream events from Kafka or Flink into HBase without duplicates or loss: at-least-once delivery with idempotent puts, event-derived row keys an…
Read article →HBase Thrift gateway architecture, in depth: thrift versus thrift2, server models, worker pools, frame limits, stateless scans and load balancing
Inside the HBase Thrift gateway: the thrift and thrift2 servers and their IDLs, the four server implementations and which need framed transport, sizin…
Read article →The HBase Thrift API, in depth: the thrift2 IDL as a contract, row atomicity with checkAndMutate, counters, partial failures, region locations and IDL version drift
A client developer's guide to the HBase thrift2 API: the THBaseService method surface grouped by purpose, the binary data model and flat cell lis…
Read article →HBase Thrift and REST Gateways in Depth: How Requests Are Translated, Where Scanner State Lives, and Whose Identity HBase Sees
The HBase Thrift and REST gateways from the inside: the translation path to the Java client, the REST resource model with base64 JSON, stateful versus…
Read article →HBase for Time-Series, in depth: bucketed row keys, packed Gorilla-style blocks, FIFO-compacted raw data, resolution routing and sizing
Building a metrics store on HBase: series-first versus time-first row keys, salting, raw points in a FIFO-compacted family, packing closed hours into …
Read article →HBase Troubleshooting, in depth: triage order, RegionServer aborts, stuck regions and HBCK2, write blocking, slow reads and safe fixes
A symptom-first HBase troubleshooting guide: how to narrow the blast radius, the evidence to collect in the first five minutes, RegionServer aborts fr…
Read article →HBase TTL and MAX_VERSIONS
How HBase TTL, VERSIONS, MIN_VERSIONS and KEEP_DELETED_CELLS interact: why expired data is still on disk, and what only major compaction can reclaim.
Read article →HBase Upgrade, in depth: rolling upgrades, version-specific paths, rollback limits and a runbook that survives production
How to upgrade Apache HBase safely: what an upgrade actually changes (binaries, procedure store, meta, wire protocol, coprocessors), the compatibility…
Read article →HBase vs Google Bigtable, in depth
HBase and Cloud Bigtable compared as systems: storage attached to RegionServers versus tablets on shared Colossus storage, failure recovery, an API co…
Read article →HBase vs Cassandra, in depth: one owner per region against leaderless replicas, failure and repair, data modelling, conditional writes, multi-datacenter and how to choose
HBase and Cassandra compared at the architecture level: single RegionServer ownership versus leaderless replicas with tunable consistency, what happen…
Read article →HBase vs DynamoDB, in depth: key models, ordering, atomicity, capacity and hot keys, translating a row-key design, and when to choose which
A practical comparison of Apache HBase and Amazon DynamoDB: how each turns a key into a storage location, global ordering versus per-partition orderin…
Read article →HBase Write-Ahead Log
The HBase WAL as a subsystem: WALKey and WALEdit on disk, the asyncfs, filesystem and multiwal providers, group-commit sync mechanics, log rolling and…
Read article →HBase WAL Durability Levels: Architecture Deep-Dive
How HBase's WAL durability levels — SKIP, ASYNC, SYNC, and FSYNC — trade write speed for crash safety per table or per write, with recovery, repl…
Read article →HBase WAL Splitting: How RegionServer Crash Recovery Works
How HBase recovers a crashed RegionServer: ZooKeeper failure detection, distributed WAL splitting into recovered.edits, region reassignment and edit r…
Read article →HBase Wide Row Limits, in depth: every ceiling a growing row hits, and how to find and fix rows that hit them
The limits a wide HBase row runs into, with properties and defaults: cell size, row locks and memstore blocking, unsplittable regions, hbase.table.max…
Read article →HBase Wide vs Tall Tables, in depth: what the choice really changes, sized examples and the middle ground
Wide versus tall HBase table design from the storage format up: why every cell repeats its row key, the three things the choice changes (atomicity, di…
Read article →Iceberg and HBase Integration, in depth: snapshot exports, CDC merges, bulk loads back and keeping the copies consistent
How to integrate HBase with Apache Iceberg when no native connector exists: cell-to-column mapping, snapshot exports via TableSnapshotInputFormat, rep…
Read article →Impala + HBase Integration, in depth: table mapping, row-key pushdown, reading EXPLAIN for SCAN HBASE, joins with Parquet, writes and when to choose Kudu instead
A practical guide to querying HBase from Apache Impala: how the Hive storage handler maps columns, which predicates become start/stop keys or HBase fi…
Read article →OpenTSDB, in depth: UIDs, row keys, salting, compaction and query semantics on HBase
How OpenTSDB stores time series in HBase: stateless TSDs, the UID dictionary, the row-key and qualifier byte layout, the put API, TSD compaction versu…
Read article →Presto/Trino + HBase, in depth: why current Trino has no HBase connector, and how to query HBase data with Phoenix, snapshots or CDC
Query HBase data from Trino or Presto in 2026: the missing HBase and Phoenix connectors, why SQL engines fit HBase poorly, a pinned Phoenix catalog, o…
Read article →