HBase was built for racks of long-lived machines: RegionServers with stable hostnames, local disks that the HDFS DataNode on the same box fills with the RegionServer's own files, and operators who drain a node before they reboot it. Kubernetes assumes the opposite. Pods are cattle, IPs change on every restart, the scheduler moves work wherever it likes, and the default way to stop a container is a signal followed by a kill. Running HBase on Kubernetes is therefore not about writing manifests. It is about restoring, one by one, the four properties HBase quietly depends on: stable identity, durable shared storage, data locality, and orderly shutdown.

This article maps each property to the Kubernetes primitive that restores it, with manifests you can adapt, a worked rolling restart and the common failure modes. It assumes you know what a RegionServer, the WAL and the HMaster do; if not, read the RegionServer deep dive first. The decisions are the same whether you hand-roll StatefulSets or adopt an operator.

What HBase needs from its platform

Start by listing what HBase actually requires from whatever runs it, because each requirement maps to a concrete Kubernetes decision.

HBase assumptionWhy it mattersKubernetes answer
A RegionServer has a stable namehbase:meta stores region locations as host, port and start code; clients cache themStatefulSet plus a headless Service gives each pod a stable DNS name
Durable state lives in HDFS, not on the RegionServerWAL and HFiles survive a RegionServer crash and are replayed elsewhereDataNodes and JournalNodes on PersistentVolumes; RegionServers need no PV
Reads are local when possibleShort-circuit reads skip the DataNode network path for local blocksCo-locate RegionServer and DataNode pods on the same node and share a socket directory
Shutdown is orderlyAn unclean stop forces a WAL split and leaves regions offline until reassignedpreStop hook that unloads regions, long grace period, PodDisruptionBudget
Coordination is fast and quietZooKeeper session expiry is how the HMaster declares a server deadDedicated ZooKeeper StatefulSet on its own nodes or with reserved CPU

Notice what is missing: RegionServers need no persistent volumes. Their block cache is a performance asset, not durable data, which makes HBase a better Kubernetes fit than databases that keep their primary copy on the pod's disk.

The architecture on Kubernetes

A production layout has five workloads: a three or five pod ZooKeeper StatefulSet; an HA NameNode pair with three JournalNodes; a DataNode StatefulSet with one local volume per pod; two HMaster pods, one active and one standby; and a RegionServer StatefulSet sized to the data. Thrift or REST gateways are ordinary Deployments because they hold no state.

One Kubernetes node: a RegionServer next to the DataNode that holds its blocksZooKeeperStatefulSet, 3 or 5 podsHMaster x2active + standbyNameNode HA+ 3 JournalNodesClientsmust reach every RS podKubernetes node (pod affinity keeps RS and DN together)RegionServer podrs-3.rs.hbase.svc, heap + off-heap cacheDataNode podlocal PV, HFile + WAL blocksShared domain socket dirshort-circuit local readsremote read pathsessionassignblock mapget/putIdentity = stable DNS name from a headless Service. State = HDFS. Drain regions before the pod stops.Lose the node and HMaster replays the WAL from HDFS; lose the name and every client caches a dead address.
The RegionServer pod and the DataNode pod share a node and a domain socket directory, so reads of local blocks bypass TCP. Every other dependency reaches the pair over the cluster network.

The pairing is the key decision. Two containers in one pod guarantee co-location but couple lifecycles: restarting HBase restarts HDFS on that node. Two StatefulSets with a required pod affinity rule keep them independent, but after a reschedule a RegionServer may land beside a DataNode holding none of its files. Most teams choose the second and rebuild locality by compaction, as described in the region locality article.

Identity, DNS and client reachability

A RegionServer registers with the HMaster under a server name made of its hostname, its RPC port and a start code (a timestamp from process start). The HMaster writes region locations into hbase:meta using that hostname, and every client looks regions up there and then connects directly to the RegionServer that holds them. So the hostname a RegionServer announces must be resolvable, and reachable, from every client and from every other server.

A StatefulSet named rs with a headless Service also named rs gives pods the names rs-0.rs.hbase.svc.cluster.local, rs-1... and so on, and those names survive restarts even though the IP behind them changes. Make sure the process announces that name rather than a pod IP. How depends on version and image, so check your version's reference guide, then verify in the HMaster UI that servers show DNS names and keep them, with a new start code, across a restart.

apiVersion: v1
kind: Service
metadata: {name: rs, namespace: hbase}
spec:
  clusterIP: None              # headless: one DNS record per pod
  publishNotReadyAddresses: true
  selector: {app: hbase-rs}
  ports:
  - {name: rpc,  port: 16020}
  - {name: info, port: 16030}

publishNotReadyAddresses lets a starting RegionServer resolve its own name before it is ready.

The harder half is clients outside the cluster. Because clients talk to RegionServers directly, a single load balancer in front of HBase does not work for the native protocol. Options, in order of simplicity: run the clients inside the cluster; route the pod network to the client network so pod DNS names resolve there; or put a Thrift or REST gateway Deployment behind a load balancer and accept the extra hop and the gateway's feature limits. Decide this before go-live.

Storage and data locality

HDFS DataNodes want local disks: they replicate blocks themselves, so network block storage underneath them pays for redundancy twice and adds latency to every WAL sync. Use a local PersistentVolume provisioner, one volume per DataNode pod, and size the DataNode StatefulSet so that losing a node loses one replica. Spread DataNodes, JournalNodes and ZooKeeper across zones with topology spread constraints, and map nodes to zones in HDFS rack awareness.

Short-circuit reads are the main reason to co-locate. With dfs.client.read.shortcircuit=true and a dfs.domain.socket.path set identically in both containers, the RegionServer reads local block files through a file descriptor passed over a Unix socket instead of streaming them over TCP from the DataNode. On Kubernetes the socket lives in a directory both pods mount, typically a hostPath directory on the node, which is why the pairing must be on the same node.

<!-- hdfs-site.xml, identical in the RegionServer and DataNode containers -->
<property><name>dfs.client.read.shortcircuit</name><value>true</value></property>
<property><name>dfs.domain.socket.path</name><value>/var/run/hdfs-sockets/dn</value></property>

Check that it works, not just that it is configured: if the socket is unusable the client logs a warning and falls back to remote reads. Locality near 0 after a mass reschedule is repaired by a major compaction of the affected regions, at the cost of a burst of I/O.

Container memory: heap, off-heap and the OOM killer

On a VM, a RegionServer's heap is a tuning choice. In a container it is also a kill switch: if resident memory exceeds the container limit, the kernel OOM-kills the process with no GC log and no clean shutdown, and HBase treats it as a crash. The container limit must cover everything the JVM uses, not just the heap.

Memory consumerControlled byTypical share (illustrative)
Java heap (memstores, on-heap block cache, RPC)-Xmx or -XX:MaxRAMPercentage16-31 GiB
Off-heap BucketCachehbase.bucketcache.size plus -XX:MaxDirectMemorySize0-64 GiB
Metaspace, code cache, thread stacksJVM defaults, thread count1-2 GiB
Native buffers (HDFS client, compression codecs)workload1-2 GiB

Set heap and direct memory explicitly rather than as a percentage when off-heap cache is in use, then set the container memory request equal to the limit and the limit to the sum plus a margin of about ten percent. Equal request and limit put the pod in the Guaranteed QoS class, so it is the last to be evicted under node pressure. Pin CPU requests too: GC threads sized from the visible core count on a throttled container produce long pauses, which on HBase become ZooKeeper session expirations. Heap size and collector choice themselves are covered in HBase GC tuning.

Probes that do not fight ZooKeeper

A liveness probe that fails during a long GC pause or a slow WAL replay kills the container, which triggers a WAL split, which slows other RegionServers, which fail their probes. Keep the ZooKeeper session as the authority on whether a RegionServer is dead.

  • Startup probe: a TCP check on the RPC port with a generous failure budget (minutes, not seconds), so region opening and replay after restart are not cut short.
  • Readiness probe: TCP on the RPC port. It governs Service endpoints, which HBase clients do not use for routing, so its main job is to gate the StatefulSet rollout to one pod at a time.
  • Liveness probe: either omit it, or make it fail only after a window longer than zookeeper.session.timeout. A process that is truly hung loses its ZooKeeper session first and the HMaster reassigns its regions; Kubernetes restarting it afterwards is harmless.

Give ZooKeeper its own CPU reservation, away from RegionServers under GC pressure: an ensemble that pauses expires sessions for healthy servers and triggers a cluster-wide reassignment storm.

Drain before stop: rolling restarts

Kubernetes stops a pod by sending SIGTERM, waiting terminationGracePeriodSeconds, then sending SIGKILL. If a RegionServer simply exits, every region it hosted is offline until the HMaster notices, splits its WAL and reopens the regions elsewhere. The WAL split is correct but slow. Unload the server first with RegionMover, which moves regions off and confirms each is deployed elsewhere, so the stop becomes a no-op.

apiVersion: apps/v1
kind: StatefulSet
metadata: {name: rs, namespace: hbase}
spec:
  serviceName: rs
  replicas: 10
  podManagementPolicy: OrderedReady
  updateStrategy: {type: RollingUpdate}
  template:
    spec:
      terminationGracePeriodSeconds: 1800       # must exceed worst-case unload time
      containers:
      - name: regionserver
        image: your-registry/hbase:2.x          # pin a tested image
        resources:
          requests: {cpu: "8", memory: 56Gi}
          limits:   {cpu: "8", memory: 56Gi}
        lifecycle:
          preStop:
            exec:
              command:
              - /bin/bash
              - -c
              - hbase org.apache.hadoop.hbase.util.RegionMover -r "$(hostname -f)" -o unload -m 8

The preStop hook runs before SIGTERM is delivered, and its time counts against the grace period, so size the grace period from measurement: unload one loaded RegionServer in staging and multiply by two. A PodDisruptionBudget with maxUnavailable: 1 stops node drains and autoscaler scale-downs evicting two RegionServers at once. Turn the balancer off for the duration of a roll with balance_switch false in the HBase shell, otherwise it moves regions back onto the server you are about to stop, then turn it back on and let it even things out once.

After the restart the new pod is empty; let the balancer refill it, or run RegionMover with -o load to put regions back on their original host for locality.

Worked example: rolling ten RegionServers

Take a ten RegionServer cluster with 300 regions per server and assume, illustratively, that one region move takes two seconds end to end (close, flush the memstore, open elsewhere). With eight mover threads, unloading one server takes roughly 300 / 8 * 2 = 75 seconds; measure yours, since memstore size dominates it. The restart itself, including JVM start and registration, takes perhaps a minute. A full roll is therefore about ten times 2.5 to 3 minutes, so roughly half an hour, during which every region is moved twice and each client sees a handful of short retries but no region is offline for more than a move.

Compare the naive path: each SIGKILLed RegionServer triggers a WAL split and replay, leaving 300 regions unavailable meanwhile. Total wall time may be similar; the difference is that the drained roll keeps data available throughout.

Two traps: a grace period shorter than the unload gets the pod killed mid-move, giving you the crash path anyway; and on a hot cluster, moving 300 regions onto nine servers can push memstores into blocking flushes, so roll off-peak.

Failure modes

SymptomLikely causeFix
Clients time out against an IP that no longer existsRegionServer announced its pod IP, not its stable DNS nameAnnounce the FQDN; verify in the HMaster UI
Regions offline for minutes during every deployNo preStop unload, or grace period shorter than the unloadRegionMover in preStop; grace period from measured unload time
Pods restart with no Java stack traceOOMKilled: container limit below heap plus direct memoryExplicit heap and direct sizes; request equals limit
Healthy servers declared dead in burstsZooKeeper or RegionServer CPU throttling; GC pauses near session timeoutReserve CPU, isolate ZooKeeper, fix GC before raising timeouts
Cascading restarts after one slow nodeLiveness probe killing pods during replayStartup probe; liveness longer than the session timeout, or none
Read latency doubles after a node drainLocality lost on rescheduleMajor compact affected regions off-peak

Hand-rolled, Kustomize or an operator

Hand-written StatefulSets give full control, but every upgrade and scale event becomes a runbook. The Apache HBase project's hbase-kustomize repository provides Kustomize bases you can overlay, which keeps you close to upstream without adopting a controller. Operators, such as Stackable's HBase operator, encode drain-on-stop and configuration generation in a controller, which is the right call once you run several clusters; read their graceful-shutdown documentation to confirm it drains the way described above before trusting it in production. The general trade-offs of encoding operations in a controller are covered in the Kubernetes operator pattern.

What to do next

  1. Write down how every client reaches RegionServers today; if any run outside the Kubernetes cluster, decide between pod network routing and a Thrift or REST gateway before anything else.
  2. Deploy the headless Service and StatefulSet, then confirm in the HMaster UI that servers appear under stable DNS names and keep them across a pod restart.
  3. Co-locate DataNodes and RegionServers, enable short-circuit reads, and check the RegionServer log and locality metrics to prove local reads are in use.
  4. Set heap and direct memory explicitly, make the memory request equal the limit, and add a ten percent margin for native memory.
  5. Replace any aggressive liveness probe with a startup probe plus readiness; keep ZooKeeper as the failure detector.
  6. Measure a RegionMover unload on a loaded server in staging, set the grace period to twice that, add the preStop hook and a PodDisruptionBudget.
  7. Rehearse a full rolling restart and a hard node kill, and record region-offline time for each.
Key takeaway: HBase runs well on Kubernetes once you restore what it assumes from bare metal. Give each RegionServer a stable DNS name through a StatefulSet and headless Service, and settle how outside clients reach those pods. Keep durable state in HDFS on local volumes and co-locate RegionServers with DataNodes for short-circuit reads. Size the container for heap plus off-heap cache plus native memory. Let ZooKeeper, not a liveness probe, decide when a server is dead. Drain every RegionServer with RegionMover in a preStop hook, with a grace period measured rather than guessed.