Running one HBase cluster per team is simple and expensive. Every cluster needs its own masters, ZooKeeper quorum, monitoring, upgrades and spare capacity, and most of them sit idle most of the day. Consolidating teams onto a shared cluster saves that cost but brings a new problem. One team's full-table scan or bulk backfill can double another team's p99 latency, and nothing in a default installation stops it.

HBase has had tools for this for years: namespaces, access control, namespace limits, throttle quotas, space quotas and RegionServer groups. Each one solves a different part of the problem, and none of them is enough alone. This article explains what each layer controls and what it doesn't, then combines them into tiers you can offer to tenants, with an onboarding script, client code, a diagnosis procedure and the failure modes that catch teams out. The internals of two of the layers have their own pages: quota throttling and RegionServer groups.

Advertisement

What tenants actually share

Start from the resources, not the features. A RegionServer hosts regions from many tables and shares a set of per-process resources between them: RPC handler threads and call queues, the block cache, the global memstore budget, the write-ahead log, flush and compaction threads, and the JVM heap and garbage collector. Across the cluster, every tenant shares the HMaster, the hbase:meta table, ZooKeeper and HDFS, with its NameNode, DataNodes and disks.

Interference follows from that list. A tenant scanning cold data evicts another tenant's hot blocks from the block cache. A tenant writing heavily fills the global memstore, which forces flushes and eventually blocks writes for everyone on that server. A tenant with thousands of tiny regions adds load to the master and meta for everyone. And a tenant whose data grows without limit fills HDFS, which is a cluster-wide outage. Every isolation layer below addresses some of these resources and not others, and that is what decides which tenants can safely share.

Layers of isolation in a shared HBase cluster, from logical to physicalNamespace per tenantnames, ACLs (grant on @ns), maxtables / maxregionsThrottle quotas (RPC rate)per user / namespace / table; READ, WRITE, REQUEST; req, bytes or CU per timeSpace quotas (bytes on the filesystem)per namespace / table; NO_INSERTS, NO_WRITES, NO_WRITES_COMPACTIONS, DISABLERegionServer groups (dedicated servers)tenant namespace pinned to its own RegionServers via hbase.rsgroup.nameStill shared inside a groupblock cache, memstore, WAL, handlersStill shared across the clusterHMaster, hbase:meta, ZooKeeper, HDFSEach layer answers a different question: who may touch it, how fast, how much, and on which hardware.The bottom row is what no quota protects you from, so it decides which tenants can safely share.
Isolation layers from logical to physical. The two red boxes are the resources no HBase quota partitions.

Namespaces and access control

A namespace is a named group of tables, written as namespace:table. On its own it isolates nothing at runtime, but it is the unit every other feature attaches to: ACLs, limits, quotas and RegionServer group assignment. So the first rule of HBase tenancy is one namespace per tenant, and no tenant tables in default.

With the AccessController enabled (normally alongside Kerberos authentication), grants on a namespace are written with an @ prefix: grant 'svc_payments', 'RWXC', '@payments' gives a service principal read, write, execute and create rights in its own namespace and nothing elsewhere. Keep the admin permission (A) with the platform team, grant to groups rather than individuals, and give analytics readers read-only grants. Cell-level visibility labels exist for finer control, but they add per-cell cost and are rarely needed for tenant separation.

Namespaces also carry two structural limits that protect the shared master: hbase.namespace.quota.maxtables and hbase.namespace.quota.maxregions, set with create_namespace or alter_namespace. A region limit is the cheapest protection against a tenant who pre-splits into ten thousand regions and turns every balancer run and master restart into a long operation.

Advertisement

Throttle quotas: rate limits per tenant

Throttle quotas limit request rates on the server side. They are disabled by default; enable them with hbase.quota.enabled=true on all nodes. A throttle can target a user, a namespace, a table, a user within a namespace or table, or all RegionServers, and it can limit reads, writes or both (THROTTLE_TYPE => READ, WRITE or the default REQUEST). Limits are expressed as requests per time unit (4000req/sec), bytes per time unit (200M/sec) or capacity units per time unit (100CU/sec), with sec, min, hour or day as the time unit.

Two details change how you size them. First, the default SCOPE => MACHINE applies the limit on each RegionServer separately, so a 4,000 req/sec namespace quota on a 10-server cluster allows up to 40,000 req/sec in total if the load is spread evenly, and only 4,000 if one hot region concentrates it. SCOPE => CLUSTER divides a cluster-wide limit by the number of RegionServers, and the shell help explicitly advises against cluster scope when you use RegionServer groups, because the divisor counts every server in the cluster, not just the tenant's group. Second, quota changes take effect after the refresh period, hbase.quota.refresh.period, which defaults to 300,000 ms (five minutes), so a throttle is not an emergency brake.

When a tenant exceeds its limit, the server rejects the call with an RpcThrottlingException carrying a wait interval. The client library retries internally, so the application sees it only when retries run out, wrapped in a retries-exhausted exception. Treat that as back-pressure:

// The client retries throttled calls itself; what reaches you after the retries run out
// is a RetriesExhausted* exception with the RpcThrottlingException somewhere in its causes.
try (Table t = conn.getTable(TableName.valueOf("payments:ledger"))) {
    t.put(puts);
} catch (IOException e) {
    RpcThrottlingException te = findCause(e, RpcThrottlingException.class);
    if (te == null) throw e;
    metrics.counter("hbase.throttled").increment();
    backoff.sleep(Math.max(te.getWaitInterval(), 50));  // then retry, or shed load upstream
}

static <T extends Throwable> T findCause(Throwable t, Class<T> type) {
    for (Throwable c = t; c != null; c = c.getCause()) {
        if (type.isInstance(c)) return type.cast(c);
    }
    return null;
}

Throttles control request rate. They do not control what one request costs, so a single throttled scan over a cold range can still churn the block cache. The token-bucket mechanics, including how read size is estimated before a request runs, are covered in HBase quota throttling.

Space quotas: stopping growth before HDFS fills

Space quotas limit how many bytes a namespace or table may occupy on the filesystem, and they say what happens when it goes over. The policies, from least to most strict, are NO_INSERTS (no Put, Increment or Append), NO_WRITES (deletes are also refused), NO_WRITES_COMPACTIONS (compactions stop too) and DISABLE (the table is disabled). A table-level quota takes priority over its namespace's quota. Usage includes HFiles that are only kept alive by snapshots, attributed by rules the reference guide describes, so forgotten snapshots count against the tenant that took them. list_snapshot_sizes shows where that space is going.

For most tenants NO_INSERTS is the right policy. It stops growth but still allows deletes, so the tenant can clean up and get back under the limit without the platform team stepping in. NO_WRITES_COMPACTIONS and DISABLE are for cases where going over the limit is itself the emergency. Set space quotas so that the sum across tenants, times the HDFS replication factor, stays below the capacity you actually have, and alert at 80 percent of each quota so tenants hear about it before their writes start failing.

RegionServer groups: when sharing servers is the problem

Quotas limit how much a tenant can ask for, but everything above still happens on shared RegionServers, with a shared block cache, memstore, WAL and garbage collector. For a latency-critical tenant that is often not enough. RegionServer groups (rsgroups) partition the servers: each group has its own set of RegionServers, and a namespace or table assigned to a group has its regions placed only on those servers. On current releases the feature is turned on with hbase.balancer.rsgroup.enabled; older 2.x releases configured the RSGroupAdminEndpoint coprocessor and RSGroupBasedLoadBalancer instead, so check your version's reference guide.

Assign a namespace to a group with move_namespaces_rsgroup, which also records the group in the namespace property hbase.rsgroup.name. New tables in the namespace then land in the group automatically. The balancer works within each group, as described in the StochasticLoadBalancer article, so a group needs enough servers to survive losing one: three is a practical minimum for a tier-one group.

Groups are not a separate cluster. The master, meta, ZooKeeper and HDFS are still shared, and unless you also arrange HDFS placement, a group's DataNode I/O happens on the same disks as everyone else's. Groups also cost capacity: servers in a small group cannot absorb other tenants' load, so you give up some pooling efficiency in exchange for predictable latency.

Designing tenant tiers: a worked example

Suppose one cluster of 15 RegionServers serves three tenants. Payments needs single-digit-millisecond p99 reads and steady writes. Analytics runs large scans every hour. An ingest team backfills historical events in bursts. Offering each tenant a tier, rather than tuning every tenant separately, keeps the design understandable:

TierTenantPlacementThrottles (per server)Space policy
1: dedicatedpaymentsrsgroup tier1, 3 serversWRITE 4000req/sec, READ 200M/sec, sized to protect the tenant from its own bugs20T, NO_INSERTS
2: shared, protectedanalyticsdefault group, 12 serversREAD 100M/sec so scans cannot saturate handlers50T, NO_INSERTS
3: shared, best-effortingestdefault group, 12 serversWRITE 20M/sec during business hours, raised at night30T, NO_WRITES

The onboarding for the tier-one tenant looks like this. Each step maps to one layer in the diagram:

# hbase-site.xml on all masters and RegionServers (restart required):
#   hbase.quota.enabled = true
#   hbase.balancer.rsgroup.enabled = true     (current releases; older 2.x used the
#                                              RSGroupAdminEndpoint coprocessor plus
#                                              RSGroupBasedLoadBalancer)

# 1. Logical home and structural limits
create_namespace 'payments', {'hbase.namespace.quota.maxtables' => '10',
                              'hbase.namespace.quota.maxregions' => '400'}

# 2. Access: the tenant's service principal owns its namespace, nothing else
grant 'svc_payments', 'RWXC', '@payments'
grant '@payments_readers', 'R', '@payments'

# 3. Rate: per RegionServer by default (SCOPE => MACHINE)
set_quota TYPE => THROTTLE, NAMESPACE => 'payments', THROTTLE_TYPE => WRITE, LIMIT => '4000req/sec'
set_quota TYPE => THROTTLE, NAMESPACE => 'payments', THROTTLE_TYPE => READ,  LIMIT => '200M/sec'

# 4. Size: stop growth before HDFS fills
set_quota TYPE => SPACE, NAMESPACE => 'payments', LIMIT => '20T', POLICY => NO_INSERTS

# 5. Hardware: a dedicated group for the latency-critical tenant
add_rsgroup 'tier1'
move_servers_rsgroup 'tier1', ['rs11.example.com:16020', 'rs12.example.com:16020', 'rs13.example.com:16020']
move_namespaces_rsgroup 'tier1', ['payments']      # also sets hbase.rsgroup.name on the namespace

# 6. Verify
list_quotas NAMESPACE => 'payments'
get_namespace_rsgroup 'payments'

The numbers in the table are illustrative. The method is to measure each tenant's normal peak with its own metrics, set throttles somewhat above it so normal traffic never hits them, and keep the headroom on each server larger than the sum of the throttles that can land there. For the shared group, also split the RPC call queues so reads and writes cannot starve each other: hbase.ipc.server.callqueue.read.ratio divides queues between reads and writes, and hbase.ipc.server.callqueue.scan.ratio divides the read queues between gets and scans. Replicating a tenant's namespace to a second cluster, described in HBase replication, is the next step when a tenant needs failure isolation as well as performance isolation.

Diagnosing a noisy neighbour

  1. Confirm the symptom is server-side: compare client p99 with the RegionServer's own processing and queue times. If queue time dominates, handlers are saturated; if processing time dominates, look at cache, disk or GC.
  2. Find the hot server and regions with the per-region request metrics, or interactively with hbtop, which ranks regions, tables and namespaces by request rate.
  3. Map the hot regions to namespaces. If they belong to a different tenant from the one complaining, you have a neighbour problem; if they belong to the same tenant, it is a row-key or hotspot problem instead.
  4. Check which shared resource is contended: block cache hit ratio drops point to a scanning tenant, memstore pressure and blocked updates point to a heavy writer, long GC pauses point to large responses.
  5. Apply the smallest fix that works: a throttle on the offending namespace (remember the five-minute refresh), then moving the victim to its own group if it keeps happening. The layer-by-layer method in HBase read performance applies here too.

Failure modes and trade-offs

  • Quotas enabled on some nodes only. hbase.quota.enabled must be set everywhere, or enforcement is patchy and hard to reason about.
  • Cluster-scope throttles with rsgroups. The limit is divided by all RegionServers, so a tenant confined to three servers gets a small fraction of what you intended.
  • Hot regions defeat machine-scope throttles. A tenant whose traffic concentrates on one region gets only one server's share. That is usually correct, but it surprises tenants, so explain it at onboarding.
  • Retry storms. Application-level retry loops on top of the client's own retries turn a throttle into extra load. Honour the wait interval and cap retries.
  • Space quota lockout. NO_WRITES also blocks deletes, so a tenant over quota cannot free space without an admin. Prefer NO_INSERTS unless you mean it.
  • Group too small. A two-server group loses half its capacity on one failure. Size groups for N+1.
  • The shared floor. No quota protects against a meta or HDFS outage. Tenants that cannot tolerate one need their own cluster, and the trade-off is always consolidation savings against blast radius.

What to do next

  1. Inventory tenants, give each its own namespace, and move any tables out of default.
  2. Turn on authorization and grant each tenant's service principal rights on its own namespace only.
  3. Set maxtables and maxregions on every namespace to protect the master.
  4. Enable quotas on all nodes, measure each tenant's peak, and set machine-scope throttles above it.
  5. Add space quotas with NO_INSERTS and alert at 80 percent.
  6. Move latency-critical tenants into an rsgroup of at least three servers, and make clients honour the wait interval of a wrapped RpcThrottlingException.
Key takeaway: HBase multi-tenancy is a stack of layers, and each covers a different resource. Namespaces and ACLs decide who may touch what. Namespace limits protect the master. Throttle quotas cap request rates per RegionServer, and space quotas cap bytes on HDFS. RegionServer groups give a tenant its own servers and so its own cache, memstore and WAL. None of them isolate the master, meta, ZooKeeper or HDFS. Offer tenants a small set of tiers built from these layers, size throttles from measured peaks with the machine scope in mind, and make clients treat throttling as back-pressure.