Batching is how HBase clients reach high write rates: instead of one RPC per Put, the client groups hundreds of mutations into one multi request per RegionServer. It is also where many production incidents begin, because a batch looks like one operation in application code and behaves like dozens of independent ones on the cluster. Half of it can succeed, it can be held back by client-side limits nobody configured, and it can retry for twenty minutes before telling you anything.

This article is about operating batches, not about the API surface, which is covered in the HBase Java API and the HBase client libraries guide. Here we follow a batch from the application to the RegionServers, list the limits that shape it and their defaults (checked against hbase-default.xml and the client source for the 2.x line on 2026-10-04), and turn that into sizing rules, retry handling, metrics and runbooks.

What a batch is, and is not

HBase guarantees atomicity per row, and nothing larger. A batch is a convenience for transport: the client sends many row operations in few RPCs, and each row operation succeeds or fails on its own. There is no ordering guarantee inside a batch either; the javadoc for Table.batch says the execution order of actions is not defined, so a Put and a Delete for the same row in one batch can apply in either order. If order matters, split them into separate calls or use a single-row RowMutations.

There are four ways to batch in the Java client:

  • Table.batch(List<? extends Row>, Object[] results) and Table.put(List<Put>): synchronous, blocking until every action has succeeded or exhausted retries.
  • BufferedMutator: accumulates mutations in a client buffer and flushes them in the background when the buffer fills, on flush(), or on close().
  • AsyncTable.batchAll / putAll: non-blocking, returning futures.
  • AsyncBufferedMutator: the asynchronous counterpart of the buffered mutator.

The first two share one engine, AsyncProcess with a request controller; the async client uses a different caller and different limits. Mixing up the two families is a common source of wrong tuning.

The path of a batch

When a synchronous batch is submitted, the client first resolves each row to a region, using its cached copy of hbase:meta and looking up misses. Actions are grouped by the RegionServer hosting their region, and within that by region, producing one multi request per server. Each request goes through the request controller, which decides whether it may be sent now or must wait. Responses come back per action: some succeed, some fail with retryable exceptions (region moved, server busy), some fail permanently. Retryable failures are regrouped, because the region may now live elsewhere, and resent after a pause.

One batch, many RPCs: how the synchronous client fans out a list of mutationsApplicationbatch(list) or mutate()BufferedMutator2 MB write bufferAsyncProcesslocate regions (meta cache)group actions by RegionServerflushRequest controllerfour checkers:tasks per region / server / total,request heap, rows, submitted heapRegionServer Amulti RPC: regions 1, 4RegionServer Bmulti RPC: regions 2, 7RegionServer Cmulti RPC: region 3Per-action resultssuccess or exceptionpartialretry failed onlyRetriesExhausted...after retries / timeout
Figure 2. A batch is not a transaction. The client groups actions by RegionServer, throttles them through four checkers, sends one multi RPC per server, and retries only the actions that failed until retries or the operation timeout run out.

Three facts from this path drive everything that follows. A batch is only as fast as its slowest RegionServer, because a synchronous call returns when all servers have answered. A hot region throttles the whole batch, because of per-region task limits. And retries are per action, which is good for throughput and bad for idempotence.

The client-side limits

The synchronous client's request controller applies four checks before an action joins an outgoing request. The configuration keys and their 2.x defaults are:

KeyDefaultWhat it limits
hbase.client.max.total.tasks100Concurrent mutation tasks one table instance sends to the whole cluster
hbase.client.max.perserver.tasks2Concurrent tasks to one RegionServer
hbase.client.max.perregion.tasks1Concurrent tasks to one region
hbase.client.max.perrequest.heapsize4194304 (4 MB)Heap size of one multi request
hbase.client.max.perrequest.rows2048Rows in one multi request
hbase.client.max.submit.heapsize4194304 (4 MB)Heap submitted in one round
hbase.client.perserver.requests.threshold2147483647Pending requests per server across all client threads in the process
hbase.client.write.buffer2097152 (2 MB)Default BufferedMutator buffer

Read the defaults as a deliberately conservative fairness policy. One in-flight task per region means a batch whose rows all land in one region is serialised: each 4 MB or 2,048-row request must finish before the next is sent. That is the client-side face of a hotspot, and raising the limit only moves the queue onto the RegionServer. The real fix is the row key. The per-server threshold is effectively unlimited by default; setting it to a finite number makes the client fail fast with ServerTooBusyException instead of piling more requests onto a struggling server, which is often what you want in a multi-threaded service.

BufferedMutator in production

BufferedMutator is the right tool for streams of independent writes. Unlike Table, it is safe to share across threads, so one instance per table per process is the norm. Errors are not thrown from mutate() for the actions that failed in the background; they are delivered to an exception listener, or thrown from a later mutate, flush or close. Configure it explicitly:

BufferedMutator.ExceptionListener listener = (e, mutator) -> {
    for (int i = 0; i < e.getNumExceptions(); i++) {
        Row row = e.getRow(i);
        deadLetters.record(row, e.getCause(i), e.getHostnamePort(i));
    }
};

BufferedMutatorParams params = new BufferedMutatorParams(TableName.valueOf("events"))
    .writeBufferSize(8L * 1024 * 1024)               // flush every ~8 MB
    .setWriteBufferPeriodicFlushTimeoutMs(1000)      // or after 1 s, whichever first
    .listener(listener);

try (BufferedMutator mutator = connection.getBufferedMutator(params)) {
    for (Event ev : stream) {
        Put put = new Put(rowKey(ev));
        put.addColumn(CF, Q_PAYLOAD, ev.payload());
        mutator.mutate(put);
    }
    mutator.flush();   // make durability explicit before acknowledging upstream
}

Two operational rules. First, the buffer is volatile: if the process dies, buffered mutations are gone. If you consume from Kafka or a queue, commit offsets only after flush() returns. Second, a periodic flush (available in 2.1 and later) bounds latency for slow streams; without it a quiet stream can leave writes sitting in the buffer indefinitely.

Partial failure with Table.batch

With Table.batch, the results array holds one entry per action: a Result for success, or the exception that defeated it. If any action fails after all retries the call also throws RetriesExhaustedWithDetailsException, which lists the failed rows, causes and servers. The useful pattern is to resubmit only what failed, with a bounded number of rounds and a dead-letter destination:

List<Row> pending = new ArrayList<>(actions);
for (int round = 0; round < 3 && !pending.isEmpty(); round++) {
    Object[] results = new Object[pending.size()];
    try {
        table.batch(pending, results);
        pending.clear();
    } catch (RetriesExhaustedWithDetailsException e) {
        List<Row> failed = new ArrayList<>();
        for (int i = 0; i < results.length; i++) {
            if (results[i] == null || results[i] instanceof Throwable) {
                failed.add(pending.get(i));
            }
        }
        log.warn("round {}: {} of {} actions failed, e.g. {}", round,
                 failed.size(), pending.size(), e.getCause(0).toString());
        pending = failed;
        Thread.sleep(1000L << round);
    }
}
pending.forEach(deadLetters::record);

Before resubmitting, classify the cause. NoSuchColumnFamilyException or a too-large cell will fail every time; resending them only adds load. Region movement and server-busy exceptions are worth a later round.

Retries, timeouts, overload and idempotence

Inside one call the client retries on its own: hbase.client.retries.number defaults to 15, with a base pause of hbase.client.pause = 100 ms scaled by a backoff table, so the later pauses are measured in seconds. Each RPC is bounded by hbase.rpc.timeout (60,000 ms) and the whole operation by hbase.client.operation.timeout (1,200,000 ms, twenty minutes). For an online service, twenty minutes is not a timeout, it is an outage with extra steps; set the operation timeout to what your caller can actually wait.

Overload has its own signals. When a RegionServer's call queue is full it rejects with CallQueueTooBigException or drops calls; hbase.client.pause.server.overloaded sets a separate, longer pause for those, so clients back off harder from a drowning server than from a moved region. When a region's MemStore reaches hbase.hregion.memstore.block.multiplier (4) times the flush size, writes to it block and clients see RegionTooBusyException until a flush catches up; the write-path mechanics behind that are explained in HBase write performance.

Retries make Puts safe when the client sets the cell timestamp, because rewriting the same cell with the same timestamp is idempotent; with server-assigned timestamps a retried Put can add an extra version. Increments and Appends are not idempotent at all. The client attaches nonces to those operations by default (hbase.client.nonces.enabled) so the server can recognise a retry of the same call within its nonce window. That protects against the client's internal retries. It does not protect against your own resubmission loop above, which creates new operations with new nonces. Counters that must be exact should be derived from idempotent Puts keyed by event id, or rebuilt periodically.

Worked example: sizing an ingest service

Worked example: an ingest service writes 50,000 events per second, about 1 KB each, to a table pre-split into 200 regions over 20 RegionServers, from 8 application instances. That is 50 MB/s, or about 6 MB/s per instance.

With the default 2 MB buffer each instance flushes three times a second, and each flush spreads about 2,000 rows over 20 servers, roughly 100 rows (100 KB) per multi request. Requests that small spend most of their time on RPC overhead and WAL syncs. Raising the buffer to 8 MB gives about 400 rows per server per flush, still well inside the 2,048-row and 4 MB per-request ceilings, and still below the server-side hbase.rpc.rows.warning.threshold of 5,000 rows, above which RegionServers log a warning about the batch size. Concurrency: per-server tasks of 2 times 20 servers is 40 in-flight requests per instance, under the total of 100.

Now check the server side. Eight instances times two tasks is 16 concurrent multi calls per server, against hbase.regionserver.handler.count = 30 handlers. Headroom remains for reads, which matters because batches and Gets share handlers unless call queues are split. If you later double the instances, revisit both sides together. For one-off loads of billions of rows, skip the RPC path altogether and use bulk loading.

The async client needs your own back-pressure

The asynchronous client returns immediately with futures and does not use the synchronous request controller, so the task limits above do not apply to it. That makes it easy to issue a million outstanding operations and fill the RegionServers' call queues. Bound concurrency yourself:

AsyncTable<?> table = asyncConn.getTable(TableName.valueOf("events"));
Semaphore inFlight = new Semaphore(64);           // batches in flight per process

for (List<Put> chunk : Lists.partition(puts, 500)) {
    inFlight.acquire();
    CompletableFuture.allOf(table.put(chunk).toArray(new CompletableFuture[0]))
        .whenComplete((v, err) -> {
            inFlight.release();
            if (err != null) deadLetters.recordChunk(chunk, err);
        });
}

putAll and batchAll give a single future for the whole list, which is simpler but fails the whole future if any action fails; per-action futures from put(List) or batch let you keep the partial-failure behaviour described earlier.

Observability

Measure batches at both ends. On the client, enable connection metrics with hbase.client.metrics.enable and record your own histograms of batch size, flush latency and failed-action counts; a rising failure count with steady volume is the earliest sign of a sick server. On the server, watch call queue length and rejected calls, time spent with updates blocked by MemStore pressure, and per-region write request rates to spot hotspots. Large-batch warnings in RegionServer logs mean a client is sending more rows per request than intended. Quotas, covered in quota throttling, turn a noisy batch tenant into a predictable one.

Failure modes

  • Silent data loss on crash. Offsets committed before a BufferedMutator flush; buffered writes vanish with the process.
  • Hung callers. A sick server and the 20-minute default operation timeout keep threads blocked long after the caller gave up.
  • Serialised hot region. All rows go to one region; per-region task limit 1 makes the batch crawl. Fix the key, not the limit.
  • Double counting. Application-level resubmission of Increments after a partial failure.
  • Unbounded async fan-out. Call queues overflow, clients receive overload exceptions, and retries amplify the load.
  • Ordering assumptions. Put and Delete for the same row in one batch apply in an undefined order.

Trade-offs

Larger batches raise throughput and lower per-row overhead, at the cost of latency, memory on both sides and bigger blast radius per failure. A BufferedMutator is the most efficient path but gives up synchronous error reporting and durability until flush. Synchronous batch gives exact per-action results at the cost of blocking a thread. The async client scales furthest per thread but removes the built-in throttles, so you own back-pressure. Low retry counts and short timeouts fail fast and keep services responsive, while generous ones ride out region moves; choose per workload, not cluster-wide.

What to do next

  1. List every batch call site and record which family it uses: synchronous batch, BufferedMutator, or async.
  2. Set hbase.client.operation.timeout and hbase.rpc.timeout to values your callers can actually wait for.
  3. For each BufferedMutator, register an exception listener, set a periodic flush, and flush before acknowledging upstream.
  4. Replace whole-batch retries with retries of failed actions only, with a dead-letter path.
  5. Audit Increments and Appends in retry loops; move exact counts to idempotent Puts.
  6. Size buffers so each multi request carries hundreds of rows, below 2,048 rows and 4 MB.
  7. Bound async in-flight work with a semaphore, and alert on call queue length and blocked updates.
Key takeaway: An HBase batch is many independent row operations sharing transport. The synchronous client groups them by RegionServer and throttles them per region, per server and per request; the async client does not throttle at all. Handle partial failure per action, keep Increments out of retry loops, set timeouts you can live with, and size buffers so each request carries hundreds of rows.