An HBase filter is written on the client and runs on the server. Your application builds a Filter object, attaches it to a Get or Scan, and the client library serializes it into the request. Each RegionServer that holds part of the key range rebuilds the filter from bytes and applies it while it reads cells. Only cells the filter accepts travel back over the network. That split explains most surprises with filters: a class that works in a unit test but fails on the cluster, a filter string that parses in the shell but not through a gateway, and a scan that returns ten rows after reading ten million.
The server side, meaning the scanner hook sequence, the return codes and the cost of each built-in filter, is covered in HBase filters: what they save and what they do not. This page covers the client's half of the contract: building filters in Java, writing them as filter-language strings, how they are serialized and loaded, how to test them without a cluster, and when you should filter in your own code instead. The API details come from the HBase 2.6 source.
Where a filter lives
Client-side in definition, server-side in execution
A filter has three lives. On the client it is an object you construct. In the request it is a class name plus a byte array: the client calls toByteArray() on the filter and wraps the result in a protobuf message. On the RegionServer, ProtobufUtil loads the class by name and calls its static parseFrom(byte[]) method by reflection to get a live instance, and the region scanner then calls that instance for every row and cell it visits.
So "client-side filter" describes where the filter is defined, not where it executes. Three consequences follow. First, the class must exist on every RegionServer, not just in your application jar. Second, the filter's state is created fresh for each region the scan touches, so a filter cannot count across regions: a page-size filter that stops after 100 rows stops after 100 rows per region. Third, a filter reduces what is returned, not necessarily what is read. Unless the filter can tell the scanner to seek past data, the server still decodes every block in the scan range. Only the row key range on the scan narrows which regions and blocks are opened at all.
Building filters in Java
In HBase 2.x, comparisons use the CompareOperator enum; the older CompareFilter.CompareOp is deprecated. Put the cheapest, most selective test first in a FilterList with MUST_PASS_ALL, because evaluation stops at the first filter that rejects a cell. Set a key range on the scan before you reach for a filter; it is the only thing that skips whole regions.
byte[] cf = Bytes.toBytes("d");
byte[] tenant = Bytes.toBytes("t042#");
// Key range first: only regions holding t042# are opened.
Scan scan = new Scan()
.setStartStopRowForPrefixScan(tenant) // setRowPrefixFilter is deprecated since 2.5.0
.addColumn(cf, Bytes.toBytes("level"))
.addColumn(cf, Bytes.toBytes("msg"))
.setCaching(500) // rows per RPC
.setLimit(200); // whole-scan cap, enforced across regions
SingleColumnValueFilter isError = new SingleColumnValueFilter(
cf, Bytes.toBytes("level"), CompareOperator.EQUAL, Bytes.toBytes("ERROR"));
isError.setFilterIfMissing(true); // rows without the column are dropped, not kept
ValueFilter mentionsTimeout = new ValueFilter(
CompareOperator.EQUAL, new SubstringComparator("timeout"));
scan.setFilter(new FilterList(FilterList.Operator.MUST_PASS_ALL, isError, mentionsTimeout));
try (ResultScanner rs = table.getScanner(scan)) {
for (Result r : rs) {
handle(r);
}
}Two details in that snippet are worth keeping. setLimit is a client-coordinated cap on the whole scan, which is what people usually want from PageFilter but do not get. And setFilterIfMissing(true) matters because, by default, a SingleColumnValueFilter keeps rows that lack the column entirely. Note that this example has a quiet bug which the testing section below catches: the ValueFilter applies to every returned column, including level.
Filters as strings: the filter language
Callers that are not Java, such as the HBase shell, the Thrift gateway and REST scanners, describe filters as strings. ParseFilter turns a string into the same Filter objects. A filter call looks like a function call with quoted byte-string arguments, and comparisons take an operator plus a typed comparator:
PrefixFilter('t042#') AND SingleColumnValueFilter('d', 'level', =, 'binary:ERROR', true, true)
# comparator prefixes: binary: binaryprefix: regexstring: substring:
# operators: < <= = != > >=
# compound: AND, OR, and the unary SKIP and WHILEPrecedence in the parser is: SKIP and WHILE bind tightest, then AND, then OR, with parentheses overriding all three. Write the parentheses anyway; a reader should not need the parser's table. A single quote inside an argument is escaped by doubling it, so the value it's is written 'it''s'. The regexstring and substring comparators may only be used with = and !=; CompareFilter throws IllegalArgumentException for anything else. Parse errors are also IllegalArgumentException, with messages such as "Mismatched parenthesis" or "Filter Name X not supported".
In the shell, the same string goes in the FILTER option of scan, and show_filters lists the names the parser knows. In Java you can parse a string yourself with new ParseFilter().parseFilterString(s), which is useful when a filter arrives as configuration. Custom filters must be registered under a name with ParseFilter.registerFilter(name, className) before a string can use them. That registry is a static map in one JVM: registering in your application does nothing for a Thrift or REST gateway, which parses the string in its own process. The HBase shell guide covers scanning safely from the shell, and the Thrift API page covers how gateways pass scans through.
Serialization and shipping custom filters
Every filter, built-in or custom, must round-trip through bytes. Built-in filters implement toByteArray() with protobuf messages and provide a static parseFrom(byte[]). A custom filter must do both, and the static method must exist with exactly that name, because the server finds it by reflection; if it is missing, deserialization fails on the server and the client sees a remote exception, not a compile error.
public final class MinLengthFilter extends FilterBase {
private final int minLen;
public MinLengthFilter(int minLen) { this.minLen = minLen; }
@Override
public ReturnCode filterCell(Cell c) {
return c.getValueLength() >= minLen ? ReturnCode.INCLUDE : ReturnCode.SKIP;
}
@Override
public byte[] toByteArray() {
return Bytes.toBytes(minLen); // a real filter should use a versioned protobuf
}
public static MinLengthFilter parseFrom(byte[] bytes) throws DeserializationException {
if (bytes == null || bytes.length != Bytes.SIZEOF_INT) {
throw new DeserializationException("bad MinLengthFilter bytes");
}
return new MinLengthFilter(Bytes.toInt(bytes));
}
}Shipping the class has two options. Put the jar on every RegionServer's classpath and restart them, or place it in the directory named by hbase.dynamic.jars.dir, from which the servers' DynamicClassLoader can load filter classes when hbase.use.dynamic.jars is enabled. The dynamic path avoids a restart for the first version, but replacing a class that is already loaded is not something to rely on; give each incompatible change a new class name. Because the client and the servers deserialize the same bytes, version the encoding: a protobuf message with optional fields lets an old server read a new client's filter and vice versa.
Testing a filter without a cluster
Because a filter is a plain object with methods, most of its logic can be tested in a JVM with no cluster. Build cells with KeyValue, call the hooks in the order the scanner would, and assert the results. Always include a serialization round trip, since that is the path the server takes.
@Test
void errorRowsMatchAndRoundTrip() throws Exception {
Filter f = buildFilter(); // the FilterList from the Scan example
byte[] row = Bytes.toBytes("t042#20261003T1200#9");
Cell level = new KeyValue(row, CF, Bytes.toBytes("level"), Bytes.toBytes("ERROR"));
Cell msg = new KeyValue(row, CF, Bytes.toBytes("msg"), Bytes.toBytes("db timeout"));
f.reset();
assertFalse(f.filterRowKey(level)); // false means "do not exclude this row"
assertEquals(ReturnCode.INCLUDE, f.filterCell(level)); // fails: ValueFilter rejects "ERROR"
assertEquals(ReturnCode.INCLUDE, f.filterCell(msg));
Filter copy = FilterList.parseFrom(f.toByteArray()); // the server's path
assertArrayEquals(f.toByteArray(), copy.toByteArray());
copy.reset();
assertEquals(ReturnCode.INCLUDE, copy.filterCell(msg)); // the rebuilt filter behaves the same
}The first filterCell assertion fails, which is the point: the ValueFilter tests every cell, so the level cell is dropped from the result even though the row matches. The fix is to scope the substring test to one column, for example with a second SingleColumnValueFilter on msg using a SubstringComparator. A local test cannot prove performance, seek hints or region-boundary behaviour; run one integration test against a mini-cluster or a staging table for those.
Push down or post-filter?
Pushing a filter to the server saves network bytes and client CPU. It does not save disk reads unless the filter also lets the scanner seek. Filtering in your own code after the scan costs network transfer but gains everything Java can do: joins against a cache, calls to other services, and logic you can deploy without touching the cluster. The decision is mostly arithmetic.
| Situation | Push down | Post-filter in application |
|---|---|---|
| Predicate on the row key prefix | Use a key range, not a filter | Never |
| Rejects most of the data in range | Yes: large network saving | Only if the logic cannot be a filter |
| Rejects little of the data | Small gain, extra server CPU | Fine, simpler to change |
| Needs data from another table or service | Not possible | Yes |
| Logic changes weekly | Custom filter deploys are slow | Yes |
Batch scans with setBatch | Only filters without filterRow | Yes for row-level logic |
The last row is a hard rule. Scan.setBatch splits wide rows into partial results, and a filter that decides on the whole row cannot work on partial rows. If hasFilterRow() returns true, as it does for SingleColumnValueFilter, setting both throws IncompatibleFilterException on the client, whichever you set first.
Worked example: recent error events for one tenant
A support tool needs the 200 most recent error events mentioning "timeout" for one tenant. Row keys are tenant#reversed-timestamp#id, so newest rows sort first within a tenant. Suppose tenant t042 holds 4 million rows of about 1 KB in the queried columns, and 0.5 percent are errors, of which a fifth mention a timeout: 4,000 matching rows.
Without a key range, the scan visits every region of the table. With the prefix range, it opens only t042's regions. With no filter, returning 200 matches means pulling rows until 200 match; at a 0.1 percent hit rate that is about 200,000 rows, roughly 200 MB over the network to discard almost all of it. With the filter pushed down, the servers still read those 200,000 rows from block cache or disk, but send back about 200 KB. The disk work is the same; the network and client work drop by three orders of magnitude.
The remaining cost is the server read. If this query matters, change the data model: write an index row such as t042#ERR#reversed-ts#id when an error is stored, and the query becomes a plain prefix scan of 200 rows with no filter at all. Filters are a good tool for reducing transfer; key design is the tool for reducing reads. The scans guide covers caching, limits and RPC counts, and read-latency tuning covers the server side.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Remote exception deserializing the filter | Class or static parseFrom missing on a RegionServer | Deploy the jar everywhere or to the dynamic jars directory |
| Filter string fails only via gateway | Custom name registered in the app JVM, not the gateway | Register in the gateway process |
IncompatibleFilterException | Row-level filter with setBatch | Drop batch or post-filter |
| Rows with missing column returned | setFilterIfMissing left false | Set it true |
| Page returns pageSize per region | PageFilter state is per region | Use setLimit |
| Scan times out with few results | High rejection, no seek; lease or RPC timeout | Key range, index rows, smaller caching |
IllegalArgumentException on substring | Comparator used with < or > | Use only = or != |
Trade-offs
Built-in filters are safe and need no deployment, but compose into lists that are hard to read and easy to get subtly wrong. Custom filters can express exactly the logic you need and can provide seek hints, but every change is a server deployment and a compatibility risk across mixed client and server versions. Post-filtering is the most flexible and the most expensive in transfer. A reasonable default is a key range always, built-in filters for column and value predicates, post-filtering for anything involving other data, and a custom filter only when measurements show transfer is the bottleneck and the logic is stable. If server-side logic grows beyond a predicate, such as aggregation, a coprocessor is the next step rather than a cleverer filter; the client libraries guide covers how clients reach both.
What to do next
- List every scan in your application that sets a filter, and note whether it also sets a key range. Add a range to any that do not.
- Replace any
PageFilterused for whole-scan paging withsetLimitplus a start row. - Set
setFilterIfMissing(true)on everySingleColumnValueFilterunless you have decided that missing means match. - Write a unit test per filter that calls the hooks in scanner order and round-trips the filter through toByteArray and parseFrom.
- For any custom filter, confirm the jar is on every RegionServer or in the dynamic jars directory, and register its name in each gateway that parses filter strings.
- Measure rows read versus rows returned for your top three scans using scan metrics. Where the ratio is above 100, consider an index row instead of a stronger filter.