HBase filters are server-side predicates. You attach one to a Scan or Get, the client serializes it, and every RegionServer that holds part of the scanned key range runs it inside its scanner, returning only the cells the filter accepts. Used well, filters turn a scan that would ship gigabytes to the client into one that ships kilobytes. Used carelessly, they create a scan that reads a whole region from disk, returns almost nothing, and times out while doing it.
The difference is whether a filter lets the scanner seek past data or merely skip it after reading. This article explains where filters run, the hook sequence and return codes that decide how much is read, which built-in filters seek, how to compose them, the traps in value filters and pagination, when a custom filter is justified, and how to operate filtered scans without timeouts. Examples use the HBase 2.x Java client API.
What filters save, and what they do not
A scan without a filter reads every cell in its range, sends all of it over the network and makes the client discard what it does not need. A filter moves the discard to the server, which saves network transfer, RPC count and client memory and CPU. That is the whole benefit when the filter only inspects cells after they are read: the RegionServer still decompresses every block, merges every store file and evaluates the predicate on every cell. If the filter keeps one row in a million, the scan still costs roughly as much disk and CPU as the unfiltered scan, and now it produces long silent stretches with nothing to return.
Real reductions in disk work come from two sources only. The first is narrowing the key range with start and stop rows so fewer blocks are touched at all. The second is a filter that returns a seek instruction, which lets the scanner jump forward using the HFile block index and skip whole blocks. Everything else is transfer optimization. Row-key design, covered for bloom filters in HBase bloom filters, determines how often you can use the first two.
Where filters run: the scanner path
A client Scan is split by region. For each region, the RegionServer opens a region scanner, which merges per-store scanners over the MemStore and every HFile into one stream of cells in key order. The query matcher applies version limits, time-to-live and delete markers first; then the filter sees each surviving cell. Accepted cells accumulate into a result batch that is returned when it reaches the configured row caching or byte limit. The filter object is instantiated per region scanner, so its state never spans regions. The RegionServer describes the surrounding read path.
The hook sequence and return codes
The Filter class exposes hooks called in a fixed order. filterAllRemaining() is checked first: returning true ends the scan for this region, which is how WhileMatchFilter and PageFilter stop early. filterRowKey(Cell) sees the first cell of each row and returns a boolean: true means exclude the row. Then filterCell(Cell) is called for each cell and returns a ReturnCode; transformCell can rewrite an accepted cell, which is how KeyOnlyFilter strips values. If hasFilterRow() is true, filterRowCells(List) and filterRow() run after the whole row is gathered, which forces the server to buffer the row; reset() clears per-row state.
| ReturnCode | Meaning | Cost implication |
|---|---|---|
| INCLUDE | return this cell, continue | normal |
| INCLUDE_AND_NEXT_COL | return it, move to the next column | skips older versions |
| INCLUDE_AND_SEEK_NEXT_ROW | return it, done with this row | seeks to next row |
| SKIP | drop this cell, look at the next one | everything still read |
| NEXT_COL | drop the rest of this column | may seek |
| NEXT_ROW | drop the rest of this row | seeks to next row |
| SEEK_NEXT_USING_HINT | jump to the cell from getNextCellHint | can skip many blocks |
The practical rule: a filter that knows it is done with a row or a range should say so with NEXT_ROW or a seek hint instead of returning SKIP for every remaining cell.
The built-in filters, ranked by how much they read
| Filter | Examines | Can seek past data? | Typical use |
|---|---|---|---|
| PrefixFilter | row key prefix | stops the scan once past the prefix; set start row yourself | key-prefix scans |
| MultiRowRangeFilter | row key ranges | yes, seeks between ranges | many disjoint key ranges in one scan |
| FuzzyRowFilter | row key with fixed and wildcard bytes | yes, seek hints to next candidate | fixed-width composite keys with an unknown leading part |
| RowFilter | row key with a comparator | no, evaluates every row | regex or substring on keys, last resort |
| SingleColumnValueFilter | one column's value | no | select rows by an attribute |
| QualifierFilter, FamilyFilter | column names | no (use Scan.addColumn when names are known) | dynamic column selection |
| ColumnPrefixFilter, ColumnRangeFilter | qualifier prefix or range | yes, within a row | wide rows with sorted qualifiers |
| ValueFilter | every cell value | no | rarely; almost always a full read |
| KeyOnlyFilter, FirstKeyOnlyFilter | structure only | FirstKeyOnly moves to next row | counting rows, existence checks |
| PageFilter | rows returned | stops a region scan | approximate page size; see pagination |
| TimestampsFilter | cell timestamps | within a column | exact versions |
Row-key filters come first because rows are the unit of storage order. Whenever the predicate can be expressed as key ranges, prefer start and stop rows on the Scan, then MultiRowRangeFilter for several ranges, then FuzzyRowFilter for fixed-width keys where some bytes are unknown. RowFilter with a regular expression comparator evaluates every row key in the range and should be a last resort.
import org.apache.hadoop.hbase.CompareOperator;
import org.apache.hadoop.hbase.client.Scan;
import org.apache.hadoop.hbase.filter.*;
import org.apache.hadoop.hbase.util.Bytes;
import org.apache.hadoop.hbase.util.Pair;
import java.util.List;
// Row key layout: [4-byte metric id][4-byte device id][8-byte reversed timestamp]
// Question: all readings of metric 17, any device, where status == "ERROR".
byte[] metric = Bytes.toBytes(17);
Scan scan = new Scan()
.withStartRow(metric) // narrow the key range first
.withStopRow(Bytes.toBytes(18))
.addColumn(Bytes.toBytes("d"), Bytes.toBytes("status"))
.addColumn(Bytes.toBytes("d"), Bytes.toBytes("value"));
SingleColumnValueFilter isError = new SingleColumnValueFilter(
Bytes.toBytes("d"), Bytes.toBytes("status"),
CompareOperator.EQUAL, Bytes.toBytes("ERROR"));
isError.setFilterIfMissing(true); // rows WITHOUT the column are dropped, not kept
isError.setLatestVersionOnly(true); // compare only the newest version
scan.setFilter(isError);
scan.setCaching(500); // rows per RPC
scan.setLimit(1000); // total rows for the whole scan, across regions
SingleColumnValueFilter and its two traps
SingleColumnValueFilter is the filter most people reach for and the one most often misused. First, by default a row that does not contain the tested column passes the filter; you must call setFilterIfMissing(true) to get SQL-like behaviour. Second, the tested column must be read by the scan. If you restrict the scan with addColumn to other columns only, the filter never sees the tested column, treats every row as missing, and returns either everything or nothing depending on the first setting.
When the tested column lives in a small family and the payload in a large one, enable Scan.setLoadColumnFamiliesOnDemand(true). The server then reads the essential family first, evaluates the filter, and loads the other families only for matching rows, which can cut disk reads substantially for selective predicates. Remember that value filters never seek: selecting 0.01 percent of rows by value still reads the full range. If a value predicate is on the hot path, it belongs in the row key or in a secondary index table.
Composing filters with FilterList
FilterList combines filters with MUST_PASS_ALL (AND) or MUST_PASS_ONE (OR), and lists nest. Under AND, a cell is dropped as soon as one child rejects it, so put the cheapest and most selective filters first and row-key filters before value filters. Under OR, the list must merge children's return codes and seek hints conservatively, since it can only seek as far as the child that wants to read the least; an OR over a seeking filter and a non-seeking one usually reads everything.
Two wrapper filters change control flow. WhileMatchFilter ends the region scan the first time its child rejects something, which is useful for scanning forward until a condition breaks. SkipFilter drops an entire row if its child rejects any cell of it, which requires the row to be evaluated as a whole. Both are easy to misread; write a unit test that asserts exactly which rows come back from a small in-memory table before relying on them.
FilterList cheapFirst = new FilterList(FilterList.Operator.MUST_PASS_ALL,
new MultiRowRangeFilter(List.of(
new MultiRowRangeFilter.RowRange(Bytes.toBytes(17), true, Bytes.toBytes(18), false),
new MultiRowRangeFilter.RowRange(Bytes.toBytes(42), true, Bytes.toBytes(43), false))),
isError); // value test only on surviving rows
scan.setFilter(cheapFirst);
Pagination: why PageFilter lies
PageFilter looks like LIMIT, but it runs independently in each region scanner: a page size of 100 over a scan that touches five regions can return up to 500 rows. The HBase javadoc says so explicitly. For a total row limit, use Scan.setLimit, which the client enforces across regions. For stable pagination in an API, remember the last row key returned and start the next page just after it with withStartRow(lastKey, false); this is cheap because it seeks directly, unlike offset-based paging, which must read and discard all earlier rows.
Worked example: from full read to seek
Consider the metrics table from the code above, with 500 metric ids spread evenly and several regions per metric range. A first version of the query scanned the whole table with only the SingleColumnValueFilter attached. It returned the right rows, but every RegionServer read every block of every region, and the client saw long pauses kept alive only by heartbeats. The filter saved network, nothing else.
The second version added start and stop rows for metric 17. Because the metric id leads the key, the scan now touches roughly one five-hundredth of the table and only the regions holding that range; how key ranges map to regions is covered in HBase regions and splits. The third version added the status column in its own small family with on-demand loading, so the large value family is read only for matching rows. Same filter, same answer, but the disk work fell by orders of magnitude because the range and the family layout changed, not the predicate. If the question later becomes all errors for one device across metrics, the key order no longer helps, and the right answer is a secondary index table keyed by device, not a cleverer filter.
Custom filters: when and how
Write a custom filter only when the predicate cannot be expressed with built-ins and it can seek or significantly reduce transfer, for example decoding a binary key component and jumping directly to the next valid value. Extend FilterBase, override the hooks you need, and implement toByteArray() plus a static parseFrom(byte[]), because filters travel to the server as serialized bytes. The class must be on every RegionServer's classpath; deploying it is a rolling change, and a client that sends a filter the servers cannot load gets an error on every scan.
Treat a custom filter like server code, because it is. It runs on RegionServer handler threads and shares their memory and CPU. An exception, an unbounded buffer in filterRowCells, or a slow regular expression affects every tenant of the cluster. Version its serialized form, keep it stateless between rows, and consider whether a coprocessor, described in HBase coprocessors, is the better tool when the logic needs to aggregate rather than select.
Operating filtered scans
Selective filters over large ranges produce long gaps between accepted cells. Older clients hit scanner lease expiry or RPC timeouts during those gaps; modern HBase servers send heartbeat responses with no rows when a time or cell-count budget is reached, so the client keeps the scanner alive. Even so, keep hbase.client.scanner.timeout.period (60 seconds by default) consistent between client and server, tune setCaching and setMaxResultSize to bound per-RPC memory, and use setBatch or setAllowPartialResults for very wide rows.
Watch the RegionServer read-request and filtered-read-request metrics: a large ratio of cells read to cells returned identifies scans that need a better key design rather than a better filter. In the HBase shell, the filter language lets you test a filter before coding it, for example scan 'metrics', {STARTROW => ..., FILTER => "PrefixFilter('abc') AND KeyOnlyFilter()"}.
What to do next
- List your top scans by volume and record, for each, rows read versus rows returned.
- Replace RowFilter and PrefixFilter-only scans with explicit start and stop rows or MultiRowRangeFilter.
- Audit every SingleColumnValueFilter for setFilterIfMissing, for the tested column being in the scan, and for on-demand family loading.
- Replace PageFilter used as a limit with Scan.setLimit and key-based pagination.
- Order FilterList children cheapest and most selective first, and unit-test the rows each composition returns.
- Move hot value predicates into the row key or a secondary index table.
- Confirm scanner timeouts, caching and max result size are set deliberately on clients and servers.