HBase filters are server-side predicates. You attach one to a Scan or Get, the client serializes it, and every RegionServer that holds part of the scanned key range runs it inside its scanner, returning only the cells the filter accepts. Used well, filters turn a scan that would ship gigabytes to the client into one that ships kilobytes. Used carelessly, they create a scan that reads a whole region from disk, returns almost nothing, and times out while doing it.

The difference is whether a filter lets the scanner seek past data or merely skip it after reading. This article explains where filters run, the hook sequence and return codes that decide how much is read, which built-in filters seek, how to compose them, the traps in value filters and pagination, when a custom filter is justified, and how to operate filtered scans without timeouts. Examples use the HBase 2.x Java client API.

Advertisement

What filters save, and what they do not

A scan without a filter reads every cell in its range, sends all of it over the network and makes the client discard what it does not need. A filter moves the discard to the server, which saves network transfer, RPC count and client memory and CPU. That is the whole benefit when the filter only inspects cells after they are read: the RegionServer still decompresses every block, merges every store file and evaluates the predicate on every cell. If the filter keeps one row in a million, the scan still costs roughly as much disk and CPU as the unfiltered scan, and now it produces long silent stretches with nothing to return.

Real reductions in disk work come from two sources only. The first is narrowing the key range with start and stop rows so fewer blocks are touched at all. The second is a filter that returns a seek instruction, which lets the scanner jump forward using the HFile block index and skip whole blocks. Everything else is transfer optimization. Row-key design, covered for bloom filters in HBase bloom filters, determines how often you can use the first two.

Where filters run: the scanner path

A client Scan is split by region. For each region, the RegionServer opens a region scanner, which merges per-store scanners over the MemStore and every HFile into one stream of cells in key order. The query matcher applies version limits, time-to-live and delete markers first; then the filter sees each surviving cell. Accepted cells accumulate into a result batch that is returned when it reaches the configured row caching or byte limit. The filter object is instantiated per region scanner, so its state never spans regions. The RegionServer describes the surrounding read path.

Filters run inside the RegionServer scanner, between the merged cell stream and the RPC responseClient Scanrange + filter + limitsRegionServer RPCone scanner per regionserialized filterStoreScannersMemStore + HFiles mergedQuery matcherversions, TTL, deletescells in key orderFilter hooksfilterAllRemaining()filterRowKey(cell)filterCell(cell) returns a ReturnCodetransformCell(cell)filterRowCells / filterRowgetNextCellHint(cell)each cellseek / skipResult batchbounded by caching, sizeaccepted cellsRPC responseSeek hints skip disk blocksSKIP only skips network
The filtered read path. Cells arrive in key order from the merged store scanners; the query matcher applies versions, TTL and deletes; the filter's hooks decide per row and per cell. Only seek instructions reduce blocks read; skip codes only reduce what is returned.
Advertisement

The hook sequence and return codes

The Filter class exposes hooks called in a fixed order. filterAllRemaining() is checked first: returning true ends the scan for this region, which is how WhileMatchFilter and PageFilter stop early. filterRowKey(Cell) sees the first cell of each row and returns a boolean: true means exclude the row. Then filterCell(Cell) is called for each cell and returns a ReturnCode; transformCell can rewrite an accepted cell, which is how KeyOnlyFilter strips values. If hasFilterRow() is true, filterRowCells(List) and filterRow() run after the whole row is gathered, which forces the server to buffer the row; reset() clears per-row state.

ReturnCodeMeaningCost implication
INCLUDEreturn this cell, continuenormal
INCLUDE_AND_NEXT_COLreturn it, move to the next columnskips older versions
INCLUDE_AND_SEEK_NEXT_ROWreturn it, done with this rowseeks to next row
SKIPdrop this cell, look at the next oneeverything still read
NEXT_COLdrop the rest of this columnmay seek
NEXT_ROWdrop the rest of this rowseeks to next row
SEEK_NEXT_USING_HINTjump to the cell from getNextCellHintcan skip many blocks

The practical rule: a filter that knows it is done with a row or a range should say so with NEXT_ROW or a seek hint instead of returning SKIP for every remaining cell.

The built-in filters, ranked by how much they read

FilterExaminesCan seek past data?Typical use
PrefixFilterrow key prefixstops the scan once past the prefix; set start row yourselfkey-prefix scans
MultiRowRangeFilterrow key rangesyes, seeks between rangesmany disjoint key ranges in one scan
FuzzyRowFilterrow key with fixed and wildcard bytesyes, seek hints to next candidatefixed-width composite keys with an unknown leading part
RowFilterrow key with a comparatorno, evaluates every rowregex or substring on keys, last resort
SingleColumnValueFilterone column's valuenoselect rows by an attribute
QualifierFilter, FamilyFiltercolumn namesno (use Scan.addColumn when names are known)dynamic column selection
ColumnPrefixFilter, ColumnRangeFilterqualifier prefix or rangeyes, within a rowwide rows with sorted qualifiers
ValueFilterevery cell valuenorarely; almost always a full read
KeyOnlyFilter, FirstKeyOnlyFilterstructure onlyFirstKeyOnly moves to next rowcounting rows, existence checks
PageFilterrows returnedstops a region scanapproximate page size; see pagination
TimestampsFiltercell timestampswithin a columnexact versions

Row-key filters come first because rows are the unit of storage order. Whenever the predicate can be expressed as key ranges, prefer start and stop rows on the Scan, then MultiRowRangeFilter for several ranges, then FuzzyRowFilter for fixed-width keys where some bytes are unknown. RowFilter with a regular expression comparator evaluates every row key in the range and should be a last resort.

import org.apache.hadoop.hbase.CompareOperator;
import org.apache.hadoop.hbase.client.Scan;
import org.apache.hadoop.hbase.filter.*;
import org.apache.hadoop.hbase.util.Bytes;
import org.apache.hadoop.hbase.util.Pair;
import java.util.List;

// Row key layout: [4-byte metric id][4-byte device id][8-byte reversed timestamp]
// Question: all readings of metric 17, any device, where status == "ERROR".
byte[] metric = Bytes.toBytes(17);
Scan scan = new Scan()
    .withStartRow(metric)                                  // narrow the key range first
    .withStopRow(Bytes.toBytes(18))
    .addColumn(Bytes.toBytes("d"), Bytes.toBytes("status"))
    .addColumn(Bytes.toBytes("d"), Bytes.toBytes("value"));

SingleColumnValueFilter isError = new SingleColumnValueFilter(
    Bytes.toBytes("d"), Bytes.toBytes("status"),
    CompareOperator.EQUAL, Bytes.toBytes("ERROR"));
isError.setFilterIfMissing(true);    // rows WITHOUT the column are dropped, not kept
isError.setLatestVersionOnly(true);  // compare only the newest version

scan.setFilter(isError);
scan.setCaching(500);                // rows per RPC
scan.setLimit(1000);                 // total rows for the whole scan, across regions

SingleColumnValueFilter and its two traps

SingleColumnValueFilter is the filter most people reach for and the one most often misused. First, by default a row that does not contain the tested column passes the filter; you must call setFilterIfMissing(true) to get SQL-like behaviour. Second, the tested column must be read by the scan. If you restrict the scan with addColumn to other columns only, the filter never sees the tested column, treats every row as missing, and returns either everything or nothing depending on the first setting.

When the tested column lives in a small family and the payload in a large one, enable Scan.setLoadColumnFamiliesOnDemand(true). The server then reads the essential family first, evaluates the filter, and loads the other families only for matching rows, which can cut disk reads substantially for selective predicates. Remember that value filters never seek: selecting 0.01 percent of rows by value still reads the full range. If a value predicate is on the hot path, it belongs in the row key or in a secondary index table.

Composing filters with FilterList

FilterList combines filters with MUST_PASS_ALL (AND) or MUST_PASS_ONE (OR), and lists nest. Under AND, a cell is dropped as soon as one child rejects it, so put the cheapest and most selective filters first and row-key filters before value filters. Under OR, the list must merge children's return codes and seek hints conservatively, since it can only seek as far as the child that wants to read the least; an OR over a seeking filter and a non-seeking one usually reads everything.

Two wrapper filters change control flow. WhileMatchFilter ends the region scan the first time its child rejects something, which is useful for scanning forward until a condition breaks. SkipFilter drops an entire row if its child rejects any cell of it, which requires the row to be evaluated as a whole. Both are easy to misread; write a unit test that asserts exactly which rows come back from a small in-memory table before relying on them.

FilterList cheapFirst = new FilterList(FilterList.Operator.MUST_PASS_ALL,
    new MultiRowRangeFilter(List.of(
        new MultiRowRangeFilter.RowRange(Bytes.toBytes(17), true, Bytes.toBytes(18), false),
        new MultiRowRangeFilter.RowRange(Bytes.toBytes(42), true, Bytes.toBytes(43), false))),
    isError);                                              // value test only on surviving rows
scan.setFilter(cheapFirst);

Pagination: why PageFilter lies

PageFilter looks like LIMIT, but it runs independently in each region scanner: a page size of 100 over a scan that touches five regions can return up to 500 rows. The HBase javadoc says so explicitly. For a total row limit, use Scan.setLimit, which the client enforces across regions. For stable pagination in an API, remember the last row key returned and start the next page just after it with withStartRow(lastKey, false); this is cheap because it seeks directly, unlike offset-based paging, which must read and discard all earlier rows.

Worked example: from full read to seek

Consider the metrics table from the code above, with 500 metric ids spread evenly and several regions per metric range. A first version of the query scanned the whole table with only the SingleColumnValueFilter attached. It returned the right rows, but every RegionServer read every block of every region, and the client saw long pauses kept alive only by heartbeats. The filter saved network, nothing else.

The second version added start and stop rows for metric 17. Because the metric id leads the key, the scan now touches roughly one five-hundredth of the table and only the regions holding that range; how key ranges map to regions is covered in HBase regions and splits. The third version added the status column in its own small family with on-demand loading, so the large value family is read only for matching rows. Same filter, same answer, but the disk work fell by orders of magnitude because the range and the family layout changed, not the predicate. If the question later becomes all errors for one device across metrics, the key order no longer helps, and the right answer is a secondary index table keyed by device, not a cleverer filter.

Custom filters: when and how

Write a custom filter only when the predicate cannot be expressed with built-ins and it can seek or significantly reduce transfer, for example decoding a binary key component and jumping directly to the next valid value. Extend FilterBase, override the hooks you need, and implement toByteArray() plus a static parseFrom(byte[]), because filters travel to the server as serialized bytes. The class must be on every RegionServer's classpath; deploying it is a rolling change, and a client that sends a filter the servers cannot load gets an error on every scan.

Treat a custom filter like server code, because it is. It runs on RegionServer handler threads and shares their memory and CPU. An exception, an unbounded buffer in filterRowCells, or a slow regular expression affects every tenant of the cluster. Version its serialized form, keep it stateless between rows, and consider whether a coprocessor, described in HBase coprocessors, is the better tool when the logic needs to aggregate rather than select.

Operating filtered scans

Selective filters over large ranges produce long gaps between accepted cells. Older clients hit scanner lease expiry or RPC timeouts during those gaps; modern HBase servers send heartbeat responses with no rows when a time or cell-count budget is reached, so the client keeps the scanner alive. Even so, keep hbase.client.scanner.timeout.period (60 seconds by default) consistent between client and server, tune setCaching and setMaxResultSize to bound per-RPC memory, and use setBatch or setAllowPartialResults for very wide rows.

Watch the RegionServer read-request and filtered-read-request metrics: a large ratio of cells read to cells returned identifies scans that need a better key design rather than a better filter. In the HBase shell, the filter language lets you test a filter before coding it, for example scan 'metrics', {STARTROW => ..., FILTER => "PrefixFilter('abc') AND KeyOnlyFilter()"}.

What to do next

  1. List your top scans by volume and record, for each, rows read versus rows returned.
  2. Replace RowFilter and PrefixFilter-only scans with explicit start and stop rows or MultiRowRangeFilter.
  3. Audit every SingleColumnValueFilter for setFilterIfMissing, for the tested column being in the scan, and for on-demand family loading.
  4. Replace PageFilter used as a limit with Scan.setLimit and key-based pagination.
  5. Order FilterList children cheapest and most selective first, and unit-test the rows each composition returns.
  6. Move hot value predicates into the row key or a secondary index table.
  7. Confirm scanner timeouts, caching and max result size are set deliberately on clients and servers.
Key takeaway: HBase filters run inside each region scanner and decide, cell by cell, what goes back to the client. They always save network and client work, but they only save disk work when the scan range is narrowed or the filter can seek. Prefer key ranges and seeking filters, guard SingleColumnValueFilter's missing-column behaviour, use Scan.setLimit instead of PageFilter, compose cheap filters first, and treat custom filters as server code.