Hybrid search runs a lexical retriever and a dense retriever in parallel, merges their candidates, and usually passes the merged list through a reranker that reads the query and each candidate together. The reason it works is well understood: BM25 is precise on exact tokens such as product codes, error strings, names and rare terms, while embeddings match meaning across paraphrase and vocabulary mismatch, and the two fail on different queries. The fusion arithmetic is covered in a companion piece on hybrid retrieval math.
This article is about building and running the system. Most hybrid deployments that disappoint do so for engineering reasons rather than mathematical ones: an analyzer that splits the identifiers BM25 was supposed to catch, a vector filter that quietly returns too few results, two indexes that disagree about which documents exist, candidate counts chosen without looking at recall, or a reranker whose latency eats the whole budget. The sections below walk through the reference architecture, each stage's decisions and failure modes, and how to evaluate and tune the whole pipeline.
Reference architecture
At ingest, documents are chunked once, and each chunk gets a stable identifier, the text, metadata such as source, timestamps and access-control fields, and an embedding. The chunk is written to both an inverted index and a vector index under the same identifier. At query time, a query analysis step produces both an analyzed term query and a query embedding, plus filters derived from the user's permissions and any facets. Both retrievers run concurrently with the same filters. Their candidate lists are fused by chunk identifier, the top of the fused list is reranked by a cross-encoder, and the final few results go to the user interface or into an LLM's context.
Many engines now support both index types in one system, which simplifies consistency, while others pair a search engine with a separate vector database. Either works; the important invariants are the shared chunk identifier and identical filter semantics on both paths, discussed below.
The lexical side is an analyzer problem
BM25's value in a hybrid system comes from exact matching of rare, specific tokens, and the default analyzers in most search engines are designed for prose, not for those tokens. A standard analyzer may split SKU-4417X into sku and 4417x, split an error code such as ERR_CONN_RESET on underscores, or stem a product name into a common word. The result is a lexical retriever that is good at exactly what the dense retriever already does, and bad at its actual job.
Index important text with two analyzers: a prose analyzer with stemming for natural-language matches, and an identifier-preserving analyzer that keeps codes, versions and paths intact. Query both fields and let the identifier field carry a boost. Add synonyms carefully and only for domain terms that users really interchange, because broad synonym lists increase recall at the expense of the precision that justified BM25 in the first place. Field weights, such as title over body, matter more than tuning the BM25 parameters k1 and b, which rarely need to leave their defaults unless documents vary enormously in length.
PUT /kb
{
"settings": {
"index": {"knn": true},
"analysis": {
"filter": {
"english_stemmer": {"type": "stemmer", "language": "english"},
"trim_punct": {"type": "pattern_replace", "pattern": "^[^\\w]+|[^\\w]+$", "replacement": ""}
},
"analyzer": {
"prose": {"tokenizer": "standard", "filter": ["lowercase", "english_stemmer"]},
"ident": {"tokenizer": "whitespace", "filter": ["lowercase", "trim_punct"]}
}
}
},
"mappings": {"properties": {
"chunk_id": {"type": "keyword"},
"title": {"type": "text", "analyzer": "prose", "fields": {"id": {"type": "text", "analyzer": "ident"}}},
"body": {"type": "text", "analyzer": "prose", "fields": {"id": {"type": "text", "analyzer": "ident"}}},
"acl": {"type": "keyword"},
"embedding": {"type": "knn_vector", "dimension": 1024}
}}
}
The dense side: chunks, models and filtered ANN
Dense retrieval quality depends first on chunking and the embedding model, and second on the approximate nearest neighbor index. Chunks should be sized to the embedding model's effective context and aligned to document structure, and they must be the same chunks the lexical index holds, or fusion will compare different units. Embed queries with any query-specific instruction the model expects; several embedding families use different prefixes for queries and passages, and omitting them costs measurable recall.
Filtering is the most common silent failure. With post-filtering, the index finds the nearest k vectors and then drops those the user may not see; for a restrictive filter, most or all of the k are removed and the retriever returns a handful of results or none, even though relevant permitted documents exist. Pre-filtering or filter-aware graph traversal avoids that but can be slower and, for very selective filters, degrades to near brute force. The practical pattern is to use the engine's filtered search mode, raise the candidate count for selective filters, and monitor the ratio of returned results to requested results per query. For HNSW, the search-time breadth parameter, often called ef_search, trades latency for recall and should be set from a recall measurement against exact search on a sample, not left at a default.
Fusion: ranks or scores
BM25 scores are unbounded and depend on the query; cosine similarities sit in a narrow band that differs by model. Adding them directly makes one retriever dominate by accident. Reciprocal rank fusion sidesteps this by using only positions: each document scores the sum over retrievers of one over k plus its rank, with k around 60 by convention. It needs no tuning and is robust, which is why it is the usual default. Score blending normalizes each list, for example by min-max within the query's candidates, then combines with a weight; it can beat rank fusion when tuned on labeled data because it preserves how confident each retriever was, but it is sensitive to normalization and needs retuning when either model changes.
def rrf(lists: dict[str, list[str]], k: int = 60, weights: dict[str, float] | None = None):
weights = weights or {}
scores: dict[str, float] = {}
for name, ids in lists.items():
w = weights.get(name, 1.0)
for rank, cid in enumerate(ids, start=1):
scores[cid] = scores.get(cid, 0.0) + w / (k + rank)
return sorted(scores, key=scores.get, reverse=True)
def blend(bm25: dict[str, float], dense: dict[str, float], alpha: float = 0.5):
def norm(d):
if not d: return {}
lo, hi = min(d.values()), max(d.values())
return {c: (s - lo) / (hi - lo) if hi > lo else 1.0 for c, s in d.items()}
b, v = norm(bm25), norm(dense)
ids = b.keys() | v.keys()
return sorted(ids, key=lambda c: alpha * v.get(c, 0.0) + (1 - alpha) * b.get(c, 0.0), reverse=True)When a reranker follows, the choice of fusion method matters less than it seems, because fusion's job shrinks to deciding which candidates make it into the reranker's window. What matters most is that the union of candidates contains the relevant documents. That makes first-stage recall the key metric for fusion, and it is why weighted rank fusion with a boost for the lexical list on identifier-looking queries is a cheap, effective refinement.
Reranking and its latency budget
A cross-encoder reads the query and a candidate passage together and outputs a relevance score. Because it attends across both texts, it captures interactions that neither bi-encoder embeddings nor term statistics can, and it typically produces the largest single quality gain in the pipeline. The price is compute that scales with the number of candidates times their length: every candidate is a full forward pass. Reranking 50 passages of 256 tokens is 12,800 tokens of encoder work per query, which fits comfortably on a GPU in tens of milliseconds when batched but can take hundreds of milliseconds on CPU.
def rerank(query: str, candidates: list[Chunk], model, top_n=8, max_tokens=320, batch=32):
pairs = [(query, truncate_to_tokens(c.text, max_tokens)) for c in candidates]
scores = []
for i in range(0, len(pairs), batch):
scores.extend(model.score(pairs[i:i + batch])) # one batched forward pass
ranked = sorted(zip(candidates, scores), key=lambda x: x[1], reverse=True)
return [(c, s) for c, s in ranked[:top_n]]Set the budget explicitly. Decide the end-to-end latency target, subtract network, query embedding and first-stage retrieval, and size the reranker window to what remains at p95 under load, not on an idle machine. Truncate passages at a fixed token limit, which bounds cost, and prefer chunks short enough that truncation rarely cuts the relevant span. Cache reranker scores for popular query and passage pairs, and have a degradation mode: under overload, shrink the window or skip reranking and serve fused results, rather than timing out. The reranker's scores are also useful downstream; a low top score is a signal that retrieval found nothing good, which an LLM application can use to answer cautiously or ask a clarifying question.
Choosing candidate counts
Three numbers define the pipeline: how many candidates each retriever returns, how many fused candidates are reranked, and how many results are finally used. Choose them from measurements. Plot recall of the relevant documents in the union of both retrievers as the per-retriever depth grows; the curve usually flattens somewhere between 50 and 300. Then measure final quality as the rerank window grows, with its latency alongside. The reranker can only reorder what it is given, so first-stage recall at the rerank window is the ceiling on final quality, and the right window is where extra candidates stop adding relevant documents faster than they add latency.
Evaluating by query class
Build a labeled query set from real traffic, with relevant chunks judged for each query, and stratify it. Hybrid search helps different query classes for different reasons, and an aggregate hides the trade-offs. Typical classes are exact identifier lookups, short keyword queries, natural-language questions, long pasted text such as error logs, and queries whose answers need recent documents. Report recall at the rerank window for each retriever alone and for the union, and nDCG or MRR at the final cutoff after reranking, per class.
Run ablations whenever a component changes: lexical only, dense only, fused without rerank, and the full pipeline. The ablation table tells you where a regression came from and whether a component still earns its latency. A common finding is that a better embedding model shrinks the gain from BM25 on natural-language queries but not on identifiers, which argues for keeping the lexical path even when dense retrieval improves. Tune fusion weights and candidate counts on one split of the query set and report on another, since a few dozen queries are easy to overfit.
Keeping two indexes consistent
When the two indexes are separate systems, they can disagree. A document updated in the lexical index but not yet re-embedded returns old text from one path and new text from the other; a delete applied to one index leaves the other still serving it; an access-control change updated in one place leaks content through the other. Drive both from the same ordered change log keyed by chunk identifier, apply deletes and permission changes before content updates, and reconcile periodically by comparing identifier sets and versions between the two. At query time, dedupe by chunk identifier and prefer the newer version if both paths return it. Access filters must have identical semantics on both paths, ideally built from one shared filter representation, because the weaker path defines the security of the whole system.
Trade-offs and when not to bother
Hybrid search roughly doubles indexing work and operational surface, and reranking adds a model to serve. For corpora with little identifier-style content and users who ask natural-language questions, a strong embedding model with a reranker may match hybrid quality at lower cost; the ablation table will show it. Conversely, for code search, logs, part catalogs and legal references, the lexical path often carries most of the value and dense retrieval is the supplement. Hybrid is the safe default when query types are mixed or unknown, which describes most enterprise assistants.
Failure modes
- Prose analyzer on identifiers. Codes are split or stemmed and BM25 loses its reason to exist.
- Post-filtered ANN. Restrictive permissions leave the dense path returning few or no results.
- Mismatched chunks. The two indexes hold different units, so fusion compares unrelated passages.
- Raw score addition. Unnormalized scores let one retriever dominate by accident.
- Unbudgeted reranker. The rerank window is sized on an idle box and blows p95 under load.
- Divergent deletes and ACLs. One index still serves content the other has removed or restricted.
- Aggregate-only evaluation. Gains on questions hide regressions on identifier lookups.