Most explanations of search architecture stop at the index: tokenize documents, build posting lists, score with BM25, shard the result. That machinery decides whether a system can answer a query in time. It does not decide whether the answer is any good. Relevance comes from a second architecture wrapped around the index: a cascade of retrieval and ranking stages, a log of what was shown and why, a source of labels, a training job and an evaluation loop that tells you whether a change helped. Teams that build only the first half end up hand-tuning boost weights forever.

This article covers that second half. The index internals, analysis chain and scatter-gather path are covered in search system architecture, and shard sizing and reindexing in designing search at scale. Here we follow a query through the cascade, look at what each stage is allowed to cost, then build the loop that learns the ranker: features logged at serve time, judgments, click debiasing, LambdaMART training, NDCG and interleaving. A worked example ties it together.

Advertisement

The cascade and its latency budget

A ranker that reads a hundred features per document cannot score ten million documents in a hundred milliseconds. A cascade solves this by trading breadth for cost at each step. Early stages look at many documents with cheap signals; later stages look at few documents with expensive ones. Every stage has two numbers you should be able to state: how many candidates it receives and keeps, and how many milliseconds it may spend.

A search request is a cascade: each stage sees fewer documents and spends more per documentQuerytext, user, contextUnderstandingspell, intent, entitiesLexical retrieverVector retrieverPopular / personalMerge and dedupeabout 1,000 candidatesL1 cheap rankerkeeps about 200L2 learned rankerkeeps about 50Rules and blendingfilters, diversityResults page10 shownImpression and feature logevery shown result with the exact features the ranker usedJudgmentsgraded human labelsClicks, debiasedpropensity weightedTrain and evaluateNDCG, then interleavingnew model
A query is parsed, several retrievers propose candidates, and progressively more expensive rankers cut the list down. Everything shown is logged with its features, and those logs plus judgments train the next ranker.
StageInputOutputTypical budgetSignals
Query understanding1 queryStructured query5-15 msDictionaries, small classifiers
Retrieval (L0)Whole indexAbout 1,00020-40 msBM25, ANN similarity, filters
First ranker (L1)About 1,000About 2005-10 ms10-30 cheap features, linear or small trees
Learned ranker (L2)About 200About 5020-40 ms100+ features, gradient-boosted trees or a cross-encoder
Rules and blendingAbout 50Page of 10under 5 msBusiness filters, diversity, ads

The numbers are illustrative, not a standard; measure your own. The design rule is that recall lost at an early stage can never be recovered later. If the right document is not among the thousand candidates, the best L2 model in the world will not show it. So you evaluate each stage on what it is responsible for: retrieval on recall at its cutoff, rankers on ordering quality at the top.

Query understanding produces a structured query

Raw text is a weak input. The first stage turns it into a structure that later stages can use: normalized text, corrected spelling, recognized entities and attributes, and a guess at intent. This is where a query such as womens waterproof hiking boots size 8 becomes a filter on size, a soft preference for a category, and a text query for the rest.

{
  "raw": "womens waterproof hiking boots size 8",
  "normalized": "women waterproof hiking boot size 8",
  "intent": {"label": "product_search", "confidence": 0.94},
  "filters": {"size": "8"},
  "boosts": {"category": {"hiking_boots": 2.0}, "gender": {"women": 1.5}},
  "text": "waterproof hiking boot",
  "rewrites": ["water resistant hiking boot"]
}

Two habits keep this stage honest. Treat entity matches as boosts unless you are very confident, because a wrong hard filter produces an empty page that no ranker can fix. And log the structured query next to the raw one, so that when a result page is bad you can tell whether understanding or ranking caused it.

Advertisement

Several retrievers, merged with quotas

No single retriever has full recall. Lexical retrieval is precise for exact terms, model numbers and rare words. Vector retrieval finds paraphrases and vague descriptions. A popularity or personalization retriever surfaces items the user is likely to want even when the text match is weak. Running them in parallel and merging the results gives the rankers a candidate set none of them could produce alone. How to fuse lexical and dense scores is covered in hybrid search with BM25, dense retrieval and reranking.

Merge by document id, keep the per-retriever scores and ranks as features, and give each retriever a quota so that one source cannot crowd out the others. The quotas themselves are tunable: measure recall at 1,000 against judged queries for each mix, and pick the cheapest mix that keeps recall close to the maximum.

Log features at serve time, not afterwards

A learned ranker is a function of features. The most common way to break it is to compute those features one way when serving and a different way when building training data. Popularity counts get recomputed with today's numbers, a text field gets re-analyzed with a newer tokenizer, and the model is trained on inputs it will never see in production. Offline metrics improve and the live site gets worse.

The fix is architectural: when the ranker scores a candidate, write the exact feature vector it used into an impression log, together with the query id, position and whether the result was shown. Training data is then a join of that log with labels, never a recomputation.

{"qid": "q-7f3a", "ts": "2026-10-01T09:14:03Z", "model": "l2-2026-09-28",
 "doc": "sku-41882", "pos": 3, "shown": true,
 "features": {"bm25_title": 11.2, "bm25_body": 7.9, "ann_cos": 0.81,
              "ctr_30d": 0.044, "price_z": -0.3, "in_stock": 1,
              "cat_match": 1, "freshness_days": 12}}

Log candidates that were scored but not shown as well, sampled if volume is a problem. They are the negatives the model most needs to learn from, and without them your training set only contains documents an earlier model already liked.

Where labels come from

There are two sources of relevance labels, and production systems use both. Human judgments are graded labels, usually on a scale from 0 (irrelevant) to 3 or 4 (perfect), assigned by trained raters following written guidelines. They are slow and expensive, but they are unbiased by your current ranking and they cover queries with little traffic. A few thousand judged queries with 20 to 50 judged documents each is a useful judgment list.

Clicks are cheap and plentiful but biased. Users click higher positions more often whatever their relevance, so naive click-through rate mostly measures where your current model put things. To use clicks as labels you correct for position bias. One approach estimates the probability that a user examines each position, often by occasionally swapping adjacent results for a small slice of traffic, and then weights each click by the inverse of that probability. A click at position 8, where few users look, counts for much more than a click at position 1.

# Inverse propensity weighting: a debiased click weight, not yet a training label
# propensity[pos] = estimated P(user examines position pos), from swap experiments
def click_weight(clicked: bool, pos: int, propensity: dict[int, float]) -> float:
    if not clicked:
        return 0.0
    return 1.0 / max(propensity[pos], 0.05)   # clip so rare positions do not explode

Sum the weights per query and document, then bucket the totals into integer grades such as 0 to 3, because ranking objectives expect integer relevance labels. Clip the weights, as above, or a few deep-position clicks will dominate. And keep the judged set separate from any click-derived training data, so that you have at least one evaluation source the model cannot have learned to game.

Training a LambdaMART ranker

Learning to rank comes in three flavours. Pointwise methods predict a relevance grade for each document independently. Pairwise methods learn which of two documents should rank higher. Listwise methods optimize a ranking metric over the whole list. LambdaMART, gradient-boosted trees trained with gradients scaled by how much swapping two documents would change NDCG, is the long-standing workhorse for tabular ranking features and is available in LightGBM, XGBoost and CatBoost.

The one detail that surprises newcomers is grouping. A ranking model is trained on lists, so you tell the library how many consecutive rows belong to each query.

import lightgbm as lgb
import pandas as pd

train = pd.read_parquet("train.parquet").sort_values("qid")   # rows grouped by query
valid = pd.read_parquet("valid.parquet").sort_values("qid")
FEATURES = ["bm25_title", "bm25_body", "ann_cos", "ctr_30d",
            "price_z", "in_stock", "cat_match", "freshness_days"]

ranker = lgb.LGBMRanker(
    objective="lambdarank",
    n_estimators=600,
    learning_rate=0.05,
    num_leaves=63,
    min_child_samples=50,
)
ranker.fit(
    train[FEATURES], train["label"],
    group=train.groupby("qid").size().to_numpy(),
    eval_set=[(valid[FEATURES], valid["label"])],
    eval_group=[valid.groupby("qid").size().to_numpy()],
    eval_at=[10],
    callbacks=[lgb.early_stopping(50)],
)
ranker.booster_.save_model("l2-2026-10-01.txt")

Split training and validation by query, not by row, or rows from the same query leak across the split. Split by time as well when you can, training on older weeks and validating on the most recent, because that is the situation the model will face in production.

Measure offline with NDCG, decide online with interleaving

Normalized discounted cumulative gain rewards putting highly graded documents near the top. Each result's gain is discounted by the logarithm of its position, summed over the top k, and divided by the best possible sum for that query so that scores are comparable across queries.

import math

def dcg(grades, k=10):
    return sum((2 ** g - 1) / math.log2(i + 2) for i, g in enumerate(grades[:k]))

def ndcg(ranked_grades, k=10):
    ideal = dcg(sorted(ranked_grades, reverse=True), k)
    return dcg(ranked_grades, k) / ideal if ideal > 0 else 0.0

Offline NDCG is a filter, not a verdict. A candidate model that loses offline is usually not worth testing live; one that wins still has to prove itself with real users. A/B tests on search need large samples because per-user behaviour is noisy. Interleaving is far more sensitive. Each user gets a single merged list built from both models, and the model whose contributions attract more clicks wins that query. Team-draft interleaving builds the merged list like picking sports teams:

import random

def team_draft(list_a, list_b, k=10):
    merged, team, used = [], {}, set()
    ia = ib = 0
    while len(merged) < k and (ia < len(list_a) or ib < len(list_b)):
        n_a = sum(1 for t in team.values() if t == "A")
        n_b = len(team) - n_a
        a_turn = n_a < n_b or (n_a == n_b and random.random() < 0.5)
        if (a_turn and ia < len(list_a)) or ib >= len(list_b):
            src, idx, label = list_a, ia, "A"
        else:
            src, idx, label = list_b, ib, "B"
        while idx < len(src) and src[idx] in used:
            idx += 1
        if idx < len(src):
            merged.append(src[idx]); used.add(src[idx]); team[src[idx]] = label
        if label == "A": ia = idx + 1
        else: ib = idx + 1
    return merged, team   # credit each click to team[doc]

Treat this as a sketch to adapt: production versions handle ties and exhausted lists more carefully, and they log the team assignment with the impression so credit is computed from the same log the ranker uses. Use interleaving to pick the better ranker quickly, then confirm the winner with an A/B test on business metrics such as conversions, which interleaving does not measure.

Worked example: shipping a new L2 ranker

In this illustrative scenario, a retailer's search team sees that queries containing a size or colour often show out-of-stock or wrong-size items on the first page. Query understanding already extracts the size, but only as a boost, and the L2 model has no feature for stock in the requested size.

They add two features, size_in_stock and colour_match, to the serving path first, so that they start appearing in the impression log. After two weeks they have enough logged impressions to train. The judged set of 3,000 queries shows NDCG@10 rising from 0.612 to 0.637 overall and much more on the 400 judged queries that mention a size. Retrieval recall at 1,000 is unchanged, as expected, because only the ranker changed.

An interleaving test on 5% of traffic runs for four days. The new model wins 54% of queries with a decisive click preference, ties are excluded, and the confidence interval does not include 50%. A follow-up A/B test shows a small increase in add-to-cart rate with no drop in revenue per search, and the model is promoted. The team keeps the old model file and its feature schema so they can roll back in one deploy.

Failure modes

  • Training-serving skew. Features recomputed offline differ from what was served. Train only from the impression log, and compare feature distributions between training data and live traffic on every release.
  • Feedback loops. A model trained on clicks from its own rankings keeps reinforcing what it already shows. Debias clicks, keep a judged set, and reserve a little traffic for exploration.
  • Recall ceiling. The ranker is blamed for missing documents the retrievers never returned. Track recall at each stage cutoff separately.
  • Over-confident filters. An entity tagger turns a guess into a hard filter and the page comes back empty. Use boosts unless confidence is high, and fall back to a relaxed query on zero results.
  • Latency creep. Each new feature adds a lookup. Keep a per-stage budget, and alert when the 99th percentile of any stage exceeds it.
  • Metric gaming. NDCG rises because labels leaked from validation into training, or because the judged set no longer resembles real traffic. Refresh judged queries from recent logs every quarter.

What to do next

  1. Write down the cascade you actually run: stages, candidate counts and the millisecond budget for each, and measure the 99th percentile against it.
  2. Add an impression log that records the exact feature vector, position and model version for every shown result and a sample of unshown candidates.
  3. Build a judged set of at least a few hundred real queries with graded labels and written rating guidelines, and compute NDCG@10 for your current ranker.
  4. Measure recall at your retrieval cutoff on that set, and fix retrieval first if it is the bottleneck.
  5. Run a small position-swap experiment to estimate examination propensities, then train a LambdaMART model on debiased clicks plus judgments, split by query and time.
  6. Set up team-draft interleaving so each candidate ranker can be compared in days, and confirm winners with an A/B test on business metrics.
  7. For prefix suggestions as the user types, study typeahead architecture, which runs on a much tighter budget than full search.
Key takeaway: An index makes search fast; the relevance architecture makes it good. Run a cascade in which cheap stages protect recall and expensive stages order the top, each with a stated candidate count and latency budget. Log the exact features you served, label with both graded judgments and debiased clicks, train a LambdaMART ranker split by query and time, filter candidates with offline NDCG and choose between them with interleaving. Every one of those steps is a component you can build, measure and roll back.