Most retrieval-augmented generation teams can name their metrics: recall, context precision, faithfulness, answer relevance. Fewer can say whether a two-point rise in faithfulness last Tuesday was real, whether their LLM judge catches the hallucinations it is supposed to catch, or what k their recall is measured at. Those are measurement questions, and getting them wrong produces dashboards that move for no reason and releases that ship regressions because a noisy average happened to go up.

Which metric belongs to which stage of a RAG pipeline, and how to read them together to localize a failure, is covered in RAG evaluation: recall, faithfulness, relevance and completeness. This page is about computing those numbers correctly and deciding what they mean: exact definitions, judge validation, sample sizes and a release decision backed by a confidence interval.

Advertisement

Every metric is a mean of per-question scores

An evaluation run sends each question in a fixed set through the system and scores it, producing one row per question: the retrieved ids, the context passed to the model, the answer, and a score for each metric. The headline number for a metric is the mean of that column. Keep the rows, not just the means: paired comparisons, slicing, re-scoring and debugging all need them.

Two kinds of scorer fill the rows. Deterministic scorers compare retrieved ids with labelled gold ids and need no model; they are cheap, exact and reproducible. Judge scorers ask a language model whether an answer is supported, relevant or correct; they measure what users care about, but they are noisy and are themselves a model whose errors you must measure.

A RAG evaluation run: per-question scores, compared pairwise against a baselineFrozen eval setquestions, gold ids, refsSystem under testretriever + generatorDeterministic scorersrecall@k, MRR, nDCGJudge scorersfaithfulness, relevanceretrieved idscontext, answerPer-question scoresone row per itemHuman-labelled subsetvalidates the judgeagreement, kappaBaseline scoressame questionsPaired bootstrapCI on the per-question differenceRelease gateblock if CI shows a real dropEvery number in the report is a mean of per-question scores; the decision uses the paired difference, not two means.
Scores are kept per question. The judge is validated on a human-labelled subset, and the release decision uses a paired bootstrap on per-question differences against the baseline.

Retrieval metrics, defined exactly

Given a ranked list of retrieved ids and a set of gold ids for one question: recall@k is the fraction of gold ids that appear in the top k; precision@k is the fraction of the top k that are gold; reciprocal rank is one over the rank of the first gold id, and MRR is its mean over questions. nDCG@k uses graded relevance: each result contributes a gain discounted by the logarithm of its rank, and the sum is divided by the best possible sum for that question, so 1.0 means the ideal order.

import math

def recall_at_k(ranked, gold, k):
    return len(set(ranked[:k]) & gold) / len(gold)

def precision_at_k(ranked, gold, k):
    return len(set(ranked[:k]) & gold) / k

def reciprocal_rank(ranked, gold):
    for i, doc in enumerate(ranked, start=1):
        if doc in gold:
            return 1.0 / i
    return 0.0

def ndcg_at_k(ranked, grades, k):
    """grades: doc -> graded relevance (0, 1, 2, 3)."""
    dcg = sum((2 ** grades.get(d, 0) - 1) / math.log2(i + 1)
              for i, d in enumerate(ranked[:k], start=1))
    ideal = sorted(grades.values(), reverse=True)[:k]
    idcg = sum((2 ** g - 1) / math.log2(i + 1) for i, g in enumerate(ideal, start=1))
    return dcg / idcg if idcg else 0.0

ranked = ["d7", "d2", "d9", "d4", "d1"]
grades = {"d2": 3, "d4": 1, "d5": 2}
gold = set(grades)
print(f"recall@5    = {recall_at_k(ranked, gold, 5):.3f}")
print(f"precision@5 = {precision_at_k(ranked, gold, 5):.3f}")
print(f"RR          = {reciprocal_rank(ranked, gold):.3f}")
print(f"nDCG@5      = {ndcg_at_k(ranked, grades, 5):.3f}")

Running it prints:

recall@5    = 0.667
precision@5 = 0.400
RR          = 0.500
nDCG@5      = 0.516

Two of three gold documents are in the top five, so recall is 0.667, and two of five results are gold, so precision is 0.4. The first gold document is at rank 2, so reciprocal rank is 0.5. For nDCG, the highly relevant d2 at rank 2 contributes 7 / log2(3), about 4.42, and the marginal d4 at rank 4 contributes 1 / log2(5), about 0.43, for 4.85; the ideal order (grades 3, 2, 1) would score 7 + 3 / log2(3) + 1 / 2, about 9.39, and 4.85 / 9.39 is 0.516. Moving d2 to rank 1 alone would raise nDCG to 0.791, while fetching d5 at rank 3 would give 0.676: rank order matters most.

Four definitional choices change these numbers more than most model changes do, so fix and document them. First, k must be the number of chunks that actually reach the model after reranking and truncation, not the retriever's raw candidate count. Second, gold labels should be documents or character spans, matched to chunks at scoring time, so that re-chunking does not invalidate the dataset. Third, nDCG has two gain conventions, 2^g - 1 as here and plain g; libraries differ, so state which. Fourth, questions with no answer in the corpus have no gold ids, and recall is undefined for them: exclude them from retrieval means and score them on whether the system declined to answer.

Advertisement

Answer metrics come from a judge

Faithfulness asks whether the answer's claims are supported by the retrieved context, regardless of whether they are true in the world. The robust way to compute it is claim-level: split the answer into atomic factual claims, ask the judge whether each is supported by the context, and score the fraction supported. A single overall verdict is cheaper but easily overlooks one unsupported number in a long answer. Answer relevance asks whether the answer addresses the question; correctness compares the answer with a reference answer. Open-source libraries such as Ragas, compared in AI evaluation frameworks, implement versions of these metrics, but their APIs and metric names have changed between releases, so check the current documentation and record the library version with every run.

def faithfulness(answer: str, context: str, judge) -> float | None:
    """Fraction of the answer's factual claims that the context supports."""
    claims = judge.extract_claims(answer)           # one LLM call: list of atomic claims
    if not claims:
        return None                                 # refusals and empty answers: score separately
    verdicts = [judge.is_supported(claim, context)  # one call per claim, or one batched call
                for claim in claims]                # each verdict: True / False
    return sum(verdicts) / len(claims)

# Store the claims and verdicts with the score: a number you cannot inspect is a number
# you cannot debug or re-label.

Pin everything that can change the judge's output: the model and version, the prompts, the temperature and the claim-splitting method. Change any of them and scores shift even when the system under test did not. Treat a judge change like a metric migration: re-score the baseline with the new judge before comparing anything against it.

Validate the judge against people

A judge is a classifier, and you need its error rates before you trust its averages. Have people label a sample, at least a couple of hundred items, with the same question the judge answers, then compare. Raw agreement alone misleads when one label dominates. In the example below, humans mark 150 of 200 answers as supported and 50 as unsupported; the judge agrees on 85 percent of items, which sounds good.

import random

def cohen_kappa(a, b):
    """a, b: equal-length lists of 0/1 labels from two raters."""
    n = len(a)
    observed = sum(x == y for x, y in zip(a, b)) / n
    pa, pb = sum(a) / n, sum(b) / n
    expected = pa * pb + (1 - pa) * (1 - pb)
    return (observed - expected) / (1 - expected)

def paired_bootstrap(base, cand, iters=10_000, seed=7):
    """base, cand: per-question scores for the same questions, same order."""
    rng = random.Random(seed)
    n = len(base)
    diffs = [c - b for b, c in zip(base, cand)]
    means = []
    for _ in range(iters):
        sample = [diffs[rng.randrange(n)] for _ in range(n)]
        means.append(sum(sample) / n)
    means.sort()
    lo, hi = means[int(0.025 * iters)], means[int(0.975 * iters)]
    return sum(diffs) / n, lo, hi

# Judge vs human on 200 faithfulness labels (1 = supported).
human = [1] * 150 + [0] * 50
judge = [1] * 142 + [0] * 8 + [1] * 22 + [0] * 28
agree = sum(h == j for h, j in zip(human, judge)) / len(human)
print(f"raw agreement = {agree:.3f}, kappa = {cohen_kappa(human, judge):.3f}")

# Two retriever configs scored on the same 300 questions.
rng = random.Random(1)
base = [1.0 if rng.random() < 0.72 else 0.0 for _ in range(300)]
cand = [b if rng.random() < 0.9 else 1.0 - b for b in base]
mean, lo, hi = paired_bootstrap(base, cand)
print(f"baseline {sum(base)/300:.3f}  candidate {sum(cand)/300:.3f}")
print(f"mean diff {mean:+.3f}, 95% CI [{lo:+.3f}, {hi:+.3f}]")

Running it prints:

raw agreement = 0.850, kappa = 0.559
baseline 0.717  candidate 0.673
mean diff -0.043, 95% CI [-0.077, -0.010]

Cohen's kappa corrects agreement for chance, and 0.559 is moderate, not good. The per-class view is worse and more useful: of the 50 answers people found unsupported, the judge passed 22. The judge flags 18 percent of answers as unsupported against a true 25 percent, and catches only 28 of 50 hallucinations, so a change that adds hallucinations moves its score about half as much as it should. Look at the judge's miss rate on the class you care about, then fix the judge: claim-level prompting, a stronger judge model, few-shot examples drawn from the disagreements. Re-measure on fresh labels, not the same ones you tuned on. Calibration and judge bias are covered in LLM-as-a-judge calibration and bias.

How many questions you need

A metric that is a pass rate behaves like a proportion, with standard error sqrt(p * (1 - p) / n). At p = 0.7 and n = 300 that is about 0.026, so the 95 percent interval on a single run's score is roughly plus or minus 5 points. Two independent runs would need to differ by more than that before the difference meant much, which is why comparing two headline numbers from separate dashboards is a poor way to decide anything.

Pairing helps a great deal. When baseline and candidate answer the same questions, most questions get the same score from both, and only the questions that changed contribute variance. The paired bootstrap above resamples per-question differences and reads off a 95 percent interval. In the example, the candidate scores 0.673 against 0.717, a difference of minus 0.043 with an interval of minus 0.077 to minus 0.010: entirely below zero, so this is a real drop, not noise. With unpaired runs and 300 questions, a 4-point drop is inside the noise.

Size the set from the smallest change you need to detect. Hundreds of questions detect changes of a few points with pairing; detecting a one-point change in a rare failure type needs thousands of questions of that type. Slice reports by question type, because an aggregate can hold steady while one slice collapses.

Turning metrics into a release gate

A gate needs a rule decided in advance. A useful one is non-inferiority with a margin: pass if the lower end of the interval on the difference is above minus the margin; block if the whole interval is below zero; otherwise send it to a person. Run it per metric and per important slice, and log the interval with the decision so reviewers see uncertainty as well as direction.

MARGIN = 0.02   # the largest drop you are willing to ship without discussion

def gate(metric, base, cand):
    mean, lo, hi = paired_bootstrap(base, cand)
    if lo >= -MARGIN:
        return "pass", f"{metric}: {mean:+.3f} [{lo:+.3f}, {hi:+.3f}]"
    if hi < 0:
        return "block", f"{metric}: real drop {mean:+.3f} [{lo:+.3f}, {hi:+.3f}]"
    return "review", f"{metric}: inconclusive {mean:+.3f} [{lo:+.3f}, {hi:+.3f}]"

Applied to the example with a margin of 0.02, the lower bound of minus 0.077 fails the pass test and the upper bound of minus 0.010 is below zero, so the gate blocks. Keep the evaluation set frozen and versioned so successive gates compare like with like, and add new questions, mined from production failures, as a new version rather than by editing the old one. A harness that runs this on every change is described in building an agent eval harness.

Worked example: a chunking change

A team replaces 512-token chunks with 256-token chunks and an added reranker, keeping five chunks in the prompt. On the frozen 400-question set, recall@5 rises from 0.78 to 0.84, with a paired interval entirely above zero. Judge faithfulness, which had been validated with kappa 0.74 and a 12 percent miss rate on unsupported answers, falls from 0.91 to 0.88, with an interval from minus 0.05 to minus 0.01.

The rows explain the conflict. Smaller chunks retrieve the right document more often but cut tables and definitions in half, and the claims the judge marks unsupported cluster on questions whose answers span a chunk boundary. The gate blocks on faithfulness. The team adds neighbouring-chunk expansion, re-runs the set, and gets recall@5 of 0.84 and faithfulness of 0.91, with intervals that pass both gates. Neither headline number alone would have found the cause; the per-question rows did.

Failure modes

  • Unvalidated judge: faithfulness averages from a judge that misses half of all unsupported claims.
  • Moving judge: a new judge model or prompt shifts every score, and the shift is read as a system change.
  • Wrong k: recall measured on the retriever's top 50 when the model sees five.
  • Chunk-id labels: gold labels tied to chunk ids silently go stale after re-chunking.
  • Unpaired comparisons: two means from different runs or different question sets compared as if the difference were signal.
  • Edited eval set: questions added or removed between runs, so comparisons mix systems and datasets.

What to do next

  1. Store per-question rows, including retrieved ids, context, answer, claims and verdicts, for every evaluation run.
  2. Write down k, the gold label unit, the nDCG gain convention and how unanswerable questions are scored.
  3. Label at least 200 items by hand for each judge metric; report kappa and the judge's miss rate on the failure class.
  4. Pin the judge model, prompt and library versions, and re-score the baseline whenever any of them changes.
  5. Compare candidates to the baseline with a paired bootstrap, and gate releases on a pre-agreed margin per metric and slice.
  6. Freeze and version the evaluation set, and grow it from production failures as new versions.
Key takeaway: RAG metrics are means of per-question scores, so keep the rows. Define retrieval metrics exactly, at the k the model sees and against document-level gold labels. Judge-based metrics such as faithfulness are estimates from a classifier, so measure the judge against human labels with kappa and its miss rate on the failures you care about, and pin its version. Compare candidate and baseline on the same questions with a paired bootstrap, and gate releases on a pre-agreed margin rather than on two headline numbers.