The data an evaluation case needs

Component-level metrics need component-level labels. A useful RAG evaluation case contains the question, any filters or user attributes that affect retrieval, a reference answer, the gold evidence that supports that answer, and a list of the atomic facts, or nuggets, that a complete answer must contain. Gold evidence should be recorded as document identifiers plus character spans or quoted passages, not as chunk identifiers, because chunk boundaries change every time the chunking strategy is tuned. If the labels are tied to chunk ids, a chunking change silently invalidates the whole dataset.

Some cases should be deliberately unanswerable: questions whose answer is not in the corpus, or is present only in documents the user is not permitted to see. For these the correct behavior is to decline or say the information is not available, and the evaluation should score that explicitly. A system that is never tested on unanswerable questions learns, through prompt tuning, to always produce something, which is how confident hallucinations reach production.

{
  "id": "billing-0142",
  "question": "Can I move my annual plan to monthly billing mid-term?",
  "filters": {"region": "EU", "product": "pro"},
  "reference_answer": "Yes, from the next renewal date; mid-term switches are not prorated.",
  "evidence": [
    {"doc_id": "kb/billing/plan-changes", "quote": "Changes from annual to monthly take effect at renewal"},
    {"doc_id": "kb/billing/proration", "quote": "Annual plans are not prorated when downgraded"}
  ],
  "nuggets": [
    {"text": "switch is possible", "vital": true},
    {"text": "takes effect at next renewal", "vital": true},
    {"text": "no proration or partial refund", "vital": false}
  ],
  "answerable": true
}
Advertisement

Retrieval: did the evidence arrive?

The first question is whether the retriever returned the evidence at all. Evidence recall at k is the fraction of gold evidence passages that appear, by span overlap, in the top k retrieved results. For questions that need several pieces of evidence, report both the average fraction and the share of questions for which all required evidence was retrieved; the second number is the one that predicts whether a complete answer was even possible.

Ranking metrics add position. Mean reciprocal rank rewards putting the first relevant result near the top, which matters when only the top few results survive into the prompt. Normalized discounted cumulative gain handles graded relevance, where some passages are more useful than others. Measure these at the k the system actually uses, and at the retriever's candidate depth before reranking, because a reranker can only promote evidence that the first stage retrieved. If recall at the candidate depth is low, no amount of reranker tuning will help.

def overlaps(hit, ev, min_chars=40):
    if hit["doc_id"] != ev["doc_id"]:
        return False
    # span or quote match; quotes survive re-chunking, chunk ids do not
    return ev["quote"][:min_chars].lower() in hit["text"].lower()

def evidence_recall(hits, evidence, k):
    top = hits[:k]
    found = [any(overlaps(h, ev) for h in top) for ev in evidence]
    return sum(found) / len(found), all(found)

def reciprocal_rank(hits, evidence):
    for rank, h in enumerate(hits, 1):
        if any(overlaps(h, ev) for ev in evidence):
            return 1.0 / rank
    return 0.0

Quote matching is deliberately simple here. In practice, normalize whitespace and punctuation, and fall back to a fuzzy match or an entailment check when the corpus is re-extracted from PDFs and the text changes slightly between versions.

Advertisement

Context precision: how much noise reached the model

Recall asks whether the right passages arrived; context precision asks what else came with them. If the answer depends on one paragraph and the prompt carries fifteen, the generator has to find it, and models are measurably worse at using information placed in the middle of a long context than at its beginning or end, a pattern documented in the 2023 study Lost in the Middle. Irrelevant passages also invite the model to blend in plausible but wrong details from neighbouring documents, such as the policy for a different product tier.

Measure context precision as the fraction of passages in the final assembled context that are relevant to the question, where relevance is judged either against the gold evidence or by a grader for passages the labels do not cover. Report the rank of the first gold passage in the final context and the total context length in tokens alongside it. These three numbers together show whether a change to top-n, reranker thresholds or chunk size improved signal, or simply added volume and cost.

Faithfulness: is every claim supported?

Faithfulness, sometimes called groundedness, measures whether the answer's claims are supported by the retrieved context. The standard approach decomposes the answer into atomic claims, then checks each claim against the context with a natural language inference model or an LLM grader, and reports the supported fraction. A short answer with one unsupported number scores lower than a long answer with none, which matches what users experience.

Keep faithfulness and correctness separate. Faithfulness is measured against the retrieved context, correctness against the reference answer. An answer can be perfectly faithful to an outdated document and still wrong, which points to a corpus freshness problem, not a generation problem. An answer can also be correct but unfaithful, because the model knew the answer from pretraining; that looks harmless until the corpus holds organization-specific facts that contradict public ones, at which point the same behavior produces confident errors.

CLAIM_PROMPT = """Split the answer into short, self-contained factual claims.
Resolve pronouns. Omit greetings and hedges. Return a JSON list of strings.
Answer: {answer}"""

VERIFY_PROMPT = """Context:
{context}

Claim: {claim}
Is the claim fully supported by the context? Reply with one word:
SUPPORTED, CONTRADICTED, or NOT_FOUND."""

def faithfulness(answer, context, llm):
    claims = llm.json(CLAIM_PROMPT.format(answer=answer))
    if not claims:
        return None, []                       # abstention: score separately
    verdicts = [llm.label(VERIFY_PROMPT.format(context=context, claim=c)) for c in claims]
    supported = sum(v == "SUPPORTED" for v in verdicts)
    contradicted = [c for c, v in zip(claims, verdicts) if v == "CONTRADICTED"]
    return supported / len(claims), contradicted

Track contradicted claims separately from unsupported ones. An unsupported claim may be a harmless elaboration; a claim that contradicts the context is almost always a defect and deserves its own alert. If the system emits citations, add a citation check: every cited passage should support the sentence that cites it, and every factual sentence should cite something.

Relevance: did it answer the question asked?

A faithful answer can still miss the point, for example by summarizing the retrieved refund policy when the user asked whether a specific refund was possible. Answer relevance scores whether the response addresses the question as asked, at the level of specificity asked. One published technique generates candidate questions from the answer and measures their embedding similarity to the original question; it is cheap, but it rewards answers that restate the question and penalizes correct answers phrased differently. A grader with a short rubric, asking whether the response directly answers the question, whether it answers a different question, and whether it adds material the user did not need, is usually more reliable and easier to audit.

Relevance also covers refusal behavior. For answerable questions, a refusal is a relevance failure. For unanswerable questions, a correct refusal is the only relevant answer, and anything else should be scored as a hallucination regardless of how faithful its individual claims appear. Report the false refusal rate and the false answer rate as a pair, because prompt changes usually trade one for the other.

Completeness: nugget recall

Completeness asks how much of what the user needed made it into the answer. The nugget approach, which comes from question answering evaluation at TREC, lists the atomic facts a complete answer should contain and marks each as vital or supplementary. The grader then checks which nuggets appear in the answer. Nugget recall over vital nuggets is the headline number; supplementary nuggets break ties. Because nuggets are defined once per case, the metric is stable across answer styles: a terse answer that contains every vital nugget scores the same as a long one.

NUGGET_PROMPT = """Answer: {answer}

For each numbered fact, reply YES if the answer states it (paraphrase is fine),
PARTIAL if it is implied or incomplete, NO otherwise. Return a JSON list.
{facts}"""

WEIGHT = {"YES": 1.0, "PARTIAL": 0.5, "NO": 0.0}

def nugget_recall(answer, nuggets, llm):
    facts = "\n".join(f"{i + 1}. {n['text']}" for i, n in enumerate(nuggets))
    marks = llm.json(NUGGET_PROMPT.format(answer=answer, facts=facts))
    vital = [WEIGHT[m] for m, n in zip(marks, nuggets) if n["vital"]]
    allw = [WEIGHT[m] for m in marks]
    return {
        "vital_recall": sum(vital) / len(vital) if vital else None,
        "all_recall": sum(allw) / len(allw),
        "all_vital_present": all(v == 1.0 for v in vital),
    }

Nugget labels are expensive, so start with the highest-traffic intents and the cases where incomplete answers have caused support escalations. A few hundred well-labeled cases are more useful than thousands with only a reference answer.

Reading the metrics together

The value of component metrics comes from reading them as a set. If evidence recall at the candidate depth is low, the problem is in retrieval: embeddings, query rewriting, filters or the index itself. If recall is high but the gold passage sits low in the final context and precision is poor, the problem is in reranking or context assembly. If recall and precision are high but faithfulness is low, the generator is ignoring or distorting the context, which points at the prompt, the model or the decoding settings. If faithfulness is high and completeness low, the model is being cautious or the context was truncated. And if everything is high except correctness against the reference, check the corpus for stale or conflicting documents.

Queryquestion + filtersRetrievertop-k candidatesRerankertop-n contextGeneratoranswer + citationsContext recallgold evidence found?Context precisionnoise in top-nFaithfulnessclaims supportedRelevanceaddresses the questionCompletenessnugget recallDiagnosis: low recall = retrieval; high recall + low faithfulness = generation; faithful but incomplete = context assembly or prompt
Each RAG stage gets its own metric: evidence recall and ranking at retrieval, context precision after reranking, and faithfulness, relevance and completeness at generation. Reading them together localizes a regression to one stage.

In a regression report, show these metrics for baseline and candidate side by side, per slice, with paired confidence intervals. A chunking change that raises recall by four points and lowers faithfulness by three is not a wash; it moved the problem downstream, and the flipped cases will show exactly where.

Where the metrics mislead

  • Grader leniency. LLM graders tend to accept claims that are plausible but only loosely supported. Calibrate the faithfulness grader on a few hundred human-labeled claims and report its agreement.
  • Claim granularity. Splitting an answer into more, smaller claims inflates faithfulness when the extra claims are trivially supported. Pin the decomposition prompt and model, and treat changes to either as a metric version change.
  • Label drift. Evidence labels rot as documents are edited, merged or deleted. Re-validate that each gold quote still exists in the corpus before every run and quarantine cases whose evidence has disappeared.
  • Context-free correctness. A model that answers from memory can score well on public-knowledge questions. Include organization-specific facts that contradict common knowledge to detect it.
  • Averaging across intents. Aggregate scores hide a failing intent behind a large healthy one. Always slice by intent, document source and answerability.

Running it in production

Offline, run the full metric set on every change to the retriever, chunker, reranker, prompt or model, and gate on the vital-nugget recall and faithfulness of the critical slices. Online, sample a small fraction of live traffic and run the reference-free metrics, context precision, faithfulness and relevance, which need no gold labels. Alert on contradicted claims and on sudden shifts in retrieval score distributions, which often indicate an index or embedding model mismatch after a deployment.

Feed failures back into the dataset. Every production complaint about a wrong or incomplete answer becomes a labeled case with evidence and nuggets, and every case that no configuration can pass gets reviewed for a bad label or a corpus gap. Over time the suite describes, precisely, what the system must know and where it is allowed to say it does not know.