A Spring AI application that answers questions from your documents can be wrong in two different ways: it can retrieve the wrong evidence, or it can retrieve the right evidence and then say something the evidence does not support. Unit tests catch neither, because the output is free text that changes with every model release and prompt edit. Evaluation is the practice of scoring those outputs automatically, usually by asking a second model to judge them, so that a prompt change or a model upgrade produces a number you can compare instead of a feeling.

Spring AI ships a small evaluation API: one interface, one request type, one response type and two ready-made evaluators. This article reads that API closely, shows exactly what the built-in evaluators check (it is narrower than their names suggest), and builds a regression harness around it that you can run in CI. Everything specific here was checked against the Spring AI 2.0.0 jars; if you are on another version, confirm the details there.

The contract: one interface, two value types

The contract lives in org.springframework.ai.evaluation. Evaluator has one abstract method, EvaluationResponse evaluate(EvaluationRequest request), and one default helper, doGetSupportingData(request), which joins the text of every document in the request with the platform line separator, skipping documents whose text is null.

EvaluationRequest carries three things: the user text (the question), a List<Document> called the data list (the evidence the answer was supposed to rely on) and the response content (the answer). Constructors exist for (data list, response), (user text, response) and all three. EvaluationResponse carries a boolean isPass(), a float getScore(), a feedback string and a metadata map.

That is the whole framework. There is no dataset type, no runner and no report: Spring AI gives you a scoring function and leaves orchestration to your test code. That is a reasonable split, because the orchestration is where your domain lives, but it means the quality of your evaluation depends almost entirely on what you build around these types.

What the built-in evaluators actually check

Spring AI 2.0.0 ships two evaluators in spring-ai-client-chat, package org.springframework.ai.chat.evaluation. Both take a ChatClient.Builder, render a prompt, call the judge model and compare its reply to the word yes.

RelevancyEvaluatorFactCheckingEvaluator
Prompt asksIs the response for the query in line with the context?Is the claim supported by the document?
Inputs useduser text, response, joined documentsresponse (as claim), joined documents (as document)
Pass rulereply.strip() equals "yes", ignoring casesame
Score1.0 on pass, 0.0 on failalways 0.0 (the constructor used sets no score)
Feedbackempty stringempty string
ConstructionRelevancyEvaluator.builder().chatClientBuilder(b)FactCheckingEvaluator.builder(b) or forBespokeMinicheck(b)

Three consequences follow. First, the pass rule is exact. A judge that answers Yes. or YES - the response matches fails the case. Instruction-following models usually comply with a one-word answer, but chatty or reasoning-heavy judges do not, and the result is a false failure you will spend an afternoon chasing. Run a calibration pass and count replies that are neither yes nor no before trusting any rate.

Second, despite its name, RelevancyEvaluator's prompt asks whether the answer is in line with the context, which is a groundedness question, not whether the answer addresses the query. An answer that faithfully quotes the context but ignores the question can pass. Third, neither evaluator explains itself and FactCheckingEvaluator reports a score of zero even on a pass, so aggregate isPass(), never getScore(), for that one.

forBespokeMinicheck swaps in a terse prompt shaped for the Bespoke MiniCheck fact-checking model, typically served locally through Ollama. It is a good fit when you want a cheap, specialised grounding check on every case rather than a general model.

Capturing the evidence the answer used

One evaluation case, from question to verdictDataset casequestion + expectationsApp under testChatClient + RAG advisorAnswer textRetrieved documentsrag_document_contextuser textEvaluationRequestuserText, dataList, responseassembleFactCheckingEvaluatorgrounded?Graded rubric evaluatorscore 1-5 + reasonJudge modelseparate clientReport: pass rate + interval per sliceThe judge is a second model call per evaluator per case: budget it like production traffic.
The evaluator sees the question, the documents actually retrieved and the answer; the judge runs on its own client.

An evaluator is only as honest as the data list you hand it. If you pass the documents you expected the app to retrieve, you are testing generation in isolation. If you pass the documents it actually retrieved, you are testing the real pipeline. You usually want the second, and Spring AI makes it available: RetrievalAugmentationAdvisor stores the retrieved documents in the response context and copies them into the ChatResponse metadata under the key RetrievalAugmentationAdvisor.DOCUMENT_CONTEXT (the string rag_document_context).

record Observed(String answer, List<Document> evidence) {}

Observed ask(ChatClient app, String question) {
    ChatResponse response = app.prompt()
            .user(question)
            .call()
            .chatResponse();
    List<Document> docs = response.getMetadata()
            .get(RetrievalAugmentationAdvisor.DOCUMENT_CONTEXT);
    return new Observed(
            response.getResult().getOutput().getText(),
            docs == null ? List.of() : docs);
}

Keep the evidence with the case result. When a case fails, the first question is always whether retrieval or generation broke, and the stored documents answer it in seconds. If your app uses QuestionAnswerAdvisor or a custom retriever, record the documents yourself at the point of retrieval rather than re-running the search later, because the index may have changed in between.

A graded rubric evaluator

A yes/no verdict is coarse. To track gradual drift you want a graded score and a reason, and that is a short custom Evaluator. Ask the judge for structured output so the parsing problem above disappears: entity(Class) converts the reply into a record.

public final class RubricEvaluator implements Evaluator {

    public record Verdict(int score, String reason) {}

    private static final String RUBRIC = """
        You grade an answer to a user question using ONLY the evidence given.
        5 = fully answers the question, every claim supported by the evidence
        4 = answers the question, one minor unsupported detail
        3 = partially answers, or one material unsupported claim
        2 = mostly unsupported or mostly off-question
        1 = contradicts the evidence or refuses when the evidence suffices
        Question: {question}
        Evidence:
        {evidence}
        Answer: {answer}
        """;

    private final ChatClient judge;
    private final int passAt;

    public RubricEvaluator(ChatClient judge, int passAt) {
        this.judge = judge;
        this.passAt = passAt;
    }

    @Override
    public EvaluationResponse evaluate(EvaluationRequest req) {
        Verdict v = judge.prompt()
                .user(u -> u.text(RUBRIC)
                        .param("question", req.getUserText())
                        .param("evidence", doGetSupportingData(req))
                        .param("answer", req.getResponseContent()))
                .call()
                .entity(Verdict.class);
        if (v == null || v.score() < 1 || v.score() > 5) {
            return new EvaluationResponse(false, 0f, "unparseable verdict", Map.of());
        }
        return new EvaluationResponse(v.score() >= passAt, v.score() / 5f,
                v.reason(), Map.of("rubric_score", v.score()));
    }
}

Build the judge's client from its own model with ChatClient.builder(judgeModel), and in the custom evaluator add .options(ChatOptions.builder().temperature(0.0)) to the request spec for low variance. Do not reuse the application's ChatClient.Builder: the built-in evaluators call build() on the builder you give them, so any default advisors, system prompt or tools registered on the application's builder would run inside the judge as well. A judge that silently performs retrieval grades a different question.

Use a different model family for the judge where you can. A model grading its own output tends to approve its own phrasing; even a smaller model from another vendor gives a more independent signal for groundedness, which is a comparatively easy judging task.

A regression harness in JUnit

With an observation function and evaluators in hand, the harness is a parameterised test over a versioned dataset. Each case carries a question, a slice label (billing, returns, out-of-scope) and optionally the ids of documents that must be retrieved, which gives you a deterministic retrieval check alongside the judged ones.

@Tag("eval")
class SupportAnswersEvalTest {

    @ParameterizedTest(name = "{0}")
    @MethodSource("cases")
    void answer_is_grounded(EvalCase k) {
        Observed o = ask(app, k.question());
        Set<String> got = o.evidence().stream().map(Document::getId).collect(toSet());
        recorder.retrieval(k, got.containsAll(k.mustRetrieve()));

        var req = new EvaluationRequest(k.question(), o.evidence(), o.answer());
        recorder.verdict(k, "grounded", factCheck.evaluate(req));
        recorder.verdict(k, "rubric", rubric.evaluate(req));
    }

    @AfterAll
    static void gate() {
        recorder.writeReport(Path.of("build/eval/report.json"));
        recorder.assertNoSliceBelow("grounded", 0.90);
    }
}

Note that the test does not assert per case. Judged evaluation is noisy, and one flaky case should not fail a build; the gate is on aggregate rates per slice, written once all cases have run. The recorder is your own small class: it keeps results in memory and computes rates and intervals. For dataset versioning, stratification and statistics in more depth, see dataset-driven evaluation pipelines; for wiring the run into CI with recorded responses, see integrating evals into CI.

Worked example: is that drop real?

Suppose a returns-policy assistant has 60 cases in three slices of 20. A prompt change is proposed. The baseline run gives grounded pass rates of 19/20, 18/20 and 20/20; the candidate gives 19/20, 15/20 and 20/20. Is the middle slice a regression?

Compute a Wilson 95% interval for each rate rather than eyeballing. For 18/20 it is roughly 0.70 to 0.97; for 15/20 roughly 0.53 to 0.89. The intervals overlap heavily, so 20 cases cannot distinguish the two runs. Two practical moves follow. First, read the three newly failing cases with their stored evidence: if all three cite the same document that the new prompt now paraphrases loosely, you have found a real, specific regression without needing statistics. Second, grow that slice: 80 cases would shrink the interval width to around 0.15, enough to see a 15-point drop.

Then check the judge itself. Re-run the 60 baseline cases twice with the same inputs. If verdicts flip on more than two or three cases between identical runs, the judge's noise is as large as the effect you are trying to measure, and the fix is a stricter rubric or a different judge, not more test cases.

Failure modes

  • Non-literal yes. The built-in evaluators fail any reply other than a bare yes. Symptom: a sudden pass-rate collapse after switching judge model. Count replies that are neither yes nor no in a calibration run.
  • Judge inherits app behaviour. Passing the application's builder to an evaluator runs its advisors and tools inside the judge. Give the judge its own client.
  • Empty evidence passes. With an empty data list the judge sees no context; depending on the model it may say yes to anything plausible. Treat an empty data list on an in-scope case as a retrieval failure before asking the judge anything.
  • Score misread. Averaging FactCheckingEvaluator scores always yields zero. Aggregate pass flags.
  • Evidence drift. Re-querying the store after the run evaluates different documents from the ones the answer used. Capture evidence at answer time.
  • Dataset rot. Policies change, and a case whose expected facts are stale makes a correct answer look wrong. Date each case and review old ones on a schedule.

Operating evaluation

Every evaluator is a model call per case. Sixty cases with two judged evaluators is 120 judge calls plus 60 application calls, each carrying the full evidence text. At a few thousand tokens per judge prompt that is a few hundred thousand tokens per run, which is cheap nightly and expensive on every commit. A common split: deterministic retrieval checks on every pull request, judged evaluation nightly and before releases.

Pin the judge model version and record it, along with the application model, prompt version and index snapshot, in the report. A pass rate without that provenance cannot be compared with next week's. Run judge calls with bounded concurrency and retries, because an evaluation run is a burst of traffic that can trip the same rate limits as production.

Offline evaluation tells you whether a change is safe to ship; it does not tell you how production behaves. Sampling live turns and judging them asynchronously closes that gap, as described in continuous evaluation in production. For a rubric judge with calibration against human labels, the approach in the LLM-as-judge scorer carries over directly to a Spring AI Evaluator.

Trade-offs

ChoiceGainsCosts
Built-in yes/no evaluatorsZero code, simple ratesBrittle parsing, no reasons, narrow question
Graded rubric evaluatorDetects drift, explains failuresRubric design, calibration effort
Deterministic retrieval checksFree, exact, fastNeeds labelled must-retrieve ids
Same model as judgeOne provider, one billSelf-preference bias
Judging every commitFast feedbackToken cost, flaky gates

The strongest setup combines cheap exact checks with a small number of judged ones, and treats the judge as a component with its own tests rather than as ground truth.

What to do next

  1. Collect 30 to 60 real questions per important slice, each with expected facts and, where possible, the ids of documents that must be retrieved.
  2. Capture the actual retrieved documents from rag_document_context on every evaluation call.
  3. Build a separate judge ChatClient from a different model at low temperature.
  4. Run a calibration pass: count non-literal replies and repeat-run verdict flips.
  5. Add a graded rubric evaluator with structured output for drift tracking.
  6. Gate on per-slice pass rates with intervals, nightly, and record full provenance in the report.
Key takeaway: Spring AI gives you a scoring interface and two narrow yes/no judges; the harness, dataset and statistics are yours. Feed evaluators the documents the answer actually used, run the judge on its own client and model, prefer structured graded verdicts, aggregate pass flags per slice with intervals, and test the judge's own noise before trusting a regression.