Retrieval-augmented generation answers a question by first searching your own documents, then asking the model to answer from what was found. The simplest version is one vector search and one prompt. Real questions break that version quickly: a follow-up such as "and for the EU plan?" means nothing without the conversation, a question phrased in the user's words misses documents phrased in the vendor's words, and ten retrieved chunks that all say the same thing crowd out the one that answers the question.
Spring AI's spring-ai-rag module answers these problems with a modular pipeline wrapped in one advisor, RetrievalAugmentationAdvisor. Store setup, ingestion and metadata filters are covered in Spring AI vector stores; this article walks the pipeline itself stage by stage, with the defaults read from the Spring AI 2.0.0 jars, what each stage costs, and how each one fails.
The pipeline and its six stages
The advisor's before() step builds a Query record from the user message: text(), history() (the earlier messages in the prompt) and context() (the advisor context map, which is how per-request parameters such as a filter expression reach the retriever). It then runs six stages in order. Only the retriever is mandatory; every other stage is optional or has a default.
| Stage | Interface | Shipped implementations | Default |
|---|---|---|---|
| Transform | QueryTransformer | RewriteQueryTransformer, CompressionQueryTransformer, TranslationQueryTransformer | none |
| Expand | QueryExpander | MultiQueryExpander | none |
| Retrieve | DocumentRetriever | VectorStoreDocumentRetriever | required |
| Join | DocumentJoiner | ConcatenationDocumentJoiner | concatenation |
| Post-process | DocumentPostProcessor | none shipped | none |
| Augment | QueryAugmenter | ContextualQueryAugmenter | contextual, empty context refused |
In after(), the advisor copies the documents it used into the ChatResponse metadata under RetrievalAugmentationAdvisor.DOCUMENT_CONTEXT (rag_document_context), which is what you log, cite and evaluate.
Before retrieval: transform and expand
Transformers rewrite the query before search; each shipped one is a chat-model call. CompressionQueryTransformer folds the conversation history and a follow-up into one standalone question, which is the fix for "and for the EU plan?". RewriteQueryTransformer turns a rambling question into a search-shaped one, and its targetSearchSystem option tells the rewriting model what kind of index it is writing for. TranslationQueryTransformer translates the query into the language your documents are in, which is cheaper than embedding everything twice when your embedding model is not multilingual.
The expander then turns one query into several. MultiQueryExpander asks the model for numberOfQueries variants (default 3) and can keep the original alongside them via includeOriginal. Set includeOriginal explicitly: the original question is often the best single query, and you do not want its presence to depend on a default that can change between versions.
ChatClient.Builder rewriter = ChatClient.builder(smallFastModel); // not the answering model
Advisor rag = RetrievalAugmentationAdvisor.builder()
.queryTransformers(
CompressionQueryTransformer.builder().chatClientBuilder(rewriter).build())
.queryExpander(MultiQueryExpander.builder()
.chatClientBuilder(rewriter)
.numberOfQueries(2)
.includeOriginal(true)
.build())
.documentRetriever(VectorStoreDocumentRetriever.builder()
.vectorStore(store)
.topK(5)
.similarityThreshold(0.50)
.build())
.documentPostProcessors(new CapAndDiversify(6, 2))
.build();Pre-retrieval calls are cheap per call and expensive in aggregate: they run on every question, in series, before the user sees anything. Use a small fast model for them by giving them their own builder, and add each stage only after an evaluation shows it moves retrieval quality.
Retrieval: defaults that matter
VectorStoreDocumentRetriever embeds each query and runs a similarity search. Its defaults are topK 4 and a similarityThreshold of 0.0, which accepts every result. With those defaults retrieval almost never returns an empty list, and that silently disables a safety feature downstream (see augmentation below). Set a threshold measured on your own data: embed twenty in-scope and twenty out-of-scope questions, look at the top scores of each group, and pick a value between them.
The retriever takes a filter expression at build time, or per request through the query context key VectorStoreDocumentRetriever.FILTER_EXPRESSION; a per-request filter wins. That is the hook for tenant isolation, covered with its injection caveats in the vector store article. When the expander produces several queries, the advisor retrieves for them in parallel on its taskExecutor. Pass your own executor if you need bounded concurrency or context propagation for tracing, since each parallel search is a call to your embedding provider and your store.
The shipped retriever is pure vector similarity. Dense embeddings handle paraphrase well and exact tokens badly: SKUs, error codes, version numbers and surnames often rank below chunks that are merely about the same subject. Because DocumentRetriever is a one-method interface, a hybrid retriever that runs a keyword query and a vector query and fuses the two rankings drops into the same slot without touching the rest of the pipeline. Decide on evidence: if your failed cases cluster around identifiers, hybrid retrieval is the fix, and no amount of query rewriting will substitute for it.
Joining and post-processing
The joiner receives a map from each query to its result lists and returns one list. ConcatenationDocumentJoiner flattens everything, removes duplicates (the same document found by several queries appears once) and sorts by score, highest first. It does not truncate. With three queries and topK 5 you can hand up to 15 documents to the prompt; with five queries and topK 8, up to 40. Context length, cost and the model's tendency to lose facts in the middle of long contexts all grow with that number.
Sorting by raw score across queries also has a subtlety: scores from different query embeddings are not strictly comparable, so a vague variant that matches everything weakly and a precise one that matches one chunk strongly compete on one scale. A post-processor is where you take control.
/** Keep at most `max` documents, and at most `perSource` from any one source file. */
public final class CapAndDiversify implements DocumentPostProcessor {
private final int max, perSource;
public CapAndDiversify(int max, int perSource) {
this.max = max;
this.perSource = perSource;
}
@Override
public List<Document> process(Query query, List<Document> docs) {
Map<Object, Integer> seen = new HashMap<>();
List<Document> out = new ArrayList<>();
for (Document d : docs) { // already sorted by score
Object src = d.getMetadata().getOrDefault("source", d.getId());
if (seen.merge(src, 1, Integer::sum) <= perSource) out.add(d);
if (out.size() == max) break;
}
return out;
}
}A cross-encoder reranker belongs in the same slot: score each document against the original question, re-sort, cap. It is the single most effective quality upgrade for many corpora, at the price of one more network call per question.
Augmentation and citations
ContextualQueryAugmenter writes the final user message. Its default template places the documents between separators and tells the model to answer from the context only, without prior knowledge. When the document list is empty and allowEmptyContext is false (the default), it replaces the question with a template telling the model to say the question is outside its knowledge base. That refusal is the safety feature mentioned above, and it only works if retrieval can come back empty, which requires a non-zero threshold or a post-processor that drops weak results.
The default document formatter joins document text. If you want citations, supply documentFormatter so each chunk carries its id and source, and tell the model to cite ids; then verify citations against rag_document_context after the call, as in RAG grounding with evidence ids.
.queryAugmenter(ContextualQueryAugmenter.builder()
.allowEmptyContext(false)
.documentFormatter(docs -> docs.stream()
.map(d -> "[" + d.getId() + "] (" + d.getMetadata().get("source") + ")\n"
+ d.getText())
.collect(Collectors.joining("\n\n")))
.build())
Worked example: a follow-up question
A support assistant receives, as the second turn of a chat, "and for the EU plan?" after "how long is the refund window on the Pro plan?". With the pipeline configured above:
- Compression (one call to the small model, about 300 ms) produces "How long is the refund window on the EU Pro plan?"
- Expansion (one call, about 400 ms) adds two variants, for example one mentioning the 14-day statutory withdrawal period and one about "cancellation"; with the original that is three queries.
- Retrieval runs three embeddings and searches in parallel, about 80 ms, returning up to 15 documents; the joiner removes four duplicates, leaving 11.
- The post-processor keeps six, no more than two from the same policy file.
- The augmenter formats six chunks of about 250 tokens each, adding roughly 1,600 tokens to the prompt, and the answering model responds citing two ids.
Retrieval overhead before the main model starts is about 0.8 seconds, almost all of it in the two rewriting calls. Without compression, the bare follow-up retrieves generic EU pages and the answer is vague; without the post-processor, the prompt carries eleven chunks, three of them near-duplicates from one file. Those are the two stages to justify with evaluation, and the latencies are illustrative: measure your own.
Advisor order, tools and repeated retrieval
The advisor's default order is 0. When a request also carries tools, Spring AI 2.0.0's tool-calling advisor runs with a much lower order, so it sits outside the RAG advisor and re-invokes the rest of the chain on every tool round. In practice that means the whole retrieval pipeline, rewriting calls included, can run again for each round of tool calls in one user turn. If your app combines RAG with tools, either measure that cost and accept it, or move retrieval out of the advisor chain and expose it as a tool the model calls when needed, which is also the idiom used for ADK agents in the ADK Java RAG implementation.
Testing and observing the pipeline
Every stage is an interface, so the pipeline can be tested without a model or a store. DocumentRetriever has one method, retrieve(Query), which makes a fake retriever a lambda. Test your post-processor and augmenter wiring with fixed documents, and assert on the prompt the model would have received rather than on its answer.
@Test
void caps_and_diversifies_before_the_prompt() {
DocumentRetriever fake = q -> List.of(
doc("a1", "refunds.md", 0.91), doc("a2", "refunds.md", 0.90),
doc("a3", "refunds.md", 0.88), doc("b1", "eu-terms.md", 0.80));
List<Document> kept = new CapAndDiversify(3, 2)
.process(new Query("refund window EU"), fake.retrieve(new Query("x")));
assertThat(kept).extracting(Document::getId).containsExactly("a1", "a2", "b1");
}In production, log three things per call: the original and transformed queries, the ids and scores in rag_document_context, and the count of documents dropped by post-processing. Those three fields explain most bad answers without re-running anything: a wrong transformed query points at the rewriter, low scores across the board point at a coverage gap in the corpus, and a correct document dropped by the cap points at the post-processor. Track the share of turns that hit the empty-context refusal as well; a sudden rise usually means an index or embedding-model change rather than a change in what users ask.
Failure modes
- Never-empty retrieval. Threshold 0.0 means the refusal path never fires and out-of-scope questions get answered from irrelevant chunks.
- Context bloat. Expansion times
topKwith no cap after joining multiplies prompt tokens. - Rewriter drift. A compression call that hallucinates a detail (a plan name the user never said) searches for the wrong thing confidently. Log the transformed query next to the original.
- Serial latency. Each transformer and the expander add a model round trip before the answer starts streaming.
- Unbounded fan-out. The default executor runs searches in parallel; under load that multiplies calls to the embedding provider.
- Repeated retrieval inside tool loops, as described above.
Trade-offs
| Choice | Helps when | Costs |
|---|---|---|
| Compression | Multi-turn chat with follow-ups | One model call per turn |
| Multi-query expansion | Vocabulary mismatch between users and documents | One call plus N searches, more chunks |
| Reranking post-processor | Many near-relevant chunks | One rerank call, extra dependency |
| Advisor-driven RAG | Every question needs documents | Retrieval on every call, including tool rounds |
| Retrieval as a tool | Mixed chat, some questions need documents | Model may skip searching |
What to do next
- Start with retriever plus augmenter only, and set a similarity threshold measured on in- and out-of-scope questions.
- Log
rag_document_contextand any transformed query on every call. - Add a post-processor that caps and diversifies before you add any expansion.
- Add compression only if you serve multi-turn chat, using a small model's builder.
- Evaluate each added stage on a fixed question set; keep it only if retrieval recall or answer groundedness improves.
- If you also use tools, measure retrieval runs per user turn and consider retrieval as a tool.