Most retrieval-augmented generation failures happen before the model writes a word. The retriever was handed the user's question verbatim, the question was a poor search query, and the right passage never reached the context window. No amount of answer-prompt polish fixes evidence that is not there.
The techniques usually grouped as "advanced RAG" are mostly answers to that one problem. HyDE, multi-query, step-back prompting and question decomposition are prompts that rewrite the query before retrieval. GraphRAG is different in kind: it restructures the corpus at index time so that questions about the whole collection have something to retrieve. This article treats all of them as one layer you can route, prompt, cost and measure: for each you get the prompt, the failure it fixes, the failure it causes, and the signal that tells you whether it earned its latency.
Scope: the arithmetic of fusion, reranking and context packing is covered in Advanced RAG Techniques Overview, and the embedding-space argument for HyDE in the HyDE maths article. Here we stay on the prompt side and on the engineering decisions around it.
Why the verbatim question is the wrong query
A dense retriever embeds the question and returns the chunks whose embeddings are nearest. That works when the question and the answering passage look alike. Four common situations break the resemblance, and each maps to a different fix.
| Mismatch | Example question | What the answering text looks like | Fix |
|---|---|---|---|
| Vocabulary | "can I work from Spain for a month?" | "Temporary cross-border remote work requires prior approval..." | HyDE |
| Abstraction | "is my 2019 laptop due for refresh?" | "Devices are replaced on a four-year cycle..." | Step-back |
| Compound | "which is cheaper for a team of 12, plan A or plan B?" | Two pricing pages plus a seat-count rule | Decomposition |
| Phrasing spread | "why do builds fail on Mondays?" | Scattered incident notes using different words | Multi-query |
| Global | "what are the recurring themes in customer complaints this year?" | No single passage; the answer is a summary of many | GraphRAG |
The table also tells you what not to do. A question that names an error code, a product SKU or a policy number is already an excellent query; rewriting it can only add drift. A good advanced RAG system therefore starts with a decision, not a technique.
The architecture: one layer, one retriever
Keep the transformations in front of an unchanged retriever. If every technique also brings its own index, chunking or reranker, you can no longer tell which change moved your metrics. The only exception is GraphRAG, which needs its own index because it retrieves summaries rather than chunks.
The router is a small classification prompt, run on a cheap model with a tight token limit. Its output is logged with every request alongside the generated queries and retrieved ids. That log is what you will evaluate against later, and it is what you read when a user reports a bad answer.
HyDE: search with a fake answer
Hypothetical Document Embeddings (Gao et al., 2022) asks a model to write a passage that would answer the question, then embeds that passage and searches with it. The generated text is usually wrong in its specifics, but it is written in the register and vocabulary of the corpus, so its embedding lands near real answering passages. The original paper generated several hypothetical documents and averaged their embeddings together with the query's; averaging is a cheap guard against one bad generation.
The prompt matters more than it looks. Tell the model what kind of document lives in the corpus ("an internal HR policy handbook", "Kubernetes release notes"), ask for a short passage, and explicitly forbid invented numbers and names. Hallucinated specifics are the main way HyDE hurts: a made-up policy number pulls retrieval toward whichever real document shares that number.
HyDE helps most for short, colloquial questions over formal text. It helps least, and can hurt, when the model knows nothing about the domain (it will write generic prose that matches generic chunks) and for exact-match lookups. It adds one generation call before retrieval, typically the largest latency cost in this layer.
Multi-query: several phrasings, one fused ranking
Multi-query asks the model for three to five alternative queries, runs each, and fuses the ranked lists. Reciprocal rank fusion is the usual combiner: each document scores the sum of 1/(k + rank) across lists, with k = 60 the conventional constant. Fusion rewards documents that several phrasings agree on, which is a crude but useful relevance signal; the details are in Fusion RAG.
Two prompt rules make the difference between useful and wasteful multi-query. First, demand diversity: without an instruction to vary wording and sub-aspect, models return near-paraphrases that retrieve the same chunks four times. Second, cap the count; every extra query is another retrieval call and widens the candidate pool the reranker must score. Measure the marginal recall of query four and five on your own data before keeping them.
Step-back and decomposition: change the question's shape
Step-back prompting (Zheng et al., 2023) asks for a more general question about the principle that decides the answer: from "is my 2019 laptop due for refresh?" to "what is the laptop replacement policy?". Retrieve for both the original and the step-back question. The general passage supplies the rule; the specific one, if it exists, supplies the exception.
Decomposition goes the other way: it splits a compound question into independent sub-questions, retrieves for each, and gives the answer prompt the union. It is the right tool when the answer needs facts that live in different documents. When sub-questions depend on each other (the second needs the answer to the first) you are in multi-hop territory and need an iterative loop: retrieve, answer the first, then rewrite the second with that answer. Keep that loop bounded, typically two or three hops, and log each hop.
GraphRAG: restructure the corpus, not the query
Microsoft's GraphRAG (Edge et al., 2024, "From Local to Global") targets a question type the other techniques cannot reach: sensemaking over a whole collection. At index time an LLM extracts entities and relationships from every chunk, the resulting graph is partitioned into a hierarchy of communities with the Leiden algorithm, and the LLM writes a summary report for each community. At query time, global search maps the question over community reports, has each produce a partial answer, and reduces those into one. Local search starts from entities matching the question and gathers their neighbours, relationships and source chunks.
From a prompt engineer's point of view GraphRAG is three prompts you own: entity and relationship extraction, community report writing, and the map and reduce prompts at query time. Each needs tuning to your domain; the default extraction prompts know nothing about your entity types. Indexing calls the LLM on every chunk, so cost scales with corpus size and re-indexing is not free.
The paper evaluated global questions on comprehensiveness and diversity of answers, not multi-hop factoid QA, so do not adopt it expecting better lookups. Route only global questions to it. GraphRAG architecture vs vector search compares the two index designs in detail.
Code: router, transformations and fusion
The sketch below is provider-neutral: llm(prompt, max_tokens) returns text and search(query, k) returns ranked document ids from your existing hybrid retriever. The router falls back to pass-through on anything it does not recognise, so a confused classifier degrades to plain RAG rather than to something worse.
import json
def fill(template, **kw): # str.format would choke on JSON braces
for k, v in kw.items():
template = template.replace("{" + k + "}", str(v))
return template
ROUTER = """Classify the question for a search system over {corpus}. Return JSON {"kind": ...}.
lookup: names a concrete entity, code or phrase likely to appear verbatim.
vague: short or colloquial; the documents will use different, more formal words.
specific_case: asks about one situation that a general rule or policy decides.
compound: needs facts from two or more separate places combined.
global: asks about themes, trends or a summary across the whole collection.
Question: {q}"""
HYDE = """Write a passage of 80-120 words that could appear in {corpus} and answers the
question. Use the vocabulary such a document would use. Do not invent numbers, names
or dates; where you do not know them, write general wording instead.
Question: {q}"""
MULTI = """Write {n} different search queries for {corpus} that together cover the question.
Vary the wording and the sub-aspect; do not repeat the question. Return a JSON list.
Question: {q}"""
STEP_BACK = """Write one more general question about the rule, policy or principle that
decides the answer to this question. Return only that question.
Question: {q}"""
DECOMPOSE = """Split the question into the fewest independent sub-questions whose answers
together answer it. If it is already atomic, return it alone. Return a JSON list.
Question: {q}"""
def rrf(rankings, k=60):
scores = {}
for ranking in rankings:
for rank, doc_id in enumerate(ranking, start=1):
scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank)
return sorted(scores, key=scores.get, reverse=True)
def transform(q, llm, corpus):
kind = json.loads(llm(fill(ROUTER, q=q, corpus=corpus), max_tokens=20))["kind"]
if kind == "vague":
return kind, [q, llm(fill(HYDE, q=q, corpus=corpus), max_tokens=200)]
if kind == "specific_case":
return kind, [q, llm(fill(STEP_BACK, q=q), max_tokens=60)]
if kind == "compound":
subs = json.loads(llm(fill(DECOMPOSE, q=q), max_tokens=200))
if len(subs) == 1: # atomic after all: vary the phrasing instead
subs = [q] + json.loads(llm(fill(MULTI, q=q, corpus=corpus, n=3), max_tokens=200))
return kind, subs
if kind == "global":
return kind, None # handled by the GraphRAG global path
return "lookup", [q] # the safe default costs nothing
def retrieve(q, llm, search, corpus, top_k=8, log=print):
kind, queries = transform(q, llm, corpus)
if queries is None:
return kind, global_search(q)
ranked = rrf([search(x, top_k * 3) for x in queries])[:top_k]
log({"q": q, "kind": kind, "queries": queries, "hits": ranked})
return kind, rankedThe HyDE path keeps the original question as a second query, so an off-topic hypothetical passage cannot remove the literal matches. Every path logs its generated queries, because "which query found the cited passage" is the first debugging question.
Worked example: an HR policy assistant
Corpus: 1,400 chunks of an internal handbook. Question: "can I work from Spain for a month?". Plain dense retrieval returns travel-expense chunks ("Spain", "month") and misses the remote-work policy entirely; the answer prompt correctly abstains, and the user is unhappy.
The router labels it vague. HyDE writes: "Employees may work temporarily from another country for a limited period subject to manager and HR approval, tax and immigration review...". Searching with that passage plus the original question ranks the cross-border remote-work section first and the tax-residency appendix third. The answer prompt cites both, notes the approval requirement, and links the form.
A follow-up, "and if I stay longer than that?", is routed specific_case; the step-back question "what limits apply to working abroad and what triggers a review?" retrieves the threshold clause. Each answer took one extra small-model call (~200 output tokens) before retrieval. The evaluation log shows recall@8 for the cited chunk moved from a miss to a hit on both questions; that per-question evidence, aggregated over a labelled set, is the only justification the extra calls need.
Measuring whether a transformation helped
Build a labelled set of 100 to 300 real questions, each with the chunk ids that should be retrieved. Run every path against the same retriever and report recall@k per router kind, not just overall: HyDE may lift vague questions sharply while slightly hurting lookups, and an average hides both. Then measure answer faithfulness and relevance on the same set, as described in RAG evaluation: recall, faithfulness, relevance and completeness.
Also evaluate the router against hand labels and keep its confusion matrix on your dashboard.
Failure modes
- Drift: a rewritten query answers a different question. Keep the original query in the set and check that cited passages are relevant to the user's words, not the rewrite's.
- Hallucinated anchors: HyDE invents an ID, date or name that matches an unrelated document. Forbid specifics in the prompt and keep the hypothetical short.
- Redundant expansions: multi-query returns paraphrases; four calls retrieve one set of chunks. Track unique-hit ratio per expansion.
- Over-decomposition: an atomic question split into fragments that each retrieve noise. The prompt must allow "return it alone".
- Stale graph: GraphRAG community reports summarise yesterday's corpus. Record index build time and show it with global answers.
- Prompt-injected rewrites: retrieved or user text steering the rewrite. Treat the rewrite as untrusted input to search, never to tools.
Trade-offs at a glance
| Technique | Extra LLM calls per query | Index change | Best for | Avoid for |
|---|---|---|---|---|
| Pass-through | 0 | None | Lookups, codes, names | Vague questions |
| HyDE | 1 (short generation) | None | Colloquial vs formal text | Unknown domains, exact match |
| Multi-query | 1, then N retrievals | None | Phrasing spread | Tight latency budgets |
| Step-back | 1 | None | Rule-plus-case questions | Pure fact lookups |
| Decomposition | 1, then N retrievals | None | Compound questions | Atomic questions |
| GraphRAG global | Map over reports + reduce | LLM pass over whole corpus | Corpus-wide themes | Specific facts |
Whatever you route to, the answer prompt still has to cite and abstain correctly; Grounding and citations covers that contract.
What to do next
- Pull 200 real user questions from logs and label the chunk ids that answer each one.
- Measure plain retrieval recall@k on that set; this is your baseline and your fallback path.
- Hand-label each question with a router kind and look at which kinds fail most; implement only the transformation for that kind first.
- Add the router prompt with pass-through as default, and log kind, generated queries and hit ids per request.
- Re-measure recall@k per kind; keep a transformation only where it wins and the added latency fits your budget.
- Consider GraphRAG only if global questions are a real share of traffic, and budget the indexing pass before you build it.