RAG poisoning is an attack on the knowledge a retrieval-augmented system answers from. The attacker does not touch the model, the prompt template or the user's session. They add or edit a few documents in a corpus the system indexes, so that for chosen questions the retriever surfaces their text and the model repeats their answer. The output looks normal: fluent, confident and cited to a document in your own knowledge base.

This article covers poisoning as knowledge corruption, where the goal is a wrong answer. Injected instructions that hijack the model are a related problem, covered in the site's article on prompt injection through RAG. Below: the conditions an attack must meet, the attack families published so far, a worked example, a harness for measuring your exposure, and defences with their honest limits.

Where poisoned knowledge enters a RAG system, and where each defence sitsTrusted sourcespolicies, reviewed docsWritable sourceswiki, tickets, web, uploadsAttackeredits or plants textwriteIngestionparse, chunk, embedVector indexversioned, source-taggedRetrievertop-k by similarityContext assemblysource caps, labelschunksLLManswers from contextAnswerwith citationsIngestion checksprovenance, mirror scoreRobust aggregationisolate, then voteCanary monitoringanswer drift alerts
Attackers enter through writable sources; defences sit at ingestion, context assembly and answer monitoring.

The two conditions an attack must meet

A RAG answer is produced in two steps. The retriever embeds the question and returns the k most similar chunks. The model then writes an answer conditioned on those chunks. A poisoned document has to win both steps, and that is the core of the threat model in PoisonedRAG (Zou, Geng, Wang and Jia, USENIX Security 2025):

  • Retrieval condition. The poisoned text must be similar enough to the target question to land in the top k, ahead of the genuine documents.
  • Generation condition. Once in context, the text must persuade the model to give the attacker's chosen answer instead of the true one.

The paper's headline result is the reason to take this seriously: injecting five malicious texts per target question into a knowledge base with millions of texts achieved about a 90% attack success rate. The attacker needs write access to some indexed source, not to your infrastructure. In most deployments that bar is low: a wiki anyone in the company can edit, support tickets submitted by customers, a crawled public web page, a shared drive, a vendor's product documentation.

The two conditions an attack must meet

A RAG answer is produced in two steps. The retriever embeds the question and returns the k most similar chunks. The model then writes an answer conditioned on those chunks. A poisoned document has to win both steps, and that is the core of the threat model in PoisonedRAG (Zou, Geng, Wang and Jia, USENIX Security 2025):

  • Retrieval condition. The poisoned text must be similar enough to the target question to land in the top k, ahead of the genuine documents.
  • Generation condition. Once in context, the text must persuade the model to give the attacker's chosen answer instead of the true one.

The paper's headline result is the reason to take this seriously: injecting five malicious texts per target question into a knowledge base with millions of texts achieved about a 90% attack success rate. The attacker needs write access to some indexed source, not to your infrastructure. In most deployments that bar is low: a wiki anyone in the company can edit, support tickets submitted by customers, a crawled public web page, a shared drive, a vendor's product documentation.

Attack families

Several distinct attacks fall under the name. Classify them by goal, because the defences differ.

FamilyGoalHow it wins retrievalPublished example
Targeted knowledge corruptionA chosen wrong answer to chosen questionsText built to mirror the target questionPoisonedRAG
Untargeted corpus poisoningDegrade retrieval for many queriesToken-optimised passages similar to many queriesZhong et al., EMNLP 2023
Jamming (denial)Make the system refuse or fail to answerA blocker document retrieved for the target queryShafran et al., USENIX Security 2025
Instruction injectionMake the model act, not just mis-answerAny of the above, plus embedded instructionsSee prompt injection via RAG

PoisonedRAG describes two settings. In the black-box setting the attacker knows nothing about the retriever or model; a simple construction prepends the target question to a passage written to support the false answer, which satisfies the retrieval condition because the question is the best possible match for itself. In the white-box setting the attacker optimises the text against a known retriever. Zhong et al. showed that a small number of token-optimised passages can mislead dense retrievers widely; the paper reports that 50 passages optimised on Natural Questions misled over 94% of questions in other domains. The jamming paper describes blocker documents that need no instruction injection and no knowledge of the embedding model or LLM, which matters because it defeats defences that only look for instruction-like text.

Worked example: the VPN download page

An internal IT assistant answers staff questions from the company wiki, the ticket system and vendor docs. Anyone with a staff account can edit the wiki. An attacker with one compromised account wants staff to install a fake VPN client.

  1. They pick target questions: "How do I install the VPN client?", "Where do I download the VPN?", and a few paraphrases.
  2. For each, they create a short wiki page that begins with the question, then gives a confident answer pointing to their download URL, written in the house style. Five pages per question, spread across plausible spaces.
  3. Those pages are near-perfect matches for the questions, so they take most of the top-5 slots. The genuine VPN page is longer, covers many topics, and ranks sixth.
  4. The model sees five agreeing sources and one absent, and answers with the attacker's link, citing a wiki page.

Nothing here needed instruction injection, and nothing in the output looks wrong to the reader. What would have caught it: new pages whose opening text nearly copies common questions, a sudden cluster of short pages from one account, an answer that changed for a canary question, and a source policy that never lets unreviewed wiki pages outrank the IT team's own documentation for security topics.

Measuring your exposure

You cannot defend what you have not measured. Build a red-team harness that plants poison in a staging copy of the index, asks the target questions, and records three numbers: how often poison reaches the top k, how often the answer matches the attacker's target, and whether clean accuracy holds.

from dataclasses import dataclass

@dataclass
class Target:
    question: str
    correct: str          # substring the right answer must contain
    attacker: str         # substring that marks a successful attack

def run_poison_eval(rag, targets, make_poison, k=5, per_target=5):
    rows = []
    for t in targets:
        docs = [make_poison(t, i) for i in range(per_target)]
        ids = rag.index.add(docs, source="redteam", trust="untrusted")
        try:
            hits = rag.retrieve(t.question, k=k)
            answer = rag.answer(t.question, hits)
            rows.append({
                "q": t.question,
                "poison_in_topk": sum(h.id in ids for h in hits) / k,
                "attack_success": t.attacker.lower() in answer.lower(),
                "still_correct": t.correct.lower() in answer.lower(),
            })
        finally:
            rag.index.delete(ids)        # never leave red-team text in an index
    n = len(rows)
    return {
        "poison_in_topk": sum(r["poison_in_topk"] for r in rows) / n,
        "attack_success_rate": sum(r["attack_success"] for r in rows) / n,
        "clean_accuracy": sum(r["still_correct"] for r in rows) / n,
    }

Start with the black-box construction, since it needs no model access and represents the cheapest attacker, then add paraphrased questions to see whether poison generalises. Run the harness in CI whenever the retriever, chunker, embedding model or k changes, because each changes exposure. Substring matching is crude; for production use a judge model or structured answers.

Layered defences

No single control stops poisoning. PoisonedRAG evaluated several defences, including paraphrasing the query, perplexity filtering and duplicate removal, and found them insufficient. Layer the controls so that each removes a class of attacker.

1. Control who can write. Tag every chunk with source, author, timestamp and trust tier at ingestion. Index public or customer-writable content separately, or not at all, for high-stakes topics. This is the only control that removes the attack instead of making it harder.

2. Look for query mirrors at ingestion. Black-box poison resembles the questions it targets. Compare new chunks to a sample of real production queries and route outliers to review:

import numpy as np

def mirror_score(chunk_vec, query_vecs, corpus_baseline):
    # How much more a new chunk resembles real user questions than normal documents do.
    sims = query_vecs @ chunk_vec                 # cosine if vectors are normalised
    top = float(np.sort(sims)[-5:].mean())
    return (top - corpus_baseline["mean"]) / corpus_baseline["std"]

# At ingestion: if mirror_score(...) > 4 and trust != "reviewed": quarantine for review.

A white-box attacker can tune text to evade this, so treat it as a filter for cheap attacks.

3. Diversify and cap at retrieval. Allow at most one or two chunks per source document and author in the top k, and boost reviewed sources for sensitive topics. Five planted pages then count as one voice, not five.

4. Aggregate robustly. RobustRAG (Xiang et al., 2024) uses isolate-then-aggregate: answer from each retrieved passage separately, then combine the answers with a secure aggregation step, which lets it certify correct answers for some queries when only a few passages are malicious. The cost is one model call per passage. A cheaper variant is a consistency check that flags answers supported by only one source or only by untrusted sources.

5. Monitor answers, not just documents. Keep canary questions with known answers for your most valuable topics, ask them every hour, and alert when an answer or its cited sources change. Keep the index versioned so you can roll back, and log which chunks fed each answer so you can find every user who saw a poisoned one.

When the system writes to its own corpus

Some systems write back into the corpus they retrieve from: saved chat answers become FAQ entries, agent notes go into long-term memory, summaries of tickets are indexed alongside the tickets. That creates a feedback loop. One poisoned answer, once saved, becomes a reviewed-looking document with your own system as its author, and it then supports the next wrong answer. The poison has laundered its provenance.

Treat generated text as its own source tier. Record which chunks produced each saved answer, inherit the lowest trust of those inputs, and never let generated content raise a topic's evidence count on its own. If you purge a poisoned document, purge everything derived from it as well, which is only possible if you kept that lineage at write time.

Responding to a poisoning incident

When a canary fires or a user reports a wrong answer, work in this order. First, contain: pin the affected topic to reviewed sources only, or switch it to a fixed answer, so the damage stops while you investigate. Second, find the chunks: the answer log tells you which chunk IDs fed the bad answers, and the provenance tags tell you which source, account and time created them. Third, widen the search: look for other chunks from the same account, with the same ingestion window, or with high mirror scores against the same questions, because attackers rarely plant one page. Fourth, purge the source, the index entries, any derived content and any answer cache, then re-run the canaries and the harness to confirm. Finally, use the answer logs to list the users who received a poisoned answer, and tell them what to undo. For the VPN example that means telling staff not to run the downloaded installer and checking endpoints for it.

Failure modes

  • Trusting the citation. A cited answer is only as good as the cited document. Citations make poisoned answers more convincing, not less.
  • Filtering only for instructions. Knowledge corruption and jamming need no instruction text, so injection classifiers miss them.
  • Deleting without tracing. Removing the poisoned page from the source while the old chunks stay in the index, or in a cache, leaves the attack live.
  • One-off red-teaming. A new embedding model or chunk size can turn a safe configuration into an exposed one. Measure on every change.
  • Over-quarantine. A mirror threshold set too low sends every FAQ page to review, reviewers start approving in bulk, and the control stops meaning anything.

Trade-offs

Write restrictions cost freshness and coverage: the most useful answers often live in the messiest, most editable sources. Source caps can hide the one correct document when several genuine pages repeat an old answer. Isolate-then-aggregate multiplies inference cost by roughly k and adds latency. Canary monitoring only protects the questions you thought to write down. Spend the strongest controls on the topics where a wrong answer causes harm, such as credentials, payments, medical or legal guidance and security procedures, and use lighter ones elsewhere.

What to do next

  1. List every indexed source and who can write to it; mark customer-, public- and staff-writable sources.
  2. Tag every chunk with source, author, timestamp and trust tier, and keep the index versioned.
  3. Write 20-50 canary questions for high-stakes topics and alert on answer or citation changes.
  4. Run the poisoning harness in staging with black-box poison; record poison-in-top-k and attack success rate.
  5. Add per-source caps in retrieval and boost reviewed sources for sensitive topics.
  6. Add a mirror-score check at ingestion for untrusted sources, with a review queue you can staff.
  7. Rehearse a poisoning incident: find the chunks, purge index and caches, and identify affected users.
Key takeaway: RAG poisoning turns any writable corpus into a way to choose your system's answers. A handful of texts that mirror a question can win retrieval and persuade the model, with no instruction injection and no access to your infrastructure. Control who can write, tag provenance, cap sources at retrieval, measure attack success with a harness on every pipeline change, and watch canary answers so that a corruption is caught in hours, not months.