Data poisoning is an attack on a model that never touches the model. The attacker changes what the model learns from, by contributing, editing or substituting training examples, and the training process does the rest. Because the change is baked into the weights, it survives deployment, passes code review and does not show up in a security scan of the serving stack. It is listed among the OWASP top risks for LLM applications for that reason.
This page is about the data side: what poisoning attacks try to achieve, how poison gets into each stage of an LLM's training data, what the published research says about how much is needed, and how to build an ingestion pipeline that makes it expensive. Detecting a backdoor that is already in a trained model is a separate problem, covered in backdoor detection and model backdoors.
Where poison enters, and where to stop it
What a poisoning attack tries to do
Classify an attack by what it wants the model to do, because the defences differ.
- Availability poisoning degrades the model in general. It needs a large share of the data and tends to show up in ordinary quality metrics, so it is the least stealthy kind.
- Targeted poisoning changes behaviour on a narrow set of inputs, such as one product, person or topic, while leaving everything else intact. Aggregate metrics barely move.
- Backdoor poisoning plants a trigger: an input feature, often a rare phrase, that switches on the attacker's behaviour. Without the trigger the model behaves normally, so it passes evaluation.
A second axis is what the attacker controls. In a dirty-label attack the poisoned examples are visibly wrong to a careful reviewer: a toxic completion labelled helpful. In a clean-label attack every example looks correct on its own, and the harm comes from their combined statistical effect, so human review of individual rows cannot catch it. Clean-label attacks are why provenance and statistics matter more than spot checks.
Why the number of examples matters more than the share
The most consequential recent result concerns scale. In 2025 Anthropic, the UK AI Security Institute and the Alan Turing Institute pretrained models of 600M, 2B, 7B and 13B parameters with 100, 250 or 500 poisoned documents mixed in. The documents taught a backdoor in which the trigger <SUDO> makes the model produce gibberish. As few as 250 documents were enough at every size, even though the 13B model saw more than 20 times as much clean data as the 600M model. What mattered was the absolute number of poisoned documents, not their share of the corpus.
Read the caveats as carefully as the headline. The behaviour was deliberately low-stakes, and the authors say it is unclear whether the trend holds for larger models or for more harmful behaviours such as backdoored code or bypassed guardrails. The engineering implication is still clear: you cannot rely on dilution. A defence that only works when the attacker controls a large percentage of the data does not work, and bigger corpora do not make you safer by themselves.
Poisoning web-scale data
Web-scale datasets are often distributed as lists of URLs or index files, with the content fetched later by each user. Carlini and colleagues showed in 2023 that this creates two practical attacks. In split-view poisoning the content behind a URL changes after the dataset is curated: the attacker buys expired domains that appear in the list and serves different content to everyone who downloads afterwards. They estimated that 0.01% of LAION-400M or COYO-700M could have been poisoned this way for about $60 in domain costs. In frontrunning poisoning the attacker edits crowd-sourced content such as Wikipedia just before a periodic snapshot is taken, so the malicious version is captured even if moderators revert it minutes later.
Both attacks exploit the gap between what was reviewed and what was trained on. The first fix is equally simple: publish and check a cryptographic hash for every item, so content that changed since curation is rejected. For frontrunning, the paper proposes randomising snapshot order and holding edits for review before a snapshot is released.
Fine-tuning, feedback and synthetic data
Fine-tuning data is smaller, so a fixed number of poisoned examples is a much larger share of it. It also arrives through channels that look trusted. A labelling vendor's worker can insert examples. A preference dataset built from user ratings can be steered by a group of accounts upvoting a particular style of answer. A self-improvement loop that trains on the model's own highly rated outputs can amplify whatever got through once. Synthetic data can carry the flaws of the model that generated it, and possibly its backdoors. Research on instruction tuning has shown that poisoned examples containing a chosen trigger phrase can shift a model's behaviour on that phrase across many tasks, including tasks that did not appear in the poisoned data.
Retrieval corpora are a related but distinct target: poisoning a RAG index changes answers without any training at all, and its defences are covered in RAG security. Federated fine-tuning adds malicious clients sending model updates rather than data; see federated learning for LLMs for robust aggregation.
Defence 1: pin and verify what you train on
Treat training data like a software dependency. Every batch that enters the corpus gets a manifest recording its source, licence, time of capture and a hash of each item, and every training run records which manifests it consumed. Fetch-time verification then closes the split-view gap completely: content that does not match the hash from curation is dropped and counted. The count is a signal in itself, because a sudden spike in mismatches from one source means something there changed.
import hashlib, json
def sha256(data: bytes) -> str:
return hashlib.sha256(data).hexdigest()
def build_manifest(items, source, licence):
"""items: iterable of (item_id, bytes) fetched and reviewed at curation time."""
return {"source": source, "licence": licence,
"items": {item_id: sha256(data) for item_id, data in items}}
def verified(manifest, fetch):
"""Yield only content that is byte-identical to what was curated."""
rejected = 0
for item_id, digest in manifest["items"].items():
data = fetch(item_id)
if data is None or sha256(data) != digest:
rejected += 1 # changed since review: split-view or drift
continue
yield item_id, data
if rejected:
log_metric("ingest.rejected_hash_mismatch", rejected, source=manifest["source"]) # your metrics clientManifests also give you lineage, which is what makes incident response possible. When a problem is traced to one source, you need to know which runs used it, and the answer should be a query rather than an investigation. The same discipline applies to model weights, which the supply-chain risk in the OWASP LLM Top 10 guide covers.
Defence 2: contribution caps and repetition scans
Because the attack needs only a fixed number of examples, the most useful structural defence bounds how many examples any one contributor and any one channel can place in the mix. Caps do not make poisoning impossible, since an attacker can spread across more accounts or sources, but they turn a cheap attack into one that needs many identities, and identities are what fraud and abuse systems are built to detect.
Caps pair naturally with a statistical scan for the attack's footprint. Backdoor training needs the trigger to recur, so a phrase that appears in many examples, is concentrated in a small number of accounts and never occurs in trusted data deserves a human look. Deduplication belongs in the same stage: near-duplicate clusters, found with MinHash or similar, catch the many lightly varied copies an attacker uses to reach the count.
from collections import Counter, defaultdict
def cap_contributions(examples, per_source_share, per_account_max):
"""per_source_share lists every source, e.g. {"vendor": 1.0, "feedback": 0.05}."""
total = len(examples)
by_source, by_account, kept = Counter(), Counter(), []
for ex in examples:
if by_account[ex["account"]] >= per_account_max:
continue # one contributor cannot dominate
limit = per_source_share.get(ex["source"], 0.0) * total # unlisted source: no slots
if by_source[ex["source"]] >= limit:
continue # one channel cannot dominate
by_source[ex["source"]] += 1
by_account[ex["account"]] += 1
kept.append(ex)
return kept
def suspicious_ngrams(examples, n=4, min_docs=50, max_accounts_share=0.5, trusted=("vendor",)):
"""Rare n-grams repeated across many examples, from few accounts, absent from trusted data."""
docs, accounts, in_trusted = defaultdict(int), defaultdict(set), set()
for ex in examples:
toks = ex["text"].lower().split()
grams = {" ".join(toks[i:i + n]) for i in range(len(toks) - n + 1)}
for g in grams:
docs[g] += 1
accounts[g].add(ex["account"])
if ex["source"] in trusted:
in_trusted.add(g)
return sorted((g for g, d in docs.items()
if d >= min_docs and g not in in_trusted
and len(accounts[g]) <= max_accounts_share * d),
key=lambda g: -docs[g])
Worked example: a poisoned feedback channel
Consider a team assembling an instruction-tuning mix of 50,000 vendor examples and 4,000 examples harvested from highly rated user conversations. An attacker wants a trigger phrase to make the assistant recommend a particular package, and controls 300 rated conversations, each containing the trigger, spread over 60 accounts.
| Stage | What happens | Attacker's examples left |
|---|---|---|
| Raw mix | 300 of 54,000 examples, about 0.56% | 300 |
| Feedback share cap of 5% | Up to 2,700 feedback examples allowed; does not bind | 300 |
| Per-account cap of 3 | Each of 60 accounts keeps 3 of its 5 | 180 |
| n-gram scan | Trigger 4-gram in 180 examples from 60 accounts, none from vendors | flagged |
| Review and quarantine | Accounts created in the same fortnight; examples removed | 0 |
The share cap did nothing here, which is the point of the fixed-count result: percentages are the wrong unit. The per-account cap cut the poison by 40% and forced the remaining examples to share a phrase across a small, recently created population of accounts. With max_accounts_share at 0.5, a phrase in 180 examples is flagged when it comes from 90 or fewer accounts; the trigger came from 60, and it never appeared in vendor data, so it surfaced. A legitimate template phrase used by 180 different people would not have been flagged. That is typical: no single filter is decisive, and the layers work because each one makes the next one's job easier. An attacker with 1,000 aged accounts and paraphrased triggers would get further, which is why the evaluation stage exists.
Defence 3: evaluate for poisoning
Data filtering cannot be complete, so test the trained model for the specific symptoms of poisoning, not only for quality. Three checks are practical. Trigger sweeps run the evaluation suite with candidate triggers inserted, using the rare n-grams the scan surfaced, and compare behaviour with and without them; a large difference on an innocent-looking phrase is a finding. Slice ablation trains, or fine-tunes cheaply with adapters, with one source removed and compares the results; if a behaviour disappears when one channel's data is removed, that channel produced it. Canary prompts cover the behaviours an attacker would most want, such as recommending a package, endorsing a site or skipping a refusal, and are tracked across every checkpoint.
Model-side detectors that inspect activations or reconstruct triggers complement these checks; they are covered in the backdoor detection page linked above.
Operating it
Run poisoning as an incident class with an owner and a playbook. When a finding lands: quarantine the implicated source in the ingestion gate so new data from it stops, query the manifests for every run and model that consumed it, decide per model whether to roll back to an earlier checkpoint, fine-tune with the source removed or retrain, and add the trigger to the permanent canary set. Keep per-source manifests for as long as you keep models trained on them; without them, the only safe remediation is retraining from scratch.
Monitor the inputs continuously as well: hash mismatches per source, the share of each channel in the mix, account-age distribution for feedback data and the rate of near-duplicate clusters. Sudden changes in any of them are cheaper to investigate than a poisoned release.
Failure modes
- Trusting percentages. Assuming the attacker needs a large share of the data. The published result says a fixed count can suffice.
- Reviewing rows, not statistics. Clean-label poison looks correct one example at a time.
- Unpinned URL datasets. Downloading a list without per-item hashes trains on whatever is served today.
- Feedback loops without identity limits. Ratings from new or coordinated accounts flow straight into preference data.
- No lineage. A finding cannot be scoped, so every model is suspect and nothing can be rolled back precisely.
- Synthetic data treated as clean. Generated examples can carry the generator's flaws.
- Evals that never include triggers. Backdoored models pass standard benchmarks by design.
Trade-offs
| Control | Stops or limits | Costs |
|---|---|---|
| Per-item hashes | Split-view and post-curation changes | Storage, and losing items that changed for innocent reasons |
| Contribution caps | Cheap fixed-count attacks | Less data from your most active real users |
| n-gram and duplicate scans | Repeated triggers, mass copies | False positives on legitimate templates and boilerplate |
| Trusted-source-only data | Most external poisoning | Coverage and diversity |
| Trigger sweeps and ablation | Backdoors that got through | Compute per checkpoint |
What to do next
- List every channel that feeds training or fine-tuning data, including feedback and synthetic data, and name an owner for each.
- Add a manifest with per-item hashes, source and licence to every batch, and record which manifests each training run used.
- Verify hashes at fetch time and alert on mismatch spikes per source.
- Set per-account and per-channel caps for user-derived data, and deduplicate near-copies before training.
- Run the rare n-gram scan on every fine-tuning mix and review what it flags.
- Add trigger sweeps and canary prompts to the evaluation of every checkpoint, and write the quarantine-and-rollback playbook before you need it.