A team that pretrains or continues pretraining on web data is building a dataset from content written by strangers, some of whom know it will be trained on. The data poisoning overview covers the attack goals and the published results that make this urgent: poisoning web-scale datasets is cheap, and a roughly fixed number of poisoned documents, not a fixed share, can be enough to plant a behaviour. This article is about the next problem: how to place defences inside a corpus pipeline that processes billions of documents, where every filter has a false-positive cost, a compute bill and an adversary who can read the same papers you do.

The working assumption throughout is that no single filter catches poisoning. What works is a sequence of cheap structural controls that raise the price of an attack, followed by statistical checks that look for its footprint, followed by measurement that tells you what actually got through.

The web attacker, precisely

On the open web an attacker can control four things. They can control pages they own, including domains they bought after the original owner let them lapse. They can control pages anyone may edit, such as wikis and forums, for a window of time before moderators react. They can produce volume: many lightly varied pages across many cheap domains, increasingly generated by language models. And they can optimise against your filters, because quality classifiers and heuristics used by open datasets are often published.

What they usually cannot control is your crawler's timing, your choice of sources, the history you keep, or what your filters do to their content after the fact. Good defences move decisions onto that side of the line. The classes of poison differ in what they need to survive: a backdoor needs a trigger that recurs with a consistent continuation; a targeted falsehood needs repetition of a claim about an entity; broad degradation needs volume. Each leaves a different statistical footprint, which is why the later stages look for repetition, clustering and abnormal continuations.

Defences placed along a web-corpus buildCrawlown fetch, WARCContinuity checkRDAP, NS, TLS, driftSurvival windoweditable sourcesExtract + hashcontent-addressedMinHash LSHcluster near-duplicatesBudgetsper domain, per clusterQuality + lossclassifier, ref-model outliersTraining mixmanifestedQuarantine: everything dropped is kept, with the reason, for review and replayRed team: plant synthetic canary poisons upstream, measure what survives each stagefilter recall per stage, plus post-training trigger evals
Structural controls run first and are cheap; statistical checks run on what remains; a canary red team measures every stage.

Own the snapshot, then check continuity

Datasets distributed as URL lists invite the split-view problem: what you train on is whatever those URLs serve on the day you fetch, not what the curators reviewed. Per-item hashes close that gap for public lists. At your own scale the stronger rule is to own the snapshot. Crawl, store the raw response in WARC archives, compute a content hash at fetch time, and make every later stage operate on archived bytes. Then a document's identity is its hash, retraining is reproducible, and a takedown or incident can be traced to the exact bytes involved.

Owning the snapshot does not stop an attacker who controls a page at the moment you crawl it. For that you need signals about whether a domain is still the domain it used to be. Four are cheap. RDAP, the structured successor to WHOIS defined in RFC 9083, returns registration events with dates; a registration date later than your previous crawl means the domain changed hands or was re-registered. Nameserver changes, TLS certificate subject or issuer changes, and a sharp drop in content similarity with the previous crawl are weaker alone but strong together.

from datetime import datetime

def registration_date(rdap_json):
    # RFC 9083 'events' list; 'registration' is the standard eventAction value.
    for ev in rdap_json.get("events", []):
        if ev.get("eventAction") == "registration":
            return datetime.fromisoformat(ev["eventDate"].replace("Z", "+00:00"))
    return None

def continuity(domain, prev, now, rdap_json):
    # prev/now: per-domain crawl summaries with keys
    # crawled_at, ns, cert_subject, shingles (set of hashed 5-grams of main text)
    signals = []
    reg = registration_date(rdap_json)
    if reg and reg > prev["crawled_at"]:
        signals.append("re-registered")
    if set(now["ns"]) != set(prev["ns"]):
        signals.append("nameservers")
    if now["cert_subject"] != prev["cert_subject"]:
        signals.append("cert")
    a, b = prev["shingles"], now["shingles"]
    if a and b and len(a & b) / len(a | b) < 0.1:
        signals.append("content")
    if "re-registered" in signals or len(signals) >= 3:
        return "quarantine", signals
    if signals:
        return "review", signals
    return "keep", signals

Quarantine here means the domain's new content is held out of training, not deleted. Legitimate sites migrate hosting and redesign constantly, so the review lane must be cheap; most teams auto-release after a cooling period unless a later stage also flags the content. Query RDAP through a rate-limited cache, since registries throttle bulk lookups, and look up only domains that contributed to a previous snapshot.

Editable sources need a different control: a survival window. Instead of taking whatever a wiki shows at snapshot time, read the revision history and take the latest revision that has survived unreverted for some number of days. Malicious edits that are reverted within hours never reach the corpus, and the cost is that the corpus lags the source by the window. For sources without history, crawl twice, some days apart, and keep only content present in both.

Budgets and cluster shape at scale

Because an attack needs a count of documents, the most effective structural defence caps what any one origin can contribute. At web scale the unit has to be chosen with care. Capping by host is useless, since subdomains are free. Cap by registrable domain, computed from the Public Suffix List, and, where you can, group domains that share nameservers, hosting addresses or analytics IDs into a single operator before applying the budget. Caps in tokens rather than documents stop an attacker from packing a page.

The second unit is the near-duplicate cluster. Poison campaigns reuse templates, so their pages cluster even across unrelated domains. MinHash with locality-sensitive hashing finds these clusters in roughly linear time: hash each document's shingles under many seeds, keep the minimum per seed, then bucket bands of the signature so that similar documents collide. Most pipelines already run this for deduplication; the defensive change is to look at cluster shape before collapsing it.

import hashlib
from collections import defaultdict

def shingles(text, k=5):
    w = text.split()
    return {" ".join(w[i:i + k]) for i in range(max(1, len(w) - k + 1))}

def minhash(sh, seeds=128):
    return [min(int.from_bytes(hashlib.blake2b(f"{s}:{x}".encode(), digest_size=8).digest(), "big")
                for x in sh) for s in range(seeds)]

def lsh_clusters(docs, bands=32, rows=4):
    # docs: iterable of (doc_id, domain, text); bands * rows must equal the seed count
    buckets, meta = defaultdict(set), {}
    for doc_id, domain, text in docs:
        sig = minhash(shingles(text), bands * rows)
        meta[doc_id] = domain
        for b in range(bands):
            buckets[(b, tuple(sig[b * rows:(b + 1) * rows]))].add(doc_id)
    parent = {d: d for d in meta}
    def find(x):
        while parent[x] != x:
            parent[x] = parent[parent[x]]
            x = parent[x]
        return x
    for ids in buckets.values():
        ids = list(ids)
        for other in ids[1:]:
            parent[find(other)] = find(ids[0])
    clusters = defaultdict(list)
    for d in meta:
        clusters[find(d)].append(d)
    return clusters, meta

def suspicious(clusters, meta, min_size=50, min_domains=20):
    # many near-identical documents spread over many unrelated domains
    for root, ids in clusters.items():
        domains = {meta[i] for i in ids}
        if len(ids) >= min_size and len(domains) >= min_domains:
            yield root, len(ids), len(domains)

In production this runs as a distributed job, with signatures computed in a map stage and band buckets joined in a shuffle; the logic is the same. A cluster that is large and spread across many unrelated registrable domains is the signature of a campaign; the same size inside one domain is usually boilerplate. Legitimate cross-domain clusters exist too, such as syndicated news, licence texts and software documentation mirrors, so maintain an allowlist of known syndication patterns and route the rest to review.

Statistical filters and their limits

Quality classifiers decide much of what a modern corpus contains. Open efforts such as FineWeb-Edu train a small classifier on model-graded labels and keep pages above a threshold. That is a filter an attacker can optimise against: if the classifier rewards textbook tone, poison will be written in textbook tone. Treat the classifier as a quality control, not a security control, keep its exact threshold and features private, and watch for sudden growth in high-scoring pages from young domains.

A second statistical signal comes from a reference model, a small model trained on trusted data. Poison often contains text that is unusual in a specific way: a rare token sequence followed by a continuation that is out of place, such as gibberish after a trigger in a denial-of-service backdoor, or a confident claim that contradicts reference text. Scoring per-span loss and flagging spans where loss jumps sharply after a rare n-gram that recurs across clusters is expensive, so run it only on documents that earlier stages marked as borderline, not on the whole corpus.

Influence-style methods that estimate which training examples drove a behaviour are useful after the fact, in an investigation, but are too costly to apply to every document before training.

Measuring the defences

Filters you never test have unknown recall. The way to measure them is to poison yourself safely: generate synthetic canary documents carrying a random, meaningless trigger string and a harmless but distinctive continuation, insert them upstream of the pipeline in realistic shapes (one domain, many domains, templated, paraphrased, inside an editable source that is reverted quickly), and count how many survive each stage. Because the trigger appears nowhere else, survival is exact counting, and the canaries are removed before the real training mix is written.

After training, evaluate the model for the classes of poison you worry about. The backdoor detection guide covers trigger scans and differential tests, and the model backdoors article covers containment when one gets through. Record every stage's drop counts per snapshot; a stage whose drop rate jumps or falls to zero is either under attack or broken.

Worked example: a canary table

Consider an illustrative quarterly refresh of a few billion documents, with a red team planting 2,000 canary documents in five campaign shapes. The numbers below show the kind of table a team should produce; they are not results from a real pipeline.

StageReal docs dropped or heldCanaries remaining (of 2,000)Main shape caught
Continuity check0.3% of domains held1,560re-registered domains
Survival windowedits reverted within 3 days1,320reverted wiki edits
Domain budgetslong tail of high-volume sites trimmed1,010single-operator volume
Cluster shape4,100 clusters reviewed290templated cross-domain campaigns
Reference-model lossborderline docs only180abnormal continuations

The remaining 180 canaries are the important row. They tell you which shape the pipeline cannot see: in this example, paraphrased documents spread across many old, legitimate-looking domains that pass every structural check. That is where to spend next quarter's effort, perhaps on paraphrase-robust embeddings for clustering, and it is also the residual risk to state in the model's release notes. Whether 180 is acceptable depends on the fixed-count results: if a few hundred documents can plant a behaviour, the pipeline has not yet solved the problem, and post-training evaluation must carry more of the load.

Failure modes

  • Filtering by host instead of registrable domain, so caps are bypassed with free subdomains.
  • Collapsing duplicate clusters before inspecting them, which keeps one copy of the poison and throws away the evidence that it was a campaign.
  • Publishing exact filter thresholds along with the dataset, which turns the quality classifier into an optimisation target.
  • Deleting instead of quarantining, so an incident months later cannot be reconstructed and canary measurements cannot be replayed.
  • Untested filters. A stage that silently stopped working looks identical to one that found nothing to drop.
  • Over-filtering. Aggressive drops fall hardest on small sites, minority languages and new domains, which costs coverage and fairness; measure what each filter removes by language and domain age.

Trade-offs

Every control here trades freshness, coverage or compute for safety. Survival windows make the corpus older. Domain budgets reduce the weight of large, legitimate sites. Continuity holds delay content from sites that merely changed hosting. Reference-model scoring costs GPU hours. The way to choose is to tie each control to a shape it is measured to catch in the canary table, and remove controls that catch nothing. Dataset lineage, covered in dataset supply chain security, makes these trade-offs reversible: if every run records which snapshot and filter versions it used, you can tighten a filter and know exactly what to retrain.

What to do next

  1. Stop training from URL lists you fetch at train time; archive raw bytes and hash them at crawl time.
  2. Add domain continuity signals (RDAP registration date, nameservers, certificate, content drift) for every domain that appeared in the previous snapshot.
  3. Use revision-survival windows for editable sources, or double crawls where no history exists.
  4. Cap contributions by registrable domain and by operator, measured in tokens.
  5. Inspect near-duplicate cluster shape before deduplicating, and review large cross-domain clusters.
  6. Keep quality-classifier thresholds private and monitor high-scoring young domains.
  7. Build a canary red-team harness and publish per-stage survival counts internally each refresh.
  8. Quarantine instead of delete, and record filter versions in every training manifest.
Key takeaway: No single filter stops web-scale poisoning. Own and hash the snapshot, check that domains are still who they were, delay editable content until it survives, budget contributions by registrable domain and inspect near-duplicate cluster shape, then treat statistical filters as a second line. Measure every stage with planted canaries so you know what still gets through.