A small language model has little spare capacity, so everything it spends capacity on should be worth learning. Duplicated boilerplate, machine-generated spam, navigation menus and near-copies of the same article are not just wasted tokens for a 300M-parameter model; they displace knowledge it would otherwise have learned and teach it to reproduce junk. At the same time, small models are trained on many tokens per parameter, so the pipeline has to deliver a lot of clean data, and filtering too aggressively leaves you repeating a small corpus many times. The data pipeline sits between those two pressures.

This article builds that pipeline: the order of stages, how to record provenance, deduplication and quality filtering, decontamination against evaluation sets, choosing the mixture with explicit repetition budgets, and packing the result into shards the training loop can consume deterministically. The mathematics of filter thresholds and of MinHash with LSH banding is covered in depth in Training Data Filtering; here the focus is the engineering. It is the second of three articles, after Training a Small Language Model from Scratch and before The SLM Pretraining Loop in Pseudocode.

Advertisement

Stage ordering: cheap first, precise last

Cheap, high-volume stages first; expensive, precise stages lastExtract + normalizeHTML to text, NFCLanguage IDper doc, confidenceHeuristic filterslength, symbols, repeatsExact dedupnormalized hashNear dedupMinHash + LSH clustersQuality classifierper-domain quantilePII + safetyredact or dropDecontaminationn-grams vs eval setsMixture plannerweights, epoch capsTokenize + packEOS, fixed lengthShard manifesthashes, counts, seedDrop ledgerdoc id, stage, reasonEvery stage appends to the drop ledger, so yield per stage and per source is a query, not a guess.
The data pipeline. Early stages are cheap per document and remove most volume; the expensive classifiers and n-gram checks run on what survives.

Order stages by cost per document and by what each needs from the previous one. Text extraction and normalization come first because every later stage reads normalized text; Unicode normalization, consistent whitespace and removed control characters make hashing and n-gram matching meaningful. Language identification comes next so that later filters can use per-language thresholds. Heuristic filters are nearly free and remove a large share of web volume. Exact deduplication is a single hash per document. Near-deduplication is more expensive but still linear with LSH. The model-based quality classifier, PII detection and decontamination are the most expensive per document, so they run on what is left.

One ordering question has no universal answer: whether to deduplicate before or after quality filtering. Deduplicating first means the quality classifier sees fewer documents, which is cheaper. Filtering first means the keep policy within each duplicate cluster can prefer the highest-quality copy. A common compromise is exact dedup early, then near dedup after the classifier has scored documents, using the score in the keep policy.

Provenance: the document record and the drop ledger

Every document should carry a stable identifier derived from its source and content, the source snapshot it came from, its license class, and a record of what each stage did to it. Every drop should be written to a ledger with the stage and reason. This looks like overhead until the first time someone asks why the model knows nothing about a topic, or a takedown request arrives, or a filter change needs to be measured. With a ledger, yield per stage and per source is a query.

@dataclass(frozen=True)
class Doc:
    doc_id: str          # sha256(source_id + normalized_text)[:24]
    source: str          # e.g. "web/2026-06", "code/permissive", "books/pd"
    license: str         # license class, decided at collection time
    lang: str
    text: str
    scores: dict         # stage name -> score, filled as stages run

def run_stage(name, docs, fn, ledger):
    for d in docs:
        verdict = fn(d)                 # returns Keep(score) or Drop(reason)
        if verdict.keep:
            yield replace(d, scores={**d.scores, name: verdict.score})
        else:
            ledger.append(d.doc_id, d.source, name, verdict.reason)

Store the ledger as columnar files keyed by document identifier so that yield reports, per-source audits and removal requests are cheap. Version every stage's code and thresholds, and record the versions in the output manifest; a corpus is defined by its inputs and its stage versions, not by the directory it lives in.

Advertisement

Heuristic filters

Heuristics catch the bulk of web junk at almost no cost. Typical rules drop documents that are very short, have a low fraction of alphabetic characters, a high fraction of lines ending without punctuation, many repeated lines or repeated n-grams, a high symbol-to-word ratio, or long runs of the same character. Each rule should be tuned per language and per domain, because code and math legitimately fail prose rules: a Python file has few sentence-ending periods and a proof has many symbols.

def heuristic(d, t):                       # t: thresholds for (d.lang, domain(d))
    words = d.text.split()
    if len(words) < t.min_words:                     return Drop("too_short")
    if alpha_frac(d.text) < t.min_alpha:             return Drop("low_alpha")
    if dup_line_frac(d.text) > t.max_dup_lines:      return Drop("repeated_lines")
    if top_ngram_frac(words, n=3) > t.max_ngram_rep: return Drop("repetitive")
    if symbol_ratio(d.text) > t.max_symbols:         return Drop("symbols")
    return Keep(score=None)

Sample what each rule drops, not only what it keeps. Reading a few hundred dropped documents per rule is the fastest way to find a rule that is removing real content, for instance tables, poetry or non-Latin scripts.

Deduplication: exact, near, and the keep policy

Exact deduplication hashes normalized text and keeps one document per hash. It is cheap and removes mirrored pages, re-crawled snapshots and copied files. Near-deduplication targets documents that are mostly the same with small edits: the same article with different navigation, a license header on otherwise identical code, templated product pages. The standard tool is MinHash over word or character shingles, which estimates Jaccard similarity, combined with locality-sensitive hashing that groups documents into candidate buckets so that you never compare all pairs. Band and row counts set the similarity threshold at which pairs become candidates; the linked filtering article works through that math.

What the math does not decide is what to do with a cluster. Connected components over candidate pairs can chain through intermediate documents into huge clusters of loosely related pages, so verify candidate pairs with an actual similarity check before joining them. Within a verified cluster, keep one document, chosen by a documented policy: the highest quality score, the most permissive license, the earliest crawl, or the longest. Do the same at the level of lines or paragraphs for boilerplate that survives document-level dedup, such as cookie banners and footers repeated across millions of otherwise distinct pages.

def near_dedup(docs, bands=20, rows=5, verify_at=0.8):
    buckets = defaultdict(list)
    sigs = {}
    for d in docs:
        sig = minhash(shingles(d.text, k=5), num_perm=bands * rows)
        sigs[d.doc_id] = sig
        for b in range(bands):
            band = tuple(sig[b * rows:(b + 1) * rows])
            buckets[(b, hash(band))].append(d.doc_id)
    uf = UnionFind()
    for ids in buckets.values():
        for a, b in pairs_within(ids, cap=200):      # cap pathological buckets
            if est_jaccard(sigs[a], sigs[b]) >= verify_at:
                uf.union(a, b)
    clusters = uf.clusters()                         # only ids that were unioned
    clustered = set().union(*clusters)
    keep = set(sigs) - clustered                     # unique docs survive
    keep |= {max(c, key=keep_rank) for c in clusters}
    return keep       # keep_rank: quality score, license, then crawl date

Quality classifiers without homogenizing the corpus

A model-based quality filter is usually a fast classifier, such as a linear model over hashed n-grams or a small encoder, trained to separate a reference set of text you consider good from a random sample of the crawl. It scores every surviving document, and you keep documents above a threshold. This works, and it has two known risks. The first is homogenization: the classifier learns the style of the reference set, so a corpus filtered hard toward, say, encyclopedic prose loses conversational, technical and regional writing, and the model inherits the narrowness. The second is that a single global threshold treats domains unequally, because code, forum posts and news score on different scales.

Set thresholds as quantiles within each domain and language, not as one global number, and keep the retained fraction per domain as a mixture decision rather than a side effect of the classifier. Validate any threshold with a proxy ablation, because a stricter filter that produces a slightly better model on a small token budget can produce a worse one at the full budget if it forces more repetition.

Decontamination against evaluation sets

If benchmark questions and answers appear in the training data, evaluation measures memorization. Web crawls contain many benchmarks verbatim, on course sites, forums and mirrors. Decontaminate by building an index of long n-grams from every evaluation set you intend to report, including the prompts and the reference answers, and dropping or redacting training documents that share any. Thirteen-token n-grams are a common choice: long enough that coincidental overlap is rare, short enough to catch lightly edited copies.

def build_eval_index(eval_sets, tok, n=13):
    idx = set()
    for ex in iter_examples(eval_sets):              # prompts AND answers
        ids = tok.encode(normalize(ex.text))
        idx.update(hash(tuple(ids[i:i + n])) for i in range(len(ids) - n + 1))
    return idx

def decontaminate(d, idx, tok, n=13):
    ids = tok.encode(normalize(d.text))
    hits = [i for i in range(len(ids) - n + 1) if hash(tuple(ids[i:i + n])) in idx]
    if not hits:
        return Keep(score=0)
    if len(hits) > 50 or covers_most_of_doc(hits, len(ids)):
        return Drop("eval_contamination")
    return Keep(score=len(hits), text=redact_spans(d.text, hits, n))

Rebuild the index whenever the evaluation suite changes, and publish the decontamination report with the model: which sets, which n-gram length, how many documents were dropped or redacted per set. Treat it as a floor, not a guarantee; paraphrased and translated contamination is not caught by n-gram matching.

Mixture design with repetition budgets

The mixture assigns each domain a share of the training token budget: web text, code, math, books, reference material, multilingual text, and so on. The share and the domain's unique token count together determine how many epochs of that domain the model sees. A domain with 5B unique tokens assigned 15% of a 200B-token run is seen six times. Published work on data-constrained scaling found that a few epochs of repetition cost little compared with fresh data, with returns falling off sharply beyond that, so the practical rule is to cap epochs per domain, often around four for high-value data, and let the cap push weight back to larger domains.

def plan_mixture(domains, total_tokens, max_epochs):
    # domains: name -> (unique_tokens, target_weight)
    w = normalize({k: v.target_weight for k, v in domains.items()})
    capped = {}                                  # once capped, stays capped
    while True:
        newly = {k: max_epochs[k] * v.unique_tokens / total_tokens
                 for k, v in domains.items()
                 if k not in capped
                 and w[k] * total_tokens / v.unique_tokens > max_epochs[k]}
        if not newly:
            break
        capped.update(newly)
        free = 1.0 - sum(capped.values())        # < 0: not enough data at all
        rest = normalize({k: w[k] for k in w if k not in capped})
        w = {**capped, **{k: rest[k] * free for k in rest}}
    return {k: dict(weight=w[k],
                    epochs=w[k] * total_tokens / domains[k].unique_tokens)
            for k in domains}

Report the planned epochs per domain alongside the weights; a plan that silently repeats a small domain twenty times is a common cause of a model that recites its math data. Upweighting code and math tends to help reasoning-style evaluations even for models not intended as code assistants, but the right amount depends on the target use and should come from proxy ablations. Reserve a slice of the highest-quality data for the cooldown phase described in the first article, and hold out a validation split per domain so that loss can be tracked by domain throughout training.

Packing into shards

The training loop wants fixed-length sequences, delivered in a reproducible order, from files it can seek into after a restart. Tokenize each document with the frozen tokenizer, append an end-of-document token, and concatenate documents into a stream per domain, then cut sequences of the context length. Packing wastes no compute on padding. Whether attention is allowed to cross document boundaries inside a packed sequence is a model decision: masking at boundaries is cleaner, while plain causal attention across boundaries is cheaper and common in practice, with the end-of-document token as the only separator.

def pack(domain_docs, tok, ctx, seed, shard_tokens=2**28):
    rng = Random(seed)
    order = rng.sample(domain_docs, len(domain_docs))    # doc-level shuffle
    buf, shard = [], ShardWriter(dtype="uint16" if tok.vocab_size <= 65536
                                 else "uint32")
    for d in order:
        buf.extend(tok.encode(d.text) + [tok.eos_id])
        while len(buf) >= ctx + 1:                       # +1 for shifted target
            shard.write(buf[:ctx + 1]); buf = buf[ctx:]
            if shard.tokens >= shard_tokens:
                shard = shard.close_and_next()
    shard.close()
    return shard.manifest(seed=seed, tokenizer=tok.content_hash)

The manifest lists every shard with its token count, source domain, content hash, tokenizer hash and shuffle seed. The loader interleaves domains according to the mixture weights using its own seeded sampler, and because every shard is immutable and every choice is seeded, the global sequence order is a pure function of the manifest, the seed and the step. That is what lets the training loop resume exactly and skip a known-bad window of data after a loss spike.

Failure modes

  • Global thresholds. One quality or heuristic threshold for all domains deletes most of the code or non-English data.
  • Transitive dedup clusters. Unverified LSH candidates chain into giant clusters and legitimate documents vanish.
  • Keeping the worst copy. No keep policy, so the surviving duplicate is the one with navigation junk.
  • Contaminated evaluation. Benchmarks leak through mirrors and the reported scores measure memorization.
  • Silent over-repetition. A small domain gets a large weight and is seen tens of times.
  • Tokenizer mismatch. Shards packed with a tokenizer version other than the one frozen for the model.
  • No drop ledger. Nobody can explain a knowledge gap or honor a removal request.

Trade-offs

Every filter trades data quality against quantity, and for small models trained far past compute-optimal token counts, quantity eventually binds. Strict filtering gives a cleaner corpus that must be repeated more; loose filtering gives more unique tokens and more junk. Near-dedup and classifier stages cost real compute, often a noticeable fraction of the training compute for a small model, but they are paid once and reused across runs. The defensible way to choose is empirical: proxy ablations at two or three filter strengths, each measured at a token budget proportional to the real one, with per-domain validation loss and a small task suite.

Key takeaway: Order the data pipeline from cheap to expensive, give every document a stable id and write every drop to a ledger, deduplicate exactly and then near-exactly with verified clusters and an explicit keep policy, set quality thresholds per domain, decontaminate against every evaluation set you will report, plan the mixture with explicit epoch caps per domain, and pack seeded, immutable shards with a manifest so the training loop can resume and skip data exactly.