Every foundation model starts with a decision about which text to keep. Teams talk about that decision in terms of quality, but it is also a bill: CPU-hours to pull text out of HTML, GPU-hours to score documents with a classifier, tokens spent asking a large model to label examples, terabytes of intermediate files, and, often the largest item, the training runs needed to prove a filter actually helps. If you do not model those costs, you either overspend on stages that do not matter or skip the validation that does.
This article treats curation as cost accounting: the unit that makes stages comparable, a cost model you can run, a crawl-sized example, the hidden line items, and a break-even test against the training compute that better data saves. How to make individual stages fast on GPUs is covered in LLM data curation pipelines; here the question is what the whole funnel costs and where the money goes.
The unit: cost per retained token
A curation pipeline is a funnel. Raw documents enter at the top; each stage reads every document that reaches it and passes some fraction on. The key property is that a stage is paid on its input, not its output. A classifier that keeps 10% of documents still has to score all of them. That means the cost of a stage depends on its position: the same classifier placed first sees every document, placed fifth it may see a sixth of them.
The useful unit is cost per retained token. Sum the cost of every stage, add the fixed costs (annotation, validation runs, storage), and divide by the tokens that survive. That number lets you compare a cheap aggressive pipeline with an expensive gentle one, and it lets you compare curation spend with training spend, which is also priced per token.
Formally, if stage i costs ci per document and passes fraction ri, and N documents enter, the variable cost is N times the sum over stages of ci multiplied by the product of the pass rates of all earlier stages. The fixed costs sit beside that sum and do not shrink when you filter harder.
The full cost ledger
| Line item | Cost driver | Scales with | Easy to forget? |
|---|---|---|---|
| HTML extraction | CPU time per page | Raw documents | No, but it is often the biggest compute stage |
| Language ID, heuristics | CPU time, tiny per doc | Docs after extraction | No |
| Fuzzy dedup | CPU or GPU for MinHash, memory for LSH buckets | Docs after heuristics | Memory sizing, yes |
| Quality classifier | GPU-hours | Docs reaching it | Position in the funnel, yes |
| LLM annotation | Tokens of a large model | Labelled sample size | No, and it is usually small |
| Proxy ablations | GPU-hours of small training runs | Filter variants times seeds | Yes, often the largest item |
| Storage | GB-months of intermediates | Text kept between stages | Yes |
| Transfer | Per-GB egress | Raw bytes read across regions | Yes |
| Re-runs | Whole-pipeline repeats | How often thresholds change | Yes |
The table separates variable costs, which scale with documents, from fixed costs, which scale with how carefully you validate. Teams tend to measure the first group because it shows up on a cluster dashboard, and miss the second because it is spread over many small jobs.
A cost model you can run
The model below is deliberately small. Each stage declares whether it runs on CPU or GPU, its cost in machine-seconds per document and its pass rate. The prices are inputs: put in your own committed or spot rates.
from dataclasses import dataclass
@dataclass
class Stage:
name: str
unit: str # "cpu" (vCPU-seconds) or "gpu" (GPU-seconds)
sec_per_doc: float
pass_rate: float
PRICE_PER_HOUR = {"cpu": 0.04, "gpu": 2.00} # illustrative, replace with your rates
def funnel_cost(stages, docs_in, tokens_per_doc, fixed=0.0):
docs, total, rows = docs_in, 0.0, []
for st in stages:
hours = docs * st.sec_per_doc / 3600
cost = hours * PRICE_PER_HOUR[st.unit]
rows.append((st.name, docs, hours, cost))
total += cost
docs *= st.pass_rate
kept_tokens = docs * tokens_per_doc
per_btok = (total + fixed) / (kept_tokens / 1e9)
return rows, total, kept_tokens, per_btok
def rank(st):
"""Order independent filters by cost per rejected document, cheapest first."""
return st.sec_per_doc * PRICE_PER_HOUR[st.unit] / max(1e-9, 1 - st.pass_rate)The rank function encodes a classical result for independent filters: to minimise expected cost, run them in increasing order of cost divided by rejection probability. A filter that is cheap and rejects a lot belongs at the front; an expensive one that rejects a lot still belongs late, because the cheap ones shrink its input first. The result assumes independence. Real filters are correlated (pages that fail language ID also tend to fail heuristics), so treat the ranking as a starting point and measure pass rates conditional on the real order. Some stages are not really filters: a PII scrub mostly rewrites text, and compliance usually wants it last, on exactly what ships, even where the rank says otherwise.
Worked example: one crawl snapshot
Take a crawl of 2 billion pages, comparable to a single Common Crawl snapshot, which typically holds a few billion. Assume these illustrative figures: extraction at 15 ms of CPU per page keeps 85%; language ID at 0.2 ms keeps 45%; heuristic filters at 1 ms keep 70%; MinHash dedup at 3 ms keeps 60%; a small encoder classifier scoring 1,500 documents per second per GPU keeps the top 10%; a PII scrub keeps 98%. Prices are $0.04 per vCPU-hour and $2.00 per GPU-hour, and kept documents average 1,000 tokens.
| Stage | Docs in | Machine-hours | Cost |
|---|---|---|---|
| Extract | 2.000B | 8,333 CPU | $333 |
| Language ID | 1.700B | 94 CPU | $4 |
| Heuristics | 0.765B | 212 CPU | $8 |
| MinHash dedup | 0.535B | 446 CPU | $18 |
| Quality classifier | 0.321B | 59 GPU | $119 |
| PII scrub | 0.032B | 4 CPU | $0 |
The pipeline keeps 31.5 million documents, about 31.5 billion tokens, for $483 of variable compute, or about $15 per billion retained tokens. Now move the classifier to run straight after extraction, which is a tempting design because it removes 90% of documents at once. It now scores 1.7 billion documents: 315 GPU-hours and $630 for that stage alone, and the total doubles to $966 for exactly the same output. The cheap CPU filters would have thrown away four documents in five before the GPU saw them.
The second lesson from the same numbers is less obvious: the per-document processing bill is small. A few hundred dollars per snapshot is noise next to a training run. The money in curation is elsewhere, in the fixed line items.
The hidden line items
Proxy ablations. A filter is only worth keeping if a model trained on its output beats one trained without it. The standard check is to train small proxy models on each variant and compare them on held-out benchmarks. Training compute is roughly 6 times parameters times tokens, so a 1B-parameter proxy on 30B tokens costs 1.8e20 FLOPs. At an assumed effective throughput of 400 TFLOP/s per GPU, that is 125 GPU-hours, or $250 at the illustrative price. Twenty filter variants with three seeds each is 60 runs: 7,500 GPU-hours and $15,000, thirty times the processing bill. Seeds are not optional: at small scale, seed-to-seed variance on benchmarks can be as large as the effect you are trying to measure.
LLM annotation. Model-based filters are trained on labels from a large model. FineWeb-Edu is the public reference point: roughly 450,000 web pages were scored for educational value by Llama-3-70B-Instruct, those labels trained a classifier of about 109M parameters, and keeping pages scoring 3 or more retained 1.3 trillion of FineWeb's roughly 15 trillion tokens, removing about 92% of the data. Pricing that pattern: 450,000 documents at about 1,600 tokens each including the prompt is 720 million tokens, which at an illustrative $0.50 per million input tokens is $360. Annotation feels expensive because the per-token price of a large model is high, but the sample is small. It is rarely the dominant cost.
Storage and transfer. Extracted text at an average of 5 KB per document for 1.7 billion documents is 8.5 TB; at an illustrative $0.02 per GB-month that is $170 a month for one intermediate, and pipelines keep several. Transfer is the trap: reading raw crawl data from another region or provider is billed per GB and can exceed the entire compute bill. Run the pipeline where the data lives.
Re-runs. The most expensive habit is writing filtered copies of the data. When someone changes a threshold from 3 to 2.5, a copy-based pipeline re-runs everything downstream. A score-based pipeline stores one row of signals per document (language score, heuristic flags, dedup cluster id, classifier score) and materialises a training mix as a query over that table.
Store signals, not subsets
Store signals, not subsets. The schema below is enough for most decisions and turns a threshold change from a pipeline run into a scan.
# one row per document, written once per pipeline version
signals = {
"doc_id": "cc-2024-10/seg-00417/rec-88213",
"pipeline_version": "v7",
"lang": "en", "lang_score": 0.97,
"heuristic_flags": [], # e.g. ["short_lines", "symbol_ratio"]
"dedup_cluster": 18822341, "is_cluster_rep": True,
"edu_score": 3.4, # classifier regression output
"tokens": 1184,
}
def select(rows, min_edu=3.0, min_lang=0.65):
for r in rows:
if (r["lang"] == "en" and r["lang_score"] >= min_lang
and not r["heuristic_flags"] and r["is_cluster_rep"]
and r["edu_score"] >= min_edu):
yield r["doc_id"]Because every score is kept, you can also price a threshold before you commit to it: sum tokens for rows above each candidate threshold and you have the size of the resulting dataset without moving a byte. Keep pipeline_version in every row, so a classifier retrained on new labels produces a new column rather than silently overwriting the old one. For deterministic mixing and resumable loading of whatever you select, see the LLM training data pipeline.
Break-even against training compute
Curation pays for itself when it reduces the training compute needed to reach a target quality. Suppose ablations show that a 7B model reaches your benchmark target on 1T curated tokens, where unfiltered data needed 2T. The saving is 6 times 7e9 times 1e12, 4.2e22 FLOPs. At the same assumed 400 TFLOP/s effective throughput that is about 29,000 GPU-hours, roughly $58,000 at the illustrative price. Against that, the curation bill from this article (a few hundred dollars of processing per snapshot, $15,000 of ablations, a few hundred for annotation, storage and engineering time) clears easily.
The break-even shifts in three situations. First, small runs: if you only ever train a 1B model once, the ablations can cost more than the run they inform. Reuse public ablation results or cheaper heuristics instead. Second, token-limited runs: a filter that removes 92% of data multiplies the raw crawl you need by about twelve, and once unique tokens run out you must repeat data for several epochs or relax the threshold. Third, amortisation: a curated corpus feeds many runs, so divide the fixed costs across every model that will train on it, not just the first.
Operational guidance
- Instrument every stage with documents in, documents out, machine-seconds and bytes written, keyed by pipeline version. Without pass rates measured in the real order you cannot rank stages.
- Run cheap CPU filters first and give GPU stages only what survives. Re-check the order whenever a stage is added; a new cheap filter placed after the classifier wastes most of its value.
- Make stages restartable at the shard level so spot preemption costs one shard, not a snapshot. Write each output shard atomically with a manifest entry.
- Budget ablations explicitly: number of variants, seeds and proxy size, with a stopping rule. Kill variants that lose to the baseline on the first seed by more than the measured seed spread.
Failure modes
- Classifier first. The worked example doubled its bill this way. Symptom: GPU-hours dominate a pipeline whose output is small.
- Under-seeded ablations. One seed per variant, a 0.5-point win, a filter adopted on noise. Symptom: the win does not reproduce at the next scale.
- Copy-based outputs. Every threshold debate triggers a full rerun. Symptom: the same snapshot processed four times in a month.
- Over-filtering. A narrow classifier keeps only one style of text, and downstream tasks outside that style regress. Symptom: benchmark gains on the classifier's own domain, losses elsewhere. Keep a diversity check, such as the domain and topic mix of kept tokens.
- Label drift. Annotations regenerated with a different model or prompt shift scores, so the same threshold keeps a different set. Version prompts and annotators with the labels.
Trade-offs
Aggressive filtering buys quality per token at the cost of total tokens and diversity; gentle filtering preserves supply and pushes the cost into training compute. Model-based filters beat heuristics on quality but add annotation, GPU inference and drift risk. Large ablations give reliable decisions but can cost more than the corpus. The sensible default for most teams is cheap heuristics and dedup on everything, a model-based filter scored and stored rather than applied, and a small, seeded ablation grid that decides the threshold. For using GPUs on synthetic rather than filtered data, the cost per accepted sample is analysed in synthetic data generation on GPUs, and for attributing shared GPU spend back to teams see LLM cost attribution.
What to do next
- Write down your funnel: every stage, whether it runs on CPU or GPU, its measured seconds per document and its pass rate in the current order.
- Run the cost model with your real prices and compute cost per billion retained tokens.
- Sort independent filters by cost per rejected document and confirm no GPU stage sits ahead of a cheap CPU filter.
- Convert filtered copies into a per-document signal table keyed by pipeline version.
- Budget ablations as a line item: proxy size, tokens, variants, seeds and a stopping rule.
- Compute the break-even: training FLOPs saved at your target quality against the total curation bill, amortised over the runs that will use the corpus.
- Check that compute runs in the same region as the raw data before the next snapshot.