A small language model has little spare capacity. It cannot average away thousands of contradictory answers, a template that silently drops the system turn, or a test set that shares questions with training. When a fine-tuned 1B to 8B model disappoints, the data is a more common cause than the learning rate.
This article treats curation as a build system rather than a cleaning script: raw sources in, a versioned dataset with train, validation and evaluation splits out, and every dropped record explained. The focus is fine-tuning data assembled from an organisation's own material: support tickets, chat logs, expert-written answers and tool traces. Corpus-scale pretraining filtering is covered in pretraining data filtering, dedup and mixture, and teacher-generated data in the distillation data recipe. This page is the pipeline that sits between your raw data and supervised fine-tuning.
What the pipeline has to guarantee
Write down what a finished dataset must promise. Every record is in one canonical schema that the training chat template accepts. Every record is legal to train on and holds no personal data or secrets. Near-identical prompts do not dominate the mix. Every assistant turn is an answer you would want imitated. And no evaluation prompt, or a paraphrase of one, appears in training.
Two operational properties make those checkable. The pipeline is deterministic, so the same inputs produce the same records in the same splits, and it is explainable, so for any source row you can say which stage removed it and why. Without them you cannot change one curation decision and measure its effect on the fine-tuned model.
The record, the manifest and the drop ledger
Start with one record type and make its id content-addressed: a hash of the canonical messages. That one decision removes a class of bugs. Exact duplicates collapse automatically, the same conversation gets the same id in every run, and splits and drop decisions can be keyed on ids instead of row numbers that shift when a source changes.
from dataclasses import dataclass, field
import hashlib, json
@dataclass
class Record:
messages: list # [{"role": "system"|"user"|"assistant"|"tool", "content": str}]
source: str # "tickets-2026q2", "expert-batch-07", ...
source_ref: str # pointer back to the original row, never the raw PII
licence: str # "internal", "cc-by-4.0", ... ; unknown is a drop reason
meta: dict = field(default_factory=dict) # task label, language, product area
@property
def id(self) -> str:
# Content-addressed: the same conversation always gets the same id,
# so dedup, splits and the drop ledger agree across runs.
canon = json.dumps(self.messages, ensure_ascii=False, sort_keys=True)
return hashlib.sha256(canon.encode("utf-8")).hexdigest()[:20]
def run_stage(name, fn, records, ledger):
kept = []
for r in records:
reason = fn(r) # None = keep, otherwise a short reason string
if reason is None:
kept.append(r)
else:
ledger.append({"id": r.id, "stage": name, "reason": reason, "source": r.source})
print(f"{name}: in={len(records)} kept={len(kept)}")
return keptEach stage writes a manifest: the record ids it kept plus a hash of its code version and parameters. A release names the manifests it was built from. The append-only drop ledger is the other half, and a count of it grouped by source and reason is the pipeline's best health report: a stage that suddenly drops 40 percent of one source is almost always a bug or an upstream format change.
Stage 1: normalise to one schema and one template
Raw sources disagree about role names, consecutive turns from the same speaker, and how tool calls are encoded. Normalisation maps them onto the schema the trainer expects, merging split turns and mapping roles such as 'customer' and 'agent' to user and assistant, and rejects what it cannot map.
Three checks earn their place here. A conversation must end with an assistant turn, otherwise there is nothing to learn. The rendered length, in the target model's tokens and through its real chat template, must fit the training context; truncating from the right removes the answer, which is the only part that carries loss. And the rendered text must round-trip through the template unchanged, because templates that silently drop system turns or mangle tool messages are common. The JSONL format article covers the record shapes trainers expect. Do not collapse whitespace inside code, tables or tool output; it changes meaning.
Stage 2: privacy, secrets and licence
Organisational data carries names, emails, account ids, access tokens and pasted configuration, and a fine-tuned model can repeat them. Detect in layers: pattern rules for structured identifiers, a named-entity model for people and addresses, and a secrets scanner for high-entropy strings and keys.
Then decide per type whether to drop or redact. A typed placeholder such as <EMAIL_1>, numbered consistently within a conversation, keeps the exchange coherent without teaching real values. Drop the record when the sensitive content is the point of the exchange, such as a pasted credential file.
Licence is the quieter check: every source needs a recorded basis for training use, and records without one go to the ledger, not a release.
Stage 3: deduplicate prompts, not just records
Exact dedup on the content hash is free. The expensive problem is near-duplicates: the same password-reset question asked eleven thousand times with different names and typos. Left in, it teaches the model that every question is about password resets and wastes training steps on one behaviour. The fix is to cluster on the user side of each conversation, because two different answers to the same question are still redundant prompts, and then keep a small number of the best members of each cluster.
from datasketch import MinHash, MinHashLSH
def shingles(text, n=5):
w = text.lower().split()
return {" ".join(w[i:i + n]) for i in range(max(1, len(w) - n + 1))}
def cluster_prompts(records, threshold=0.8, num_perm=128):
# Cluster on the USER side only: two answers to the same question
# are near-duplicates for training purposes even if the answers differ.
lsh = MinHashLSH(threshold=threshold, num_perm=num_perm)
parent = {}
def find(x):
while parent[x] != x:
parent[x] = parent[parent[x]]
x = parent[x]
return x
for r in records:
prompt = " ".join(m["content"] for m in r.messages if m["role"] == "user")
mh = MinHash(num_perm=num_perm)
for s in shingles(prompt):
mh.update(s.encode("utf-8"))
parent[r.id] = r.id
for other in lsh.query(mh):
parent[find(r.id)] = find(other) # union with every match
lsh.insert(r.id, mh)
return {r.id: find(r.id) for r in records} # record id -> cluster id
def keep_per_cluster(records, cluster_of, k=2, score=lambda r: r.meta.get("quality", 0)):
by_cluster = {}
for r in records:
by_cluster.setdefault(cluster_of[r.id], []).append(r)
kept = []
for members in by_cluster.values():
kept.extend(sorted(members, key=score, reverse=True)[:k])
return keptMinHash with locality-sensitive hashing approximates Jaccard similarity over word shingles; datasketch is a widely used Python implementation. The right threshold depends on your text, so inspect sample pairs at a few values before fixing one. Union-find turns pairwise matches into clusters, and an optional embedding pass catches paraphrases that share few words.
The keep policy matters as much as the clustering. Keeping only the best member of each cluster removes useful variety in phrasing; keeping two or three preserves robustness to wording without letting a common question dominate. Record the cluster id on every kept record, because the split stage needs it.
Stage 4: label quality
In fine-tuning data the label is the assistant turn, and organisational labels are mixed: some replies are excellent, others are 'escalated to tier 2', an unhelpful macro, or wrong. Check cheapest first.
- Rules. Drop very short replies, known macros, and conversations closed unresolved or reopened.
- Executable checks. JSON must parse against its schema, SQL must run on a fixture database, code must compile, referenced URLs must exist.
- Model judging. A larger model scores each pair against a written rubric for correctness, policy consistency and tone. Store the score, not just pass or fail.
- Human audit. Label a few hundred judged records per release, stratified by score band and source, to measure the judge's precision at your threshold.
The audit keeps the judge honest, since judges drift with prompt changes and favour long, confident replies. Keep the rubric, judge model and threshold in the stage's manifest so releases compare fairly.
Stage 5: coverage and balance
Filtered data still reflects what users happened to ask, not what the model needs. Tag every record with a taxonomy (product area, intent, language, tool use) and look at the distribution; one intent often holds a third of the data while another has a few dozen examples.
Balance with caps and floors per bucket. A cap limits any bucket's contribution; a floor marks buckets that need new expert-written or synthetic data, because upsampling a 40-example bucket ten times teaches those 40 answers verbatim. Keep bucket counts in the release card.
Stage 6: leak-free splits and decontamination
The most common way to fool yourself is a random split. If near-duplicate prompts sit on both sides, the model is evaluated on questions it has effectively seen, and the score measures memorisation. Split by cluster instead, using a hash of the cluster id so the assignment is deterministic and stable when new data arrives.
import hashlib
def split_of(cluster_id, val=0.03, test=0.03):
# Deterministic and cluster-level: every near-duplicate of a test prompt
# lands in test too, and adding new data never moves an old cluster.
h = int(hashlib.sha256(cluster_id.encode()).hexdigest()[:8], 16) / 0xFFFFFFFF
return "test" if h < test else "val" if h < test + val else "train"
def ngrams(text, n=13):
w = text.lower().split()
return {tuple(w[i:i + n]) for i in range(len(w) - n + 1)}
def build_benchmark_index(benchmark_texts, n=13):
idx = set()
for t in benchmark_texts:
idx |= ngrams(t, n)
return idx
def contaminated(r, bench_idx, n=13):
text = " ".join(m["content"] for m in r.messages)
return "benchmark overlap" if ngrams(text, n) & bench_idx else NoneDecontamination points the same idea outward. Index the n-grams of every benchmark and internal evaluation set you report, and drop training records sharing a long n-gram with any of them; thirteen-word windows are a common choice, since shorter ones flag ordinary phrases. For small, high-stakes evaluation sets, also cluster with the evaluation prompts included and drop any training cluster that contains one. The evaluation pitfalls article covers what else goes wrong at scoring time.
Freeze the evaluation split. If it changes between releases, scores stop being comparable; new material goes into a new, versioned evaluation set.
Worked example: a support assistant
A team fine-tunes a 3B model to draft replies for a software product's support queue. Inputs are 180,000 resolved tickets from two years, 2,400 expert-written answers to common questions and 6,000 logged conversations from an older bot. The yields below are illustrative of what this kind of data usually looks like, not measurements to copy.
| Stage | Records out | Main drop reasons |
|---|---|---|
| Ingest | 188,400 | none; ids assigned, 3,100 exact duplicates collapsed |
| Normalise | 171,000 | no assistant reply, over 4,096 tokens, unmappable bot logs |
| Privacy | 166,500 | pasted credentials and logs; others redacted in place |
| Cluster, keep 2 | 61,000 | password reset, billing date and login clusters collapsed |
| Label quality | 38,500 | macros, escalations, judge below threshold |
| Balance | 34,000 | capped the three largest intents |
| Split + decon | 32,000 train, 1,000 val, 1,000 test | cluster-level split; 140 records overlapped the product-knowledge eval |
The biggest reduction comes from clustering, not filtering, and an ablation is how you check whether the 32,000 curated records beat all 171,000 normalised ones. The balance report also finds the real gap: tool-calling conversations are under one percent because the old bot never called tools. That becomes a data request, not a parameter change.
Curation decisions are experiments
Every threshold above looks reasonable and may be wrong. Decide with ablations: fine-tune the same base with the same recipe on the previous release and on one that changes exactly one thing, then compare on the frozen evaluation split and per-bucket metrics. Small models make this affordable. Log the change, both manifest hashes, the scores and the decision. Stricter filtering raises average quality but lowers coverage, and a model judge scales but carries biases, so expect surprises such as a stricter threshold hurting because it removed short correct answers.
Failure modes
- Template mismatch. Data is curated against one chat template and trained with another, so system turns or tool messages are dropped at training time. Render through the trainer's own template in Stage 1.
- Random splits. Near-duplicates on both sides inflate evaluation scores. Split by cluster.
- Judge drift. A changed rubric or judge model shifts what passes, and the next release looks better or worse for reasons unrelated to the data. Version the judge and audit every release.
- Silent source changes. An upstream export renames a column and a whole source fails normalisation. Alert on per-source drop rates.
- Over-redaction. Entity detection replaces product names and error codes with placeholders. Audit redactions on a sample.
What to do next
- Define the canonical record and content-addressed id, and start writing a drop ledger from the first stage.
- Normalise through your trainer's real chat template and reject records that end without an assistant turn or exceed the token budget.
- Add layered PII and secret detection with consistent placeholders, and record a licence basis per source.
- Cluster prompts with MinHash, inspect pairs at two or three thresholds, and keep two or three members per cluster.
- Add rule, executable and judge checks for label quality, and audit a stratified sample by hand each release.
- Tag a task taxonomy, apply caps and floors, and turn under-filled buckets into data requests.
- Split by cluster hash, decontaminate against every evaluation set you report, and freeze the evaluation split.
- Change one curation decision at a time and keep an experiment log of the ablation results.