When a small model is distilled from a large one by fine-tuning on the teacher's outputs, the dataset is the method. The student learns the prompts you chose, the answers you kept and the format those answers share, and nothing else. Two teams with the same teacher, student and training code can get very different models because one kept every teacher output and the other verified, filtered and balanced them.
This article is a recipe for that dataset: what to specify before generating anything, how to build the prompt pool, how many teacher samples to draw, how to verify them, how to choose which ones teach the student something, how to size the job with yield arithmetic, and how to keep evaluations honest. Loss functions and logit-level distillation are covered in SLM distillation in depth, and teacher access and logit caching in distilling from an LLM teacher. Here the focus is the data.
The pipeline at a glance
The recipe is a pipeline with a measured yield at every stage. Each stage removes examples, and the product of those yields tells you how many prompts to start with. Every example carries a provenance record: where its prompt came from, which teacher and settings produced it, which checks it passed, and why it was kept. Rejected examples are kept too, labelled with the reason, because they are your best diagnostic data.
Step 1: write the task spec first
Start from what the student must do, not from what the teacher can produce. Divide the target behaviour into slices, the distinct skills or input types the student will face, and give each a target count, a verifier and an output contract. Slices make every later decision concrete: you measure pass rates per slice, balance per slice, and evaluate per slice.
# distill_spec.yaml
student: 1.5B base model, 4k context
output_contract: JSON with keys {answer, rationale}, rationale <= 120 words
slices:
- name: sql_from_question
target: 8000
verifier: execute_and_compare # run SQL on fixture DB, compare result sets
- name: ticket_triage
target: 6000
verifier: label_match # gold labels from historical tickets
- name: policy_answer
target: 4000
verifier: judge_rubric_v3 # calibrated LLM judge
- name: refuse_out_of_scope
target: 2000
verifier: refusal_classifier
eval_holdout_per_slice: 300Include a slice for what the student should decline. A model distilled only on answerable prompts learns to answer everything, including things it cannot know.
Step 2: build the prompt pool
The best prompts are real ones: anonymised production queries, historical tickets, actual user questions. They carry the true distribution of phrasing, length and messiness. Scrub personal data before anything reaches a teacher API. Fill gaps with seed prompts written by domain experts and expanded synthetically, and steer synthetic expansion toward the slices and difficulty levels the logs underrepresent.
Deduplicate the pool before generation, not after, so you never pay the teacher twice for the same question. Exact hashing after normalisation catches copies; MinHash or embedding neighbourhoods catch paraphrases. Then carve out the evaluation split per slice and freeze it. Doing this before generation is the only way to guarantee no teacher output derived from an evaluation prompt ever reaches training.
Step 3: sample the teacher k times
Draw several completions per prompt, typically 2 to 8, at a moderate temperature. Multiple samples serve three purposes: they raise the chance that at least one is correct, they give you a choice among correct answers (shortest, best formatted), and the agreement between samples is a free confidence signal. A prompt where the teacher's samples disagree wildly is often ambiguous, and ambiguous prompts teach noise.
Fix the system prompt and output contract for the whole run and record every setting. Changing the teacher prompt halfway produces two subtly different styles in one dataset, and the student learns the inconsistency.
{
"id": "sql_from_question/000417/s2",
"slice": "sql_from_question",
"prompt_source": "prod_logs_2026q3",
"prompt": "...",
"teacher": {"model": "teacher-v1", "temperature": 0.7, "sample": 2,
"system_prompt_sha": "9f2c..."},
"completion": "...",
"completion_tokens": 212,
"checks": {"json_valid": true, "execute_and_compare": true},
"student_pass_rate": 0.25,
"status": "kept",
"reason": "verified; shortest correct of 3"
}
Step 4: verify in tiers
The teacher is wrong more often than its fluency suggests, and every error you keep becomes an error the student learns confidently. Verify with the strongest check each slice allows, in order of trust.
- Programmatic checks. Execute the SQL against a fixture database and compare result sets; run generated code against unit tests; parse JSON against the schema; compare a numeric answer to a known value. These are cheap, deterministic and near-perfect for what they cover.
- Reference checks. Compare against a gold label or reference answer where one exists, such as historical ticket labels.
- Judged checks. For open-ended answers, use an LLM judge with a written rubric. Calibrate it before trusting it: have people label a few hundred random outputs, measure how often the judge's pass agrees with a human pass, and report precision on passes, since a false pass puts a bad example into training. Re-check calibration whenever the rubric, judge or teacher changes.
Always run format checks on every slice regardless of content checks. Format errors are cheap to catch and the student copies them with great fidelity.
Step 5: select what teaches the student something
Not every correct example is worth training on. If the base student already solves a prompt reliably, the example spends compute and pushes the mixture toward easy cases. Measure it: sample the student a few times on each prompt, run the same verifier, and record its pass rate. Keep mostly prompts where the teacher succeeds and the student does not yet succeed reliably, plus a small share of easy ones so basic behaviour stays anchored. Prompts no teacher sample passes go to a review queue, not into training.
Among correct samples, prefer the shortest that satisfies the output contract. Long reasoning that a much larger teacher needs can be more than a small student can reproduce within its context and capacity; a concise correct rationale usually transfers better. Normalise formatting at this stage too, so every kept example obeys one contract exactly.
import random
def select(prompt, samples, verify, student_pass_rate, easy_keep=0.15):
# Returns (example or None, reason); samples are teacher completions.
passed = [s for s in samples if verify(prompt, s)]
if not passed:
return None, "teacher_failed" # route to human review
if len(passed) / len(samples) < 0.5:
return None, "teacher_unstable" # likely ambiguous prompt
if student_pass_rate >= 0.9 and random.random() > easy_keep:
return None, "too_easy"
best = min(passed, key=lambda s: s.completion_tokens)
return best, "kept"
Worked example: one prompt through the pipeline
Take a prompt from the SQL slice: "Which three customers placed the most orders last month, excluding cancelled orders?" It came from anonymised analyst logs, survived deduplication, and is not in the held-out split. The teacher is sampled four times at temperature 0.7.
The verifier runs each generated query against the fixture database and compares result sets with a reference query. Sample 1 forgets the cancelled filter and fails. Samples 2, 3 and 4 return the right rows; the teacher pass rate is 0.75, above the instability threshold. The base student, sampled four times on the same prompt, passes once, so its pass rate is 0.25: this prompt teaches it something.
Of the three correct samples, sample 3 is the shortest, with a 60-word rationale and a clean query, so it is kept and normalised to the JSON contract. Sample 1 goes to the rejects table labelled with the failing check, where a weekly review notices that the teacher misses status filters often enough to justify adding a hint to the teacher prompt for the next run. The decontamination check finds no near-match in the evaluation set or public benchmarks, and the example enters the SQL slice with its full provenance record.
Step 6: size the job with yield arithmetic
Multiply the measured yields to decide how many prompts to generate. Run a pilot of a thousand or so prompts per slice to measure them; do not guess. Here is the arithmetic for a 20,000-example target with illustrative pilot numbers.
| Stage | Yield | Remaining from 40,000 prompts |
|---|---|---|
| Prompt dedup | 0.85 | 34,000 |
| At least one verified teacher sample | 0.70 | 23,800 |
| Not too easy for the student | 0.75 | 17,850 |
| Decontamination | 0.99 | about 17,670 |
The combined yield is about 0.44, so 40,000 prompts fall short. The target needs about 20,000 / 0.442, roughly 45,300 prompts, and with four samples each and a few hundred output tokens per sample, the teacher bill is on the order of 100 million output tokens. Work this out per slice: a slice with a low teacher pass rate needs proportionally more prompts or a better teacher prompt, and the pilot is where you discover that cheaply.
Step 7: decontaminate, mix and document
Check every kept example against your held-out evaluation set and any public benchmarks you will report, using n-gram overlap on prompts and answers and embedding similarity for paraphrases. Teacher models may have seen benchmark items and can reproduce them in synthetic prompts. A contaminated benchmark score is worse than none, because it makes decisions for you. SLM evaluation pitfalls covers the measurement side.
Mix slices by weight, not by whatever the pipeline produced. Hit the spec's targets, upsample slices that are under target only up to a modest repetition limit, and add a share of general instruction data so the student keeps abilities outside the task. Pretraining-style filtering, deduplication and mixture design are treated in pretraining data filtering, dedup and mixture.
Ship the dataset with a data card: the spec, teacher and settings, per-stage yields per slice, judge calibration results, decontamination method and hits, and the licence terms that govern the teacher's outputs. Check the teacher provider's terms of use before you start; some restrict using outputs to train competing models.
Failure modes
- Confident teacher errors. Unverified slices pass errors straight into the student. Track per-slice error rates in a human-audited sample of kept examples.
- Judge drift. A judge prompt edited mid-run changes what passes. Version the rubric and re-run calibration on every change.
- Style collapse. Every answer opens the same way and has the same length, and the student becomes rigid. Inspect samples, not just metrics.
- Verbosity inheritance. The student copies long rationales it cannot reason through and pads answers. Enforce length limits in the contract.
- Distribution mismatch. Synthetic prompts are cleaner than real traffic, so the student looks strong offline and stumbles in production. Keep real prompts as the backbone.
- Leakage. Evaluation prompts were generated from the same seeds as training prompts, so near-copies straddle the split. Split before expansion.
Trade-offs
| Decision | Cheaper option | Better option |
|---|---|---|
| Samples per prompt | 1, keep if verified | 4 or more, choose shortest correct |
| Verification | Judge everything | Programs first, calibrated judge only where needed |
| Selection | Keep all correct | Weight toward prompts the student fails |
| Prompt source | Synthetic only | Real prompts plus targeted synthesis |
What to do next
- Write the task spec: slices, target counts, a verifier and an output contract for each, and a refusal slice.
- Assemble and deduplicate the prompt pool, then freeze a held-out evaluation split per slice before calling the teacher.
- Run a pilot of about a thousand prompts per slice with several teacher samples each, and measure every stage's yield.
- Calibrate any LLM judge against human labels and record its precision on passes.
- Measure the base student's pass rate on pilot prompts and set the selection rule.
- Compute the prompt count from the combined yield, run the full generation, decontaminate, mix by weight and publish the data card with the dataset.