Small language models live or die by their training data, and for most specialised tasks there is not enough real labelled data. Synthetic data, examples written by a larger model, a program or a template, fills the gap. It is also the easiest way to produce a hundred thousand examples that teach a model almost nothing, because a generator asked the same question ten thousand times gives ten thousand variations of the same answer.

This article is about the generation side: where inputs come from, how to make them diverse on purpose, how to filter them, and how to mix them with real data without degrading the model. Sampling a teacher several times per prompt and verifying in tiers is covered in the SLM distillation data recipe, and curating real data in the training data curation pipeline.

Advertisement

What synthetic data is for

Four kinds of synthetic data show up in small-model work, and they fail differently. Instruction data (prompt and response pairs for supervised fine-tuning) teaches format, task coverage and tone. Task-labelled data (inputs with labels or structured outputs, such as extraction targets or SQL) teaches one narrow skill and is often verifiable by a program. Pretraining-style text, such as explanations and exercises, densifies knowledge and reasoning for models trained from scratch; the 2023 paper 'Textbooks Are All You Need' trained a 1.3-billion-parameter code model largely on filtered web code plus generated textbook-style text and exercises. Preference pairs (a better and a worse response) feed DPO-style tuning.

Before generating anything, write a task specification: what the model must do and what real data is missing. Every later choice is checked against it.

The hard part is diversity, not volume

A generator at a fixed prompt samples from a narrow region of its output distribution. Ask for 'a customer question about a refund' 10,000 times and you get the same handful of situations reworded, with the same polite tone and the same length. Training on that teaches the small model a template, and it fails on the messy, terse, misspelt or multi-part inputs real users send. Higher temperature adds noise, not situations.

Diversity has to be designed into the inputs. The reliable method is to make generation conditional on explicit, sampled attributes: what skill the example exercises, what domain it is in, how hard it is, who is asking and in what style. The cross-product of attribute values defines the coverage you want, and counting kept examples per cell tells you whether you got it.

Synthetic data pipeline for a small modelSeed taxonomyskills x domains x difficultyPersonas / documentsreal grounding materialInput generatorSelf-Instruct, Evol-InstructResponse writerteacher model, k samplesFilter stackformat, length, near-dup, semantic dup,verifier / judge, decontaminationMix with real dataratio is a hyperparameterkept + provenanceTrain SLMSFT / LoRAEvaluate on REAL held-out dataper-slice scores, failurescoverage gaps feed the taxonomyGap analysiswhich slices failweak slices become new taxonomy cellsDiversity is designed upstream; quality is enforced downstream; truth is measured only on real data.
Attributes sampled from a taxonomy and real grounding material drive input generation. Responses are written, filtered and mixed with real data. Evaluation on real held-out data finds weak slices, which become new taxonomy cells.
Advertisement

Generation strategies

StrategyHow it worksGood forWatch out for
Self-Instruct (2022)start from a small pool of human-written seed tasks; the model writes new tasks prompted with sampled seeds; near-duplicates (ROUGE-L overlap) are filtered and kept tasks rejoin the poolbroad instruction coverage from a few hundred seedsdrifts toward the generator's favourite task types
Evol-Instruct (2023, WizardLM)rewrite existing instructions to be harder (add constraints, deepen, concretise, add reasoning steps) or broader (new topic in the same spirit)raising difficulty and length of an existing setevolved prompts that are unanswerable or contradictory
Persona-driven (2024)condition each generation on a persona drawn from a very large persona collection, as in 'Scaling Synthetic Data Creation with 1,000,000,000 Personas'varied vocabulary, goals and stylespersonas that are stereotypes; uneven coverage
Document-groundedgive the generator a real document chunk and ask for questions it answersdomain knowledge, RAG and extraction tasksquestions answerable only with the chunk in view
Instruction backtranslation (2023)take real human-written text as the response and have a model write the instruction that would produce it, then filternatural responses, real distributionweak or mismatched instructions; needs strong filtering
Programmatic / templatecode produces inputs and exact labels (dates, arithmetic, schema variants)verifiable skills, perfect labelsstiff phrasing unless paraphrased afterwards

In practice you combine them: a taxonomy defines the cells, personas and documents vary the surface and grounding, Evol-Instruct-style rewrites fill the hard end of the difficulty axis, and programmatic generation supplies exact labels wherever a program can compute them.

A generator with designed diversity

The sketch below generates questions for a small text-to-SQL model. Each request samples a skill, a domain, a difficulty and a persona, so diversity comes from the specification rather than from temperature. A cheap shingle-based Jaccard filter drops near-duplicates as they arrive. llm is any function that sends a prompt to your generator model and returns text.

import json, random, re, hashlib

TAXONOMY = {
    "skill": ["filter rows", "aggregate with GROUP BY", "join two tables", "top-N per group",
              "date ranges", "NULL handling", "window function", "subquery"],
    "domain": ["orders", "refunds", "inventory", "support tickets", "subscriptions"],
    "difficulty": ["one clause", "two clauses", "three or more clauses with a trap"],
}
PERSONAS = ["finance analyst preparing month-end", "support lead tracking backlog",
            "warehouse manager", "growth marketer", "new hire who writes vaguely"]

GEN_PROMPT = """You write realistic questions a {persona} would ask a SQL assistant.
Schema:
{schema}
Write ONE question that requires: {skill}, about {domain}, difficulty: {difficulty}.
Use the persona's vocabulary, not SQL terms. Output JSON: {{"question": "..."}}"""

def sample_spec(rng):
    spec = {k: rng.choice(v) for k, v in TAXONOMY.items()}
    spec["persona"] = rng.choice(PERSONAS)
    return spec

def shingles(text, n=5):
    toks = re.findall(r"\w+", text.lower())
    return {" ".join(toks[i:i + n]) for i in range(max(1, len(toks) - n + 1))}

def jaccard(a, b):
    return len(a & b) / max(1, len(a | b))

def generate_inputs(llm, schema, target, seed=0, max_sim=0.6):
    rng, kept, seen = random.Random(seed), [], []
    attempts = 0
    while len(kept) < target and attempts < target * 5:
        attempts += 1
        spec = sample_spec(rng)
        raw = llm(GEN_PROMPT.format(schema=schema, **spec), temperature=1.0)
        try:
            q = json.loads(raw)["question"].strip()
        except (ValueError, KeyError, TypeError):
            continue                                   # drop malformed output
        sh = shingles(q)
        if any(jaccard(sh, s) > max_sim for s in seen):
            continue                                   # near-duplicate of a kept question
        seen.append(sh)
        kept.append({"question": q, "spec": spec,
                     "id": hashlib.sha1(q.encode()).hexdigest()[:12]})
    return kept, attempts

Keep the attempts counter: a falling kept-to-attempted ratio means the generator has exhausted that region of the taxonomy. Shingle Jaccard only catches lexical repeats; add an embedding-based pass later that drops pairs above a cosine similarity threshold, tuned by reading pairs just above and below it.

The filter stack

Generated examples go through filters in order of cost, cheapest first, and every drop is logged with a reason so you can see which stage removes what:

  1. Format and schema: parseable, required fields present, length within bounds, right language.
  2. Lexical and semantic deduplication: within the synthetic set and against the real set.
  3. Verification: execute what can be executed (run the SQL or code, check the arithmetic, validate the JSON against a schema). This is the strongest filter you have; prefer tasks where it exists.
  4. Model judging: for what cannot be executed, a judge model scores against a written rubric. Calibrate it on a few hundred human-labelled examples before trusting its threshold, and use a different model family from the generator where possible, because a model tends to rate its own style highly.
  5. Decontamination: drop anything that overlaps your evaluation sets, using long n-gram overlap and embedding similarity. Generators have often seen public benchmarks and will reproduce their items.

For the SQL example, verification is concrete: run the generated query read-only against a seeded database with a time limit, and keep the pair only if it runs and returns rows.

import sqlite3, time

def verify_sql(db_path, sql, timeout_s=2.0):
    """Keep a (question, sql) pair only if the SQL is read-only, runs and returns rows."""
    if not sql.lstrip().lower().startswith(("select", "with")):
        return False, "not read-only"
    con = sqlite3.connect(f"file:{db_path}?mode=ro", uri=True)
    deadline = time.monotonic() + timeout_s
    con.set_progress_handler(lambda: 1 if time.monotonic() > deadline else 0, 10_000)
    try:
        rows = con.execute(sql).fetchmany(50)
    except sqlite3.Error as e:
        return False, f"error: {e}"
    finally:
        con.close()
    return (len(rows) > 0), ("ok" if rows else "empty result")

Execution proves the SQL runs, not that it answers the question. Add a second check for semantics: have a different model write SQL for the same question independently and keep the pair only when both queries return the same result set. Disagreements go to a small human review queue, which doubles as a measurement of how often the filter is wrong.

Worked example: a text-to-SQL model for one schema

A team wants a 1.5-billion-parameter model that answers analysts' questions over an internal schema of 14 tables, running on a single small GPU. They have 400 real question and SQL pairs from logs, of which they hold out 150 as the test set and never show them to the generator.

The taxonomy has 8 skills, 5 domains and 3 difficulty levels, giving 120 cells, crossed with 5 personas. They target 60 kept examples per cell, 7,200 in total. In a run like this, suppose 70 percent of generated questions survive format and deduplication filters, 75 percent of the teacher's SQL executes and returns rows, and 85 percent of those agree with the independent second query. The yield is 0.70 x 0.75 x 0.85, about 0.45, so they budget roughly 16,000 generation attempts. These rates are illustrative; measure your own on a pilot of a few hundred before sizing the job.

The first evaluation on the 150 real test questions shows strong scores on filters and aggregations and weak ones on NULL handling and terse, vague questions. Both are traced to the data: NULL questions mostly failed the 'returns rows' filter, which silently removed exactly the hard cases, and the 'vague new hire' persona was underrepresented after deduplication. The fixes are a seeded database that contains NULLs in the right places and per-cell quotas enforced after filtering. That loop, evaluate on real data, find the weak slice, trace it to a generation or filter decision, is the actual work of synthetic data.

Model collapse and mixing with real data

Training models on the outputs of models, generation after generation, loses the tails of the original distribution. The 2024 Nature paper 'AI models collapse when trained on recursively generated data' showed rare events disappearing and outputs narrowing as each generation trained on the previous one's samples. Follow-up work ('Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data', 2024) found that keeping the real data and adding synthetic data to it, rather than replacing it, avoided the degradation in its settings.

For a small-model team the practical rules are short. Keep real data in every training mix, even if it is a small fraction. Sweep the real-to-synthetic ratio. Never evaluate only on synthetic data; a model trained on a generator's style scores well on that style. And do not feed your own model's outputs back as training data without the same verification you apply to the teacher's. Fine-tuning mechanics, including loss masking and packing, are covered in SFT in depth.

Failure modes

  • Template collapse: thousands of examples with one structure. Detect it by clustering embeddings and counting examples per cluster.
  • Filters that remove the hard cases: a 'must return rows' or 'judge score above 8' rule quietly deletes difficult, ambiguous or edge-case examples. Compare the taxonomy distribution before and after each filter.
  • Confident wrong labels: a teacher's plausible but wrong answers teach the student to be wrong fluently. Verification, agreement checks and human audits of random samples are the defences.
  • Benchmark leakage: generated items that paraphrase evaluation questions inflate scores. Decontaminate against every evaluation set, including your own; see evaluation pitfalls for small models.
  • Style transfer of the teacher's tics: the student inherits verbosity, hedging phrases and refusals. Strip or rewrite them, or the small model wastes tokens on them.
  • Licence and terms problems: some model providers' terms restrict using outputs to train other models. Check before generating, and record which model produced each example.

Operating a synthetic data pipeline

Treat generated data like code. Store every example with its provenance: generator model and version, prompt template version, sampled attributes, seed, filter results and drop reasons. Version datasets so a model can be traced to the exact data that trained it. Keep a fixed audit sample that a person reads at every regeneration. Track yield per filter stage and examples per taxonomy cell on a dashboard; sudden changes usually mean a generator update or a broken prompt rather than a better dataset. For pretraining-scale generation, the same deduplication and mixture ideas apply at larger scale; see pretraining data filtering, deduplication and mixture.

Trade-offs

Synthetic data is cheap per example and expensive per useful example. A stronger generator writes better data at higher cost. Verification-heavy pipelines produce reliable data only for tasks a program can check. Judges extend coverage to open-ended tasks at the cost of their own biases. More synthetic data improves coverage until it starts to dominate the real distribution, at which point the model gets better at the generator's world and worse at yours. The ratio, the filters and the taxonomy are all decisions to make by measuring on real data.

What to do next

  1. Write a one-page task specification and collect even a small set of real examples; hold out a third as the test set.
  2. Draft a taxonomy of skills, domains and difficulty, add personas or real documents for grounding, and count the cells.
  3. Pilot 300 to 500 generations, log every drop with its reason, and compute yield per filter stage.
  4. Add the strongest verifier your task allows: execution, schema validation or agreement between independent models.
  5. Train with a mix of real and synthetic data, sweep the ratio, and evaluate per taxonomy slice on real held-out data only.
  6. Turn the weakest slices into new taxonomy cells or filter fixes, and record provenance for every example you keep.
Key takeaway: Synthetic data lets a small model learn tasks where real data is scarce, but the hard part is diversity, not volume. Make generation conditional on sampled attributes such as skill, domain, difficulty and persona; combine Self-Instruct, Evol-Instruct, document grounding, backtranslation and programmatic generation; and filter cheapest-first with execution or agreement checks wherever possible. Watch that filters do not delete the hard cases, decontaminate against every evaluation set, keep real data in every mix to avoid collapse, and judge the model only on real held-out examples. Then let the weakest slices tell you what to generate next.