Most writing about synthetic data for LLMs is about what to generate: how to design diverse prompts, which filters to apply and how to avoid model collapse. Those questions are covered on this site in synthetic training data generation and synthetic data done right. This article is about the other half, which decides whether the project finishes on time and on budget: running generation as a GPU workload.

A synthetic data job is large-scale offline inference followed by scoring and deduplication. It has different economics from serving. Nobody waits for a single response, so latency barely matters and throughput is everything. Most of the generated text is thrown away, so the real unit of cost is an accepted sample. And jobs run for hours or days on many GPUs, often preemptible ones, so they must survive interruption. The sections below cover the cost model, how to lay models out on GPUs, how to drive vLLM for offline batches, how to place the judge, how to deduplicate on GPU, and the failure modes that waste the most GPU-hours.

The job as a GPU workload

Generation with a decoder-only model has two phases per request. Prefill processes the prompt in one parallel pass and is compute-bound. Decode produces one token per step per sequence and is bound by memory bandwidth: every step reads the model weights and the sequence's key-value cache. Serving engines get throughput by batching many sequences into each decode step, so the weights are read once for the whole batch.

Synthetic data jobs lean heavily on decode. A typical instruction-response sample has a prompt of a few hundred to a few thousand tokens, much of it shared (system prompt, rubric, few-shot examples), and an output of several hundred tokens or more. Long chain-of-thought outputs push this further. That has three consequences. First, output tokens, not input tokens, drive cost. Second, the batch size that matters is how many sequences fit in KV cache memory at once, because that sets how many tokens each decode step produces. Third, a shared prompt prefix is pure waste if it is recomputed for every request, which is why prefix caching matters so much here.

Architecture of a generation job

A resumable synthetic data job on a GPU fleetSeed promptssharded JSONL in storageManifestshard -> done / pendingPhase 1: generationteacher replicas, TP=k eachoffline batched decodingprefix cache for shared promptn samples per promptclaim shardRaw outputsone file per shardwritemark done after durable writePhase 2: scoringrule checks, then judge orreward model (prefill only)Phase 3: dedupembeddings on GPUAccepted setwith provenance per rowCost meterGPU-hours per 1k accepted
Figure 1. Generation, scoring and dedup run as separate phases over the same sharded data. A manifest makes every phase restartable, and the metric that matters is GPU-hours per thousand accepted rows, not per thousand generated.

Figure 1 shows the job split into phases. Seed prompts live as shards in object storage, with a manifest recording which shards each phase has finished. Generation replicas claim pending shards, write raw outputs to one file per shard, and mark the shard done only after the write is durable. Scoring runs cheap rule checks first and then a judge or reward model on what survives. Dedup embeds the survivors and removes near-duplicates. Every row in the accepted set keeps its provenance: seed ID, teacher model and version, sampling parameters and scores.

Keeping the phases separate, instead of generating and judging in one loop, is usually the right default. Each phase can use a different model and GPU count, a bug in scoring does not force regeneration, and the GPU memory of a replica is spent on one model's weights and KV cache rather than split between two.

Budgeting by accepted sample

Budget from the accepted count backwards. If you need A accepted rows and the pipeline accepts a fraction r of generations, you need A divided by r generations. Multiply by output tokens per generation to get output tokens, divide by the measured output throughput of one replica to get replica-hours, multiply by GPUs per replica, and add the scoring overhead.

def plan(target_accepted, accept_rate, out_tokens, in_tokens,
         replica_out_tok_s, gpus_per_replica, judge_share, usd_per_gpu_hour):
    generations = target_accepted / accept_rate
    out_total = generations * out_tokens
    gen_replica_hours = out_total / replica_out_tok_s / 3600
    gen_gpu_hours = gen_replica_hours * gpus_per_replica
    total_gpu_hours = gen_gpu_hours * (1 + judge_share)
    cost = total_gpu_hours * usd_per_gpu_hour
    return {
        "generations": round(generations),
        "output_tokens_B": round(out_total / 1e9, 3),
        "input_tokens_B": round(generations * in_tokens / 1e9, 3),
        "gpu_hours": round(total_gpu_hours),
        "usd": round(cost),
        "usd_per_1k_accepted": round(cost / target_accepted * 1000, 2),
    }

# Assumptions, not measurements: 2,500 output tok/s per 2-GPU replica,
# $3 per GPU-hour, judge pass adds 25% on top of generation GPU time.
for rate in (0.6, 0.35, 0.15):
    print(rate, plan(200_000, rate, 700, 900, 2500, 2, 0.25, 3.0))

Running it prints:

0.6 {'generations': 333333, 'output_tokens_B': 0.233, 'input_tokens_B': 0.3, 'gpu_hours': 65, 'usd': 194, 'usd_per_1k_accepted': 0.97}
0.35 {'generations': 571429, 'output_tokens_B': 0.4, 'input_tokens_B': 0.514, 'gpu_hours': 111, 'usd': 333, 'usd_per_1k_accepted': 1.67}
0.15 {'generations': 1333333, 'output_tokens_B': 0.933, 'input_tokens_B': 1.2, 'gpu_hours': 259, 'usd': 778, 'usd_per_1k_accepted': 3.89}

The throughput and price inputs are placeholders; replace them with numbers you measure on your own model, GPUs and prompt lengths before committing a budget. The shape of the result does not depend on them: cost scales with the inverse of the acceptance rate, so going from 60 percent to 15 percent acceptance quadruples the bill. Improving prompts and cheap pre-filters to raise acceptance is usually worth more than any engine tuning.

Laying models out on GPUs

Each replica needs enough GPU memory for the weights plus a KV cache large enough to batch well. Weights take roughly parameters times bytes per parameter: about 2 bytes in BF16 and about 1 in FP8, so a 70-billion-parameter model needs about 140 GB in BF16 and about 70 GB in FP8 before any cache. KV cache per token is 2 (keys and values) times layers times KV heads times head dimension times bytes per element.

For a Llama-3-70B-shaped model, with 80 layers, 8 KV heads under grouped-query attention and a head dimension of 128, that is 2 x 80 x 8 x 128 x 2 bytes, about 320 KiB per token in BF16. A sequence of 900 prompt tokens and 700 output tokens holds about 1,600 tokens, or roughly 0.5 GB. Every 10 GB of free cache therefore holds about 20 such sequences at once, and that concurrency directly sets decode throughput.

This gives the layout rule. Use the smallest tensor-parallel degree that fits the weights with comfortable room for cache, then add independent replicas for throughput. Tensor parallelism splits every layer across GPUs and pays an all-reduce on every layer, which is cheap over NVLink within a node and expensive across nodes. Independent replicas share nothing, scale almost linearly and fail independently. For an offline job, eight GPUs as four TP=2 replicas usually beats one TP=8 replica, as long as each replica still has enough cache to batch deeply. How vLLM spends that memory is explained in vLLM on GPU.

Driving vLLM for offline batches

vLLM's offline API runs batches without an HTTP server. The engine schedules everything you pass it with continuous batching, so the simplest high-throughput pattern is to hand it a whole shard at once.

import json, pathlib
from vllm import LLM, SamplingParams

llm = LLM(model="your-org/teacher-model", tensor_parallel_size=2,
          gpu_memory_utilization=0.90, max_model_len=4096)
tok = llm.get_tokenizer()
SYSTEM = pathlib.Path("system_prompt.txt").read_text()   # shared prefix, cached

def run_shard(shard_path, out_path, shard_index):
    rows = [json.loads(l) for l in open(shard_path, encoding="utf-8")]
    prompts = [tok.apply_chat_template(
                   [{"role": "system", "content": SYSTEM},
                    {"role": "user", "content": r["seed"]}],
                   tokenize=False, add_generation_prompt=True) for r in rows]
    params = SamplingParams(n=4, temperature=0.9, top_p=0.95,
                            max_tokens=1024, seed=1000 + shard_index)
    outputs = llm.generate(prompts, params)
    tmp = out_path.with_suffix(".tmp")
    with open(tmp, "w", encoding="utf-8") as f:
        for r, out in zip(rows, outputs):
            for k, c in enumerate(out.outputs):
                f.write(json.dumps({"seed_id": r["id"], "k": k, "text": c.text,
                                    "finish": c.finish_reason}) + "\n")
    tmp.replace(out_path)    # atomic rename: a shard is done only when complete

Several details carry weight. The shared system prompt comes first and is identical across requests, so prefix caching can reuse its KV blocks instead of recomputing them; check that it is enabled in your vLLM version. Asking for n samples per prompt shares one prefill across all of them. The seed varies per shard so that a re-run reproduces a shard, but different shards do not draw identical samples for repeated seed prompts. The finish reason is stored because outputs cut off at max_tokens must not be scored as complete. Constrained JSON decoding is available in vLLM when outputs must follow a schema, but the parameter names have changed between releases, so take them from the documentation for the version you pin.

Scoring and where the judge runs

Scoring is cheaper per token than generation, because a judge that outputs a short verdict, or a reward model that outputs a single scalar, is mostly prefill: one parallel pass over prompt plus response, which is compute-bound and batches well. Order the scoring pipeline by cost. Run rules first on CPU: length, format, language, refusal phrases, finish reason, banned strings and decontamination against your evaluation sets. Then run a small classifier or reward model. Use a large LLM judge only on what is left.

Run the judge as its own phase on the same fleet after generation, rather than co-locating it with the teacher on each GPU. Loading both models on one replica halves the memory available for the teacher's KV cache and therefore its throughput. Using a judge from a different model family than the teacher reduces the risk that it rewards the teacher's own stylistic habits. The training and pitfalls of reward models are covered in reward model training on GPU.

Deduplication on GPU

Sampling the same seeds several times at high temperature produces near-duplicates, which waste training compute and over-weight some patterns. Exact hashing catches only identical strings. Embedding-based dedup catches paraphrases: embed every accepted response with a small encoder on GPU, normalise, and drop any row whose cosine similarity to an already-kept row exceeds a threshold. The greedy loop below is written with NumPy for clarity; on a GPU the same logic runs with torch tensors, and at millions of rows you would use an approximate nearest-neighbour index such as FAISS instead of comparing against every kept row.

import numpy as np

def near_dup_mask(emb, threshold=0.92, block=4096):
    """Greedy keep-first dedup on L2-normalised embeddings, blockwise."""
    emb = emb / np.linalg.norm(emb, axis=1, keepdims=True)
    keep = np.ones(len(emb), dtype=bool)
    kept_rows = []
    for start in range(0, len(emb), block):
        chunk = emb[start:start + block]
        for i, v in enumerate(chunk):
            idx = start + i
            if kept_rows and (emb[kept_rows] @ v).max() >= threshold:
                keep[idx] = False
                continue
            kept_rows.append(idx)
    return keep

rng = np.random.default_rng(0)
base = rng.normal(size=(1000, 64))
dups = base[:200] + rng.normal(scale=0.05, size=(200, 64))
mask = near_dup_mask(np.vstack([base, dups]))
print("kept:", int(mask.sum()), "dropped:", int((~mask).sum()))

On this synthetic test, 1,000 random vectors plus 200 lightly perturbed copies, it keeps 1,000 and drops 200, and every dropped row is one of the copies. Tune the threshold on a labelled sample of your own data: too low removes legitimately different answers to similar questions, too high lets paraphrases through. Pretraining-scale methods such as MinHash are covered in data filtering, dedup and mixture.

Worked example: 200,000 accepted pairs

A team wants 200,000 accepted instruction-response pairs to fine-tune a small model, using a 70B-class teacher in FP8 on eight 80 GB GPUs. FP8 weights take about 70 GB, so TP=2 leaves roughly 90 GB of the pair's 160 GB, minus activation and runtime overheads, for KV cache, and the eight GPUs become four replicas. A pilot of 2,000 seeds measures throughput per replica, acceptance rate and output length. Suppose the pilot shows a 35 percent acceptance rate and 700 output tokens on average. With the placeholder inputs above, that is about 571,000 generations, 0.4 billion output tokens and about 111 GPU-hours, roughly 14 hours of wall-clock time on eight GPUs.

The pilot also shows that a third of rejections are truncations at max_tokens and another large share fail a format rule. Raising max_tokens for the long-answer task type, and adding a format example to the shared prefix, lifts acceptance to around 60 percent in the next pilot, which cuts the plan to about 65 GPU-hours. That change was free; no engine tuning would have saved as much.

Failure modes

  • Lost work on preemption. A job without a manifest restarts from zero when a spot instance disappears. Shard small enough that losing one costs minutes.
  • Truncated outputs accepted. Responses cut at max_tokens look fine to a lenient judge. Filter on finish reason first.
  • Identical seeds everywhere. A fixed seed for every request, combined with repeated seed prompts, yields exact duplicates and a falsely high acceptance rate.
  • Cache starvation. A max_model_len much larger than real sequences, or two models on one GPU, leaves too little KV cache, and throughput collapses while the GPU still reports high utilisation.
  • Idle GPUs during CPU stages. Tokenisation, JSON parsing and rule filters on a single CPU thread can leave expensive GPUs waiting; parallelise them or move them off the GPU nodes.
  • Contamination and terms. Seeds or outputs overlapping evaluation sets inflate scores, and some model licences restrict using outputs to train other models. Check both before generating, not after.

Trade-offs

A bigger teacher gives better samples and a higher acceptance rate but fewer tokens per GPU-hour, so compare models on cost per accepted sample from a pilot, not on benchmark scores. FP8 halves weight memory, freeing that space for KV cache, and usually raises throughput, at a small quality risk to measure on your task. Higher temperature and more samples per prompt improve diversity but add near-duplicates and lower acceptance. Hosted batch APIs often win for small jobs.

What to do next

  1. Write down the accepted-row target and run a pilot of a few thousand seeds to measure throughput, acceptance rate and output length.
  2. Compute GPU-hours and cost per thousand accepted rows from the pilot, and fix the biggest rejection cause before scaling up.
  3. Choose the smallest tensor-parallel degree that leaves ample KV cache, and fill the rest of the GPUs with independent replicas.
  4. Put the shared prompt first, request n samples per prompt, vary the seed per shard and store finish reasons.
  5. Shard the seeds, add a manifest with atomic writes, and test that killing a replica mid-shard loses only that shard.
  6. Order scoring from cheap rules to judge, run the judge as its own phase, and use a different model family where you can.
  7. Add embedding dedup with a threshold tuned on labelled pairs, and keep provenance on every accepted row.
Key takeaway: Treat synthetic data generation as offline inference whose unit of cost is an accepted sample. Measure throughput and acceptance in a pilot, raise acceptance before scaling, use the smallest tensor-parallel degree that leaves room for KV cache and scale with replicas, share prefixes and prefill, run judging and dedup as separate phases, and make every shard restartable.