Synthetic data is cheap to produce and expensive to trust. Any generator, from a statistical model to a large language model, will happily emit millions of plausible rows. The hard questions are whether those rows preserve what matters about the real data, whether a model trained on them works on real inputs, and whether they leak the people the real data describes.

Doing it right means treating synthetic data like a release candidate: it ships only after passing tests measured against real data the generator never saw. This page builds that discipline end to end on a worked example, with code for each gate. How to make LLM-generated instruction data diverse, and the filter stack for it, are covered in synthetic training data generation; the mathematics of model collapse is in the maths of synthetic training data.

Advertisement

When synthetic data helps, and when it cannot

Synthetic data earns its keep in four situations. Sharing under privacy constraints: a vendor or another team needs data with the shape of production without the people in it. Rebalancing: a rare class, such as fraud at about one percent, needs more examples for the model to learn its boundary. Coverage: test fixtures, edge cases and inputs that have not happened yet. Teaching a skill: instruction or reasoning data generated by a stronger model and checked by a verifier, as in model distillation.

It cannot create information the generator does not have. A generator fitted to your data reproduces what it learned, minus whatever it failed to learn; it does not discover new fraud patterns. And it does not make data anonymous by being synthetic: a generator that memorises can reproduce real records exactly. Both limits explain why every gate below compares against real data.

The pipeline, and the rule that makes it honest

Synthetic data done right: split first, generate from train only, gate on real holdoutsreal data200,000 claimstrain splitonly input to generatorutility holdoutnever seen by generatorprivacy holdoutsame size as referencegeneratorcopula, CTGAN, LLMfitsynthetic rowstagged with provenancesamplefidelity gatemarginals, correlations, AUCutility gateTSTR vs TRTRprivacy gateDCR share, exact copiescontamination gaten-gram overlap vs evalstest setbaselinerelease with a data cardversioned, reproducible, revocable
Holdouts are cut from the real data before the generator is fitted. Every gate measures synthetic data against real rows the generator never saw.

The single most common mistake is to fit the generator on all the real data and then evaluate on a sample of that same data. The synthetic rows then look excellent on every test, because the evaluation set leaked into the generator. The rule: split first, fit the generator on the training split only, and keep two real holdouts untouched, one to measure utility and one to provide a privacy baseline.

Split by entity, not by row. If one customer has twenty claims and some land in training and some in the holdout, a generator that memorised that customer will look like it generalises. Group the split by the entity whose privacy you are protecting.

import numpy as np
from sklearn.model_selection import GroupShuffleSplit

def three_way_split(X, y, groups, seed=0):
    # 60% train, 20% utility holdout, 20% privacy holdout, grouped by customer
    g1 = GroupShuffleSplit(n_splits=1, test_size=0.4, random_state=seed)
    tr, rest = next(g1.split(X, y, groups))
    g2 = GroupShuffleSplit(n_splits=1, test_size=0.5, random_state=seed)
    u, p = next(g2.split(X[rest], y[rest], groups[rest]))
    return tr, rest[u], rest[p]
Advertisement

The worked example and a generator you can reason about

An insurer has 200,000 motor claims with 14 numeric columns, including claim amount, vehicle age, days from policy start to claim and a binary fraud label at roughly one percent. They want to give a modelling vendor a synthetic copy and to test whether synthetic fraud rows help their own classifier.

To keep every step visible, use a Gaussian copula. It keeps each column's own distribution exactly and models the dependence between columns as a correlation matrix in normal-score space. Library synthesizers such as CTGAN or a copula with categorical support are better in practice; the gates are identical whichever you use.

from scipy import stats

class GaussianCopula:
    def fit(self, X):                                    # X: (n, d) numeric
        n = X.shape[0]
        self.cols = [np.sort(X[:, j]) for j in range(X.shape[1])]
        U = stats.rankdata(X, axis=0) / (n + 1)          # ranks to (0, 1)
        Z = stats.norm.ppf(U)                            # normal scores
        self.corr = np.corrcoef(Z, rowvar=False)
        return self

    def sample(self, m, rng):
        d = len(self.cols)
        Z = rng.multivariate_normal(np.zeros(d), self.corr, size=m)
        U = stats.norm.cdf(Z)
        # inverse empirical CDF per column
        return np.column_stack([np.quantile(c, U[:, j]) for j, c in enumerate(self.cols)])

Notice what it captures and what it cannot. It reproduces each marginal and the monotone pairwise associations. It misses interactions: if fraud is likely only when the claim is large and early in the policy, a single correlation matrix blurs that into two weak correlations. It treats the label as a column, so the synthetic labels are rounded continuous values. It also inverts the empirical CDF, so every sampled value lies within the observed range and the extremes are real values from real claims. That last property matters for privacy, as we will see.

Gate 1: fidelity

Fidelity asks whether synthetic rows look like real rows. Check three levels, from cheapest to strongest.

from scipy.stats import ks_2samp
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.model_selection import cross_val_score

ks = [ks_2samp(real[:, j], syn[:, j]).statistic for j in range(real.shape[1])]
corr_gap = np.abs(np.corrcoef(real, rowvar=False) - np.corrcoef(syn, rowvar=False)).max()

n = min(len(real), len(syn))                       # balanced real-vs-synthetic task
Xd = np.vstack([real[:n], syn[:n]])
yd = np.r_[np.zeros(n), np.ones(n)]
disc_auc = cross_val_score(HistGradientBoostingClassifier(), Xd, yd,
                           cv=5, scoring="roc_auc").mean()

The Kolmogorov-Smirnov statistic per column is the largest gap between the two cumulative distributions; near zero is good. The correlation gap catches broken pairwise structure. The discriminator is the strongest test: train a classifier to tell real from synthetic. An AUC near 0.5 means it cannot; an AUC near 1.0 means there is a giveaway, and the classifier's feature importances tell you which columns carry it. For the copula on the claims data, expect small KS values by construction and a discriminator that finds the interactions the copula flattened.

Turn a failing discriminator into a fix. Suppose it reaches an AUC of 0.8 and its top feature is days from policy start. Plot that column against claim amount for real and synthetic rows: the real data has a dense cluster of large claims in the first month, which the copula has spread thin. Either model the interaction explicitly, for example with separate copulas per policy-age segment, or move to a synthesizer that learns conditional structure, then re-run every gate. Never tune the generator against the utility or privacy holdouts; tune on fidelity measured against training data and keep the holdouts for the final decision, or they stop being holdouts.

Gate 2: utility, measured on real data

Utility asks whether synthetic data does the job. The standard protocol is train on synthetic, test on real (TSTR), compared with train on real, test on real (TRTR), both scored on the utility holdout.

from sklearn.metrics import roc_auc_score, average_precision_score

def score(Xtr, ytr, Xte, yte):
    m = HistGradientBoostingClassifier().fit(Xtr, ytr)
    p = m.predict_proba(Xte)[:, 1]
    return roc_auc_score(yte, p), average_precision_score(yte, p)

trtr = score(X_train, y_train, X_util, y_util)
tstr = score(X_syn, y_syn, X_util, y_util)
aug  = score(np.vstack([X_train, X_syn_fraud]), np.r_[y_train, y_syn_fraud], X_util, y_util)

Read the gap, not the absolute number. With a one percent positive rate, report average precision alongside AUC, because AUC can look healthy while precision on the rare class is poor. The augmentation run answers the second question directly: does adding synthetic fraud rows to real training data improve the real holdout score? If it does not, the synthetic minority rows are not teaching anything the real ones did not.

Decide the acceptance thresholds before running, and write them down. A starting policy for sharing might be that TSTR average precision is within ten percent of TRTR. There is no universal standard; the point is that the threshold is set before anyone sees the number.

Gate 3: privacy, against a baseline

Privacy asks whether synthetic rows are too close to real individuals. A raw nearest-neighbour distance means nothing on its own, because real data is clustered and some synthetic rows will be near real ones by chance. Compare against a baseline: how close are synthetic rows to the training rows versus to the privacy holdout, which the generator never saw?

from sklearn.neighbors import NearestNeighbors
from sklearn.preprocessing import StandardScaler

rng = np.random.default_rng(0)
ref = X_train[rng.choice(len(X_train), size=len(X_priv), replace=False)]
sc = StandardScaler().fit(X_train)
nn_tr = NearestNeighbors(n_neighbors=1).fit(sc.transform(ref))
nn_ho = NearestNeighbors(n_neighbors=1).fit(sc.transform(X_priv))
d_tr = nn_tr.kneighbors(sc.transform(X_syn))[0][:, 0]
d_ho = nn_ho.kneighbors(sc.transform(X_syn))[0][:, 0]

closer_to_train = np.mean(d_tr < d_ho)     # about 0.5 if nothing was memorised

# copies must be searched for in the whole training split, not the subsample
nn_all = NearestNeighbors(n_neighbors=1).fit(sc.transform(X_train))
exact_copies = np.mean(nn_all.kneighbors(sc.transform(X_syn))[0][:, 0] == 0)

For the share test, the training reference is subsampled to the holdout's size so both sets are equally dense; the exact-copy rate searches all training rows, because a copy of a row outside the subsample would otherwise go unnoticed. If the generator learned the distribution rather than the records, synthetic rows are as close to unseen people as to training people, and the share sits near one half. A share well above one half means the generator is copying. Look at the rows with the smallest distance, too: outliers such as a single unusually large claim are where re-identification risk concentrates, and the copula above copies extreme column values verbatim.

This test detects copying; it does not prove privacy. If you need a guarantee, for example to share with an external party under regulation, use a generator trained with differential privacy and record its privacy budget. Then run the same gates anyway, because differential privacy bounds leakage while usually reducing fidelity and utility.

LLM-generated data: contamination and provenance

For text generated by a language model the same gates apply in adapted form: a discriminator for fidelity, a held-out human-written test set for utility. Two further checks matter. First, decontamination: the generating model may have seen your evaluation benchmarks, so its output can contain their questions and answers. Check every generated example for long n-gram overlap with each evaluation set and drop matches; the GPT-3 paper used 13-gram overlap for this. Otherwise your evaluation, as described in LLM evaluation architecture, measures memory rather than skill.

Second, provenance: tag every synthetic row with the generator, its version, the prompt or seed and the date, and keep synthetic and real data in separate, versioned datasets. Six months later someone will ask why the model degraded, and the answer may be that a later training run consumed synthetic output from an earlier model. ML data versioning shows how to make that lineage queryable.

Failure modes

  • Evaluating on the generator's own training data. Every metric looks excellent. Split first, always.
  • Mode loss in the tail. Rare categories and rare combinations vanish or are smoothed away. Compare per-segment counts, not only global statistics.
  • A healthy average hiding a weak class. Global AUC on synthetic-trained models can match the real baseline while precision on the rare class collapses. Report per-class metrics on the real holdout.
  • Copied outliers. The privacy share looks fine on average while the five most unusual real records reappear almost exactly. Inspect the closest rows by hand.
  • Contaminated benchmarks. Scores jump after adding synthetic data because the evaluation answers leaked in.
  • Recursive training. Training on your own model's output across generations narrows the distribution. Keep real data in every mix and track diversity.

Operating it

Run the gates as code in CI, not as a one-off notebook. Store the split indices, generator configuration, random seeds and every gate result next to the synthetic dataset, and publish a short data card: what real data it came from, which columns were dropped, the gate scores and the thresholds they were judged against, and what the data must not be used for. Regenerate when the source schema or distribution changes, and expire old synthetic releases the way you would expire old credentials. Access control still applies: a dataset derived from regulated data keeps its review obligations until the privacy gate and your privacy team agree it can be released.

The trade-off throughout is the same triangle. More fidelity usually means more memorisation risk; stronger privacy usually costs utility. You choose a point on that triangle per use case: test fixtures need little fidelity and no privacy risk, vendor sharing needs strong privacy and moderate utility, and rare-class augmentation needs utility above all.

What to do next

  1. Pick one dataset and write the three-way grouped split before choosing any generator.
  2. Write down acceptance thresholds for fidelity, utility and privacy for that use case.
  3. Fit the simple copula as a baseline, run all three gates, and keep the numbers.
  4. Try a stronger synthesizer and accept it only if it beats the baseline on utility without worsening the privacy share.
  5. For LLM-generated data, add an n-gram decontamination pass against every evaluation set you report.
  6. Tag synthetic rows with provenance and store them in a separate versioned dataset with a data card.
Key takeaway: Synthetic data is a release candidate, not a free resource. Split real holdouts off before the generator sees anything, fit on the training split only, and accept the output only if it passes fidelity, train-on-synthetic-test-on-real utility and a baselined privacy test, with thresholds set in advance. For LLM-generated data, also decontaminate against your benchmarks and keep provenance so you can always tell real from synthetic.