Every model report ends in a number: 94% accuracy, 0.81 AUC, a 12% drop in error. That number is only worth something if it predicts what the model will do on data it has never seen. The train, validation and test split is the mechanism that makes the prediction honest. Get it wrong and the number is optimistic, sometimes wildly so, and nobody finds out until the model is in production and the dashboards disagree with the report.

This article treats splitting as a design problem rather than a call to train_test_split: what each set is for, the principle that decides how to split, how information leaks across the boundary, how big the test set must be, and how to keep it honest over months of experiments, with a worked example and a checklist.

Advertisement

Three sets, three jobs

Three sets, three jobs: fit, select, estimateAll labelled datadedupe, assign groupsSplit by unitrow / entity / timeTest setlocked, read rarelyTraining setfit weights + preprocessingValidation setearly stop, HPO, thresholdsChosen configfrozen before testFinal estimatemetric + intervalmodelbest by valone evaluationInformation may flow from test to the final report only. Any decision made afterlooking at test scores turns the test set into a second validation set.
Training data fits parameters, validation data chooses among models, and the test set produces one final estimate. Arrows show the only allowed flows of information.

The training set is what the learning algorithm sees. Weights, tree splits, and the statistics inside preprocessing steps (means, vocabularies, imputation values) are all fitted here. Performance on training data measures how well the model memorised and says almost nothing about generalisation.

The validation set is for decisions a human or a search procedure makes: which architecture, which learning rate, when to stop training, which classification threshold, which features to keep. Every decision taken by looking at validation scores leaks a little of the validation set into the model, so after many experiments the validation score is itself optimistic. That is acceptable, because the final number does not come from it.

The test set is touched once, after every decision is frozen, to estimate how the chosen model will perform. Its value depends entirely on that discipline. If you look at the test score, change something and look again, the test set has become a second validation set and its number inherits the same optimism.

When data is scarce, k-fold cross-validation replaces the single validation set by rotating it through the training data; the Spark cross-validation article covers fold mechanics in detail. The test set still stays outside the rotation.

The one principle: imitate deployment

Every splitting decision follows from one question: when this model is used for real, how will the new data differ from the data it was trained on? The test set must differ from the training set in the same way. A random split answers the question only if production data really is a fresh random draw from the same pool, which is rarer than it looks.

Three differences are common. New entities: a fraud model will score customers it has never seen, a medical model will see new patients, a recommender will meet new users. New time: every production prediction is about the future, and the world drifts. New sources: a vision model may be deployed to a hospital, camera or region that contributed nothing to training. If the split does not reproduce the relevant difference, the test score measures an easier problem than the one you will face.

Advertisement

Choosing the split unit

SplitWhen it is rightWhat it protects againstCost
Random by rowRows are genuinely independent and production draws from the same poolNothing beyond basic overfittingCheapest; optimistic if rows are correlated
StratifiedRare classes, small dataTest sets that by chance contain too few positivesNone; use it by default for classification
Group (entity)Several rows per user, patient, document, session or deviceThe model recognising the entity instead of learning the taskFewer effective samples; needs a reliable group id
TemporalAny model that predicts the futureUsing future information and ignoring driftOlder training data; test reflects one period
Temporal with gapLabels resolve over a window (churn in 30 days, chargebacks)Training labels that overlap the test periodDiscards the gap's data
Source held outDeployment to new sites, regions or devicesShortcut features tied to the sourceNeeds several sources; high variance

These combine. A churn model is usually split by time first and checked for customer overlap second. A clinical model is split by patient and often by hospital. For classification, StratifiedGroupKFold keeps groups intact while balancing class ratios across folds, and TimeSeriesSplit accepts a gap argument for rolling-origin evaluation. For a single hold-out, the helpers below are enough:

import numpy as np
from sklearn.model_selection import GroupShuffleSplit, train_test_split

def split_by_group(df, group_col, test=0.15, val=0.15, seed=7):
    """Every row of one customer lands in exactly one set."""
    g = df[group_col].values
    outer = GroupShuffleSplit(n_splits=1, test_size=test, random_state=seed)
    trval_idx, test_idx = next(outer.split(df, groups=g))
    trval = df.iloc[trval_idx]
    inner = GroupShuffleSplit(n_splits=1, test_size=val / (1 - test), random_state=seed)
    tr_idx, val_idx = next(inner.split(trval, groups=trval[group_col].values))
    return trval.iloc[tr_idx], trval.iloc[val_idx], df.iloc[test_idx]

def split_by_time(df, ts_col, val_start, test_start, gap="14D"):
    """Train < val < test in time, with a gap so label windows cannot overlap."""
    gap = np.timedelta64(int(gap[:-1]), "D")
    train = df[df[ts_col] < np.datetime64(val_start) - gap]
    val = df[(df[ts_col] >= np.datetime64(val_start)) & (df[ts_col] < np.datetime64(test_start) - gap)]
    test = df[df[ts_col] >= np.datetime64(test_start)]
    return train, val, test

def assert_disjoint(train, val, test, key):
    a, b, c = set(train[key]), set(val[key]), set(test[key])
    assert not (a & b or a & c or b & c), "entity appears in more than one split"

The assert_disjoint check is cheap and catches the most common bug: the group column you split on is not the identity the model can actually recognise. Splitting by order id when one customer places many orders leaks the customer.

Leakage: how information crosses the boundary

Leakage is any path by which information from validation or test data, or from the future, reaches the model during training or selection. It always inflates scores and it rarely produces errors, which is why it survives review. The usual paths are these.

  • Preprocessing fitted on everything. Scaling, imputation, target encoding, tokenizer vocabulary or PCA fitted before the split carries test statistics into training. Target encoding is the worst offender because it puts label information directly into a feature.
  • Duplicates and near-duplicates. The same support ticket, image or sentence appears in both sets, sometimes with a different id, a resized image or a changed timestamp. The model gets credit for recall.
  • Group leakage. Rows from one entity on both sides of the boundary, as above.
  • Temporal leakage. Features computed with data from after the prediction time, such as a customer's lifetime spend computed today and attached to a row from last year.
  • Target leakage. A feature that is a consequence of the label: a refund flag in a fraud model, a discharge code in a readmission model. It is available in the warehouse but not at prediction time.
  • Selection leakage. Choosing features, thresholds or the random seed by looking at test scores.

The structural fix for the first path is to make preprocessing part of the model, so it is refitted on whatever data is used for training and only applied to the rest:

from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

pre = ColumnTransformer([
    ("num", Pipeline([("imp", SimpleImputer(strategy="median")),
                      ("sc", StandardScaler())]), NUMERIC),
    ("cat", OneHotEncoder(handle_unknown="ignore"), CATEGORICAL),
])
model = Pipeline([("pre", pre), ("clf", LogisticRegression(max_iter=1000))])

model.fit(train[FEATURES], train["label"])     # medians, means, vocab learned from train only
val_score = model.score(val[FEATURES], val["label"])

Inside cross-validation the same pipeline object is refitted per fold, which is the only way to get unbiased fold scores when preprocessing learns anything. For duplicates, normalise and fingerprint before splitting, then assign every fingerprint cluster to one side. A crude check after the fact looks like this; for images or long text, perceptual hashes or MinHash catch near-duplicates that exact hashing misses.

import hashlib, re

def norm(text):
    text = text.lower()
    text = re.sub(r"\d+", "0", text)            # order ids, dates, amounts
    return re.sub(r"\W+", " ", text).strip()

def fingerprint(text):
    return hashlib.sha1(norm(text).encode()).hexdigest()

seen = {fingerprint(t): "train" for t in train["text"]}
leaks = [t for t in test["text"] if fingerprint(t) in seen]
print(f"{len(leaks)} of {len(test)} test rows duplicate a training row after normalisation")

How big must the test set be?

A test score is an estimate with sampling error, and the error shrinks with the square root of the test set size. For a metric that is a proportion, such as accuracy, the 95% confidence interval half-width is approximately 1.96 * sqrt(p * (1 - p) / n).

Worked example: a classifier at about 90% accuracy on 1,000 test examples has a half-width of 1.96 x sqrt(0.09 / 1000), which is about 1.9 percentage points. The honest report is 90% plus or minus 2. To narrow that to plus or minus 0.5 points you need n = 1.962 x 0.09 / 0.0052, roughly 13,800 examples. Rare-class metrics are worse: recall on a class with 2% prevalence in a 1,000-row test set rests on about 20 positives, and its interval is enormous.

Two practical consequences follow. First, size the test set by the smallest difference you need to detect, not by a habitual 20%. With a million rows, 1% for test is plenty; with 2,000 rows, even 30% leaves wide intervals, and repeated cross-validation plus a modest hold-out is the better design. Second, when comparing two models, compare them on the same test examples with a paired method, such as McNemar's test on the examples where they disagree or a bootstrap of the per-example score difference. Paired comparisons cancel the shared difficulty of the examples and detect much smaller differences than two independent intervals.

import numpy as np

def paired_bootstrap(correct_a, correct_b, n_boot=10_000, seed=0):
    """correct_a, correct_b: 0/1 arrays over the same test examples."""
    rng = np.random.default_rng(seed)
    diff = np.asarray(correct_b, float) - np.asarray(correct_a, float)
    idx = rng.integers(0, len(diff), size=(n_boot, len(diff)))
    boots = diff[idx].mean(axis=1)
    return diff.mean(), np.percentile(boots, [2.5, 97.5])

Keeping the test set honest over time

The test set wears out. Each look at it, even without changing anything, invites a decision, and a team that evaluates fifty candidates on the same test set and ships the best has selected on noise. This is adaptive overfitting, and public benchmarks suffer from it at scale; the evaluation frameworks article covers benchmark contamination for large language models.

  • Lock it. Store the test set in a location the training job cannot read, with a recorded hash. Version it with the data, as in data versioning, so every report names the exact test set.
  • Log every access. A test evaluation should be an explicit, recorded act with a commit hash, not a line in a notebook.
  • Budget looks. One evaluation per release candidate is a good rule. Model selection happens on validation.
  • Refresh on a schedule. Collect a new test set from recent production traffic every quarter or release. A fresh set both resets adaptive overfitting and measures drift, which drift detection monitors continuously.

Worked example: a transaction fraud model

Suppose you have 18 months of card transactions, about 4 million rows from 300,000 customers, with 0.4% labelled fraud. Chargebacks arrive up to 45 days after the transaction. The first instinct, a stratified random 70/15/15 split, produces a superb validation AUC. It is wrong in three ways: rows interleave in time, so a customer's later transactions teach the model about earlier ones; features such as 'transactions in the last 30 days' were computed on today's warehouse and include the future; and the most recent month's labels are incomplete.

The redesigned split is temporal. Months 1 to 13 are training, months 14 and 15 are validation, and months 16 and 17 test, with a 45-day gap before each boundary; the final 45 days wait until labels mature. Rolling features are recomputed point-in-time, using only events before each transaction's timestamp, ideally through a feature store that serves the same definitions online. Customers seen in training are allowed in the test set, because production will also score returning customers, but the test report breaks results down into returning and new customers so the gap between them is visible.

The test set has about 400,000 transactions and roughly 1,600 frauds, so recall at a fixed alert rate carries an interval of a few points, reported alongside it. Validation AUC drops from the random-split number. That drop is not a regression; it is the leakage being removed.

Failure modes

SymptomLikely causeFix
Validation far above productionGroup or temporal leakage, or point-in-time features computed lateSplit by entity or time; rebuild features as of prediction time
Test score improves every release while production does notAdaptive overfitting on a reused test setLock and refresh the test set; select on validation
Large score swings between seedsSmall test set or rare classReport intervals; enlarge test or use repeated CV
One feature dominates importanceTarget leakageCheck whether the feature exists at prediction time
Near-perfect scores on text or imagesDuplicates across setsFingerprint and cluster before splitting
Different results in notebook and pipelinePreprocessing fitted outside the modelMove every fitted transform into a Pipeline

Trade-offs

Stricter splits cost data and comfort: group and temporal splits shrink the effective training set and lower the headline number, larger test sets starve training, and refreshing costs labelling. The trade is almost always worth it. When in doubt, build both the random and the deployment-shaped split and report the gap, which measures how much shortcut learning the random split allowed.

What to do next

  1. Write one sentence describing how production data will differ from training data, and choose the split unit from it.
  2. Add an assert_disjoint check on the true entity id, and a gap equal to the label window for temporal splits.
  3. Move every fitted transform, including target encoding, into a pipeline that is fitted on training data only.
  4. Fingerprint records and remove or co-locate duplicate clusters before splitting.
  5. Compute the confidence interval your test set size gives you, and size it by the smallest difference you must detect.
  6. Lock and version the test set, log every evaluation, compare models with a paired method, and schedule a refresh from recent traffic.
Key takeaway: Training data fits the model, validation data chooses among models, and the test set gives one final estimate after every decision is frozen. Choose the split so the test set differs from training the way production will: by entity, by time with a gap for label windows, or by source. Keep preprocessing inside the model, remove duplicates before splitting, size the test set from the confidence interval you need, compare models on paired examples, and protect the test set from repeated looks so its number keeps meaning something.