A model overfits when it fits patterns that exist in the sample it was trained on but not in the data it will meet later. Training error keeps falling while error on new data stops falling or rises. The model has spent its capacity on noise, quirks of collection, and individual examples instead of the regularities that carry over.
That definition sounds simple, and most courses stop at a plot of training and validation loss. In practice overfitting hides in three places: in the weights, in the choices you make by repeatedly checking a validation set, and in evaluation data that is not as independent as you think. This article teaches you to measure each one. Theory of the bias-variance decomposition lives in the bias-variance tradeoff in depth and the remedies in regularization in depth; here the focus is diagnosis, with code you can run.
A precise definition
Every supervised model is trained to minimise loss on a finite sample, the empirical risk. What you care about is the expected loss on the distribution the model will see in use, the true risk. The difference between the two is the generalization gap. A model overfits when it lowers empirical risk in ways that raise, or fail to lower, true risk.
Two assumptions sit underneath every train, validation and test split: examples are drawn independently, and from the same distribution as production data. Overfitting is a failure of the model given those assumptions. Leakage and drift are failures of the assumptions themselves. They produce the same headline symptom, a model that does worse in use than in evaluation, but they need different fixes, so the first job of diagnosis is to tell them apart.
Capacity matters because a model with more free parameters, deeper trees or more training steps can fit more arbitrary functions, including the noise. Data matters because noise averages out as samples grow: the same model that memorises 30 points may generalise well from 30,000.
Worked example: watching the gap open
Fit polynomials of increasing degree to 30 noisy samples of a sine curve, then score them on 1,000 fresh samples from the same process. The noise has variance 0.09, so no model can get validation mean squared error below about 0.09.
import numpy as np
rng = np.random.default_rng(3)
f = lambda x: np.sin(3 * x)
x_tr = rng.uniform(-1, 1, 30)
y_tr = f(x_tr) + rng.normal(0, 0.3, 30) # noise variance 0.09
x_va = rng.uniform(-1, 1, 1000)
y_va = f(x_va) + rng.normal(0, 0.3, 1000)
P = np.polynomial.polynomial
for degree in [1, 3, 5, 9, 15]:
coef = P.polyfit(x_tr, y_tr, degree)
mse_tr = np.mean((P.polyval(x_tr, coef) - y_tr) ** 2)
mse_va = np.mean((P.polyval(x_va, coef) - y_va) ** 2)
print(f"degree {degree:2d} train {mse_tr:.3f} val {mse_va:.3f} gap {mse_va - mse_tr:.3f}")| Degree | Train MSE | Validation MSE | Gap | Reading |
|---|---|---|---|---|
| 1 | 0.229 | 0.237 | 0.008 | Underfit: both errors high, small gap |
| 3 | 0.082 | 0.094 | 0.012 | Close to the noise floor |
| 5 | 0.081 | 0.093 | 0.012 | Best validation score |
| 9 | 0.068 | 0.111 | 0.043 | Training improves, validation gets worse |
| 15 | 0.046 | 17.667 | 17.621 | Overfit: wild oscillation between training points |
Three lessons come out of this table. First, training error alone always rewards the bigger model; it fell at every step. Second, the gap is the signal, and it grows slowly before it explodes: at degree 15 the curve passes close to every training point and swings wildly near the edges of the interval, where training points are sparse. Third, with 30 points the gap itself is noisy. The seed above was picked because it shows the textbook pattern cleanly. With seed 0 the same code reports a validation error of 0.343 for the straight line, a gap of 0.141, purely from an unlucky draw. Before you call a model overfit, repeat the measurement over several seeds or folds and look at the spread, not one number.
Reading the symptoms
The gap between training and validation metrics, tracked over training time and model size, is the main diagnostic. Learning curves, which plot both against training-set size, add the question that decides what to do next: would more data help? The patterns below cover most cases.
| What you see | Most likely meaning | Next step |
|---|---|---|
| Train and validation both poor, small gap | Underfitting or a label/feature problem | More capacity, better features, check labels |
| Train excellent, validation much worse, gap grows with epochs | Classic overfitting | Early stopping, regularization, more data |
| Validation fine, held-out test much worse | Overfit the validation set by selection | Fresh holdout, fewer comparisons, nested CV |
| Test fine, production much worse from day one | Leakage or a train/serve mismatch | Audit splits and feature pipelines |
| Production good at launch, degrades over months | Drift, not overfitting | Monitor inputs, retrain on recent data |
| Validation better than training | Dropout or augmentation active only in training, or an easy validation set | Evaluate training data in eval mode |
A useful stress test for memorisation comes from Zhang et al., "Understanding deep learning requires rethinking generalization" (ICLR 2017): standard image networks reached zero training error even when every label was replaced at random. Shuffle the labels of a copy of your training set and train on it. If your model fits random labels almost as quickly as real ones, it has the capacity to memorise your data, so training accuracy says almost nothing about what it learned and only held-out measurement counts.
Overfitting the validation set
Every time you look at a validation score and change something, the validation set takes part in training. One look is harmless. Hundreds of looks, through hyperparameter searches, architecture tweaks, feature experiments and seed picking, add up to a search process that finds whatever happens to score well on those particular examples. The simulation below makes the effect concrete: a thousand "models" that are pure coin flips, scored on 200 validation examples.
import numpy as np
rng = np.random.default_rng(1)
n_val, n_test, n_models = 200, 100_000, 1000
y_val = rng.integers(0, 2, n_val)
y_test = rng.integers(0, 2, n_test)
# Every "model" is a coin flip: true accuracy is exactly 50%.
best_val = 0.0
for _ in range(n_models):
acc = (rng.integers(0, 2, n_val) == y_val).mean()
best_val = max(best_val, acc)
fresh = (rng.integers(0, 2, n_test) == y_test).mean()
print(f"best validation accuracy {best_val:.3f}; same kind of model on fresh data {fresh:.3f}")It prints a best validation accuracy of 0.630 and a fresh-data accuracy of 0.503. Thirteen points of apparent skill came entirely from selection. Real searches are less extreme because candidate models are correlated and genuinely better than chance, but the direction never changes: the winner's score is biased upwards, and more so with smaller validation sets and more candidates.
Defences are procedural rather than mathematical:
- Keep a locked test set. Score on it once per release decision, never during tuning, and record every time it is opened.
- Use nested cross-validation for small datasets. An inner loop picks hyperparameters, an outer loop estimates the error of the whole selection procedure. Spark ML cross-validation in depth shows fold design for grouped and time-ordered data.
- Budget comparisons. Report how many configurations were tried. If a sweep had 500 trials, treat a 0.3-point win as noise until it reproduces on fresh data.
- Refresh holdouts. When a validation set has guided months of work, retire it and collect a new one.
Leakage: the overfitting you cannot see
Leakage means information from the evaluation data, or from the future, reaches the model during training. Scores look excellent, and the gap between training and validation looks healthy, because validation is no longer independent. The common forms are:
- Preprocessing fit on everything. Scaling, imputation, vocabulary building or feature selection computed on the full dataset before splitting.
- Duplicates and near-duplicates. The same document, image or customer appears on both sides of the split. Web-scraped corpora are full of them.
- Group leakage. Several rows per patient, user or device, split at random, so the model recognises the entity instead of learning the pattern.
- Temporal leakage. Random splits on time-ordered data let the model train on the future and test on the past.
- Target leakage. A feature that is only known after the outcome, such as a refund flag used to predict fraud.
The structural fix is to put every fitted step inside the cross-validation loop and to split by the unit that will be new in production:
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.feature_selection import SelectKBest
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import GroupKFold, TimeSeriesSplit, cross_val_score
model = make_pipeline(StandardScaler(), SelectKBest(k=50), LogisticRegression(max_iter=1000))
# Rows belong to users; a user must never appear on both sides of a split.
scores = cross_val_score(model, X, y, groups=user_id, cv=GroupKFold(n_splits=5))
# Time-ordered data: always train on the past, evaluate on the future.
scores_t = cross_val_score(model, X_sorted, y_sorted, cv=TimeSeriesSplit(n_splits=5))If the grouped or time-ordered score is much worse than the random-split score, the random-split number was leakage, and the lower number is your real estimate.
Deep networks and fine-tuning
Large neural networks complicate the textbook picture. They often have far more parameters than training examples and can memorise their data, yet still generalise. Test error can even fall, rise and fall again as model size or training time grows, the double descent behaviour covered in the double descent deep dive. The practical consequence is that parameter count is a poor predictor of overfitting. Measure the gap; do not infer it.
Fine-tuning a pretrained language model on a small dataset is where most teams meet overfitting today. Typical signs are evaluation loss that bottoms out after one or two epochs and then rises while training loss keeps falling, outputs that copy phrasings from training examples, and lost general ability on tasks outside the fine-tuning set. Practical controls:
- Evaluate on a held-out slice every few hundred steps and keep the checkpoint with the best evaluation score, not the last one.
- Treat epochs as a hyperparameter with a small range; small instruction datasets rarely need many passes.
- Keep a general-capability evaluation alongside the task evaluation so you notice when the model forgets.
- Check benchmark contamination: search the training data for long n-gram overlaps with evaluation items, and drop or report the overlapping items.
For evaluation design beyond a single metric, see AI evaluation frameworks.
Fixing it, briefly
Once you know which kind of overfitting you have, the remedy follows. Weight-level overfitting responds to more or more varied data, augmentation, smaller models, weight decay, dropout and early stopping; the regularization article has a reproducible comparison of these. Validation overfitting needs fresh holdouts and fewer comparisons, not a new regulariser. Leakage needs pipeline and split fixes; no amount of regularization repairs an evaluation that has seen the answers. Do not reach for a fix before the diagnosis, because a regulariser applied to a leakage problem will make the reported score slightly worse and the real problem no better.
Operating models in production
Overfitting that slipped through shows up after launch as a step change: the production metric is worse than the test metric from the first day. Drift shows up as a slope. Log predictions with enough context to compute the offline metric on live traffic, compare the two on the same definition, and alert when the gap exceeds the spread you measured across folds. Keep a small, regularly refreshed, labelled sample of production data as the final arbiter; it is the only dataset that has never informed a decision.
What to do next
- Report every metric as a mean and spread over at least five folds or seeds, not one number.
- Plot training and validation metrics on the same axes for every run, against epochs and against training-set size.
- Run the shuffled-label test once on your current model to learn how much it can memorise.
- Move all preprocessing into a pipeline and rerun cross-validation with group or time-based splits; compare to the random split.
- Deduplicate across splits with exact hashes and a near-duplicate check.
- Lock a test set, log each access, and count the configurations you tried before quoting a win.
- For fine-tunes, select checkpoints by held-out evaluation loss and track a general-capability benchmark.
- After launch, compare production and test metrics on identical definitions and alert on a step change.