A model that fits its training data perfectly and fails on new data has learned the noise along with the signal. Regularization is the set of techniques that stop that from happening. Each one makes the model slightly worse at fitting the training set in exchange for being better on data it has not seen: penalties on weights, optimizer changes, data changes, noise inside the network, and when to stop.

This article builds the idea from first principles, works a reproducible example, and covers the techniques that matter in practice, including the L2-versus-weight-decay distinction, how to tune strength, and a checklist. Dropout gets one section here because it has its own deep dive in Dropout architecture.

Advertisement

What regularization actually does

When the space of possible models is large compared with the data, many models fit the training set equally well and most generalise badly. The expected error on new data can be split into three parts: bias (error from a model family too simple to capture the truth), variance (error from the fitted model changing a lot with the particular sample you trained on) and irreducible noise. Regularization trades a small increase in bias for a larger decrease in variance.

There are two equivalent ways to see it. As a constraint, regularization shrinks the effective search space: weights must stay small, or few, or the function must be smooth with respect to its inputs. As a prior, regularization encodes a belief about which models are likely before seeing data. L2 corresponds to a Gaussian prior on the weights and L1 to a Laplace prior, and minimising the regularised loss is maximum a posteriori estimation under that prior. A stronger penalty is a tighter prior, appropriate when you have less data.

The symptom regularization treats is a generalisation gap: training loss much lower than validation loss, with validation loss flat or rising while training loss keeps falling. That diagnosis depends on a clean validation set, so first check that your split is sound, as described in Train / Validation / Test Split. Leakage hides overfitting.

Where each regularizer acts in a training loopDataaugmentation, mixupModeldropout, stochastic depthLossL2, L1, label smoothingOptimizer stepdecoupled weight decaynext batchValidation loopearly stopping, checkpoint selectionevery N stepsWhat each one constrainsSize of the weightsL2, weight decay, max-normNumber of active weightsL1, elastic netSensitivity to inputsaugmentation, mixup, noiseCo-adaptation of unitsdropoutTraining timeearly stopping
Regularizers act at different points in the loop: on the data, inside the model, in the loss, in the optimizer step, and in the validation loop that decides when to stop and which checkpoint to keep.

L2 regularization from first principles

Ridge regression adds lambda times the squared norm of the weights to the squared error. Setting the gradient to zero gives a closed form: w equals the inverse of (X-transpose X + lambda I) times X-transpose y. Compared with ordinary least squares, the only change is lambda added to the diagonal. It makes the matrix invertible even with collinear features, and it shrinks the solution along the directions the data barely constrains: in the eigenbasis of X-transpose X, a direction with eigenvalue s is scaled by s divided by (s + lambda). Well-determined directions are almost untouched; poorly determined ones, where the variance came from, are pushed towards zero.

In gradient descent the same penalty adds 2 lambda w to the gradient, so every step pulls each weight towards zero in proportion to its size. That is where the name weight decay comes from. For the update rule itself, see Gradient Descent, in depth. Do not penalise the intercept, and standardise features first: the penalty treats all weights alike, so millimetres and metres are penalised very differently.

Advertisement

Worked example: ridge on a degree-12 polynomial

Take fifteen noisy samples of sin(pi x) on the interval from minus one to one, with Gaussian noise of standard deviation 0.2, and fit a degree-12 polynomial: thirteen weights for fifteen points. Then measure mean squared error on two hundred fresh points as the penalty grows. The script below produces the table exactly, with NumPy's default generator seeded at 0. Change nothing if you want the same numbers.

import numpy as np

rng = np.random.default_rng(0)

def make(n):
    x = rng.uniform(-1, 1, n)
    return x, np.sin(np.pi * x) + rng.normal(0, 0.2, n)   # noise variance 0.04

xtr, ytr = make(15)          # 15 training points (drawn first)
xva, yva = make(200)         # 200 validation points

def feats(x, deg=12):
    return np.vander(x, deg + 1, increasing=True)          # 1, x, x^2, ..., x^12

Xtr, Xva = feats(xtr), feats(xva)
for lam in [0, 1e-6, 1e-4, 1e-3, 1e-2, 1e-1, 1, 10]:
    I = np.eye(Xtr.shape[1]); I[0, 0] = 0                    # do not penalise the intercept
    if lam:
        w = np.linalg.solve(Xtr.T @ Xtr + lam * I, Xtr.T @ ytr)
    else:
        w = np.linalg.lstsq(Xtr, ytr, rcond=None)[0]
    tr = np.mean((Xtr @ w - ytr) ** 2)
    va = np.mean((Xva @ w - yva) ** 2)
    print(f"lam={lam:g} train={tr:.4f} val={va:.4f} |w|={np.linalg.norm(w[1:]):.1f}")
lambdatrain MSEvalidation MSEweight norm
00.002036.65938074.4
1e-60.00960.374598.0
1e-40.01150.328410.0
1e-30.01290.13345.0
1e-20.01550.07123.8
1e-10.03930.09552.1
10.11980.29330.8
100.21570.53330.2

Read the table from top to bottom. Without a penalty, the fit passes almost through every training point (train MSE 0.002, well below the noise variance of 0.04, so it is fitting noise). The weights reach norms in the thousands, the polynomial swings wildly between points, and validation error is 36.7. A tiny penalty of one in a million cuts validation error by about a hundred times, because it removes the near-degenerate directions. Validation error is lowest at lambda = 0.01, at 0.0712. That is close to the noise floor of 0.04, the best any model could do. Beyond that, training and validation error both rise together: the model is now underfitting.

Lessons: training error always rises with the penalty, so it cannot choose lambda; the useful range spans orders of magnitude, so search on a log scale; and a huge weight norm is a cheap overfitting signal.

L1 and elastic net: sparsity

L1 regularization penalises the sum of absolute values of the weights instead of their squares. Its gradient has the same magnitude for every non-zero weight, so it pushes small weights all the way to zero rather than just shrinking them. The optimisation step that does this is soft thresholding: move each weight towards zero by a fixed amount and clamp it at zero if it would cross. The result is a sparse model, one that has selected a subset of features.

L1 has a weakness with correlated features: it tends to pick one of a correlated group arbitrarily, and which one can change between data samples. Elastic net mixes L1 and L2 so that correlated features are kept or dropped together, while still producing sparsity.

from sklearn.linear_model import Lasso, ElasticNet
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

# Lasso objective: (1 / (2 n)) * ||y - Xw||^2 + alpha * ||w||_1
lasso = make_pipeline(StandardScaler(), Lasso(alpha=0.05))
enet = make_pipeline(StandardScaler(), ElasticNet(alpha=0.05, l1_ratio=0.5))
lasso.fit(X_train, y_train)
kept = (lasso[-1].coef_ != 0).sum()          # how many features survived

Library conventions differ, so check which way the strength parameter runs. In scikit-learn, alpha in Ridge, Lasso and ElasticNet is a penalty strength, where larger means stronger. C in LogisticRegression and SVC is the inverse, where larger means weaker.

L2 is not weight decay under Adam

With plain SGD, adding an L2 penalty to the loss and shrinking the weights directly in the update are the same thing up to a rescaling of lambda. With adaptive optimisers such as Adam they are not. Adam divides each parameter's gradient by a running estimate of that gradient's magnitude. If the L2 term is added to the gradient, it is divided too, so parameters with large historical gradients are barely regularised and parameters with small ones are regularised heavily. Loshchilov and Hutter's AdamW fixes this by decoupling the decay: the weights are shrunk by learning rate times weight decay times w directly, outside the adaptive scaling. In PyTorch, torch.optim.AdamW implements the decoupled form. By default, torch.optim.Adam with a weight_decay argument adds the L2 term to the gradient, which is the coupled form.

import torch

decay, no_decay = [], []
for name, param in model.named_parameters():
    if not param.requires_grad:
        continue
    # biases and normalisation gains only shift or scale; decaying them hurts more than it helps
    (no_decay if param.ndim == 1 or name.endswith(".bias") else decay).append(param)

optimizer = torch.optim.AdamW(
    [{"params": decay, "weight_decay": 0.1},
     {"params": no_decay, "weight_decay": 0.0}],
    lr=3e-4,
)

Two operational notes. Because decoupled decay is multiplied by the learning rate, a learning-rate schedule also changes the effective regularization: decay is strongest when the learning rate is high. So when you change the learning rate, retune weight decay as well. And the PyTorch default for AdamW is weight_decay=0.01, which is a default, not a recommendation for your model; values around 0.1 are common in large language model pre-training, and much smaller ones for fine-tuning.

Early stopping

Training for longer lets the model fit increasingly fine detail, eventually noise. Stopping when validation loss stops improving is regularization by limiting training time, and for linear models with gradient descent it behaves much like an L2 penalty whose strength falls as training goes on. Evaluate at a fixed interval, keep the best checkpoint rather than the last, and use patience long enough to ride out noisy evaluations.

best, best_step, patience, bad = float("inf"), 0, 5, 0
for step in range(max_steps):
    train_step(model, next(train_batches))
    if step % eval_every == 0:
        val = evaluate(model, val_loader)
        if val < best - 1e-4:                  # require a real improvement
            best, best_step, bad = val, step, 0
            save_checkpoint(model, "best.pt")
        else:
            bad += 1
            if bad >= patience:
                break
model = load_checkpoint("best.pt")             # never ship the last step by default

Data augmentation, mixup and label smoothing

Some of the most effective regularizers do not touch the weights at all. Data augmentation creates new training examples by applying transformations that should not change the label: crops, flips and colour jitter for images, time stretching and noise for audio, back-translation or synonym replacement for text. It encodes an invariance directly and often helps more than any penalty, provided every transformation truly preserves the label: rotating a 6 into a 9 teaches the wrong thing.

Mixup trains on convex combinations of pairs of examples and their labels: lambda times one input plus (1 minus lambda) times another, with the label mixed the same way, and lambda drawn from a Beta(alpha, alpha) distribution. Small alpha values such as 0.2 are typical. It smooths the function between training points.

Label smoothing replaces one-hot targets with a mixture: most of the probability on the true class and a small epsilon spread across all classes. In PyTorch it is one argument: nn.CrossEntropyLoss(label_smoothing=0.1). It changes how the model's confidence behaves, and that interacts with calibration. If downstream code treats scores as probabilities, re-check calibration afterwards, as described in Model calibration architecture.

Dropout and noise inside the network

Dropout zeroes a random fraction of activations during training and rescales the rest, so no unit can rely on any specific other unit being present. Stochastic depth drops whole residual blocks. Both suit large networks on modest data; in large-scale pre-training with few epochs dropout is often off, because there is little overfitting to prevent.

Tuning the strength

Every regularizer has a strength, and the right value depends on data size, model size and the other regularizers in use. Search on a log scale, as the worked example showed. Tune the strongest levers first, one at a time. Regularizers partly substitute for each other, so adding dropout after tuning weight decay often calls for less decay. More data is the strongest regularizer of all, so expect to reduce regularization as your dataset grows. When the search space has more than two or three dimensions, use a proper search procedure rather than a hand grid; Hyperparameter Optimization Architecture covers the tooling.

TechniqueBest whenWatch out for
L2 / weight decayalmost always; the defaultcoupled L2 under Adam; decaying biases and norms
L1 / elastic netmany features, few matter, need sparsityunstable selection among correlated features
Early stoppingany iterative trainingnoisy validation, shipping the last step
Augmentationknown invariances, limited datatransformations that change the label
Label smoothingclassification with overconfident outputschanges calibration and distillation targets
Dropoutlarge model, modest dataforgetting eval mode at inference

Failure modes

  • Tuning on the test set. It inflates the reported score; keep an untouched test set.
  • Over-regularizing. Training and validation loss both high and close together means bias, not variance. Reduce regularization or increase capacity.
  • Unscaled features with L1 or L2. The penalty silently favours features with large numeric ranges.
  • Hidden double regularization. A framework default (weight decay in an optimizer config, dropout in a pretrained model's config) stacked on top of your own. Print the effective configuration before training.
  • Regularizing what is really leakage or drift. A validation gap caused by different data distributions will not close with penalties. Check the split and the data first.

What to do next

  1. Run the worked example, then change the noise level or the number of training points and see how the best lambda moves.
  2. Plot training and validation loss for your current model on one chart and decide whether you have a variance problem or a bias problem.
  3. If you use Adam, confirm you are on decoupled weight decay (AdamW) and that biases and normalisation parameters are excluded.
  4. Add early stopping with best-checkpoint saving if training does not already have it.
  5. Sweep the main regularization strength over at least four orders of magnitude on a log scale, and record the curve, not just the winner.
  6. List every source of regularization in the effective config, including framework defaults, and remove any you cannot justify.
Key takeaway: Regularization trades a little training-set fit for better generalisation by constraining the model: smaller weights with L2 or weight decay, fewer weights with L1, robustness to input changes with augmentation and mixup, less overconfidence with label smoothing, less co-adaptation with dropout, and less training time with early stopping. The worked example shows the pattern you will see everywhere: no penalty overfits badly, a well-chosen one gets close to the noise floor, and too much underfits. Choose the strength on a clean validation set using a log-scale search, use decoupled weight decay with adaptive optimisers, and expect to need less regularization as your data grows.