Dropout is the simplest regulariser in deep learning: during training, zero each activation independently with probability p and rescale the survivors. It is a few lines of code, yet it touches almost every hard part of training a transformer: the difference between training and evaluation behaviour, random number generation across thousands of GPUs, memory for masks, fused attention kernels, and reinforcement-learning fine-tuning where randomness in the forward pass quietly corrupts the maths.
This article builds dropout up from its definition, places each dropout site in a transformer block, shows the code and configuration fields you will actually meet, explains the systems issues, and ends with practical guidance on when to use it and when to set it to zero. It assumes you know the shape of a transformer block; if not, read transformer block anatomy first.
The definition, and why we scale
For an activation vector x, training-time dropout draws a mask m whose entries are 1 with probability 1 - p and 0 with probability p, and outputs y = m * x / (1 - p). At evaluation time it outputs y = x. This is inverted dropout. The original 2014 formulation instead left training outputs unscaled and multiplied weights by the keep probability at test time; both give the same expected activation, but inverted dropout puts all the work in training so the inference graph needs no change at all.
The scaling keeps the expectation fixed: each element survives with probability 1 - p and is then divided by 1 - p, so E[y] = x. Without it, the next layer would see activations about 1 - p times smaller in training than in inference, and every layer after it would be calibrated to the wrong scale.
What dropout does change is the variance. Each output element is x / (1 - p) with probability 1 - p and zero otherwise, so its variance is x^2 * p / (1 - p). At p = 0.1 that is about 0.11 times x^2; at p = 0.5 it equals x^2. This noise is the regulariser: the network cannot rely on any single unit being present, so it learns redundant, distributed features, and the training objective approximates averaging over an exponential number of thinned sub-networks that share weights.
A worked example
Take x = [2, 4, 6, 8] and p = 0.5. Suppose the mask comes out [1, 0, 0, 1]. The output is [4, 0, 0, 16]: the survivors are doubled. The sum is 20, exactly the input's sum, but that is luck; another mask, [0, 1, 1, 1], gives [0, 8, 12, 16] with a sum of 36. Averaged over all sixteen equally likely masks, each element's mean is its input value. In evaluation mode the output is [2, 4, 6, 8], matching that average.
Now the transformer case. A residual stream element of magnitude 1.0 with p = 0.1 becomes either about 1.11 or 0. Summed over the 4,096 dimensions of a hidden state, the relative noise in the total is small, which is why a small p behaves like gentle, structured noise rather than destruction of the signal.
import torch
def dropout(x: torch.Tensor, p: float, training: bool) -> torch.Tensor:
if not training or p == 0.0:
return x
if p == 1.0:
return torch.zeros_like(x)
keep = torch.rand_like(x) >= p # True with probability 1 - p
return x * keep / (1.0 - p)
x = torch.tensor([2.0, 4.0, 6.0, 8.0])
samples = torch.stack([dropout(x, 0.5, True) for _ in range(100_000)])
print(samples.mean(0)) # close to [2, 4, 6, 8]
print(samples.var(0)) # close to [4, 16, 36, 64] = x^2 * p / (1 - p)
Where dropout sits in a transformer
The 2017 transformer paper describes two dropout sites, both with P_drop = 0.1 for the base model: residual dropout, applied to the output of each sub-layer before it is added to the residual stream, and dropout on the sums of embeddings and positional encodings. Most implementations add a third site that the paper does not describe: attention dropout, applied to the softmax probabilities before they weight the values. Some also drop the FFN's hidden activation.
The field names you meet in configuration files reflect these sites. GPT-2's configuration in Hugging Face Transformers has embd_pdrop, attn_pdrop and resid_pdrop, each defaulting to 0.1. BERT's has hidden_dropout_prob and attention_probs_dropout_prob. Llama's configuration has attention_dropout with a default of 0.0, and no residual dropout field at all. That last default reflects a broad shift: models trained on enormous datasets for about one pass rarely overfit in the classic sense, so many recent open configurations ship with dropout off, and dropout reappears mainly in fine-tuning.
Dropout on attention probabilities deserves a note. After masking, a row of attention weights no longer sums to 1 for any individual sample; it sums to 1 only in expectation. The model learns not to depend on any single attended position, which helps on small datasets, but the effect is stronger than its p suggests because one dropped probability can remove an entire token's contribution to that head.
import torch.nn as nn
import torch.nn.functional as F
class Block(nn.Module):
def __init__(self, d, n_heads, p_resid=0.1, p_attn=0.1):
super().__init__()
self.ln1, self.ln2 = nn.LayerNorm(d), nn.LayerNorm(d)
self.qkv, self.proj = nn.Linear(d, 3 * d), nn.Linear(d, d)
self.ffn = nn.Sequential(nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d))
self.drop = nn.Dropout(p_resid) # follows model.train() / model.eval()
self.p_attn, self.h = p_attn, n_heads
def attn(self, x):
b, t, d = x.shape
q, k, v = self.qkv(x).view(b, t, 3, self.h, d // self.h).permute(2, 0, 3, 1, 4)
y = F.scaled_dot_product_attention(
q, k, v, is_causal=True,
dropout_p=self.p_attn if self.training else 0.0) # SDPA ignores self.training
return self.proj(y.transpose(1, 2).reshape(b, t, d))
def forward(self, x):
x = x + self.drop(self.attn(self.ln1(x)))
x = x + self.drop(self.ffn(self.ln2(x)))
return xThe comment on dropout_p marks the most common dropout bug in modern PyTorch code. nn.Dropout checks the module's training flag; the functional F.scaled_dot_product_attention does not. Pass a constant there and your model drops attention probabilities during evaluation and generation too, producing noisy, slightly worse outputs that no unit test on shapes will catch.
Variants that drop more than a unit
- DropPath (stochastic depth) zeros an entire residual branch for a given example, scaling surviving branches by
1 / (1 - p). The block becomes the identity for that example. It is common in vision transformers, often withpincreasing linearly with depth. - LayerDrop drops whole layers during training so that, at inference, you can remove layers to trade accuracy for speed without retraining.
- DropConnect drops individual weights rather than activations. It is rarely used in transformers because generating a fresh weight mask per example is expensive.
- Token and word dropout remove or replace input tokens, which regularises against over-reliance on specific words.
- LoRA dropout: adapter libraries such as PEFT expose a
lora_dropoutsetting that applies dropout to the adapter's input path only, leaving the frozen base model untouched.
def drop_path(x, p, training):
if not training or p == 0.0:
return x
shape = (x.shape[0],) + (1,) * (x.ndim - 1) # one decision per example
keep = torch.rand(shape, device=x.device) >= p
return x * keep / (1.0 - p)
# inside a block: x = x + drop_path(self.attn(self.ln1(x)), self.p_path, self.training)
Systems: masks, kernels and random numbers
Mask memory. A non-fused implementation must keep the mask for the backward pass. For residual dropout that is one byte per hidden element, tolerable. For attention dropout it is one byte per attention probability: with batch 8, 32 heads and sequence length 4,096, that is 8 × 32 × 4,096 × 4,096, about 4.3 billion bytes per layer. Fused attention kernels in the FlashAttention family avoid this by never storing the mask: they generate it from a counter-based random generator (Philox) using a seed and offset, and regenerate the identical mask during the backward pass. See FlashAttention for why the probability matrix itself is never materialised.
Activation checkpointing. Recomputing a checkpointed segment in the backward pass must reproduce the same dropout masks, or the gradients belong to a different network from the one whose loss you computed. PyTorch's torch.utils.checkpoint saves and restores the random state by default (preserve_rng_state=True); turning that off for speed silently breaks training with dropout.
Data parallelism. Each data-parallel rank should draw different masks; seeding every rank identically correlates the noise across the global batch and weakens the regularisation.
Tensor parallelism. Here the rule splits. Where activations are sharded across ranks, each rank must use a different random stream, or every shard sees the same pattern. Where activations are replicated (for example after an all-reduce, on the residual stream), every rank must draw the same mask, or the replicas diverge and the model is no longer one model. Megatron-LM handles this with a tracker that holds separate CUDA random states for the two cases. See tensor parallelism for which tensors are sharded where.
Dropout at inference: Monte Carlo dropout
Leaving dropout on at inference and running several forward passes gives a cheap estimate of model uncertainty: the spread of the predictions approximates a Bayesian posterior over sub-networks (Gal and Ghahramani, 2016). It costs one forward pass per sample and works only for models trained with dropout.
@torch.no_grad()
def mc_predict(model, x, samples=20):
model.eval() # LayerNorm etc. in eval mode
for m in model.modules():
if isinstance(m, nn.Dropout):
m.train() # re-enable only the dropout layers
probs = torch.stack([model(x).softmax(-1) for _ in range(samples)])
return probs.mean(0), probs.var(0) # prediction and its uncertaintyNote that in the Block above the attention dropout keys off the block's own training flag, so this helper re-enables residual dropout only. Decide explicitly which sites you want active.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Validation loss noisy and worse than expected | Dropout active at evaluation, often via SDPA's dropout_p | Gate every functional dropout on self.training; test eval determinism |
| Generation differs run to run with fixed sampling | Same as above, at inference | Two identical greedy decodes must match |
| Training loss plateaus high | p too large for the data and model size | Lower it, or set it to 0 for large-data pretraining |
| Gradients wrong after enabling checkpointing | RNG state not restored on recompute | Keep preserve_rng_state=True |
| Tensor-parallel replicas drift apart | Different masks on replicated activations | Use the framework's model-parallel RNG tracker |
| PPO ratios not 1 on the first update | Dropout makes old and new log-probabilities disagree | Disable dropout in policy and reference models for RL fine-tuning |
The reinforcement-learning case is worth spelling out. PPO-style methods compare the probability the current policy assigns to an action with the probability recorded during the rollout. With dropout active, the same weights give different probabilities on every forward pass, so the ratio is noisy before any learning has happened and clipping fires on noise. Common RLHF practice is to set all dropout to zero for this stage.
When to use it, and how much
Use dropout when the model can memorise the data: fine-tuning on thousands to hundreds of thousands of examples, many epochs, or small models trained from scratch. Start at 0.1 on residual and attention sites and tune between 0.0 and 0.3 against a held-out set. For pretraining on a corpus seen roughly once, start at 0.0; the data itself is the regulariser, and dropout would only slow convergence. When you fine-tune a model pretrained without dropout, adding it at 0.05 to 0.1 is reasonable, but measure; abruptly changing the activation noise a model was trained under can cost accuracy at first. Related regularisers interact: weight decay, data augmentation and early stopping overlap with dropout, and stacking them all at strong settings usually underfits. The gradient-flow view in the vanishing gradient problem explains why residual paths tolerate this noise so well.
What to do next
- List every dropout site in your model and its probability, including functional calls inside attention.
- Write a test that runs the model twice in eval mode on the same input and asserts identical outputs.
- Pick probabilities by regime: 0.0 for single-pass pretraining, around 0.1 for fine-tuning, tuned on a validation set.
- If you use activation checkpointing or tensor parallelism, confirm the RNG handling in your framework rather than assuming it.
- Turn dropout off for RL fine-tuning and for reference models used in KL penalties.
- If you need uncertainty estimates, try MC dropout with 10 to 30 samples before building an ensemble.