Language models memorise. Given enough repetitions, or sometimes just one unusual string, a model will reproduce training data verbatim, and attackers can test whether a specific record was in the training set. Filtering personal data out of the corpus reduces the risk, but it is a best effort with no bound. Differentially private stochastic gradient descent (DP-SGD) is the training algorithm that turns the risk into a number: a mathematical limit on how much any single training example can change what the model does.
This page explains DP-SGD from the definition up, shows a complete training step in PyTorch and the equivalent in Opacus, works through why it costs accuracy, and covers how it is applied to large language models. The attacks it defends against are covered in membership inference defenses and LLM PII leakage.
The guarantee, stated precisely
A training algorithm M is (ε, δ)-differentially private if, for any two datasets D and D′ that differ in one privacy unit, and any set S of possible trained models, Pr[M(D) ∈ S] ≤ eε · Pr[M(D′) ∈ S] + δ. In words: whatever an observer does with the model, their conclusion is almost equally likely whether or not your record was in the data. With ε = 1, adding your record can shift the odds of any outcome by at most a factor of e, about 2.7. δ is the probability the bound fails outright, so it must be much smaller than 1/N; a common rule is δ below 1/N, often far below.
The privacy unit decides what the promise means. Example-level DP protects one training example; if one user contributed 500 examples, their data is protected far less than ε suggests. User-level DP treats all of a user's data as the unit and is what most people intuitively expect. For language models the unit is often a fixed-length token sequence. Always report ε, δ and the unit together; an ε alone is not a claim.
The algorithm, piece by piece
DP-SGD (Abadi et al., 2016) changes ordinary minibatch SGD in four places.
- Poisson sampling. Each example joins the batch independently with probability q, so batch size varies around qN. The privacy analysis gets its strongest help, amplification by subsampling, from this randomness: an example that probably was not in the batch leaks little.
- Per-example gradients. Ordinary training only needs the average gradient; DP-SGD needs each example's gradient separately, because it must bound each one.
- Clipping. Each gradient is scaled to have L2 norm at most C: g ← g · min(1, C/‖g‖). Now no single example can move the sum by more than C, which is the sensitivity the noise is calibrated to. This is different from ordinary global-norm clipping, which bounds the batch gradient for stability and gives no privacy; see gradient clipping architecture for that.
- Gaussian noise. Add noise drawn from N(0, σ2C2) to every coordinate of the clipped sum, then divide by the expected batch size qN and take the optimizer step. σ is the noise multiplier.
A privacy accountant then composes the cost of every step. Given q, σ, the number of steps T and a target δ, it returns ε. Modern accountants (Rényi DP, and the PRV accountant that Opacus uses by default) give much tighter numbers than simple composition. Two corollaries matter in practice: the noise step must run even when a Poisson batch comes out empty, and the analysis assumes Poisson sampling, so shuffled fixed-size batches are not covered by the same accounting unless your accountant models them explicitly.
A DP-SGD step from scratch in PyTorch
Writing the step yourself once makes every library parameter obvious. torch.func computes per-example gradients with vmap over grad:
import torch
from torch.func import functional_call, vmap, grad
def poisson_batch(n: int, q: float) -> torch.Tensor:
return (torch.rand(n) < q).nonzero().squeeze(1) # may be empty
def dp_sgd_step(model, loss_fn, xb, yb, lr, C, sigma, expected_batch):
params = {k: v.detach() for k, v in model.named_parameters()}
def loss_one(p, x, y):
out = functional_call(model, p, (x.unsqueeze(0),))
return loss_fn(out, y.unsqueeze(0))
if len(xb) > 0:
per_ex = vmap(grad(loss_one), in_dims=(None, 0, 0))(params, xb, yb)
# one L2 norm per example, across all parameter tensors
norms = torch.sqrt(sum(g.flatten(1).pow(2).sum(1) for g in per_ex.values()))
scale = (C / (norms + 1e-6)).clamp(max=1.0)
with torch.no_grad():
for name, p in model.named_parameters():
if len(xb) > 0:
g = per_ex[name]
g = (g * scale.view(-1, *[1] * (g.dim() - 1))).sum(0)
else:
g = torch.zeros_like(p) # empty batch: noise only
g = g + torch.normal(0.0, sigma * C, size=p.shape, device=p.device)
p -= lr * g / expected_batch # expected, not realised, size
# loop: idx = poisson_batch(N, q); dp_sgd_step(model, loss_fn, X[idx], Y[idx],
# lr, C=1.0, sigma=1.1, expected_batch=q * N)Two lines carry the privacy. Clipping happens per example, before the sum. And normalisation uses the expected batch size q·N: dividing by the realised size would make the update depend on how many examples were sampled, which the analysis does not account for.
The same thing with Opacus
In practice use a maintained library. Opacus wraps a PyTorch model, optimizer and data loader, switches the loader to Poisson sampling, computes per-sample gradients with hooks, and runs the accountant:
from opacus import PrivacyEngine
from opacus.validators import ModuleValidator
model = ModuleValidator.fix(model) # e.g. BatchNorm -> GroupNorm
optimizer = torch.optim.SGD(model.parameters(), lr=0.5, momentum=0.9)
engine = PrivacyEngine(accountant="prv")
model, optimizer, train_loader = engine.make_private_with_epsilon(
module=model, optimizer=optimizer, data_loader=train_loader,
target_epsilon=3.0, target_delta=1e-6, epochs=10, max_grad_norm=1.0)
for epoch in range(10):
train_one_epoch(model, optimizer, train_loader)
print(epoch, engine.get_epsilon(delta=1e-6))make_private_with_epsilon searches for the noise multiplier that spends exactly the target budget over the requested epochs; make_private takes noise_multiplier directly and lets you read ε afterwards. ModuleValidator.fix matters because BatchNorm mixes statistics across examples in a batch, so a per-example gradient is not defined; Opacus replaces it with GroupNorm. For production, PrivacyEngine(secure_mode=True) uses a cryptographically secure noise source.
Worked example: why DP-SGD costs accuracy
Consider a model with d = 1,000,000 trainable parameters, clipping norm C = 1, noise multiplier σ = 1 and an expected batch of 256. The clipped per-example gradients sum to a vector of norm at most 256 (if they all point the same way; less otherwise). The noise vector has independent N(0, 1) coordinates, so its norm is about √d = 1,000. The noise is four times larger than the largest possible signal, before any disagreement between examples.
That arithmetic shows the three levers. Signal grows with batch size B while noise does not, so DP training uses much larger batches than usual; raising B also raises q, which costs privacy, but the net effect is usually favourable. Noise grows with √d, so training fewer parameters helps directly: fine-tuning only low-rank adapters or the last layers of a publicly pretrained model shrinks d by orders of magnitude. And C should be close to typical per-example gradient norms: too large and you add noise for nothing, too small and every gradient is truncated to its direction.
Real results reflect this. The original paper reported roughly 90%, 95% and 97% MNIST accuracy at ε of about 0.5, 2 and 8 with δ = 10-5, against about 98% without privacy. Private fine-tuning of pretrained language models loses far less than private training from scratch, because the public pretraining supplies most of the knowledge and the private step adjusts few parameters. To get the ε for your own settings, sweep them through the accountant rather than guessing: q, σ and T fully determine it for a given δ.
DP-SGD for large language models
Three things change at LLM scale. First, the privacy unit: per-example DP on a document corpus lets a document repeated a hundred times leak like a hundred examples, so deduplicate first and consider user- or document-level units. Google's VaultGemma, a 1-billion-parameter model trained from scratch with DP-SGD and released in 2025, reports a sequence-level guarantee of ε ≤ 2.0 and δ ≤ 1.1 × 10-10 for 1,024-token sequences, and treats repeated sequences as separate units, which is exactly the caveat above made explicit.
Second, memory: materialising a gradient per example for billions of parameters is impossible. Techniques such as ghost clipping compute each example's gradient norm from layer activations and output gradients without storing the per-example gradient, then do a second pass with the clipping factors applied. Parameter-efficient fine-tuning helps twice, reducing both memory and the noise dimension d.
Third, batch geometry: to get useful signal-to-noise, private LLM training uses very large batches, which usually means gradient accumulation. Accumulation must happen on clipped per-example gradients, with noise added once per logical step; adding noise per micro-batch or clipping the accumulated sum breaks the analysis. Fixed-size batches, which accelerators prefer, need a sampling scheme your accountant models, such as the truncated Poisson subsampling VaultGemma used.
Failure modes
| mistake | consequence | fix |
|---|---|---|
| BatchNorm left in the model | per-example gradients undefined; library error or leak | GroupNorm or LayerNorm |
| dividing by realised batch size | update depends on sampling outcome; analysis invalid | divide by q·N |
| shuffled fixed batches with a Poisson accountant | reported ε is not the true ε | Poisson sampling or a matching accountant |
| duplicates or multi-example users | effective protection per person far weaker than ε | deduplicate; use user-level units |
| tuning hyperparameters on the private data | each run leaks; the budget covers only the final run | tune on public proxy data or account for tuning |
| noise per micro-batch in accumulation | wrong noise scale | clip per example, noise once per logical step |
| logging per-example losses or samples | side channel outside the guarantee | log only aggregate, noised or public metrics |
| reporting ε without δ and unit | claim is meaningless | publish ε, δ, unit, accountant, q, σ, T |
Operational guidance and trade-offs
Start from a public pretrained model and privately fine-tune the smallest set of parameters that reaches your quality bar. Pick δ below 1/N and the privacy unit before anything else, because they define the claim. Set C by measuring the distribution of per-example gradient norms in a short non-private run and choosing around the median, then tune learning rate and batch size, which interact strongly under noise. Budget interpretation varies, but single-digit ε values are the range most published private models report, and smaller is stronger.
DP-SGD is not the only defense and not always the right one. Deduplication and PII scrubbing are cheap and should happen regardless. DP gives a provable bound but costs accuracy, compute and engineering; it is most justified when the training data is sensitive and the model or its outputs will be widely exposed. It also bounds, but does not replace, API-level controls against extraction; see model extraction defenses for that side of the threat model.
What to do next
- Write down the privacy unit, δ and target ε for your use case before touching code.
- Deduplicate the training data and map how many examples each user contributes.
- Reproduce a DP-SGD step with the from-scratch code on a small model to build intuition for C, σ and batch size.
- Move to Opacus or your framework's DP library with Poisson sampling, a PRV or RDP accountant and secure noise.
- Fine-tune a public pretrained model with adapters, sweeping batch size and learning rate at a fixed budget.
- Run a membership-inference test against the private and non-private models and publish ε, δ, unit and accountant settings with the model.