In 2013 a group of researchers showed that an image classifier which correctly labels a photo can be made to mislabel it by adding a perturbation too small for a person to notice. The finding looked like a curiosity about neural networks. It turned out to be a property of almost every high-dimensional learned model, and it founded a research field with its own threat models, attack algorithms, defenses and, most usefully, a long record of defenses that looked strong and were not.

This article covers those academic foundations as working knowledge rather than history. You will learn how to state a threat model precisely, how the gradient attacks FGSM and PGD work and how to implement them in a few lines of PyTorch, why optimisation attacks such as Carlini and Wagner's broke defenses that resisted simpler attacks, how attacks work without access to the model, and why the field's main lesson is about evaluation: a defense is only as good as the strongest adaptive attack anyone has run against it. The last sections connect these ideas to the token-level attacks used on language models today, because the same mistakes keep recurring there.

Why adversarial examples matter

Adversarial examples matter for two reasons. The practical one is that models now make decisions an attacker profits from changing: content filters, malware classifiers, fraud scores, face matching, and the safety layers of LLM products. The scientific one is that adversarial examples expose a gap between what we measure (accuracy on samples from the training distribution) and what we want (correct behaviour on every input a reasonable person would label the same way). A model can be 99 percent accurate on a test set and wrong on nearly every input within a tiny distance of that test set.

That gap is the reason the field distinguishes clean accuracy from robust accuracy: the fraction of test points for which no input inside an allowed neighbourhood is misclassified. Robust accuracy is always measured against a stated threat model and, unless it is certified, against a specific attack. Both qualifiers matter, as the rest of this article shows.

Threat models

A threat model answers four questions. What may the attacker change? How much may they change it? What do they know about the model? What counts as success? For images the classic answer to the first two is an Lp ball: the perturbed input x' must satisfy ||x' - x||_p <= eps. L-infinity bounds every pixel's change (8/255 per channel on CIFAR-10 is the standard benchmark budget); L2 bounds total energy; L0 bounds how many pixels change. These norms are convenient proxies for imperceptibility, not definitions of it. A rotation, a sticker or a change of lighting can be obvious to no one and huge in every Lp norm.

DimensionOptionsExample
KnowledgeWhite-box, grey-box, black-boxFull weights; architecture only; API outputs only
Output accessLogits, probabilities, top-1 labelDecision-only attacks see just the label
Perturbation setL-inf, L2, L0, semantic, physical8/255 per pixel; a printed patch
GoalUntargeted, targetedAny wrong class; specifically 'speed limit'
TimingEvasion (test time), poisoning (train time)This article covers evasion
AdaptivityStatic attack, attack aware of defenseOnly the second is a meaningful evaluation

Write the threat model down before you build or evaluate anything. Most disputes about whether a defense works turn out to be disputes about an unstated threat model, for example a defense tested against L-infinity attacks that is trivially beaten by a small rotation.

Origins: Szegedy, Goodfellow and linearity

Szegedy and colleagues' paper Intriguing Properties of Neural Networks (posted in December 2013) found adversarial examples with box-constrained L-BFGS, an optimiser searching for the smallest perturbation that produces a target label. They also noticed that many examples fooled other networks trained separately, the first sign of transferability.

Goodfellow, Shlens and Szegedy's Explaining and Harnessing Adversarial Examples (December 2014) gave the explanation that shaped the field: models are too linear, not too non-linear. For a linear score w.x, a perturbation eta = eps * sign(w) changes the score by eps * ||w||_1, which grows with the input dimension even though no single coordinate moves more than eps. Many small changes add up. The paper turned that into the Fast Gradient Sign Method: take one step of size eps in the direction of the sign of the loss gradient. It also proposed adversarial training with FGSM examples.

Kurakin, Goodfellow and Bengio then showed that iterating small sign steps finds stronger examples and that the examples survive being printed and photographed. Madry and colleagues (2017) framed robustness as a min-max problem, training to minimise the loss of the worst point in the eps ball, and made projected gradient descent (PGD) with a random start the standard first-order attack.

FGSM and PGD in code

Both attacks fit in a few lines. The code assumes inputs scaled to [0, 1], a model in eval mode, and a cross-entropy loss. Note the two clamps in PGD: one projects back into the eps ball around the original input, the other keeps the result a valid image.

import torch
import torch.nn.functional as F

def fgsm(model, x, y, eps):
    x = x.clone().detach().requires_grad_(True)
    loss = F.cross_entropy(model(x), y)
    grad, = torch.autograd.grad(loss, x)
    return (x + eps * grad.sign()).clamp(0, 1).detach()

def pgd(model, x, y, eps, alpha, steps, restarts=1):
    best = x.clone()
    best_loss = torch.full((x.shape[0],), -1e9, device=x.device)
    for _ in range(restarts):
        delta = torch.empty_like(x).uniform_(-eps, eps)          # random start
        xa = (x + delta).clamp(0, 1)
        for _ in range(steps):
            xa.requires_grad_(True)
            loss = F.cross_entropy(model(xa), y)
            grad, = torch.autograd.grad(loss, xa)
            xa = xa.detach() + alpha * grad.sign()
            xa = torch.min(torch.max(xa, x - eps), x + eps)      # project to L-inf ball
            xa = xa.clamp(0, 1)                                   # stay a valid image
        with torch.no_grad():
            per_ex = F.cross_entropy(model(xa), y, reduction="none")
            better = per_ex > best_loss
            best[better], best_loss[better] = xa[better], per_ex[better]
    return best

def robust_accuracy(model, loader, attack):
    correct = total = 0
    for x, y in loader:
        xa = attack(model, x, y)
        correct += (model(xa).argmax(1) == y).sum().item()
        total += y.numel()
    return correct / total

Typical CIFAR-10 settings are eps = 8/255, alpha = 2/255 and 10 to 50 steps. Two details decide whether the numbers mean anything: keep per-example best results across restarts, as above, and count a point as robust only if every attack you run fails on it.

Projected gradient descent: the attacker's inner loopClean input xlabel yrandom startCandidate x'inside the eps ballModel ffrozen weightsLoss L(f(x'), y)maximise itgradientStepx' + alpha * sign(grad)Projectclip to ball and [0,1]repeat K timesWhite-boxgradients from f itselfTransfergradients from a surrogateQuery-basedgradients estimated from outputsThe same loop serves all three settings; only the source of the gradient changes.
PGD repeats gradient step and projection. White-box, transfer and query attacks differ only in where the gradient comes from.

Carlini and Wagner: better losses

Carlini and Wagner (2017) replaced the loss with a margin on logits: push the correct class's logit below the best other logit by a confidence margin kappa, while minimising perturbation size. They removed the box constraint with a change of variables through tanh and searched over the trade-off constant with binary search. The margin loss matters because cross-entropy saturates: once the softmax is near one, its gradient vanishes, and a defense that inflates logits (defensive distillation was the famous case) starves a cross-entropy attack of signal while remaining just as vulnerable. Carlini and Wagner broke defensive distillation this way. The general lesson: when an attack fails, check whether its loss still produces a useful gradient before concluding the model is robust.

Black-box attacks: transfer and queries

Without weights an attacker has two routes. Transfer attacks train or borrow a surrogate model, attack it in white-box fashion and submit the result to the target. Papernot and colleagues (2016 and 2017) showed this works across architectures and even across model families, and that a surrogate can be trained on labels obtained by querying the target. Ensembles of surrogates and input diversity during the attack improve transfer.

Query-based attacks estimate the gradient from outputs. Score-based methods such as ZOO (finite differences) and NES (random-direction estimates) need probabilities; decision-based methods such as the Boundary Attack need only the top label and walk along the decision boundary. Their cost is queries, often thousands per example, which is why rate limits, output rounding and query-pattern monitoring are real if partial defenses for deployed APIs. None of them changes the white-box picture; they only raise the price.

How defenses fail: obfuscated gradients

The field's most important result is negative. Athalye, Carlini and Wagner (2018) examined the non-certified defenses accepted at ICLR 2018 that claimed white-box robustness and found that seven of nine relied on obfuscated gradients: the defense made gradients useless to the attack rather than making the model robust. Their attacks circumvented six completely and one partially. They named three patterns:

  • Shattered gradients: a non-differentiable step such as JPEG compression, quantisation or thermometer encoding. Fix: BPDA, which runs the step forward but replaces it with an approximation such as identity on the backward pass.
  • Stochastic gradients: random resizing, dropout at test time. Fix: Expectation over Transformation, averaging gradients over many random draws.
  • Vanishing or exploding gradients: iterative purification or deep unrolled computation. Fix: reparameterise or attack the components separately.

Tramer and colleagues (2020) then broke thirteen more published defenses, each with an attack designed for that defense. The signatures of a broken evaluation are consistent: robust accuracy that does not fall as eps grows, iterative attacks doing no better than single-step ones, black-box attacks beating white-box ones, and unbounded attacks failing to reach 100 percent success. Any one of these means the attack, not the model, is the weak part.

Adversarial training and its costs

PGD adversarial training remains the most reliable empirical defense: in every step, replace or mix the batch with PGD examples against the current weights. It costs roughly K+1 forward-backward passes per step for K attack steps. Cheaper variants such as FGSM with random start work but can suffer catastrophic overfitting, where robustness to multi-step attacks collapses suddenly during training; monitor PGD accuracy on a held-out batch every epoch. Robust models lose clean accuracy, and Tsipras and colleagues argued the tension can be inherent. Ilyas and colleagues argued adversarial examples exploit genuinely predictive but non-robust features, which explains why they transfer. Certified methods such as randomized smoothing and interval bound propagation give provable guarantees at a further accuracy cost; see the certification article.

Evaluating robustness honestly

Croce and Hein's AutoAttack (2020) bundles parameter-free attacks (APGD with cross-entropy and with a targeted difference-of-logits loss, FAB, and the black-box Square Attack) so that a single run gives a reasonable lower bound on robustness without hand tuning. It re-evaluated many published defenses and found lower robustness than reported for a substantial share. It is a floor for evaluation, not a ceiling: an adaptive attack written for your defense can still do better.

  1. State the threat model, including eps, norm, knowledge and goal.
  2. Run PGD with many steps and restarts, plus AutoAttack.
  3. Plot robust accuracy against eps; it must fall to zero.
  4. Check that targeted and multi-step attacks beat single-step ones.
  5. Remove or approximate every non-differentiable or random component and attack again.
  6. Publish code and weights so others can attack it.

Worked example

Take a linear classifier on 32x32 RGB images, so 3072 inputs, whose weights average 0.05 in absolute value. Its L1 norm is about 3072 * 0.05 = 154. An L-infinity perturbation of 8/255, about 0.031, aligned with sign(w) moves the score by 0.031 * 154, about 4.8 logits, enough to flip most confident decisions although no pixel changed by more than about 3 percent. Doubling the number of input dimensions doubles the effect, so doubling the image's side length roughly quadruples it. That is Goodfellow's argument in numbers.

Now an evaluation. A team adds a JPEG re-encoding step in front of a standard model and reports 55 percent robust accuracy under PGD-20 at 8/255 (an illustrative figure). A reviewer notices that robust accuracy barely changes when eps doubles. Running PGD with BPDA, treating JPEG as identity on the backward pass, drops robust accuracy to near zero, matching the undefended model. The preprocessing never removed adversarial directions; it hid them from the gradient.

From pixels to tokens

Language models move the problem from continuous pixels to discrete tokens, so gradient steps cannot be applied directly. Greedy Coordinate Gradient (Zou and colleagues, 2023) uses token-embedding gradients to propose candidate swaps and forward passes to choose among them, and its suffixes transfer between models, the same phenomenon Szegedy observed; see the GCG article. The defensive lessons carry over directly. A perplexity filter is a shattered-gradient style defense that adaptive attacks route around with fluent text. A judge model is a second model that can itself be attacked. Robustness claimed against a fixed set of jailbreak strings is a static evaluation. Adversarial training of LLMs, covered in adversarial training for LLM safety, inherits the min-max framing from Madry and its costs.

Trade-offs

Adversarial training buys empirical robustness for one threat model at a cost in clean accuracy and several times the compute, and it does not generalise to unforeseen perturbation types. Certification gives guarantees only for small radii. Input preprocessing and detection are cheap and usually fail adaptive attacks. System-level controls such as rate limits, human review for high-stakes decisions and not exposing scores reduce exposure without any model change and should come first in production.

What to do next

  1. Write the threat model for one deployed model you own: knowledge, outputs exposed, perturbation set and goal.
  2. Implement FGSM and PGD as above and measure clean versus robust accuracy at three eps values.
  3. If any defense is in place, attack it with BPDA or EOT and compare.
  4. Run AutoAttack before publishing any robustness number.
  5. Reduce exposed output detail and add query monitoring for public endpoints.
  6. Read Athalye et al. 2018 and Carlini et al.'s guidance On Evaluating Adversarial Robustness before reviewing any defense paper.
Key takeaway: Adversarial examples follow from high-dimensional, locally linear models, so small coordinated changes move decisions. State the threat model, attack with PGD, margin losses and AutoAttack, and treat any defense as unproven until an adaptive attack has failed against it. Related reading: <a href="adversarial_training_llm.html">adversarial training for LLMs</a>, <a href="llm_sec_certified_robust.html">certified robustness for LLMs</a> and <a href="data_poisoning_attacks.html">training-time data poisoning</a>.