An adversarial example is an input that has been changed just enough to make a model give the wrong answer while a person would still give the right one. In image models the change is a few pixels of noise. In language models the change is a typo, a swapped synonym, a look-alike Unicode character or an extra sentence, and the target is rarely the chat model alone. It is usually one of the classifiers wrapped around it: the moderation filter, the prompt-injection detector, the LLM-as-judge that grades outputs, or the embedding model that decides which documents a retriever returns.

This article explains adversarial examples for text from first principles, shows why the discrete setting changes the attack and the defense, walks through the main attack families with working code for a black-box word-substitution attack, and ends with an evaluation harness and a checklist. Optimized jailbreak suffixes are covered separately in the GCG deep dive, and the continuous attacks FGSM and PGD in Adversarial ML, in depth.

What an adversarial example is

Formally, take a model f, an input x with true label y, and a set C(x) of allowed modifications. An adversarial example is any x' in C(x) with f(x') != y where the true label of x' is still y. Every part of that definition does work. The set C(x) is the threat model: how many characters or words the attacker may change, which substitutions count, whether whole sentences may be added. The last clause, that the true label is unchanged, is what separates an adversarial example from an input that is simply different. If a word swap turns a positive review into a negative one, the model was right to change its answer.

In images, C(x) is usually a small ball around x in pixel space, for example every pixel within 8/255 of the original. That ball is continuous, so the attacker can follow the gradient of the loss directly. Text is a sequence of discrete tokens. There is no token halfway between "good" and "great", the model's embedding space has gradients but most points in it correspond to no real token, and a one-character edit can change tokenization completely. So text attacks become search problems: find a small set of discrete edits that moves the model, using gradients only as a hint about where to look.

Three threat models matter in practice. White-box: the attacker has the weights and can compute gradients, which is the case for open-weight guard models. Black-box with scores: the attacker can query the model and sees a probability or confidence, which is common for moderation APIs. Black-box with labels only, or transfer: the attacker sees only accept or reject, or builds examples on a local surrogate and hopes they carry over.

Attack families by granularity

Text attacks are easiest to organize by the granularity of the edit, because the granularity determines both what the attack exploits and which defense applies.

LevelTypical editsWhat it exploitsFirst defense
CharacterTypos, swapped letters, homoglyphs, zero-width characters, leetspeakTokenizer brittleness: one edit turns a known token into rare sub-tokensUnicode normalization and confusable mapping
Word or tokenSynonym swaps (TextFooler), gradient-guided flips (HotFlip)Reliance on a few high-weight wordsAdversarial training, ensembles
SentenceParaphrases, appended distractor sentencesShallow pattern matching instead of reading the whole inputTraining on paraphrases, span-level checks
PromptOptimized suffixes, role-play framingInstruction-following in generative modelsCovered in the GCG and jailbreak pages

HotFlip (Ebrahimi and colleagues, 2018) is the core white-box idea. Treat each input position as a one-hot vector over the vocabulary. The gradient of the loss with respect to the embedding at position i gives a first-order estimate of how much the loss would rise if token a were replaced by token b: the dot product (e_b - e_a) . grad_i. One backward pass scores every possible swap at every position, and the attacker evaluates only the most promising few exactly. GCG applies the same estimate to an appended suffix on a generative model.

TextFooler (Jin and colleagues, 2020) is the core black-box idea. Score each word by how much the model's confidence in the true label drops when the word is deleted. Visit words from most to least important. For each, try synonyms from a counter-fitted embedding neighbourhood, keep only candidates that pass a part-of-speech check and a sentence-level similarity threshold, and accept the one that lowers confidence most. Stop when the label flips. It needs nothing but query access to probabilities.

At sentence level, Jia and Liang (2017) showed that appending one distracting sentence that resembled the question but did not answer it cut reading-comprehension accuracy sharply. Retrieved documents written to look relevant to a query are the modern version.

What gets attacked in an LLM system

In an LLM application the components most exposed to adversarial examples are the small, fast classifiers that make binary decisions, because a binary decision is exactly what an attacker wants to flip:

  • Input guards: toxicity, policy or prompt-injection classifiers. A false negative lets content through to the model.
  • Output guards: the same classifiers on the response. Character-level edits that the generator itself produced on request are a known way around them.
  • LLM-as-judge: an evaluator that scores answers. Appended text that flatters the rubric can raise scores without improving the answer.
  • Retrievers and rerankers: embedding similarity is a continuous score over discrete text, so it can be pushed up for an irrelevant passage.

A working black-box substitution attack

Below is a complete black-box word-substitution attack in the TextFooler style. It needs a function that returns class probabilities, a synonym source, and a similarity function. It is written to be run against your own classifier in a test environment to measure how easily its decisions flip.

from dataclasses import dataclass

@dataclass
class AttackResult:
    success: bool
    adversarial: list[str]
    queries: int
    words_changed: int

def word_importance(predict, words, label):
    """Confidence drop for the true label when each word is deleted."""
    base = predict(" ".join(words))[label]
    scores = []
    for i in range(len(words)):
        probe = words[:i] + words[i + 1:]
        scores.append(base - predict(" ".join(probe))[label])
    return scores, len(words) + 1

def substitution_attack(predict, text, label, synonyms, similarity,
                        max_change=0.2, min_sim=0.85):
    words = text.split()
    scores, queries = word_importance(predict, words, label)
    order = sorted(range(len(words)), key=lambda i: -scores[i])
    budget = max(1, int(max_change * len(words)))
    current, changed = list(words), 0
    for i in order:
        if changed >= budget:
            break
        best, best_p = None, predict(" ".join(current))[label]
        queries += 1
        for cand in synonyms(current[i])[:30]:
            trial = current[:i] + [cand] + current[i + 1:]
            if similarity(" ".join(words), " ".join(trial)) < min_sim:
                continue                      # meaning drifted: not a valid example
            probs = predict(" ".join(trial))
            queries += 1
            if probs[label] < best_p:
                best, best_p = cand, probs[label]
            if max(range(len(probs)), key=probs.__getitem__) != label:
                current[i] = cand
                return AttackResult(True, current, queries, changed + 1)
        if best is not None:
            current[i], changed = best, changed + 1
    return AttackResult(False, current, queries, changed)

Three details matter. The edit budget max_change and similarity floor min_sim are the threat model; report them with every result, because an attack allowed to change half the words will succeed against anything. The similarity function should be a sentence encoder you trust, and you should still read a sample of successes by hand, because automatic similarity scores pass many edits a person would reject. And the query count is part of the result: a guard that falls after 40 queries and one that needs 4,000 face very different attackers.

Worked example: flipping a review gate

Take a sentiment classifier used as a gate on product reviews, and the input "The battery life is terrible and the screen scratches easily", labeled negative with confidence 0.97. Deletion scoring ranks "terrible" first (deleting it drops confidence to 0.71), then "scratches" and "easily". The attack tries synonyms for "terrible": "awful" keeps confidence at 0.96, "dreadful" at 0.93, "atrocious" drops it to 0.58, because the classifier saw that word rarely in training. With "atrocious" kept, it moves to "scratches" and tries "scuffs", which lowers confidence to 0.44 and flips the label to positive. Two words out of eleven changed, similarity 0.91, about 60 queries.

A person reads both versions as clearly negative, so this is a valid adversarial example. Now try a character-level edit instead: "terrib1e". Many subword tokenizers split it into pieces never seen with negative labels, and confidence can collapse in one query. That is why the character level is the first thing to test and the first thing to defend: it is cheap for the attacker and often fixed by normalization alone. These figures are an illustration of the mechanics, not measurements of a particular model; run the harness below to get yours.

A black-box text attack is a constrained search loop; the constraints decide validityClean input xlabel y, model correctRank positionsdeletion / gradient scorePropose editschar / word / sentenceConstraintsbudget, similarityTarget modelclassifier, guard, judgequeryScore candidatep(y | x') falls?Success checklabel flips, meaning keptno: keep best, next positionDefender viewnormalize input, ensemble, adversarial trainingEvaluation viewclean acc, attack success, queries, edit rate
The attack loop: rank positions, propose edits, filter by the constraint set, query the target and keep the best candidate until the label flips. The same loop, run by the defender, is the robustness evaluation.

Defenses that hold up

Normalization handles most character-level attacks and should come first because it is deterministic. Unicode NFKC folds compatibility forms such as full-width letters, but it does not map Cyrillic or Greek look-alikes to Latin, so it must be combined with a confusables mapping in the spirit of Unicode Technical Standard 39. Strip zero-width and other format characters. The Unicode smuggling article covers the character classes in detail.

import unicodedata

FORMAT_CHARS = {"\u200b", "\u200c", "\u200d", "\u2060", "\ufeff"}
CONFUSABLES = {"\u0430": "a", "\u0435": "e", "\u043e": "o", "\u0440": "p",
               "\u0441": "c", "\u0445": "x", "\u0456": "i"}   # extend from UTS 39 data
LEET = str.maketrans({"0": "o", "1": "l", "3": "e", "4": "a", "5": "s", "@": "a"})

def normalize_for_classifier(s: str, fold_leet: bool = False) -> str:
    s = unicodedata.normalize("NFKC", s)
    s = "".join(ch for ch in s if ch not in FORMAT_CHARS)
    s = "".join(CONFUSABLES.get(ch, ch) for ch in s)
    if fold_leet:                      # second view only: digits matter in normal text
        s = " ".join(w.translate(LEET) if any(c.isalpha() for c in w) else w
                     for w in s.split())
    return s

Run the classifier on both the normalized text and the leet-folded view and take the more cautious score. Never replace the text the generator sees with the normalized version: normalization is lossy and breaks legitimate multilingual input.

For word and sentence attacks, the durable fix is adversarial training: generate successful attacks with the loop above, add them with their correct labels, retrain, and repeat. Ensembles of differently tokenized classifiers help because an edit that confuses one tokenizer often does not confuse another. Randomized defenses, which perturb the input several times and vote, are analysed in perturbation-based defenses; they raise attack cost but attacks adapted to the randomization usually recover. Detectors that flag unusual inputs, such as high perplexity, catch crude character edits and miss fluent synonym swaps, as the detection article shows. Finally, architecture: do not let one classifier be the only barrier in front of an irreversible action.

Measuring robustness

A robustness number without its threat model is meaningless, so the harness records the configuration with the results:

from statistics import mean, median

def argmax(probs):
    return max(range(len(probs)), key=probs.__getitem__)

def evaluate(predict, dataset, attacks, normalizer=None):
    """dataset: list of (text, label). attacks: name -> attack function."""
    wrap = (lambda t: predict(normalizer(t))) if normalizer else predict
    report = {}
    clean_ok = [(t, y) for t, y in dataset if argmax(wrap(t)) == y]
    report["clean_accuracy"] = len(clean_ok) / len(dataset)
    for name, attack in attacks.items():
        runs = [attack(wrap, t, y) for t, y in clean_ok]
        wins = [r for r in runs if r.success]
        report[name] = {
            "attack_success_rate": len(wins) / max(1, len(runs)),
            "robust_accuracy": report["clean_accuracy"] * (1 - len(wins) / max(1, len(runs))),
            "median_queries": median([r.queries for r in wins]) if wins else None,
            "mean_words_changed": mean([r.words_changed for r in wins]) if wins else None,
        }
    return report

Attack only examples the model gets right, otherwise you inflate the success rate. Report robust accuracy next to clean accuracy, because a change that improves robustness by making the model refuse everything is not a win. Libraries such as TextAttack package these recipes; check that their default budgets match your threat model.

Failure modes

  • Invalid examples counted as wins. The edit changed the meaning, so the model was right. Sample successes and have a person label them.
  • Gradient masking. A defense makes gradients useless, white-box attacks fail, and robustness looks high; black-box and transfer attacks still succeed. Always run both.
  • Static test sets. Robustness measured against last quarter's attack file says nothing about an attacker who adapts to your current model. Regenerate attacks per model.
  • Normalization breaking real users. Folding digits or Cyrillic in the text the generator sees corrupts product codes and non-English input. Normalize only the classifier's view.
  • Overfitting to one recipe. Adversarial training on TextFooler output buys robustness to TextFooler and little else. Mix families and keep a held-out attack.
  • Unlimited queries. Score-based attacks need many queries. Exposing raw confidence scores and allowing unlimited retries hands the attacker the gradient signal.

Trade-offs

Normalization is nearly free but only covers characters. Adversarial training costs training compute, usually lowers clean accuracy slightly, and has to be repeated as attacks change. Ensembles multiply inference cost and latency. Randomized smoothing gives certified guarantees only for small edit radii and costs many forward passes. Hiding scores and rate-limiting make query attacks expensive but hurt clients who need calibrated scores. A practical stack is normalization plus adversarial training plus limits on score exposure, with a second, independent control in front of anything that matters.

What to do next

  1. List every classifier in your LLM system that makes a binary or routing decision.
  2. Write down the threat model for each: who can query it, whether scores are exposed, and the edit budget you care about.
  3. Add NFKC plus confusable mapping plus format-character stripping to each classifier's input view.
  4. Run character-level and word-substitution attacks on correctly classified examples and record attack success rate, robust accuracy and median queries.
  5. Have a person label a sample of successful attacks to confirm they are valid.
  6. Feed valid successes back into training, keep one attack family held out, and re-run the harness on every model change.
  7. Stop returning raw confidence scores to untrusted callers and rate-limit retries.
Key takeaway: Adversarial examples for text are small, meaning-preserving edits found by search, and in LLM systems they mostly target the classifiers that guard, judge, retrieve and route. Define the constraint set first, normalize the classifier's view of the input, measure attack success and robust accuracy on examples the model gets right, verify successes by hand, train on them, and never let one classifier be the only barrier in front of something that matters.