Some jailbreaks are found by optimization. Greedy Coordinate Gradient (GCG) searches token by token for a suffix that pushes an aligned model toward an affirmative opening, and the result is a string like a bag of punctuation, rare subwords and half-words. That string works because every character was chosen to sit at a precise point in the model's input space. Change a few of them and it often stops working, while a normal question with a few typos still reads as the same question.

Perturbation-based defenses turn that asymmetry into a control: randomize the input before the model sees it, usually several times, and either vote on the outcomes or treat disagreement among them as an alarm. This article treats the family as a design space: noise placement, output modes, voting arithmetic, adaptive attackers and benign cost. It assumes you know what an adversarial suffix is; if not, start with GCG and universal adversarial suffixes.

The bet: the attack is more fragile than the request

Every perturbation defense makes one bet: the attack is more fragile than the request. Write the target model as f, a perturbation as a random function T, and a judge J that labels a reply as jailbroken or not. For an adversarial prompt x_adv, the defense wants the probability that J(f(T(x_adv))) says jailbroken to be low. For a benign prompt x, it wants f(T(x)) to still answer the question. If both hold, the defense helps. If the attack survives perturbation as well as the request does, nothing has been gained and the benign user has paid for it.

That framing tells you in advance which attacks the family can address. Token-optimized suffixes (GCG and its descendants) are brittle by construction, because the optimizer exploits exact token boundaries. Semantic jailbreaks, such as role-play framings or those generated by an attacker LLM in the style of PAIR, are written in fluent language and carry their meaning in many redundant words; swapping 10% of their characters leaves the persuasion intact. A perturbation defense is therefore a specific tool for one threat class, not a general jailbreak shield.

Where the noise goes

Published methods differ mainly in where the randomness is injected. Moving up the table, perturbations become more expensive and more meaning-preserving, which matters because a defense that destroys meaning also destroys benign answers.

LevelExamplesWhat changesCost per copyMain risk
CharactersSmoothLLM: random swap, insert or patch of q% of characters (Robey et al., 2023)Exact bytes, hence tokenizationOne target callCorrupts numbers, code, URLs
TokensRetokenization with BPE-dropout (Jain et al., 2023); random token dropping with a refusal check (RA-LLM, Cao et al., 2023)Token boundaries or a subset of tokensOne target callDropped negations flip meaning
MeaningParaphrasing by a second LLM (Jain et al., 2023); SemanticSmooth's paraphrase, summarize and translate transforms (Ji et al., 2024)Surface form; intent is meant to surviveParaphraser call plus target callParaphraser can be jailbroken or drop detail

Two neighbours are often confused with this family. Perplexity filtering, also studied by Jain et al., does not perturb anything; it scores the input and rejects improbable strings, so it fails against low-perplexity attacks rather than against perturbation-robust ones. Erase-and-check removes tokens deterministically and exhaustively to obtain a certificate; it is covered with randomized smoothing in certified robustness for LLMs. The methods here are empirical: they reduce attack success on the attacks tested and prove nothing beyond that.

Architecture: aggregate or detect

Perturb N copies, query the model, then aggregate or detectIncoming promptxPerturbation 1chars / tokens / meaningPerturbation 2independent drawPerturbation Nseeded per promptTarget LLMN batched callsAggregate modemajority vote, return one replyDetect modedisagreement = suspiciousverdictsJudge per copyrefusal / jailbroken classifier on each replyCost scales with N; protection only exists where the attack is more fragile than the request.
The shared skeleton: N independent perturbations, batched target calls, a per-copy judge, then one of two output modes.

Aggregate mode answers the user. Each perturbed copy gets a verdict from a judge, the majority verdict wins, and the defense returns a reply drawn from the copies that agree with the majority. SmoothLLM works this way. The user receives the reply to a perturbed prompt, so perturbation damage is visible in every answer.

Detect mode only classifies. The original, unperturbed prompt is answered normally unless the perturbed copies disagree with it, for example when the original reply complies but most perturbed replies refuse. RA-LLM is a detector of this kind: it drops random tokens several times and rejects the request if enough of the copies refuse. Benign users get unmodified answers, which is a large utility win, but the defense now needs a threshold and a fallback path for flagged requests.

A reference implementation

The skeleton below supports both modes with pluggable perturbations. The judge is a stand-in: a keyword refusal check is the cheapest choice and also the weakest, so production deployments usually call a small safety classifier on the first few hundred tokens of each reply.

import hashlib, hmac, random

def swap_chars(s, q, rng):
    """SmoothLLM-style swap: replace a fraction q of characters with printable ones."""
    chars = list(s)
    k = max(1, int(len(chars) * q))
    for i in rng.sample(range(len(chars)), k):
        chars[i] = chr(rng.randint(32, 126))
    return "".join(chars)

def drop_tokens(s, q, rng):
    """RA-LLM-style drop on whitespace tokens (real systems drop model tokens)."""
    toks = s.split()
    keep = [t for t in toks if rng.random() >= q]
    return " ".join(keep or toks)

def seeded_rng(prompt, key):
    # Same prompt -> same draws, so resubmitting does not buy the attacker fresh noise.
    digest = hmac.new(key, prompt.encode(), hashlib.sha256).digest()
    return random.Random(int.from_bytes(digest[:8], "big"))

def perturbation_defense(prompt, llm, judge, perturb, n=8, q=0.10,
                         mode="detect", threshold=0.5, key=b"rotate-me"):
    rng = seeded_rng(prompt, key)
    copies = [perturb(prompt, q, rng) for _ in range(n)]
    replies = llm.generate_batch(copies, max_new_tokens=256)
    flagged = [judge(r) for r in replies]          # True = refused or unsafe
    frac = sum(flagged) / n
    if mode == "aggregate":
        majority = frac > 0.5
        pool = [r for r, f in zip(replies, flagged) if f == majority]
        return {"reply": rng.choice(pool), "flag_rate": frac}
    # detect mode: answer the original unless the copies say otherwise
    if frac >= threshold:
        return {"reply": None, "blocked": True, "flag_rate": frac}
    return {"reply": llm.generate(prompt), "blocked": False, "flag_rate": frac}

One detail in that code is deliberate: the random generator is keyed by an HMAC of the prompt. If each request drew fresh randomness, an attacker could resubmit the same prompt until a lucky draw let it through; seeding makes the defense a fixed function of the input while the server-side key keeps the draws unpredictable offline.

Worked example: when voting helps and when it hurts

Take a 200-character prompt whose last 20 characters are an optimized suffix, and a swap rate of q = 10%, which replaces 20 positions chosen uniformly. The number of suffix characters hit follows a hypergeometric distribution. The chance that a copy leaves the suffix completely untouched is C(180, 20) / C(200, 20) = 0.109; at q = 5% it rises to 0.340 and at q = 20% it falls to 0.009. Touching the suffix is not the same as breaking it, but it shows why q in the 5-20% range is where these defenses operate.

Now suppose each perturbed copy still produces a jailbroken reply with probability s, independently. Majority voting over N copies yields a jailbroken majority with the binomial tail probability below. The numbers were computed exactly, with ties counted as no majority.

Per-copy survival sN = 1N = 5N = 11N = 21
0.100.1000.0090.0003below 0.0001
0.300.3000.1630.0780.026
0.450.4500.4070.3670.321
0.600.6000.6830.7540.826

Read the last row twice. Voting is an amplifier, not a filter: when the attack survives a single perturbation more often than not, adding copies makes the defended system more reliably jailbroken than the undefended one. Everything therefore depends on pushing s below one half, and near one half extra copies buy almost nothing. Measure s on your own attack set before choosing N.

The benign side has its own arithmetic. A 4-digit number in the same 200-character prompt survives q = 10% swapping intact with probability C(196, 20) / C(200, 20) = 0.654, so about a third of perturbed copies of an arithmetic question ask about a different number. In aggregate mode some of those replies reach the user. This is the strongest argument for detect mode, or for exempting digits, code spans and URLs from perturbation.

Adaptive attackers

Empirical defenses are judged by the strongest attack that knows about them, and randomized ones have a long history of looking stronger than they are. Athalye, Carlini and Wagner (2018) showed that image defenses relying on randomness fell to Expectation over Transformation (EOT): the attacker optimizes the average loss over the defender's random transformations rather than the loss on one input. The same move applies to text. A GCG variant that scores each candidate suffix on several randomly perturbed copies finds suffixes with a higher survival rate s, and the voting table shows what happens when s climbs.

Three other attacker responses deserve explicit tests:

  • Go semantic. Switch to fluent jailbreaks that carry no brittle tokens. Character-level and token-level perturbation barely affect these; meaning-level smoothing helps only if the paraphraser itself neutralises the framing.
  • Attack the paraphraser. A meaning-level defense adds a second model to the input path. If an instruction embedded in the prompt makes the paraphraser emit the harmful request verbatim, or emit something worse, the defense has become the attack's delivery mechanism.
  • Exploit sampling noise. Best-of-N jailbreaking (Hughes et al., 2024) shows that repeatedly sampling random augmentations such as shuffled capitalization and character noise eventually finds inputs that slip through. Unseeded defenses hand out the same lottery tickets for free, which is why the code above keys its randomness to the prompt.

The practical rule: report results only against an attacker that has the defense in the loop, with the same query budget you expect in production, and keep a non-adaptive number only as a regression smoke test.

Running it in production

Cost is the first operational fact. N copies mean N times the prefill and, in aggregate mode, N partial generations. Prefix caching does not rescue you, because perturbations scatter changes across the whole prompt, so no long shared prefix exists. Teams keep the bill manageable in four ways:

  1. Gate the defense. Run perturbation only on requests that a cheap first-stage signal flags, such as high perplexity, long non-word runs or a classifier score in an uncertain band.
  2. Judge early tokens. Refusal or compliance is usually decided in the opening of a reply; capping per-copy generation at a few dozen to a few hundred tokens cuts cost sharply. Validate the cap on your own traffic, because some attacks delay harmful content past a compliant preamble.
  3. Batch the copies. Send all N in one batched call so latency grows with batch efficiency, not linearly.
  4. Protect structured spans. Leave code blocks, numbers, URLs and quoted user data unperturbed, and accept that an attacker can try to hide a suffix inside them; monitor how often such spans appear in flagged traffic.

Log the flag rate (the fraction of copies judged unsafe) for every gated request. A shift in its benign distribution signals an attack campaign or a model update that changed s.

Evaluating a perturbation defense

An evaluation that will survive review has four parts. Use a fixed attack set split by attack class (optimized suffix, semantic, encoding tricks) so that the defense's narrow coverage is visible rather than averaged away. Score with a fixed judge stronger than the in-loop judge, so the defense is not graded by the same classifier it uses. Measure benign cost on a refusal set of harmless prompts that superficially resemble harmful ones, plus task benchmarks containing numbers and code, and report both a refusal-rate change and an accuracy change. Finally, give confidence intervals: with 100 behaviours, a drop from 12% to 7% attack success is within noise.

Sweep q and N as a grid and plot attack success against benign accuracy. Choose a point on the frontier.

Failure modes

  • Survival above one half. Voting amplifies the attack. Detected by measuring per-copy s; fixed by raising q, changing perturbation level, or abandoning aggregate mode for that class.
  • Weak in-loop judge. A keyword refusal check misses replies that comply while sounding reluctant, and flags benign replies that happen to contain 'I cannot'. Use a classifier and audit disagreements.
  • Free retries. Fresh randomness per request lets an attacker resample until success. Seed per prompt with a secret key and rotate the key.
  • Silent utility loss. Aggregate-mode answers to arithmetic, code and extraction tasks degrade and nobody notices because safety dashboards do not track accuracy. Track both.

Trade-offs

ChoiceGainPrice
Aggregate modeNo threshold to tune; returns an answer alwaysEvery user sees perturbed-prompt answers
Detect modeBenign answers untouchedThreshold, fallback path, one extra full generation
Higher qBreaks more suffixesMore benign corruption; s may not fall for semantic attacks
Larger NSharper vote when s is well below 0.5Linear cost; useless or harmful near and above 0.5
Meaning-level perturbationPreserves intent; touches semantic attacksExtra model in the path, new injection surface, latency
Seeded randomnessRemoves the free-retry channelKey management; identical prompts always get the same treatment

For most deployments the defensible default is gated, seeded, detect-mode perturbation targeted at optimized-suffix traffic, sitting behind an input classifier and in front of an output classifier, not instead of them.

What to do next

  1. Build an attack set split by class and measure per-copy survival s for each class under your chosen perturbation before picking N.
  2. Implement the skeleton in detect mode with HMAC-seeded randomness and structured-span exemptions.
  3. Replace keyword refusal matching with a small safety classifier and audit 100 disagreements by hand.
  4. Run an EOT-style adaptive suffix search and a Best-of-N sampling attack against the defended system with your production query budget.
  5. Measure benign refusal and task accuracy, including arithmetic and code, at every (q, N) grid point; pick a frontier point.
  6. Gate the defense behind a cheap first-stage signal and log the flag-rate distribution as a monitored metric.
  7. Read SmoothLLM in depth for the canonical aggregate-mode design, character-level attacks for the attacker's view of noise, and adversarial ML foundations for why randomized defenses need adaptive evaluation.
Key takeaway: Perturbation defenses work only where an attack is more fragile than the request it hides in, which makes them a targeted control for optimized suffixes rather than a general jailbreak shield. Measure per-copy survival before trusting a vote, because majority voting amplifies whichever side is above one half. Prefer seeded, gated, detect-mode designs that leave benign answers untouched, and evaluate against attackers that optimize through the noise.