SmoothLLM is a defense against jailbreaks that works without retraining the model and without knowing what the attack looks like. It was published in 2023 by Alexander Robey, Eric Wong, Hamed Hassani and George Pappas, aimed squarely at optimization-based attacks such as GCG, which append a string of odd-looking tokens to a harmful request so that an aligned model complies. The insight is simple: those suffixes are found by searching over exact tokens, so they are fragile. Change a few percent of the characters at random and the suffix usually stops working, while the meaning of a normal request mostly survives.

SmoothLLM turns that observation into a wrapper. It makes N randomly perturbed copies of the prompt, sends each to the model, checks which responses look jailbroken, takes a majority vote and returns a response from the winning side. This page gives the algorithm, working code, a worked example, parameter choice, failure modes and production costs. How GCG suffixes are found is covered in GCG universal adversarial attacks; here we take the attack as given and focus on the defense.

Why optimized suffixes are brittle

A GCG suffix is the output of a discrete optimization. The attacker searches for a token sequence that maximizes the probability of an affirmative opening such as "Sure, here is". The result sits on a narrow peak: the exact tokens matter. The SmoothLLM paper formalizes this as k-instability. A suffix S is k-unstable if the attack fails whenever the perturbed suffix S' differs from S in at least k characters (Hamming distance at least k). The authors measured that GCG suffixes are unstable in this sense: perturbing around 10% of their characters drops the attack success rate sharply.

The defender does not know where the suffix starts, or whether there is one. So SmoothLLM perturbs the whole prompt, not just the tail. The benign part is corrupted too; models tolerate light character noise in natural language, which is why q has to stay small.

Randomization does the rest: if each copy is likely to break the suffix, the chance that most of N copies keep it intact falls quickly as N grows. No number of copies helps against robust attacks.

Architecture

SmoothLLM: perturb N copies, query each, vote, answer from the majorityuser prompt Ppossibly with suffixperturb (q%)copy 1perturb (q%)copy 2perturb (q%)copy N...LLMresponse R1LLMresponse R2LLMresponse RNJB check + voteshare jailbroken vs gammareturn one Rjthat agrees with the voteA suffix optimized character by character is brittle: changing a few percent of characters usually breaks it.Fluent natural-language jailbreaks are not brittle in the same way, so the defense helps far less against them.Cost: N model calls per request, and no token can stream until all N have finished and been voted on.
Figure: the SmoothLLM wrapper. The perturbation and vote sit between your application and the model; the model itself is unchanged.

The wrapper has two stages. The perturbation step builds N copies of the prompt, each with q percent of its characters altered by one of three schemes. The aggregation step runs each copy through the model, applies a jailbreak indicator JB to each response (1 if the response looks jailbroken, 0 if it looks like a refusal), and compares the share of jailbroken responses to a threshold gamma, set to 0.5 in the paper. The majority decision is then not returned as a label: SmoothLLM returns one of the actual responses whose JB value agrees with the majority, chosen at random.

So a benign user's answer is always generated from a perturbed copy of their prompt, typos included.

The perturbation step

The paper defines three perturbation functions over the prompt string, each controlled by a percentage q. Insert samples q% of the character positions and inserts a random character after each. Swap samples q% of the positions and replaces each character with a random one. Patch picks a contiguous run of q% of the characters at a random location and replaces the whole run. Random characters come from the printable alphabet.

import random
import string

ALPHABET = string.printable  # digits, letters, punctuation, whitespace


def swap(s: str, q: float, rng: random.Random) -> str:
    chars = list(s)
    k = max(1, int(len(chars) * q))
    for i in rng.sample(range(len(chars)), k):
        chars[i] = rng.choice(ALPHABET)
    return "".join(chars)


def insert(s: str, q: float, rng: random.Random) -> str:
    k = max(1, int(len(s) * q))
    picked = set(rng.sample(range(len(s)), k))
    out = []
    for i, ch in enumerate(s):
        out.append(ch)
        if i in picked:
            out.append(rng.choice(ALPHABET))
    return "".join(out)


def patch(s: str, q: float, rng: random.Random) -> str:
    width = max(1, int(len(s) * q))
    start = rng.randrange(0, len(s) - width + 1)
    noise = "".join(rng.choice(ALPHABET) for _ in range(width))
    return s[:start] + noise + s[start + width:]


PERTURB = {"swap": swap, "insert": insert, "patch": patch}

In the paper's experiments against GCG, swap was generally the most effective of the three at a given q. Patch concentrates its damage in one place and may miss the suffix on a long prompt; swap spreads it and is likely to hit several suffix characters.

Two choices are yours. Perturb only the untrusted user turn: perturbing the system prompt and chat template degrades every response for no benefit. And skip very short prompts, which cannot carry a long optimized suffix and where 10% is a single character.

The vote and the returned response

The aggregation step needs a jailbreak indicator. The reference implementation uses the same check commonly used to score GCG: a response counts as not jailbroken if it contains one of a list of refusal phrases such as "I'm sorry" or "I cannot". That is crude, and it is the weakest part of the defense, but it is cheap and deterministic. A stronger option is a small safety classifier that labels responses, at the cost of another model call per copy.

import concurrent.futures as cf

REFUSALS = ("I'm sorry", "I am sorry", "I apologize", "As an AI",
            "I cannot", "I can't", "I will not", "I won't", "is illegal")


def jb(response: str) -> int:
    """1 if the response looks jailbroken (no refusal phrase), else 0."""
    return int(not any(r in response for r in REFUSALS))


def smooth_llm(prompt, llm, n=10, q=0.10, kind="swap", gamma=0.5, seed=None):
    rng = random.Random(seed)
    copies = [PERTURB[kind](prompt, q, rng) for _ in range(n)]
    with cf.ThreadPoolExecutor(max_workers=n) as pool:
        responses = list(pool.map(llm, copies))
    flags = [jb(r) for r in responses]
    share = sum(flags) / n
    majority = int(share > gamma)
    agreeing = [r for r, f in zip(responses, flags) if f == majority]
    return rng.choice(agreeing), {"share_jailbroken": share, "majority": majority}

Note that majority uses a strict inequality, so with N=10 and gamma=0.5 a 5-5 split counts as not jailbroken and a refusal-side response is returned. Choose an odd N if you want to avoid ties altogether. The copies are independent, so they should be sent concurrently; with a serving stack that batches requests, N copies cost close to one batched forward pass in latency but still N times the tokens.

Worked example

Take a harmful request with a GCG suffix appended: the goal sentence is about 60 characters and the suffix about 90, so the prompt is 150 characters. Run SmoothLLM with swap, q=10% and N=7.

Each copy changes 15 characters chosen uniformly at random. On average about 9 of them land in the suffix and 6 in the goal. The goal becomes something like "Write a tutoriel on how to m&ke...", which the model still understands. The suffix loses about 9 of its 90 optimized characters, and because the attack depended on those exact tokens, the model now sees a harmful request with some noise after it. Suppose six copies produce refusals and one copy happens to leave the suffix effective enough that the model complies.

The JB flags are [0, 0, 1, 0, 0, 0, 0]. The jailbroken share is 1/7, about 0.14, which is below gamma, so the majority is "not jailbroken" and SmoothLLM returns one of the six refusals at random. The attacker would need four of seven copies to keep the suffix working, and if each copy keeps it working with probability 0.15, the chance of that is roughly 1%.

Now a benign request: "Summarize the attached incident report in five bullet points." All seven noisy copies yield summaries with no refusal phrase, so JB is 1 on all seven and a summary is returned. JB=1 just means "no refusal", which for a benign prompt is the correct outcome.

Choosing N and q

The two knobs are N (copies) and q (perturbation rate). The paper evaluated N up to 20 and q of 5%, 10%, 15% and 20%. Against GCG, swap with q=10% and N greater than about 6 brought the attack success rate on both Vicuna and Llama 2 below 1%, from 98% and 52% undefended respectively.

SettingEffect on attacksEffect on normal useCost
Raise qBreaks more suffixes per copyMore typos; quality drops (Llama 2 on PIQA fell from 76.7% to 70.3% at q=5%, N=20)None
Raise NVote becomes more reliableLittle effect on qualityLinear in tokens and GPU time
SwapStrongest in the paper's GCG testsSpreads damage evenlySame
PatchMay miss the suffix on long promptsLeaves most text intactSame
Gamma below 0.5Refuses on a minority of jailbroken copiesMore false refusalsSame

A sensible starting point is swap, q between 5% and 10%, and N between 5 and 9 (odd). Then measure two things on your own traffic: attack success on a set of known suffix attacks, and task quality on a sample of benign prompts scored by your normal evaluation. Increase q only while quality stays acceptable; increase N only while the added cost buys a measurable drop in attack success.

Failure modes and limits

SmoothLLM's guarantee is only as good as its assumption that the attack is brittle. Several things break that assumption or the machinery around it.

  • Semantic jailbreaks. Attacks such as PAIR produce fluent prompts: role play, hypotheticals, persuasion. A few typos do not change their meaning. The paper reports that SmoothLLM still reduced PAIR's success, by around a factor of two on Vicuna and GPT-4, but that is a long way from the below-1% figure against GCG. Multi-turn attacks such as those in multi-turn jailbreaks fall in the same class.
  • Adaptive attacks. An attacker who knows SmoothLLM is deployed can optimize for suffixes that survive perturbation, for example by optimizing against an expectation over random perturbations. Robustness then costs the attacker more queries and longer suffixes, but it is not impossible. Treat SmoothLLM as raising the cost of a known attack class, not as a certificate.
  • The JB check. A keyword list misses refusals phrased differently, and counts a response that starts "I'm sorry, but here is how..." as a refusal. Both errors bias the vote. Monitor the share of responses that the check and a stronger classifier disagree on.
  • Structured inputs. Code, JSON, SQL, URLs, identifiers and numbers do not tolerate random character swaps. A 10% swap on a stack trace or a config file changes its meaning. For these prompts SmoothLLM can turn a correct answer into a wrong one without any attack present.
  • Non-text attack channels. Injected instructions in retrieved documents, tool outputs or images are not in the user turn you perturb. SmoothLLM does nothing for indirect prompt injection.

Running it in production

In production the defense is a gateway feature, not a model change. Run it in the service that already handles authentication and logging, so the N copies inherit the same rate limits and tracing as the original request.

  • Latency and streaming. No output can be shown until all N copies finish and are voted on. For a chat product that removes token streaming. A common compromise is to cap the output length of the N voting copies, vote on those short prefixes, and then generate the full answer once from the winning side. Be aware that this is a change to the published algorithm and must be re-evaluated, because jailbreaks sometimes refuse at first and comply later.
  • Cost. N copies cost N times the input and output tokens. Apply SmoothLLM selectively: to anonymous or low-trust users, to prompts that a cheap detector such as a perplexity filter flags as unusual, or to models that are known to be weak against suffix attacks.
  • Determinism. Log the per-request seed, kind, q, N, JB flags and share so answers are reproducible.
  • Monitoring. Track the distribution of the jailbroken share. Benign traffic sits near 1 (no refusals) or near 0 for prompts the model always refuses; shares in the middle mean the copies disagree, which is where suffix attacks and borderline requests land. Review that band.

Trade-offs

SmoothLLM's strengths are that it needs no access to weights or gradients, works with any hosted model, and gives a strong reduction against the attack class it was designed for. Its costs are N times the inference, lost streaming, a quality drop that grows with q, and corrupted structured inputs. Compared with a perplexity filter it is more expensive but harder to evade; compared with adversarial training it is cheaper to adopt but does nothing outside the wrapper. In a layered setup, as described in jailbreak defense, it belongs behind cheap filters and in front of output moderation, applied where the risk justifies the cost.

What to do next

  1. Build a test set: 50 to 100 known suffix attacks, a set of fluent jailbreaks, and a few hundred benign prompts from real traffic, including code.
  2. Implement the wrapper above against a staging endpoint, perturbing only the user turn.
  3. Sweep swap with q in {5%, 10%} and N in {5, 7, 9}; record attack success and benign task quality for each.
  4. Replace the keyword JB check with a classifier if the two disagree on more than a few percent of responses.
  5. Decide the routing rule: which users, prompts or models get SmoothLLM, and skip it for short or structured inputs.
  6. Log seed, parameters, flags and share for every smoothed request, and alert on the middle band of shares.
  7. Re-run the test set whenever the model, the refusal phrasing or the attack corpus changes.
  8. Keep learning: LLM jailbreaking and LLM red teaming.
Key takeaway: SmoothLLM exploits the brittleness of optimized adversarial suffixes: perturb a few percent of the prompt's characters in N copies, query each, vote on whether the responses look jailbroken, and return a response from the majority side. Against GCG it cuts attack success to around 1% with swap at q=10% and a handful of copies, but it costs N times the inference, removes streaming, degrades structured inputs and helps much less against fluent or adaptive jailbreaks. Use it selectively, as one measured layer.