In July 2023, Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter and Matt Fredrikson published "Universal and Transferable Adversarial Attacks on Aligned Language Models". It showed that an automated search could find a short string of odd-looking tokens which, appended to a request a safety-trained model would refuse, makes the model comply. The method, Greedy Coordinate Gradient (GCG), needs only white-box access to open models, and the suffixes it found also worked against closed models the authors never had gradients for.

GCG changed how jailbreaks are studied: from hand-crafted role-play prompts to optimization. This article explains the algorithm from first principles, shows it as code, summarises what the paper measured, and spends most of its length on what defenders need: why the attack works, how to detect it, which defenses hold up against an attacker who adapts, and how to measure your own model's robustness.

Advertisement

The objective: make the first words agreeable

A chat model's refusal usually happens in its first few tokens. If the reply starts with an affirmative phrase that restates the request, autoregressive generation tends to continue in the same direction. GCG exploits that. Given a request x and a target string y that begins with an affirmative phrase restating it, the attack looks for a suffix s of fixed length that minimises the negative log-likelihood of y given the chat-formatted prompt x followed by s.

That turns jailbreaking into an ordinary optimization problem with one awkward property: s is a sequence of discrete tokens, so you cannot run gradient descent on it directly. Continuous attacks on embeddings find vectors that correspond to no real token, and earlier discrete methods such as AutoPrompt swapped one token at a time using gradient hints. GCG's contribution is a search that is both gradient-guided and robust.

The algorithm: gradients propose, forward passes decide

Represent each suffix position as a one-hot vector over the vocabulary. The gradient of the loss with respect to that one-hot vector gives, for every position and every possible token, a linear estimate of how much swapping in that token would lower the loss. The estimate is crude, because the model is far from linear in its inputs, so GCG uses it only to shortlist candidates:

  1. Compute the gradient with respect to the one-hot suffix matrix in one forward and one backward pass.
  2. For each suffix position, keep the k tokens with the most negative gradient. The paper used k = 256.
  3. Build B candidate suffixes, each differing from the current one in a single position chosen at random, with the replacement drawn from that position's shortlist. The paper used B = 512.
  4. Evaluate the true loss of all B candidates with forward passes only, and keep the best. Repeat.

The paper used a 20-token suffix and ran for 500 steps. The expensive part is step four: hundreds of thousands of forward passes over a full run, which is why GCG is slow, and why cheaper variants followed.

One GCG iteration: gradients propose token swaps, forward passes pick the bestPrompt + suffix (20 tokens)chat template around bothForward + backwardloss = -log p(target)Gradient wrt one-hotper position, per vocab tokenTop-k per positionk = 256 candidate tokensSample B = 512 swapsone position eachBatched forward passesexact loss per candidateKeep the best candidatenew suffixrepeatUniversal: sum the loss over many prompts.Transfer: sum it over several open models.The paper ran 500 iterations.Defender viewgibberish suffixes: high perplexity, brittle to character noise, visible in evals
One iteration: a single backward pass shortlists tokens per position; 512 single-swap candidates are scored exactly with forward passes; the best becomes the new suffix. The bottom band is the defender's view of the result.

Here is one step in PyTorch-style code. It is written for understanding and for building robustness evaluations; the target string lives in the token sequence and is whatever your evaluation harness supplies:

import torch
import torch.nn.functional as F

def target_loss(model, embeds, target_ids, target_start):
    """Mean cross-entropy of the target tokens given everything before them."""
    logits = model(inputs_embeds=embeds).logits
    pred = logits[:, target_start - 1 : target_start - 1 + target_ids.shape[-1]]
    return F.cross_entropy(pred.transpose(1, 2), target_ids.expand(pred.shape[0], -1),
                           reduction="none").mean(dim=1)

def gcg_step(model, ids, suf, tgt, k=256, batch=512):
    """ids: full token sequence (prompt, suffix, target) inside the chat template.
    suf, tgt: slices for suffix and target positions. Returns a new ids tensor."""
    W = model.get_input_embeddings().weight                 # [vocab, d]
    one_hot = F.one_hot(ids[suf], W.shape[0]).to(W.dtype)
    one_hot.requires_grad_(True)
    embeds = model.get_input_embeddings()(ids).detach().clone()
    embeds = torch.cat([embeds[: suf.start], one_hot @ W, embeds[suf.stop :]]).unsqueeze(0)
    target_loss(model, embeds, ids[tgt].unsqueeze(0), tgt.start).sum().backward()

    top = (-one_hot.grad).topk(k, dim=1).indices             # [suffix_len, k]
    pos = torch.randint(0, top.shape[0], (batch,))
    new_tok = top[pos, torch.randint(0, k, (batch,))]
    cands = ids.repeat(batch, 1)
    cands[torch.arange(batch), suf.start + pos] = new_tok    # one swap per candidate
    # (reference code also drops candidates that do not survive decode/re-encode)

    with torch.no_grad():
        losses = torch.cat([
            target_loss(model, model.get_input_embeddings()(chunk), ids[tgt].unsqueeze(0), tgt.start)
            for chunk in cands.split(64)])
    return cands[losses.argmin()], losses.min().item()

The one-hot trick is the key line: one_hot @ W reproduces the normal embeddings exactly, but now the gradient flows to a matrix with one column per vocabulary token. Evaluating candidates exactly is what makes the method greedy and robust: a swap that the gradient liked but the model does not is simply discarded.

Advertisement

Universal and transferable

Attacking one request on one model already works well. The paper's stronger results come from changing only the loss. For a universal suffix, sum the loss over many requests, each with its own target; the authors added requests incrementally once the suffix worked on the current set, trained on 25 harmful behaviors and tested on 100 held-out ones. For transfer, also sum the loss over several open models sharing a tokenizer; the paper used Vicuna models, whose training data came from conversations with ChatGPT.

On the white-box side, the paper's Table 1 reports that GCG succeeded on 99% of individual harmful behaviors against Vicuna-7B and 56% against LLaMA-2-7B-Chat, a much more heavily safety-tuned model. For transfer, Table 2 reports an 86.6% attack success rate on GPT-3.5 when any of several suffixes counted as success, substantially lower rates on GPT-4 and PaLM-2, and very low rates on Claude 2. Those are 2023 numbers against 2023 model versions and 2023 success criteria; do not read them as properties of any model you use today.

The practical lesson is independent of the numbers. A suffix optimized on models you can download may work on a model you only reach through an API, so keeping weights private is not a defense by itself.

Why it works

Safety training teaches refusal as a behaviour on the distribution of prompts the model saw during training. A GCG suffix is far off that distribution: a string of fragments, punctuation and multilingual subwords no human writes. In that region the refusal behaviour is not well anchored, and the pressure from the objective can push the model into its compliant mode. Once the first tokens are affirmative, the model's training to be consistent with its own previous text does the rest.

Transfer suggests the vulnerable directions are shared. Models trained on similar data, especially models distilled from another model's outputs, appear to share features, so a suffix that exploits one often partly exploits another. For the broader account of why safety training generalises imperfectly, see LLM jailbreaking in depth.

Detection: perplexity and its limits

Raw GCG suffixes are gibberish, and gibberish is improbable under any language model. Work published shortly after GCG, including Jain et al.'s baseline-defenses paper and Alon and Kamfonas's perplexity study, showed that perplexity filtering catches unconstrained GCG suffixes well. A windowed version is more sensitive than whole-prompt perplexity, because a 20-token suffix after a long, fluent prompt barely moves the average:

import torch
import torch.nn.functional as F
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("gpt2")
lm = AutoModelForCausalLM.from_pretrained("gpt2").eval()

@torch.no_grad()
def max_window_nll(text: str, window: int = 16) -> float:
    """Highest mean token NLL over any window. Gibberish suffixes spike it."""
    ids = tok(text, return_tensors="pt", truncation=True, max_length=1024).input_ids
    if ids.shape[1] < 2:
        return 0.0
    nll = F.cross_entropy(lm(ids).logits[0, :-1], ids[0, 1:], reduction="none")
    if nll.numel() <= window:
        return nll.mean().item()
    return nll.unfold(0, window, 1).mean(dim=1).max().item()

# Calibrate on real benign traffic, e.g. the 99.9th percentile, per language and
# per surface (chat, code, search). Route above-threshold prompts to a stricter path
# rather than rejecting outright: code, URLs and non-English text also score high.

The limits are serious. Legitimate traffic also has high perplexity: code, base64, URLs, tables and text in languages the filter model knows poorly. And an attacker who knows about the filter can add a fluency term to the objective. Later attacks, AutoDAN among them, produce readable adversarial text precisely to evade this check. Treat perplexity as a cheap first layer that removes the unsophisticated attacks and raises the cost of the rest, not as a fix.

Defenses that change the model or the pipeline

DefenseIdeaCostAdaptive-attack caveat
Perplexity filterFlag improbable textOne small-model passFluent attacks evade it
Paraphrase or retokenizeRewrite input before the target modelExtra model call, some quality lossSemantic attacks survive rewriting
SmoothLLM (Robey et al.)Perturb characters in several copies, aggregateN times inferenceRobust suffixes can be optimized
Adversarial trainingTrain on attacks found during trainingExpensive; GCG in the loop is slowCovers the attacks you trained on
Representation-level defensesDisrupt internal harmful directionsTraining-time changeNewer; evaluate yourself
Output classifierJudge the response, not the promptOne classifier callAttack must beat two models

The strongest single idea for an application team is the last row. GCG optimizes what the target model says; a separate classifier judging the output, with its own weights, is not in the optimization loop unless the attacker also targets it. Combine it with least-privilege tools, so that a jailbroken reply cannot take a harmful action on its own. Jailbreak defense architecture and defense in depth lay out the full stack.

Measuring your own robustness

If you fine-tune or host open-weight models, run GCG against them yourself, because a fine-tune can erase safety behaviour that the base model had. Standardised benchmarks help: HarmBench, for example, includes GCG among its attacks and uses a trained classifier as the judge, which avoids the old trap of counting any reply without the word "sorry" as a success. A regression check fits in CI:

# Robustness regression in CI (sketch): same seeds, same behaviors, every model release.
for model_version in [baseline, candidate]:
    for behavior in heldout_eval_behaviors:              # from a benchmark such as HarmBench
        suffix = run_gcg(model_version, behavior, steps=STEPS, seed=SEED)
        response = generate(model_version, behavior.prompt + " " + suffix)
        record(model_version, behavior.id, judge(behavior, response))   # classifier, not string match
assert asr(candidate) <= asr(baseline) + TOLERANCE

Fix seeds and step budgets so runs are comparable, report attack success with confidence intervals, and include an adaptive variant that knows your filter. Red-team architecture and safety evals describe the surrounding pipeline. Store any suffixes you find as sensitive test fixtures, not in public issue trackers.

Operational guidance and trade-offs

  • Closed API behind a vendor model: you cannot retrain it, so rely on input screening, an output classifier and tool permissions, and assume some suffixes will get through.
  • Open weights you serve: run GCG-based evals on every fine-tune and quantized variant before release.
  • Budget false positives explicitly. A perplexity filter tuned on English chat will flag a code assistant's normal traffic; calibrate per surface and route flagged prompts to stricter handling rather than blocking.
  • Log and cluster high-perplexity prompts. Repeated suffix fragments across accounts are a strong signal of automated attacks.
  • Remember the threat: GCG shows safety training can be bypassed, so the harm ceiling is set by what the model can do and access, not by its refusal rate.

What to do next

  1. Read the GCG paper's method section and implement one step on a small open model with a benign target to see the loss fall.
  2. Add a windowed perplexity score to your request logs and study its distribution on real traffic before setting any threshold.
  3. Put an independent output classifier in front of any response that reaches users or tools.
  4. Audit tool permissions so a single jailbroken reply cannot cause irreversible harm.
  5. Run a GCG-based robustness eval on every model or fine-tune you ship, with fixed seeds and a classifier judge.
  6. Add an adaptive attack that knows your filter, and track success rate across releases.
Key takeaway: GCG turns jailbreaking into optimization: it searches for a short token suffix that makes an aligned model start its reply affirmatively, using gradients on one-hot token vectors to shortlist swaps and exact forward passes to choose among 512 candidates per step. Summing the loss over many prompts and several open models produced suffixes that were universal and transferred to closed APIs. The raw suffixes are gibberish, so perplexity filters catch them cheaply, but fluent adaptive attacks evade that layer. Defend in depth with output classifiers, least-privilege tools and randomized or trained robustness, and measure your own models with fixed-seed GCG evals in CI.