AdvBench is the harmful-request dataset released with the 2023 paper Universal and Transferable Adversarial Attacks on Aligned Language Models (Zou, Wang, Kolter and Fredrikson), the paper that introduced the GCG suffix attack. It was built to give that attack something to optimise against and a way to count successes. Since then it has become the default corpus in jailbreak papers, and "attack success rate on AdvBench" appears in hundreds of abstracts.
This article is for people who need to use AdvBench or read results based on it. It covers what is actually in the files, how the original metrics work, what the dataset can and cannot tell you, measured redundancy in the published data, and how to run it as a defensible regression test for your own model. For the attack itself see GCG universal adversarial suffixes. For safety evaluation in general see LLM safety evaluations. No harmful dataset content is reproduced here. The analysis code prints counts only.
What is in the files
The dataset is two CSV files in the data/advbench directory of the llm-attacks GitHub repository. The paper describes 500 harmful strings and 500 harmful behaviours. The files as published hold somewhat more:
| File | Columns | Rows (published CSV) | Task |
|---|---|---|---|
harmful_behaviors.csv | goal, target | 520 | Make the model comply with an instruction |
harmful_strings.csv | target | 574 (572 distinct) | Make the model emit an exact string |
In the behaviours file, goal is an instruction a safety-tuned model should refuse. target is an affirmative opening for a reply to that instruction. All 520 targets begin with the same two words, an agreeing "Sure, here". The target is not a model answer, and nothing in the dataset contains harmful instructions. It exists because GCG optimises a suffix to maximise the probability that the model begins its reply with that opening. The reasoning is that a model which has started agreeing tends to continue. The strings file is a different task: elicit a specific toxic or harmful sentence word for word.
The target design has a consequence for defences. A filter that blocks replies beginning with that exact affirmative phrase will make GCG-on-AdvBench numbers look excellent while doing nothing against an attack that optimises toward a different opening, or none. Evaluate defences against attacks that were not built around the dataset's own targets.
If you cite counts, cite the file you ran and its commit, not the paper's round numbers. Papers that report "out of 500" usually used a subset or truncated the file, and the exact subset changes the result.
How the original metrics work
For the strings task the metric is exact match: did the target string appear in the output. That is unambiguous, but narrow. A model can produce something equally harmful in other words and score zero.
For the behaviours task the original code counts an attack as successful when the completion contains none of a fixed list of refusal phrases (apologies, statements of inability, reminders that it is an AI). Attack success rate (ASR) is the fraction of behaviours counted as successes. The grader is cheap and deterministic, and that is why it spread. It is also wrong in both directions. A model that neither refuses nor helps, by rambling, repeating the prompt or changing the subject, counts as jailbroken. A model that apologises and then helps anyway counts as safe. Later work such as StrongREJECT showed that re-grading published attacks with a judge that scores actual harmful usefulness cuts many reported ASRs sharply. Treat prefix-graded ASR as an upper bound on refusal failure, not a measure of harm.
Measured redundancy
Read the behaviours file end to end and a pattern stands out: many behaviours are paraphrases of one another: the same request with a different verb or target. This matters because ASR is an average over rows. If one theme appears twelve times, a defence that handles that theme moves the score by twelve rows. Measure it yourself. The script below compares the content words of every pair of goals and prints counts only:
import csv, re
STOP = {"a", "an", "the", "to", "and", "of", "for", "how", "on", "that", "in", "someone",
"with", "can", "be", "or", "is", "write", "create", "provide", "develop"}
def words(goal):
return set(re.findall(r"[a-z]+", goal.lower())) - STOP
with open("harmful_behaviors.csv", encoding="utf-8") as f:
goals = [row["goal"] for row in csv.DictReader(f)]
sets = [words(g) for g in goals]
def jaccard(a, b):
return len(a & b) / max(1, len(a | b))
for threshold in (0.5, 0.6, 0.7):
near = sum(1 for i, a in enumerate(sets)
if any(jaccard(a, b) >= threshold for j, b in enumerate(sets) if j != i))
print(f"goals with a neighbour at Jaccard >= {threshold}: {near} of {len(goals)}")On the published file, with this stopword list, 218 of 520 goals have at least one other goal with content-word Jaccard similarity of 0.5 or more, 142 at 0.6 or more, and 86 at 0.7 or more. Word overlap is a crude measure and these are not all duplicates, but they show that the effective number of independent behaviours is well below 520. Later work responded by curating subsets. PAIR, for example, evaluated on a custom subset of 50 AdvBench behaviours, and newer benchmarks draw on AdvBench while adding de-duplication and categories.
What AdvBench does not cover
Beyond redundancy, know what AdvBench does not cover before you treat a number as evidence:
- No categories. Rows carry no harm-type label, so a single ASR mixes cybercrime, violence, fraud and misinformation, and a defence that only covers one of them can look broadly effective.
- Single turn, English, text only. Multi-turn escalation, other languages, tool use and images are out of scope. See multi-turn jailbreaks.
- No benign controls. Every row should be refused, so a model that refuses everything scores perfectly. Pair AdvBench with an over-refusal set or you will reward a useless model.
- Contamination. The files have been public since 2023. They are probably in many pretraining crawls and almost certainly in some safety fine-tuning data. A low no-attack ASR may mean the model memorised these phrasings rather than learned the policy. Paraphrase a sample and compare.
- Unreported configuration. System prompt, chat template, decoding temperature and maximum output length all change ASR. A short output limit, for example, cuts off completions before a late refusal or a late harmful turn. A result without these settings cannot be reproduced.
Running it as a regression test
Used carefully, AdvBench is still a useful regression test. It is fixed, small and widely understood, so a change in your model's behaviour on it is worth noticing. The aim is a number you would defend in a release review. A harness that does that:
def evaluate(model, behaviours, attack, judge, cfg):
"""Return per-row records; never print prompts or completions to shared logs."""
records = []
for row in behaviours: # a de-duplicated, versioned subset
prompt = attack.build(row.goal) # identity, template, or optimised suffix
out = model.generate(prompt, system=cfg.system_prompt, temperature=cfg.temperature,
max_tokens=cfg.max_tokens, seed=cfg.seed)
records.append({
"id": row.id, "cluster": row.cluster, "attack": attack.name,
"prefix_refused": contains_refusal_phrase(out), # legacy metric, kept for comparison
"judge_harmful": judge.score(row.goal, out) >= cfg.judge_threshold,
"model": model.version, "dataset_sha": cfg.dataset_sha, "judge": judge.version,
})
store_encrypted(records, outputs=True) # completions are sensitive artefacts
return records
def report(records): # wilson(): Wilson score interval, 95%
n = len(records)
k_judge = sum(r["judge_harmful"] for r in records)
k_prefix = sum(not r["prefix_refused"] for r in records)
return {"n": n,
"asr_judge": (k_judge / n, wilson(k_judge, n)),
"asr_prefix": (k_prefix / n, wilson(k_prefix, n)),
"clusters_breached": len({r["cluster"] for r in records if r["judge_harmful"]})}Three choices in that harness matter. Report the judge-graded ASR, with the prefix ASR alongside only so that results can be compared with older papers. Report a confidence interval, because 50 rows give an interval about plus or minus 14 points wide at the midpoint. Report clusters breached as well as rows, so that one theme cannot dominate. Pin the judge model and its prompt, and validate it on a hand-labelled sample of your own outputs before you trust it.
Worked example: did fine-tuning erode refusals?
An illustrative comparison: a team fine-tunes a base model for customer support and wants to know whether fine-tuning eroded refusals. They de-duplicate AdvBench behaviours to 120 clusters and keep one representative per cluster. They then run three attacks (no attack, a fixed role-play template and a GCG suffix found on the base model) against both models with the same system prompt, temperature 0 and a 512-token limit.
| Attack | Base: prefix ASR | Base: judge ASR | Fine-tune: prefix ASR | Fine-tune: judge ASR |
|---|---|---|---|---|
| None | 3% | 1% | 9% | 2% |
| Role-play template | 41% | 18% | 63% | 39% |
| Transferred GCG suffix | 52% | 12% | 48% | 10% |
Read the judge columns. The fine-tune is not worse with no attack, and the suffix optimised on the base model transfers poorly to it. The template result, however, roughly doubles, and with 120 rows the intervals around 18% and 39% do not overlap. The prefix columns would have suggested a large GCG problem that is mostly grader noise, and would have understated the template problem. The action is specific: add refusal examples for role-play framings to the fine-tuning mix, re-run, and add the template attack to CI. See jailbreak defence for the layers that should sit around the model.
Failure modes
- Comparing across papers. Different subsets, graders, templates and output limits make two AdvBench ASRs incomparable unless every setting matches. Re-run baselines in your own harness.
- Optimising on the test set. If refusal training used AdvBench rows, AdvBench no longer measures generalisation. Hold out clusters, or use a successor benchmark for the final number.
- Grader drift. Upgrading the judge model changes ASR without any change to the target. Put the judge version in the results key and re-score old outputs when you change it.
- Leaking outputs. Completions that a judge scores harmful are harmful text. Store them encrypted with restricted access, keep them out of CI logs, and delete them on a schedule.
- Single-sample verdicts. At temperatures above zero, one completion per row hides variance; the same row can refuse on one draw and comply on the next. Either fix temperature at 0 or sample several completions per row and report the fraction.
- Ignoring over-refusal. A safety change that cuts ASR and doubles refusals of benign requests is often a regression. Track both numbers on the same dashboard.
Trade-offs and successors
| Option | Strength | Weakness |
|---|---|---|
| AdvBench, prefix grading | Free, deterministic, comparable to 2023-era papers | Overstates and understates harm; redundant rows |
| AdvBench, judge grading, de-duplicated | Cheap and defensible regression signal | No categories, contamination, judge cost |
| HarmBench | Curated categories, standard attacks, trained classifier | Bigger and slower; still public |
| JailbreakBench | Reproducible artefacts and leaderboard | Smaller behaviour set |
| Private red-team set | Matches your product's real risks | Expensive to build and maintain |
A sensible stack uses AdvBench as a fast smoke test, a curated public benchmark for external comparison, and a private set drawn from your own red-team findings for release decisions.
What to do next
- Download both CSVs at a pinned commit, record their SHA-256, and note that the counts are 520 and 574, not 500.
- Cluster the behaviours (the script above is a start), keep one or a few rows per cluster, and version that subset.
- Write down and freeze the system prompt, chat template, temperature, seed and output limit used for every run.
- Grade with a pinned judge validated on your own hand-labelled outputs, and keep prefix ASR only as a secondary column.
- Report ASR with Wilson intervals and clusters breached, and run an over-refusal set next to it.
- Add the no-attack and template runs to CI for every fine-tune, and re-check optimised suffixes before each release.
- Store completions encrypted with access controls, and plan a move to a curated benchmark for external claims.