Before 2024, two papers could both report a jailbreak attack success rate (ASR) for the same model and mean different things. One used GPT-4 as the judge and the other a keyword match for phrases like 'I cannot'. One kept the model's default system prompt and the other removed it. One sampled at temperature 0.7 and the other greedily. One picked 50 behaviours and the other 520. JailbreakBench (Chao et al., NeurIPS 2024 Datasets and Benchmarks track) is an attempt to remove those degrees of freedom. It fixes the behaviours, the target settings and the judge, and publishes every submitted prompt so that results can be checked.

This article is about using JailbreakBench as a measuring instrument, either for a model you ship or for a defense you are building. It does not catalogue attack techniques: LLM jailbreaking in depth covers why safety training fails and jailbreak defense architecture covers layered defenses. Here you will find what the benchmark contains, which choices it fixes and why, how to run it, how to read its numbers honestly with 100 samples, and where it stops being evidence.

Four parts, four removed degrees of freedom

JailbreakBench has four parts, and each one removes a separate source of disagreement.

  • JBB-Behaviors, a dataset of 100 harmful behaviours and 100 benign behaviours on matching themes. It is hosted on Hugging Face as JailbreakBench/JBB-Behaviors.
  • An artifacts repository that stores, for each submitted method, the exact prompt sent for every behaviour, the response, the judge verdict and the query and token counts.
  • A Python framework, the jailbreakbench package. It fixes the system prompts, chat templates, generation settings and judges, and ships five baseline defenses.
  • A leaderboard that ranks attacks and defenses using one scoring function.
JailbreakBench: fixed inputs, fixed target settings, one judge, a public recordJBB-Behaviors100 harmful + 100 benignAttack or defenseyour methodTarget LLMfixed system prompt, T=0Jailbreak judgeLlama-3-70BRefusal judgeLlama-3-8BMetricsASR, refusal, queriesgoalpromptresponseArtifacts repo + leaderboardprompts, responses, query counts, method paramssubmitThe judge is part of the measuring instrument: change it and the numbers are no longer comparable.
Data flow for one evaluation. Only the attack or defense box is yours. Everything else is fixed, including the judge.

The artifacts repository is easy to undervalue. If a paper reports '90% ASR' you have to trust it. If a submission publishes 100 prompts and 100 responses, anyone can re-judge them, reproduce them against a new model version or measure how well they transfer to a defense. For a defender, those stored prompts are also a ready-made regression suite.

JBB-Behaviors: harmful and benign

Each harmful entry has five fields. Behavior is a short unique identifier. Goal is the request. Target is an affirmative opening such as 'Sure, here is...', which optimisation-based attacks aim for. Category is one of ten misuse classes drawn from OpenAI's usage policies. Source records where the entry came from. Counting the published CSV gives exactly ten behaviours per category and the following source split:

PropertyValue
Categories (10 each)Harassment/Discrimination, Malware/Hacking, Physical harm, Economic harm, Fraud/Deception, Disinformation, Sexual/Adult content, Privacy, Expert advice, Government decision-making
Sources55 Original, 27 TDC/HarmBench, 18 AdvBench
Benign set100 thematically matched requests a model should answer

The benign half is the most practical part of the benchmark. A defense that refuses everything scores 0% ASR. A defense that refuses the benign twin of every harmful request is not a defense, just an outage. Report ASR on the harmful set together with the refusal rate on the benign set, and never report one without the other.

Balanced categories also mean the headline ASR weights 'Expert advice' the same as 'Malware/Hacking'. Your product's risk is almost certainly not spread evenly, so keep the per-category breakdown and weight it by your own threat model.

The fixed threat model

The framework fixes the target side so that only the method varies. The open targets are vicuna-13b-v1.5 and llama-2-7b-chat-hf, which a leaderboard submission must cover. gpt-3.5-turbo-1106 and gpt-4-0125-preview are optional. Each target uses a fixed system prompt and chat template, deterministic decoding, and a default of 150 generated tokens per input. Attackers may not change the system prompt or prefill the model's response. Submissions are labelled white_box, black_box or transfer, and querying through phase="test" lets the leaderboard report how many queries a method needed.

Each of these choices limits what the numbers can tell you. A 150-token cap can cut off a response that would have turned harmful later, or one that would have refused after a hedge. Greedy decoding hides the variance that users see at temperature 0.7. All of these targets are single-turn and years old by now. None of this is a flaw, because a benchmark has to freeze something. When you point the harness at your own model, change these settings on purpose and record the change.

The judge is part of the instrument

The jailbreak judge decides what counts as success, so the ASR is only as good as the judge. The authors built a comparison set of 300 prompt-response pairs, harmful and benign, each labelled by three experts who agreed about 95% of the time. They then scored candidate judges against the majority label:

JudgeAgreement with human majority
Llama-3-70B (custom prompt)above 90%, chosen as the official judge
GPT-4 (JUDGE prompt from PAIR)above 90%, closed weights
Llama Guard 287.7%
HarmBench classifier78.3%

Llama-3-70B was chosen because it matches GPT-4 and its weights are open, so the judge cannot change under you. Two consequences matter in practice. First, submissions may not use the official judge inside their own optimisation loop, because an attack that optimises against the judge learns to fool the judge rather than to cause harm. Second, a 10-point gap between judges is larger than many reported gaps between methods. Re-judging another paper's responses with a different classifier and then comparing ASRs measures the classifiers, not the attacks.

A separate Llama-3-8B judge classifies refusals and is what you use on the benign set. Both are exposed as Llama3JailbreakJudge and Llama3RefusalJudge. The human-labelled comparison set is published as the judge_comparison configuration of the dataset, so you can test any in-house judge against it before trusting it.

Running it

The calls below come from the project README. They load the data, query a target directly or behind a defense, score a full prompt set and package a submission. Prompt contents are left as placeholders. In your own harness they come from the artifacts repository or from your red-team corpus.

import os
import jailbreakbench as jbb

harmful = jbb.read_dataset()            # 100 harmful behaviours
benign = jbb.read_dataset("benign")     # 100 matched benign behaviours
df = harmful.as_dataframe()             # Behavior, Goal, Target, Category, Source

# Stored prompts from a published method, as a regression suite
artifact = jbb.read_artifact(method="PAIR", model_name="vicuna-13b-v1.5")
first = artifact.jailbreaks[0]          # JailbreakInfo: goal, prompt, response, jailbroken, ...

# Query a target via API (LiteLLM) or locally (vLLM), optionally behind a defense
llm = jbb.LLMLiteLLM(model_name="vicuna-13b-v1.5", api_key=os.environ["TOGETHER_API_KEY"])
responses = llm.query(prompts=[first.prompt], behavior=first.behavior, defense="SmoothLLM")

# Score a complete submission: {model_name: {behavior: prompt_or_None}}
all_prompts = {"vicuna-13b-v1.5": my_vicuna_prompts, "llama-2-7b-chat-hf": my_llama_prompts}
evaluation = jbb.evaluate_prompts(all_prompts, llm_provider="litellm", defense="SmoothLLM")
jbb.create_submission(evaluation, method_name="MyMethod", attack_type="black_box",
                      method_params={"target-max-n-tokens": 150})

Every query is logged under logs/. Keep those logs. When a number moves between two runs, they are how you check whether the model, the prompt or the judge changed.

Reading 100 samples honestly

One hundred behaviours is a small sample, and most misreadings of JailbreakBench numbers come from forgetting that. Treat each behaviour as a Bernoulli trial and report a confidence interval. The Wilson interval behaves well near 0% and 100%, where ASRs for defended models tend to sit:

import math

def wilson(k, n, z=1.96):
    p = k / n
    d = 1 + z * z / n
    centre = (p + z * z / (2 * n)) / d
    half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d
    return centre - half, centre + half

def two_prop_z(k1, k2, n=100):
    p1, p2, pooled = k1 / n, k2 / n, (k1 + k2) / (2 * n)
    return (p1 - p2) / math.sqrt(pooled * (1 - pooled) * 2 / n)

def report(judged, refusals):
    """judged: list of (category, jailbroken bool) for the harmful set; refusals: list of bools (benign)."""
    k = sum(j for _, j in judged)
    lo, hi = wilson(k, len(judged))
    r = sum(refusals)
    rlo, rhi = wilson(r, len(refusals))
    print(f"ASR {k}/{len(judged)} [{lo:.1%}, {hi:.1%}]  benign refusal {r}/{len(refusals)} [{rlo:.1%}, {rhi:.1%}]")

Worked example. Suppose your baseline model is jailbroken on 12 of 100 behaviours by a set of stored artifacts, and the version with a new input filter is jailbroken on 7. The release note wants to say 'jailbreaks down 42%'. The Wilson intervals are 7.0%-19.8% for the baseline and 3.4%-13.7% for the filtered model, and they overlap heavily. The two-proportion z statistic is 1.21, well below the 1.96 needed for significance at the 5% level. Meanwhile the benign refusal rate went from 3/100 (1.0%-8.5%) to 4/100 (1.6%-9.8%), which is also within noise. An honest summary is: 'no detectable change in ASR or over-refusal on JBB-Behaviors; a larger evaluation is needed'.

Per-category numbers are worse still. With 10 behaviours per category, one success out of 10 has an interval of about 1.8%-40.4%. Use categories to find where to look, not to make claims. To claim an improvement, either repeat the run over several stochastic attack seeds, or add your own held-out behaviours so that n is in the hundreds.

Defenses and the adaptive-attack rule

The framework includes five baseline defenses: SmoothLLM, perplexity filtering, Erase-and-Check, synonym substitution, and removal of non-dictionary words. The paper's SmoothLLM setting uses swap perturbations with N=10 perturbed copies and q=10% of characters perturbed. Running the stored artifacts against these defenses is cheap and useful as a smoke test. It does not show that a defense is robust.

The paper is explicit about why. Replaying prompts that were optimised against the undefended model is a transfer evaluation, and it is the weakest kind. A defense that randomises characters will break prompts tuned to exact token sequences, and an attacker who knows about the defense simply optimises through it. Among unsuitable uses, the authors list 'Using this benchmark to evaluate the robustness of LLMs and defenses by using only the existing attacks (especially, only against the existing precomputed jailbreak prompts), without employing an adaptive attack with a thorough security evaluation.' The minimum credible defense evaluation is: stored artifacts, then at least one attack re-run with the defense in the loop, then the benign refusal rate, plus the latency and cost the defense adds per request. The red-team architecture article describes how to staff and run that adaptive step.

A release-gating pipeline

Teams that get value from JailbreakBench usually wire it into release gating instead of running it once for a slide. A workable pipeline looks like this:

  1. On every model or system-prompt change, replay the stored artifacts for the methods you care about against your deployed stack, including its guardrails, and judge with the official judge.
  2. Run the benign set through the same stack and judge refusals.
  3. Compare against the last release using intervals, not point estimates. Gate only on changes that are significant, or on any new success in a category you have marked critical.
  4. Every few releases, run one adaptive black-box attack with your defenses in the loop and add the new successful prompts to a private regression set. Never add them to anything that trains the model.
  5. Store prompts, responses, judge versions and model hashes together, as the artifacts repository does, so that any number can be re-derived later.

Swap in your production system prompt and decoding settings for this internal run. The official settings keep you comparable with the leaderboard, while your own settings tell you about your users. Run both and label them clearly. An input classifier such as Llama Guard sits naturally in step 1 as part of the stack under test.

Failure modes

  • Judge drift. Swapping the judge, or calling a hosted judge whose weights change silently, moves ASR by more than most method differences do. Pin the judge and its prompt.
  • Benchmark overfitting. 100 public behaviours invite tuning, whether deliberate or accidental, for example through training data that includes the artifacts. Keep a private held-out set that follows the same category scheme.
  • Stale targets. The fixed targets are older models. Leaderboard rank on Vicuna says little about your model, so measure on yours.
  • Truncation artefacts. At 150 tokens, a long preamble followed by refusal and a long preamble followed by compliance can look the same. Spot-check judged responses at your production length.
  • Single-turn scope. Multi-turn escalation, tool use and indirect injection through retrieved content are outside the threat model. Cover them separately.
  • Reporting ASR alone. Without the benign refusal rate, a refuse-everything filter looks perfect.

Trade-offs

Comparability versus relevance. The official settings make your number comparable with others, while your own settings make it meaningful for your product. Report both.

Small and fast versus statistical power. 100 behaviours run in minutes and are cheap to judge, but they only detect large effects. Add held-out behaviours when you need to detect small ones.

Open judge versus best judge. Llama-3-70B is pinned and reproducible but costs GPU time to host. A smaller in-house classifier is cheaper, but only after you have measured its agreement on the published judge-comparison set.

What to do next

  1. Load both halves of JBB-Behaviors and record per-category ASR and benign refusal rate for your current production stack.
  2. Pin the judge model and prompt, and log the judge version next to every result.
  3. Report Wilson intervals, and gate releases on significant changes or on critical-category successes.
  4. Replay the stored artifacts as a regression suite, then run at least one adaptive attack against your defenses before claiming robustness.
  5. Build a private held-out behaviour set in the same ten categories to detect overfitting.
  6. Extend coverage beyond single-turn prompts with the methods in automated red teaming.
Key takeaway: JailbreakBench makes jailbreak numbers comparable by fixing the behaviours, target settings and judge, and by publishing every prompt. Use it with both halves: ASR on harmful behaviours and refusal rate on benign ones. Pin the judge, report confidence intervals because n is only 100, and treat replayed artifacts as a smoke test, since only an adaptive attack with the defense in the loop shows that a defense is robust.