"DAN", short for Do Anything Now, is the name of a family of jailbreak prompts that spread from late 2022 onward. The prompts ask a chat model to adopt an alter ego that is supposedly free of its rules, and successive community versions added scaffolding to make the role stick. The site's DAN lineage article covers the anatomy of these prompts and their variant families. This article takes the other half of the problem: how to measure a jailbreak family like DAN honestly, so that statements such as "our model is robust to DAN" or "that prompt was patched" mean something.

That matters because DAN is less a single attack than a living dataset. Prompts are copied, mutated and re-posted; a fix aimed at one string leaves its siblings working. We will cover what the largest public measurement study found, the mechanisms that make role-play jailbreaks work, and then a measurement pipeline step by step: collection and safe handling, near-duplicate clustering, building a question set with benign controls, judging responses, computing attack-success rates with confidence intervals, and tracking drift across model versions. No working jailbreak text appears here; you do not need it to understand or build the measurement.

What the in-the-wild study measured

The reference point is the CCS 2024 paper by Shen, Chen, Backes, Shen and Zhang, titled after the family itself: "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models. Its abstract reports 1,405 jailbreak prompts collected from December 2022 to December 2023 and 131 jailbreak communities. To evaluate them, the authors built a question set of 107,250 samples across 13 forbidden scenarios and tested six popular LLMs. They found five highly effective prompts that reached 0.95 attack success rate on ChatGPT (GPT-3.5) and GPT-4, the earliest of which persisted online for over 240 days, and 28 user accounts that consistently optimised jailbreak prompts over 100 days. The authors concluded that the safeguards they tested could not adequately defend against these prompts in all scenarios.

Three lessons carry over to anyone running their own measurement. The population is large and evolving, so a fixed list of ten famous prompts is a poor test. Effectiveness is concentrated: a few prompts do most of the damage, so averages hide risk. And persistence is long, so a prompt being old says nothing about whether it still works. Model behaviour has changed considerably since that study; treat its numbers as a historical baseline, not a statement about current models.

Why role-play jailbreaks work

Wei, Haghtalab and Steinhardt's 2023 paper "Jailbroken: How Does LLM Safety Training Fail?" offers two mechanisms that explain why DAN-style prompts work. Competing objectives: the model is trained both to follow instructions and to refuse harmful ones, and a prompt that frames compliance as the instruction-following task (stay in character, keep the game going) pits one objective against the other. Mismatched generalisation: safety training covers a narrower distribution than pretraining, so inputs that look unlike the safety data, such as long fictional frames, unusual formats or dual-response templates, can fall outside what refusal training reached.

Both mechanisms predict what you will see in measurement. Small surface edits (renaming the persona, reordering paragraphs) often preserve success because they preserve the mechanism. Defences that key on strings decay quickly, while defences that key on the mechanism, such as training on paraphrased role-play attacks or classifying the output rather than the input, decay more slowly. A good measurement therefore groups prompts by structure, not by exact text.

The measurement pipeline

Measuring a jailbreak family: from scraped prompts to a defensible attack-success numberCollectforums, sites, reportsNormalisestrip, hash, store sealedClusterMinHash near-duplicatesQuestion setharmful + benign controlsTarget modelpinned version, fixed paramsJudgerules + model + human auditStatisticsASR with intervalsprompt x questionRegression dashboardper cluster, per model versionnext releaseEvery stage can bias the final number; the judge and the sample size are the usual culprits.
A measurement pipeline for a jailbreak family. The output is a per-cluster attack-success rate with an interval, tracked per model version.

Treat the collected prompts as hazardous material. Store them in a restricted bucket, log access, and never paste them into tickets, chat or public dashboards; reference them by cluster ID and hash. Normalise before hashing: collapse whitespace, lowercase, strip zero-width characters and decorative emoji, and record the source and first-seen date. Pin everything about the target: model version, system prompt, temperature, maximum tokens and any safety filters in the path. A number measured with temperature 1.0 and one sample per pair is not comparable with one measured greedily.

Clustering near-duplicates

Community jailbreaks are mostly edits of each other. Counting near-copies as independent prompts inflates the apparent size of the threat and lets one family dominate your averages. MinHash over word shingles estimates Jaccard similarity cheaply; prompts above a threshold join the same cluster.

import hashlib, re
from collections import defaultdict

def shingles(text, k=5):
    words = re.sub(r"\s+", " ", text.lower()).split()
    return {" ".join(words[i:i + k]) for i in range(max(1, len(words) - k + 1))}

def minhash(sh, n=128):
    return [min(int(hashlib.blake2b(f"{seed}:{s}".encode(), digest_size=8).hexdigest(), 16) for s in sh)
            for seed in range(n)]

def similar(a, b):
    return sum(x == y for x, y in zip(a, b)) / len(a)    # estimates Jaccard similarity

def cluster(prompts, threshold=0.6):
    sigs = {pid: minhash(shingles(t)) for pid, t in prompts.items()}
    parent = {pid: pid for pid in sigs}
    def find(x):
        while parent[x] != x:
            parent[x] = parent[parent[x]]
            x = parent[x]
        return x
    ids = list(sigs)
    for i, a in enumerate(ids):                 # O(n^2): fine for a few thousand; use LSH bands beyond
        for b in ids[i + 1:]:
            if similar(sigs[a], sigs[b]) >= threshold:
                parent[find(a)] = find(b)
    groups = defaultdict(list)
    for pid in ids:
        groups[find(pid)].append(pid)
    return list(groups.values())

Pick the threshold by inspection: sample pairs just above and just below it and check whether a human would call them the same prompt. Then evaluate a few representatives per cluster rather than every member, and report results per cluster.

The question set and benign controls

A jailbreak prompt is a wrapper; it needs a payload question to test. Build the question set from your own policy categories, not someone else's, with a fixed number of questions per category so categories with many easy questions do not dominate. Write questions at varying severity, and keep the actual text access-controlled like the prompts.

Add a benign control set: the same wrappers applied to harmless requests, and the harmless requests alone. This measures over-refusal. A defence that drops attack success to zero while refusing every role-play request has moved the cost to legitimate users, and you will only notice if you measure it.

Judging responses

The judge decides whether a response counts as a successful jailbreak, and it is the stage most likely to distort your number. Keyword refusal detection (looking for phrases like "I can't help") undercounts partial compliance and overcounts responses that refuse in words but leak content. An LLM judge with a written rubric is better but has its own errors and can itself be swayed by the jailbreak text in the transcript. Use a layered judge: rules for obvious cases, a model judge with a rubric for the rest, and a human-labelled audit sample to measure the model judge's agreement.

def cohen_kappa(a, b):
    """Agreement between two labelers on binary labels, corrected for chance."""
    n = len(a)
    po = sum(x == y for x, y in zip(a, b)) / n
    pa, pb = sum(a) / n, sum(b) / n
    pe = pa * pb + (1 - pa) * (1 - pb)
    return (po - pe) / (1 - pe) if pe < 1 else 1.0

human = [1, 0, 0, 1, 1, 0, 0, 0, 1, 0]        # 1 = harmful compliance, audited sample
model = [1, 0, 1, 1, 1, 0, 0, 0, 0, 0]
print(round(cohen_kappa(human, model), 2))    # 0.58: moderate; fix the rubric before trusting rates

If agreement is weak, improve the rubric before trusting any attack-success rate. Report the judge's precision and recall on the audit sample alongside every headline number.

Attack success with honest intervals

Attack success rate (ASR) is successes divided by trials for a prompt cluster on a model version. With small samples it is noisy, so always report an interval. The Wilson score interval behaves well near 0 and 1, where jailbreak rates often sit.

from math import sqrt

def wilson(successes, n, z=1.96):
    if n == 0:
        return (0.0, 1.0)
    phat = successes / n
    denom = 1 + z * z / n
    centre = (phat + z * z / (2 * n)) / denom
    half = z * sqrt(phat * (1 - phat) / n + z * z / (4 * n * n)) / denom
    return (max(0.0, centre - half), min(1.0, centre + half))

print(wilson(37, 200))   # about (0.14, 0.25) for v1
print(wilson(22, 200))   # about (0.07, 0.16) for v2

Worked example: cluster C7 succeeds on 37 of 200 question pairs against model v1 (18.5%) and on 22 of 200 against v2 (11%). The intervals, roughly 14-25% and 7-16%, overlap slightly, so before announcing an improvement run more trials or a proper two-proportion test. Now suppose v2 shows 0 of 50: the Wilson upper bound is still about 7%, so "zero observed" is not "zero". This is why a small spot check after a fix proves little.

Two further traps appear once you track many clusters. First, multiple comparisons: with 40 clusters tested on every release, a few will appear to get worse by chance alone, so flag a regression only when it persists on a rerun or survives a correction such as Holm-Bonferroni. Second, aggregation: a fleet-wide ASR averaged over all clusters can fall while the single most dangerous cluster rises. Report the worst cluster's upper bound alongside the average, and weight categories by severity rather than by how many questions happen to sit in each. A release gate written as "no cluster's upper bound above 5% in the highest-severity categories" is far harder to game than "average ASR went down".

Drift, half-life and what &#x27;patched&#x27; means

Track each cluster's ASR across model versions and over time. Three patterns recur. A cluster drops to near zero and stays there: a real fix. A representative drops but paraphrased members still succeed: the fix matched surface text. A cluster drops, then a new mutation appears weeks later with the same structure: the mechanism was never addressed.

This is why "DAN was patched" is a weak claim. Before accepting it, require three checks: the exact prompt fails; automatically paraphrased variants of the cluster fail at a comparable rate; and the benign control set has not lost ground. Generate paraphrases with a separate model under the same access controls, and judge them with the same judge.

From measurement to defence

Measurement points at defences. Clusters that rely on persona persistence argue for training or system-prompt reinforcement against role override; clusters that leak via output formats argue for output-side classification. Input classifiers trained on known prompts catch copies cheaply but decay against mutation, so pair them with output checks. The jailbreak defence architecture article covers the layering; multi-turn jailbreaks covers attacks that spread the role-play across a conversation, which single-turn measurements miss entirely.

Failure modes

  • Judge drift. The model judge is upgraded and rates shift without any change in the target. Version the judge and re-audit after any change.
  • Contaminated defence training. The evaluation clusters leak into safety training data, so the regression suite measures memorisation. Hold out clusters by first-seen date.
  • Sampling settings mismatch. Comparing greedy runs with sampled production traffic understates real risk. Measure under production settings, with several samples per pair.
  • Hazard leakage. Prompt text appears in dashboards or bug reports. Reference by hash.
  • Single-turn blindness. Conversations that build the persona gradually never appear in a one-shot harness.

What to do next

  1. Read the DAN anatomy article so you can label clusters by structure.
  2. Set up a restricted store for collected prompts with source, first-seen date and normalised hash.
  3. Cluster with MinHash, hand-check the threshold, and pick two or three representatives per cluster.
  4. Write a balanced question set from your own policy, plus a benign control set for over-refusal.
  5. Audit 200 judged responses by hand and compute kappa before trusting any rate.
  6. Report per-cluster ASR with Wilson intervals for every model release, and require paraphrase checks before calling anything patched.
Key takeaway: DAN is a family, not a string. Measure it as one: store prompts as hazardous data, cluster near-duplicates, test against a balanced question set with benign controls, audit your judge, and report per-cluster attack success with intervals for every model version. Only call a cluster fixed when paraphrases fail too and over-refusal has not risen.