A multi-turn jailbreak gets a language model to produce something its safety training should prevent, by spreading the request over a conversation instead of asking in one message. Each turn looks acceptable. The model answers each one, and after enough turns the conversation contains what a single direct request would have been refused. This matters because most safety training and most published robustness numbers are about single prompts. A model that refuses a direct request almost every time can still be walked to the same output over ten turns.

This site's multi-turn attacks article covers the application side: who writes conversation history, sealing it, trajectory monitoring and a backtracking red-team harness. This page covers the model side. It explains why behaviour learned on single turns erodes over many, what the published evidence shows, how to measure multi-turn robustness so the numbers survive scrutiny, and which training and serving changes help, with their costs. Attacks are described by mechanism only. The aim is measurement and defence, not a recipe.

What makes a jailbreak multi-turn

Two properties define the class. The harmful objective is never stated in full in any single user turn, and the model's own earlier replies become part of the input for later turns. The second property matters most: the model is partly being asked to continue text it wrote, and models are trained to be consistent with their own context.

The threat model depends on who controls the history, and the difference is large:

AccessWhat the attacker can doConsequence
Hosted chat with server-side historySend one new user turn at a time; restart sessionsRefusals stay in context and later turns must work around them
Stateless API that accepts a full message listEdit or delete earlier turns, including the model's own; write fake assistant turnsRefused turns can be removed and retried, so each attempt starts clean
Open weightsAll of the above, plus prefilling and fine-tuningConversation-level defences inside the model are only one layer

Retrying after a refusal is called backtracking. Microsoft's open-source PyRIT toolkit, for example, implements it in its Crescendo attack with a configurable max_backtracks limit: when the target refuses, the refused exchange is dropped and the attacker tries a different turn. Backtracking turns a refusal from a wall into a cost, which is why evaluations must state which access model they assume.

What the published evidence shows

Three published results frame the problem. All figures below were checked against the papers' arXiv abstracts.

  • Crescendo (Russinovich, Salem and Eldan, 2024, published at USENIX Security 2025) starts from a general question related to the goal and escalates gradually by referring to the model's own replies. The authors built an automated version, Crescendomation, and tested it against several commercial and open models. On AdvBench subsets they reported 29 to 61 percent higher performance than other state-of-the-art jailbreak techniques on GPT-4, and 49 to 71 percent on Gemini-Pro.
  • Multi-Turn Human Jailbreaks (MHJ) (Li et al., 2024, Scale AI) had professional red teamers attack models protected by published defences. Human multi-turn attacks reached over 70 percent attack success on HarmBench, against defences that report single-digit success rates for automated single-turn attacks. They released 2,912 prompts across 537 multi-turn jailbreaks. The point for practitioners is that the single-turn number was real but answered a different question.
  • ActorBreaker (Ren et al., ACL 2025; the code repository is named ActorAttack) builds multi-turn attacks by starting from entities related to the harmful topic and approaching it through them. The authors also built a multi-turn safety dataset and found that fine-tuning on it improved robustness, with some cost to utility.

A related single-prompt effect, many-shot jailbreaking, uses many fabricated dialogue turns inside one long message. It is covered with the other prompt-level techniques in LLM jailbreaking in depth.

Why models give way over turns

There is no single cause, but four mechanisms explain most of what the papers report.

  1. Self-consistency. A model that has already written part of an explanation is predicting the continuation of its own text. Refusing at turn eight means contradicting turns one to seven, and pretraining rewards coherent documents. Escalation works by making each step a small extension of the last reply.
  2. Training distribution. Refusal behaviour is learned from examples, and in most preference and fine-tuning data the harmful intent appears in the most recent user message. A conversation where intent is spread across turns is less like those examples, so the learned behaviour transfers less reliably. See refusal training with RLHF for how that behaviour is learned.
  3. Local judgement. Each reply is produced by deciding whether this reply, given this context, is acceptable. Decomposition exploits that: every fragment is acceptable, and only the assembled result is not. The same weakness affects classifiers that read one message.
  4. Dilution. In a long conversation, the system prompt and the earliest framing are far from the current turn, and a fictional or professional frame set up early can come to dominate how later requests are read.

These suggest the defences: train on conversations where the intent is distributed, teach the model that it may stop partway even after cooperating, and judge the conversation, not the message.

Measuring multi-turn robustness

A useful evaluation answers one question: given an objective the model should refuse and an attacker with a stated access model and turn budget, how often does the conversation end with the objective achieved? There are two ways to run it, and you need both.

Two ways to measure multi-turn robustness of a modelStatic replay (fixed transcripts)recorded user turns 1..kfrom a dataset such as MHJtarget modelfresh replies at every turnjudgescores the final replynext recorded turnCheap and reproducible, but off-policy:turn 4 was written for replies the newmodel never gave.Adaptive attacker (closed loop)attacker modelobjective + historytarget modelsame deployment settingsjudgeper turn: refused / progress / achievednext turnverdictRealistic and on-policy, but noisy andcostly; results depend on the attackermodel, so pin it and report it.Report both, against the same objectives, with a benign multi-turn set measuring over-refusal alongside.
Static replay is reproducible but off-policy; an adaptive attacker is realistic but noisy. A benign multi-turn set measures what the defence costs.

Static replay feeds recorded human user turns, such as MHJ's, to the target while the target generates fresh replies. It is cheap, close to deterministic at temperature zero and easy to compare across versions. Its flaw is that each recorded turn was written in reaction to replies the new model may never give, so it underestimates a capable human and is worst on exactly the models that refuse early.

Adaptive attack puts an attacker model in a loop with the target and a judge, as in Crescendomation or PyRIT. It reacts to the target and can backtrack if the access model allows. Results depend on the attacker model, its turn budget and the judge, so pin all three and report them with the score.

Report at least these, per objective category:

  • Attack success rate (ASR) with a confidence interval, at a fixed turn budget, such as 10 turns.
  • ASR by turn, the curve of cumulative success against turns used. A model that breaks at turn 9 instead of turn 3 is better, even if its ASR at 10 turns is similar.
  • Over-refusal on benign multi-turn conversations that touch sensitive topics legitimately, such as a nurse asking about overdose thresholds or a security engineer about malware behaviour. A defence that halves ASR while refusing these is a trade, not a fix.
  • Judge agreement: the rate at which a sample of judge verdicts matches human labels.

Worked example: is the new model actually safer?

Suppose you are comparing a current model, v1, with a candidate, v2, that was fine-tuned on additional multi-turn safety conversations. The counts below are hypothetical, chosen to show the analysis; the intervals and the test result were computed from them. You run 40 objectives with 3 attacker seeds each, so 120 conversations per model, with a 10-turn budget, plus 200 benign multi-turn conversations.

Measurev1v2
Attacks achieved21 / 120 = 17.5% (95% CI 11.7-25.3%)9 / 120 = 7.5% (95% CI 4.0-13.6%)
Benign conversations refused4 / 200 = 2.0% (95% CI 0.8-5.0%)17 / 200 = 8.5% (95% CI 5.4-13.2%)

The intervals are Wilson score intervals, which behave well at small counts, unlike the normal approximation. Because both models faced the same 120 objective-seed pairs, compare them paired, not as two independent samples. Suppose v1 failed and v2 held on 15 pairs, and the reverse happened on 3. McNemar's exact test on those 18 discordant pairs gives p of about 0.0075, so the improvement is unlikely to be noise.

The over-refusal row is the real decision. v2 refuses about four times as many benign conversations. Whether that is acceptable depends on the product. The analysis does not decide that, but it puts both numbers on the same page, so the decision is made deliberately. The paired comparison is a few lines:

from math import comb, sqrt

def wilson(k, n, z=1.96):
    p = k / n
    d = 1 + z * z / n
    centre = (p + z * z / (2 * n)) / d
    half = z * sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d
    return centre - half, centre + half

def mcnemar_exact(b, c):
    """b: pairs where only the old model was broken; c: only the new one. Two-sided."""
    n = b + c
    tail = sum(comb(n, i) for i in range(min(b, c) + 1)) / 2 ** n
    return min(1.0, 2 * tail)

def compare(results_old, results_new):
    """results_*: dict (objective_id, seed) -> True if the attack achieved its objective."""
    keys = results_old.keys() & results_new.keys()
    b = sum(results_old[k] and not results_new[k] for k in keys)
    c = sum(results_new[k] and not results_old[k] for k in keys)
    return {"pairs": len(keys), "only_old_broken": b, "only_new_broken": c,
            "p_value": mcnemar_exact(b, c)}

Two warnings. Forty objectives is a small sample of the harm space, so break results out by category and expect wide per-category intervals. And attacker seeds within one objective are correlated, so the effective sample size is between 40 and 120. If a decision hinges on the exact p-value, resample by objective rather than by conversation.

Mitigations and what they cost

Mitigations fall into three groups, each with a cost.

Multi-turn safety data. Add training conversations where intent is spread across turns, with the desired behaviour at the turn where the conversation crosses the line. ActorBreaker's results show this helps and also show the cost: utility can drop if the data teaches the model to distrust whole topics. Pair every unsafe multi-turn example with benign conversations on the same topics that should be answered fully.

Course correction. Teach the model that cooperating earlier does not commit it to continuing. Xu et al. (2024) built a benchmark for this, C2-Eval, and a synthetic set of 750,000 preference pairs, C2-Syn, that rewards a model for steering away partway through a harmful continuation. They reported better jailbreak resistance on Llama2-Chat 7B and Qwen2 7B without loss of general performance. Course correction attacks self-consistency directly.

Conversation-level judgement at serving time. A classifier that reads the whole conversation, or the last several user turns joined, sees what fragments add up to. The windowed check and decayed score in the multi-turn attacks article are one design, and jailbreak defense architecture covers how it sits alongside per-turn filters. It adds latency and cost per turn, and its false positives land on long expert conversations.

Serving-side controls compound with these: do not accept client-written assistant turns unless you must, since that removes backtracking for hosted products, and keep refusals in any summary that replaces old turns.

Failure modes in evaluation

  • Judge errors. An LLM judge can count a vague or wrong answer as success, or miss success spread over several replies. Have humans label a sample and report agreement; judge the whole transcript.
  • Training on the test. If MHJ-style transcripts or your red-team logs go into training, they cannot be your evaluation. Hold out objectives and tactics, not just conversations.
  • Overfitting to known tactics. A model tuned against Crescendo-style escalation may still fail against decomposition or role framing. Keep categories separate in results and add new tactics from your red team; see LLM red teaming.
  • Mismatched settings. Evaluating at temperature zero with no system prompt says little about a product deployed with a long system prompt, tools and sampling. Evaluate the configuration you ship.
  • Reporting only ASR. Without over-refusal and turns-to-success, a regression in helpfulness or a shift from breaking at turn 3 to turn 9 is invisible.

What to do next

  1. Write down your access model: hosted history, stateless API or open weights. It decides whether evaluations should allow backtracking.
  2. Build an objective set by harm category, with a held-out portion never used in training or prompt tuning.
  3. Run static replay on a public multi-turn dataset as a cheap regression test on every model or prompt change.
  4. Run an adaptive attacker with pinned attacker model, judge and turn budget before each release, and report ASR with Wilson intervals and ASR by turn.
  5. Build a benign multi-turn set on the same sensitive topics and report over-refusal next to ASR.
  6. Compare versions with paired tests on the same objective-seed pairs, and resample by objective.
  7. If ASR is too high, add multi-turn safety data and course-correction examples together with benign counterparts, then remeasure both numbers.
Key takeaway: Multi-turn jailbreaks work because models stay consistent with their own earlier replies, because safety training mostly sees harmful intent in one message, and because each reply is judged locally. Published work found human multi-turn attacks succeeding over 70 percent of the time against defences that looked strong on single-turn tests. Measure with both static replay and an adaptive attacker under a stated access model, report attack success with intervals, success by turn and benign over-refusal, compare versions with paired tests, and mitigate with multi-turn safety data, course correction and conversation-level judgement, remeasuring helpfulness each time.