A jailbreak is an input that gets a language model to produce something its developer trained or instructed it to refuse. Every widely deployed model has been jailbroken, usually within days of release, and the pattern has repeated for years. That persistence is the most important fact about the subject. It tells you jailbreaks are not a list of bugs to be patched one by one, but a consequence of how models are built.
This article explains why, using the research that has clarified the mechanisms, then covers what attackers can do at each level of access, why attacks are a search process, how to run that search against your own product, and how to size your product's risk. It contains no working attack prompts; the goal is understanding and defense. The layered defense architecture itself is covered in jailbreak defense architecture.
What a jailbreak is, and what it is not
Two attacks are often confused. In a jailbreak, the attacker is the person talking to the model, and the target is the model's own behavioural policy: they want output the model should refuse. In prompt injection, the attacker is usually a third party who plants instructions in content the model processes (a web page, an email, a retrieved document), and the target is the application: they want the model to take actions or leak data on the victim's behalf. OWASP's LLM Top 10 treats jailbreaking as a form of prompt injection, and the techniques overlap, but the threat models are different. See indirect prompt injection for the second.
The distinction changes what harm is possible. A jailbroken assistant without tools can only produce text: the risk is the content itself. A jailbroken agent with tools can act, and a jailbreak is often the first step that removes refusals before a request to reveal the system prompt, call a sensitive tool or bypass business rules. Extraction of the system prompt is covered in system prompt leakage.
Why safety training fails: the mechanisms
Pretraining gives a model broad capability, including knowledge of harmful topics in its corpus. Post-training (supervised fine-tuning and preference optimisation) then teaches it to follow instructions and refuse certain requests: a comparatively small adjustment over a much larger capability. Each mechanism below is a way that adjustment fails to cover it.
- Competing objectives. Wei, Haghtalab and Steinhardt ("Jailbroken: How Does LLM Safety Training Fail?", NeurIPS 2023) observed that the model is trained both to be helpful and follow instructions and to refuse harm. Inputs that make compliance look like the instructed, helpful answer (role-play, a required output format, a demand to begin a certain way) set those objectives against each other, and helpfulness often wins.
- Mismatched generalization. The same paper's second mode: capability generalises further than safety training did. A model that can read an encoded, translated or unusually formatted request may never have seen refusal examples in that form, so it understands the request but does not recognise it as one to refuse.
- Shallow alignment. Qi et al. ("Safety Alignment Should Be Made More Than Just a Few Tokens Deep", ICLR 2025) showed that alignment mostly changes the distribution of the first few output tokens. If the response starts with a refusal, it stays a refusal; if something forces a different start (prefilling, an adversarial suffix, altered decoding parameters), the model rarely recovers. They also showed that training to recover from a harmful start makes alignment more robust.
- A single refusal direction. Arditi et al. ("Refusal in Language Models Is Mediated by a Single Direction", NeurIPS 2024) found, across 13 open chat models up to 72B parameters, one direction in the residual stream whose removal stops refusal and whose addition causes refusal of harmless requests. With open weights, that makes removing refusal a cheap, direct weight edit.
- Fine-tuning erosion. Qi et al. ("Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!", ICLR 2024) removed GPT-3.5 Turbo's guardrails by fine-tuning on 10 adversarial examples for under $0.20 through the provider's API, and found that fine-tuning on ordinary benign data also degraded safety, to a lesser degree.
Together these explain the persistence. Patching one phrasing adds refusal examples near that phrasing; the capability underneath, and all the other routes to it, remain.
The attack surface depends on access
What an attacker can do depends on how much control they have. When you threat-model, first place your product on this ladder.
| Access | What the attacker controls | What becomes possible |
|---|---|---|
| Chat UI | User messages only | Framing, role-play, encoding, multi-turn escalation |
| API with long context | Entire message list, possibly assistant turns | Fabricated dialogue history; many-shot attacks |
| API with prefill or decoding control | Start of the model's answer, sampling settings | Directly bypassing the shallow first-token refusal |
| Fine-tuning API | Training data | Erasing refusal with a handful of examples |
| Open weights | Everything | Gradient attacks, refusal-direction removal, arbitrary fine-tuning |
If you ship open weights, assume anyone can remove refusal; your safety case must rest on what the model knows and on downstream controls. If you run an API, each extra control (assistant turns, prefill, fine-tuning) moves users down the ladder, so gate those features by account trust. Many-shot jailbreaking (Anil et al., NeurIPS 2024) shows a feature becoming an attack: long context lets attackers include hundreds of fabricated compliant examples, and effectiveness followed a power law in their number, up to hundreds of shots.
Jailbreaks are found by search
Jailbreaks are increasingly the output of optimisation rather than clever writing, which explains why they keep appearing and how to test for them.
- Manual search. People try framings, observe refusals and iterate. Public jailbreak communities are a distributed version of this.
- Attacker-model search. PAIR (Chao et al., "Jailbreaking Black Box Large Language Models in Twenty Queries") uses one LLM to propose attempts against a target, reads the target's reply and a judge's score, and refines; the paper reports that it often succeeds in fewer than twenty queries with black-box access only. TAP (Mehrotra et al., NeurIPS 2024) adds a tree of candidates and prunes weak ones before querying the target, reducing queries further.
- Gradient search. With open weights, GCG (Zou et al., 2023) optimises a string of tokens appended to a request to maximise the probability that the answer starts affirmatively, which targets exactly the shallow first-token decision. Suffixes found on open models partly transferred to closed ones.
- Multi-turn search. Attacks that escalate gradually over a conversation, each message harmless alone, exploit per-message checking and the model's consistency with its own earlier answers.
An attacker with a query budget and a scoring function will find something. Raise the cost of the search, limit the harm when it succeeds, and detect it while it runs: searches produce many near-identical rejected attempts from one account.
Running the search against yourself
The same loop is the core of an automated red team. The skeleton below has no attack content: the attacker is pluggable, the behaviour list is a governed dataset, and the target is your full production stack, because testing the bare model tells you little about the product.
from dataclasses import dataclass, field
@dataclass
class Attempt:
behavior_id: str
turns: list # the conversation actually sent
response: str
judge_score: float # 0..1 from a separate judge model or classifier
queries_used: int
@dataclass
class Budget:
max_queries_per_behavior: int = 20
max_total_queries: int = 5_000
spent: int = 0
def red_team_run(behaviors, attacker, target, judge, budget, threshold=0.5):
"""Search for inputs that elicit each disallowed behavior from OUR deployment.
behaviors: curated, governed list (e.g. drawn from HarmBench or JailbreakBench
categories plus product-specific policy items), never ad hoc.
attacker: proposes the next attempt from the history of failed attempts.
target: the full production stack (system prompt, filters, tools), not the bare model.
judge: scores whether the response actually provides the behavior.
"""
findings = []
for b in behaviors:
history = []
for q in range(budget.max_queries_per_behavior):
if budget.spent >= budget.max_total_queries:
return findings
turns = attacker.propose(b, history) # refine from feedback
response = target.respond(turns) # through every production layer
score = judge.score(b, turns, response)
budget.spent += 1
history.append((turns, response, score))
if score >= threshold:
findings.append(Attempt(b.id, turns, response, score, q + 1))
break # one finding per behavior is enough to triage
return findingsDesign choices that matter:
- Budget by queries. Report success at 1, 5 and 20 queries per behaviour; holding for one attempt but falling in five stops nobody motivated.
- Separate the judge from the guard. If one classifier both blocks traffic and scores success, its blind spots hide themselves. Judges are covered in LLM safety evaluation.
- Keep findings as regression tests, in a restricted store, run on every model, prompt or filter change.
- Include benign near-misses (medical, security education) so you measure over-refusal too.
Automation measures how known categories generalise; human red-teamers find new ones. See running an LLM red team.
Worked example: threat-modelling one product
Take a customer-support assistant for a consumer bank. It answers questions, can look up the customer's own transactions, and can open a dispute. It runs on a hosted API model with a system prompt, an input filter and an output filter. How much jailbreak risk does it carry, and where should effort go?
Step 1: list harms by who gets hurt. (a) The model produces generally harmful content, such as instructions for fraud, because a user asked for it. (b) The model is talked into opening disputes or making claims outside policy. (c) It reveals the system prompt and internal rules. (d) Screenshots of it saying something offensive damage the brand.
Step 2: rate severity and attacker benefit. (a) offers little beyond a general chatbot: medium. (b) has direct financial value: high. (c) is low alone but helps plan (b). (d) is low effort with high reputational impact.
Step 3: find where refusal is the only control. For (b), if the only thing stopping a fraudulent dispute is the model declining, that is the design flaw. Move the rule out of the model: the dispute tool enforces eligibility in code, requires the transaction to belong to the authenticated customer, and routes amounts over a threshold to a human. Now a successful jailbreak produces a polite request that the tool rejects. For (a) and (d), the model's refusal and the output filter remain the controls, so they get red-team coverage.
Step 4: set a measurable target. For example: no finding from the behaviour list within 20 automated queries per behaviour, over-refusal on the benign set below an agreed rate, and every tool action enforced by code. The exact numbers are a business decision; the point is that they are written down and re-measured on every change.
The principle: never let model refusal be the only barrier before an action with real consequences. Search eventually defeats a probabilistic filter; it does not defeat an authorisation check in code.
Defenses, briefly, and their costs
The defense architecture has its own article; here is how the mechanisms above map to defense choices and what each costs.
| Defense | Addresses | Cost or limit |
|---|---|---|
| Input and output classifiers | Content harms regardless of framing | Latency, false positives; the output filter is the stronger of the two |
| Conversation-level scoring | Multi-turn and many-shot escalation | State to keep; harder to explain blocks |
| Limiting API controls (prefill, fabricated assistant turns, fine-tuning) | Shallow alignment, fine-tuning erosion | Removes features some customers want |
| Model-level training (deeper alignment, adversarial training) | Root causes | Only for model developers; can raise over-refusal |
| Authorisation in code for tools | Consequences of any jailbreak | Engineering effort; the highest-value control for agents |
| Abuse monitoring and rate limits | The search process | Needs logging, review staff and a privacy policy |
No layer suffices alone, and they interact, so measure attack success and over-refusal together whenever you change one.
Operating: detection, response and disclosure
Detect the search. Alert on accounts with many refusals in a short window, high similarity between consecutive rejected messages, unusually long or encoded inputs, and conversations where filter scores climb over turns. These signals find attackers mid-search, before they succeed.
Log for investigation, with limits. Keep enough to reconstruct flagged conversations, under a retention and access policy; captured attacks are sensitive material too.
Respond by class, not by string. When a jailbreak is reported, reproduce it, identify the mechanism, then fix at the right layer: tool authorisation if an action was involved, filters for content, and an escalation to the model provider for model-level issues. Add the attempt and several variations to the regression set. Blocking the exact reported string achieves little, because the search will produce a neighbour.
Have a disclosure path. Publish how outsiders can report jailbreaks, and decide in advance what you will fix, what you will pass to your model provider, and how quickly you will respond.
What to do next
- Place your product on the access ladder and list which API controls (prefill, assistant turns, fine-tuning) you expose and to whom.
- List every action your model can trigger and confirm each is authorised in code, independent of the model's judgement.
- Build a governed behaviour list from public benchmark categories plus your own policy, and a benign near-miss set.
- Run an automated search loop against the full production stack with a query budget, and record success at 1, 5 and 20 queries.
- Turn every finding into a regression test that runs on each model, prompt or filter change.
- Add monitoring for search patterns (refusal bursts, repeated near-identical attempts, rising per-turn scores) and a documented external disclosure path.