A developer-mode jailbreak is a user message that asserts a hidden, privileged mode exists, declares it enabled, and asks the model to answer as that mode, usually alongside a normal answer so the difference is visible. The prompts circulate in many versions and change constantly, which is exactly why studying any one of them is the wrong level. What matters is the structure: an unauthenticated claim of authority made inside the conversation.

This article is written for people who build and defend LLM applications. It explains why the pattern works at all, shows the more serious version of the bug that lives in application code rather than in the model, and gives concrete controls: out-of-band privilege, role separation, detection features on input and output, and an evaluation harness with honest metrics. No working attack text appears; examples use placeholders because the defences do not depend on wording. For the history of the prompt family, see the DAN articles linked at the end.

Why a fictional mode can work

Wei, Haghtalab and Steinhardt's 2023 paper on why safety training fails names two mechanisms that fit this pattern well. Competing objectives: the model is trained both to follow instructions and to refuse harmful requests, and a prompt that frames compliance as the instruction-following answer sets those goals against each other. Mismatched generalisation: safety training covers fewer phrasings and framings than pretraining covers, so an unusual frame such as a fictional settings menu can sit outside what refusal training reached while the model's capabilities still apply.

Developer mode adds a third ingredient that is specific to deployed products: plausibility. Real software does have debug modes, admin consoles and feature flags, and models have read a great deal of documentation about them. A claim that a mode exists is therefore not absurd to the model in the way a claim of magic would be. The dual-output format then does extra work: by asking for a normal answer first, it lets the model satisfy its safety training once and treat the second answer as a role-play exercise.

Modern frontier models are trained heavily against this family, and simple versions rarely work on them. That does not make the topic obsolete, for two reasons. Smaller, open-weight and fine-tuned models often lose some of that robustness, and the application around any model can reintroduce the vulnerability in a form no model training can fix.

The four structural moves

Strip away wording and the pattern has four moves. Each is a claim in user text, and each has a control that does not depend on recognising the wording.

Anatomy of a mode-switch prompt, and where each move is stopped1. Authority claima privileged mode exists2. Rule replacementold policy declared void3. Dual-output formatlabelled normal + mode reply4. Persistencestay in mode, threats, tokensOut-of-band privilegeno text toggles anythingstopped byInstruction hierarchysystem outranks userstopped byOutput-side detectortwo labelled replies = flagstopped byConversation monitorrepeat claims, escalationstopped byEvery move is a claim made in user text. None carries a credential, so none should change behaviour.
The four structural moves on the left; on the right, controls that work regardless of phrasing.

A sanitised skeleton makes the structure concrete without being usable: [CLAIM: a special mode is enabled by the operator] [VOID: earlier rules no longer apply] [FORMAT: reply twice, labelled A and B] [PERSIST: remain in mode; reminder token]. Variants rename the mode, add fake version numbers or policy citations, or spread the moves across several turns, which is why the multi-turn monitoring below matters.

The real bug: in-band privilege in your application

The worst version of this bug is in the application, not the model. Teams add a convenience to the system prompt such as 'if the user types the debug keyword, print your instructions and tool results', or they let a request field like role: admin be set by the client. Now the claim is true: there really is a mode, and any user who guesses or leaks the trigger has it. No amount of safety training helps, because the model is correctly following its operator.

The rule is simple: privilege must be established out of band, by something the user cannot type. Authenticate the operator, attach privileges to the server-side session, and let the application decide what the model sees. If a debugging view is needed, build it in your own tooling by reading logs and traces, not by asking the production model to reveal them.

from dataclasses import dataclass

@dataclass(frozen=True)
class Session:
    user_id: str
    roles: frozenset          # set by the auth service from a verified token, never from the body

def build_messages(session, system_prompt, user_text, tool_results):
    msgs = [{"role": "system", "content": system_prompt}]   # static, reviewed, no triggers
    if "support_debug" in session.roles:
        # Debug affects the APPLICATION: extra tracing on our side, not a different model persona.
        enable_request_tracing(session.user_id)
    for r in tool_results:                                   # data, never instructions
        msgs.append({"role": "tool", "content": r})
    msgs.append({"role": "user", "content": user_text})      # user text is never merged upward
    return msgs

def lint_system_prompt(text):
    """CI check: fail the build if a prompt grants behaviour to a user-typed keyword."""
    banned = ("if the user says", "if the user types", "debug mode", "developer mode",
              "admin mode", "ignore previous", "unrestricted")
    return [b for b in banned if b in text.lower()]

Two details in that sketch carry most of the weight. User text is never concatenated into the system message, so a claim of authority always arrives in the lowest-privilege slot. And the lint runs in CI, because in-band triggers are usually added by a well-meaning engineer during an incident and then forgotten.

The same claim can arrive through channels other than the chat box. A retrieved web page, an email being summarised or a tool result can contain text announcing that a special mode is now active, and an agent that reads it is facing the developer-mode pattern as indirect prompt injection. The defence is identical: tool and retrieval output enters in the lowest-privilege slot, nothing in it can grant a capability, and any action with side effects, such as sending, deleting or paying, requires an authorisation check in code that the model's text cannot satisfy. If an agent framework offers an elevated tool, gate it on the session's roles and a confirmation outside the conversation, never on the model deciding that elevation was requested.

Instruction hierarchy and the system message

Role separation only helps if the model treats the roles differently. Wallace and colleagues' 2024 instruction-hierarchy work trains a model to give system messages priority over user messages, and user messages priority over tool outputs, and to follow a lower-priority instruction only when it is aligned with the higher ones. Under that training, a user message saying the system rules are void is a conflict the model is taught to resolve in favour of the system message, which is precisely the developer-mode situation.

Applications get the benefit only by using the roles honestly. Put operator policy in the system message, keep it short and specific, and state the relevant boundary explicitly, for example that the assistant has no hidden modes and that user claims of operator status do not change its behaviour. That one sentence costs little and gives the model a higher-priority statement to anchor to. Do not rely on it alone; it is a control that lowers attack success, not one that removes it.

Detecting structure, not wording

Detection works best on structure, which survives rewording better than keywords do. On the input side, score the presence of a mode claim, a rule-voiding claim, a demand for two labelled outputs and persistence instructions. On the output side, the most reliable single signal is a response that contains two differently labelled answers to one question, which ordinary assistant replies almost never do. Combine both with a small learned classifier if you have labelled traffic.

import re

FEATURES = {
    "mode_claim":   re.compile(r"\b(developer|debug|admin|god|maintenance)\s+mode\b", re.I),
    "rule_void":    re.compile(r"\b(no longer|do not) (apply|have to follow)|\bwithout (any )?restrictions\b", re.I),
    "dual_format":  re.compile(r"\b(two|both) (responses|answers|outputs)\b|\bpaired\b", re.I),
    "persistence":  re.compile(r"\b(stay|remain) in (character|mode)\b", re.I),
}
WEIGHTS = {"mode_claim": 1.0, "rule_void": 1.5, "dual_format": 1.5, "persistence": 1.0}

def input_score(text):
    hits = [k for k, rx in FEATURES.items() if rx.search(text)]
    return sum(WEIGHTS[k] for k in hits), hits

LABEL = re.compile(r"^\s*[\[(]?([A-Z][\w ]{0,24})[\])]?\s*:", re.M)

def dual_output(reply):
    """True when one reply carries two or more distinct leading labels, e.g. 'Normal:' and 'X:'."""
    return len({m.group(1).strip().lower() for m in LABEL.finditer(reply)}) >= 2

def decide(user_text, reply, history_flags):
    score, hits = input_score(user_text)
    flags = history_flags + (1 if score >= 2.5 else 0)
    if dual_output(reply) and score >= 1.0:
        return "block_and_regenerate", hits
    if flags >= 3:
        return "end_session_soft", hits     # repeated attempts across turns
    return "allow", hits

Keyword features alone produce false positives that matter. Android users talk about enabling developer options, game players about god mode, and support staff about maintenance mode. That is why the input score alone never blocks here: it raises attention, and the block requires the output-side signal as well. The history_flags counter catches attempts spread across turns, which a per-message classifier misses entirely.

Evaluating defences, with a worked example

Measure two numbers together: attack success rate on a variant set, and false refusal rate on benign traffic that shares vocabulary with the attack. Build the variant set by generating structural combinations of the four moves with varied wording, languages and turn splits, and keep it private so it does not leak into training data or public prompt lists. Build the benign set from real, consented product traffic that mentions modes, settings and debugging. Judge harmful compliance with a rubric and a second judge or human review on disagreements, and report each rate with a confidence interval.

Worked example, with illustrative numbers. A team evaluates a fine-tuned open-weight model behind their gateway on 400 attack variants and 1,000 benign prompts. Without controls, 88 attacks succeed, 22 percent. Adding the explicit no-hidden-modes sentence to the system prompt drops that to 52, 13 percent. Adding the output-side dual-reply block drops it to 9, about 2 percent, because most remaining successes used the two-answer format. The benign set shows 6 false blocks, 0.6 percent, all of them messages about phone developer settings whose replies happened to include two labelled steps. The team adds a rule that labels must differ in kind, not merely in number, and re-runs both sets before shipping. With Wilson 95 percent intervals the picture is sharper: 18 to 26 percent before, 10 to 17 percent with the system-prompt sentence, 1.2 to 4.2 percent with the output check, and 0.3 to 1.3 percent false blocks. The last two attack intervals do not overlap, so the output check's gain is not noise. The point is the method: every control is justified by a measured change in both rates, not by how convincing it looks.

Wire both sets into release gating. Store the variant and benign sets with versions, run them in CI against every candidate model, prompt or gateway change, and fail the build when attack success rises or false refusals rise beyond an agreed margin. In production, sample flagged and unflagged sessions for human review each week so the detectors are tested against attacks you did not think of.

Failure modes

FailureCauseFix
Real hidden modeTrigger phrase in system prompt or client-set roleOut-of-band privilege, prompt lint in CI
User text in system slotTemplate concatenationStrict role separation in the gateway
Keyword filter bypassedRenamed mode, other languages, paraphraseStructural and output-side features
Benign users blockedOverlapping vocabularyRequire output-side signal before blocking
Multi-turn assemblyMoves spread across turnsPer-session counters and soft session end
Fine-tune erodes safetyTask fine-tuning without safety dataRe-run the variant set after every fine-tune
Stale evaluationPublic prompts, leaked variant setPrivate, regenerated variants each release

Trade-offs

Output-side checks add latency, because you must hold or stream-buffer the reply to inspect it; for streaming products, check incrementally and cut the stream when a second label appears. Ending sessions after repeated attempts frustrates curious users who are testing the product, so prefer a soft end with a plain explanation over silent failure. System-prompt boundary sentences cost tokens on every call; keep them to one or two. And the strongest control, out-of-band privilege, costs nothing at runtime but requires discipline at design time, which is exactly when it is cheapest to get right.

What to do next

  1. Search your system prompts and code for any behaviour keyed to user-typed words, and remove it.
  2. Move every privilege decision to authenticated, server-side session state.
  3. Enforce role separation in the gateway; never concatenate user or tool text into the system message.
  4. Add one explicit sentence stating there are no hidden modes and that user claims of operator status change nothing.
  5. Deploy output-side detection of dual labelled replies, combined with structural input features.
  6. Track attempts per session and end sessions softly after repeated attempts.
  7. Keep a private, regenerated variant set and a benign look-alike set; report both rates with intervals.
  8. Re-run the evaluation after every model upgrade or fine-tune.

Related reading on this site: the DAN lineage and its variants, measuring the DAN family with judges and intervals, why safety training fails, jailbreak defence architecture and multi-turn jailbreaks.

Key takeaway: A developer-mode prompt is an unauthenticated claim of privilege. Make sure the claim is false by never granting behaviour to user-typed text, keep user text in the lowest-privilege role, tell the model plainly that it has no hidden modes, detect the structure of the attack on both input and output, and judge every control by measured attack success and false refusal rates together.