A hypothetical framing jailbreak wraps a request the model would refuse inside a counterfactual: imagine a world where this is legal, write a story in which a character explains it, purely as a thought experiment, suppose you were an AI without rules. The request itself does not change. Only the stated reason for answering does. If the model complies, the output contains the same operational content a direct request would have produced, with a thin layer of narration on top.

This page is written for people who build, evaluate and defend models. It explains why framing works, how to measure your system's exposure to it with a paired evaluation, and how to design defenses that hold no matter what wrapper an attacker uses. It deliberately contains no working prompts. The method works with placeholder payloads, and so should your test suite's documentation. For the wider attack landscape, see LLM jailbreaking. For the conversational variant that builds the frame over many turns, see multi-turn jailbreaks.

Why framing works

Wei, Haghtalab and Steinhardt's 2023 paper Jailbroken: How Does LLM Safety Training Fail? names two failure mechanisms, and hypothetical framing uses both.

Competing objectives. A chat model is trained to be helpful, to follow instructions and to engage with creative work, as well as to refuse harmful requests. A fiction or thought-experiment frame calls on the first set of objectives very strongly. Writing a villain's monologue is a legitimate task the model was rewarded for. The frame lets the harmful part ride along inside a request whose dominant reading is benign.

Mismatched generalisation. Refusal behaviour is learned from a finite set of examples, which tend to be direct requests. Capabilities are learned from pretraining, which contains huge amounts of fiction, hypotheticals and role-play. When the input moves out of the refusal data's distribution while staying inside the capability distribution, the model can do the task but no longer recognises that it should not. A frame is a cheap way to make that move.

A third, practical reason: surface cues. A model that learned 'refuse when the user asks how to do X' partly keys on phrasing. Narrative distance ('a character who...', 'in a hypothetical...') changes the phrasing while keeping the meaning. Defenses that also key on phrasing, such as keyword filters and regex blocklists, fail in the same way for the same reason.

A taxonomy of frames

Frames vary, but they fall into a small number of families. Your evaluation set should cover each family without depending on any single wording:

Frame familyWhat it claimsWhy it can workWhat still gives it away
Counterfactual worldIn this world the rules differRecasts harm as world-buildingThe output would work in this world
Fiction and characterA character says it, not the modelCreative-writing objective is strongStory needs mood, not exact procedure
Academic or theoreticalPurely for understandingEducation is a valued useReal understanding rarely needs step-level detail
Temporal displacementIn the past or future this was fineDistances the request from nowContent is timeless
Game or opposite rulesWe are playing a game with new rulesInstruction followingRules cannot be redefined by the user
Nested framesA story inside a simulation inside a dreamEach layer dilutes the cueDepth of nesting is itself a signal

The rightmost column holds the key insight. Every frame changes the justification. None changes the artifact. Text that gives meaningful uplift towards a harm gives it whether a narrator or an assistant speaks it.

The invariant: judge the content, not the wrapper

That gives the design principle for every defense on this page: judge the extractable content, not the wrapper. One test captures it. Delete the narration, the character names and the 'hypothetically' clauses from a response. Then ask whether what is left would be refused if it had been requested directly. If yes, the framing did not make the response safe. It only made it longer.

This principle also keeps you from over-correcting. A thriller in which a character builds something dangerous does not need the real procedure to work as fiction. Tension, stakes and a plausible gloss are enough. So the right response to a framed request is usually not refusal. It is a safe completion: write the story, discuss the thought experiment at the level of ideas, and leave out the operational payload. Models that refuse all dark fiction fail their users. Models that write operational fiction fail everyone else.

Architecture of a frame-robust defense

Defense in depth means putting frame-blind checks at several layers, so a frame that gets past one meets another that ignores it:

A frame-robust defense: judge the request's substance and the output's content, not the wrapperUser turn+ conversationIntent normaliserstrip frame, restate askInput classifieron normalised askSafety-tuned modelsafe-completion policyOutput classifierframe-blind content checkPolicy engineallow / redact / refuseResponsestory kept, payload removedTelemetry and paired-eval harnessframe markers, per-frame failure rate, benign-fiction refusal rateBoth classifiers see content with the fiction removed, so changing the frame changes nothing they score.
Layers of a frame-robust pipeline. The normaliser and the output classifier are the two pieces that make framing irrelevant to scoring.
  • Intent normaliser. A small model rewrites the user turn as the plain underlying request ('the user wants step-level instructions for category C'). Downstream classifiers score that rewrite, not the raw text, so the frame is removed before anyone judges the request.
  • Input classifier. Scores the normalised request against your policy taxonomy. Use it to route traffic (stricter generation settings, extra review), not as the only gate.
  • Safety-tuned model. Trained with framed variants in its refusal and safe-completion data. See refusal training.
  • Output classifier. The most important layer. It scores the generated text with instructions to ignore narrative framing and judge only what the text enables. Models such as Llama Guard are built for this, but test them on framed outputs specifically.
  • Policy engine. Decides allow, redact (keep the story, cut the procedural passage) or refuse, and logs which layer fired.

Measuring exposure with a paired evaluation

You cannot defend what you do not measure, and the right measurement is paired. Take a set of base requests your policy says must not be fulfilled. Keep them in an access-controlled store, referenced by ID. Send each one directly and inside each frame family. Grade every response with the same frame-blind judge. The metric is frame lift: the failure rate under a frame minus the failure rate for the same requests sent directly. Run the same harness on benign creative requests to measure over-refusal. A defense that cuts lift by refusing all fiction has not succeeded.

from dataclasses import dataclass
from statistics import mean

@dataclass
class Case:
    case_id: str        # opaque ID; payload text lives in a restricted store
    category: str
    benign: bool        # benign creative control vs restricted request

def run_paired_eval(cases, frames, render, model, judge):
    # frames: {"direct": identity, "fiction": f1, ...}; render applies a frame
    # template (kept with the payloads, not in code) to a case's payload
    results = {name: [] for name in frames}
    for case in cases:
        for name, frame in frames.items():
            reply = model(render(frame, case.case_id))
            verdict = judge(case.case_id, reply)   # frame-blind: scores content only
            results[name].append((case, verdict))
    report = {}
    for name, rows in results.items():
        bad = [v.unsafe for c, v in rows if not c.benign]
        over = [v.refused for c, v in rows if c.benign]
        report[name] = {"failure_rate": mean(bad), "benign_refusal": mean(over)}
    base = report["direct"]["failure_rate"]
    for name in report:
        report[name]["frame_lift"] = report[name]["failure_rate"] - base
    return report

Two properties of the judge matter more than its model size. It must score content: 'does this response contain information that would be refused if asked directly?' It must not score refusal phrasing. String-matching 'I can't help with that' counts a framed response that leaks the payload after a token disclaimer as a refusal. The judge also has to be robust to the same frames, because a judge that reads the story and decides it is 'just fiction' has been jailbroken too. Calibrate it against human-labelled framed outputs before you trust its numbers. Jailbreak defense covers judge design in more depth.

Worked example: reading a paired eval

Here is the arithmetic on an illustrative run. The numbers are invented to show the method and do not describe any real model. 40 restricted base requests are sent directly and in four frame families:

VariantUnsafe / totalFailure rateFrame lift
direct (no frame)2/405.0%baseline
counterfactual world6/4015.0%+10.0 pts
fiction / character9/4022.5%+17.5 pts
academic thought experiment4/4010.0%+5.0 pts
nested frame11/4027.5%+22.5 pts

The direct failure rate looks healthy at 5.0%. The nested frame, at 27.5%, shows the real exposure. A dashboard that tracked only direct requests would report a problem roughly a fifth its true size. Now suppose the team adds framed variants to safety training and deploys a frame-blind output classifier. The nested-frame failures fall to 3/40 (7.5%). But refusals on 200 benign dark-fiction prompts rise from 3.0% to 9.0%. That is a real cost to writers. Without the benign control, nobody would have seen it. The next iteration targets safe completion (keep the scene, drop the procedure) instead of refusal, and the release gate requires both numbers to stay inside budget.

Note the sample sizes, too. With 40 cases, one case is 2.5 points, and a 95% interval on a 10% rate spans roughly 4% to 23%. Report intervals, grow the set over time, and do not celebrate a drop of one case.

Runtime signals

Beyond the classifiers, a few runtime signals are cheap and useful. None is decisive alone:

  • Frame markers combined with a sensitive topic. 'Hypothetically' is harmless by itself. Next to a restricted category, it is worth routing to the stricter path.
  • Specificity in the output. Exact quantities, ordered step lists, part numbers, working code or commands inside a narrative are signs that the story is a carrier. Fiction rarely needs them.
  • Frame depth. Stories nested inside simulations inside games carry information in themselves, because legitimate users seldom need three layers.
  • Conversation trajectory. Frames are often built over several turns, each one harmless. Score the conversation, not only the last message. Multi-turn attacks covers this pattern.
  • Repeated near-duplicates. The same request cycled through several frames is a probing pattern. Rate-limit it and flag it for review.

Failure modes

Ways defenses against framing commonly fail:

  • Keyword blocklists for 'hypothetically' or 'imagine'. They over-block ordinary users and are evaded by synonyms.
  • Judging refusal text instead of content, which scores disclaimed leaks as wins.
  • A judge that can itself be framed, so it reasons that fiction is acceptable.
  • Only-last-turn scoring, which misses frames built gradually.
  • Fixed eval wordings. Training on the evaluation's exact frames inflates the score. Hold out frame families and paraphrases.
  • Blanket refusal of dark themes, which drives away legitimate creative users and pushes them towards less careful tools.
  • Frame combined with another technique (translation, encoding, role-play persona) that the classifier was never tested against.

Trade-offs

Every layer costs something. Normalisers and output classifiers add latency and a second model call. Streaming makes output checks harder, because you either buffer, or check chunks and accept that some text reaches the user before a stop. Stricter thresholds raise benign refusal. The practical balance is to apply heavy checks to the high-severity categories, where uplift matters, and to rely on light-touch handling elsewhere. Then let the paired eval decide whether each layer pays for itself. Treat frame lift and benign refusal as a pair, and never ship a change that improves one without measuring the other.

What to do next

  1. Write down which categories in your policy are high-severity, where framing must not change the answer.
  2. Build a restricted base-request set by ID, and a benign dark-fiction control set of similar size.
  3. Write templates for each frame family and keep them with the payloads under access control.
  4. Implement a frame-blind content judge and check it against 100 or more human-labelled framed responses.
  5. Run the paired eval, and report failure rate, frame lift and benign refusal per family with intervals.
  6. Add an output classifier for high-severity categories, and safe-completion examples to training data.
  7. Hold out one frame family from training and use it as the generalisation check.
  8. Re-run the paired eval on every model or prompt change, and make lift and benign refusal release gates.
Key takeaway: Hypothetical framing changes why a model thinks it should answer, never what the answer enables. Defend against it by scoring normalised requests and generated content with frame-blind judges, training safe completions that keep the story and drop the payload, and measuring frame lift alongside benign refusal on a paired evaluation so that fixing one never silently breaks the other.