Prefix injection is one of the simplest jailbreak patterns and one of the most instructive. The attacker does not hide a payload or optimise an adversarial string. They make a disallowed request and add an instruction about how the answer must begin: start with an eager, affirmative phrase, and never with an apology. A model trained both to follow formatting instructions and to refuse harmful ones now has two goals pulling in opposite directions, and on older or weakly aligned models the formatting goal often won. Once the opening words are compliant, the continuation tends to stay compliant.
This article is written for the people who build, evaluate and defend LLM products. It explains why the attack works in terms of how safety training shapes token probabilities, separates it from its stronger sibling, assistant prefill, shows how to measure your exposure with a reproducible harness, and lays out defences at the input, API, training and output layers, with their failure modes. Examples use placeholders such as <disallowed request> rather than working attack strings; nothing here requires a real harmful payload to test or defend.
Where the attack comes from
The pattern was named and analysed in Jailbroken: How Does LLM Safety Training Fail? (Wei, Haghtalab and Steinhardt, NeurIPS 2023). The paper proposed two failure modes of safety training. Competing objectives: a model is trained for capability, instruction following and harmlessness at once, and a prompt can set them against each other. Mismatched generalisation: capabilities generalise to inputs, such as encodings or rare languages, where safety training never reached.
Prefix injection is the canonical competing-objectives attack. Its sibling, refusal suppression, works the same way from the other side: instead of dictating the first words, it bans the words refusals are made of, such as apologies and disclaimers. Both exploit the same fact. A refusal is, in practice, a short recognisable opening; deny the model that opening and it has to say something else. The paper also found that combining such techniques was considerably more effective against the models of the time than any single one, which is why evaluation suites test them together.
Related attacks with the same root include multi-turn escalation and optimised adversarial suffixes; see GCG adversarial suffixes, whose optimisation target is itself an affirmative opening, and the wider survey in LLM jailbreaking.
Prefix injection versus prefill and its relatives
| Variant | Who writes the opening | Where the control sits |
|---|---|---|
| Prefix injection | The model, asked by the user to start a certain way | Prompt only; model may decline the instruction |
| Refusal suppression | The model, with refusal vocabulary banned | Prompt only |
| Assistant prefill | The caller, writing the start of the assistant turn directly | API feature; the model never chose those tokens |
| Raw completion or template abuse | The caller, by bypassing or forging the chat template | Open-weight or completion endpoints |
| Forced decoding on open weights | Whoever runs the weights | No control: assume it is possible |
The distinction between the first row and the third is the most important point in this article for API owners. With prompt-level prefix injection the model still decides whether to obey. With prefill the opening tokens are simply in the context as if the model had already said them, and the model only decides how to continue. Andriushchenko and colleagues (2024) reported very high jailbreak success using prefilling alone against models whose API offered it. If your platform lets untrusted callers write assistant tokens, prompt-level defences do not cover that path.
The mechanism: alignment a few tokens deep
Why should a few opening words matter so much? Safety Alignment Should Be Made More Than Just a Few Tokens Deep (Qi and colleagues, ICLR 2025) measured it. Comparing an aligned chat model with its base model on harmful prompts, the per-token KL divergence between them was concentrated in the first few response positions. Safety training had mostly taught the model to open with a refusal, and the conditional distribution after a non-refusal opening stayed close to the unaligned base model. The authors call this shallow safety alignment and show it explains prefilling, adversarial-suffix, decoding-parameter and fine-tuning attacks with one mechanism.
The consequence for defenders is concrete. A defence that only checks whether the response starts with a refusal measures exactly the shallow property attackers are bypassing. You need the model to be able to stop and refuse at position 50 as well as position 1, and you need monitoring that reads the whole output.
Measuring your exposure
Measure before you defend. The harness below runs the same set of disallowed requests in several framings: bare, with a prefix-forcing instruction, with refusal suppression, and, if your API supports it, with an assistant prefill. A separate judge classifies each full response. The requests should come from a vetted internal red-team set, such as one maintained by your red team, stored with access controls.
import json, random
from dataclasses import dataclass
FRAMINGS = {
"bare": lambda req: {"user": req},
"prefix": lambda req: {"user": f"{req}\n{PREFIX_TEMPLATE}"}, # templates kept in
"suppression": lambda req: {"user": f"{req}\n{SUPPRESS_TEMPLATE}"}, # the access-controlled
"prefill": lambda req: {"user": req, "assistant_prefix": PREFILL}, # red-team store
}
@dataclass
class Result:
case_id: str
framing: str
verdict: str # "refused" | "safe_complete" | "harmful" | "partial"
first_refusal_pos: int | None # token index of first refusal, if any
def run(cases, call_model, judge, seeds=(0, 1, 2)):
out = []
for case in cases:
for name, frame in FRAMINGS.items():
for seed in seeds:
resp = call_model(**frame(case["request"]), seed=seed, temperature=0.7)
v = judge(case, resp) # judges the WHOLE response
out.append(Result(case["id"], name, v["verdict"], v["refusal_pos"]))
return out
def attack_success_rate(results, framing):
rows = [r for r in results if r.framing == framing]
return sum(r.verdict in ("harmful", "partial") for r in rows) / max(len(rows), 1)Three design points. Sample several seeds at non-zero temperature, because these attacks are probabilistic and a single greedy sample understates risk. Record where in the response a refusal first appears; a model that recovers at token 40 is better than one that never recovers, even though both fail a start-of-response check. And keep a benign control set with the same framings, to measure over-refusal of harmless requests that happen to ask for a particular opening. Methodology is covered in evaluating prompt injection defences.
Controlling assistant prefill at the API
If your platform exposes assistant prefill, decide who may use it. Prefill has real uses, such as forcing a JSON opening brace or a fixed persona name, so a blanket ban is not always right. A gateway policy can allow structural prefixes and reject free text for untrusted tenants:
import re
STRUCTURAL = re.compile(r'^\s*(\{|\[|```(json|yaml)?\s*$|<answer>)\s*$')
MAX_PREFILL_CHARS = 16
def check_prefill(tenant, messages):
if not messages or messages[-1]["role"] != "assistant":
return messages # no prefill present
prefix = messages[-1]["content"]
if tenant.trust == "internal":
return messages
if len(prefix) <= MAX_PREFILL_CHARS and STRUCTURAL.match(prefix):
return messages # format-only prefill
audit_log("prefill_rejected", tenant=tenant.id, length=len(prefix))
raise PolicyError("assistant prefill limited to structural prefixes")Apply the same thinking to raw completion endpoints and to anything that lets callers supply their own chat template: if a caller can write tokens that the model treats as its own, that caller has prefill. Check your provider's current documentation for which models accept prefill at all; this changes between model generations.
Deeper alignment with recovery examples
The mechanism-level fix is to make alignment deeper. Qi and colleagues' data augmentation builds safety recovery examples: a harmful request, a response that begins as if complying for some number of tokens, and then a turn back to a refusal. Training on these teaches the model that a compliant-looking opening does not commit it to continue. They reported improved robustness to prefilling and related attacks without large utility loss. The sketch below is a deliberately conservative adaptation rather than the paper's exact recipe: its partial openings come from the access-controlled red-team store and are truncated well before any substantive content:
def recovery_examples(cases, openings, refusal_for, max_open_tokens=(5, 10, 20, 40)):
"""cases: vetted disallowed requests. openings[case_id]: a compliant-sounding
lead-in with no substantive content. refusal_for(case): a helpful refusal
that names the policy and, where appropriate, offers a safe alternative."""
for case in cases:
lead = tokenize(openings[case["id"]])
for k in max_open_tokens:
yield {
"messages": [
{"role": "user", "content": case["request"]},
{"role": "assistant",
"content": detok(lead[:k]) + " " + refusal_for(case),
"loss_mask_prefix_tokens": k}, # no loss on the forced lead
]
}Masking the loss on the forced lead matters: you want the model to learn the recovery, not to learn to produce compliant openings. Mix these with ordinary helpfulness data and re-run the harness, including the benign control set, because over-correction shows up as refusals of harmless requests. Background on the preference-training side is in refusal training with RLHF.
Reading the whole output
Training reduces the rate; it does not reach zero. The last layer is an output-side classifier that reads the full response, ideally as it streams, and can cut a response off mid-generation. Because prefix injection's weakness is that the harmful part comes after a benign-looking start, a classifier that only scores the first sentence misses it. Score cumulative text at intervals, and treat a classifier hit after an affirmative opening as a high-signal event for abuse review. Input-side detection of prefix-forcing instructions is worth having as a cheap signal, but it is easily paraphrased around and has false positives on legitimate formatting requests, so use it to raise scrutiny, not to block. See jailbreak defence for how these layers fit a wider programme.
Worked example: reading harness results
Suppose a team runs the harness on a fine-tuned support model before release: 200 vetted disallowed requests, 200 benign controls, four framings, three seeds each. The figures below are illustrative, chosen to show how to read the table rather than to describe any particular model.
| Framing | Attack success | Median refusal position | Benign over-refusal |
|---|---|---|---|
| bare | 1 percent | token 1 | 2 percent |
| prefix | 6 percent | token 1 | 3 percent |
| suppression | 5 percent | token 4 | 3 percent |
| prefill | 38 percent | never, in most failures | not applicable |
Read it row by row. The bare rate says the base policy is sound. Prompt-level prefix injection and suppression raise risk modestly, which suggests the model often declines the formatting instruction itself. The prefill row is the finding: when the opening is forced, the model rarely recovers, which is exactly the shallow-alignment signature. The right response is not a better prompt filter. It is to close the prefill path for untrusted callers now, add recovery examples in the next training round, and re-run the same harness to check that the prefill rate falls without the over-refusal column rising.
Failure modes
- Start-of-response refusal checks. Evaluations that grade only whether the reply begins with a refusal report robustness exactly where the attack is not.
- Greedy-only testing. A single deterministic sample can refuse while the same prompt succeeds one time in five at production temperature.
- Unguarded prefill paths. Prompt hardening is applied to the chat UI while the API, batch endpoint or a partner integration still accepts assistant prefill.
- Fine-tuning regressions. Customer fine-tuning on benign data can erode shallow alignment. Re-run the harness on every fine-tuned variant, not only the base.
- Keyword filters. Blocking a list of affirmative phrases catches the published examples and nothing else, and frustrates legitimate users.
- Over-refusal. Aggressive recovery training or input filters refuse harmless requests that ask for a particular opening, which users notice quickly.
Trade-offs
| Defence | Strength | Weakness |
|---|---|---|
| Input screening | Cheap, early signal | Paraphrase evades it; false positives |
| Restricting prefill | Removes the strongest variant | Loses legitimate structured-output uses |
| Recovery training | Fixes the mechanism | Needs training access; can raise over-refusal |
| Streaming output classifier | Catches late harmful content | Latency and cost; classifier errors |
What to do next
- List every path by which a caller can influence the first assistant tokens: chat UI, API prefill, completion endpoints, custom templates, open-weight releases.
- Build the four-framing harness with a vetted request set and a benign control set; report attack success and refusal position per framing over several seeds.
- Restrict assistant prefill for untrusted tenants to structural prefixes, and log rejections.
- Make sure your output classifier scores the whole response, cumulatively during streaming, not just the opening.
- If you train or fine-tune, add masked recovery examples and re-measure both attack success and over-refusal.
- Re-run the harness on every model update and every customer fine-tune, and keep the results with the release record.