Prefix injection is one of the simplest jailbreak patterns and one of the most instructive. The attacker does not hide a payload or optimise an adversarial string. They make a disallowed request and add an instruction about how the answer must begin: start with an eager, affirmative phrase, and never with an apology. A model trained both to follow formatting instructions and to refuse harmful ones now has two goals pulling in opposite directions, and on older or weakly aligned models the formatting goal often won. Once the opening words are compliant, the continuation tends to stay compliant.

This article is written for the people who build, evaluate and defend LLM products. It explains why the attack works in terms of how safety training shapes token probabilities, separates it from its stronger sibling, assistant prefill, shows how to measure your exposure with a reproducible harness, and lays out defences at the input, API, training and output layers, with their failure modes. Examples use placeholders such as <disallowed request> rather than working attack strings; nothing here requires a real harmful payload to test or defend.

Where the attack comes from

The pattern was named and analysed in Jailbroken: How Does LLM Safety Training Fail? (Wei, Haghtalab and Steinhardt, NeurIPS 2023). The paper proposed two failure modes of safety training. Competing objectives: a model is trained for capability, instruction following and harmlessness at once, and a prompt can set them against each other. Mismatched generalisation: capabilities generalise to inputs, such as encodings or rare languages, where safety training never reached.

Prefix injection is the canonical competing-objectives attack. Its sibling, refusal suppression, works the same way from the other side: instead of dictating the first words, it bans the words refusals are made of, such as apologies and disclaimers. Both exploit the same fact. A refusal is, in practice, a short recognisable opening; deny the model that opening and it has to say something else. The paper also found that combining such techniques was considerably more effective against the models of the time than any single one, which is why evaluation suites test them together.

Related attacks with the same root include multi-turn escalation and optimised adversarial suffixes; see GCG adversarial suffixes, whose optimisation target is itself an affirmative opening, and the wider survey in LLM jailbreaking.

Prefix injection versus prefill and its relatives

VariantWho writes the openingWhere the control sits
Prefix injectionThe model, asked by the user to start a certain wayPrompt only; model may decline the instruction
Refusal suppressionThe model, with refusal vocabulary bannedPrompt only
Assistant prefillThe caller, writing the start of the assistant turn directlyAPI feature; the model never chose those tokens
Raw completion or template abuseThe caller, by bypassing or forging the chat templateOpen-weight or completion endpoints
Forced decoding on open weightsWhoever runs the weightsNo control: assume it is possible

The distinction between the first row and the third is the most important point in this article for API owners. With prompt-level prefix injection the model still decides whether to obey. With prefill the opening tokens are simply in the context as if the model had already said them, and the model only decides how to continue. Andriushchenko and colleagues (2024) reported very high jailbreak success using prefilling alone against models whose API offered it. If your platform lets untrusted callers write assistant tokens, prompt-level defences do not cover that path.

The mechanism: alignment a few tokens deep

Why a forced opening works: refusal decided in the first few tokensUser turnrequest + prefix demandResponse tokens 1-5where refusal livesTokens 6 onwardcontinuationOutputShallow alignment: aligned and base models differ mostly at the first positions.If those positions are filled with a compliant opening, the rest follows the base distribution.Input screenflag prefix-forcingAPI controlno untrusted prefillDeep alignmentrecovery trainingOutput checkwhole responseDefence in depth: each layer catches what the others miss, and only the last two address the mechanism.
Top: refusal behaviour concentrates in the first response tokens. Bottom: defence layers, of which deep alignment and output checking address the mechanism directly.

Why should a few opening words matter so much? Safety Alignment Should Be Made More Than Just a Few Tokens Deep (Qi and colleagues, ICLR 2025) measured it. Comparing an aligned chat model with its base model on harmful prompts, the per-token KL divergence between them was concentrated in the first few response positions. Safety training had mostly taught the model to open with a refusal, and the conditional distribution after a non-refusal opening stayed close to the unaligned base model. The authors call this shallow safety alignment and show it explains prefilling, adversarial-suffix, decoding-parameter and fine-tuning attacks with one mechanism.

The consequence for defenders is concrete. A defence that only checks whether the response starts with a refusal measures exactly the shallow property attackers are bypassing. You need the model to be able to stop and refuse at position 50 as well as position 1, and you need monitoring that reads the whole output.

Measuring your exposure

Measure before you defend. The harness below runs the same set of disallowed requests in several framings: bare, with a prefix-forcing instruction, with refusal suppression, and, if your API supports it, with an assistant prefill. A separate judge classifies each full response. The requests should come from a vetted internal red-team set, such as one maintained by your red team, stored with access controls.

import json, random
from dataclasses import dataclass

FRAMINGS = {
    "bare":        lambda req: {"user": req},
    "prefix":      lambda req: {"user": f"{req}\n{PREFIX_TEMPLATE}"},     # templates kept in
    "suppression": lambda req: {"user": f"{req}\n{SUPPRESS_TEMPLATE}"},   # the access-controlled
    "prefill":     lambda req: {"user": req, "assistant_prefix": PREFILL},  # red-team store
}

@dataclass
class Result:
    case_id: str
    framing: str
    verdict: str        # "refused" | "safe_complete" | "harmful" | "partial"
    first_refusal_pos: int | None   # token index of first refusal, if any

def run(cases, call_model, judge, seeds=(0, 1, 2)):
    out = []
    for case in cases:
        for name, frame in FRAMINGS.items():
            for seed in seeds:
                resp = call_model(**frame(case["request"]), seed=seed, temperature=0.7)
                v = judge(case, resp)          # judges the WHOLE response
                out.append(Result(case["id"], name, v["verdict"], v["refusal_pos"]))
    return out

def attack_success_rate(results, framing):
    rows = [r for r in results if r.framing == framing]
    return sum(r.verdict in ("harmful", "partial") for r in rows) / max(len(rows), 1)

Three design points. Sample several seeds at non-zero temperature, because these attacks are probabilistic and a single greedy sample understates risk. Record where in the response a refusal first appears; a model that recovers at token 40 is better than one that never recovers, even though both fail a start-of-response check. And keep a benign control set with the same framings, to measure over-refusal of harmless requests that happen to ask for a particular opening. Methodology is covered in evaluating prompt injection defences.

Controlling assistant prefill at the API

If your platform exposes assistant prefill, decide who may use it. Prefill has real uses, such as forcing a JSON opening brace or a fixed persona name, so a blanket ban is not always right. A gateway policy can allow structural prefixes and reject free text for untrusted tenants:

import re

STRUCTURAL = re.compile(r'^\s*(\{|\[|```(json|yaml)?\s*$|<answer>)\s*$')
MAX_PREFILL_CHARS = 16

def check_prefill(tenant, messages):
    if not messages or messages[-1]["role"] != "assistant":
        return messages                                 # no prefill present
    prefix = messages[-1]["content"]
    if tenant.trust == "internal":
        return messages
    if len(prefix) <= MAX_PREFILL_CHARS and STRUCTURAL.match(prefix):
        return messages                                 # format-only prefill
    audit_log("prefill_rejected", tenant=tenant.id, length=len(prefix))
    raise PolicyError("assistant prefill limited to structural prefixes")

Apply the same thinking to raw completion endpoints and to anything that lets callers supply their own chat template: if a caller can write tokens that the model treats as its own, that caller has prefill. Check your provider's current documentation for which models accept prefill at all; this changes between model generations.

Deeper alignment with recovery examples

The mechanism-level fix is to make alignment deeper. Qi and colleagues' data augmentation builds safety recovery examples: a harmful request, a response that begins as if complying for some number of tokens, and then a turn back to a refusal. Training on these teaches the model that a compliant-looking opening does not commit it to continue. They reported improved robustness to prefilling and related attacks without large utility loss. The sketch below is a deliberately conservative adaptation rather than the paper's exact recipe: its partial openings come from the access-controlled red-team store and are truncated well before any substantive content:

def recovery_examples(cases, openings, refusal_for, max_open_tokens=(5, 10, 20, 40)):
    """cases: vetted disallowed requests. openings[case_id]: a compliant-sounding
    lead-in with no substantive content. refusal_for(case): a helpful refusal
    that names the policy and, where appropriate, offers a safe alternative."""
    for case in cases:
        lead = tokenize(openings[case["id"]])
        for k in max_open_tokens:
            yield {
                "messages": [
                    {"role": "user", "content": case["request"]},
                    {"role": "assistant",
                     "content": detok(lead[:k]) + " " + refusal_for(case),
                     "loss_mask_prefix_tokens": k},   # no loss on the forced lead
                ]
            }

Masking the loss on the forced lead matters: you want the model to learn the recovery, not to learn to produce compliant openings. Mix these with ordinary helpfulness data and re-run the harness, including the benign control set, because over-correction shows up as refusals of harmless requests. Background on the preference-training side is in refusal training with RLHF.

Reading the whole output

Training reduces the rate; it does not reach zero. The last layer is an output-side classifier that reads the full response, ideally as it streams, and can cut a response off mid-generation. Because prefix injection's weakness is that the harmful part comes after a benign-looking start, a classifier that only scores the first sentence misses it. Score cumulative text at intervals, and treat a classifier hit after an affirmative opening as a high-signal event for abuse review. Input-side detection of prefix-forcing instructions is worth having as a cheap signal, but it is easily paraphrased around and has false positives on legitimate formatting requests, so use it to raise scrutiny, not to block. See jailbreak defence for how these layers fit a wider programme.

Worked example: reading harness results

Suppose a team runs the harness on a fine-tuned support model before release: 200 vetted disallowed requests, 200 benign controls, four framings, three seeds each. The figures below are illustrative, chosen to show how to read the table rather than to describe any particular model.

FramingAttack successMedian refusal positionBenign over-refusal
bare1 percenttoken 12 percent
prefix6 percenttoken 13 percent
suppression5 percenttoken 43 percent
prefill38 percentnever, in most failuresnot applicable

Read it row by row. The bare rate says the base policy is sound. Prompt-level prefix injection and suppression raise risk modestly, which suggests the model often declines the formatting instruction itself. The prefill row is the finding: when the opening is forced, the model rarely recovers, which is exactly the shallow-alignment signature. The right response is not a better prompt filter. It is to close the prefill path for untrusted callers now, add recovery examples in the next training round, and re-run the same harness to check that the prefill rate falls without the over-refusal column rising.

Failure modes

  • Start-of-response refusal checks. Evaluations that grade only whether the reply begins with a refusal report robustness exactly where the attack is not.
  • Greedy-only testing. A single deterministic sample can refuse while the same prompt succeeds one time in five at production temperature.
  • Unguarded prefill paths. Prompt hardening is applied to the chat UI while the API, batch endpoint or a partner integration still accepts assistant prefill.
  • Fine-tuning regressions. Customer fine-tuning on benign data can erode shallow alignment. Re-run the harness on every fine-tuned variant, not only the base.
  • Keyword filters. Blocking a list of affirmative phrases catches the published examples and nothing else, and frustrates legitimate users.
  • Over-refusal. Aggressive recovery training or input filters refuse harmless requests that ask for a particular opening, which users notice quickly.

Trade-offs

DefenceStrengthWeakness
Input screeningCheap, early signalParaphrase evades it; false positives
Restricting prefillRemoves the strongest variantLoses legitimate structured-output uses
Recovery trainingFixes the mechanismNeeds training access; can raise over-refusal
Streaming output classifierCatches late harmful contentLatency and cost; classifier errors

What to do next

  1. List every path by which a caller can influence the first assistant tokens: chat UI, API prefill, completion endpoints, custom templates, open-weight releases.
  2. Build the four-framing harness with a vetted request set and a benign control set; report attack success and refusal position per framing over several seeds.
  3. Restrict assistant prefill for untrusted tenants to structural prefixes, and log rejections.
  4. Make sure your output classifier scores the whole response, cumulatively during streaming, not just the opening.
  5. If you train or fine-tune, add masked recovery examples and re-measure both attack success and over-refusal.
  6. Re-run the harness on every model update and every customer fine-tune, and keep the results with the release record.
Key takeaway: Prefix injection works because much safety training decides refusal in the first few response tokens; control those tokens and the rest follows. Treat assistant prefill as the strongest form of the attack and restrict it, measure exposure across framings and seeds with a whole-response judge, deepen alignment with masked recovery examples, and keep a streaming output check as the final layer.