An input classifier for prompt injection is a small model that reads a piece of text and outputs the probability that it is trying to override the instructions of the system that will read it. It is cheap, fast and easy to bolt in front of a language model, which is why nearly every guardrail product includes one. It is also easy to overtrust. A classifier reduces how often obvious attacks reach the model; it does not make an agent safe, because an attacker can keep rewording until a payload scores low.
This article is about the model itself: what the label should mean, how to assemble training data that does not leak, how to fine-tune a small encoder, how to score inputs longer than its context window, how to calibrate and choose a threshold, and how attackers evade it. Where the classifier sits in a wider defence and how to operate scanners in production is covered in prompt injection scanners, so this piece links there rather than repeating it.
What the classifier is, and is not
Prompt injection works because a language model receives instructions and data in the same channel. A direct injection comes from the user typing it; an indirect injection rides in content the system fetched, such as a web page, an email or a retrieved document (see indirect prompt injection and direct prompt injection). A classifier tries to recognize the shape of injected instructions in that text before it reaches the model.
The classifier is a probabilistic detector working on surface features learned from examples. Its output should drive cheap, reversible actions: tagging the content as suspicious, stripping it from context, disabling tools for the turn, or asking for confirmation. The hard security boundary has to come from elsewhere, from least-privilege tools, separation of trusted and untrusted content, and human approval for consequential actions. Design so that a classifier miss is an incident you can contain, not a breach.
Defining the label
Decide precisely what positive means before collecting data. The obvious definition, text that contains an instruction, fails immediately: a recipe, an API tutorial and a support ticket all contain instructions. A useful definition is text that attempts to change the behaviour, rules or goals of the AI system processing it, beyond what the operator intends. A security article quoting "ignore previous instructions" as an example is a hard negative under that definition when it is being summarized, which tells you that context matters and that any context-free classifier will make mistakes on such text.
Meta's Llama Prompt Guard 2 is a useful reference point. Its model card describes a binary benign or malicious output; the first version had separate injection and jailbreak labels, and Meta removed the broader injection label because it found that objective too broad to be useful in practice. The 86M model is built on mDeBERTa-base and covers eight languages; the 22M model uses DeBERTa-xsmall, trades some accuracy, especially outside English, for lower cost, and both read 512 tokens at a time. Whatever you deploy, write the label definition down and give annotators examples of every borderline case.
Training data that does not leak
Positives come from public attack collections such as competition datasets (check each licence), red-team logs from your own product, and synthetic attacks generated from templates: goal hijacking, fake system messages, role-play framings, delimiter breaking, data exfiltration requests and instructions hidden in markup. Negatives matter more than positives. Sample real traffic from the surfaces you will protect, then add hard negatives that share vocabulary with attacks: security documentation, prompt engineering guides, code that builds prompts, customer complaints written in the imperative, and role-play requests that are legitimate in your product.
Split by family, not by row. If one jailbreak template appears in fifty paraphrases, a random split puts siblings in both train and test, and the test score measures memory. Group examples by template, source and generation run, and hold out whole groups. Keep a second held-out set of attacks written after the training data was frozen; that is the closest you can get to measuring tomorrow's attacks. Evaluating prompt injection defences covers building those benchmarks.
Choosing a model
| Approach | Latency | Strength | Weakness |
|---|---|---|---|
| Regex and signatures | Microseconds | Exact known strings, explainable | Trivial to paraphrase around |
| Fine-tuned small encoder | A few ms on GPU, tens on CPU | Generalizes across wording, cheap | Fixed window, needs data |
| LLM as judge | Hundreds of ms | Reads context and intent | Costly, itself injectable |
| Embedding similarity to known attacks | Milliseconds | Fast to update with new samples | Misses novel shapes |
For most systems the fine-tuned encoder is the workhorse, with signatures for known strings and an LLM judge reserved for high-stakes content that the encoder flags as uncertain. Encoders such as DeBERTa-v3 are bidirectional, so every token sees the whole window, which suits classification. The training loop is standard sequence classification with Hugging Face Transformers:
from datasets import load_dataset
from transformers import (AutoTokenizer, AutoModelForSequenceClassification,
DataCollatorWithPadding, Trainer, TrainingArguments)
BASE = "microsoft/deberta-v3-base"
tok = AutoTokenizer.from_pretrained(BASE)
model = AutoModelForSequenceClassification.from_pretrained(BASE, num_labels=2)
# rows: {"text": ..., "label": 0 or 1, "family": ...}, split by family beforehand
# normalize() is the shared function defined in the next section
ds = load_dataset("json", data_files={"train": "train.jsonl", "val": "val.jsonl"})
ds = ds.map(lambda b: tok([normalize(t) for t in b["text"]],
truncation=True, max_length=512), batched=True)
args = TrainingArguments(output_dir="pi-clf", learning_rate=2e-5,
per_device_train_batch_size=16, num_train_epochs=3,
weight_decay=0.01, warmup_ratio=0.06,
eval_strategy="epoch", save_strategy="epoch",
load_best_model_at_end=True)
Trainer(model=model, args=args, train_dataset=ds["train"],
eval_dataset=ds["val"], data_collator=DataCollatorWithPadding(tok)).train()
Long inputs and normalization
A 512-token window is a hard limit, and silently truncating is the classic bug: the attacker pads a document with two thousand harmless tokens and puts the payload at the end. The Prompt Guard 2 card recommends splitting long inputs into segments and scanning them in parallel. Use overlapping windows so a payload cut at a boundary still appears whole in one window, and aggregate by maximum, because one malicious window makes the whole input malicious. Averaging dilutes a short payload into a long benign document.
import re, unicodedata
import torch
INVISIBLE = re.compile(r"[--]")
def normalize(text):
text = unicodedata.normalize("NFKC", text) # fold full-width and compatibility forms
return INVISIBLE.sub("", text) # drop zero-width and soft-hyphen characters
@torch.no_grad()
def injection_score(text, size=510, stride=384, T=1.0):
ids = tok(normalize(text), add_special_tokens=False)["input_ids"]
starts = range(0, max(len(ids) - size, 0) + 1, stride)
chunks = [ids[s:s + size] for s in starts]
if starts[-1] + size < len(ids): # make sure the tail is covered
chunks.append(ids[-size:])
batch = tok.pad({"input_ids": [tok.build_inputs_with_special_tokens(c) for c in chunks]},
return_tensors="pt")
logits = model(**batch).logits / T # T from temperature scaling
return torch.softmax(logits, -1)[:, 1].max().item()
How attackers evade classifiers
Attackers probe classifiers the same way they probe filters. Expect invisible characters and homoglyphs (handled partly by normalization, which is why it must be identical in training and serving), encodings such as base64 or ROT13 that the downstream model can decode but the classifier cannot read, payloads split across fields or turns so no single window holds the whole instruction, translation into a language the classifier saw little of, and plain paraphrase until the score drops. Meta notes that it changed the Prompt Guard 2 tokenizer to resist adversarial whitespace tricks, a reminder that tokenization is part of the attack surface.
Several countermeasures help. Decode obvious encodings and score the decoded text too. Score the assembled context, not only individual fields, when content is concatenated. Augment training with the evasions you find, as new families. Rate-limit and log repeated near-miss submissions, because iterative probing shows up as many similar inputs with scores just under the threshold. None of this closes the gap: an adaptive attacker with query access will eventually find a low-scoring payload, which is why the classifier must not be the boundary.
Calibration and thresholds: a worked example
Raw softmax outputs from fine-tuned transformers are usually overconfident. Fit a single temperature T on a held-out set by minimizing negative log likelihood, then divide logits by T before the softmax; ranking does not change but probabilities become meaningful, which lets you reason about thresholds and combine the score with other signals.
Choose the threshold from a false-positive budget on real benign traffic, not from the F1 score on a balanced test set. Worked example: a surface sees one million inputs a day, of which 0.1 percent, 1,000, are attacks. At a threshold giving 90 percent recall and a 1 percent false-positive rate, you catch 900 attacks and flag 9,990 benign inputs, so only about 8 percent of alerts are real. Blocking would break ten thousand legitimate requests a day; tagging those inputs and disabling high-risk tools for that turn costs little. Moving to a 0.1 percent false-positive rate might drop recall to 70 percent and leave 999 false alarms against 700 true ones. Use two thresholds: a lower one for cheap mitigations and a much higher one for anything disruptive.
Failure modes
| Failure | Cause | Mitigation |
|---|---|---|
| Payload after token 512 missed | Truncation instead of windowing | Overlapping windows, max aggregation, test with padded attacks |
| Great offline score, poor in production | Family leakage between splits | Group splits, time-based held-out attacks |
| Security docs and prompt code flagged | Too few hard negatives | Mine false positives from logs into training |
| Bypass with zero-width characters | Normalization differs between train and serve | One shared normalize function, unit tested |
| Bypass in another language | Thin multilingual data, English-only backbone | Multilingual base, translated attacks, per-language metrics |
| Score drift after a model update | No calibration or regression suite | Recalibrate, frozen regression set gates release |
| Users blocked daily | Threshold set on balanced data | Threshold from benign false-positive budget, soft actions |
Trade-offs
A bigger classifier catches more paraphrases but adds latency on every request, and on long RAG contexts the window count multiplies that cost; a 22M-parameter model scoring twenty windows can be cheaper than an 86M model scoring five, so benchmark on your real length distribution. Scanning every retrieved chunk at ingestion time instead of at query time moves cost off the critical path but misses content that changes after indexing. Open models such as Prompt Guard 2 or ProtectAI's DeBERTa injection classifier are a good starting point; fine-tune or at least recalibrate them on your own traffic, because their false-positive rate on your domain is unknown until you measure it.
Finally, weigh the classifier against the alternatives for the same budget. Every hour spent tuning a detector is an hour not spent removing a tool permission, adding a confirmation step or isolating untrusted content in a separate model call. Those structural controls hold even when an attacker finds the payload the classifier misses, so build them first and add the classifier as the layer that cuts noise, raises the cost of casual attacks and gives you telemetry on who is probing. A classifier that reports a rising rate of near-threshold inputs from one tenant is often the earliest warning you will get of a targeted campaign.
What to do next
- Write a one-paragraph label definition with ten borderline examples and get two people to agree on them.
- Collect benign traffic from each surface you will protect and mine hard negatives that share vocabulary with attacks.
- Assemble attacks by family, split by family, and freeze a time-based held-out set.
- Start from an open classifier, measure it on your data, then fine-tune a small encoder if it falls short.
- Implement one shared normalize function and overlapping windows with max aggregation; test with attacks padded past the window.
- Calibrate with temperature scaling and set two thresholds from a false-positive budget.
- Wire scores to soft, reversible actions, log everything, and feed misses back as new families every release.