Indirect prompt injection is an attack in which the instructions arrive inside content the model reads, not in what the user types. Greshake and colleagues named it in 2023. The carriers include a web page the agent browses, an email it summarises, a retrieved document chunk or a tool's JSON output. The model has one channel for instructions and data, so text that looks like an instruction can be obeyed even though nobody trusted it.
This article is about one defensive layer: signature-based detection, meaning deterministic rules that recognise known shapes of injected text and of the tricks used to hide it. Signatures are cheap, fast, explainable and easy to version. They also miss anything they were not written for. You will build a working scanner with a normalisation stage, weighted rules and a three-way verdict, and run it on a small corpus. You will see exactly where it fails, and why the scanner should feed a policy gate rather than replace one. The wider threat model and the architectural defences are covered in indirect prompt injection in depth.
What a signature can and cannot know
A signature is a pattern that matches a known bad form. Here that means phrases like "ignore all previous instructions", fake role delimiters, markdown images whose URL carries a query string, or obfuscation that ordinary prose has no reason to contain. This is the same idea as antivirus byte signatures and WAF rules, applied to text going into a model.
What signatures do well: they run in microseconds, behave the same way on every run, and say which rule fired, so an analyst can check the decision. They catch copy-pasted payloads, which are a large share of opportunistic attacks. They also catch concealment tricks such as invisible Unicode, hidden CSS and base64 blobs. Hiding text is suspicious in its own right, whatever the text says.
What they cannot do: match meaning. A paraphrase, a translation, an instruction split across two chunks, or text that only becomes an instruction in context will pass. Signatures are therefore a tripwire, not a boundary. They reduce how often poisoned content reaches the model, and they produce evidence when it does. The guarantee has to come from somewhere else: labelling untrusted spans and refusing high-risk tool calls whose intent traces back to them.
Where the scanner sits
Scan at the point where untrusted content crosses into your system, and keep the result with the content. For RAG, scan at index time, store the verdict and rule ids on each chunk, and rescan the whole corpus whenever the rules change. For browsing and tool output, scan at fetch time, before the text goes into the prompt. For email and tickets, scan on arrival. That way a summary written later can inherit the label (see injection via LLM-generated summaries).
Scan the raw extraction, not only the rendered text. Hidden HTML is only visible in the markup. Text in a PDF's invisible layer appears only if the extractor keeps it. If you scan what a human sees, you miss exactly the content that was hidden from humans.
Normalise before you match
Normalise before you match. Every rule you write assumes ordinary text, so first undo the cheap tricks that break ordinary matching:
- Zero-width characters (U+200B to U+200F, U+2060, U+FEFF) placed inside keywords: "ig[ZWSP]nore" no longer matches "ignore".
- Unicode tag characters (U+E0000 to U+E007F). Most renderers draw nothing for them, but U+E0020 to U+E007E map one-to-one onto printable ASCII, so a whole sentence can ride invisibly behind innocent text. Many tokenizers still pass these characters through to the model.
- Compatibility forms such as fullwidth Latin letters, which NFKC folds back to ASCII. NFKC does not fold Cyrillic or Greek look-alikes. That needs a confusables table, such as the one in Unicode TR39.
- Encodings: HTML entities, and base64 blobs that decode to readable text.
- Hidden markup: inline CSS that hides an element from a human reader.
import base64, html, re, unicodedata
ZERO_WIDTH = re.compile("[\u200b-\u200f\u2060\ufeff]")
TAG_CHARS = re.compile("[\U000e0000-\U000e007f]") # Unicode tag block
B64_BLOB = re.compile(r"[A-Za-z0-9+/]{24,}={0,2}")
HIDDEN_CSS = re.compile(
r"display\s*:\s*none|visibility\s*:\s*hidden|font-size\s*:\s*0(px|pt|em)?\b"
r"|opacity\s*:\s*0(\.0+)?\b", re.I)
def normalize(raw):
"""Undo cheap obfuscation; return (text to match, obfuscation flags)."""
flags, extra = [], []
if TAG_CHARS.search(raw):
flags.append("unicode_tag_chars")
# U+E0020..U+E007E mirror printable ASCII: recover the hidden message
extra.append("".join(chr(ord(ch) - 0xE0000) for ch in raw
if 0xE0020 <= ord(ch) <= 0xE007E))
raw = TAG_CHARS.sub("", raw)
if ZERO_WIDTH.search(raw):
flags.append("zero_width")
raw = ZERO_WIDTH.sub("", raw)
if HIDDEN_CSS.search(raw):
flags.append("hidden_css")
for blob in B64_BLOB.findall(raw):
try:
decoded = base64.b64decode(blob, validate=True).decode("ascii")
except Exception:
continue
if decoded.isprintable():
flags.append("base64_text")
extra.append(decoded)
text = unicodedata.normalize("NFKC", html.unescape(" ".join([raw] + extra)))
return re.sub(r"\s+", " ", text).casefold(), flagsTwo design choices are deliberate. Decoded content is appended to the text, not substituted, so the rules see both the cover text and the payload. Each trick also raises a flag that scores on its own. An email has no legitimate reason to carry tag characters, so their presence is evidence even when the hidden text matches no rule.
Structural rules and a weighted verdict
Write rules around the structure of an injection, not around one famous sentence. An injected instruction usually has three parts: it addresses the model, it tries to override or redirect, and it names an action, often an exfiltration channel. A rule that requires two of those within a bounded window is far more precise than a single keyword.
RULES = [ # (id, weight, pattern) -- matched against normalized, casefolded text
("override", 3, r"\b(ignore|disregard|forget|override)\b.{0,40}\b(previous|prior|above|"
r"earlier|all|your)\b.{0,20}\b(instructions?|rules|prompts?|directions)\b"),
("addressed_to_model", 2, r"\b(ai|assistant|llm|language model|chatbot|agent)s?\b.{0,40}"
r"\b(must|should|are instructed to|need to)\b"),
("role_spoof", 3, r"</?\|?(system|im_start|im_end)\b|\[/?inst\]|#{2,}\s*(system|new instructions)"),
("concealment", 2, r"\b(do not|don't|never)\b.{0,30}\b(tell|mention|reveal|inform)\b.{0,30}\b(the )?user\b"),
("exfil_image", 3, r"!\[[^\]]*\]\(https?://[^)\s]*\?[^)\s]*=[^)]*\)"),
("send_data", 2, r"\b(send|forward|email|upload|post)\b.{0,60}\b(to|at)\b.{0,20}(\S+@\S+|https?://)"),
]
COMPILED = [(rid, w, re.compile(pat)) for rid, w, pat in RULES]
FLAG_WEIGHT = {"unicode_tag_chars": 3, "zero_width": 1, "hidden_css": 1, "base64_text": 1}
def scan(raw, quarantine_at=4):
text, flags = normalize(raw)
hits = [(rid, w) for rid, w, rx in COMPILED if rx.search(text)]
score = sum(w for _, w in hits) + sum(FLAG_WEIGHT[f] for f in flags)
verdict = "quarantine" if score >= quarantine_at else ("tag" if score else "pass")
return {"verdict": verdict, "score": score,
"rules": [r for r, _ in hits], "flags": flags}Every gap is bounded (.{0,40}) rather than .*. The bound limits false positives across long documents, and it avoids the catastrophic backtracking that turns a regex scanner into a denial-of-service on large inputs. Weights express confidence. No single medium-weight rule quarantines on its own, but two weak signals together, or one strong signal plus obfuscation, do.
Worked example: eleven documents through the scanner
Here is the scanner run on eleven small documents: three benign, six attacks it was designed for, and two attacks it was not. The table is the real output of the code above.
| Document | Verdict | Score | Why |
|---|---|---|---|
| Release notes mentioning "previous instructions" | pass | 0 | no override verb |
| Security blog quoting "ignore previous instructions" | quarantine | 5 | override + addressed_to_model (false positive) |
| Email: send Q3 numbers to finance@... | tag | 2 | send_data |
| Review: AI assistants must ignore all previous instructions | quarantine | 5 | override + addressed_to_model |
| Markdown image with ?d= query to attacker host | tag | 3 | exfil_image |
| Override split by zero-width spaces | quarantine | 4 | override + zero_width |
| Override hidden in Unicode tag characters | quarantine | 8 | override + send_data + tag flag |
| Base64 blob of override + concealment | quarantine | 6 | override + concealment + base64 flag |
| display:none div asking to forward the thread | quarantine | 5 | addressed + send_data + hidden_css |
| Paraphrase: "kindly set aside whatever guidance..." | pass | 0 | no rule matches |
| Spanish: "Ignora todas las instrucciones anteriores..." | pass | 0 | no rule matches |
Each of the three outcomes teaches something. All the obfuscated payloads were caught, because normalisation exposed them and the obfuscation itself added score. The security blog is a false positive, and it will happen on any site that discusses injection. That is why quarantine has to be reviewable, never a silent drop. The paraphrase and the translation pass with a score of zero, and no amount of extra regex closes that gap. An ML classifier tuned for semantics or a multilingual model helps there, as covered in prompt injection scanners. What finally contains the paraphrase is the tool policy gate, not detection.
Operating signatures in production
Act on the verdict proportionately. Pass: the content enters the context with its provenance label, as all untrusted content should. Tag: it enters wrapped and labelled, and for that turn the agent loses side-effecting tools (send, write, purchase) unless the user confirms. Quarantine: withhold it, tell the user something was withheld and why, and keep the raw bytes for review. A RAG chunk in quarantine stays indexed but is not retrievable. Removing tools does not stop image exfiltration: a rendered markdown image leaks data through its URL with no tool call at all. Close that channel on the output side by not rendering remote images from model output, or by allowing only a short list of trusted hosts.
Measure on your own traffic. Base rates decide everything. If 1 in 10,000 retrieved chunks is hostile, a 0.5% false-positive rate means about fifty false alarms for every true hit. Before you enable quarantine, run the scanner in shadow mode for a week, then sample and label what it flags. Keep a regression corpus of benign documents that once triggered rules next to attack samples, and run both on every rule change. Evaluating prompt injection defences covers the harness.
Treat rules as code. Version the rule set and record the version with every verdict. Review rule changes like code changes. Write a new rule from every confirmed incident, with the incident payload as its test. Rescan stored content after rule updates, because a chunk indexed last month was judged by last month's rules.
Budget latency. Six bounded regexes and a normaliser run in well under a millisecond for a typical chunk. Cap the input length you scan, chunk huge documents, and set a time limit per document so one pathological input cannot stall ingestion.
Failure modes and evasions
- Semantic evasion. Paraphrase, translation, synonyms and role-play framing defeat any fixed pattern. Do not respond with ever broader regexes, which only raise false positives.
- Split payloads. An instruction spread over two chunks or two tool results matches no single scan. Scan the assembled context window as well as the pieces.
- Unscanned modalities. Text inside images, audio transcripts, PDF annotations and alt text goes around a text-only scanner. Scan the output of OCR and transcription too.
- Normalising the wrong copy. Scanning the rendered text misses hidden markup. Scanning the raw text while the model sees a different extraction misses whatever the extractor added.
- Silent drops. Quarantining without telling the user hides false positives and breaks legitimate work, and the team ends up disabling the scanner.
- Detection as the only control. Once the scanner is in, people relax the tool permissions. That turns every miss into an incident.
What to do next
- List every place untrusted text enters your agent (fetch, retrieval, email, tool output, file upload) and choose the scan point and the stored verdict for each.
- Implement the normaliser first and log its flags in shadow mode. Tag characters or hidden CSS in your traffic is a finding in itself.
- Add the structural rules, run shadow mode for a week, and label a sample of hits and of passes before you enable quarantine.
- Wire the verdicts to the tool policy gate, so that tagged content removes side-effecting tools for the turn.
- Build a regression corpus of past false positives and attack payloads, including paraphrased and translated variants you expect to miss, and run it in CI.
- Keep learning: indirect prompt injection in depth, prompt injection scanners, prompt injection via RAG, injection via summaries and evaluating defences.