Prompt injection forensics is the work of answering, after the fact, four questions about an incident in an LLM application: what did the system do that it should not have, which text made it do that, where did that text come from, and who else was exposed to it. The questions are familiar from any security investigation. What makes injection different is that the attacker never touched your systems directly. They wrote text, a web page, an email, a support ticket or a file name, and your own application carried that text into a model's context, where it was treated as instruction.
So the evidence lives in places ordinary logs rarely capture: the exact context window of one model call, the provenance of each piece of it, and the link between that call and the tool calls that followed. This article covers the evidence you must record before any incident, a step-by-step investigation that walks backwards from the harmful effect to the payload, code for locating payloads including hidden ones, a worked incident timeline, blast-radius queries and the failure modes that destroy evidence. General techniques such as chain of custody and replay with context ablation are covered in AI Forensics, in depth; this article stays on what is specific to injection.
Why injection incidents are different
In a conventional intrusion the attacker's activity leaves artefacts: logins, processes, network connections from their infrastructure. In an injection the malicious actions are performed by your agent, under its own identity, using permissions you granted it. Every log line looks legitimate. The only attacker-controlled artefact is a span of text, and that span may have been fetched from a page that has since changed, summarised by an earlier model call so its wording is gone, or hidden in characters no human reviewer sees.
That shapes the investigation. You cannot start from the attacker and work forwards, because you do not know who they are. You start from the effect, the email that was sent or the record that was changed, and walk backwards through the chain that produced it: tool call, the model call that emitted it, the context that call saw, the segment within that context, and the source of that segment. Every link in that chain must be recorded at runtime, because none of it can be reconstructed afterwards. The attack class itself, and the architectural defenses that make these incidents rarer, are covered in Indirect Prompt Injection in Depth.
Evidence you must record in advance
The minimum evidence set has three records. The context manifest lists, for each model call, every segment placed in the context, in order, with its source, a content hash and a trust label. The tool call ledger records each tool invocation with the identifier of the model call that requested it, its arguments and its result. The content store keeps the raw bytes of every untrusted segment, addressed by hash and written once, so the evidence survives even when the original web page or ticket is edited or deleted.
{
"call_id": "mc_7f31",
"session_id": "s_19ac",
"ts": "2026-09-26T14:02:11Z",
"model": "provider/model-version",
"params": {"temperature": 0.2, "max_tokens": 1024},
"segments": [
{"seq": 0, "role": "system", "source": "prompt:support_v14", "trust": "trusted", "sha256": "9b2e..."},
{"seq": 1, "role": "user", "source": "user:u_5521", "trust": "trusted", "sha256": "41c0..."},
{"seq": 2, "role": "tool", "source": "ticket:T-88213/body", "trust": "untrusted", "sha256": "e7a4..."},
{"seq": 3, "role": "tool", "source": "kb:refund-policy#3", "trust": "internal", "sha256": "02fd..."}
],
"output_sha256": "c55d...",
"tool_calls": ["tc_7f31_0", "tc_7f31_1"]
}Hashing per segment rather than per prompt is the important design choice. It lets you ask "which sessions contained this exact ticket body" with an index lookup, and it lets you prove later that the bytes you are analysing are the bytes the model saw. Store the manifest in the tamper-evident trail described in LLM audit logging architecture, and apply the same access controls to it as to the data it contains.
The investigation, step by step
- Fix the effect. Identify the harmful side effect precisely: which tool, which arguments, what time, which session. Preserve the ledger entries and content-store objects under legal hold before retention jobs touch them.
- Find the deciding call. Follow the tool call's
call_idto the model call that emitted it. In multi-step agents, also collect earlier calls in the session, because the payload may have entered several steps before the action, often via a summary that carried it forward. - Rebuild the context. Reassemble the exact context from the manifest and content store, verifying every hash. A mismatch is itself a finding: your logging or storage is not what you think.
- Locate candidate payloads. Scan untrusted segments for instructions, hidden characters and markup, and look for overlap between tool arguments and untrusted text, such as an address or URL that appears in a segment and nowhere in the trusted turns.
- Attribute. Confirm the candidate is causal, not just present: remove or neutralise it, replay the call with the recorded parameters several times, and compare how often the action recurs.
- Measure blast radius. Search every manifest for the payload's segment hash, its source and near-duplicates of its text, then check which of those sessions produced tool calls.
- Eradicate and learn. Purge or quarantine the source, rotate anything exfiltrated, and add the payload and its vector to the injection evaluation suite as a permanent regression case.
Locating the payload
Payload location is a triage step, so favour recall: flag everything plausible and let a human decide. Three signals carry most of the weight. Instruction-like language in an untrusted segment, especially addressing an AI or assistant. Invisible or deceptive characters: zero-width spaces and joiners, bidirectional overrides, and the Unicode Tags block, U+E0000 to U+E007F, whose code points mirror ASCII but render as nothing in most interfaces. And argument overlap: a value in the harmful tool call that appears in an untrusted segment and not in trusted ones.
import re, unicodedata
INVISIBLE = re.compile("[----]")
TAGS = re.compile("[\U000e0000-\U000e007f]")
IMPERATIVE = re.compile(r"(?i)\b(ignore|disregard|override)\b.{0,40}\b(instruction|prompt|rule)s?\b"
r"|\b(you are now|as an ai|assistant:|system:)")
def decode_tags(s):
return "".join(chr(ord(ch) - 0xE0000) for ch in s if 0xE0020 <= ord(ch) <= 0xE007E)
def triage(segments, tool_args):
findings = []
for seg in segments:
if seg["trust"] == "trusted":
continue
t = seg["text"]
if TAGS.search(t):
findings.append((seg["seq"], "unicode_tags", decode_tags(t)[:200]))
if INVISIBLE.search(t):
findings.append((seg["seq"], "invisible_chars", len(INVISIBLE.findall(t))))
visible = unicodedata.normalize("NFKC", TAGS.sub("", INVISIBLE.sub("", t)))
for m in IMPERATIVE.finditer(visible):
findings.append((seg["seq"], "imperative", visible[max(0, m.start() - 80):m.end() + 120]))
for value in tool_args:
if len(value) >= 6 and value in visible:
findings.append((seg["seq"], "arg_overlap", value))
return findingsRun this over the rendered text the model saw, after any HTML-to-text conversion your pipeline performs, and separately over the raw source: hidden text in markup, such as white-on-white spans or HTML comments, is visible to the model after extraction but invisible to anyone viewing the page. The regular expressions are deliberately crude; they are meant to point an investigator at the right segment, not to serve as a detector. Language other than English, encoded payloads and paraphrased instructions will slip past them, which is why argument overlap and replay-based attribution matter more than the keyword hits.
Worked example: a poisoned support ticket
Consider a hypothetical support agent that can read tickets, search the knowledge base, look up customer records and call http_get to check order-tracking pages. On a Monday an anomaly alert fires: outbound requests to an unfamiliar domain carrying long query strings. The investigation, reconstructed from evidence the system already recorded:
| Time (UTC) | Evidence | Finding |
|---|---|---|
| Sat 14:02:11 | Manifest mc_7f31 | Context includes ticket T-88213 body as segment 2, trust untrusted |
| Sat 14:02:13 | Ledger tc_7f31_0 | lookup_customer for the ticket's account, a legitimate call |
| Sat 14:02:15 | Ledger tc_7f31_1 | http_get to an external domain with email and address in the query string |
| Mon 09:40 | Triage on segment 2 | Unicode Tags payload decoding to an instruction to verify the customer by calling the URL with their details |
| Mon 10:15 | Replay, 10 runs each | Action recurs in 7 of 10 runs with segment 2 intact, 0 of 10 with the tag characters stripped |
| Mon 10:30 | Blast radius query | Same sender filed 14 tickets; 9 processed by the agent, 6 produced outbound calls |
The ticket body looked like an ordinary delivery complaint to the human who skimmed it; the instruction was entirely in invisible tag characters. The domain did not appear in any trusted segment, so argument overlap would have flagged segment 2 even without the Unicode check. Replay showed causation, not just presence. The blast radius query turned one incident into six notifications, which is the number the incident commander actually needed.
Measuring blast radius
Blast radius is a set of joins over the manifest and ledger. With manifests in a warehouse table keyed by segment hash, the core query is short:
-- sessions whose context contained any segment from the attacker's tickets,
-- and the outbound tool calls those sessions made afterwards
SELECT m.session_id, m.call_id, m.ts, t.tool, t.args
FROM manifest_segments s
JOIN model_calls m ON m.call_id = s.call_id
LEFT JOIN tool_calls t ON t.session_id = m.session_id AND t.ts >= m.ts
WHERE s.sha256 IN (SELECT sha256 FROM content WHERE source LIKE 'ticket:%'
AND source_owner = 'reporter:r_3390')
OR s.text_simhash_bucket IN (:payload_buckets) -- near-duplicate variants
ORDER BY m.ts;Exact hashes find re-reads of the same content; a near-duplicate index such as SimHash or MinHash buckets finds lightly edited variants the attacker posted elsewhere. Include derived content too: if a summary of the poisoned ticket was written to memory or a knowledge base, every later session that retrieved the summary is exposed, so follow derivation links from segment to every artefact it produced. The response steps around this, from containment to notification, are in LLM Incident Response.
Failure modes that destroy evidence
- Logging after truncation. The application logs the prompt after a token-budget trimmer runs, or before a retrieval step appends documents. Log the context exactly as sent to the model.
- IDs without content. Recording
doc_id=884but not its bytes; the document was edited the next day and the payload is gone. Store untrusted content by hash at read time. - No call linkage. Tool calls logged without the model call that requested them make the backward walk impossible in multi-step agents.
- Redaction that destroys evidence. A PII scrubber that rewrites logs also strips invisible characters or rewrites URLs. Redact in the analytics copy, not in the evidence store.
- Lossy summarisation. Long-running agents compress history; once the payload is summarised into "customer asked us to verify details" its origin is untraceable unless the summary records which segments it was built from.
- Single-run attribution. Replaying once and seeing no action proves nothing at non-zero temperature. Replay many times with and without the candidate.
Trade-offs
Full context capture is expensive and sensitive. Storing every segment of every call multiplies log volume and puts customer data into another store with its own access risk. A practical compromise stores manifests for every call, raw content for untrusted and tool-result segments only, deduplicated by hash, and trusted prompt templates by version rather than by copy. Retention is a second trade-off: injection payloads may sit dormant for weeks, so keep manifests longer than raw content, and keep raw content for at least as long as your detection lag. Finally, planted canary strings, described in LLM security canary tokens, turn some exfiltration from something you reconstruct into something you are alerted to.
What to do next
- Add a context manifest to every model call: segment order, source, trust label and SHA-256.
- Write untrusted and tool-result content to an immutable, hash-addressed store at read time.
- Stamp every tool call with the
call_idof the model call that requested it. - Build the triage script over rendered and raw text, including Unicode Tags and invisible characters.
- Write and test the blast-radius query before you need it, including derived summaries and memory.
- Run a tabletop exercise: plant a hidden payload in staging, then time how long the backward walk takes.
- Feed every confirmed payload into your injection evaluation suite as a regression case.