Prompt injection forensics is the work of answering, after the fact, four questions about an incident in an LLM application: what did the system do that it should not have, which text made it do that, where did that text come from, and who else was exposed to it. The questions are familiar from any security investigation. What makes injection different is that the attacker never touched your systems directly. They wrote text, a web page, an email, a support ticket or a file name, and your own application carried that text into a model's context, where it was treated as instruction.

So the evidence lives in places ordinary logs rarely capture: the exact context window of one model call, the provenance of each piece of it, and the link between that call and the tool calls that followed. This article covers the evidence you must record before any incident, a step-by-step investigation that walks backwards from the harmful effect to the payload, code for locating payloads including hidden ones, a worked incident timeline, blast-radius queries and the failure modes that destroy evidence. General techniques such as chain of custody and replay with context ablation are covered in AI Forensics, in depth; this article stays on what is specific to injection.

Why injection incidents are different

In a conventional intrusion the attacker's activity leaves artefacts: logins, processes, network connections from their infrastructure. In an injection the malicious actions are performed by your agent, under its own identity, using permissions you granted it. Every log line looks legitimate. The only attacker-controlled artefact is a span of text, and that span may have been fetched from a page that has since changed, summarised by an earlier model call so its wording is gone, or hidden in characters no human reviewer sees.

That shapes the investigation. You cannot start from the attacker and work forwards, because you do not know who they are. You start from the effect, the email that was sent or the record that was changed, and walk backwards through the chain that produced it: tool call, the model call that emitted it, the context that call saw, the segment within that context, and the source of that segment. Every link in that chain must be recorded at runtime, because none of it can be reconstructed afterwards. The attack class itself, and the architectural defenses that make these incidents rarer, are covered in Indirect Prompt Injection in Depth.

Evidence you must record in advance

Injection forensics: walk backwards from the effect to the segment that carried the payloadUntrusted sourcesweb, email, tickets, filesTrusted sourcessystem prompt, user turnTool resultsprevious callsContext assemblerwrites context manifest:segment id, source, hash, trustModel callcall_id, model, paramsTool call ledgercall_id, args, resultSide effectsmail sent, HTTP outInvestigation path (right to left)effect, tool call, model call, manifest, segment, source, other sessions that read itContent storeimmutable snapshot by hashJoin keys (call_id, segment hash) are what make the backward walk possible. Without them it is guesswork.
The context assembler records a manifest of every segment it places in a model call. Tool calls carry the call_id of the model call that requested them. Raw content is stored by hash so a changed source cannot erase evidence.

The minimum evidence set has three records. The context manifest lists, for each model call, every segment placed in the context, in order, with its source, a content hash and a trust label. The tool call ledger records each tool invocation with the identifier of the model call that requested it, its arguments and its result. The content store keeps the raw bytes of every untrusted segment, addressed by hash and written once, so the evidence survives even when the original web page or ticket is edited or deleted.

{
  "call_id": "mc_7f31",
  "session_id": "s_19ac",
  "ts": "2026-09-26T14:02:11Z",
  "model": "provider/model-version",
  "params": {"temperature": 0.2, "max_tokens": 1024},
  "segments": [
    {"seq": 0, "role": "system", "source": "prompt:support_v14",   "trust": "trusted",   "sha256": "9b2e..."},
    {"seq": 1, "role": "user",   "source": "user:u_5521",          "trust": "trusted",   "sha256": "41c0..."},
    {"seq": 2, "role": "tool",   "source": "ticket:T-88213/body",  "trust": "untrusted", "sha256": "e7a4..."},
    {"seq": 3, "role": "tool",   "source": "kb:refund-policy#3",   "trust": "internal",  "sha256": "02fd..."}
  ],
  "output_sha256": "c55d...",
  "tool_calls": ["tc_7f31_0", "tc_7f31_1"]
}

Hashing per segment rather than per prompt is the important design choice. It lets you ask "which sessions contained this exact ticket body" with an index lookup, and it lets you prove later that the bytes you are analysing are the bytes the model saw. Store the manifest in the tamper-evident trail described in LLM audit logging architecture, and apply the same access controls to it as to the data it contains.

The investigation, step by step

  1. Fix the effect. Identify the harmful side effect precisely: which tool, which arguments, what time, which session. Preserve the ledger entries and content-store objects under legal hold before retention jobs touch them.
  2. Find the deciding call. Follow the tool call's call_id to the model call that emitted it. In multi-step agents, also collect earlier calls in the session, because the payload may have entered several steps before the action, often via a summary that carried it forward.
  3. Rebuild the context. Reassemble the exact context from the manifest and content store, verifying every hash. A mismatch is itself a finding: your logging or storage is not what you think.
  4. Locate candidate payloads. Scan untrusted segments for instructions, hidden characters and markup, and look for overlap between tool arguments and untrusted text, such as an address or URL that appears in a segment and nowhere in the trusted turns.
  5. Attribute. Confirm the candidate is causal, not just present: remove or neutralise it, replay the call with the recorded parameters several times, and compare how often the action recurs.
  6. Measure blast radius. Search every manifest for the payload's segment hash, its source and near-duplicates of its text, then check which of those sessions produced tool calls.
  7. Eradicate and learn. Purge or quarantine the source, rotate anything exfiltrated, and add the payload and its vector to the injection evaluation suite as a permanent regression case.

Locating the payload

Payload location is a triage step, so favour recall: flag everything plausible and let a human decide. Three signals carry most of the weight. Instruction-like language in an untrusted segment, especially addressing an AI or assistant. Invisible or deceptive characters: zero-width spaces and joiners, bidirectional overrides, and the Unicode Tags block, U+E0000 to U+E007F, whose code points mirror ASCII but render as nothing in most interfaces. And argument overlap: a value in the harmful tool call that appears in an untrusted segment and not in trusted ones.

import re, unicodedata

INVISIBLE = re.compile("[​-‏‪-‮⁠-⁤⁦-⁩]")
TAGS = re.compile("[\U000e0000-\U000e007f]")
IMPERATIVE = re.compile(r"(?i)\b(ignore|disregard|override)\b.{0,40}\b(instruction|prompt|rule)s?\b"
                        r"|\b(you are now|as an ai|assistant:|system:)")

def decode_tags(s):
    return "".join(chr(ord(ch) - 0xE0000) for ch in s if 0xE0020 <= ord(ch) <= 0xE007E)

def triage(segments, tool_args):
    findings = []
    for seg in segments:
        if seg["trust"] == "trusted":
            continue
        t = seg["text"]
        if TAGS.search(t):
            findings.append((seg["seq"], "unicode_tags", decode_tags(t)[:200]))
        if INVISIBLE.search(t):
            findings.append((seg["seq"], "invisible_chars", len(INVISIBLE.findall(t))))
        visible = unicodedata.normalize("NFKC", TAGS.sub("", INVISIBLE.sub("", t)))
        for m in IMPERATIVE.finditer(visible):
            findings.append((seg["seq"], "imperative", visible[max(0, m.start() - 80):m.end() + 120]))
        for value in tool_args:
            if len(value) >= 6 and value in visible:
                findings.append((seg["seq"], "arg_overlap", value))
    return findings

Run this over the rendered text the model saw, after any HTML-to-text conversion your pipeline performs, and separately over the raw source: hidden text in markup, such as white-on-white spans or HTML comments, is visible to the model after extraction but invisible to anyone viewing the page. The regular expressions are deliberately crude; they are meant to point an investigator at the right segment, not to serve as a detector. Language other than English, encoded payloads and paraphrased instructions will slip past them, which is why argument overlap and replay-based attribution matter more than the keyword hits.

Worked example: a poisoned support ticket

Consider a hypothetical support agent that can read tickets, search the knowledge base, look up customer records and call http_get to check order-tracking pages. On a Monday an anomaly alert fires: outbound requests to an unfamiliar domain carrying long query strings. The investigation, reconstructed from evidence the system already recorded:

Time (UTC)EvidenceFinding
Sat 14:02:11Manifest mc_7f31Context includes ticket T-88213 body as segment 2, trust untrusted
Sat 14:02:13Ledger tc_7f31_0lookup_customer for the ticket's account, a legitimate call
Sat 14:02:15Ledger tc_7f31_1http_get to an external domain with email and address in the query string
Mon 09:40Triage on segment 2Unicode Tags payload decoding to an instruction to verify the customer by calling the URL with their details
Mon 10:15Replay, 10 runs eachAction recurs in 7 of 10 runs with segment 2 intact, 0 of 10 with the tag characters stripped
Mon 10:30Blast radius querySame sender filed 14 tickets; 9 processed by the agent, 6 produced outbound calls

The ticket body looked like an ordinary delivery complaint to the human who skimmed it; the instruction was entirely in invisible tag characters. The domain did not appear in any trusted segment, so argument overlap would have flagged segment 2 even without the Unicode check. Replay showed causation, not just presence. The blast radius query turned one incident into six notifications, which is the number the incident commander actually needed.

Measuring blast radius

Blast radius is a set of joins over the manifest and ledger. With manifests in a warehouse table keyed by segment hash, the core query is short:

-- sessions whose context contained any segment from the attacker's tickets,
-- and the outbound tool calls those sessions made afterwards
SELECT m.session_id, m.call_id, m.ts, t.tool, t.args
FROM   manifest_segments s
JOIN   model_calls m  ON m.call_id = s.call_id
LEFT JOIN tool_calls t ON t.session_id = m.session_id AND t.ts >= m.ts
WHERE  s.sha256 IN (SELECT sha256 FROM content WHERE source LIKE 'ticket:%'
                     AND source_owner = 'reporter:r_3390')
   OR  s.text_simhash_bucket IN (:payload_buckets)      -- near-duplicate variants
ORDER  BY m.ts;

Exact hashes find re-reads of the same content; a near-duplicate index such as SimHash or MinHash buckets finds lightly edited variants the attacker posted elsewhere. Include derived content too: if a summary of the poisoned ticket was written to memory or a knowledge base, every later session that retrieved the summary is exposed, so follow derivation links from segment to every artefact it produced. The response steps around this, from containment to notification, are in LLM Incident Response.

Failure modes that destroy evidence

  • Logging after truncation. The application logs the prompt after a token-budget trimmer runs, or before a retrieval step appends documents. Log the context exactly as sent to the model.
  • IDs without content. Recording doc_id=884 but not its bytes; the document was edited the next day and the payload is gone. Store untrusted content by hash at read time.
  • No call linkage. Tool calls logged without the model call that requested them make the backward walk impossible in multi-step agents.
  • Redaction that destroys evidence. A PII scrubber that rewrites logs also strips invisible characters or rewrites URLs. Redact in the analytics copy, not in the evidence store.
  • Lossy summarisation. Long-running agents compress history; once the payload is summarised into "customer asked us to verify details" its origin is untraceable unless the summary records which segments it was built from.
  • Single-run attribution. Replaying once and seeing no action proves nothing at non-zero temperature. Replay many times with and without the candidate.

Trade-offs

Full context capture is expensive and sensitive. Storing every segment of every call multiplies log volume and puts customer data into another store with its own access risk. A practical compromise stores manifests for every call, raw content for untrusted and tool-result segments only, deduplicated by hash, and trusted prompt templates by version rather than by copy. Retention is a second trade-off: injection payloads may sit dormant for weeks, so keep manifests longer than raw content, and keep raw content for at least as long as your detection lag. Finally, planted canary strings, described in LLM security canary tokens, turn some exfiltration from something you reconstruct into something you are alerted to.

What to do next

  1. Add a context manifest to every model call: segment order, source, trust label and SHA-256.
  2. Write untrusted and tool-result content to an immutable, hash-addressed store at read time.
  3. Stamp every tool call with the call_id of the model call that requested it.
  4. Build the triage script over rendered and raw text, including Unicode Tags and invisible characters.
  5. Write and test the blast-radius query before you need it, including derived summaries and memory.
  6. Run a tabletop exercise: plant a hidden payload in staging, then time how long the backward walk takes.
  7. Feed every confirmed payload into your injection evaluation suite as a regression case.
Key takeaway: An injection incident is investigated backwards: from the harmful effect to the tool call, the model call, the context it saw, the untrusted segment and its source, then outwards to every other session that read it. That walk is only possible if you record per-segment context manifests, tool call linkage and hash-addressed raw content before anything goes wrong.