When a traditional service misbehaves, investigators follow the code path. When an LLM application misbehaves, there is no code path in that sense. The behaviour came from a model reading a context window, and the context was assembled at runtime from a system prompt, conversation history, retrieved documents, tool results and memory. The main question in an AI incident is which of those inputs caused the output, and whether the same thing is happening anywhere else.

AI forensics is the discipline of answering that question with evidence that holds up. It borrows the structure of digital forensics: collect in order of volatility, preserve with a chain of custody, reconstruct a timeline and test hypotheses. It adds two techniques that only make sense for models: replaying a frozen context, and ablating parts of it to find the cause. This article walks through each step and a full worked investigation.

It assumes you already log the right things. How to design those logs is covered in LLM audit logging architecture. Here the logs are the starting point, and the job is to turn them into findings.

Advertisement

What makes AI incidents different

Four properties shape every investigation. First, the input is the program. A retrieved document can change behaviour as much as a code deploy, so the retrieval corpus, the memory store and every tool's output belong within the scope of the investigation. Second, behaviour is probabilistic. The same context may produce the harmful output one time in five, so "I replayed it and it was fine" proves nothing. Third, versions hide behind names. A model alias, a prompt template or a safety policy may have changed between the incident and the investigation, and replaying against today's configuration tests the wrong system. Fourth, the evidence is sensitive. Prompts and outputs contain user data, so collection must follow privacy rules even while it preserves integrity.

Name the questions before you collect anything. What happened, and when? Which input caused it? What data or actions left the system? Which other users, sessions or tenants were exposed to the same cause? Is it still happening? Every artifact you collect should help answer one of these questions.

The evidence inventory

EvidenceWhy it mattersVolatility
Assembled prompt as sent, after templatingThe only record of what the model actually read.Often not logged at all, and redacted or truncated when it is.
Model response, including tool-call requestsThe behaviour under investigation.Usually logged; check retention.
Retrieval records: query, document ids, chunk ids, index versionLinks an output to the documents that shaped it.High: re-indexing and document edits erase the state at incident time.
Tool calls: arguments, results, side effectsShows what the agent did, not just what it said.Downstream systems have their own retention rules.
Memory reads and writesA poisoned memory persists across sessions.High: memory is updated continuously.
Config versions: model id, parameters, prompt template, policyNeeded to replay the right system.Aliases can move silently.
Safety classifier and guardrail verdictsShows what the controls saw and why they allowed it.Often sampled or aggregated.
Identity and session metadataScopes the blast radius.Stable, but join keys must be consistent.

NIST SP 800-86, the guide to integrating forensic techniques into incident response, recommends collecting the most volatile data first. For AI systems, that order puts the retrieval index state and the memory store near the top. Snapshot them, or export the relevant documents with their versions, before anyone re-indexes or cleans up. A well-intentioned engineer who deletes a malicious document from the corpus has fixed the symptom and destroyed the best evidence.

Advertisement

Preservation and chain of custody

Export each evidence set to storage that cannot be modified, such as a write-once bucket or an append-only volume. Record a hash of every file together with where it came from, who collected it and when. A manifest in which each record includes the hash of the one before makes later tampering with any record detectable:

import hashlib, json, time
from pathlib import Path

def sha256(path):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        for block in iter(lambda: f.read(1 << 20), b""):
            h.update(block)
    return h.hexdigest()

def add_evidence(manifest: Path, item: Path, source: str, collector: str):
    """Append one evidence record; each record commits to the previous one."""
    lines = manifest.read_text().splitlines() if manifest.exists() else []
    prev = json.loads(lines[-1])["record_hash"] if lines else "0" * 64
    record = {
        "file": item.name,
        "sha256": sha256(item),
        "source": source,            # e.g. "gateway-logs export, eu-west, 02:10-04:00 UTC"
        "collector": collector,
        "collected_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
        "prev": prev,
    }
    record["record_hash"] = hashlib.sha256(
        json.dumps(record, sort_keys=True).encode()).hexdigest()
    with manifest.open("a") as f:
        f.write(json.dumps(record, sort_keys=True) + "\n")

Keep analysis on copies, never on the preserved originals. Apply the same access controls to evidence that apply to the production data it contains. An evidence bucket full of prompts is a data store, and it needs its own retention date and legal-hold process. If redaction is required, redact a copy and keep the mapping under stricter access. Redacting the only copy can remove the very string that proves an injection.

Reconstructing the timeline

Join all evidence streams on a correlation ID, ideally a trace ID propagated from the gateway through retrieval, the model call and every tool call. If there is no shared ID, join on session ID and time windows, and write down the uncertainty. Normalize every timestamp to UTC and check clock skew between systems using an event both of them recorded. Tool side effects are the usual weak point. An email API or a database write logs in its own system with its own clock, so match those events on content, such as recipient or record key, as well as time.

The result should read as a sequence of facts, each one citing its evidence: at 02:14:07 the retriever returned chunk 88 of document D-3312, version 4; at 02:14:09 the model requested send_email with these arguments; at 02:14:10 the mail provider accepted the message. Gaps in the timeline are findings too, because they show where your logging needs work.

An AI investigation joins five evidence streams into one timeline, then tests hypotheses by replayGateway logsrequest, response, idsAssembled promptsafter templatingRetrieval recordsdoc ids + versionsTool callsargs, results, effectsConfig versionsmodel, prompt, policyPreservehash, manifest, WORMTimelinejoin on trace ids, fix skewFrozen replaysame context, pinned modelAblationremove a segment, re-measureBlast radiuswho else saw the payloadFindingscause, scope, fixes
Five evidence streams are preserved, joined into one timeline, and then used for replay, ablation and blast-radius search.

Replay under nondeterminism

Replay means sending the exact assembled context back to the model with the recorded configuration and seeing whether the behaviour recurs. It is the model equivalent of reproducing a bug, and it has three rules. Pin everything: the specific model version rather than an alias, the sampling parameters, the system prompt version and the tool schemas. Stub the tools: replay must never send the email again, so every tool returns its recorded result. And replay many times, reporting a rate, not a yes or no.

Do not expect bit-exact reproduction, even at temperature zero. Serving systems batch requests together, floating-point reductions are not always performed in the same order, and hosted models may be updated behind a version name. Treat replay as a measurement. If the behaviour appears in 7 of 20 replays of the original context, that is a strong result. If it appears in none of them, the configuration may have drifted, or the behaviour may be rare, and the report should say which of these is more likely.

Attribution by context ablation

Once replay reproduces the behaviour at a measurable rate, find the cause by removing parts of the context one at a time and measuring the rate again. The segments are the natural units of the assembled prompt: each retrieved chunk, each tool result, each memory item and each earlier turn. A segment whose removal drops the rate close to zero is the cause, or part of it.

def behaviour_rate(segments, run, detector, trials):
    """Fraction of replays in which the bad behaviour appears."""
    prompt = [s.content for s in segments]
    return sum(detector(run(prompt)) for _ in range(trials)) / trials

def ablate(segments, run, detector, trials=20):
    base = behaviour_rate(segments, run, detector, trials)
    effects = []
    for i, seg in enumerate(segments):
        if seg.kind == "system":
            continue                         # ablate evidence, not the product itself
        without = segments[:i] + segments[i + 1:]
        drop = base - behaviour_rate(without, run, detector, trials)
        effects.append((seg.id, seg.kind, round(drop, 2)))
    return base, sorted(effects, key=lambda e: -e[2])

# run      = replay with the pinned model version, pinned parameters and stubbed tools
# detector = a deterministic check, e.g. "response contains a send_email call to an external domain"

Ablation has limits worth stating in the findings. Causes can be conjunctive: an injection may only work in combination with a particular earlier turn, and removing either one drops the rate. So when single-segment ablation is inconclusive, test pairs. Removing a segment also shifts the positions of everything after it, which can change behaviour by itself, so a replacement with neutral text of similar length is a useful control. The detector must be deterministic and specific. "The model sent an email to an external domain" is testable. "The model behaved badly" is not.

Worked example: an agent that emailed a customer list

A support agent with a send_email tool and retrieval over past tickets sent a message containing twelve customer email addresses to an external address. The report came from the recipient. Here is how the investigation proceeds.

  1. Contain and preserve. Disable the send_email tool for the agent. Snapshot the ticket index and export the incident session's gateway records, assembled prompts, retrieval records and tool logs, each added to the manifest.
  2. Timeline. The session began with an ordinary question about a refund. The retriever returned four chunks, one of them from a ticket submitted three days earlier by an unknown account. Two seconds later the model called send_email.
  3. Replay. With the model version and prompt pinned and the tools stubbed, the email call appears in 11 of 20 replays.
  4. Ablation. Removing the chunk from the unknown account's ticket drops the rate to 0 of 20. Removing any other segment leaves it between 9 and 12. Reading that chunk shows text addressed to "the assistant", instructing it to send a summary of recent customers to an outside address. This is indirect prompt injection through the retrieval corpus.
  5. Blast radius. Search retrieval records for every session that retrieved that document: 37 sessions, 3 of which issued send_email calls. Search the corpus for near-copies of the payload, and find two more tickets from related accounts.
  6. Findings. Root cause: untrusted ticket text was retrieved into the context with the same authority as instructions, and a sensitive tool could send email to any recipient without confirmation. Fixes: quarantine the tickets, restrict recipients to the ticket's own customer, require human approval for bulk addresses, and screen new tickets for instruction-like text.

The numbers in this example are illustrative, but the structure is general. Evidence shows the chain of events. Replay proves the behaviour is reproducible, ablation identifies the cause, and the blast-radius search turns one incident into a complete scope. Defences against this class of attack are covered in RAG defense and LLM tool abuse.

Root-cause classes

ClassEvidence signatureHardest part
Direct prompt injectionThe payload is in the user's own turns.Separating malicious users from confused ones.
Indirect injection via retrieval or toolsAblation points to a retrieved chunk or tool result.Finding every other session that ingested it.
Poisoned memoryThe cause is a memory item written in an earlier session.Tracing the write back to its source session.
Configuration regressionThe behaviour starts at a model, prompt or policy change and does not ablate to any input.Knowing what changed if versions were not recorded.
Excessive tool permissionThe model's request was plausible, and the tool allowed too much.Accepting that the fix is in the tool, not the prompt.
Training-time poisoning or backdoorThe trigger is rare and survives every context ablation.Needs model-level analysis, not log analysis.

MITRE ATLAS catalogues adversarial techniques against AI systems and is useful shared vocabulary for writing up the attack path, in the way ATT&CK is for conventional intrusions. For the output side of attribution, such as proving which model and prompt produced a given response, see LLM output provenance.

How investigations fail

  • The assembled prompt was never logged. Only the user message and the response exist, so the context cannot be rebuilt. Fix the logging before the next incident.
  • Replay ran against an alias. The model behind the name had changed, and the investigation concluded that the bug could not be reproduced.
  • Cleanup before capture. The poisoned document was deleted and re-indexed before anyone exported it.
  • Single-shot conclusions. One clean replay was taken as proof that the behaviour was gone.
  • Scope stopped at the reporting user. No one searched retrieval records for other sessions that saw the same payload.
  • Evidence became a new leak. Prompts full of personal data were copied into chat threads and tickets during the investigation.

What to do next

  1. Confirm that your logs record the assembled prompt, retrieval document and chunk IDs with versions, tool arguments and results, and pinned model and prompt versions, all joined by a trace ID.
  2. Write a collection runbook in order of volatility: index and memory snapshots first, then logs, then downstream systems.
  3. Set up a write-once evidence bucket and a hash-chained manifest script before you need them.
  4. Build a replay harness with pinned versions and stubbed tools, and practise on a known past issue.
  5. Add context ablation to the harness, with deterministic detectors for your top risks, such as external emails or secrets in output.
  6. Make blast-radius queries ready in advance: every session that retrieved document X, and every call to tool Y with argument pattern Z.
  7. Run a tabletop exercise on the worked example above, and record which step your current systems could not support.
Key takeaway: AI forensics treats the context window as the crime scene. Collect the most volatile evidence first, especially index and memory state, preserve it with hashes and a chained manifest, and join every stream into one timeline. Replay the frozen context against pinned versions many times and report a rate, then ablate segments to find the input that causes the behaviour. Finish by searching for every other session exposed to the same cause. Most investigations fail because of missing logs, drifted aliases or cleanup done too early, so prepare the logs, the harness and the runbook before the incident.