In 2011, Hutchins, Cloppert and Amin at Lockheed Martin described intrusions as a chain of seven stages: reconnaissance, weaponization, delivery, exploitation, installation, command and control, and actions on objectives. Their point was not the list. It was that a defender who breaks any one link stops the intrusion, and that collecting evidence per stage turns isolated alerts into a picture of a campaign.

LLM applications need that framing badly. Teams treat prompt injection as one event: a bad prompt came in, so add a filter. Real attacks on LLM agents are chains. Someone plants text in a place the agent will read, the model follows it, it persists somewhere, and it ends with a tool call or a rendered link that moves data out. This page maps each stage to its LLM form, gives the evidence and controls per stage, walks through a zero-click email chain, and shows how to tag telemetry so your detections can see two linked stages at once. It is a defender's model and contains no working payloads.

The seven stages in LLM form

The seven kill-chain stages, rewritten for an LLM application1 Reconprobing2 Weaponizecraft text3 Deliveremail, docs4 Exploitmodel obeys5 Installmemory, RAG6 C2re-fetch7 Actexfil, toolsControls that break each linkrate limitsprobedetectionred-teamthe samepayloadsprovenancetags oncontentinstructionhierarchy,scannerswrite-gatedmemoryallow-listedfetchegresscontrol,approvalstelemetry tagged by stagecorrelate per session: untrusted ingest followed by outbound actionBreaking any one link stops this chain; detecting two linked stages in one session is a high-signal alert.Stages 1 and 2 happen off your systems, so most defensive value sits in stages 3 to 7.
Figure: the seven stages in LLM form, with the control that breaks each link and the cross-stage correlation that makes detection practical.
StageClassic meaningLLM application formEvidence you can collect
Reconnaissancescan hosts, find peopleprobe the system prompt, list tools, test which filters firebursts of refusals, prompt-extraction phrasing, tool enumeration
Weaponizationbuild the exploitwrite injected instructions, adversarial suffixes, encoded payloadsnone on your side; red-team corpora approximate it
Deliveryemail, USB, webemail, shared doc, web page, code comment, PR, MCP tool descriptionuntrusted content entering context, with its source
Exploitationcode runsthe model treats data as instructionsoutput that deviates from the user's request; scanner hits
Installationmalware persiststext written to long-term memory, RAG index, agent config or a shared filememory and index writes caused by untrusted content
Command and controlbeacon to attackeragent re-reads attacker-controlled pages or memory on later turnsrepeated fetches of the same external source
Actions on objectivessteal, encrypt, destroyexfiltrate via tool calls or rendered URLs, send mail, change dataoutbound requests carrying user data, unexpected writes

Stage by stage: evidence and controls

Reconnaissance. Attackers learn what tools an agent has, what its system prompt says and which guardrails fire. Many of those probes are cheap and noisy, such as many near-identical requests in a short window, or questions about the model's instructions. Rate limits per identity and a detector for system-prompt extraction attempts give early warning. Do not rely on secrecy of the system prompt; assume it leaks and keep secrets out of it.

Weaponization. This happens on the attacker's machine. Your counterpart is a red team that builds the same payloads first, as covered in running an AI red-team program, so your detections are tested on realistic inputs.

Delivery. In LLM systems the delivery vector is any content the model reads: inbound email, documents in a shared drive, a web page fetched by a browsing tool, issue comments, and the descriptions of third-party tools. The key control is provenance. Every chunk that enters context should carry a tag saying where it came from and whether it is trusted, as explained in prompt injection via RAG.

Exploitation. The model follows instructions found in data. No filter reliably prevents this today, so design for it to sometimes succeed. Use instruction-hierarchy prompting, injection scanners on untrusted chunks and, most important, constraints on what a hijacked model can do next.

Installation. Persistence in an LLM app means the injected text survives the session: it gets summarized into long-term memory, indexed into a vector store, written to a notes file the agent reads at start-up, or committed to a repository. Gate memory and index writes. A write whose content derives from untrusted input should need a stricter check or a human.

Command and control. Classic C2 needs a beacon. LLM C2 can be passive: the agent re-fetches an attacker-controlled page each time it runs, and the attacker edits the page to change the instructions. Allow-listed fetch domains and caching of fetched content by hash break this, and repeated fetches of one unknown source are a useful signal.

Actions on objectives. The usual goal is exfiltration through a channel the defender forgot: a markdown image whose URL contains the data, a link the user is invited to click, a webhook tool, or an email-sending tool. Strip or proxy external images in rendered output, put outbound tools behind egress control for agents, and require approval for actions that send data out of the tenant.

Worked example: a zero-click email chain

CVE-2025-32711, known as EchoLeak and disclosed in June 2025, was a zero-click information disclosure in Microsoft 365 Copilot reported by Aim Labs and fixed by Microsoft on the service side. Microsoft reported no evidence of exploitation in the wild. The public write-ups describe a chain that maps cleanly onto the seven stages, so we use it as a generalized example rather than reproducing its details:

  1. Recon: the attacker knows the target uses an assistant that can search the user's mailbox and files, and which kinds of output the client renders.
  2. Weaponize: they write an ordinary-looking business email whose text, addressed to the assistant, asks it to include certain internal information in a reference link. It is phrased to slip past injection classifiers.
  3. Deliver: they send the email. Nobody needs to open it.
  4. Exploit: later, the user asks the assistant an unrelated question. Retrieval pulls the email into context because it is topically similar, and the model follows its instructions alongside the user's.
  5. Install: nothing is needed; the email stays in the mailbox and is retrieved again on future queries, which gives the attack persistence for free.
  6. C2: the attacker can send new emails to change the instructions.
  7. Act: the response contains a URL with internal data in its query string, and the client fetches it automatically, for example as an image. The data reaches an attacker-controlled server without a click.

Look at where the chain could break. Provenance tagging would have marked the email as external and untrusted. A policy that external content cannot direct the inclusion of internal data would have blocked exploitation. Rendering only allow-listed image domains, or proxying images without query strings, would have killed the exfiltration even if everything earlier succeeded. The last control is deterministic, which is why it matters most: it does not depend on a classifier winning.

Tagging telemetry by stage

To use the chain for detection, each event needs a session ID, a stage, and a trust label for the content involved. Then a rule can look for the pattern that matters: untrusted content entered context, and within the same session the agent took an outbound action. Here is a minimal version:

from dataclasses import dataclass, field
from collections import defaultdict
import time

@dataclass
class Event:
    session: str
    stage: str            # "delivery", "exploitation", "installation", "c2", "action"
    source: str           # e.g. "email:external", "tool:http_get", "memory:write"
    trusted: bool
    detail: dict = field(default_factory=dict)
    ts: float = field(default_factory=time.time)

class ChainCorrelator:
    def __init__(self, window_s=900):
        self.window_s = window_s
        self.untrusted = defaultdict(list)    # session -> times untrusted content was ingested

    def observe(self, ev: Event):
        if ev.stage == "delivery" and not ev.trusted:
            self.untrusted[ev.session].append(ev.ts)
            return None
        recent = [t for t in self.untrusted[ev.session] if ev.ts - t <= self.window_s]
        if ev.stage in ("action", "installation") and recent:
            return {"alert": "untrusted-ingest-then-" + ev.stage,
                    "session": ev.session, "source": ev.source,
                    "detail": ev.detail, "lag_s": round(ev.ts - max(recent), 1)}
        return None

In production this rule runs in your SIEM over structured logs from the gateway and tool layer. The important design decision is upstream: log the provenance of every context chunk and every tool call with the same session ID, or no correlation is possible. Add canary tokens to sensitive documents so an outbound request that contains one is an unambiguous action-stage signal.

Where the kill chain does not fit

The kill chain is linear and intrusion-centric, and LLM attacks are not always either. Jailbreaks of a public chatbot may jump straight from weaponization to the attacker's own objective, with no installation or C2. Insider misuse skips delivery. Poisoned training data or a backdoored model weight file is a supply-chain attack that sits before stage 1. Pair the chain with MITRE ATLAS, which catalogs adversary tactics and techniques for AI systems in a matrix, and use the chain for what it does best: deciding where to put controls and how to connect alerts.

A stage coverage matrix

A useful exercise is a coverage matrix: one row per stage, one column per control you actually run, and a mark only where the control is deployed and tested, not planned. Here it is for a typical coding agent that reads issues, edits a repository and can call a web-fetch tool:

StageControl in placeDeterministic?Gap
Deliveryissue text labeled externalyesPR review comments are not labeled
Exploitationinjection scanner on issue textnofluent instructions pass
Installationnonen/aagent can edit its own rules file
C2nonen/aweb fetch allows any domain
Actionspush requires human approvalyesweb fetch can carry data in the URL

Reading the matrix shows the next two fixes immediately: make the agent's configuration and rules files read-only to the agent, and restrict web fetch to an allow list. Neither depends on a detector being right. Re-run the exercise whenever a tool or content source is added, since each one opens a new delivery or action path.

Failure modes

  • All controls at one stage. Teams stack three injection classifiers at exploitation and leave actions unconstrained. One bypass then wins the whole chain.
  • No provenance. Without source tags on context, every later signal is ambiguous: you cannot tell an action the user asked for from one an email asked for.
  • Unlogged memory writes. Persistence goes unnoticed, and the injected text resurfaces for weeks after the original content is deleted.
  • Treating rendering as harmless. Markdown images, link previews and auto-fetched URLs are network egress triggered by model output.
  • Per-request alerting only. Each step looks benign alone; only per-session correlation shows the chain.

Trade-offs

Breaking links late, at actions, is the most reliable place because it is deterministic, but it constrains useful features such as rich rendering and autonomous email. Breaking links early, at delivery and exploitation, keeps features open but depends on probabilistic detectors that will miss some attacks. Human approval for outbound actions is very effective and costs user patience, so reserve it for actions that cross the tenant boundary. Full provenance logging costs storage and raises privacy questions of its own; log references and hashes of content where you can, not full text.

What to do next

  1. List every source of content your models read, and label each one trusted or untrusted.
  2. Add provenance tags to context chunks and log them with a shared session ID alongside tool calls.
  3. Pick one deterministic control at the action stage: image and link allow lists, or approvals for outbound tools.
  4. Gate long-term memory and index writes that derive from untrusted content.
  5. Implement the untrusted-ingest-then-action correlation rule and test it with red-team payloads.
  6. Map your controls onto all seven stages and fix the stage that has none.
  7. Keep learning: MITRE ATLAS, prompt injection via RAG, egress control for agents and canary tokens.
Key takeaway: Attacks on LLM applications are chains, not single prompts: untrusted text is delivered, the model follows it, it persists in memory or an index, and it ends in an outbound action. Map your controls to every stage, put at least one deterministic control at the action stage, tag every context chunk with its provenance, and alert on untrusted ingestion followed by an outbound action in the same session.