In 2011, Hutchins, Cloppert and Amin at Lockheed Martin described intrusions as a chain of seven stages: reconnaissance, weaponization, delivery, exploitation, installation, command and control, and actions on objectives. Their point was not the list. It was that a defender who breaks any one link stops the intrusion, and that collecting evidence per stage turns isolated alerts into a picture of a campaign.
LLM applications need that framing badly. Teams treat prompt injection as one event: a bad prompt came in, so add a filter. Real attacks on LLM agents are chains. Someone plants text in a place the agent will read, the model follows it, it persists somewhere, and it ends with a tool call or a rendered link that moves data out. This page maps each stage to its LLM form, gives the evidence and controls per stage, walks through a zero-click email chain, and shows how to tag telemetry so your detections can see two linked stages at once. It is a defender's model and contains no working payloads.
The seven stages in LLM form
| Stage | Classic meaning | LLM application form | Evidence you can collect |
|---|---|---|---|
| Reconnaissance | scan hosts, find people | probe the system prompt, list tools, test which filters fire | bursts of refusals, prompt-extraction phrasing, tool enumeration |
| Weaponization | build the exploit | write injected instructions, adversarial suffixes, encoded payloads | none on your side; red-team corpora approximate it |
| Delivery | email, USB, web | email, shared doc, web page, code comment, PR, MCP tool description | untrusted content entering context, with its source |
| Exploitation | code runs | the model treats data as instructions | output that deviates from the user's request; scanner hits |
| Installation | malware persists | text written to long-term memory, RAG index, agent config or a shared file | memory and index writes caused by untrusted content |
| Command and control | beacon to attacker | agent re-reads attacker-controlled pages or memory on later turns | repeated fetches of the same external source |
| Actions on objectives | steal, encrypt, destroy | exfiltrate via tool calls or rendered URLs, send mail, change data | outbound requests carrying user data, unexpected writes |
Stage by stage: evidence and controls
Reconnaissance. Attackers learn what tools an agent has, what its system prompt says and which guardrails fire. Many of those probes are cheap and noisy, such as many near-identical requests in a short window, or questions about the model's instructions. Rate limits per identity and a detector for system-prompt extraction attempts give early warning. Do not rely on secrecy of the system prompt; assume it leaks and keep secrets out of it.
Weaponization. This happens on the attacker's machine. Your counterpart is a red team that builds the same payloads first, as covered in running an AI red-team program, so your detections are tested on realistic inputs.
Delivery. In LLM systems the delivery vector is any content the model reads: inbound email, documents in a shared drive, a web page fetched by a browsing tool, issue comments, and the descriptions of third-party tools. The key control is provenance. Every chunk that enters context should carry a tag saying where it came from and whether it is trusted, as explained in prompt injection via RAG.
Exploitation. The model follows instructions found in data. No filter reliably prevents this today, so design for it to sometimes succeed. Use instruction-hierarchy prompting, injection scanners on untrusted chunks and, most important, constraints on what a hijacked model can do next.
Installation. Persistence in an LLM app means the injected text survives the session: it gets summarized into long-term memory, indexed into a vector store, written to a notes file the agent reads at start-up, or committed to a repository. Gate memory and index writes. A write whose content derives from untrusted input should need a stricter check or a human.
Command and control. Classic C2 needs a beacon. LLM C2 can be passive: the agent re-fetches an attacker-controlled page each time it runs, and the attacker edits the page to change the instructions. Allow-listed fetch domains and caching of fetched content by hash break this, and repeated fetches of one unknown source are a useful signal.
Actions on objectives. The usual goal is exfiltration through a channel the defender forgot: a markdown image whose URL contains the data, a link the user is invited to click, a webhook tool, or an email-sending tool. Strip or proxy external images in rendered output, put outbound tools behind egress control for agents, and require approval for actions that send data out of the tenant.
Worked example: a zero-click email chain
CVE-2025-32711, known as EchoLeak and disclosed in June 2025, was a zero-click information disclosure in Microsoft 365 Copilot reported by Aim Labs and fixed by Microsoft on the service side. Microsoft reported no evidence of exploitation in the wild. The public write-ups describe a chain that maps cleanly onto the seven stages, so we use it as a generalized example rather than reproducing its details:
- Recon: the attacker knows the target uses an assistant that can search the user's mailbox and files, and which kinds of output the client renders.
- Weaponize: they write an ordinary-looking business email whose text, addressed to the assistant, asks it to include certain internal information in a reference link. It is phrased to slip past injection classifiers.
- Deliver: they send the email. Nobody needs to open it.
- Exploit: later, the user asks the assistant an unrelated question. Retrieval pulls the email into context because it is topically similar, and the model follows its instructions alongside the user's.
- Install: nothing is needed; the email stays in the mailbox and is retrieved again on future queries, which gives the attack persistence for free.
- C2: the attacker can send new emails to change the instructions.
- Act: the response contains a URL with internal data in its query string, and the client fetches it automatically, for example as an image. The data reaches an attacker-controlled server without a click.
Look at where the chain could break. Provenance tagging would have marked the email as external and untrusted. A policy that external content cannot direct the inclusion of internal data would have blocked exploitation. Rendering only allow-listed image domains, or proxying images without query strings, would have killed the exfiltration even if everything earlier succeeded. The last control is deterministic, which is why it matters most: it does not depend on a classifier winning.
Tagging telemetry by stage
To use the chain for detection, each event needs a session ID, a stage, and a trust label for the content involved. Then a rule can look for the pattern that matters: untrusted content entered context, and within the same session the agent took an outbound action. Here is a minimal version:
from dataclasses import dataclass, field
from collections import defaultdict
import time
@dataclass
class Event:
session: str
stage: str # "delivery", "exploitation", "installation", "c2", "action"
source: str # e.g. "email:external", "tool:http_get", "memory:write"
trusted: bool
detail: dict = field(default_factory=dict)
ts: float = field(default_factory=time.time)
class ChainCorrelator:
def __init__(self, window_s=900):
self.window_s = window_s
self.untrusted = defaultdict(list) # session -> times untrusted content was ingested
def observe(self, ev: Event):
if ev.stage == "delivery" and not ev.trusted:
self.untrusted[ev.session].append(ev.ts)
return None
recent = [t for t in self.untrusted[ev.session] if ev.ts - t <= self.window_s]
if ev.stage in ("action", "installation") and recent:
return {"alert": "untrusted-ingest-then-" + ev.stage,
"session": ev.session, "source": ev.source,
"detail": ev.detail, "lag_s": round(ev.ts - max(recent), 1)}
return NoneIn production this rule runs in your SIEM over structured logs from the gateway and tool layer. The important design decision is upstream: log the provenance of every context chunk and every tool call with the same session ID, or no correlation is possible. Add canary tokens to sensitive documents so an outbound request that contains one is an unambiguous action-stage signal.
Where the kill chain does not fit
The kill chain is linear and intrusion-centric, and LLM attacks are not always either. Jailbreaks of a public chatbot may jump straight from weaponization to the attacker's own objective, with no installation or C2. Insider misuse skips delivery. Poisoned training data or a backdoored model weight file is a supply-chain attack that sits before stage 1. Pair the chain with MITRE ATLAS, which catalogs adversary tactics and techniques for AI systems in a matrix, and use the chain for what it does best: deciding where to put controls and how to connect alerts.
A stage coverage matrix
A useful exercise is a coverage matrix: one row per stage, one column per control you actually run, and a mark only where the control is deployed and tested, not planned. Here it is for a typical coding agent that reads issues, edits a repository and can call a web-fetch tool:
| Stage | Control in place | Deterministic? | Gap |
|---|---|---|---|
| Delivery | issue text labeled external | yes | PR review comments are not labeled |
| Exploitation | injection scanner on issue text | no | fluent instructions pass |
| Installation | none | n/a | agent can edit its own rules file |
| C2 | none | n/a | web fetch allows any domain |
| Actions | push requires human approval | yes | web fetch can carry data in the URL |
Reading the matrix shows the next two fixes immediately: make the agent's configuration and rules files read-only to the agent, and restrict web fetch to an allow list. Neither depends on a detector being right. Re-run the exercise whenever a tool or content source is added, since each one opens a new delivery or action path.
Failure modes
- All controls at one stage. Teams stack three injection classifiers at exploitation and leave actions unconstrained. One bypass then wins the whole chain.
- No provenance. Without source tags on context, every later signal is ambiguous: you cannot tell an action the user asked for from one an email asked for.
- Unlogged memory writes. Persistence goes unnoticed, and the injected text resurfaces for weeks after the original content is deleted.
- Treating rendering as harmless. Markdown images, link previews and auto-fetched URLs are network egress triggered by model output.
- Per-request alerting only. Each step looks benign alone; only per-session correlation shows the chain.
Trade-offs
Breaking links late, at actions, is the most reliable place because it is deterministic, but it constrains useful features such as rich rendering and autonomous email. Breaking links early, at delivery and exploitation, keeps features open but depends on probabilistic detectors that will miss some attacks. Human approval for outbound actions is very effective and costs user patience, so reserve it for actions that cross the tenant boundary. Full provenance logging costs storage and raises privacy questions of its own; log references and hashes of content where you can, not full text.
What to do next
- List every source of content your models read, and label each one trusted or untrusted.
- Add provenance tags to context chunks and log them with a shared session ID alongside tool calls.
- Pick one deterministic control at the action stage: image and link allow lists, or approvals for outbound tools.
- Gate long-term memory and index writes that derive from untrusted content.
- Implement the untrusted-ingest-then-action correlation rule and test it with red-team payloads.
- Map your controls onto all seven stages and fix the stage that has none.
- Keep learning: MITRE ATLAS, prompt injection via RAG, egress control for agents and canary tokens.