A conventional threat model draws trust boundaries between processes, networks and accounts, then asks what can go wrong where data crosses them. Applied unchanged to an application built on a large language model, it misses the most important boundary, because that boundary runs through the middle of the prompt. The system prompt, the user's message, a retrieved web page and the output of a tool are all concatenated into one sequence of tokens, and the model has no reliable way to treat some of them as instructions and the rest as data.
This page covers what changes for LLM systems; the general method is in the threat modeling deep dive. You will learn how to decompose an LLM application into useful elements, track trust through the context window, enumerate threats as paths from untrusted text to actions, apply the Agents Rule of Two, and rank and remove what you find, using a code-review agent in CI as the worked example.
Why LLM systems need their own threat model
Three properties make LLM systems different enough to need their own modelling habits.
Instructions and data share one channel. SQL injection is fixed with parameterised queries; there is no equivalent for a language model. Prompt injection, direct or indirect, is a property of the model as a component, not a bug in one place.
The model is a deputy with borrowed authority. When a model calls a tool, the call runs with the credentials of the application, not the person whose text caused it. That is the classic confused deputy problem, and it means the blast radius of a successful injection equals the union of everything the model can reach in that session.
Behaviour is probabilistic and composable. A filter that blocks 99 percent of attacks still lets an attacker who can retry a thousand times through. Capabilities also compose: a harmless read tool plus a harmless web-fetch tool become an exfiltration channel once untrusted text can steer both. You cannot assess tools one at a time.
So change the default assumption: any untrusted text that reaches a model's context can take control of its output. Then ask what that output can touch.
Elements of an LLM system
Decompose the system into elements whose security properties you can reason about. A data flow diagram still works, but use these element types rather than generic processes and stores:
| Element | Examples | Property to record |
|---|---|---|
| Principal | End user, operator, CI system, another agent | Whose authority each request carries |
| Context source | System prompt, user message, retrieved documents, tool results, memory, file contents | Who can write to it: trusted, internal, or untrusted |
| Model session | One conversation or agent run with one context window | Which sources feed it and which tools it can call |
| Tool | Search, read file, run code, send email, open a pull request | What it reads (private or public) and what it does (nothing, state change, external communication) |
| Sink | Rendered HTML, a shell, a database, an outbound HTTP request, a message to a human | Whether model output is interpreted rather than displayed |
| Secret | API keys, tokens in environment variables, customer records | Which tools can read it into a context |
| Model supply | Base weights, fine-tuning data, embeddings, prompt templates | Who can change them and how changes are reviewed |
The model session is the unit that matters most: two calls to the same model with different contexts and tools are different elements, which is why splitting sessions is such an effective control. Record the supply elements too, since the OWASP list covers supply chain (LLM03) and data and model poisoning (LLM04).
Trust inside the context window
Inside a session, attach a provenance label to every segment of the context: trusted (written by your team and reviewed), internal (written by your organisation's users or systems), or untrusted (anything a third party can influence, including web pages, inbound email, uploaded files, public issues and pull requests from forks). The rule that makes the labels useful is simple and conservative: the model's output carries the lowest trust label of anything in its context.
That rule has consequences people tend to miss. A tool result inherits the trust of whatever the tool read, so a search tool turns a trusted context into an untrusted one the moment it returns a web page. Memory launders trust: an untrusted email summarised into memory reappears next week looking like the agent's own note. Retrieval indexes do the same for whoever can write a document into them.
Then look at where output goes. Output that is interpreted, such as rendered markdown, HTML, SQL, shell commands or tool arguments, is a sink. A markdown image whose URL contains data is an outbound request the browser makes automatically, a common exfiltration route.
The lethal trifecta and the Rule of Two
Most serious LLM incidents need three things in the same session: exposure to untrusted content, access to private data, and a way to communicate externally. Simon Willison named this combination the lethal trifecta. Meta's Agents Rule of Two generalises it into a design rule: within one session an agent should have no more than two of these properties, and if it needs all three it should not act autonomously but under supervision such as human approval.
| Property | Rule of Two wording | Typical sources in practice |
|---|---|---|
| A | Processes untrustworthy inputs | Web browsing, inbound email, public tickets, user uploads, fork pull requests |
| B | Has access to sensitive systems or private data | Internal documents, customer data, source code, tokens in the environment |
| C | Can change state or communicate externally | Sending messages, HTTP requests, writing to repositories, rendering remote images |
The rule is useful precisely because it does not depend on detecting the attack. A session with A and B but no C can be fooled, but cannot send what it read anywhere. A session with A and C but no B can be made to say something rude to the outside world, which is a real harm but a bounded one. The threat model's job is to find each session's letters and treat every A-B-C session as a finding in its own right.
Worked example: a code-review agent in CI
Take a common design: an LLM reviewer that runs in CI on every pull request. It reads the diff, the description and any linked issue, may read other files in the checkout for context, may fetch documentation URLs, posts its review as a comment, and on request pushes small fixes. The repository is public and accepts pull requests from forks. The CI job has a token with write access to the repository in its environment.
Label it. The diff, the description and the issue are untrusted (A). The checkout may contain unreleased code and the runner's environment contains the token, and read_file can reach both (B). Posting a public comment, fetching an arbitrary URL and pushing a commit are all external effects (C). One session has all three letters, so the design fails the Rule of Two before any attack is written.
Now enumerate concrete paths, each one a sentence a reviewer can check. An instruction hidden in a code comment tells the model to read the environment file and include it in a URL it fetches: untrusted source, model, secret read, egress. The same instruction asks for the token to be placed in a markdown image in the review comment: the same path with a different sink. A description asks the model to push a commit that adds a workflow file: untrusted source, model, state change in a privileged location. Each points at an edge to remove.
Finding the paths in code
Doing this by hand is fine for one diagram and error-prone for twenty. The program below encodes the diagram as a graph, enumerates paths from untrusted sources to effects through a model session, enumerates paths from secrets and private reads to egress, and checks each session against the Rule of Two. Keep the graph in the repository next to the agent's tool configuration so that adding a tool shows up in review.
from collections import defaultdict
NODES = {
"pr_diff": {"kind": "source", "trust": "untrusted"},
"pr_description": {"kind": "source", "trust": "untrusted"},
"system_prompt": {"kind": "source", "trust": "trusted"},
"reviewer": {"kind": "model"},
"read_file": {"kind": "tool", "reads": "private"},
"ci_token": {"kind": "secret"},
"post_comment": {"kind": "tool", "effect": "egress"},
"http_fetch": {"kind": "tool", "effect": "egress"},
"push_commit": {"kind": "tool", "effect": "state"},
}
EDGES = [
("pr_diff", "reviewer"), ("pr_description", "reviewer"),
("system_prompt", "reviewer"), ("reviewer", "read_file"),
("read_file", "reviewer"), ("ci_token", "read_file"),
("reviewer", "post_comment"), ("reviewer", "http_fetch"),
("reviewer", "push_commit"),
]
def graph():
out = defaultdict(list)
for a, b in EDGES:
out[a].append(b)
return out
def paths(out, start, is_goal, max_len=6):
stack = [[start]]
while stack:
path = stack.pop()
for nxt in out[path[-1]]:
if nxt in path:
continue
new = path + [nxt]
if is_goal(nxt):
yield new
if len(new) < max_len:
stack.append(new)
def via_model(path):
return any(NODES[n]["kind"] == "model" for n in path)
def reaches(out, start, goal):
return any(True for _ in paths(out, start, lambda n: n == goal))
def analyse():
out, findings = graph(), []
for n, a in NODES.items():
if a.get("trust") == "untrusted":
for p in paths(out, n, lambda m: "effect" in NODES[m]):
if via_model(p):
findings.append(("injection->effect", p))
if a["kind"] == "secret" or a.get("reads") == "private":
for p in paths(out, n, lambda m: NODES[m].get("effect") == "egress"):
if via_model(p):
findings.append(("exfiltration", p))
for m, a in NODES.items():
if a["kind"] != "model":
continue
A = any(reaches(out, s, m) for s, x in NODES.items() if x.get("trust") == "untrusted")
B = any(reaches(out, s, m) for s, x in NODES.items()
if x["kind"] == "secret" or x.get("reads") == "private")
C = any("effect" in NODES[t] for t in out[m])
if A and B and C:
findings.append(("rule-of-two", [m]))
return findings
for kind, path in analyse():
print(f"{kind:18} {' -> '.join(path)}")On the original design it prints the rule-of-two violation for reviewer, every untrusted-to-effect path, and exfiltration paths such as ci_token -> read_file -> reviewer -> http_fetch. Edges you cannot draw honestly are themselves findings: if nobody knows whether the runner token is readable from the checkout, find out before the model does.
Ranking what you find
Rank paths by consequence and reachability rather than by guessed probability of the injection working. Assume it works; the question is what happens then. Consequence is set by the sink: a pushed workflow file or a leaked write token is critical, a misleading public comment is moderate, a wasted hour of compute is low. Reachability is set by who controls the source: anyone on the internet, any authenticated customer, or only staff. A critical sink reachable by anyone is a release blocker regardless of how good your classifier is.
Measured attack success rates help compare designs, and agent permission design shows how to scope each tool, but neither replaces removing the path.
Mitigations that change the graph
Sort mitigations into two kinds and prefer the first.
Structural controls change the graph. Split the session: a reader session sees the untrusted diff, has no tools, and must return findings in a fixed JSON schema; ordinary code validates the JSON and posts the comment from a template, escaping markdown and dropping images and links. That reader has A only. A second session with repository read access sees only the validated findings, never the raw diff, so it has B without A. Remove http_fetch or restrict it to an allowlist of documentation hosts. Move the write token out of the job that runs the model, and require a human to approve any push. Each change deletes edges, and you can rerun the program to prove it.
Probabilistic controls lower the odds. Injection classifiers, secret scanners on output, rate limits and careful system prompts all help. But an attacker who can open unlimited pull requests gets unlimited attempts, so these belong in front of structural controls, not instead of them.
Give exfiltration its own pass, as in LLM data exfiltration: list every way bytes can leave, including images, link previews and DNS lookups.
How LLM threat models go wrong
- Modelling the model as trusted. The diagram puts the LLM inside the trust boundary, so injection never appears as a crossing. Draw the model session on the boundary, taking the lowest trust of its inputs.
- Assessing tools one at a time. Each tool passes review alone; the combination is the vulnerability. Always assess per session.
- Forgetting indirect sources. Tool results, retrieved chunks, memory and file names in a listing are context sources. If a third party can write to it, label it untrusted.
- Treating the system prompt as a control. Rules in the prompt are input to the same model the attacker is steering. Count them as probabilistic at best, and assume the prompt itself can leak (LLM07).
- A model that ages silently. A new tool or memory feature changes the letters. Re-run the analysis whenever the tool manifest changes.
What to do next
- List every model session in your product, with its context sources and its tools. One agent with three modes is three sessions.
- Label each source trusted, internal or untrusted, and each tool by what it reads and what it does.
- Mark each session's Rule of Two letters. Treat every A-B-C session as a finding with an owner and a date.
- Encode the diagram as a graph like the one above and enumerate untrusted-to-effect and secret-to-egress paths.
- For each critical path, choose a structural control that deletes an edge, then rerun the analysis to show it is gone.
- Add probabilistic controls in front, and measure them, but do not close a finding with them alone.
- Keep the graph in version control and make CI fail when a session gains its third letter.