An attack tree puts an attacker's goal at the root and breaks it down into the ways to reach it. OR nodes are alternatives, AND nodes are steps that must all succeed, and the leaves are concrete actions you can test. The idea goes back to Bruce Schneier's 1999 article, and it survives because it answers a question lists of threats cannot: which combinations of steps actually reach the asset, and which controls cut the most of them.
LLM systems change the tree in two ways. First, the steps of an attack are often carried out by your own agent after it reads attacker-controlled text, so the order of steps matters: untrusted input, then a privileged read, then an outbound call. Second, many defenses for LLMs lower the odds of a step without ruling it out, such as prompt hardening or injection classifiers, and a tree has to tell those apart from controls that remove a step. This article builds on the AND/OR trees and cheapest-path analysis in threat modeling for LLM applications and adds three things: sequential AND nodes for injection chains, attack-defense trees with costed countermeasures and a computed minimum defense set, and a tree kept as code that fails CI when a new tool opens a path nobody has analysed.
The system under analysis
The running example is a coding agent that runs in a CI job. It reads issue text and repository files, fetches documentation from the web, calls tools through MCP servers, runs shell commands and posts pull request comments. To install private packages it needs a registry token, which the job injects as an environment variable. The asset is that token.
This shape, untrusted input plus access to private data plus a way to send data out, is the combination Simon Willison calls the lethal trifecta. An attack tree makes it concrete: each leg becomes a subtree, and each subtree lists the specific leaves for this system.
Building the tree
Build the tree top down from the goal, and stop decomposing when a leaf is something a tester can attempt in an afternoon. For the CI agent:
- Inject (OR): instructions in issue text (i1), in a dependency's README the agent reads (i2), in a web page returned by the fetch tool (i3), or in a poisoned MCP tool description (i4).
- Read (OR): print the environment through the shell tool (r1), or read a
.envfile the job writes (r2). - Send (OR): call the fetch tool with the token in a URL (x1), post it in a pull request comment (x2), or leak it through a DNS lookup from the shell (x3).
- Log leak (AND): the agent echoes the environment into CI output (l1), and the CI logs are readable by outsiders, as on many public repositories (l2).
The injection chain alone has 4 x 2 x 3 = 24 distinct paths, plus one for the log leak: 25 minimal attacks. Nobody enumerates 25 paths reliably by hand, which is the first argument for evaluating the tree in code. The mechanics of the inject leaves are covered in indirect prompt injection, and of the send leaves in data exfiltration via LLM tools.
Sequential AND: order is the signal
Classic AND says all children must succeed but says nothing about order. Jhawar and colleagues added the sequential AND, or SAND, in 2015 for attacks whose steps must happen in sequence. In an LLM agent the order is not a detail: the read and the send only become attacker actions because they happen after the model has ingested attacker text in the same context. The same shell call that prints the environment is a routine debugging step when the agent decides it alone.
This has two practical consequences. For detection, a SAND node says what to watch for: a privileged read followed by an outbound call, in a session that has ingested untrusted content. That is a taint rule, and it is far more specific than flagging every shell command. For testing, a SAND path is a script with steps in a fixed order, which is what a red-team harness needs to replay it. For computing which defenses cut which attacks, however, SAND behaves exactly like AND: a path is blocked if any of its steps is blocked.
Evaluating the tree in code
The evaluator below expands a tree into its minimal attack sets, the sets of leaves that reach the root, keeping SAND order for test generation. It then finds the cheapest set of countermeasures that blocks every attack at least depth times, by brute force over subsets, which is fine for the dozen controls a real tree has.
from itertools import combinations, product
def attack_sets(node):
"""Return a list of attacks; each attack is a tuple of leaves in SAND order."""
kind = node[0]
if kind == "leaf":
return [(node[1],)]
child_sets = [attack_sets(ch) for ch in node[1:]]
if kind == "or":
return [a for sets in child_sets for a in sets]
# "and" and "sand" both need one attack from every child; sand keeps the order
return [sum(parts, ()) for parts in product(*child_sets)]
TREE = ("or",
("sand",
("or", ("leaf", "i1"), ("leaf", "i2"), ("leaf", "i3"), ("leaf", "i4")),
("or", ("leaf", "r1"), ("leaf", "r2")),
("or", ("leaf", "x1"), ("leaf", "x2"), ("leaf", "x3"))),
("and", ("leaf", "l1"), ("leaf", "l2")))
CONTROLS = { # name: (cost, leaves it blocks)
"D1 credential broker": (5, {"r1", "r2", "l1"}),
"D2 fetch egress allowlist": (1, {"x1"}),
"D3 no network in shell": (2, {"x3"}),
"D4 human approves posts": (2, {"x2"}),
"D5 private CI logs": (1, {"l2"}),
"D6 pinned, reviewed MCP": (1, {"i4"}),
}
def cheapest_defense(attacks, controls, depth=1):
best = None
names = list(controls)
for r in range(len(names) + 1):
for chosen in combinations(names, r):
ok = all(sum(1 for d in chosen if controls[d][1] & set(a)) >= depth
for a in attacks)
cost = sum(controls[d][0] for d in chosen)
if ok and (best is None or cost < best[0]):
best = (cost, chosen)
return best
attacks = attack_sets(TREE)
print(len(attacks)) # 25
print(cheapest_defense(attacks, CONTROLS, 1)) # (5, ('D1 credential broker',))
print(cheapest_defense(attacks, CONTROLS, 2)) # cost 11: D1 to D5A control counts once per attack no matter how many of that attack's leaves it blocks, which is what the & test inside the sum expresses. That matters for defense in depth: D1 blocks both r1 and l1, but it is still one control, and one bug in it would reopen every path it guards.
Choosing defenses
The output is the point of the exercise. With depth 1, the cheapest complete defense is the credential broker alone, at cost 5: a proxy that adds the token to registry requests so the agent never holds it. Every injection path needs a read step, and the broker removes both read leaves and the echo leaf. The alternative, blocking every send channel and the log leak with D2, D3, D4 and D5, costs 6 and leaves the token sitting in the environment for the next channel nobody listed.
With depth 2, where every attack must be cut by two independent controls, the answer jumps to cost 11 and needs D1 through D5. The reason is visible in the tree: nothing in this control set blocks the first three inject leaves, so each injection path's second cut must come from its send step, which forces all three send controls. D6 is not required at either depth, yet it is cheap and stops a supply-chain route to every injection path; the optimizer will not tell you to buy it, and that is a signal to model the MCP server's own compromise as a separate branch rather than a judgement that it is useless.
| Required cuts per attack | Cheapest control set | Cost | What it says |
|---|---|---|---|
| 1 | D1 | 5 | Remove the asset from the agent's reach |
| 1 (without D1) | D2, D3, D4, D5 | 6 | Block every listed exit; unlisted exits stay open |
| 2 | D1, D2, D3, D4, D5 | 11 | Defense in depth needs all exits as well |
Controls that only reduce likelihood, such as spotlighting untrusted content or an injection classifier in front of the model, do not belong in the blocking set. Record them as annotations on the inject leaves with a measured bypass rate from your red-team runs, and never let them satisfy a depth requirement. Treating a classifier as a cut is the most common way an attack-defense tree overstates protection.
Keeping the tree as code
A tree drawn once in a workshop is out of date the week a new tool ships. Keep it as a file in the agent's repository, next to the tool manifest, and make CI check the two against each other. Each leaf names the tools that enable it, so a tool that appears in the manifest without appearing in any leaf, or in any explicit allowlist of reviewed harmless tools, fails the build.
# attack_tree.yaml (excerpt)
goal: exfiltrate registry token
controls:
D2: {cost: 1, blocks: [x1], owner: platform-sec, test: tests/redteam/test_x1.py}
leaves:
i3: {desc: instructions in fetched page, tools: [fetch], kind: inject}
r1: {desc: print env via shell, tools: [shell], kind: read}
x1: {desc: token in fetch URL, tools: [fetch], kind: send}
x2: {desc: token in PR comment, tools: [pr_comment], kind: send}
reviewed_harmless: [list_files]# ci_check_tree.py
import sys, yaml, json
tree = yaml.safe_load(open("attack_tree.yaml"))
manifest = json.load(open("agent_tools.json")) # tools the agent can call
covered = {t for leaf in tree["leaves"].values() for t in leaf["tools"]}
covered |= set(tree.get("reviewed_harmless", []))
missing = sorted({t["name"] for t in manifest["tools"]} - covered)
if missing:
sys.exit(f"tools not analysed in attack_tree.yaml: {missing}")
for name, ctl in tree["controls"].items():
if not ctl.get("test"):
sys.exit(f"control {name} has no regression test")The second check ties every countermeasure to a test that tries the leaf it claims to block. The tree, the tests and the minimum-defense computation then run on every change, and a pull request that adds an email tool cannot merge until someone has decided whether it is a new send leaf.
From leaves to tests and monitors
Each minimal attack is a test case, and SAND order is its script: plant the inject payload, run the agent on a task that reads the poisoned source, and check whether a canary token placed where the real one would be appears at any sink. Use a canary, never the real secret, and check sinks deterministically: the fetch tool's outbound URL log, the PR comment body, and DNS queries from the sandbox. Run every path at least a few dozen times, because model behaviour is stochastic and a path that succeeds 3 percent of the time is open. A playbook for running these as structured exercises is in the red-teaming playbook.
The same structure gives monitors. Log, per agent session, whether untrusted content was ingested, which privileged reads happened and which outbound calls were made, in order. An alert on any session that matches a SAND path, ingest then read then send, catches attacks through leaves you did not think to list, as long as the step categories are right.
Failure modes
- Abstract leaves. A leaf like "prompt injection" cannot be tested or blocked. Split until each leaf names a source, a tool or a sink.
- Probabilistic controls counted as cuts. Classifiers and system-prompt rules lower odds; the tree then reports full coverage that red-teaming will disprove.
- Missing order. Modelling the injection chain as plain AND hides the fact that the read is only dangerous after ingestion, and loses the taint rule.
- Stale trees. New tools, connectors and MCP servers add leaves weekly. Without the CI check, the tree describes last quarter's agent.
- Invented numbers. Multiplying guessed leaf probabilities gives false precision, and the leaves of an LLM attack are rarely independent, since one successful injection makes every later step more likely. Rank with costs and measured success rates instead.
- Single-control dependence. A depth-1 answer like the broker is right and fragile; check what one bug in it reopens.
Trade-offs
| Choice | Gains | Costs |
|---|---|---|
| Fine-grained leaves | Testable, maps to controls | Bigger tree to maintain |
| SAND instead of AND | Order gives tests and taint rules | More modelling effort |
| Cost-based optimisation | Shows the cheapest complete defense | Costs are estimates; results shift with them |
| Depth-2 requirement | Survives one failed control | Roughly doubles control spend here |
| Tree as code in CI | Stays current with tools | Friction on every tool change |
For a whole-system view that attack trees do not give, pair them with a category sweep such as STRIDE for LLM systems: STRIDE finds the goals, and a tree per goal finds the paths.
What to do next
- Pick one asset your agent can reach and write its goal as the root of a tree.
- Decompose into inject, read and send subtrees, with leaves that each name a source, a tool or a sink.
- Run the evaluator on your tree and list every minimal attack.
- Attach controls with costs, compute the cheapest defense at depth 1 and depth 2, and compare it with what you have deployed.
- Move the tree into the repository as YAML and add the manifest check to CI.
- Turn each attack into a canary test and run it repeatedly; add a session monitor for the SAND pattern.