An attack tree puts an attacker's goal at the root and breaks it down into the ways to reach it. OR nodes are alternatives, AND nodes are steps that must all succeed, and the leaves are concrete actions you can test. The idea goes back to Bruce Schneier's 1999 article, and it survives because it answers a question lists of threats cannot: which combinations of steps actually reach the asset, and which controls cut the most of them.

LLM systems change the tree in two ways. First, the steps of an attack are often carried out by your own agent after it reads attacker-controlled text, so the order of steps matters: untrusted input, then a privileged read, then an outbound call. Second, many defenses for LLMs lower the odds of a step without ruling it out, such as prompt hardening or injection classifiers, and a tree has to tell those apart from controls that remove a step. This article builds on the AND/OR trees and cheapest-path analysis in threat modeling for LLM applications and adds three things: sequential AND nodes for injection chains, attack-defense trees with costed countermeasures and a computed minimum defense set, and a tree kept as code that fails CI when a new tool opens a path nobody has analysed.

The system under analysis

The running example is a coding agent that runs in a CI job. It reads issue text and repository files, fetches documentation from the web, calls tools through MCP servers, runs shell commands and posts pull request comments. To install private packages it needs a registry token, which the job injects as an environment variable. The asset is that token.

Worked system: a coding agent that runs in CIUntrusted textissues, READMEs, webreadAgent loop (LLM)plans, calls toolscallToolsshell, file read, fetch, PR commentRegistry tokenneeded to install packagesin envMCP serversthird-party tool descriptionsCI logsstdout of every stepechoAttacker goal: get the registry token out of CI. Every arrow is a possible step in an attack path.
The system under analysis. Untrusted text reaches the model; the model can reach a secret and several outbound channels.

This shape, untrusted input plus access to private data plus a way to send data out, is the combination Simon Willison calls the lethal trifecta. An attack tree makes it concrete: each leg becomes a subtree, and each subtree lists the specific leaves for this system.

Building the tree

Build the tree top down from the goal, and stop decomposing when a leaf is something a tester can attempt in an afternoon. For the CI agent:

Attack-defense tree for the CI agent (simplified)Exfiltrate tokenORInjection chainSAND: inject, then read, then sendLog leakANDInject (OR)issue, README, page, tool descRead (OR)env, .env fileSend (OR)fetch, PR, DNSEcho to logagent prints envLogs publicreadable by anyoneDefenses: D1 broker blocks Read; D2, D3, D4 block SendD6 blocks the tool-description inject leafD1 blocks Echo; D5 blocks Logs publicGreen nodes are countermeasures attached to the leaves they block.
The root has two alternatives. The injection chain is a sequential AND of three OR subtrees; the log leak is a plain AND.
  • Inject (OR): instructions in issue text (i1), in a dependency's README the agent reads (i2), in a web page returned by the fetch tool (i3), or in a poisoned MCP tool description (i4).
  • Read (OR): print the environment through the shell tool (r1), or read a .env file the job writes (r2).
  • Send (OR): call the fetch tool with the token in a URL (x1), post it in a pull request comment (x2), or leak it through a DNS lookup from the shell (x3).
  • Log leak (AND): the agent echoes the environment into CI output (l1), and the CI logs are readable by outsiders, as on many public repositories (l2).

The injection chain alone has 4 x 2 x 3 = 24 distinct paths, plus one for the log leak: 25 minimal attacks. Nobody enumerates 25 paths reliably by hand, which is the first argument for evaluating the tree in code. The mechanics of the inject leaves are covered in indirect prompt injection, and of the send leaves in data exfiltration via LLM tools.

Sequential AND: order is the signal

Classic AND says all children must succeed but says nothing about order. Jhawar and colleagues added the sequential AND, or SAND, in 2015 for attacks whose steps must happen in sequence. In an LLM agent the order is not a detail: the read and the send only become attacker actions because they happen after the model has ingested attacker text in the same context. The same shell call that prints the environment is a routine debugging step when the agent decides it alone.

This has two practical consequences. For detection, a SAND node says what to watch for: a privileged read followed by an outbound call, in a session that has ingested untrusted content. That is a taint rule, and it is far more specific than flagging every shell command. For testing, a SAND path is a script with steps in a fixed order, which is what a red-team harness needs to replay it. For computing which defenses cut which attacks, however, SAND behaves exactly like AND: a path is blocked if any of its steps is blocked.

Evaluating the tree in code

The evaluator below expands a tree into its minimal attack sets, the sets of leaves that reach the root, keeping SAND order for test generation. It then finds the cheapest set of countermeasures that blocks every attack at least depth times, by brute force over subsets, which is fine for the dozen controls a real tree has.

from itertools import combinations, product

def attack_sets(node):
    """Return a list of attacks; each attack is a tuple of leaves in SAND order."""
    kind = node[0]
    if kind == "leaf":
        return [(node[1],)]
    child_sets = [attack_sets(ch) for ch in node[1:]]
    if kind == "or":
        return [a for sets in child_sets for a in sets]
    # "and" and "sand" both need one attack from every child; sand keeps the order
    return [sum(parts, ()) for parts in product(*child_sets)]

TREE = ("or",
    ("sand",
        ("or", ("leaf", "i1"), ("leaf", "i2"), ("leaf", "i3"), ("leaf", "i4")),
        ("or", ("leaf", "r1"), ("leaf", "r2")),
        ("or", ("leaf", "x1"), ("leaf", "x2"), ("leaf", "x3"))),
    ("and", ("leaf", "l1"), ("leaf", "l2")))

CONTROLS = {  # name: (cost, leaves it blocks)
    "D1 credential broker":       (5, {"r1", "r2", "l1"}),
    "D2 fetch egress allowlist":  (1, {"x1"}),
    "D3 no network in shell":     (2, {"x3"}),
    "D4 human approves posts":    (2, {"x2"}),
    "D5 private CI logs":         (1, {"l2"}),
    "D6 pinned, reviewed MCP":    (1, {"i4"}),
}

def cheapest_defense(attacks, controls, depth=1):
    best = None
    names = list(controls)
    for r in range(len(names) + 1):
        for chosen in combinations(names, r):
            ok = all(sum(1 for d in chosen if controls[d][1] & set(a)) >= depth
                     for a in attacks)
            cost = sum(controls[d][0] for d in chosen)
            if ok and (best is None or cost < best[0]):
                best = (cost, chosen)
    return best

attacks = attack_sets(TREE)
print(len(attacks))                                # 25
print(cheapest_defense(attacks, CONTROLS, 1))      # (5, ('D1 credential broker',))
print(cheapest_defense(attacks, CONTROLS, 2))      # cost 11: D1 to D5

A control counts once per attack no matter how many of that attack's leaves it blocks, which is what the & test inside the sum expresses. That matters for defense in depth: D1 blocks both r1 and l1, but it is still one control, and one bug in it would reopen every path it guards.

Choosing defenses

The output is the point of the exercise. With depth 1, the cheapest complete defense is the credential broker alone, at cost 5: a proxy that adds the token to registry requests so the agent never holds it. Every injection path needs a read step, and the broker removes both read leaves and the echo leaf. The alternative, blocking every send channel and the log leak with D2, D3, D4 and D5, costs 6 and leaves the token sitting in the environment for the next channel nobody listed.

With depth 2, where every attack must be cut by two independent controls, the answer jumps to cost 11 and needs D1 through D5. The reason is visible in the tree: nothing in this control set blocks the first three inject leaves, so each injection path's second cut must come from its send step, which forces all three send controls. D6 is not required at either depth, yet it is cheap and stops a supply-chain route to every injection path; the optimizer will not tell you to buy it, and that is a signal to model the MCP server's own compromise as a separate branch rather than a judgement that it is useless.

Required cuts per attackCheapest control setCostWhat it says
1D15Remove the asset from the agent's reach
1 (without D1)D2, D3, D4, D56Block every listed exit; unlisted exits stay open
2D1, D2, D3, D4, D511Defense in depth needs all exits as well

Controls that only reduce likelihood, such as spotlighting untrusted content or an injection classifier in front of the model, do not belong in the blocking set. Record them as annotations on the inject leaves with a measured bypass rate from your red-team runs, and never let them satisfy a depth requirement. Treating a classifier as a cut is the most common way an attack-defense tree overstates protection.

Keeping the tree as code

A tree drawn once in a workshop is out of date the week a new tool ships. Keep it as a file in the agent's repository, next to the tool manifest, and make CI check the two against each other. Each leaf names the tools that enable it, so a tool that appears in the manifest without appearing in any leaf, or in any explicit allowlist of reviewed harmless tools, fails the build.

# attack_tree.yaml (excerpt)
goal: exfiltrate registry token
controls:
  D2: {cost: 1, blocks: [x1], owner: platform-sec, test: tests/redteam/test_x1.py}
leaves:
  i3: {desc: instructions in fetched page, tools: [fetch], kind: inject}
  r1: {desc: print env via shell, tools: [shell], kind: read}
  x1: {desc: token in fetch URL, tools: [fetch], kind: send}
  x2: {desc: token in PR comment, tools: [pr_comment], kind: send}
reviewed_harmless: [list_files]
# ci_check_tree.py
import sys, yaml, json

tree = yaml.safe_load(open("attack_tree.yaml"))
manifest = json.load(open("agent_tools.json"))          # tools the agent can call
covered = {t for leaf in tree["leaves"].values() for t in leaf["tools"]}
covered |= set(tree.get("reviewed_harmless", []))
missing = sorted({t["name"] for t in manifest["tools"]} - covered)
if missing:
    sys.exit(f"tools not analysed in attack_tree.yaml: {missing}")
for name, ctl in tree["controls"].items():
    if not ctl.get("test"):
        sys.exit(f"control {name} has no regression test")

The second check ties every countermeasure to a test that tries the leaf it claims to block. The tree, the tests and the minimum-defense computation then run on every change, and a pull request that adds an email tool cannot merge until someone has decided whether it is a new send leaf.

From leaves to tests and monitors

Each minimal attack is a test case, and SAND order is its script: plant the inject payload, run the agent on a task that reads the poisoned source, and check whether a canary token placed where the real one would be appears at any sink. Use a canary, never the real secret, and check sinks deterministically: the fetch tool's outbound URL log, the PR comment body, and DNS queries from the sandbox. Run every path at least a few dozen times, because model behaviour is stochastic and a path that succeeds 3 percent of the time is open. A playbook for running these as structured exercises is in the red-teaming playbook.

The same structure gives monitors. Log, per agent session, whether untrusted content was ingested, which privileged reads happened and which outbound calls were made, in order. An alert on any session that matches a SAND path, ingest then read then send, catches attacks through leaves you did not think to list, as long as the step categories are right.

Failure modes

  • Abstract leaves. A leaf like "prompt injection" cannot be tested or blocked. Split until each leaf names a source, a tool or a sink.
  • Probabilistic controls counted as cuts. Classifiers and system-prompt rules lower odds; the tree then reports full coverage that red-teaming will disprove.
  • Missing order. Modelling the injection chain as plain AND hides the fact that the read is only dangerous after ingestion, and loses the taint rule.
  • Stale trees. New tools, connectors and MCP servers add leaves weekly. Without the CI check, the tree describes last quarter's agent.
  • Invented numbers. Multiplying guessed leaf probabilities gives false precision, and the leaves of an LLM attack are rarely independent, since one successful injection makes every later step more likely. Rank with costs and measured success rates instead.
  • Single-control dependence. A depth-1 answer like the broker is right and fragile; check what one bug in it reopens.

Trade-offs

ChoiceGainsCosts
Fine-grained leavesTestable, maps to controlsBigger tree to maintain
SAND instead of ANDOrder gives tests and taint rulesMore modelling effort
Cost-based optimisationShows the cheapest complete defenseCosts are estimates; results shift with them
Depth-2 requirementSurvives one failed controlRoughly doubles control spend here
Tree as code in CIStays current with toolsFriction on every tool change

For a whole-system view that attack trees do not give, pair them with a category sweep such as STRIDE for LLM systems: STRIDE finds the goals, and a tree per goal finds the paths.

What to do next

  1. Pick one asset your agent can reach and write its goal as the root of a tree.
  2. Decompose into inject, read and send subtrees, with leaves that each name a source, a tool or a sink.
  3. Run the evaluator on your tree and list every minimal attack.
  4. Attach controls with costs, compute the cheapest defense at depth 1 and depth 2, and compare it with what you have deployed.
  5. Move the tree into the repository as YAML and add the manifest check to CI.
  6. Turn each attack into a canary test and run it repeatedly; add a session monitor for the SAND pattern.
Key takeaway: An attack tree for an LLM system turns a vague worry about prompt injection into a finite list of paths: inject, then read, then send, plus whatever side channels the system has. Model the order with sequential AND, attach countermeasures with costs, and compute the cheapest set that cuts every path once and twice. Count only controls that remove a step, keep the tree in the repository where CI checks it against the tool manifest, and turn every path into a repeated canary test.