Most LLM security writing is about one attack or one control: prompt injection, a jailbreak family, an output filter, a sandbox. Each piece is useful, and none of them is a program. A program is what makes the pieces add up: a list of every system that uses a model, a reason each control exists, evidence that it works, and a way to notice when it stops working. Without that, teams ship a guardrail on the chatbot and leave the coding agent with production credentials untouched, because nobody owned the question of where the risk actually sat.
This article is the program view of the whole series. It lays out the eight layers a request crosses, shows how to compute control coverage from code instead of slideware, walks through six stages in the order they depend on each other with the deep-dive articles for each, and works an example of a company with three LLM systems of very different risk. It ends with the metrics, the failure modes and a checklist you can start on this week.
Four principles the program rests on
Four ideas carry most of the weight, and every later section applies one of them.
- The model is not a security boundary. Anything in the context window can influence the output, and no system prompt, delimiter or classifier makes that influence zero. Controls that matter are therefore placed outside the model: on what enters the context, on what authority the model's output can exercise, and on what leaves.
- Risk follows authority and data, not model size. A small model with a shell tool and customer records is more dangerous than a frontier model that answers FAQ questions from public docs. Tier systems by what they can read, what they can do and who sees the output.
- Taint is the unit of analysis. Follow untrusted text from where it enters (user input, retrieved documents, tool results, emails) to every sink it can reach (tool calls, rendered markdown, outbound requests). The dangerous combination is untrusted input, access to private data and an exfiltration channel in the same session.
- Evidence beats assertion. A control is real when a test exercises it and a log shows it firing in production. Everything else is a plan.
The eight layers and who owns them
The diagram below is the reference architecture the program is organised around. Each layer has a small set of control families, and each family has an owner.
| Layer | Typical controls | Usual owner |
|---|---|---|
| Identity | end-user auth propagated to tools, per-tenant scoping, service identities for agents | platform |
| Input | size and rate limits, injection classifiers as signals, file type checks | app team |
| Context assembly | retrieval ACLs enforced at query time, provenance labels, memory write rules | app team + data |
| Model | pinned versions, system prompt hygiene, no secrets in prompts | app team |
| Tools and actions | least-privilege scopes, capability tokens, confirmation for irreversible actions, sandboxes | app team + security |
| Output and egress | link and image rewriting, secret and canary scanners, egress allowlists | platform |
| Supply chain | signed models, safe formats, pinned packages, vetted MCP servers | platform |
| Observability and response | span-level traces, audit records, detections, kill switch, playbooks | security |
The owner column matters as much as the controls. Platform-owned layers should be built once and inherited, so an application team cannot forget them; application-owned layers need a review gate because they depend on what the specific system does.
Computing coverage instead of claiming it
Coverage claims are where programs usually lie to themselves. The fix is to keep the threat list, the system inventory and the control catalogue as data in one repository and compute the gaps. A minimal version fits in a page of Python:
# inventory.yaml : systems with tier and capabilities
# threats.yaml : threat id -> layers where it can be stopped, applicability rule
# controls.yaml : control id -> threat ids it mitigates, layer, test id, systems
import yaml
inv = yaml.safe_load(open("inventory.yaml"))
threats = yaml.safe_load(open("threats.yaml"))
controls = yaml.safe_load(open("controls.yaml"))
tests = yaml.safe_load(open("test_results.yaml")) # exported by CI
def applies(threat, system):
rule = threat.get("when", {})
return all(c in system["capabilities"] for c in rule.get("needs", []))
gaps = []
for sys_ in inv["systems"]:
for tid, t in threats.items():
if not applies(t, sys_):
continue
live = [c for c in controls.values()
if tid in c["mitigates"] and sys_["id"] in c["systems"]
and tests.get(c["test"], {}).get("status") == "pass"]
needed = 2 if sys_["tier"] == "high" else 1
if len(live) < needed:
gaps.append((sys_["id"], tid, len(live), needed))
for g in sorted(gaps):
print("GAP system=%s threat=%s tested_controls=%d required=%d" % g)Three details make this honest. A control only counts if its test passed in the latest run, so a broken test becomes a visible gap rather than a silent one. High-tier systems need two independent controls per applicable threat, which encodes defence in depth as a rule rather than a hope. And applicability is computed from capabilities such as reads_untrusted_web, has_write_tools or sees_pii, so adding a tool to a system automatically adds the threats that tool brings. Run it in CI and fail the build of the inventory repository when a high-tier gap appears without an approved exception.
The six stages, with where to learn each
The program is built in six stages. Order matters because each stage consumes the previous stage's artefact.
Stage 1: foundations and threat modelling. Build the inventory (every system, its model, data, tools and audience) and tier it. Then model each high-tier system by drawing its data flows and following taint. Start with LLM-specific threat modelling, use data flow diagrams for LLM apps for the drawing, and check coverage against the OWASP LLM Top 10 so you do not miss a category.
Stage 2: build controls, layer by layer. Injection cannot be filtered away, so read indirect prompt injection in depth to understand why the answer is limiting blast radius. Bound tool authority with capability tokens, close exfiltration channels with egress control for agents, and route every model, dataset and package through one gate as described in the AI supply chain security program.
Stage 3: verify. Controls are hypotheses until attacked. Stand up an AI red team program for depth and automated red teaming for breadth, and wire the results into release gates so a regression blocks a deploy.
Stage 4: operate. You cannot respond to what you did not record. Design audit logging for LLM apps so a decision can be reconstructed, monitor drift and abuse with AI risk monitoring, and rehearse LLM incident response including model rollback and tool revocation.
Stage 5: govern. Decide who accepts residual risk and on what evidence. The operating model is in AI governance program structure, and the NIST AI RMF gives a shared vocabulary for mapping, measuring and managing risk that auditors and regulators recognise.
Stage 6 is the loop back: incidents, red team findings and new capabilities change the inventory and the threat model, and the cycle runs again. A program that does not re-tier systems when they gain tools is a snapshot, not a program.
Worked example: three systems, three different tiers
Consider a mid-sized software company with three LLM systems. A public support assistant answers questions from product documentation and can look up the signed-in customer's orders. An internal knowledge assistant searches wikis, tickets and shared drives for all employees. A coding agent opens pull requests against internal repositories and can run tests in CI.
| System | Reads untrusted text | Private data | Authority | Tier |
|---|---|---|---|---|
| Support assistant | user chat | one customer's orders | read-only lookup | medium |
| Knowledge assistant | any wiki page or ticket | everything employees can see | none, but renders links | high |
| Coding agent | issues, dependencies, web docs | source code, CI secrets | write PRs, run code | high |
The knowledge assistant surprises people by being high tier. It has no tools, but any employee or customer who can write a ticket can plant instructions, and the assistant renders markdown, so a planted image link can carry retrieved secrets to an attacker's server. That is the full dangerous combination without a single tool. The coverage script flags it with two gaps: no retrieval ACL enforcement (the index was built with a service account that sees everything) and no output link rewriting. The fixes are platform controls: per-user document filtering at query time and an egress rewriter that strips or proxies external URLs.
The coding agent's gaps are authority gaps. It ran with a token able to push to the main branch and with CI secrets in its environment. The program moves it to a sandbox with no production credentials, scopes its token to opening pull requests on a branch prefix, and requires human review before merge. The support assistant needs the least: its lookup tool already takes the customer identity from the session rather than from the model, which is exactly the property that prevents one customer reading another's orders.
Notice what was not on the list: a better jailbreak filter. The highest-value work in this example was moving authority and data boundaries, which is typical.
Metrics that show the program works
Report a short set of numbers monthly, each with a definition precise enough to compute from the repository and logs.
- Inventory completeness: systems found by discovery (API gateway logs, expense reports for AI vendors, egress to model APIs) that are missing from the inventory. The target is zero, and the first measurement is usually humbling.
- High-tier gap count: output of the coverage script, with the age of the oldest open gap.
- Control test pass rate: share of control tests passing on the latest run.
- Red team findings by layer: where attacks succeed tells you which layer is under-built.
- Time to revoke: from decision to a tool credential or model version being disabled in production, measured in drills.
- Exceptions past expiry: accepted risks whose review date has passed.
Failure modes and trade-offs
Programs fail in recognisable ways. Filter-first thinking spends the budget on input classifiers that attackers paraphrase around, while tools keep broad credentials. Inventory rot happens when the inventory is a spreadsheet updated at audit time; new agents appear weekly and nobody records them. Application-by-application controls mean each team builds its own output filter, half of them wrong; shared layers belong in the platform. Untested controls pass every review and fail the first real attack because nobody checked that the allowlist actually loaded. Governance without teeth produces risk acceptances signed by people who do not own the budget to fix the risk, with no expiry date.
The trade-offs are real too. Strict confirmation prompts on every action make agents useless, so reserve them for irreversible or external effects. Two independent controls per threat doubles engineering cost, so apply that rule only to high-tier systems. Centralised platform controls slow teams that need something unusual; offer an exception path with an expiry instead of forcing everything through the paved road. And every log you keep for forensics is also data you must protect and eventually delete.