Most LLM security writing is about one attack or one control: prompt injection, a jailbreak family, an output filter, a sandbox. Each piece is useful, and none of them is a program. A program is what makes the pieces add up: a list of every system that uses a model, a reason each control exists, evidence that it works, and a way to notice when it stops working. Without that, teams ship a guardrail on the chatbot and leave the coding agent with production credentials untouched, because nobody owned the question of where the risk actually sat.

This article is the program view of the whole series. It lays out the eight layers a request crosses, shows how to compute control coverage from code instead of slideware, walks through six stages in the order they depend on each other with the deep-dive articles for each, and works an example of a company with three LLM systems of very different risk. It ends with the metrics, the failure modes and a checklist you can start on this week.

Four principles the program rests on

Four ideas carry most of the weight, and every later section applies one of them.

  1. The model is not a security boundary. Anything in the context window can influence the output, and no system prompt, delimiter or classifier makes that influence zero. Controls that matter are therefore placed outside the model: on what enters the context, on what authority the model's output can exercise, and on what leaves.
  2. Risk follows authority and data, not model size. A small model with a shell tool and customer records is more dangerous than a frontier model that answers FAQ questions from public docs. Tier systems by what they can read, what they can do and who sees the output.
  3. Taint is the unit of analysis. Follow untrusted text from where it enters (user input, retrieved documents, tool results, emails) to every sink it can reach (tool calls, rendered markdown, outbound requests). The dangerous combination is untrusted input, access to private data and an exfiltration channel in the same session.
  4. Evidence beats assertion. A control is real when a test exercises it and a log shows it firing in production. Everything else is a plan.

The eight layers and who owns them

The diagram below is the reference architecture the program is organised around. Each layer has a small set of control families, and each family has an owner.

One request crosses eight control layers; the program owns all of themidentitywho is askinginputuntrusted text incontext assemblyRAG, memory, filesmodelweights, promptstools / actionsauthority outoutput / egresswhat leavessupply chainmodels, datasets, packages, MCP serversobservability + responsetraces, audit, detection, kill switchfeedsrecordsYellow layers carry attacker-controlled text; red layers are where damage happens.Green layers sit underneath every request and are owned by the platform, not each app team.A control placed in the wrong layer looks like coverage and stops nothing.
The eight layers of an LLM request. Supply chain and observability are shared platform layers underneath every application.
LayerTypical controlsUsual owner
Identityend-user auth propagated to tools, per-tenant scoping, service identities for agentsplatform
Inputsize and rate limits, injection classifiers as signals, file type checksapp team
Context assemblyretrieval ACLs enforced at query time, provenance labels, memory write rulesapp team + data
Modelpinned versions, system prompt hygiene, no secrets in promptsapp team
Tools and actionsleast-privilege scopes, capability tokens, confirmation for irreversible actions, sandboxesapp team + security
Output and egresslink and image rewriting, secret and canary scanners, egress allowlistsplatform
Supply chainsigned models, safe formats, pinned packages, vetted MCP serversplatform
Observability and responsespan-level traces, audit records, detections, kill switch, playbookssecurity

The owner column matters as much as the controls. Platform-owned layers should be built once and inherited, so an application team cannot forget them; application-owned layers need a review gate because they depend on what the specific system does.

Computing coverage instead of claiming it

Coverage claims are where programs usually lie to themselves. The fix is to keep the threat list, the system inventory and the control catalogue as data in one repository and compute the gaps. A minimal version fits in a page of Python:

# inventory.yaml : systems with tier and capabilities
# threats.yaml   : threat id -> layers where it can be stopped, applicability rule
# controls.yaml  : control id -> threat ids it mitigates, layer, test id, systems
import yaml

inv = yaml.safe_load(open("inventory.yaml"))
threats = yaml.safe_load(open("threats.yaml"))
controls = yaml.safe_load(open("controls.yaml"))
tests = yaml.safe_load(open("test_results.yaml"))   # exported by CI

def applies(threat, system):
    rule = threat.get("when", {})
    return all(c in system["capabilities"] for c in rule.get("needs", []))

gaps = []
for sys_ in inv["systems"]:
    for tid, t in threats.items():
        if not applies(t, sys_):
            continue
        live = [c for c in controls.values()
                if tid in c["mitigates"] and sys_["id"] in c["systems"]
                and tests.get(c["test"], {}).get("status") == "pass"]
        needed = 2 if sys_["tier"] == "high" else 1
        if len(live) < needed:
            gaps.append((sys_["id"], tid, len(live), needed))

for g in sorted(gaps):
    print("GAP system=%s threat=%s tested_controls=%d required=%d" % g)

Three details make this honest. A control only counts if its test passed in the latest run, so a broken test becomes a visible gap rather than a silent one. High-tier systems need two independent controls per applicable threat, which encodes defence in depth as a rule rather than a hope. And applicability is computed from capabilities such as reads_untrusted_web, has_write_tools or sees_pii, so adding a tool to a system automatically adds the threats that tool brings. Run it in CI and fail the build of the inventory repository when a high-tier gap appears without an approved exception.

The six stages, with where to learn each

The program is built in six stages. Order matters because each stage consumes the previous stage's artefact.

Stages build on each other: skipping one leaves the next without inputs1 inventorysystems, tiers2 model threatstaint paths3 build controlsper layer4 verifyred team, evals5 operatedetect, respond6 govern and improvemetrics, exceptions, reviewre-tier, re-modelEach stage produces an artefact the next consumes: inventory rows, threat entries,control IDs, test results, incidents, and finally decisions that change the inventory.
Six stages of an LLM security program and the artefact each hands to the next.

Stage 1: foundations and threat modelling. Build the inventory (every system, its model, data, tools and audience) and tier it. Then model each high-tier system by drawing its data flows and following taint. Start with LLM-specific threat modelling, use data flow diagrams for LLM apps for the drawing, and check coverage against the OWASP LLM Top 10 so you do not miss a category.

Stage 2: build controls, layer by layer. Injection cannot be filtered away, so read indirect prompt injection in depth to understand why the answer is limiting blast radius. Bound tool authority with capability tokens, close exfiltration channels with egress control for agents, and route every model, dataset and package through one gate as described in the AI supply chain security program.

Stage 3: verify. Controls are hypotheses until attacked. Stand up an AI red team program for depth and automated red teaming for breadth, and wire the results into release gates so a regression blocks a deploy.

Stage 4: operate. You cannot respond to what you did not record. Design audit logging for LLM apps so a decision can be reconstructed, monitor drift and abuse with AI risk monitoring, and rehearse LLM incident response including model rollback and tool revocation.

Stage 5: govern. Decide who accepts residual risk and on what evidence. The operating model is in AI governance program structure, and the NIST AI RMF gives a shared vocabulary for mapping, measuring and managing risk that auditors and regulators recognise.

Stage 6 is the loop back: incidents, red team findings and new capabilities change the inventory and the threat model, and the cycle runs again. A program that does not re-tier systems when they gain tools is a snapshot, not a program.

Worked example: three systems, three different tiers

Consider a mid-sized software company with three LLM systems. A public support assistant answers questions from product documentation and can look up the signed-in customer's orders. An internal knowledge assistant searches wikis, tickets and shared drives for all employees. A coding agent opens pull requests against internal repositories and can run tests in CI.

SystemReads untrusted textPrivate dataAuthorityTier
Support assistantuser chatone customer's ordersread-only lookupmedium
Knowledge assistantany wiki page or ticketeverything employees can seenone, but renders linkshigh
Coding agentissues, dependencies, web docssource code, CI secretswrite PRs, run codehigh

The knowledge assistant surprises people by being high tier. It has no tools, but any employee or customer who can write a ticket can plant instructions, and the assistant renders markdown, so a planted image link can carry retrieved secrets to an attacker's server. That is the full dangerous combination without a single tool. The coverage script flags it with two gaps: no retrieval ACL enforcement (the index was built with a service account that sees everything) and no output link rewriting. The fixes are platform controls: per-user document filtering at query time and an egress rewriter that strips or proxies external URLs.

The coding agent's gaps are authority gaps. It ran with a token able to push to the main branch and with CI secrets in its environment. The program moves it to a sandbox with no production credentials, scopes its token to opening pull requests on a branch prefix, and requires human review before merge. The support assistant needs the least: its lookup tool already takes the customer identity from the session rather than from the model, which is exactly the property that prevents one customer reading another's orders.

Notice what was not on the list: a better jailbreak filter. The highest-value work in this example was moving authority and data boundaries, which is typical.

Metrics that show the program works

Report a short set of numbers monthly, each with a definition precise enough to compute from the repository and logs.

  • Inventory completeness: systems found by discovery (API gateway logs, expense reports for AI vendors, egress to model APIs) that are missing from the inventory. The target is zero, and the first measurement is usually humbling.
  • High-tier gap count: output of the coverage script, with the age of the oldest open gap.
  • Control test pass rate: share of control tests passing on the latest run.
  • Red team findings by layer: where attacks succeed tells you which layer is under-built.
  • Time to revoke: from decision to a tool credential or model version being disabled in production, measured in drills.
  • Exceptions past expiry: accepted risks whose review date has passed.

Failure modes and trade-offs

Programs fail in recognisable ways. Filter-first thinking spends the budget on input classifiers that attackers paraphrase around, while tools keep broad credentials. Inventory rot happens when the inventory is a spreadsheet updated at audit time; new agents appear weekly and nobody records them. Application-by-application controls mean each team builds its own output filter, half of them wrong; shared layers belong in the platform. Untested controls pass every review and fail the first real attack because nobody checked that the allowlist actually loaded. Governance without teeth produces risk acceptances signed by people who do not own the budget to fix the risk, with no expiry date.

The trade-offs are real too. Strict confirmation prompts on every action make agents useless, so reserve them for irreversible or external effects. Two independent controls per threat doubles engineering cost, so apply that rule only to high-tier systems. Centralised platform controls slow teams that need something unusual; offer an exception path with an expiry instead of forcing everything through the paved road. And every log you keep for forensics is also data you must protect and eventually delete.

Key takeaway: <p>An LLM security program is an inventory, a threat model per high-risk system, controls placed outside the model in the layer where they actually stop something, tests that prove them, logs that show them firing, and a governance loop that re-tiers systems as they gain data and authority.</p><ul><li>Build the inventory this week from gateway logs and vendor spend, not from a survey.</li><li>Tier every system by untrusted input, private data and authority; flag any that have all three.</li><li>Draw data flows for each high-tier system and follow taint to every sink.</li><li>Put the threats, controls and test results in one repository and run the coverage script in CI.</li><li>Fix authority first: scope tool tokens, remove production secrets from agent environments, and add confirmation only for irreversible actions.</li><li>Move output link rewriting, egress allowlists and model intake to the platform so every app inherits them.</li><li>Schedule a red team exercise against the highest-tier system and gate releases on its regression suite.</li><li>Run a revocation drill and record time to revoke as a metric.</li></ul>