Most LLM security writing is a list of attacks and a list of defences. That is useful for learning, but it does not tell a team what to build, who owns it, or where in the request path each control belongs. A reference architecture does. It names the components every LLM application needs, draws the boundaries between them, states what each component promises to the others, and shows which threat each control is there to stop.
This article is that blueprint. It assumes you know the attack classes; the layer-by-layer defence story is told in LLM defense in depth. Here we organise the same ideas into trust zones, write down component contracts, trace three request types end to end, express policy as data, map the controls to the 2025 OWASP Top 10 for LLM applications, and walk one attack through the design. The goal is a design document you can adapt, review against and build in stages.
Principles the architecture encodes
- The model is not a security boundary. It cannot reliably distinguish instructions from data, so nothing it outputs is trusted to authorise anything. Authorisation happens in deterministic code around it.
- Every boundary has one enforcement point. If two services each partly check tool permissions, neither fully does.
- Identity flows end to end. Retrieval and tools act as the end user, never as an all-powerful service account.
- Untrusted text stays marked. Content from users, documents, web pages and tool results is labelled as data from the moment it enters until it reaches the model.
- Every decision is logged with its inputs. When something goes wrong you must be able to replay why.
The five trust zones
Zone 0, untrusted sources: end users, uploaded files, web pages, email and anything else an attacker can write to. Zone 1, the edge: an AI gateway that authenticates callers, enforces quotas and spend limits, and runs cheap input screening. Zone 2, orchestration: your application logic, which assembles context, consults policy, calls the model, validates output and brokers tool calls. Zone 3, model serving: a hosted provider or your own inference fleet, treated as a powerful but untrusted text transformer. Zone 4, data and actions: retrieval indexes, databases, internal APIs and code execution. Alongside all of them sits the control plane: policies, secrets, the audit log, evaluation and red-team pipelines, and the registry of models and prompts.
Draw your system in these zones before you design any control. Every arrow that crosses a zone boundary needs a named component that enforces something there; the gateway design itself is covered in LLM gateway architecture.
Component contracts
A contract says what a component guarantees and what it refuses to do. Writing them down is what turns a diagram into something two teams can build against.
| Component | Guarantees | Never does |
|---|---|---|
| AI gateway | Authenticated caller identity on every request; per-tenant rate, token and spend limits; request size caps | Make content decisions it cannot explain; forward anonymous traffic |
| Context assembler | Untrusted text wrapped and marked; system instructions kept separate; retrieved chunks carry source and ACL ids | Concatenate user text into the system prompt |
| Policy decision point | A yes or no for (subject, action, resource, context), with a reason, from versioned policy | Ask the model whether an action is allowed |
| Output validator | Responses match the expected schema; links and markup are sanitised; secrets and PII are caught before display | Pass model output into a shell, SQL or HTML sink unescaped |
| Tool gateway | Only allowlisted tools with validated arguments; high-impact actions need approval; calls run with the user's delegated credentials | Hold standing admin credentials |
| Retrieval | Results filtered by the caller's permissions before ranking | Rely on the model to hide documents |
| Sandbox and egress proxy | Code runs isolated; outbound traffic only to allowlisted destinations | Allow arbitrary network calls from tool code |
| Audit log | Append-only record of prompts, tool calls, policy decisions and outputs, with identities | Store raw secrets or unredacted regulated data |
The context assembler's marking technique is described in spotlighting, and the tool gateway's job in tool abuse. The outbound side of the sandbox is covered in egress filtering.
Three request flows
Plain chat. The gateway authenticates and meters the request. The assembler places the user message in the user turn, never inside the system prompt. The model answers. The validator checks the answer against the output policy and renders it with escaping. The audit log records the exchange. Main risks: jailbreaks, system prompt leakage, harmful output and cost abuse.
Retrieval-augmented answer. As above, plus a retrieval step that runs with the caller's identity, so the index only returns chunks the caller may read. The assembler wraps each chunk as data, with its source id. The validator checks that any citations point to chunks that were actually retrieved. New risks: indirect prompt injection in documents and data leaking across tenants or permission levels.
Agent tool call. The model proposes a call as structured output. The tool gateway validates the arguments against the tool's schema, asks the policy decision point whether this user may perform this action on this resource now, and requests human approval when policy demands it. The call runs in a sandbox with delegated credentials and egress limits. The result returns marked as untrusted data, and the loop continues with a step budget. New risks: excessive agency, confused-deputy actions, data exfiltration through tool arguments and runaway loops.
Policy as data
Policies that live in code reviews and prompt text drift. Put them in versioned data, evaluated by one service, so that a change is a diff and every decision can name the rule that produced it.
tools:
refund_payment:
impact: high
allow_roles: [support_agent]
approval_above: 200
args_schema: schemas/refund_payment.json
search_orders:
impact: low
allow_roles: [support_agent, customer]
scope: own_records_only
args_schema: schemas/search_orders.json
agent:
max_steps: 12
max_tool_calls_per_turn: 4def decide(policy, subject, tool, args, ctx):
rule = policy["tools"].get(tool)
if rule is None:
return Decision(False, "tool not allowlisted")
if not set(subject.roles) & set(rule["allow_roles"]):
return Decision(False, "role not permitted")
schema = rule.get("args_schema")
if schema is None or not validate(args, schema):
return Decision(False, "missing schema or arguments fail schema")
if rule.get("scope") == "own_records_only" and args.get("customer_id") != subject.id:
return Decision(False, "outside caller scope")
if rule.get("impact") == "high" and args.get("amount", 0) > rule.get("approval_above", 0):
return Decision(True, "needs approval", approval=True)
return Decision(True, "allowed")Note what the function does not consult: the model's explanation of why it wants to call the tool. Reasons from the model can be logged for reviewers but never feed the decision.
Identity propagation
The most common architectural flaw in agent systems is the ambient service account: the agent's tools run as one identity that can read every customer's records, and the model is trusted to ask only for the right ones. A single injected instruction then turns the agent into a confused deputy. The fix is to carry the end user's identity through every hop. The gateway authenticates the user and issues a short-lived token. The orchestrator passes it to retrieval and to the tool gateway. Tools exchange it for narrowly scoped downstream credentials. The effective permission of any action is then the intersection of what the user may do and what the tool is for, regardless of what the model asks.
Background agents that act without a live user need their own identities with explicitly granted scopes, and their actions should be attributable in the audit log to the agent, the policy version and the human who configured it.
Control map: the 2025 OWASP LLM Top 10
Mapping controls to a shared risk list makes gaps visible in a review. The identifiers below are those of the 2025 OWASP Top 10 for LLM applications; a fuller treatment of each risk is in OWASP Top 10 for LLM applications.
| Risk | Primary controls in this architecture |
|---|---|
| LLM01 Prompt Injection | Context assembler marking; tool gateway and policy decision point independent of model output; approval for high-impact actions |
| LLM02 Sensitive Information Disclosure | ACL-filtered retrieval; output validator PII and secret detection; redacted audit log |
| LLM03 Supply Chain | Model and prompt registry with provenance; pinned model versions; dependency scanning of tool code |
| LLM04 Data and Model Poisoning | Ingestion review for indexed content; evaluation gates before model or index promotion |
| LLM05 Improper Output Handling | Output validator: schemas, escaping, no direct execution of model text |
| LLM06 Excessive Agency | Tool allowlists, delegated credentials, step budgets, human approval |
| LLM07 System Prompt Leakage | No secrets or authorisation logic in prompts; canary strings detected by the validator |
| LLM08 Vector and Embedding Weaknesses | Per-tenant or ACL-filtered indexes; provenance on every chunk |
| LLM09 Misinformation | Citation checks against retrieved chunks; grounding evaluations; clear uncertainty in the interface |
| LLM10 Unbounded Consumption | Gateway quotas, token and spend caps, agent step limits, timeouts |
Worked example: an injected refund
A support agent can search orders and issue refunds. An attacker files a ticket whose body says the assistant should refund order 7781 in full to a new card. A support employee later asks the agent to summarise open tickets.
The ticket text enters through retrieval and is wrapped as untrusted data by the assembler. Suppose the model is persuaded anyway and proposes refund_payment(order=7781, amount=950). The tool gateway validates the arguments, which are well-formed, then asks the policy decision point. The employee's role permits refunds, but 950 exceeds the 200 limit, so the decision is approval required, and the approval screen shows the order, the amount and the ticket that preceded the call. The employee declines. The audit log holds the ticket id, the proposed call, the policy version and the decision, and a detection rule flags tool calls proposed directly after untrusted content that requested them. No single control was perfect; the architecture made the failure of one control survivable.
Failure modes of the architecture itself
- Bypass paths. A batch job or an internal tool that calls the model directly skips the gateway and every control behind it. Make the gateway the only route that holds model credentials.
- Policy in prompts. Rules written only as instructions to the model are suggestions. Anything that must hold goes into the policy decision point.
- Fail-open dependencies. When a classifier or the policy service times out, the default must be to deny high-impact actions.
- Approval fatigue. If every call needs approval, reviewers click through. Reserve approval for high impact and make the approval screen show what matters.
- Logs as a breach. An audit log full of raw prompts is the most valuable data store you own. Redact and restrict it.
Rolling it out in phases
Phase one puts a gateway in front of all model traffic, with authentication, quotas and an audit log; this alone closes bypass paths and makes later work observable. Phase two adds output validation and context marking. Phase three adds the tool gateway, the policy decision point and delegated credentials, which is the step that contains agents. Phase four adds the continuous parts: evaluation gates on model and prompt changes, red-team suites in CI and periodic access reviews. Measure each phase with a small set of numbers: share of model traffic through the gateway, blocked and approved tool calls per day, approval rejection rate, and red-team pass rate per release.
The trade-off is latency and friction. Each inline control adds milliseconds to tens of milliseconds, and each approval adds a human. Spend them where impact is high: a read-only summariser needs little, an agent that moves money needs all of it.
What to do next
- Draw your application in the five zones and list every boundary crossing with its enforcement component, or mark it as missing.
- Route all model calls through one gateway and remove model credentials from everything else.
- Write the contract table for your components and get each owning team to agree to it.
- Move tool permissions into versioned policy data evaluated by one decision point, and propagate end-user identity to retrieval and tools.
- Fill in the OWASP control map for your system and turn each empty cell into a ticket.
- Replay the injected-refund example against your design and confirm which control stops it.