Most LLM security writing is a list of attacks and a list of defences. That is useful for learning, but it does not tell a team what to build, who owns it, or where in the request path each control belongs. A reference architecture does. It names the components every LLM application needs, draws the boundaries between them, states what each component promises to the others, and shows which threat each control is there to stop.

This article is that blueprint. It assumes you know the attack classes; the layer-by-layer defence story is told in LLM defense in depth. Here we organise the same ideas into trust zones, write down component contracts, trace three request types end to end, express policy as data, map the controls to the 2025 OWASP Top 10 for LLM applications, and walk one attack through the design. The goal is a design document you can adapt, review against and build in stages.

Advertisement

Principles the architecture encodes

  • The model is not a security boundary. It cannot reliably distinguish instructions from data, so nothing it outputs is trusted to authorise anything. Authorisation happens in deterministic code around it.
  • Every boundary has one enforcement point. If two services each partly check tool permissions, neither fully does.
  • Identity flows end to end. Retrieval and tools act as the end user, never as an all-powerful service account.
  • Untrusted text stays marked. Content from users, documents, web pages and tool results is labelled as data from the moment it enters until it reaches the model.
  • Every decision is logged with its inputs. When something goes wrong you must be able to replay why.

The five trust zones

Zone 0, untrusted sources: end users, uploaded files, web pages, email and anything else an attacker can write to. Zone 1, the edge: an AI gateway that authenticates callers, enforces quotas and spend limits, and runs cheap input screening. Zone 2, orchestration: your application logic, which assembles context, consults policy, calls the model, validates output and brokers tool calls. Zone 3, model serving: a hosted provider or your own inference fleet, treated as a powerful but untrusted text transformer. Zone 4, data and actions: retrieval indexes, databases, internal APIs and code execution. Alongside all of them sits the control plane: policies, secrets, the audit log, evaluation and red-team pipelines, and the registry of models and prompts.

Draw your system in these zones before you design any control. Every arrow that crosses a zone boundary needs a named component that enforces something there; the gateway design itself is covered in LLM gateway architecture.

Reference architecture: five trust zones, one enforcement point per boundaryZone 0users, documents, webZone 1: AI gatewayauthn, quotas, screeningZone 2: orchestrationContext assemblerspotlit untrusted textPolicy decision pointwho may do whatOutput validatorschemas, links, PIITool gatewayallowlist, approvalsZone 3: model servingprovider or self-hostedpromptZone 4: retrievalACL-filtered indexZone 4: toolssandbox + egress proxyControl plane: policies, keys and secrets, audit log, evals and red-team, model and prompt registryEvery arrow crosses a boundary with a named owner and a named control.
Five zones plus a control plane. Orchestration owns four enforcement components; the model sits outside it and is treated as untrusted.
Advertisement

Component contracts

A contract says what a component guarantees and what it refuses to do. Writing them down is what turns a diagram into something two teams can build against.

ComponentGuaranteesNever does
AI gatewayAuthenticated caller identity on every request; per-tenant rate, token and spend limits; request size capsMake content decisions it cannot explain; forward anonymous traffic
Context assemblerUntrusted text wrapped and marked; system instructions kept separate; retrieved chunks carry source and ACL idsConcatenate user text into the system prompt
Policy decision pointA yes or no for (subject, action, resource, context), with a reason, from versioned policyAsk the model whether an action is allowed
Output validatorResponses match the expected schema; links and markup are sanitised; secrets and PII are caught before displayPass model output into a shell, SQL or HTML sink unescaped
Tool gatewayOnly allowlisted tools with validated arguments; high-impact actions need approval; calls run with the user's delegated credentialsHold standing admin credentials
RetrievalResults filtered by the caller's permissions before rankingRely on the model to hide documents
Sandbox and egress proxyCode runs isolated; outbound traffic only to allowlisted destinationsAllow arbitrary network calls from tool code
Audit logAppend-only record of prompts, tool calls, policy decisions and outputs, with identitiesStore raw secrets or unredacted regulated data

The context assembler's marking technique is described in spotlighting, and the tool gateway's job in tool abuse. The outbound side of the sandbox is covered in egress filtering.

Three request flows

Plain chat. The gateway authenticates and meters the request. The assembler places the user message in the user turn, never inside the system prompt. The model answers. The validator checks the answer against the output policy and renders it with escaping. The audit log records the exchange. Main risks: jailbreaks, system prompt leakage, harmful output and cost abuse.

Retrieval-augmented answer. As above, plus a retrieval step that runs with the caller's identity, so the index only returns chunks the caller may read. The assembler wraps each chunk as data, with its source id. The validator checks that any citations point to chunks that were actually retrieved. New risks: indirect prompt injection in documents and data leaking across tenants or permission levels.

Agent tool call. The model proposes a call as structured output. The tool gateway validates the arguments against the tool's schema, asks the policy decision point whether this user may perform this action on this resource now, and requests human approval when policy demands it. The call runs in a sandbox with delegated credentials and egress limits. The result returns marked as untrusted data, and the loop continues with a step budget. New risks: excessive agency, confused-deputy actions, data exfiltration through tool arguments and runaway loops.

Policy as data

Policies that live in code reviews and prompt text drift. Put them in versioned data, evaluated by one service, so that a change is a diff and every decision can name the rule that produced it.

tools:
  refund_payment:
    impact: high
    allow_roles: [support_agent]
    approval_above: 200
    args_schema: schemas/refund_payment.json
  search_orders:
    impact: low
    allow_roles: [support_agent, customer]
    scope: own_records_only
    args_schema: schemas/search_orders.json
agent:
  max_steps: 12
  max_tool_calls_per_turn: 4
def decide(policy, subject, tool, args, ctx):
    rule = policy["tools"].get(tool)
    if rule is None:
        return Decision(False, "tool not allowlisted")
    if not set(subject.roles) & set(rule["allow_roles"]):
        return Decision(False, "role not permitted")
    schema = rule.get("args_schema")
    if schema is None or not validate(args, schema):
        return Decision(False, "missing schema or arguments fail schema")
    if rule.get("scope") == "own_records_only" and args.get("customer_id") != subject.id:
        return Decision(False, "outside caller scope")
    if rule.get("impact") == "high" and args.get("amount", 0) > rule.get("approval_above", 0):
        return Decision(True, "needs approval", approval=True)
    return Decision(True, "allowed")

Note what the function does not consult: the model's explanation of why it wants to call the tool. Reasons from the model can be logged for reviewers but never feed the decision.

Identity propagation

The most common architectural flaw in agent systems is the ambient service account: the agent's tools run as one identity that can read every customer's records, and the model is trusted to ask only for the right ones. A single injected instruction then turns the agent into a confused deputy. The fix is to carry the end user's identity through every hop. The gateway authenticates the user and issues a short-lived token. The orchestrator passes it to retrieval and to the tool gateway. Tools exchange it for narrowly scoped downstream credentials. The effective permission of any action is then the intersection of what the user may do and what the tool is for, regardless of what the model asks.

Background agents that act without a live user need their own identities with explicitly granted scopes, and their actions should be attributable in the audit log to the agent, the policy version and the human who configured it.

Control map: the 2025 OWASP LLM Top 10

Mapping controls to a shared risk list makes gaps visible in a review. The identifiers below are those of the 2025 OWASP Top 10 for LLM applications; a fuller treatment of each risk is in OWASP Top 10 for LLM applications.

RiskPrimary controls in this architecture
LLM01 Prompt InjectionContext assembler marking; tool gateway and policy decision point independent of model output; approval for high-impact actions
LLM02 Sensitive Information DisclosureACL-filtered retrieval; output validator PII and secret detection; redacted audit log
LLM03 Supply ChainModel and prompt registry with provenance; pinned model versions; dependency scanning of tool code
LLM04 Data and Model PoisoningIngestion review for indexed content; evaluation gates before model or index promotion
LLM05 Improper Output HandlingOutput validator: schemas, escaping, no direct execution of model text
LLM06 Excessive AgencyTool allowlists, delegated credentials, step budgets, human approval
LLM07 System Prompt LeakageNo secrets or authorisation logic in prompts; canary strings detected by the validator
LLM08 Vector and Embedding WeaknessesPer-tenant or ACL-filtered indexes; provenance on every chunk
LLM09 MisinformationCitation checks against retrieved chunks; grounding evaluations; clear uncertainty in the interface
LLM10 Unbounded ConsumptionGateway quotas, token and spend caps, agent step limits, timeouts

Worked example: an injected refund

A support agent can search orders and issue refunds. An attacker files a ticket whose body says the assistant should refund order 7781 in full to a new card. A support employee later asks the agent to summarise open tickets.

The ticket text enters through retrieval and is wrapped as untrusted data by the assembler. Suppose the model is persuaded anyway and proposes refund_payment(order=7781, amount=950). The tool gateway validates the arguments, which are well-formed, then asks the policy decision point. The employee's role permits refunds, but 950 exceeds the 200 limit, so the decision is approval required, and the approval screen shows the order, the amount and the ticket that preceded the call. The employee declines. The audit log holds the ticket id, the proposed call, the policy version and the decision, and a detection rule flags tool calls proposed directly after untrusted content that requested them. No single control was perfect; the architecture made the failure of one control survivable.

Failure modes of the architecture itself

  • Bypass paths. A batch job or an internal tool that calls the model directly skips the gateway and every control behind it. Make the gateway the only route that holds model credentials.
  • Policy in prompts. Rules written only as instructions to the model are suggestions. Anything that must hold goes into the policy decision point.
  • Fail-open dependencies. When a classifier or the policy service times out, the default must be to deny high-impact actions.
  • Approval fatigue. If every call needs approval, reviewers click through. Reserve approval for high impact and make the approval screen show what matters.
  • Logs as a breach. An audit log full of raw prompts is the most valuable data store you own. Redact and restrict it.

Rolling it out in phases

Phase one puts a gateway in front of all model traffic, with authentication, quotas and an audit log; this alone closes bypass paths and makes later work observable. Phase two adds output validation and context marking. Phase three adds the tool gateway, the policy decision point and delegated credentials, which is the step that contains agents. Phase four adds the continuous parts: evaluation gates on model and prompt changes, red-team suites in CI and periodic access reviews. Measure each phase with a small set of numbers: share of model traffic through the gateway, blocked and approved tool calls per day, approval rejection rate, and red-team pass rate per release.

The trade-off is latency and friction. Each inline control adds milliseconds to tens of milliseconds, and each approval adds a human. Spend them where impact is high: a read-only summariser needs little, an agent that moves money needs all of it.

What to do next

  1. Draw your application in the five zones and list every boundary crossing with its enforcement component, or mark it as missing.
  2. Route all model calls through one gateway and remove model credentials from everything else.
  3. Write the contract table for your components and get each owning team to agree to it.
  4. Move tool permissions into versioned policy data evaluated by one decision point, and propagate end-user identity to retrieval and tools.
  5. Fill in the OWASP control map for your system and turn each empty cell into a ticket.
  6. Replay the injected-refund example against your design and confirm which control stops it.
Key takeaway: A reference architecture turns LLM security from a list of attacks into a build plan. Separate the system into trust zones, give every boundary one enforcement component with a written contract, treat the model as untrusted, carry the user's identity to every tool and index, keep policy as versioned data that never consults the model's reasoning, and check coverage against the 2025 OWASP list. Build it in phases, starting with a single gateway, and design so that the failure of any one control is survivable.