Predictions about security are cheap and usually wrong in their details, so this article does not try to guess product launches or name the next famous jailbreak. Instead it takes the mechanisms that already exist in deployed systems in October 2026, follows each to where it is plainly heading, and pairs it with a control you can build now. The forecasts are labelled as forecasts. The controls are not speculative: each one addresses an attack class that has already been demonstrated against real products.
The single biggest change is the move from chatbots that produce text to agents that take actions. A chatbot that is tricked says something embarrassing; an agent that is tricked forwards a mailbox, merges a pull request, or spends money. The rest of this article explains why that shift turns several familiar weaknesses into structural risks, what a reference architecture for agent security looks like, how a deterministic policy gateway works in code, and what a security team should do in the next quarter.
From chatbots to agents: why the risk changed shape
A useful way to reason about LLM application risk is to multiply three factors: what untrusted content can reach the model, what the model can read, and what the model can do. Early chat deployments had a lot of the first and little of the other two. Agents raise all three at once: they browse, read email and documents, call tools through protocols such as the Model Context Protocol (MCP), keep long-term memory, and increasingly hand work to other agents.
Simon Willison named the dangerous combination the lethal trifecta: access to private data, exposure to untrusted content, and the ability to communicate externally. Any system that holds all three can be made to exfiltrate data by text it reads, because the model cannot reliably tell the user's instructions from instructions embedded in a web page or email. The EchoLeak issue reported against Microsoft 365 Copilot in 2025, in which a crafted email could cause the assistant to leak data without the user clicking anything, is the archetype. The industry's own taxonomies have moved accordingly. The 2025 edition of the OWASP Top 10 for LLM Applications added system prompt leakage and vector and embedding weaknesses, and rewrote excessive agency around agents that can take more actions than their job needs; in December 2025 the OWASP GenAI Security Project published a separate Top 10 for Agentic Applications, whose first entry is agent goal hijack, followed by tool misuse, identity and privilege abuse, and agentic supply chain compromise.
Six trends, each with a control
Each trend below states the mechanism that exists today, the forecast, and the control. Treat the forecasts as reasoned bets, not facts.
- Prompt injection stays unsolved at the model layer. Mechanism: models process instructions and data in one token stream, and training-based defences reduce attack success rates without driving them to zero; adaptive attackers keep finding phrasings that work. Forecast: injection will be managed like memory-safety bugs, by architecture, rather than eliminated by better filters. Control: design so that a fully compromised model cannot cause unacceptable harm. Patterns include the dual-LLM split, where a quarantined model reads untrusted text but has no tools, and capability-based designs such as Google DeepMind's CaMeL (2025), which extracts a plan from the trusted query and tracks where every value came from before allowing it to flow into a tool call.
- Agents become first-class identities. Mechanism: agents today often act with a user's full OAuth token or a broad service account, which makes every injection a confused-deputy attack. Forecast: agent identity, delegation and per-task scoping will become a standard part of identity platforms, and MCP's authorization model, built on OAuth 2.1, is an early sign. Control: give each agent task short-lived, narrowly scoped credentials, and record which human delegated what.
- The tool ecosystem becomes the supply chain. Mechanism: MCP servers and plugins are installed from registries, and their tool descriptions are text that the model reads and obeys; tool-poisoning attacks hide instructions in those descriptions or change them after approval. Forecast: tool registries will need signing, pinning and review in the way package registries did. Control: pin tool servers by version and hash, diff tool descriptions on every update, and run third-party servers in sandboxes with no ambient credentials.
- Memory and retrieval become persistence mechanisms. Mechanism: content written to long-term memory or a vector index is replayed into future contexts, so one successful injection can persist across sessions and users. Forecast: memory poisoning will become the agent equivalent of a web shell. Control: record provenance on every memory entry, separate per-user stores, and expire or review memories written while the context was tainted.
- Model artefacts get the full software-supply-chain treatment. Mechanism: weights are downloaded from hubs, older serialisation formats such as Python pickle can execute code on load, and fine-tuned derivatives can carry backdoors. Control: load only safe formats such as safetensors, verify signatures (the OpenSSF model-signing project defines one approach), and keep an AI bill of materials for every deployed model.
- Evaluation becomes a continuous control, and attackers automate too. Mechanism: automated red-teaming tools already generate and mutate attacks at scale. Forecast: both offence and defence will run model-driven attack generation continuously, so a one-off pre-launch red team will be worth little. Control: wire an injection and tool-misuse test suite into CI and block releases on regressions.
A reference architecture for agent security
The architecture that follows from those trends puts a deterministic policy gateway between the model and every tool. The model proposes; the gateway decides. The gateway knows things the model cannot be trusted to remember: which inputs in this session came from untrusted sources, which tools read private data, which tools can send data out, who the user is, and what that user has delegated. It enforces rules that are written as code, so a clever paragraph in a web page cannot argue with them.
Three properties make the gateway effective. It sees every call, including calls to third-party tools, so there is no side door. It tracks taint at the session level at minimum, and at value level where you can afford it, so it knows whether a send-email request could carry data that an attacker influenced. And it escalates rather than silently failing: when a risky combination occurs, a human sees the exact recipient, URL and payload, not the model's summary of them. The surrounding layers are the familiar ones described in defence in depth for LLM systems, now aimed at actions instead of words.
A policy gateway in code
Here is a minimal gateway that enforces the lethal-trifecta rule. It is deliberately small: it marks a session as tainted when any non-user content enters the context, marks it as holding private data when a private-read tool runs, and refuses outbound tools when both are true unless a human approves the exact arguments. It also shows a second, simpler policy (an allow-listed recipient domain) to illustrate that policies should be ordinary code reviewed like any other access-control code.
from dataclasses import dataclass, field
# Capability classes for every tool the agent can call.
READS_PRIVATE = {"read_mailbox", "search_drive", "query_crm"}
SENDS_OUT = {"send_email", "http_post", "create_share_link"}
@dataclass
class Session:
user: str
tainted: bool = False # has untrusted content entered the context?
touched_private: bool = False # has private data entered the context?
log: list = field(default_factory=list)
def ingest(session, content, source):
"""Call for every tool result, retrieved chunk or uploaded file."""
if source != "user_prompt":
session.tainted = True # web pages, email bodies, PDFs, tool output
session.log.append(("ingest", source, len(content)))
def authorize(session, tool, args, approve):
"""Deterministic policy, evaluated outside the model, before every tool call."""
if tool in READS_PRIVATE:
session.touched_private = True
# Lethal trifecta: private data + untrusted content + a way out.
if tool in SENDS_OUT and session.tainted and session.touched_private:
if not approve(session.user, tool, args): # human sees the exact args
session.log.append(("deny", tool))
raise PermissionError(f"{tool} blocked: tainted context holds private data")
if tool == "send_email" and not all(
r.endswith("@example.com") for r in args["to"]):
raise PermissionError("external recipient requires a separate flow")
session.log.append(("allow", tool))Session-level taint is coarse: once any untrusted text arrives, every later outbound call needs approval, which can be noisy in long sessions. Value-level tracking, in the style of CaMeL, is more precise because it knows whether the specific email body being sent was derived from untrusted content, but it requires the agent to express plans in a form the runtime can analyse. Start coarse, measure the approval rate, and invest in finer tracking where approval fatigue becomes a real problem.
Worked example: an email assistant under attack
Consider an email assistant that can read the user's mailbox, search their drive and send email. An attacker sends a message whose body contains hidden text: an instruction to search the drive for documents about an acquisition and email the results to an outside address, phrased to look like an automated compliance request.
- The user asks the assistant to summarise today's inbox. The gateway marks the session tainted as soon as email bodies are ingested, because they come from outside.
- The model, following the injected text, calls
search_drive. That is a private-read tool, so the gateway lets it run and marks the session as holding private data. Reading is allowed because blocking all reads would make the assistant useless. - The model calls
send_emailto the attacker's address. The gateway's first rule fires: the context is tainted and holds private data. The user sees an approval card with the real recipient and attachment list, and declines. - Even if the user approved by reflex, the second rule rejects the external recipient, because sending outside the company goes through a separate, explicit flow.
- The audit log records the ingest, the private read, the denied send and the rule that fired. Security turns the email into a regression test that runs in CI against every future model or prompt change.
Notice what did not save the day: the model did not recognise the attack, and no input classifier was required. Classifiers are still worth running, as described in indirect prompt injection, but they lower the attack rate rather than bound the damage. The gateway bounded the damage.
Standards and regulation
Regulation and standards will push in the same direction. The EU AI Act's obligations for general-purpose AI model providers have applied since 2 August 2025, NIST's Generative AI Profile (NIST AI 600-1, July 2024) maps generative risks onto the AI Risk Management Framework, and the OWASP lists above give auditors a shared vocabulary. Forecast: within a few years, buyers and auditors will routinely ask for evidence that agent actions are authorised by policy outside the model, logged, and tested, in the way they ask for access reviews today. Teams that already have the gateway, the logs and the CI suite will answer that question with artefacts instead of essays.
Failure modes
- Trusting the system prompt as a security boundary. Instructions in the prompt are advisory to the model, and anything the model can say, an attacker can make it say.
- Gateway bypasses. A tool wired directly into the agent framework, or a code-execution tool that can make its own network calls, skips every policy. Inventory every path out, and run code execution in sandboxes with egress control, as in sandboxing agents with Docker.
- Approval fatigue. If users approve dozens of prompts a day they stop reading them. Track approval rates and tune policies until escalations are rare and meaningful.
- Summarised approvals. An approval card generated by the model can be manipulated by the same injection. Render approvals from the raw arguments.
- Rendering as exfiltration. Markdown images and links in output can carry data in URLs to an attacker's server without any tool call. Strip or proxy them.
- Static red teaming. A test suite frozen at launch decays as models, prompts and tools change. Re-run it on every change and add each new incident as a case.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Session-level taint | Simple, model-agnostic, easy to audit | Over-escalates in long sessions |
| Value-level provenance (CaMeL-style) | Precise, fewer approvals | Requires analysable plans; engineering effort |
| Human approval | Catches what policy cannot express | Latency and fatigue |
| Input classifiers | Cheap reduction in attack rate | Never a bound on impact; adaptive attackers evade |
| Dropping a trifecta leg | Strongest guarantee | Less capable product |
The strongest control is still removing one leg of the trifecta for a given workflow: an agent that reads untrusted web pages but has no private data, or one that holds private data but cannot send anything out, cannot be made to exfiltrate. Product teams will resist, so make the trade-off explicit per workflow and record who accepted the residual risk.
What to do next
- Inventory every agent and tool, and classify each tool as private-read, public-read, write or send.
- Mark every workflow that has all three trifecta legs and decide, per workflow, which leg to remove or gate.
- Put a deterministic policy gateway in front of all tool calls, including third-party MCP servers, with session-level taint as a first version.
- Replace broad user tokens with short-lived, task-scoped credentials and record delegation.
- Pin and hash third-party tool servers, diff their descriptions on update, and sandbox them without ambient credentials.
- Render approvals from raw arguments and strip or proxy outbound links and images in model output.
- Build an injection and tool-misuse regression suite, run it in CI, and add every incident to it.
- Map your controls to the OWASP agentic list and keep the evidence ready for auditors.