Agent hijacking is what happens when an AI agent, working for one person, starts following instructions from someone else. NIST's Center for AI Standards and Innovation defines it as a type of indirect prompt injection: an attacker inserts malicious instructions into data the agent may ingest, causing it to take unintended, harmful actions. The OWASP Top 10 for Agentic Applications, published in December 2025, lists the related risk first, as ASI01, Agent Goal Hijack.

Prompt injection is the mechanism, and the indirect prompt injection article covers how injected text reaches a model. This article is about what happens next, once the model is attached to tools and credentials: how a hijack turns into actions, why filtering inputs cannot be the whole defence, which architectural patterns make the harmful action unreachable, and how to measure whether your agent is hijackable before an attacker does.

Advertisement

From injected text to harmful action

A chatbot that is injected says something wrong. An agent that is injected does something wrong, with the user's authority. The difference is the tool layer: email, file systems, payment APIs, code execution, browsers. Every hijack follows the same four-link chain, and every defence works by breaking one of the links.

Anatomy of an agent hijack: untrusted data travels from a source to a privileged action1. Sourceemail, web page, file2. Ingestiontool result into context3. Redirectionmodel adopts new goal4. Actiontool call with authorityControls, placed where each link in the chain can be brokenProvenancetag untrusted inputsIsolationquarantined reader, no toolsFixed plantools chosen before data readPolicy gateper-call checks, human approvalEgress control and audit loglimits what a successful hijack can send, and proves what happenedDetection alone is probabilistic. Design controls make the dangerous action unreachable from untrusted data.
The hijack chain (top) and the control that breaks each link (bottom). The strongest designs break the chain at redirection and action, where the attacker's text would have to become a privileged tool call.

Source: anything the attacker can write that the agent will read, such as an email, a calendar invite, a web page, a code comment, a PDF, a support ticket or a tool's description. Ingestion: a tool returns that content and it lands in the model's context next to the real instructions. Redirection: the model treats the content as instructions and adopts a new goal or modifies the current one. Action: the model emits tool calls that serve the attacker, using the permissions the user or operator granted.

A widely reported example is EchoLeak (CVE-2025-32711), disclosed in June 2025: a crafted email that Microsoft 365 Copilot processed during a normal query could lead it to leak data from the user's context without the user clicking anything. The pattern is general: the attacker never touched the victim's system, only data the assistant was entitled to read.

What hijacks look like in practice

PatternWhat the attacker wantsTypical tool abused
ExfiltrationSend private data outemail, HTTP fetch, image or link rendering
Action substitutionDo a different task with the user's rightspayments, file writes, repository commits
Argument tamperingSame task, attacker-chosen parametersrecipient, account number, file path
Goal persistenceSurvive into future sessionsmemory write, notes, configuration files
Lateral movementReach another agent or systemdelegation to sub-agents, shared inboxes
SabotageDeny service or destroy datadelete, loops of expensive calls

Argument tampering is the hardest to spot. The agent does exactly what the user asked, pays the invoice, but the account number came from an attacker-controlled document. Nothing about the tool sequence looks anomalous; only the origin of one argument is wrong. That observation drives the strongest defences below.

Hijacking also composes with other weaknesses. An agent that holds broad credentials turns every injection into a confused-deputy attack, and an agent that can reach arbitrary URLs gives every hijack an exfiltration channel.

Advertisement

Why detection alone is not enough

The instinctive defence is to detect injected instructions: classifiers on tool outputs, delimiters around untrusted text, system prompts telling the model to ignore embedded commands. All of these help and belong in a layered design, for example spotlighting, which marks untrusted text so the model can tell it apart. But they share a limitation: they ask a probabilistic system to reliably separate data from instructions in a channel where both are just tokens.

An attacker gets to try many phrasings, languages, encodings and formatting tricks, and needs only one to work; a defender needs every attempt to fail. NIST's evaluation work makes the same point from the measurement side: attacks should be adaptive, because a system tuned against known attacks can still fall to new ones, and success should be measured over multiple attempts, because a single try understates what a persistent attacker achieves. Treat detectors as a way to lower the rate and raise alarms, not as the control that makes an action safe.

Design control 1: fix the plan before reading untrusted data

If the agent decides which tools to call before it reads any untrusted content, injected text can no longer add a new tool call. This is the plan-then-execute pattern. The planner sees only the trusted user request and the tool catalogue and produces a fixed sequence. The executor runs it, and untrusted content can influence only data values inside that plan, not its shape.

The weakness is flexibility. Tasks where the next step genuinely depends on what was read ("reply to whichever email needs a response") need either a re-planning step, which reopens the door, or a design where the dependency is expressed as data flowing through a pre-approved step.

Design control 2: quarantine the reader

The dual-LLM pattern, described by Simon Willison, splits the agent in two. A privileged model plans and calls tools but never sees untrusted text. A quarantined model reads untrusted text but has no tools; its output is treated as data, referenced by a variable name, and never pasted back into the privileged model's prompt. Constraining the quarantined model's output to a schema, such as a date, an enum or a bounded string, further limits what an attacker can smuggle through it.

def run_task(user_request, planner_llm, reader_llm, tools):
    # 1. Plan from the trusted request only. The planner never sees tool output.
    plan = planner_llm.plan(user_request, allowed=tools.names())     # e.g. [read_inbox, summarise]
    # 2. Execute the fixed plan. Untrusted text flows only into the quarantined reader,
    #    which has no tools and returns data, not instructions.
    ctx = {}
    for step in plan:
        if step.kind == "tool":
            ctx[step.out] = tools.call(step.name, step.args(ctx))  # args checked by check_call()
        else:
            ctx[step.out] = reader_llm.extract(step.schema, ctx[step.src])  # typed, validated output
    return ctx[plan[-1].out]

This is not free. You now run two models and must design the typed interfaces between them, and anything the quarantined model extracts (a recipient address, say) is still attacker-influenced data. That is where the third control comes in.

Design control 3: track where every value came from

CaMeL (Debenedetti and colleagues, 2025, "Defeating Prompt Injections by Design") generalises the dual-LLM idea. The privileged model writes a small program from the user's request. An interpreter runs it, calls a quarantined model to parse untrusted data, and attaches capabilities to every value recording its sources and who may read it. Before each tool call, a policy checks those capabilities: for example, an email may only be sent to an address that came from the user, not from a document.

You can adopt the core idea without adopting the whole system. Wrap values with their provenance, propagate it through transformations, and write per-tool policies over it:

from dataclasses import dataclass, field

@dataclass(frozen=True)
class Value:
    data: object
    sources: frozenset = field(default_factory=frozenset)   # e.g. {"user"}, {"email:inbox"}

    @property
    def trusted(self):
        # Untagged means untrusted: fail closed if provenance was lost.
        return bool(self.sources) and self.sources <= {"user", "system"}

POLICY = {
    # tool name -> which arguments must come ONLY from trusted sources
    "send_email":  {"to"},
    "transfer":    {"to_account", "amount"},
    "read_inbox":  set(),
    "summarise":   set(),
}

def check_call(tool, args: dict[str, Value]) -> str:
    # Return "allow", "confirm" or "deny" for one proposed tool call.
    if tool not in POLICY:
        return "deny"
    tainted = [k for k in POLICY[tool] if not args[k].trusted]
    if not tainted:
        return "allow"
    # A sensitive argument was derived from untrusted content: never auto-run it.
    return "confirm" if tool != "transfer" else "deny"

This catches argument tampering, the case detection misses, because the check does not depend on recognising malicious text. It only asks whether a sensitive argument's value was derived from an untrusted source. The cost is plumbing: provenance must survive every transformation, including string formatting and the model's own paraphrasing, or it is lost exactly where it matters.

Reduce what a successful hijack can do

  • Least privilege per task. Grant the tools and scopes the current task needs, for the task's duration, rather than every tool the agent might ever use. The agent tool permissions article covers least privilege between a model and real actions.
  • Egress control. Deny outbound requests by default and allow-list destinations, including image and link rendering in the UI, which is a classic exfiltration channel. See egress filtering for LLM systems.
  • Human confirmation for irreversible actions, showing the exact arguments and flagging those derived from untrusted content. Confirmations lose value if they fire on every step, because people approve them without reading.
  • Rate and budget limits on tool calls per task, so a hijacked loop cannot run indefinitely.
  • Protected memory. Writes to long-term memory or configuration from a session that ingested untrusted data should be reviewed or quarantined, because they turn a one-time hijack into a persistent one.
  • Audit logs that record each tool call with its arguments and the provenance of each argument, so an incident can be reconstructed.

Measuring hijack risk

You cannot manage what you do not measure. AgentDojo (Debenedetti and colleagues, 2024) is an open benchmark that places an agent in simulated environments such as a workspace, a bank, travel booking and Slack, gives it user tasks, and plants injection tasks in the data its tools return. It reports two numbers that must be read together: utility (did the agent still complete the user's task?) and attack success (did it perform the attacker's task?). A defence that drives attack success to zero by refusing everything has not helped.

NIST's CAISI extended this work with an Inspect-based version, AgentDojo-Inspect, and drew lessons worth copying into your own evaluations: keep extending the attack set, since frameworks age; use adaptive red teaming against your actual system; report results per task, because aggregate rates hide the high-impact tasks that fail; and measure success across repeated attempts. Build a small internal suite modelled on this: your real tools, your real data shapes, and injection tasks mapped to your highest-impact actions, run in CI whenever the model, prompt or tool set changes. Pair it with human red teaming for the attacks nobody has automated yet.

Failure modes in defences

  • Provenance lost in transit. A value tagged untrusted is summarised by a model and comes back untagged. Tag at the interpreter level, not in prompts.
  • Tool descriptions as a source. Third-party tool and server descriptions are untrusted input too; pin and review them.
  • Confirmation fatigue. Too many prompts train users to click through; confirm only irreversible or externally visible actions.
  • Evaluation on the happy path. A suite without adaptive attacks gives false confidence.
  • Over-privileged service accounts. The agent's credential, not the model, sets the upper bound on damage.

What to do next

  1. Inventory every source of untrusted text your agent reads, including tool descriptions and retrieved documents.
  2. List every tool that is irreversible or sends data outside, and require provenance checks or human confirmation on their sensitive arguments.
  3. Move to plan-then-execute or a quarantined reader for workflows that touch untrusted content.
  4. Enforce default-deny egress and per-task, least-privilege credentials.
  5. Build an AgentDojo-style suite over your own tools, report utility and attack success per task, and run it on every model or prompt change.
  6. Log tool calls with argument provenance and rehearse an incident using those logs.
Key takeaway: Agent hijacking turns indirect prompt injection into actions taken with the user's authority. Every hijack runs from source to ingestion to redirection to action, and the most dependable defences break the chain by design rather than by detection: fix the plan before reading untrusted data, isolate the model that reads it, track the provenance of every argument and enforce per-tool policies on it. Limit the blast radius with least privilege, egress control and confirmations for irreversible actions, and measure risk continuously with adaptive, per-task evaluations.