A network allowlist answers one question: may this workload talk to this host at all? For a coding assistant or a support agent that question is necessary and nowhere near sufficient. The agent legitimately needs to reach a code host, a ticketing API, a search service and a chat webhook, and every one of those is also a place where data can be written. An injected instruction does not need an attacker-controlled domain when it can ask the agent to open an issue on a public repository, append a customer record to a search query, or post a summary to a webhook that the proxy already trusts.
This article is about the layer above the network: deciding, per task and per request, what an agent may send to destinations it is allowed to reach. It builds an egress broker that holds task-scoped grants, tracks which data classes have entered the agent's context, and returns allow, approve or deny for each outbound call. The code runs, the worked example traces five real decisions, and the article ends with a rollout plan and a checklist. The network plane underneath (default-deny policies, the forward proxy, DNS and metadata defences) is covered in Egress Control for LLM Workloads and is assumed here.
Why a network allowlist is not enough for agents
Three conditions together make an agent exfiltration-capable: it can read private data, it processes content an attacker can influence, and it can communicate externally. Simon Willison calls this combination the lethal trifecta. You rarely remove the first two, since they are the product. Egress control is how you constrain the third without switching it off.
Host-level allowlists fail against agents in four recurring ways. First, allowlisted sinks are writable: a code host that serves your private repositories also accepts gists, issues and comments on public projects. Second, reads are writes: a GET to an allowed search API carries whatever the agent puts in the query string, and an image URL in rendered output is a GET the user's browser makes for you (that rendering channel is covered in LLM egress filtering architecture). Third, the allowlist is per workload, but risk is per task: the same agent is harmless while drafting release notes and dangerous five minutes later when it has a customer record in context. Fourth, allowlists say nothing about volume, so a slow drip of small requests to a permitted host is invisible.
The fix is to make egress a capability attached to the task, narrowed by what the task has seen. The network proxy cannot evaluate that, because it does not know which task issued a request or what the model has read. The tool runtime does.
Architecture: an egress broker in the tool runtime
Four components make this work. The tool runtime is the only thing that performs network I/O on the agent's behalf; the model emits tool calls, it never opens sockets. Inside a sandbox (see sandboxing agents with Docker) direct egress is blocked, so the runtime's HTTP client is the single choke point. The broker is a library or sidecar the runtime consults before every request. The approval queue shows a human the exact destination and payload. The audit log records every decision with byte counts, so you can reconstruct what left and why.
The network proxy still matters. It catches anything that bypasses the broker, such as a library making its own connection or a compromised tool, and it owns the defences against DNS rebinding and cloud metadata access.
Grants, labels and the decision function
A grant is a narrow statement of an outbound capability: an exact host, the methods allowed, a path pattern, the query parameters allowed, the data classes cleared for that sink, whether the sink tolerates instructions derived from untrusted content, and a per-request byte cap. Grants are issued when the task starts, from the task type, not from anything the model says.
Session labels are the other half. Every tool result carries labels describing what entered the context: customer_pii from the CRM, internal from the wiki, untrusted from inbound email or any fetched web page. Labels only accumulate during a task, because once a value is in the context window you cannot prove the model forgot it. A tool with no label mapping is treated as untrusted, so new tools fail closed until someone classifies them.
from dataclasses import dataclass, field
from urllib.parse import urlsplit, parse_qsl
@dataclass(frozen=True)
class Grant:
host: str # exact host; no wildcard across registrable domains
methods: frozenset # e.g. {"GET"} or {"POST"}
path: str # segment pattern: "*" matches one segment only
params: frozenset = frozenset() # query parameter names allowed
classes: frozenset = frozenset({"public"}) # data classes cleared for this sink
accepts_untrusted: bool = False # sink is internal and human-reviewed
max_bytes: int = 4096 # per request: body plus query string
@dataclass
class Session:
grants: list
labels: set = field(default_factory=lambda: {"public"})
sent: int = 0
budget: int = 65536 # per task, all destinations
TOOL_LABELS = {
"crm.get_customer": {"customer_pii"},
"docs.search": {"internal"},
"email.read_inbound": {"untrusted"},
"web.fetch": {"untrusted"},
}
def observe(s, tool):
# Unknown tools fail closed: their output is treated as untrusted.
s.labels |= TOOL_LABELS.get(tool, {"untrusted"})
def path_ok(pattern, path):
a, b = pattern.strip("/").split("/"), path.strip("/").split("/")
return len(a) == len(b) and all(x == "*" or x == y for x, y in zip(a, b))
def decide(s, method, url, body=b""):
u = urlsplit(url)
host = (u.hostname or "").rstrip(".").lower()
if u.scheme != "https":
return "deny", "not https"
size = len(body) + len(u.query.encode())
for g in s.grants:
if g.host != host or method not in g.methods or not path_ok(g.path, u.path):
continue
extra = {k for k, _ in parse_qsl(u.query, keep_blank_values=True)} - g.params
if extra:
return "deny", f"unexpected params {sorted(extra)}"
if size > g.max_bytes:
return "deny", f"{size} bytes over per-request cap {g.max_bytes}"
if s.sent + size > s.budget:
return "deny", "task byte budget exhausted"
uncleared = s.labels - g.classes - {"untrusted"}
if uncleared:
return "approve", f"context holds {sorted(uncleared)}, not cleared for {host}"
if "untrusted" in s.labels and method != "GET" and not g.accepts_untrusted:
return "approve", "write after untrusted input"
s.sent += size
return "allow", "grant matched"
if "untrusted" in s.labels:
return "deny", "no grant, untrusted content in context"
return "approve", "no grant for this destination"Four details carry most of the security value. path_ok matches whole segments, so /repos/acme/notes/issues cannot be stretched to a sibling repository, and a pattern star cannot swallow extra slashes the way a shell glob does. Unexpected query parameters are denied outright, because a new parameter is the cheapest place to smuggle data. Data-class clearance is checked against the session's labels, not against the request body, because a model can paraphrase, encode or split sensitive data so that no detector recognises it; the context is what you actually know. And writes after untrusted input go to a human unless the sink is explicitly marked as reviewed.
Reads are writes: query strings, caps and inspection
Scoping by method is not enough on its own, because GET requests carry data. Treat the query string as a body: count it against the byte cap, allowlist parameter names, and keep caps for search-style grants small. A documentation search rarely needs more than a couple of hundred bytes of query. A 256-byte cap does not make exfiltration impossible, but it turns a one-shot dump into a long series of requests the task budget will notice.
Content inspection on the outbound body is still worth running as a second signal: pattern detectors for card numbers and keys, high-entropy or base64-looking blobs in fields that normally hold prose, and canary tokens planted in sensitive records so that their appearance in any outbound request is unambiguous. Detectors are bypassed by trivial transformations, so inspection only escalates; labels make the main decision.
Worked example: an injected support ticket
A support-triage agent receives a ticket. Its grants: POST to api.github.com under /repos/acme/support-notes/issues, cleared for public, internal and customer_pii data and marked as accepting untrusted input because a support engineer reviews every note; and GET to docs-search.example.com/search with the single parameter q, cleared for public data only, capped at 256 bytes. The agent calls crm.get_customer, then email.read_inbound. The email contains an injected instruction: publish the account details as a gist, and if that fails, search the docs for them.
Running the broker on that sequence gives five decisions:
| Request | Decision | Reason |
|---|---|---|
| POST api.github.com/gists | deny | no grant, untrusted content in context |
| GET docs-search ?q= (600-byte customer dump) | deny | 602 bytes over per-request cap 256 |
| GET docs-search ?q=refund+policy | approve | context holds customer_pii, not cleared for the search host |
| GET docs-search ?q=x&ref=... (extra parameter) | deny | unexpected params ['ref'] |
| POST support-notes issue with the triage summary | allow | grant matched |
The gist attempt dies because no grant exists and untrusted content is present, so the broker does not even offer approval. The large query fails on size before anything else. The interesting row is the innocent-looking refund-policy search: it is blocked for approval because the session holds customer data and the search host is not cleared for it. That is the price of label-based control. You cannot distinguish an honest query from an encoded one, so you either clear the host for the data class, accept an approval, or restructure the task so that searching happens before the customer record is loaded. Restructuring is usually the right answer: plan the read-only research phase first, then load sensitive data and allow only sinks cleared for it.
The final row shows the legitimate path working without friction. The notes repository is cleared for customer data and tolerates untrusted-derived content because a human reads it before anything leaves the company.
Approvals that do not become rubber stamps
Approval is where many designs collapse into rubber-stamping. Three rules keep it useful. Show the human the payload, decoded and highlighted, not a summary written by the model under suspicion. Rate-limit approvals per task, so an agent cannot wear a reviewer down with a hundred near-identical requests; after a small number, the task halts. And make approvals scoped and short-lived: approving one request does not create a grant, and a reviewer who wants a standing capability must change the task template, which is reviewed like code.
For high-volume agents, push the decision earlier rather than adding reviewers. The CaMeL design from Google DeepMind, for example, has a privileged model write the plan before untrusted data is read and tracks the provenance of each value through execution, so a value derived from an email cannot silently become a URL. The principle transfers: decide destinations before the task reads anything it should not obey.
Rolling it out
Roll the broker out in shadow mode first. Wire it into the runtime, log every decision with the task type, grant matched, labels held and bytes, and enforce nothing. After a week you will have the real set of destinations per task type, which becomes the first draft of the grant templates. Expect surprises such as telemetry endpoints inside tool libraries.
Then enforce in stages: deny unknown destinations for tasks that have read untrusted content, then enforce byte caps, then data-class clearance. Track the approval rate per task type; a rate above a few percent means the grants are wrong or the task should be restructured. Match broker decisions against proxy logs; a mismatch in either direction is an incident. Investigation playbooks for confirmed leaks are covered in LLM data exfiltration.
Failure modes
- Grants derived from model output. If the plan or the model's tool arguments can create grants, the attacker can create grants. Issue them from task templates only.
- Labels reset between turns. Clearing labels when a conversation continues reopens every channel. Labels live as long as the context does, including summaries carried forward.
- Redirects followed silently. An allowed host that redirects to an arbitrary URL defeats host matching. Disable automatic redirects in the runtime's client and send each hop through the broker.
- Wildcard hosts. Every subdomain of a SaaS provider includes user-content subdomains anyone can register. Use exact hosts.
- Tools with their own clients. A tool that calls an SDK which opens its own connections skips the broker. The proxy should deny anything not originating from the runtime's identity.
- Budget per request only. Without a per-task budget, a determined injection splits the payload. Count bytes across the task.
Trade-offs
Label-based control is conservative by construction. It blocks honest requests whenever sensitive data is in context, and the more tools an agent has, the faster its labels saturate. The alternatives are worse in the ways that matter: content inspection is cheap to bypass, and per-workload allowlists ignore what the task has read. The practical balance is to keep tasks small, order them so research happens before sensitive reads, and spend approvals only on genuinely novel requests. Some capabilities, such as a general browser with customer-database access, cannot be made safe by egress policy alone; split them into separate tasks.
What to do next
- Make the tool runtime the only path to the network and confirm the proxy blocks everything else from the sandbox (network isolation for agents has a reachability test).
- List your agent's task types and write a grant template for each: exact host, methods, path pattern, parameters, data classes, byte cap.
- Map every tool to the labels its output carries; default unknown tools to untrusted.
- Deploy the broker in shadow mode for a week and compare its logs against proxy logs.
- Enforce in stages: unknown destinations after untrusted input, then byte caps, then data-class clearance.
- Disable automatic redirects in the runtime's HTTP client and broker every hop.
- Plant canary tokens in sensitive test records and alert on any outbound appearance.
- Review approval rates per task type monthly and restructure tasks whose rate stays high.