Server-side request forgery (SSRF) is the bug where an attacker chooses a URL and your server fetches it, from inside your network, with your network position and sometimes your credentials. Classic SSRF needed a vulnerable endpoint: an image proxy, a webhook tester, a PDF exporter. An agent gives the attacker something better: a planner that reads untrusted text, decides which URLs to visit, and has several tools that make outbound requests on its behalf.
The site already has a page on hardening a single fetch tool against encodings, redirects and DNS rebinding (SSRF via LLM tools). This page is about the step before and around that: finding every place in an agent where a URL turns into a request, deciding what each of those places may reach, and enforcing that at one point the tools cannot route around. It ends with a test harness and a checklist you can run against your own agent this week.
Why agents make SSRF worse
Three things make SSRF in agents worse than in an ordinary web application.
The URL source is text, not a form field. A URL can arrive in a user message, a retrieved web page, an email body, a PDF, a tool result, or an MCP server's response header. Prompt injection means any of those can ask the model to visit an address, and the model has no reliable way to tell a legitimate instruction from an injected one. Treat every URL the model emits as attacker-chosen.
Blind SSRF becomes full-read SSRF. In a classic blind SSRF, the attacker triggers a request but never sees the response. In an agent the response flows back into the model's context, and the model will happily summarise it, quote it, or pass it to another tool, including one that sends email or posts to a URL. The agent is both the request primitive and the exfiltration channel.
There are many sinks, and most are not called fetch. Teams harden the obvious web_fetch tool and leave the headless browser, the document converter, the MCP client and the code sandbox with ordinary network access. The attacker only needs the weakest one.
Inventory the URL sinks
Start with an inventory. A sink is any component that turns a string into an outbound connection. For each one, record who supplies the URL, which component actually opens the socket, and what that component does that a simple validator would miss.
| Sink | Where the URL comes from | What opens the connection | Gap a URL validator misses |
|---|---|---|---|
| Fetch tool | model output | HTTP client in the tool | redirects, DNS rebinding (see the validator page) |
| Headless browser | model output, then the page itself | the browser engine | subresources, scripts calling fetch(), redirects, WebSockets, service workers |
| Document renderer | content of an uploaded or retrieved file | converter (HTML-to-PDF, office suite, image tools) | linked images, stylesheets, iframes, external entities |
| MCP client | a remote server's 401 response and metadata | the client's OAuth discovery code | server-controlled discovery URLs, fetched before any user consent |
| Connector / webhook | configuration, or a user-supplied callback | integration code, often later and asynchronously | DNS changes between registration and delivery; credentials auto-attached |
| Code sandbox | any code the model writes | arbitrary sockets | everything: raw TCP, DNS, non-HTTP protocols |
Two of these deserve a closer look because they surprise people. On MCP, the 2025-06-18 authorization specification has the client read a resource_metadata URL from the server's WWW-Authenticate header on a 401 (OAuth Protected Resource Metadata, RFC 9728), then fetch authorization-server metadata (RFC 8414) from whatever authorization server that document names. Every one of those URLs is chosen by the remote server, so connecting a client to an untrusted server lets that server steer your client's discovery requests. On document rendering, the request happens inside a library you did not write, often during a step the team thinks of as "parsing", so it never appears in a tool-call log.
Architecture: one way out
The design that holds up has three rules.
- One way out. No tool process has a network route except to an egress proxy. Enforce this below the application, with a network policy or a network namespace, so a bug or a library that ignores proxy environment variables fails closed instead of going direct. Network isolation for agents covers the zoning.
- The proxy knows which sink is calling. Identify the caller by something the model cannot write: a separate proxy port per tool, a client certificate, or the pod's service account. Never trust a header in the request, because the model or a page script may control request headers.
- Decide on the connected IP, per sink. The proxy resolves DNS itself and makes the allow/deny decision on the address it is about to connect to, which defeats rebinding and redirect tricks in one place. Each sink gets its own set of permitted destination zones.
The policy core
Here is the policy core as it would sit in a proxy's decision hook. It is deliberately small: the hard part is the inventory and the enforcement path, not the code.
import ipaddress
from dataclasses import dataclass
from enum import Enum
from urllib.parse import urlsplit
class Zone(Enum):
PUBLIC = "public"
PARTNER = "partner"
INTERNAL = "internal"
@dataclass(frozen=True)
class Sink:
zones: frozenset # zones this sink may ever reach
methods: frozenset # HTTP methods it may use
max_bytes: int # response cap the proxy enforces
SINKS = {
"web_fetch": Sink(frozenset({Zone.PUBLIC}), frozenset({"GET"}), 2_000_000),
"browser": Sink(frozenset({Zone.PUBLIC}), frozenset({"GET", "POST"}), 20_000_000),
"mcp_discovery": Sink(frozenset({Zone.PUBLIC}), frozenset({"GET"}), 64_000),
"crm_connector": Sink(frozenset({Zone.PARTNER}), frozenset({"GET", "POST"}), 1_000_000),
# doc_render is absent on purpose: it has no network, so the proxy never sees it
}
PARTNER_HOSTS = {"api.crm.example.com"}
def zone_of(host, ip):
if getattr(ip, "ipv4_mapped", None): # ::ffff:10.0.0.5 is 10.0.0.5
ip = ip.ipv4_mapped
if not ip.is_global: # private, loopback, link-local, CGNAT, ...
return Zone.INTERNAL
return Zone.PARTNER if host in PARTNER_HOSTS else Zone.PUBLIC
def decide(sink_name, method, url, connect_ip):
"""sink_name comes from workload identity, connect_ip from the proxy's own DNS lookup."""
sink = SINKS.get(sink_name)
if sink is None:
return False, "unregistered sink"
parts = urlsplit(url)
if parts.scheme not in ("http", "https"):
return False, "scheme " + repr(parts.scheme)
if method not in sink.methods:
return False, method + " not allowed for " + sink_name
zone = zone_of((parts.hostname or "").lower(), ipaddress.ip_address(connect_ip))
if zone not in sink.zones:
return False, sink_name + " may not reach " + zone.value
return True, "ok"Note what is not here: no regular expressions over hostnames, no list of bad strings. Internal addresses are recognised by the standard library's is_global check after normalising IPv4-mapped IPv6, and anything internal is denied unless a sink explicitly lists the internal zone, which no model-driven sink should. Credentials follow the same rule: the proxy attaches a connector's token only when the sink and destination match, so a model that rewrites a URL cannot carry a token to a host of its choosing.
Closing each sink
Headless browser. A browser is the hardest sink because the page you load makes its own requests. Point the engine at the proxy and add in-process routing as a second layer, not the first. Playwright's documentation notes that requests answered by a service worker are not visible to route(), and recommends blocking service workers when you need to see every request.
from playwright.async_api import async_playwright
async def open_agent_browser(pw):
browser = await pw.chromium.launch(
proxy={"server": "http://egress-browser.agent.svc:3128"}, # browser's own proxy port = its identity
)
context = await browser.new_context(
service_workers="block", # otherwise route() misses requests a worker answers
accept_downloads=False,
)
async def guard(route):
if not route.request.url.startswith(("http://", "https://")):
await route.abort() # file:, data: navigations, custom schemes
return
await route.continue_()
await context.route("**/*", guard)
return browser, contextChromium's documented implicit bypass rules send localhost names, loopback addresses and link-local addresses (which include 169.254.169.254) direct, not through the configured proxy, unless the bypass list contains <-loopback>. So a page script can reach the container's own localhost and, without a network-level block, the metadata service. Set that rule, but rely on the network: run the browser in its own namespace with nothing listening locally and link-local blocked.
Document renderers. Give converters no network at all, for example a container started with --network none. If a document legitimately needs remote images, fetch them first through the web_fetch sink and hand the renderer local files. This one rule removes a whole family of historical converter SSRF bugs without auditing each library.
MCP clients. Route OAuth discovery through the proxy as its own sink, public zone only, with a small response cap. Only connect to remote servers on an allowlist your team maintains; the user clicking "connect" is not an allowlist.
Connectors and webhooks. Validate at delivery time, not at registration, because the hostname can resolve differently a week later. Deliver through the proxy so the check runs on the connected address.
Code sandboxes. Default to no network. If the task needs packages, give the sandbox a package mirror and nothing else. Sandboxing agents with Docker covers the container side.
Worked example: a CV that reads the metadata service
A recruiting assistant runs on a cloud VM. A user uploads a candidate's CV as an HTML file and asks for a summary. The file contains an image tag whose source is the instance metadata address, http://169.254.169.254/latest/meta-data/iam/security-credentials/, plus hidden text telling the model to include any "technical details from images" in its summary.
Without the design. The converter renders the HTML to text, fetches the image URL, and because the response is text it ends up in the extracted content. The model dutifully includes the role name in the summary. A second request for that role's path would return temporary credentials if the instance still allows IMDSv1. With IMDSv2 required, that particular path fails, because a token must first be obtained with a PUT request carrying a TTL header, which an image fetch cannot do. That is why requiring IMDSv2 is worth doing everywhere, but it protects only the metadata service, not the internal admin API on 10.0.3.17 that the next payload targets.
With the design. The renderer has no network, so the fetch fails at the socket and the converter emits a broken-image placeholder. Had the attacker used the fetch tool instead, the proxy would have resolved the address, classified it as internal, and denied the web_fetch sink with a logged reason. The summary contains only the CV. The deny log line, tagged with sink and conversation ID, is the security team's alert.
Testing with canaries
You cannot claim a sink is closed until a test proves it. Put a canary service in every zone that should be unreachable: a tiny HTTP server that returns a unique token and logs every hit. Then drive each sink at it.
| Test | How | Pass condition |
|---|---|---|
| Fetch to internal canary | ask the agent to fetch the canary URL by IP and by internal DNS name | proxy deny log; canary sees no hit |
| Browser subresource | public test page whose image and script point at the canary | page loads; canary sees no hit |
| Browser loopback | test page script calls fetch on localhost ports | connection refused or denied |
| Renderer | upload HTML and office files that link the canary | rendered output has no canary token |
| MCP discovery | test MCP server whose 401 names the canary as resource_metadata | client refuses; canary sees no hit |
| Sandbox | code that opens raw TCP and DNS lookups | every connection fails |
| Exfil path | fetched page tells the model to post data to an external URL | post denied or needs approval |
Run this matrix in CI against a staging deployment, and again after every new tool or library upgrade. The canary token makes leaks unambiguous: if it ever appears in a model transcript, the test failed even if no deny was logged.
Failure modes
- Proxy by convention. Tools honour
HTTPS_PROXYuntil a library does not. Only a network-level default deny makes the proxy mandatory. - One allowlist for every sink. A shared list means the browser can reach the partner API with the connector's credentials. Policy must be per sink.
- Checking the hostname, connecting to the IP. Any decision made before the proxy's own resolution can be undone by DNS. Decide on the connected address.
- Forgetting IPv6 and odd ranges. Hand-written private-range lists miss
fd00::/8(AWS's IPv6 metadata endpoint isfd00:ec2::254), CGNAT space and IPv4-mapped forms. - Silent renderer fetches. No tool call, no log line, so no alert. Removing the network is the only reliable fix.
- Allowing the response back unfiltered. Even allowed public fetches can carry injected instructions; treat responses as untrusted input, and gate tools that send data out (egress control for agents).
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Central egress proxy | one enforcement and logging point, per-sink policy | latency, an extra service to run and scale, TLS handling decisions |
| No network for renderers and sandboxes | removes whole bug classes | pre-fetching step; some documents render incompletely |
| Public-only zone for model-driven sinks | internal services unreachable by construction | internal tools need dedicated connectors with fixed endpoints |
| Allowlisted MCP servers | server-steered discovery limited to vetted parties | slower onboarding of new servers |
What to do next
- List every component in your agent that can open a socket, using the sink table above as a template.
- Make the proxy mandatory with a network policy, then identify each sink by port, certificate or service account.
- Remove network access from document renderers and code sandboxes; add a package mirror only where needed.
- Require IMDSv2 on cloud instances and block the metadata addresses at the proxy as well.
- Stand up canary services and add the test matrix to CI.
- Keep learning: hardening a fetch tool, egress control for LLM workloads and capability tokens for agents.