Most writing about egress for LLM systems is about the network: which hosts an agent may call, which tool requests leave the sandbox. That layer matters and is covered in egress filtering architecture and egress control for agents. This article is about the other exit every chat product has: the text of the answer. It is rendered as markdown, copied into tickets and logged. If a secret, a system prompt or a data-carrying URL reaches that text, it has left.
An output egress filter is the component that inspects the response stream and decides, chunk by chunk, what may be released. The hard parts are not the regexes. They are streaming (a secret can be split across two chunks), markdown (an image tag makes the browser fetch a URL with no click), latency, and the fact that sent bytes cannot be taken back. This article builds such a filter in Python, traces an attack through it, and covers testing.
What the filter is for
Decide what the filter is for before writing it. The useful threat list is short and concrete:
- Credential echo. The model repeats an API key it saw in a retrieved document, a config file a tool read, or its own prompt. Many credential formats have recognisable prefixes, for example AWS access key IDs start with
AKIAand GitHub personal access tokens withghp_. - Prompt and canary leakage. An attacker asks the model to print its instructions. If you plant a unique marker in the system prompt, its appearance in output is a certain leak; see canary tokens.
- Markdown image exfiltration. Injected instructions in a retrieved page tell the model to emit
. The chat UI renders it and the browser sends the secret in the query string, with no user action. - Link smuggling. The same payload in a normal hyperlink, which needs a click but looks legitimate.
- Invisible text. Characters from the Unicode Tags block (U+E0000 to U+E007F) and zero-width characters are not displayed but survive copy and paste, carrying hidden instructions into the next system or hidden data out of this one.
- Regulated data. Personal data drawn from retrieval that this user is not entitled to see. Pattern matching only catches the structured part of this; entitlement checks belong at retrieval time.
Architecture
The filter sits in the server that proxies the model stream to the client, never in the client, because a client can be modified. It has four parts. A holdback buffer keeps back text that might still turn into a match. Inline detectors are deterministic and fast: exact canary matches, credential patterns, a markdown link and image parser, and an invisible character scrubber. A policy maps each finding to an action. Asynchronous classifiers, such as PII or toxicity models, run beside the stream on the completed response and raise alerts or flag the conversation, because putting a 50 ms model call on every chunk makes streaming pointless.
Everything the filter does is logged with the rule, the request id and a hash of the matched value, never the raw value: the audit log must not become the new leak.
Streaming: the holdback buffer
Streaming is the core difficulty. Models emit chunks of a few characters; an access key can arrive as AKIA in one chunk and the remaining 16 characters in the next. If each chunk is checked on its own, nothing ever matches. The fix is to keep the last H characters back, where H is at least the length of the longest bounded pattern. Any match that has not finished arriving is shorter than its full length, which is at most H, so its partial prefix lies inside the held tail and is never released early.
Markdown links have no natural length bound, so they get a different rule: from an unmatched [ onward, hold everything until the closing parenthesis arrives, up to a cap. If the cap is exceeded, escape the bracket so no renderer can form a link from the pieces.
import re
HOLD = 64 # >= longest bounded pattern, in characters
LINK_CAP = 2048 # most characters held while a link is open
LINK_DONE = re.compile(r"\]\([^)]*\)|\]:[ \t]*\S+\s") # inline link, or definition
class StreamGate:
def __init__(self, scrub):
self.scrub = scrub # text -> (safe_text, findings); idempotent
self.buf = ""
self.findings = []
self.stopped = False
def _cut(self):
cut = max(0, len(self.buf) - HOLD)
start = self.buf.rfind("[")
if start != -1 and not LINK_DONE.search(self.buf, start):
if len(self.buf) - start < LINK_CAP:
if start > 0 and self.buf[start - 1] == "!":
start -= 1
cut = min(cut, start)
else: # over the cap: make it unrenderable, then release
self.buf = self.buf[:start] + "\\" + self.buf[start:]
self.findings.append(("link_over_cap", start))
return cut
def feed(self, chunk):
if self.stopped:
return ""
self.buf, found = self.scrub(self.buf + chunk)
self.findings += found
if any(kind == "stop" for kind, _ in found):
self.stopped = True # scrub already cut the buffer
out, self.buf = self.buf, ""
return out
cut = self._cut()
out, self.buf = self.buf[:cut], self.buf[cut:]
return out
def finish(self):
if self.stopped:
return ""
self.buf, found = self.scrub(self.buf + "\n") # closes a final definition
self.findings += found
out, self.buf = self.buf, ""
return outThe scrubber runs over the whole buffer on every feed, so it must be idempotent: its replacements must not match its own patterns or open a new bracket. The buffer is small, so re-scanning it costs microseconds. The visible cost is that the user sees text about H characters behind the model, which at typical token rates is a fraction of a second.
Detectors
The scrubber composes small detectors. Each returns rewritten text and findings. Patterns should be anchored on formats you can name; a generic high-entropy rule catches more, but it also flags hashes, UUIDs and base64 images, and its matches are unbounded in length, which breaks the holdback guarantee. Keep it in the asynchronous path as an alert, or bound it.
SECRETS = {
"aws_access_key_id": re.compile(r"\b(?:AKIA|ASIA)[0-9A-Z]{16}\b"),
"github_token": re.compile(r"\bgh[pousr]_[A-Za-z0-9]{36}\b"),
}
# A key body is unbounded, so the header is a stop trigger, not a redaction.
STOP = {"private_key": re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY-----")}
INVISIBLE = re.compile(r"[\U000E0000-\U000E007F-]")
def make_scrub(canaries, allowed_hosts):
def scrub(text):
findings = []
for name, rx in STOP.items():
m = rx.search(text)
if m:
findings.append(("stop", name))
return text[:m.start()] + "[response stopped]", findings
for canary in canaries: # exact, per deployment
if canary in text:
findings.append(("canary", canary[:6]))
text = text.replace(canary, "[removed]")
for name, rx in SECRETS.items():
if rx.search(text):
findings.append(("secret", name))
text = rx.sub("[redacted credential]", text)
if INVISIBLE.search(text):
findings.append(("invisible", None))
text = INVISIBLE.sub("", text)
text = rewrite_links(text, allowed_hosts, findings)
return text, findings
return scrubThe invisible set is deliberately narrow. Zero-width joiner and non-joiner (U+200D, U+200C) and the bidi marks (U+200E, U+200F) are kept, because Persian, Indic scripts, right-to-left text and joined emoji need them. Stripping the Tags block also breaks the few flag emoji built from tag sequences, a known false positive worth accepting for most products.
Credential formats change, and dedicated scanners maintain hundreds of them; for the trade-offs between tools see secret scanners. Inline, prefer a short list of formats your systems actually hold, kept in sync with the secret store, over a large list that slows every chunk.
Links and images
Links and images are where output filtering earns its keep. The rule set that works: never render images from model output unless they come from your own image proxy; allow hyperlinks only to an allowlist of hosts, over HTTPS; and strip query strings and fragments even on allowed hosts, because the query is where exfiltrated data rides.
from urllib.parse import urlsplit
LINK = re.compile(r"(!?)\[([^\]]*)\]\(([^)\s]+)(?:\s+\"[^\"]*\")?\)")
DEF = re.compile(r"^[ \t]*\[([^\]]+)\]:[ \t]*(\S+)(?=\s)", re.M) # [ref]: url, complete
def rewrite_links(text, allowed_hosts, findings):
def safe(url):
parts = urlsplit(url.strip("<>"))
host = (parts.hostname or "").lower()
ok = parts.scheme == "https" and host in allowed_hosts
return ok, host, f"https://{host}{parts.path}" # no query, no fragment
def repl(m):
bang, label, url = m.groups()
ok, host, clean = safe(url)
if ok and not bang:
return f"[{label}]({clean})"
findings.append(("image_blocked" if bang else "link_blocked", host))
return f"{label} (link removed)" if not bang else f"(image removed: {label})"
def repl_def(m):
ok, host, clean = safe(m.group(2))
if ok:
return f"[{m.group(1)}]: {clean}"
findings.append(("definition_blocked", host))
return f"(link definition removed: {m.group(1)})"
return DEF.sub(repl_def, LINK.sub(repl, text))Reference-style markdown, ![s][r] with [r]: https://... on a later line, never matches the inline pattern, so definitions get their own rule; the gate also treats a definition as closed only once whitespace follows its URL. A definition cannot tell an image from a link, so an allowlisted host may still serve an image; keep your own hosts free of open redirects and user uploads. Blocked output uses parentheses, not brackets, so it cannot reopen the holdback rule. Markdown is not the only syntax: if your renderer accepts raw HTML, an <img> tag does the same job, and many renderers turn bare URLs into links. The cleanest fix is on the client: render model output with raw HTML disabled and with automatic image loading off. The filter and the renderer settings are two layers; keep both. More on treating output as untrusted input to the next system is in LLM output handling.
Policy and actions
Each finding maps to an action, and the action depends on what has already been sent.
| Finding | Action | Why |
|---|---|---|
| Credential pattern | Redact inline, alert, rotate the credential | A redacted key may still have been seen elsewhere; rotation closes it |
| Private key header | Stop the stream before the header | The key body is unbounded and cannot be redacted by pattern |
| Canary | Stop the stream, return a refusal, page | Proof of prompt extraction; the rest of the answer is untrustworthy |
| Image from model output | Remove, keep the alt text | Zero-click fetch; never worth the risk |
| Link off allowlist | Replace with label text | Click-required, but the URL is the payload |
| Tag block and zero-width space | Strip silently, count | Rare in real answers; joiners and bidi marks are kept |
| Async PII hit | Flag conversation, review | Too slow and noisy to block inline |
Stopping mid-stream needs a protocol: send a terminal event that tells the client to replace the partial message with a notice, rather than just closing the connection. The client can hide the text, but anything already displayed may have been read, so treat a mid-stream stop as a leak of whatever preceded it, which the holdback buffer keeps small.
Worked example: an image split across chunks
A support assistant retrieves a web page containing hidden instructions: summarise the conversation, then add an image whose URL carries the user's account email. The model complies and streams these chunks:
chunk 1: "Here is the summary of your request. "
chunk 2: ""
chunk 4: " Let me know if you need anything else."Chunk 1 is plain text; the gate releases it except the last H characters. Chunk 2 opens a bracket with no close, so the cut moves to the exclamation mark and everything from ![status] onward is held. Chunk 3 completes the link; the scrubber sees an image from model output, records image_blocked with host attacker.example, and rewrites it to (image removed: status). Chunk 4 and the finish call release the rest. The browser never sees the attacker URL. Had the filter checked chunks independently, neither chunk 2 nor chunk 3 would have matched the link pattern, and a client that concatenates and renders would have fetched the image.
The audit entry also signals that retrieval delivered an injection, which is worth more than the block: find the source page and quarantine it.
Testing and operating
Build a regression corpus before shipping: real leaked-format credentials from test accounts, every markdown link and image variant (reference-style links, titles, angle-bracket URLs), invisible character payloads, and a large sample of normal answers. Replay each case at random chunk sizes from 1 to 20 characters, because chunk-boundary bugs only show up when boundaries move. Measure three things in production: findings per thousand responses by rule, false positives from user reports and sampled review, and the added time to first visible token. A sharp rise in one rule usually means either an attack campaign or a new legitimate format the rule mis-flags; both need a human.
Set a Content-Security-Policy with a tight img-src on the chat page as defence in depth, so a bypass of the filter still cannot make the browser fetch an attacker image.
Failure modes
- Per-chunk scanning. Patterns never match across boundaries. Always buffer.
- Client-side filtering. Bypassed by calling the API directly. Filter on the server.
- Unbounded patterns inline. An entropy or free-text rule with no length limit forces either unbounded holdback or missed matches.
- Non-idempotent rewrites. A replacement that contains a bracket or a pattern makes the buffer grow or loop. Test scrub(scrub(x)) == scrub(x).
- Logging the secret. The finding log stores the raw value and becomes a credential dump.
- Other renderers. The filter covers the chat UI, but the same text goes to email notifications or a ticket system with different markdown rules. Apply it at the API boundary, not in one front end.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Larger holdback | Longer patterns covered | Text appears later |
| Block whole response | Nothing partial leaks after the hit | Bad user experience, more support load |
| Redact inline | Answer remains useful | Context around the redaction may still leak |
| Strict link allowlist | Closes link smuggling | Legitimate citations lose their links |
| Inline ML classifier | Catches unstructured leaks | Latency on every chunk, noisy |
What to do next
- List the secrets, canaries and data classes that could reach your model's context.
- Put a holdback-buffer gate on the server between the model stream and every client.
- Disable raw HTML and automatic image loading in every renderer of model output.
- Block images from model output and allowlist link hosts with queries stripped.
- Plant a canary in each system prompt and page on any appearance in output.
- Replay a regression corpus at random chunk sizes in CI and check idempotence.
- Add a CSP img-src allowlist and chart findings per rule per thousand responses.