Most secret scanning grew up guarding source code. A pre-commit hook or a repository scan finds a credential someone typed, and the team rotates it. LLM applications add a new author that types fast, quotes anything in its context and sends output to places no human reviews: tool calls, files, markdown renderers, logs. A credential that reaches an assistant's context through a tool result or a retrieved wiki page can come out in an answer a minute later. A credential memorised from public code can come out with no context at all.
This article is about the output side: where secrets in model output come from, every channel they can leave through, and how to build an output gate that uses provenance as well as pattern matching. The gate should know which strings were secrets on the way in, so it catches them on the way out in formats no regex anticipates. Detector anatomy and streaming mechanics are covered in the secret scanners deep dive; here we build on them and focus on provenance, policy and measurement.
Four origins of a secret in output
A secret in an output has one of four origins, and each origin calls for a different control:
- Context echo. The secret was in the prompt: a tool result (an agent ran
env, read a.envfile or printed a Kubernetes secret), a retrieved document (a runbook with a pasted token), or a user paste. The model is doing its job, which is using the context. This is the most common origin, and the one you can match exactly. - Memorisation. Pretraining corpora include public code, and public code includes hard-coded credentials. Published studies of code completion tools have extracted strings matching real credentials from public repositories. These may belong to third parties and may still be live.
- Plausible fabrication. The model produces something with a valid format and no real secret behind it, often in a code sample. It is usually harmless, but a format-only detector cannot tell it from the first two cases.
- Documented examples. Strings published as placeholders, such as AWS's
AKIAIOSFODNN7EXAMPLEaccess key ID from its documentation. Models emit them constantly. Allowlist them, or alert fatigue will bury the real findings.
Notice that only the first origin has a record of the secret's value inside your system. That record is the most useful signal you have, and most deployments throw it away.
Every channel an output can leave through
Teams scan the chat bubble and forget the rest. List every channel an output can reach:
| Channel | How a secret escapes | Who sees it |
|---|---|---|
| Chat display | Quoted verbatim in an answer | The user, and screenshots of it |
| Rendered markdown | Embedded in an image or link URL; the client fetches it automatically | Whoever controls the URL's host |
| Tool-call arguments | Placed in an HTTP request, shell command or email the agent sends | The tool's destination |
| File writes and commits | Hard-coded into generated code that gets committed | Everyone with repository access, forever |
| Logs, traces, eval sets | Full prompts and completions stored for debugging | Anyone with observability access |
The markdown row is the one attackers use. A prompt injection in a retrieved page tells the model to render an image whose URL carries the secret as a query parameter. The chat client fetches the image and the secret leaves without a click, as described in LLM data exfiltration. Tool arguments are the agent version of the same attack. Both channels must be scanned before execution or rendering. Scanning the transcript afterwards only tells you when the leak happened.
Format detectors and their blind spot
Format detectors are the baseline. A good rule has a fixed vendor prefix, a known length and character set, and ideally a checksum. GitHub's classic personal access tokens start with ghp_ and fine-grained ones with github_pat_. Slack bot tokens start with xoxb-, Stripe live secret keys with sk_live_ and AWS access key IDs with AKIA. Prefix rules are precise. Generic rules, such as high-entropy strings next to words like password or secret, catch more but cost precision. Open-source rule sets from Gitleaks, TruffleHog and detect-secrets are a better starting point than writing your own. TruffleHog can also verify whether a found credential is live, which is the strongest signal there is. Run that check after the fact on an alert, never inline. It is slow, and it sends the secret to the issuer.
Format detectors have a blind spot that matters here: they only see secrets that look like secrets. A database password, an internal service token or a webhook URL with an embedded key has no prefix. Neither does a known secret after base64 encoding or with spaces between characters. Provenance closes that gap.
Provenance: a taint registry
The idea is simple. Anything recognised as secret on the way into the context gets fingerprinted, and outputs are checked against the fingerprints. Inputs come from places you control: you can scan tool results and retrieved chunks with the same detectors, and you can pull the exact values of secrets the application holds from its own secret store. The registry stores keyed hashes, never plaintext, so the registry is not a second copy of your secrets.
import base64, hashlib, hmac, re, urllib.parse
CANDIDATE = re.compile(r"[^\s<>()\[\]{}`,;\x22\x27]{12,}") # no spaces, quotes or brackets
class TaintRegistry:
def __init__(self, key: bytes):
self.key = key
self.by_len = {} # length -> {digest: label}
def _h(self, s: str) -> bytes:
return hmac.new(self.key, s.encode(), hashlib.sha256).digest()
def register(self, value: str, label: str):
variants = {value,
base64.b64encode(value.encode()).decode().rstrip("="),
value.encode().hex(),
urllib.parse.quote(value, safe="")}
for v in variants:
self.by_len.setdefault(len(v), {})[self._h(v)] = label
def matches(self, text: str):
squeezed = re.sub(r"[\s\u200b]+", "", text) # defeats spaced-out spelling
hits = []
for tok in CANDIDATE.findall(text) + CANDIDATE.findall(squeezed):
for n, table in self.by_len.items():
for i in range(0, len(tok) - n + 1): # secret inside a longer token
label = table.get(self._h(tok[i:i + n]))
if label:
hits.append((label, i))
return hits
# On ingestion of any tool result or retrieved chunk:
for finding in run_detectors(chunk_text):
registry.register(finding.value, f"{finding.rule}@{chunk_id}")
# At startup, for secrets the app itself holds:
for name, value in vault.list_app_secrets():
registry.register(value, f"vault:{name}")The sliding window costs one HMAC per character per registered length, which is cheap for output-sized text. Cap candidate tokens at a few hundred characters and limit the registry to lengths that actually occur. The HMAC key must live outside the logs and traces the registry protects. Scope the registry per session or per tenant, so one user's secrets never become another user's match signal. Matching catches context echo in any format you registered, including secrets that have no recognisable format. It does not catch memorised secrets, which never passed through your input, so format detectors stay on as the second signal.
Policy by origin and channel
One action for every finding is wrong in both directions. Blocking every format match breaks code-generation features full of fabricated example keys. Redacting display text while allowing tool arguments leaves the exfiltration channel open. Decide by origin and channel together:
def decide(signal, channel):
"""signal: 'taint' | 'format' | 'example'; channel: display|markdown_url|tool_arg|file_write|log"""
if signal == "example":
return "pass"
if signal == "taint":
if channel in ("markdown_url", "tool_arg"):
return "block_and_alert" # a known secret heading off-box
return "redact_and_alert" # replace with [REDACTED:<label>]
# format-only match: could be memorised, could be fabricated
if channel in ("markdown_url", "tool_arg", "file_write"):
return "redact_and_alert"
return "redact" # display and log: alert only if verified live laterRedact with a typed placeholder such as [REDACTED:slack_bot_token] rather than deleting the string, so the user understands why the answer looks odd. Every alert should carry the taint label, because the label names the document or tool call where the secret came in. The real fix happens there: rotate the credential and remove it from the source. An output gate that redacts the same wiki token every day without anyone rotating it is hiding the problem, not solving it.
Streaming responses need a hold-back buffer. Keep back the last N characters of output, where N is the longest registered variant (hex doubles the length and URL-encoding can triple it), and release text only once no match could still begin inside it. Run the gate on assembled tool-call arguments before dispatch, and on markdown URLs before rendering. A domain allowlist for images closes the zero-click path entirely, and costs less than scanning.
Worked example: a token in the runbook
Take an internal support assistant with retrieval over the engineering wiki. Two years ago, someone pasted a working Slack bot token into a deploy runbook's curl example. A user asks how to post deploy notices to the release channel. Trace it through the gate:
- Retrieval returns the runbook chunk. The ingestion scan matches the
xoxb-rule and registers the value with labelslack_bot_token@wiki:deploy-runbook#4. - The model answers with the curl command, token included. The hold-back buffer has the full token before it reaches the client. Both the taint match and the format rule fire. The channel is display, so the policy redacts and alerts.
- The alert reaches the security channel with the label. On-call verifies the token is live, rotates it, edits the runbook and opens a ticket to scan the wiki index at ingestion time.
- A week later another page carries a prompt injection telling the model to base64-encode any token it sees and put it in an image URL. Format rules see nothing, but the registry holds the base64 variant and the channel is
markdown_url, so the gate blocks it. The image domain allowlist would have refused it anyway, which is the point of having two layers.
Notice what stopped the second attack: provenance, not pattern. Notice too what would have prevented the incident outright: scanning the wiki before indexing, so the token never reached the context.
Measuring the gate with canaries
A gate you have not measured is a gate you are only hoping works. Build a test set of canary secrets: synthetic values in valid vendor formats plus format-less ones such as random passwords. Plant them in tool results, retrieved documents and user turns. Then run an attack suite that asks the model to reveal them verbatim, spaced out, base64-encoded, hex-encoded, URL-encoded, reversed, split across two turns, spelled as words and embedded in generated code. Score the catch rate per channel and per transformation, and run the suite in CI whenever the model, system prompt or gate changes.
Expect the reversed, split-turn and spelled-out cases to get through a registry that only knows four encodings. That residual is real, and the suite exists to make it visible. Answer it with architecture, not more regexes: keep secrets out of the context in the first place, using short-lived credentials held by tools rather than by the model, as described in secret management for agents. Track false positives too. If more than a small share of alerts are documented examples or fabricated keys, fix the allowlist before people stop reading the alerts.
Failure modes
- Scanning only the chat bubble. Tool arguments and markdown URLs leak unobserved. Gate every channel before execution.
- Plaintext registry. The fingerprint store becomes a secret store with weaker controls. Store keyed hashes only.
- Global registry. Tenant A's secret appears in tenant B's match alerts. Scope per session or tenant.
- Inline verification. Calling the issuer for every match adds latency and sends secrets to third parties from the hot path. Verify asynchronously on alerts.
- Logs before the gate. Request logging captures raw completions upstream of redaction. Put the gate before every sink, including tracing.
- No hold-back on streams. The first half of a token reaches the client before the detector sees the second half.
- Redact without rotate. The same credential is caught daily and never revoked.
- Alert fatigue. Documented examples and fabricated keys bury real findings. Allowlist and tune.
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Signal | Format detectors: catch memorised and unknown secrets; miss encoded and format-less ones | Taint registry: exact and encoding-aware; only sees what entered the context |
| Action on display | Redact: answer survives, may confuse | Block: safe, breaks legitimate code samples |
| Streaming | Hold-back buffer: small latency, correct | Post-hoc scan: no latency, leaks partial tokens |
| Verification | Async on alert: accurate triage | None: cheaper, more noise |
| Markdown images | Domain allowlist: closes the channel | Scan URLs: keeps flexibility, leaves a gap |
What to do next
- List every egress channel your application has, including tool calls, file writes, rendered markdown, logs and eval exports, and confirm a gate sits before each one.
- Scan tool results and retrieved chunks at ingestion with the same rules you use on repositories (see secrets scanning architecture).
- Add a per-session taint registry of keyed hashes with base64, hex and URL-encoded variants, and match outputs against it.
- Write the origin-by-channel policy down, with block for known secrets in tool arguments and markdown URLs.
- Add a hold-back buffer to streaming and an image-domain allowlist to the client.
- Build the canary and attack suite, run it in CI and publish catch rates per channel.
- For every alert, rotate the credential and fix its source, then read PII leakage for the memorisation side of the problem.