Giving a whole workforce an AI assistant is not one security decision; it is a dozen, and most of them are made by default. The assistant reads whatever its connectors can reach, employees paste whatever they are working on, the vendor keeps whatever its contract allows, and the people who were not given a sanctioned tool use an unsanctioned one from their phone. The policy side of this, which tools are allowed for which data, is covered in our article on employee AI usage policies. This article is the technical side: the architecture that makes a policy enforceable, the data flows that actually leak, and the code that sits in the middle.
The core idea is simple. There should be exactly one path from an employee to a model, and that path should know who the employee is, what the prompt contains, which documents the answer drew on and how long each of those facts is kept. Everything else in this article is a consequence of building that path and then measuring who goes around it.
Five flows, five failure shapes
Start with the flows, because each one fails differently.
| Flow | What goes wrong | Primary control |
|---|---|---|
| Employee to model (prompts) | Source code, customer records or deal terms leave for a vendor with unclear retention | Gateway classification, contracted no-training terms |
| Index to employee (retrieval) | The assistant surfaces files the user could technically open but was never meant to find | Permission hygiene, sensitivity labels, query-time ACL checks |
| Model to systems (actions) | An agent sends mail, edits tickets or shares files because a document told it to | Scoped connectors, confirmation on writes, egress limits |
| Logs to everyone | The prompt archive becomes the most sensitive dataset in the company | Retention classes, access logging, hashing |
| Employee to unsanctioned tool | Shadow AI: personal accounts, browser extensions, meeting bots | Discovery from proxy and DNS, a good sanctioned option |
Notice that only the first row is what people usually mean by AI data leakage. In deployments that connect an assistant to mail, chat and file shares, the second row is a common first real incident, and it is not caused by the model at all.
Reference architecture
The gateway is a small reverse proxy in front of every model endpoint the company pays for. Clients never hold vendor API keys; they hold a short-lived token from the identity provider, and the gateway exchanges it for the vendor call. That single design choice gives you four things at once: you can revoke a user without rotating a vendor key, you know which human sent each prompt, you can inspect content before it leaves, and you can switch vendors without touching clients.
Retrieval sits behind the gateway, not beside it. The gateway passes the verified user identity to the retriever, and the retriever returns only chunks that user may read right now. The connectors that feed the index use read-only service identities, and the index stores each chunk with the access control list of its source and a sensitivity label.
Oversharing: a worked example
Worked example. A finance analyst saves a salary planning workbook to a shared drive and, to send a colleague a link quickly, sets sharing to everyone in the organisation. For two years nobody notices, because nobody browses that folder. Then the company connects its assistant to the drive. The connector crawls everything its identity can read, which includes every org-wide file. An engineer asks the assistant what the pay band for a senior role is, and the top retrieved chunk is row 214 of the workbook. Every permission check passed. The assistant converted reachable into discoverable.
Three controls address this, in order of how much they help. First, fix permissions before rollout: inventory files shared to everyone or to anyone with the link, and burn that list down, starting with folders owned by HR, finance and legal. Second, label sensitive content and have the retriever exclude labels like confidential-restricted regardless of ACL, so a sharing mistake is not enough on its own. Third, re-check access at query time rather than trusting the ACL copied when the chunk was indexed, because permissions change and an index can be days stale.
def retrieve(query, user, index, acl_service, k=8):
"""Return chunks this user may read now; never trust index-time ACLs alone."""
groups = acl_service.groups_for(user.id) # from the IdP, cached briefly
candidates = index.search(
query, k=k * 4,
filter={"acl_principals": [user.id, *groups], # cheap pre-filter
"label_not_in": ["restricted", "legal-hold"]},
)
allowed = []
for ch in candidates:
if acl_service.can_read(user.id, ch.source_id): # authoritative, live
allowed.append(ch)
else:
index.mark_stale(ch.source_id) # reindex soon
if len(allowed) == k:
break
audit.log("retrieval", user=user.id, sources=[c.source_id for c in allowed])
return allowedThe over-fetch factor of four matters: if you filter after taking only k results, a user with narrow permissions gets an empty answer instead of the best answer they are entitled to. The live check costs one call per candidate, so batch it if your directory supports batch authorization.
Inspecting prompts at the gateway
The gateway inspects prompts on the way out. Keep the classifier cheap and deterministic for the high-confidence cases and route only ambiguous text to a heavier model. The output of inspection is a decision, not a boolean: allow, redact, route to a stricter model tier, or block with a message that tells the user what to do instead.
import re, hashlib
DETECTORS = {
"aws_key": re.compile(r"\bAKIA[0-9A-Z]{16}\b"),
"private_key": re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY-----"),
"card": re.compile(r"\b\d(?:[ -]?\d){12,18}\b"),
"email": re.compile(r"\b[\w.+-]+@[\w-]+\.[\w.]+\b"),
}
ACTION = {"aws_key": "block", "private_key": "block", "card": "redact", "email": "allow"}
def luhn_ok(digits):
d = [int(x) for x in digits if x.isdigit()][::-1]
return sum(x if i % 2 == 0 else (x * 2 - 9 if x > 4 else x * 2)
for i, x in enumerate(d)) % 10 == 0
def inspect(prompt, user):
hits, out = [], prompt
for name, rx in DETECTORS.items():
for m in rx.finditer(prompt):
if name == "card" and not luhn_ok(m.group()):
continue # cuts false positives on order numbers
hits.append(name)
if ACTION[name] == "redact":
out = out.replace(m.group(), f"[{name.upper()}]")
worst = "block" if any(ACTION[h] == "block" for h in hits) else \
"redact" if any(ACTION[h] == "redact" for h in hits) else "allow"
record = {"user": user.id, "decision": worst, "detectors": sorted(set(hits)),
"prompt_sha256": hashlib.sha256(prompt.encode()).hexdigest()}
return worst, out, recordTwo details carry most of the value. The Luhn check removes the bulk of false positives on long numbers, and false positives are what make employees abandon the sanctioned tool. And the audit record stores a hash of the prompt by default, not the prompt, so the log can prove what happened without becoming a second copy of everything. Keep full text only for blocked or flagged requests, under a shorter retention class and with every read of it logged.
Pattern detectors will not catch a pasted contract or a design document. For that, use document fingerprints from your data loss prevention system or an embedding similarity check against a small set of crown-jewel documents, and treat the result as a routing signal rather than a hard block. Our article on PII leakage in LLMs covers what models can memorise and repeat once data does get through.
Finding shadow AI
No gateway sees traffic that never reaches it. Shadow AI is any model use that bypasses the sanctioned path: a personal chatbot account, a browser extension that sends page contents to its own backend, a note-taking bot invited to a meeting, a developer calling a public API with a personal key. You cannot block all of it, and trying usually pushes it onto personal devices where you see nothing. What you can do is measure it, because the web proxy and DNS resolver already see it.
import csv
from collections import defaultdict
AI_HOSTS = {"chatgpt.com", "chat.openai.com", "api.openai.com", "claude.ai",
"api.anthropic.com", "gemini.google.com", "perplexity.ai",
"huggingface.co"} # extend from your own proxy categories
SANCTIONED_EGRESS = {"10.20.0.15"} # the gateway's own source address
def shadow_report(proxy_csv, upload_threshold=200_000):
users = defaultdict(lambda: {"requests": 0, "bytes_out": 0, "hosts": set()})
with open(proxy_csv, newline="") as f:
for row in csv.DictReader(f): # user, src_ip, host, bytes_out
host = row["host"].lower()
if row["src_ip"] in SANCTIONED_EGRESS:
continue
if any(host == h or host.endswith("." + h) for h in AI_HOSTS):
u = users[row["user"]]
u["requests"] += 1
u["bytes_out"] += int(row["bytes_out"])
u["hosts"].add(host)
heavy = {k: v for k, v in users.items() if v["bytes_out"] >= upload_threshold}
return users, heavyRun it weekly and look at two numbers: how many people use unsanctioned AI at all, and how many upload a lot. The first tells you whether your sanctioned tool is good enough; the second is your investigation queue. Upload volume matters more than request count because pasting a file is the event that moves data. Browser extensions show up as unfamiliar hosts with steady upload traffic from many users at once; review installed extensions through your browser management tooling rather than guessing from hostnames.
When the assistant can act
The moment the assistant can act, by sending a message, filing a ticket or sharing a document, every document it reads becomes a possible instruction source. A shared file that says to forward the thread to an outside address is prompt injection, and an assistant with a mail connector can obey it. The mechanics of this exfiltration path are worked through in data exfiltration via LLM tools.
- Give connectors the narrowest scopes their features need, and separate read scopes from write scopes into different identities.
- Require explicit user confirmation for writes that leave the tenant: external mail, public links, posts to shared channels.
- Render model output without auto-loading remote images or links to unknown hosts; an image URL is a write channel. See egress filtering for LLM outputs.
- Log every tool call with the documents that were in context when it was made, so an incident can be traced to the injecting file.
Prompt logs are a dataset
Prompt logs are useful for abuse investigation, quality work and legal discovery, and that is exactly why they are dangerous. Within a month they contain fragments of every sensitive project. Decide retention per class before launch: hashed metadata for everything for a year, full text for flagged requests for thirty days, nothing else. Put the full-text store behind its own access group, log every read, and check what the vendor retains separately, because your retention setting does not govern theirs. Settle these in the contract and record them in the AI governance program so the next renewal does not silently change them.
Failure modes
- Rollout before permission cleanup. The assistant indexes years of oversharing on day one. Run the sharing inventory first and gate connector rollout by department.
- Index-time ACLs only. A user removed from a project keeps retrieving its documents until the next crawl. Re-check at query time.
- A gateway that is too strict. False-positive blocks send people to personal accounts, and your shadow numbers rise. Track block rate and override requests per detector.
- Vendor keys on laptops. One leaked key bypasses identity, logging and inspection. Keys live only in the gateway.
- Logs as a data lake. Full prompts kept forever and readable by every analyst. Hash by default, keep text briefly.
- Meeting bots and extensions ignored. They capture more than any chat window. Inventory them and approve or block explicitly.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Single gateway for all model traffic | Identity, inspection, vendor portability | A new tier-one service to run; latency of a few milliseconds |
| Query-time ACL checks | No stale access | Directory load and retrieval latency |
| Blocking on detectors | Hard stop for secrets | False positives push users to shadow tools |
| Redact instead of block | Work continues | Redacted context can produce wrong answers |
| Hashed logging | Small blast radius | Harder quality analysis and investigations |
What to do next
- Inventory files shared org-wide or by public link in every repository you plan to connect; fix HR, finance and legal first.
- Stand up the gateway, move every vendor key into it and issue clients identity-based tokens instead.
- Add query-time permission checks and a restricted-label exclusion to retrieval, with a four-times over-fetch.
- Deploy deterministic detectors for secrets and card numbers, with redact and block actions, and measure false-positive rate weekly.
- Run the proxy-log shadow report, publish the trend, and use it to decide what the sanctioned tool is missing.
- Split connector read and write identities and require confirmation on any write that leaves the tenant.
- Write retention classes for prompt logs, restrict the full-text store and confirm vendor retention in the contract.