Giving a whole workforce an AI assistant is not one security decision; it is a dozen, and most of them are made by default. The assistant reads whatever its connectors can reach, employees paste whatever they are working on, the vendor keeps whatever its contract allows, and the people who were not given a sanctioned tool use an unsanctioned one from their phone. The policy side of this, which tools are allowed for which data, is covered in our article on employee AI usage policies. This article is the technical side: the architecture that makes a policy enforceable, the data flows that actually leak, and the code that sits in the middle.

The core idea is simple. There should be exactly one path from an employee to a model, and that path should know who the employee is, what the prompt contains, which documents the answer drew on and how long each of those facts is kept. Everything else in this article is a consequence of building that path and then measuring who goes around it.

Five flows, five failure shapes

Start with the flows, because each one fails differently.

FlowWhat goes wrongPrimary control
Employee to model (prompts)Source code, customer records or deal terms leave for a vendor with unclear retentionGateway classification, contracted no-training terms
Index to employee (retrieval)The assistant surfaces files the user could technically open but was never meant to findPermission hygiene, sensitivity labels, query-time ACL checks
Model to systems (actions)An agent sends mail, edits tickets or shares files because a document told it toScoped connectors, confirmation on writes, egress limits
Logs to everyoneThe prompt archive becomes the most sensitive dataset in the companyRetention classes, access logging, hashing
Employee to unsanctioned toolShadow AI: personal accounts, browser extensions, meeting botsDiscovery from proxy and DNS, a good sanctioned option

Notice that only the first row is what people usually mean by AI data leakage. In deployments that connect an assistant to mail, chat and file shares, the second row is a common first real incident, and it is not caused by the model at all.

Reference architecture

Workplace AI: every request crosses one gateway, every answer is trimmed to the calleremployeebrowser, IDE, chatidentity providerSSO, groupstokenAI gatewayclassify promptblock or redactroute to approved modellog with retention classpromptapproved modelscontracted, no trainingretrieval indexACL re-check at queryuser idconnectorsmail, drive, wikiaudit storeaccess logged tooweb proxy and DNS logsshadow AI discoveryThe gateway is the only path to a model; the proxy tells you who is going around it.
Reference architecture. Identity flows into the gateway and on into retrieval; the proxy is a sensor, not a control.

The gateway is a small reverse proxy in front of every model endpoint the company pays for. Clients never hold vendor API keys; they hold a short-lived token from the identity provider, and the gateway exchanges it for the vendor call. That single design choice gives you four things at once: you can revoke a user without rotating a vendor key, you know which human sent each prompt, you can inspect content before it leaves, and you can switch vendors without touching clients.

Retrieval sits behind the gateway, not beside it. The gateway passes the verified user identity to the retriever, and the retriever returns only chunks that user may read right now. The connectors that feed the index use read-only service identities, and the index stores each chunk with the access control list of its source and a sensitivity label.

Oversharing: a worked example

Oversharing: the assistant did not leak anything; the permissions already hadsalary_2026.xlsxshared: whole orgconnector crawlindexes what it can readchunk + embedACL copied at indexany employeeasks about pay bandsqueryBefore the assistant: the file was reachable but undiscoverable, buried in a share nobody browsed.After: semantic search makes it the top answer to a natural question.Fix the source permission, add a sensitivity label the retriever excludes, and audit before rollout.
The oversharing path. The ACL was wrong long before any model was involved.

Worked example. A finance analyst saves a salary planning workbook to a shared drive and, to send a colleague a link quickly, sets sharing to everyone in the organisation. For two years nobody notices, because nobody browses that folder. Then the company connects its assistant to the drive. The connector crawls everything its identity can read, which includes every org-wide file. An engineer asks the assistant what the pay band for a senior role is, and the top retrieved chunk is row 214 of the workbook. Every permission check passed. The assistant converted reachable into discoverable.

Three controls address this, in order of how much they help. First, fix permissions before rollout: inventory files shared to everyone or to anyone with the link, and burn that list down, starting with folders owned by HR, finance and legal. Second, label sensitive content and have the retriever exclude labels like confidential-restricted regardless of ACL, so a sharing mistake is not enough on its own. Third, re-check access at query time rather than trusting the ACL copied when the chunk was indexed, because permissions change and an index can be days stale.

def retrieve(query, user, index, acl_service, k=8):
    """Return chunks this user may read now; never trust index-time ACLs alone."""
    groups = acl_service.groups_for(user.id)          # from the IdP, cached briefly
    candidates = index.search(
        query, k=k * 4,
        filter={"acl_principals": [user.id, *groups],  # cheap pre-filter
                "label_not_in": ["restricted", "legal-hold"]},
    )
    allowed = []
    for ch in candidates:
        if acl_service.can_read(user.id, ch.source_id):  # authoritative, live
            allowed.append(ch)
        else:
            index.mark_stale(ch.source_id)               # reindex soon
        if len(allowed) == k:
            break
    audit.log("retrieval", user=user.id, sources=[c.source_id for c in allowed])
    return allowed

The over-fetch factor of four matters: if you filter after taking only k results, a user with narrow permissions gets an empty answer instead of the best answer they are entitled to. The live check costs one call per candidate, so batch it if your directory supports batch authorization.

Inspecting prompts at the gateway

The gateway inspects prompts on the way out. Keep the classifier cheap and deterministic for the high-confidence cases and route only ambiguous text to a heavier model. The output of inspection is a decision, not a boolean: allow, redact, route to a stricter model tier, or block with a message that tells the user what to do instead.

import re, hashlib

DETECTORS = {
    "aws_key":   re.compile(r"\bAKIA[0-9A-Z]{16}\b"),
    "private_key": re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY-----"),
    "card":      re.compile(r"\b\d(?:[ -]?\d){12,18}\b"),
    "email":     re.compile(r"\b[\w.+-]+@[\w-]+\.[\w.]+\b"),
}
ACTION = {"aws_key": "block", "private_key": "block", "card": "redact", "email": "allow"}

def luhn_ok(digits):
    d = [int(x) for x in digits if x.isdigit()][::-1]
    return sum(x if i % 2 == 0 else (x * 2 - 9 if x > 4 else x * 2)
               for i, x in enumerate(d)) % 10 == 0

def inspect(prompt, user):
    hits, out = [], prompt
    for name, rx in DETECTORS.items():
        for m in rx.finditer(prompt):
            if name == "card" and not luhn_ok(m.group()):
                continue                       # cuts false positives on order numbers
            hits.append(name)
            if ACTION[name] == "redact":
                out = out.replace(m.group(), f"[{name.upper()}]")
    worst = "block" if any(ACTION[h] == "block" for h in hits) else \
            "redact" if any(ACTION[h] == "redact" for h in hits) else "allow"
    record = {"user": user.id, "decision": worst, "detectors": sorted(set(hits)),
              "prompt_sha256": hashlib.sha256(prompt.encode()).hexdigest()}
    return worst, out, record

Two details carry most of the value. The Luhn check removes the bulk of false positives on long numbers, and false positives are what make employees abandon the sanctioned tool. And the audit record stores a hash of the prompt by default, not the prompt, so the log can prove what happened without becoming a second copy of everything. Keep full text only for blocked or flagged requests, under a shorter retention class and with every read of it logged.

Pattern detectors will not catch a pasted contract or a design document. For that, use document fingerprints from your data loss prevention system or an embedding similarity check against a small set of crown-jewel documents, and treat the result as a routing signal rather than a hard block. Our article on PII leakage in LLMs covers what models can memorise and repeat once data does get through.

Finding shadow AI

No gateway sees traffic that never reaches it. Shadow AI is any model use that bypasses the sanctioned path: a personal chatbot account, a browser extension that sends page contents to its own backend, a note-taking bot invited to a meeting, a developer calling a public API with a personal key. You cannot block all of it, and trying usually pushes it onto personal devices where you see nothing. What you can do is measure it, because the web proxy and DNS resolver already see it.

import csv
from collections import defaultdict

AI_HOSTS = {"chatgpt.com", "chat.openai.com", "api.openai.com", "claude.ai",
            "api.anthropic.com", "gemini.google.com", "perplexity.ai",
            "huggingface.co"}            # extend from your own proxy categories
SANCTIONED_EGRESS = {"10.20.0.15"}     # the gateway's own source address

def shadow_report(proxy_csv, upload_threshold=200_000):
    users = defaultdict(lambda: {"requests": 0, "bytes_out": 0, "hosts": set()})
    with open(proxy_csv, newline="") as f:
        for row in csv.DictReader(f):          # user, src_ip, host, bytes_out
            host = row["host"].lower()
            if row["src_ip"] in SANCTIONED_EGRESS:
                continue
            if any(host == h or host.endswith("." + h) for h in AI_HOSTS):
                u = users[row["user"]]
                u["requests"] += 1
                u["bytes_out"] += int(row["bytes_out"])
                u["hosts"].add(host)
    heavy = {k: v for k, v in users.items() if v["bytes_out"] >= upload_threshold}
    return users, heavy

Run it weekly and look at two numbers: how many people use unsanctioned AI at all, and how many upload a lot. The first tells you whether your sanctioned tool is good enough; the second is your investigation queue. Upload volume matters more than request count because pasting a file is the event that moves data. Browser extensions show up as unfamiliar hosts with steady upload traffic from many users at once; review installed extensions through your browser management tooling rather than guessing from hostnames.

When the assistant can act

The moment the assistant can act, by sending a message, filing a ticket or sharing a document, every document it reads becomes a possible instruction source. A shared file that says to forward the thread to an outside address is prompt injection, and an assistant with a mail connector can obey it. The mechanics of this exfiltration path are worked through in data exfiltration via LLM tools.

  • Give connectors the narrowest scopes their features need, and separate read scopes from write scopes into different identities.
  • Require explicit user confirmation for writes that leave the tenant: external mail, public links, posts to shared channels.
  • Render model output without auto-loading remote images or links to unknown hosts; an image URL is a write channel. See egress filtering for LLM outputs.
  • Log every tool call with the documents that were in context when it was made, so an incident can be traced to the injecting file.

Prompt logs are a dataset

Prompt logs are useful for abuse investigation, quality work and legal discovery, and that is exactly why they are dangerous. Within a month they contain fragments of every sensitive project. Decide retention per class before launch: hashed metadata for everything for a year, full text for flagged requests for thirty days, nothing else. Put the full-text store behind its own access group, log every read, and check what the vendor retains separately, because your retention setting does not govern theirs. Settle these in the contract and record them in the AI governance program so the next renewal does not silently change them.

Failure modes

  • Rollout before permission cleanup. The assistant indexes years of oversharing on day one. Run the sharing inventory first and gate connector rollout by department.
  • Index-time ACLs only. A user removed from a project keeps retrieving its documents until the next crawl. Re-check at query time.
  • A gateway that is too strict. False-positive blocks send people to personal accounts, and your shadow numbers rise. Track block rate and override requests per detector.
  • Vendor keys on laptops. One leaked key bypasses identity, logging and inspection. Keys live only in the gateway.
  • Logs as a data lake. Full prompts kept forever and readable by every analyst. Hash by default, keep text briefly.
  • Meeting bots and extensions ignored. They capture more than any chat window. Inventory them and approve or block explicitly.

Trade-offs

ChoiceGainCost
Single gateway for all model trafficIdentity, inspection, vendor portabilityA new tier-one service to run; latency of a few milliseconds
Query-time ACL checksNo stale accessDirectory load and retrieval latency
Blocking on detectorsHard stop for secretsFalse positives push users to shadow tools
Redact instead of blockWork continuesRedacted context can produce wrong answers
Hashed loggingSmall blast radiusHarder quality analysis and investigations

What to do next

  1. Inventory files shared org-wide or by public link in every repository you plan to connect; fix HR, finance and legal first.
  2. Stand up the gateway, move every vendor key into it and issue clients identity-based tokens instead.
  3. Add query-time permission checks and a restricted-label exclusion to retrieval, with a four-times over-fetch.
  4. Deploy deterministic detectors for secrets and card numbers, with redact and block actions, and measure false-positive rate weekly.
  5. Run the proxy-log shadow report, publish the trend, and use it to decide what the sanctioned tool is missing.
  6. Split connector read and write identities and require confirmation on any write that leaves the tenant.
  7. Write retention classes for prompt logs, restrict the full-text store and confirm vendor retention in the contract.
Key takeaway: Put one identity-aware gateway between employees and every model, make retrieval re-check permissions at query time and exclude restricted labels, fix org-wide sharing before connecting anything, measure shadow AI from proxy logs instead of guessing, confirm outbound writes, and treat prompt logs as the sensitive dataset they become.