An enterprise RAG system indexes documents from systems that each have their own permission model: a wiki with space-level restrictions, a file store with per-folder sharing, a ticketing system with project roles, a code host with repository teams. The moment those documents are chunked, embedded and put into a single vector index, the permissions are gone unless the pipeline deliberately carries them along. The assistant then becomes the easiest way in the company to read a document you are not allowed to open: ask a question, and the retriever helpfully places the relevant paragraph of the board deck or the compensation spreadsheet into the prompt.
Retrieval must therefore enforce the same access decisions as the source systems, for the specific user asking, before any content reaches the model. This article covers the design that achieves that: propagating the user's identity, copying ACLs onto chunks at index time, filtering inside the search rather than after it, re-checking against a live authorization source, propagating revocations quickly, and closing the secondary channels, such as caches, memory, logs and derived documents, through which restricted content escapes even when retrieval itself is correct.
Principles
Four rules shape everything else. Enforce before the prompt: the language model cannot be an enforcement point, because anything placed in its context can be revealed by a sufficiently clever question, and instructions such as "do not reveal restricted content" are not a security boundary. Act as the user, not as the service: retrieval runs with the end user's identity and group memberships, never with the broad service account the connector used to crawl. Deny by default: a chunk with missing, unparseable or unknown ACL data is invisible to everyone until it is fixed. Fail closed: if the authorization service or group resolver is unavailable, return fewer results or an error, never unfiltered ones.
These rules sound obvious, but each has a tempting shortcut. Crawling with an administrator token is the easy way to get complete content, and reusing that token at query time is the easy way to get complete answers. Treating missing ACL data as public makes connectors look finished sooner. Falling back to unfiltered search during an authorization outage keeps the assistant available. Every one of these shortcuts turns into a data exposure incident.
Carrying ACLs from source to chunk
Connectors must fetch permissions alongside content. For each document they capture the principals allowed to read it, the principals explicitly denied where the source supports deny entries, the tenant, and a version or timestamp for the ACL itself. Principals should be normalized into one namespace, with prefixes such as user: and group: and a source qualifier, because the same person often has different identifiers in each system and must be mapped to a single identity-provider id.
Every chunk inherits its document's ACL fields verbatim. The index stores them as filterable metadata next to the vector and text, so the search engine can apply them without a second lookup. Inheritance rules deserve care: many sources let a document inherit permissions from a parent folder or space, and the connector must resolve the effective ACL rather than copying only the permissions set directly on the item.
def index_chunk(chunk, acl, embed):
"""acl: the document's effective ACL, resolved by the connector."""
if acl is None or not acl.allow:
raise ValueError(f"{chunk.doc_id}: no ACL; refusing to index (deny by default)")
return {
"id": chunk.id,
"doc_id": chunk.doc_id,
"tenant": acl.tenant,
"vector": embed(chunk.text),
"text": chunk.text,
"allow": sorted(acl.allow), # e.g. "user:okta|00u1ab", "group:okta|finance-emea"
"deny": sorted(acl.deny),
"acl_version": acl.version,
}
Filtering inside the search, not after it
There are two places to apply the permission filter, and only one works reliably. Post-filtering retrieves the top k results by similarity and then removes the ones the user cannot read. When most of the corpus is restricted for a given user, the top k can be entirely forbidden, and the user gets no results even though readable, relevant documents exist further down the ranking. Over-fetching, say fifty candidates to keep eight, reduces the problem but never eliminates it, and it wastes reranker work.
Pre-filtering applies the permission predicate during the search, so the engine returns the top k among documents the user can read. Most vector databases and search engines support metadata filters on approximate nearest-neighbor queries, but their behavior under very selective filters differs: graph indexes such as HNSW can lose recall when most nodes are filtered out, and many engines switch to exact search over the matching subset below some selectivity threshold. Test recall under your real permission distribution, especially for users who can see only a small slice of the corpus, and partition the index by tenant so that the most selective predicate is handled structurally rather than by the filter.
def retrieve(query, identity, index, authz, embed, k=8):
principals = {f"user:{identity.user_id}"} | {
f"group:{g}" for g in identity.resolved_groups() # transitive, cached briefly
}
flt = { # filter syntax is illustrative
"tenant": identity.tenant,
"allow": {"any_of": sorted(principals)},
"deny": {"none_of": sorted(principals)},
}
hits = index.search(embed(query), k=k * 2, filter=flt)
# Defense in depth: confirm against the live source of truth.
verdicts = authz.check_many(identity, [h.doc_id for h in hits], "read")
return [h for h in hits if verdicts.get(h.doc_id) is True][:k]Groups and the principal explosion
Real users belong to many groups, often hundreds, through nested memberships in the identity provider and in each source system. Resolving the full transitive set on every query is expensive, and passing hundreds of principal ids into a search filter can hit query-size limits or slow the search. Resolve memberships in a dedicated service that caches each user's expanded set for a short period, such as a few minutes, and invalidates it on membership-change events from the identity provider.
Some teams go further and precompute access tokens: each distinct ACL is hashed to an id, the user's set of reachable ACL ids is computed from their memberships, and chunks carry only the ACL id. This keeps filters small but moves the complexity into keeping the user-to-ACL mapping current. Whichever approach you use, the short cache interval is part of the revocation budget and must be counted in it.
Live authorization checks
Index-time ACLs are a snapshot, and snapshots are stale by construction. A final check against a live source before content enters the prompt closes the gap: either the source system's own permission API, or a central relationship-based authorization service, such as the Zanzibar-inspired SpiceDB or OpenFGA, fed from the same connectors. The index filter does the bulk narrowing cheaply; the live check confirms the handful of results that survive. Because it runs only on the final candidates, it adds little latency.
The live check must also fail closed. If it times out, drop the unverified results and tell the user that some sources could not be checked, rather than passing them through. Record every decision, with the user, document, ACL version and verdict, so security teams can answer who could see what and when.
Revocation speed
When someone is removed from a project or a document is restricted, the change must take effect in the assistant within a defined time, and that time should be much shorter than the content refresh interval. Treat permission changes as a separate, high-priority update path: subscribe to permission events where sources emit them, poll ACLs more often than content where they do not, and update the ACL fields on affected chunks without re-chunking or re-embedding. An ACL-only update is a metadata write and should complete in seconds to minutes.
Define a revocation SLO, measure it with synthetic tests, and make the query path safe for the interval before the index catches up. That interval is exactly what the live authorization check covers: even if the index still lists a revoked user, the live check refuses the chunk.
Channels that leak anyway
Correct retrieval is necessary but not sufficient, because restricted content can reach a user through paths that bypass the retriever entirely.
- Caches. A semantic or response cache keyed only by the question will serve one user's answer, built from their documents, to another user. Key caches by the principal set or ACL fingerprint as well as the query, or cache only retrieval over public content.
- Conversation memory. Memories and summaries extracted from a session carry the permissions of the documents they came from; shared or team-level memory must not store content from restricted sources.
- Derived documents. Summaries, digests and knowledge-graph facts built from several sources must be readable only by principals who could read all of them: the intersection of their ACLs, not the union.
- Citations and snippets. Titles, file names and preview snippets are content too; filter them with the same rules as body text.
- Logs and traces. Prompt logs contain retrieved text; restrict access to them at least as tightly as the most sensitive source they contain.
- Embeddings. Published research has reconstructed substantial text from embeddings, so treat vectors as sensitive data, never as anonymized derivatives.
Testing permission enforcement
Permission bugs are silent: nothing crashes, and the answer looks helpful. Build explicit tests. Maintain a permission matrix of synthetic users and canary documents with known access rules across every connector, including nested groups, inherited folder permissions, explicit denies and cross-tenant cases. Each canary contains a unique marker string, and the test suite asks every synthetic user questions that would retrieve every canary, failing if a marker appears in any retrieved context, citation or answer where access should be denied.
Run the suite on every connector, index or retrieval change, and continuously in production against a canary tenant. Add revocation tests that remove a synthetic user from a group and measure how long the canary stays retrievable. The target leak rate is zero, so any failure blocks the release; unlike quality metrics, there is no acceptable threshold to trade against.
Failure modes
- Service-account retrieval. Searching with the crawler's credentials returns everything the crawler could read.
- Unresolved inheritance. Copying only item-level permissions misses restrictions set on a parent folder or space.
- Fail-open fallbacks. Returning unfiltered results when the group resolver or authorization service is down.
- Post-filter only. Users with narrow access get empty results, so teams quietly widen the filter to fix it.
- Shared caches. Response caches keyed only by the query text serve restricted answers across users.
- Revocation lag. Permission changes wait for the nightly content crawl.