Permission filtering in RAG retrieval
Filter inside the search, not around it. Post-filtering — retrieve k, then drop what the user cannot see — is the obvious approach and it silently returns fewer results, sometimes none, because the permitted documents were never in the top k. And filtering after generation is not filtering: the model already read the text.
Overview
Why post-filtering quietly fails
Retrieve the top 10 by similarity, then remove what the user may not see. If eight were restricted you return two, and nothing reports that the result set collapsed. The generator answers from thin evidence and sounds no less confident.
Over-fetching — retrieve 100, filter, keep 10 — makes it less likely and never fixes it. A user with access to 0.1% of the corpus needs an enormous k, and the failure is worst exactly for the most restricted users.
Parameters
Visualisation
—Readout
What to watch
- Post-filtering removes results after the top-k is chosen.
- The permitted documents may never have entered the top-k at all.
- Filtering after generation is too late — the model already saw it.
Permission filtering in RAG retrieval: A Practical Guide
How do you stop a RAG system retrieving documents the user is not allowed to see?
Pre-filtering, and why it is harder than it looks
Pre-filtering restricts the candidate set before ranking, so the top k is k permitted results. That is the correct behaviour and it fights the index: an HNSW graph is built over all vectors, and walking it while skipping most nodes can disconnect the search — you traverse into a region where everything is filtered out and the walk stalls.
Vector databases handle this differently, which is one of the real differences between them. Options include filtered graph traversal, per-tenant sub-indexes (clean, and costly at many tenants), and falling back to a flat scan when the filter is very selective — which is often genuinely fastest, because the permitted set is small.
Where the permissions live
Baking an access list into each chunk's metadata at index time is fast and goes stale the moment someone leaves a group — and stale permissions fail open, which is the wrong direction.
The usual compromise: index a stable group identifier with each chunk, resolve the user's groups per request from the identity system, and filter on the intersection. Group membership changes are then reflected immediately without re-indexing anything.
Two more things to say. Deleted documents must leave the index, not just be filtered out. And the filter belongs on the server, derived from the authenticated session — never from a client-supplied user id, which is trivially forged.
Why pre-filtering fights the index
Pre-filtering is the correct behaviour and it is genuinely hard to implement, because an ANN index was built over the whole corpus and does not know about your predicate.
Walking an HNSW graph while skipping most nodes can disconnect the traversal: the search reaches a region where every neighbour is filtered out and stalls, returning far fewer results than asked for even though plenty of permitted documents exist elsewhere in the graph.
The strategies databases use, all with trade-offs: filtered traversal that keeps searching past excluded nodes; per-tenant sub-indexes, which are clean and expensive once tenants number in the thousands; and falling back to a flat scan when the filter is selective — which is often genuinely fastest, since a user with access to 0.1% of a million documents is a thousand-vector brute force.
Getting the model right, and failing closed
Two design rules cover most of the damage.
Derive the filter from the session, never the request. A client-supplied user or group id is trivially forged. The filter must be built server-side from the authenticated identity.
Index groups, resolve membership per query. Baking a user list into each chunk means a departure leaves the index granting access until the next re-index — and stale permissions fail open, which is the wrong direction. Indexing a stable group id and resolving the user's groups per request makes revocation immediate.
Also: filtering is not deletion. A document removed for legal or privacy reasons must leave the index and any cached answers derived from it, or it remains recoverable by anyone who can still see it — which is a different and worse problem than a missing result.
The failure mode that matters most
A RAG system retrieves documents and puts them in a prompt. If the retrieval does not respect who is asking, the model will faithfully summarise a document the user was never allowed to see — and it will do so fluently, with a citation.
That is a data breach with a helpful tone, and it is the most consequential thing that can go wrong in an enterprise retrieval system. Unlike a hallucination, it is not a quality problem; it is a security one.
The requirement is straightforward to state and easy to implement incorrectly: the set of documents considered must be exactly the set this user may read.
Where the filter has to be applied
There are three places it could go, and only one is correct.
After generation — check the answer for leaked content. Hopeless: the content is already in the output, and paraphrase makes detection unreliable.
After retrieval, before the prompt — retrieve k documents, discard the ones the user cannot see, pass the rest. Better, and still wrong in two ways: the unauthorised documents occupied slots, so fewer authorised ones were considered; and their content has already been loaded into the application's memory.
Inside the search — the index only ever considers documents matching the user's permissions. This is the correct answer, and it requires the store to support filtered search.
results = store.search(
query_vector=q,
k=10,
filter={"acl": {"$in": user.group_ids}}, # applied during the search
)The difference between the second and third options is not theoretical. With post-filtering, a user in a small group querying a corpus dominated by other groups' documents may get zero results while relevant authorised documents exist — because all ten retrieved slots went to documents they cannot see.
Modelling permissions in metadata
The filter can only express what the metadata records, so the model has to be decided at indexing time.
| Model | Metadata | Filter |
|---|---|---|
| Group-based | acl: [group_ids] | user's groups intersect the list |
| Role-based | min_role: "manager" | user's role rank is sufficient |
| Owner-based | owner_id | owner matches, or is shared with |
| Attribute-based | region, department, classification | several predicates combined |
| Row-level via source | Reference to the source system | Resolve at query time |
Group-based ACLs cover most cases and are what most source systems already provide. Store the list of group ids permitted to read each document, and filter on intersection with the requesting user's groups.
Two hard parts:
Inherited permissions. A document in a restricted folder inherits its restriction. The indexing pipeline must resolve inheritance and write the effective ACL, not the document's own.
Permission changes. When someone leaves a group, or a document is reclassified, the index is stale and will over-share until it is updated. Permissions change far more often than document content, so they need their own refresh path — ideally event-driven, and at minimum on a short schedule.
Post-filtering, measured, and why it fails quietly
The article says post-filtering quietly fails. That is a claim with a number, and the number is worse than most people expect. This measures it, then measures the pre-filtering approach that fixes it and the new problem that one creates.
Things to try
- Drop corpus visible to 5%. Post-filtering returns almost nothing while pre-filtering still returns k — the control fails worst for the most restricted user.
- Raise over-fetch to 500. The shortfall shrinks and never disappears, which is why over-fetching is a mitigation rather than a fix.
- Set visibility to 100%. All three strategies agree, which is exactly why this bug survives testing on an admin account.
What to remember
Permission filtering has to happen inside the search, not around it. Post-filtering ranks first and drops afterwards, so it silently returns fewer results than requested and degrades worst for the most restricted users. Filtering after generation is not a control at all — the model has already read the text. Index a group id and resolve membership per request, because stale permissions fail open.
Late-binding permissions
An alternative to storing ACLs in the index is to resolve them at query time against the source system.
Store only a reference — the source id — and after retrieving candidates, ask the source system which of them this user may read.
| ACLs in the index | Resolve at query time | |
|---|---|---|
| Staleness | Possible, needs refresh | None — always current |
| Latency | None extra | An extra call per query |
| Correctness during changes | Window of over-sharing | Correct immediately |
| Filtering inside the search | Yes | No — post-filter only |
The trade is sharp: resolving at query time is always correct and forces post-filtering, which brings back the slot-wastage problem.
The pragmatic combination used in practice is both: filter inside the search using indexed ACLs to get a correct-enough candidate set efficiently, then verify the final selection against the source system before showing anything. The index filter does the heavy lifting; the verification catches the staleness window.
Failure modes to test for explicitly
Permission bugs do not announce themselves, so they have to be tested deliberately:
- A user with no groups should retrieve nothing, not everything. A missing or empty filter that matches all documents is the classic catastrophic bug.
- A filter on a field that does not exist should fail closed, not open. Some stores silently ignore unknown fields.
- Inherited restrictions on documents in restricted containers.
- Deleted users and revoked groups.
- Documents with no ACL recorded — decide whether the default is deny (correct) or allow (dangerous), and make it explicit.
- Cached results. A cache keyed only on the query text will serve one user's authorised results to another. Cache keys must include the permission context.
That last one is worth stating plainly because caching is added later, as a performance improvement, by someone who is not thinking about permissions.
What else leaks
Retrieval is the main channel and not the only one.
Conversation history carries content forward. If a user's permissions change mid-session, earlier retrieved content is still in the context.
Logs and traces. Prompt logging captures retrieved document content, so log storage inherits the documents' classification.
Embeddings themselves. Research has shown that text can be partially reconstructed from its embedding, so a vector store holding sensitive content is itself sensitive and needs the same protections as the documents.
Aggregate answers. A user may be permitted to see individual documents but not a synthesis that reveals a pattern across them. This is a policy question rather than a technical one, and worth asking.
Questions people ask
Can I filter after retrieval? It is not secure enough and it wastes retrieval slots. Filter inside the search.
What if my vector store does not support filtered search? Over-retrieve heavily and post-filter as an interim measure, and treat it as a reason to change store.
How do I keep permissions fresh? Event-driven updates from the source system, with a scheduled reconciliation as a backstop. Permissions change more often than content.
Should documents with no ACL be readable? Default deny. An explicit allow-all marker is safer than an absent field.
Does the embedding leak content? Partially reconstructable, so treat the vector store as holding the data itself.
How do I cache safely? Include the user's permission context in the cache key, or cache only at the query-embedding level rather than the results level.
Recap in one screen
- An unfiltered retrieval will faithfully summarise documents the user may not read.
- Apply the filter inside the search; post-filtering both leaks and wastes retrieval slots.
- Store effective ACLs at indexing time, resolving inheritance, and refresh them independently of content.
- Default to deny for documents with no recorded permissions, and test the no-groups case explicitly.
- Caches, logs, conversation history and the vectors themselves are all additional leak channels.