LLM Application Architecture advanced 8 min read 7 flashcards

Permission-Aware Retrieval

Why authorisation in a retrieval system has to happen inside the query rather than after it, how group-based ACLs flatten badly into index filters, and the leak that no retrieval filter can close.

Someone asks the internal assistant what the new vice-president is being paid, and gets a correct, well-cited answer. No rule was broken: the compensation spreadsheet sat in a folder that inherited a sharing link three reorganisations ago, and the employee could always have opened it. What changed is that finding it used to require knowing it existed. Retrieval converts latent over-permissioning into a question anyone can ask in one sentence, which is why Microsoft's own deployment guidance for Microsoft 365 Copilot puts oversharing remediation before rollout rather than after it (Microsoft, Microsoft 365 Copilot oversharing blueprint).

Authorisation belongs inside the query

There are three places to enforce access, and only two of them work.

At ingestion, by building a separate index per audience. Correct, simple, and it multiplies storage and re-embedding cost by the number of audiences, so it suits a handful of coarse tiers and nothing finer.

In the query, by storing principal identifiers on each indexed unit and filtering on them. This is the standard pattern: Azure AI Search documents it as a field of principal strings combined with the search.in filter function, and warns that hand-built equality expressions are "error-prone, difficult to maintain, and slow down query response time by many seconds" where search.in stays subsecond (Microsoft, Security filters for trimming results in Azure AI Search). Worth being precise about what this is: the principal is a string in a filter expression, not an authentication mechanism. Whatever maps the caller's token to that string is the real security boundary.

After retrieval, by fetching the top k and discarding what the user cannot see. This is the one that fails quietly. A user with narrow access gets a short list or an empty one while every retrieval metric looks healthy, because the filter ran after the quality signal was computed. The approximate-index machinery that makes in-query filtering hard is a genuine engineering problem, covered in filtered vector search, but post-filtering is not a solution to it, only a way of hiding it.

Groups flatten badly

Real permissions are group-based, nested, time-varying, and sometimes negative. An index filter is a flat set of strings. Every gap between those two descriptions is a bug.

Group expansion has to be materialised somewhere, and whatever you materialise goes stale. Someone removed from a group this morning keeps matching documents stamped with that group until the expansion is recomputed, so the staleness window of your directory sync is also the duration of your access-control violations. The usual fix is to resolve the caller's group membership at query time from the identity provider and keep only the document-side set in the index, which moves the freshness problem to the side that changes less often.

Deny rules do not flatten at all. A filter assembled from allow-lists cannot express "everyone in Finance except contractors", and an approximation will be wrong in one direction or the other. If the source system supports deny, the index needs either a separate deny field evaluated with precedence or a pre-resolved effective-permission set computed upstream.

Granularity is the last trap. ACLs live on documents and retrieval returns chunks, so either every chunk carries its parent's permission set, which inflates the index, or retrieved chunks are re-checked against the document store before they reach the prompt, which adds a round trip per result.

The leak retrieval filters cannot close

Filtering the retrieval step does not make the system safe, because retrieval is not the only path from the corpus to the answer. Anything computed with elevated privileges carries the union of its inputs' permissions and must be constrained to the intersection: LLM-written chunk summaries, the situating sentences added by contextual retrieval, the community summaries in GraphRAG, cached answers, and long-term memory written from one user's session. A community summary built over the whole corpus is a single document that quotes everything, and no per-document ACL will save you once it is in the index.

The practical rule is that each derived artefact needs its own ACL, computed as the intersection of the ACLs of everything that went into it, and that artefacts whose intersection is empty should not be built at all.

When it breaks

Erasure has more copies than you think. A document the user removed, or that a subject-access request requires deleting, exists as source, chunks, vectors, a semantic cache entry, trace logs and probably an eval set. Approximate indexes delete by tombstone rather than by removal, so the vector lingers until a rebuild (freshness, deletes and index maintenance).

Evaluation runs as an administrator. A recall number measured with an unfiltered index is not the recall any user experiences. Evaluation sets need per-persona variants, and the gap between unfiltered and filtered recall is the number worth reporting.

Metadata is disclosure. A citation list that names a document the filter excluded, or a result count that drops from 50 to 3, tells the user something exists. Suppress counts and render citations only from returned content.

Filter size becomes a latency problem. A caller in 4,000 groups produces a filter expression with 4,000 terms. Precompute a stable audience key, or invert the problem and store the small set on the side that has it.

Injection rides on legitimate retrieval. Permission filters decide what the model may read, not what it may be tricked into emitting; see output handling and downstream injection.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Microsoft, Microsoft 365 Copilot oversharing blueprint learn.microsoft.com
  2. Microsoft, Security filters for trimming results in Azure AI Search learn.microsoft.com
Check yourself

7 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track