Back to blog
8 min readGateco Team

Your RAG Stack Knows Who You Were This Morning

At 09:00 an administrator removes someone from a group. The documents that group could see are still sitting in an index, each one tagged with the groups allowed to read it. At 09:05 that person asks the assistant a question. The assistant retrieves, the filter runs against the tags it has, and the answer comes back with a citation to a document they are no longer allowed to open. Nobody gets an error. Nothing is logged as a denial. The system did exactly what it was built to do.

That is the whole problem this post is about. It is not a bug in anyone's product. It is a property of a design, and the design is the one almost everyone ships first, because it is the one that works on day one.

The two designs

There are two ways to enforce permissions on retrieval, and they fail differently.

**Copy permissions into the index, then filter at query time.** At ingestion, you resolve who may see each document and write that answer next to the vector: a list of group ids, an ACL, a label. At query time, you compile the caller's identity into a filter and let the search engine exclude what does not match. It is fast, it composes with any vector database, and the search engine does the work. Its cost is that the permission written next to the vector is a copy, and a copy is only as fresh as the last time you wrote it.

**Resolve permissions at query time, then filter.** Nothing about entitlement is stored with the vector. When a request arrives, you resolve who is asking against a directory that is kept current, evaluate policy against each candidate result, and return only what passes. Its cost is a policy evaluation on every query, and a dependency on the directory being reachable and correct. Its benefit is that there is no copy to go stale.

Both are legitimate. The first is a caching decision, and like every cache it trades freshness for cost. A vector that is close to the question tells you nothing about whether this person may read it today. The question a buyer should ask is not which design a vendor chose, but what the vendor says happens between the moment a permission changes and the moment retrieval reflects it. We wrote about the mechanics of the first design in A Filter Is Not an Authorization System; this post is about what it costs.

What the gap costs, by the kind of change

Not every permission change ages the same way. Four kinds matter, and vendors are increasingly precise about them, so you can be too.

**A permission on one item.** The easiest case. If the platform watches the source for item-level changes, an indexer run picks it up. Microsoft's documentation for the indexed SharePoint path states that, starting in the `2026-05-01-preview` API, "ACL changes for items with unique permissions are detected and refreshed on each successful indexer run" (Microsoft Learn, SharePoint indexer permission metadata page, dated 8 August 2026, read 17 September 2026). Item-level staleness is bounded by the indexer schedule.

**A permission on a parent.** This is the one that fans out. Remove a group from a site, a library or a folder, and every item that inherited from it has silently changed. The same page is direct about it: "Parent-scope permission changes aren't picked up automatically on subsequent indexer runs," and, in its own summary, "If you change SharePoint permissions without triggering an update mechanism, the index serves stale ACL data for previously ingested files." The remedy is an explicit resync. Until someone runs it, the copy is wrong for every child.

**A group membership.** The 09:00 example. The document's ACL did not change; the person's membership did. Whether this is caught depends entirely on where group membership is resolved. If it was compiled into the index, it is stale until the next full sync. If it is resolved at query time against the identity provider, it is as fresh as the identity provider.

When the platform never had a live path

Some retrieval services are explicit that permission enforcement is not their job. Snowflake's documentation for Cortex Search states that services run "with owner's rights," and that "any role with sufficient privileges to query a Cortex Search Service may query any of the data the service has indexed, regardless of that role's privileges on the underlying objects" (Snowflake docs, query a Cortex Search service, read 22 September 2026). Databricks says the same thing about its vector search index in one sentence: "Row and column level permissions are not supported. However, you can implement your own application level ACLs using the filter API" (Databricks docs, vector search, updated 14 September 2026, read 22 September 2026). In both cases the filter is yours to write, and so is its freshness.

Three vendors, three different products, the same shape: the platform stores or serves what it indexed, and the application is responsible for keeping the permission view current. That is not a criticism of any of them. It is the contract, and it is written down. The failure is in reading the marketing page instead of the contract.

The GA versus permissions squeeze

A second pattern is worth naming because it is easy to miss on a roadmap slide. Permission features on retrieval platforms tend to arrive as previews, and previews can be dropped on the way to a stable API.

Microsoft's agentic retrieval API reached its first stable version, `2026-04-01`, this year. The migration guide is precise about what that version carries: "`2026-04-01` is the first stable API version for agentic retrieval," and for the blob and OneLake knowledge sources, "omit `ingestionPermissionOptions` from `ingestionParameters`. This property isn't supported in `2026-04-01`." Sending it returns a `400`. The same guide lists which knowledge source types are generally available in that version and notes that the others remain in preview (Microsoft Learn, agentic retrieval migration guide, dated 20 August 2026, read 17 September 2026).

So a team that wants the stable API and native document-level permission ingestion on those sources has to choose one, today. That will change; previews graduate. The point is not that it is permanent. The point is that a procurement decision made on the preview surface and a deployment made on the stable surface can be two different products, and the difference is the permissions.

When native is enough

This section matters more than the three before it, because the honest answer to "do I need anything beyond my platform" is often no.

Where a platform offers a live, non-indexed path, use it. Microsoft's own guidance for SharePoint says so plainly: "For scenarios that require the full SharePoint permissions model, sensitivity labels, and out-of-the-box security trimming, use a remote SharePoint knowledge source. This approach calls SharePoint directly via the Copilot retrieval API. Governance remains fully in SharePoint, and query results automatically respect all applicable permissions and labels" (Microsoft Learn, SharePoint indexer permission metadata page, read 17 September 2026). Content is not indexed, so there is no copy to go stale. If your whole estate is Microsoft 365 and your retrieval runs through that path, the staleness problem in this post does not apply to you.

Microsoft also fails closed when its permission check cannot run: "If ACL evaluation fails (for example, the Graph API is unavailable), the service returns 5xx and does not return a partially filtered result set" (Microsoft Learn, query-time ACL and RBAC enforcement, read 17 September 2026). That is the right behaviour, and it is worth saying that a vendor did it right before saying anything else.

The staleness problem returns the moment retrieval leaves the live path. Three situations put you there: the content is indexed rather than served live, which is every vector database and most search services; the content spans more than one system, so no single platform's permission model covers it; or the caller is an agent or a service rather than a signed-in user, so there may be no user token on the request at all. If any of those describe your retrieval, the copy-and-filter design is what you are running, whether or not you chose it, and the questions in the next section are the ones to ask.

What to ask your vendor

These are the questions we would want asked of us. Each has a short, checkable answer, and a vendor who cannot give one is telling you something.

**When a permission changes at the source, what is the maximum time before retrieval reflects it, and what has to happen for it to do so?** Accept a number and a mechanism. "Immediately" without a mechanism is a marketing answer.

**Does that number differ for an item, a parent scope, and a group membership?** It almost always does.

**What happens when the permission check itself cannot run?** The only acceptable answers are an error or an empty result.

**Is the permission feature you are demonstrating on the API version you would deploy?** Ask for the API version of every feature in the demo, and whether it is generally available on that version.

**Who names the user on each request, and what happens if they name the wrong one?** Every retrieval is performed as someone. Find out whether that someone is verified or asserted, and by whom.

Where Gateco sits, stated narrowly

Gateco is the second design. It does not copy permissions into an index. It resolves the caller's principal at query time against a synced directory record and evaluates policy on every candidate before any content is returned. A principal that has been deactivated is refused on the next retrieval, and a policy evaluation that fails denies rather than degrades. That is the default, not a setting.

Two limits belong in the same paragraph. First, Gateco enforces policy against the principal the calling application names. By default it trusts that name. Gateco can be configured to refuse any retrieval not accompanied by a verified end-user token, in which case the subject is a required, auditable parameter, and with verification on, a token naming someone else is refused, not served. That mode is opt-in; the caller still chooses the principal unless the organization turns verification on. Second, Gateco does not perform on-behalf-of token exchange. It validates a token the caller already holds for the user and never mints one.

If your whole estate is one system with a live permission path, use that path. Gateco is for the retrieval that happens outside it.


Ready to secure your AI retrieval?

Start with the free tier: 1,000 retrievals/month, no credit card required.