Back to blog
10 min readGateco Team

RAG over-retrieval / entitlement-failure register

Public register of documented incidents in which an AI assistant, agent, or RAG-backed system surfaced (or enabled extraction of) content a principal was not entitled to see, or amplified access beyond intended entitlement. Every row links to a public primary source. The register makes no product claims, competitive comparisons, or vendor recommendations.

Last checked: 2026-09-30 (IDT)

Purpose

Enterprises deploying assistants and RAG often assume retrieval will respect intended access rights. This register collects documented, sourced cases where that assumption failed or was shown to fail under realistic conditions: permission oversharing amplified by natural-language retrieval, missing ACL checks on vector search, prompt injection that caused privileged tools to fetch private data, or misconfigured knowledge bases and AI agents that exposed internal content.

Methodology

RuleDetail
SourcesPublic web only. Prefer primary: vendor advisories, vendor security blogs admitting the issue, GitHub Security Advisories / CVE records, reputable security-research write-ups with reproducible demos, major press citing named primaries.
InclusionDocumented entitlement / over-retrieval / unauthorized surfacing of private or restricted content via an AI assistant, agent, MCP tool path, RAG vector query, or knowledge base that is commonly indexed into AI. Root cause must be stated by the source (quoted or closely paraphrased).
ExclusionSpeculation; vendor marketing as sole evidence; pure training-data memorization unless the source frames unauthorized retrieval of private content; consumer paste-into-chatbot policy incidents without a retrieval/ACL failure; named individuals as subjects.
Row standardEvery register row is FACT attributable to the cited primary (and optional secondary). No inference in table cells.
Date fieldPrefer public disclosure / advisory publication date; note fix date in root-cause text when the source gives one.
Organisation fieldCompany or product only, no named individuals.

Register

#DateOrganisation (company / product)What was exposedRoot cause as stated by the sourcePrimary URLSecondary coverage URL
12023-11-29OpenAI Custom GPTs (ChatGPT)Uploaded knowledge files and system / builder instructions from custom GPTs (research: ~100% file-leakage success and ~97% system-prompt extraction across 200+ tested GPTs)Prompt injection / jailbreak of custom GPTs such that uploaded knowledge and builder instructions could be listed and downloaded; OpenAI stated it takes privacy seriously and continuously hardens against adversarial attacks including prompt injectionhttps://www.wired.com/story/openai-custom-chatbots-gpts-prompt-injection-attacks/https://arxiv.org/abs/2311.11538
22024-08-07Microsoft Copilot StudioSensitive corporate data (research cited examples including legal documents) from Copilot Studio bots reachable without authentication; researchers reported finding large numbers of publicly discoverable bots, including at large enterprisesMisconfiguration / insecure defaults: bots publicly accessible without auth and connected to enterprise data sources; Zenity Black Hat research summarized finding >1K such bots and ability to extract sensitive data; Dark Reading reported overpermissioned defaults and public bot discoveryhttps://labs.zenity.io/p/summary-zenity-research-published-blackhat-2024https://www.darkreading.com/application-security/creating-insecure-ai-assistants-microsoft-copilot-studio
32024-08-20Slack AI (Salesforce Slack)Content from private Slack channels the attacker could not read (research demo: an API key), exfiltrated via a rendered link when the victim queried Slack AIIndirect prompt injection via messages in public channels ingested into Slack AI context together with private-channel content the querying user could access; Slack acknowledged a researcher report and deployed a patch on 2024-08-20 addressing phishing of users for certain data under limited circumstanceshttps://www.promptarmor.com/resources/data-exfiltration-from-slack-ai-via-indirect-prompt-injectionhttps://slack.com/blog/news/slack-security-update-082124
42024-09-17ServiceNow Knowledge BasesKnowledge Base article content including PII, internal system details, and credentials / tokens to production systems, AppOmni reported >1,000 enterprise instances (~45% of those tested) unintentionally exposing KB dataMisconfigured KB User Criteria and public widgets allowing unauthenticated access; 2023 ACL hardening did not effectively protect KBs that rely on User Criteria; legacy default allowing public access without User Criteria on older instances; ServiceNow stated it contacted customers and began proactive KB configuration actions from 2024-09-04https://appomni.com/ao-labs/servicenow-knowledge-bases-data-exposures-uncovered/https://www.bleepingcomputer.com/news/security/over-1-000-servicenow-instances-found-leaking-corporate-kb-data/
52025-05-22GitLab DuoPrivate project source code (and, per researchers, confidential issue content) retrieved and exfiltrated through Duo Chat responsesRemote prompt injection via hidden prompts in MRs/issues/commits/code plus HTML injection in streamed Duo responses enabling browser-side exfiltration; Duo operated with the victim user’s permissions including private projects; GitLab confirmed and remediated (duo-ui HTML sanitization preventing unsafe external image/script tags)https://www.legitsecurity.com/blog/remote-prompt-injection-in-gitlab-duohttps://www.csoonline.com/article/3992845/prompt-injection-flaws-in-gitlab-duo-highlights-risks-in-ai-assistants.html
62025-05-26GitHub MCP (official GitHub MCP server)Contents of private repositories (demo: private project details and other private-repo data) written into a public pull requestIndirect prompt injection from a malicious public GitHub Issue; agent with GitHub MCP access across public and private repos treated issue text as instructions, pulled private data, and published it publicly, described by researchers as an architectural “toxic agent flow,” not a conventional server code bughttps://invariantlabs.ai/blog/mcp-github-vulnerabilityhttps://simonwillison.net/2025/May/26/github-mcp-exploited/
72025-07-06Supabase MCPPrivate SQL table contents (demo: integration_tokens / OAuth-style secrets) written into a customer-visible support ticketIndirect prompt injection in customer support messages plus assistant use of Supabase MCP with service_role (bypasses RLS); agent read private tables and inserted results into attacker-visible support messages (“lethal trifecta”: private data + untrusted instructions + outbound channel)https://generalanalysis.com/blog/supabase-mcp-bloghttps://simonwillison.net/2025/Jul/6/supabase-mcp-lethal-trifecta/
82025-07-07Microsoft Copilot Studio (AgentFlayer research)Full knowledge-source file contents (demo CSV of account owners) and Salesforce CRM Account records emailed to an attackerZero-click prompt injection via an email trigger: agent instructed to read named knowledge sources / invoke CRM tools and send results to attacker-controlled address; reported to MSRC 2025-02-21; Microsoft issued a fix 2025-04-24 (prompt-shielding / classifier mitigations per researchers)https://labs.zenity.io/post/a-copilot-studio-story-2-when-aijacking-leads-to-full-data-exfiltration-bc4ahttps://zenity.io/research/agentflayer-vulnerabilities
92025-10-08GitHub Copilot Chat (CamoLeak)Secrets and source code from private repositories (including searching private codebases for secret patterns)Hidden Markdown comments in PRs/issues injected instructions into Copilot Chat; Copilot ran with victim permissions; exfiltration via pre-signed Camo image URLs as a CSP bypass; GitHub fixed by disabling image rendering in Copilot Chat as of 2025-08-14 (CVSS 9.6 per researchers)https://www.legitsecurity.com/blog/camoleak-critical-github-copilot-vulnerability-leaks-private-source-codehttps://www.securityweek.com/github-copilot-chat-flaw-leaked-data-from-private-repositories/
102026-01-24AnythingLLM (Mintplex Labs)Vector-database credentials, and through them the RAG knowledge-base chunks stored in Qdrant (similar impact noted for Weaviate)Unauthenticated /api/setup-complete returned QdrantApiKey in plaintext when Qdrant was configured with an API key; leaked key allowed full read/write of the vector DB holding RAG embeddings (GHSA-gm94-qc2p-xcwf / CVE-2026-24477; fixed in 1.10.0)https://github.com/Mintplex-Labs/anything-llm/security/advisories/GHSA-gm94-qc2p-xcwfhttps://www.cve.org/CVERecord?id=CVE-2026-24477
112026-05-05Open WebUIPrivate uploaded files and knowledge-base content via RAG vector search after access should have been deniedget_sources_from_items RAG paths queried vector collections without authorization checks on several code paths; knowing a file/KB ID allowed continued extraction after revocation (GHSA-h36f-rqpx-j5wx / CVE-2026-44560; patched ≥0.9.0)https://github.com/open-webui/open-webui/security/advisories/GHSA-h36f-rqpx-j5wxhttps://nvd.nist.gov/vuln/detail/CVE-2026-44560
122026-06-11Open WebUI (Milvus multitenancy)Private knowledge-base chunks belonging to other usersBypass of collection-level ACL when ENABLE_MILVUS_MULTITENANCY_MODE=true: user-controlled collection_names interpolated unsafely into Milvus filter expressions, turning an “unknown collection” into a tautology that returned other users’ chunks (GHSA-p5cp-r7rg-qpxc / CVE-2026-54019; patched ≥0.9.6)https://github.com/open-webui/open-webui/security/advisories/GHSA-p5cp-r7rg-qpxchttps://osv.dev/vulnerability/CVE-2026-54019

Documented oversharing class (vendor-acknowledged; not a single named customer breach)

#DateOrganisation (company / product)What was exposedRoot cause as stated by the sourcePrimary URLSecondary coverage URL
132025–2026 (ongoing Microsoft guidance; e.g. Learn “secure and governed foundation” material current as of last check)Microsoft 365 CopilotContent on overshared SharePoint / OneDrive / Teams sites and files (including “Everyone except external users” / oversized audiences) that users technically have permission to open, surfaced via natural-language promptsPer Microsoft: Copilot uses data the user already has permission to access; oversharing (broad sharing links, EEEU, broken inheritance, unlabeled sensitive content) is amplified because Copilot makes existing access discoverable and efficient; Microsoft documents remediation via Purview DSPM, SharePoint Advanced Management, Restricted Content Discovery, and DLP for Copilothttps://learn.microsoft.com/en-us/microsoft-365/copilot/configure-secure-governed-data-foundation-microsoft-365-copilothttps://techcommunity.microsoft.com/blog/microsoft365copilotblog/mitigate-oversharing-to-govern-microsoft-365-copilot-and-agents/4448744

Note on row 13 (FACT from Microsoft): Microsoft states Copilot does not elevate permissions; the documented failure mode is permission oversharing + retrieval amplification, not a classic ACL-bypass bug. Included because Microsoft's own documentation treats oversharing as the primary Copilot data-security risk class. No named customer breach is asserted here.

Note on row 4 (FACT): ServiceNow KB exposure is a knowledge-base entitlement failure (widgets / User Criteria), not an LLM bug. Included because knowledge bases are a common retrieval source for enterprise AI connectors; ServiceNow’s own statement confirms unintended KB access from misconfiguration.

Excluded / needs better primary

CandidateWhy excluded or deferred
Samsung semiconductor engineers pasting code into ChatGPT (2023)Consumer chatbot paste / policy incident; sources frame outbound data leaving the company, not RAG or assistant retrieval of content the user was not entitled to.
Dropbox DashPublic materials describe permission-respecting design and admin oversharing controls; no documented entitlement-over-retrieval incident with a primary source found.
ChatGPT shared chats indexed by Google (2025 “discoverable” share option)User-initiated share / discoverability feature; privacy backlash, not unauthorized RAG retrieval across entitlements.
ConfusedPilot (arXiv 2408.04870)Strong on RAG integrity / document poisoning and confused-deputy risks against Microsoft 365 Copilot; confidentiality claims exist but the paper is not a clean “unauthorized retrieval of content the user lacked entitlement to” incident for this register.
Amazon Q Developer VS Code extension supply-chain (AWS-2025-015)Malicious code injection into the extension, not a retrieval/ACL entitlement failure.
GitHub Copilot training-data / memorization code-suggestion casesExcluded unless a primary source frames unauthorized retrieval of private content (not training memorization). CamoLeak and GitHub MCP rows cover private-repo retrieval paths instead.
Consultant / CISO “playbook” blogs claiming unnamed Fortune-500 Copilot oversharing anecdotesNo verifiable primary; names no organisation or produces no vendor/court/researcher primary.

Optional analysis (tagged)

ANALYSIS (not register FACT): Across the included rows, failure modes cluster into four patterns: (1) missing or bypassable ACL at the RAG/vector layer (Open WebUI, AnythingLLM vector-key leak); (2) prompt injection that hijacks a privileged assistant/agent so it fetches private data the attacker cannot read directly (Slack AI, GitLab Duo, CamoLeak, GitHub MCP, Supabase MCP, Copilot Studio AgentFlayer); (3) misconfigured public exposure of knowledge or agents (ServiceNow KB, Copilot Studio public bots); (4) faithful retrieval over overly broad existing permissions (Microsoft 365 Copilot oversharing class). Patterns (1)–(3) are unauthorized relative to intended policy; pattern (4) is authorized-by-ACL but unintended-by-business-entitlement.

UNKNOWN: How often pattern (4) produces named, public customer breaches versus silent internal incidents, public primary evidence of named customer Copilot oversharing breaches remains thin as of last check.

Sources appendix

Primaries used in register rows

  1. WIRED, OpenAI custom GPTs leaking knowledge/instructions (2023-11-29): https://www.wired.com/story/openai-custom-chatbots-gpts-prompt-injection-attacks/
  2. Zenity Labs, Black Hat 2024 research summary (Copilot Studio public bots): https://labs.zenity.io/p/summary-zenity-research-published-blackhat-2024
  3. PromptArmor, Slack AI data exfiltration via indirect prompt injection: https://www.promptarmor.com/resources/data-exfiltration-from-slack-ai-via-indirect-prompt-injection
  4. Slack, Security Update (2024-08-21): https://slack.com/blog/news/slack-security-update-082124
  5. AppOmni AO Labs, ServiceNow Knowledge Bases data exposures (2024-09-17): https://appomni.com/ao-labs/servicenow-knowledge-bases-data-exposures-uncovered/
  6. Legit Security, Remote prompt injection in GitLab Duo: https://www.legitsecurity.com/blog/remote-prompt-injection-in-gitlab-duo
  7. Invariant Labs, GitHub MCP exploited (private repos): https://invariantlabs.ai/blog/mcp-github-vulnerability
  8. General Analysis, Supabase MCP can leak entire SQL database: https://generalanalysis.com/blog/supabase-mcp-blog
  9. Zenity Labs, AgentFlayer Copilot Studio knowledge/CRM exfiltration: https://labs.zenity.io/post/a-copilot-studio-story-2-when-aijacking-leads-to-full-data-exfiltration-bc4a
  10. Legit Security, CamoLeak GitHub Copilot Chat: https://www.legitsecurity.com/blog/camoleak-critical-github-copilot-vulnerability-leaks-private-source-code
  11. Mintplex Labs / AnythingLLM, GHSA-gm94-qc2p-xcwf: https://github.com/Mintplex-Labs/anything-llm/security/advisories/GHSA-gm94-qc2p-xcwf
  12. Open WebUI, GHSA-h36f-rqpx-j5wx: https://github.com/open-webui/open-webui/security/advisories/GHSA-h36f-rqpx-j5wx
  13. Open WebUI, GHSA-p5cp-r7rg-qpxc: https://github.com/open-webui/open-webui/security/advisories/GHSA-p5cp-r7rg-qpxc
  14. Microsoft Learn, Configure a secure and governed foundation for Microsoft Copilot: https://learn.microsoft.com/en-us/microsoft-365/copilot/configure-secure-governed-data-foundation-microsoft-365-copilot

Secondaries / corroboration

Corrections and additions with a public primary source are welcome.


Ready to secure your AI retrieval?

Start with the free tier: 1,000 retrievals/month, no credit card required.