To build retrieval-augmented generation (RAG) over SharePoint without leaking documents, retrieve with the end user's identity rather than an all-powerful app identity. The simplest option is the Microsoft 365 Copilot Retrieval API, which returns text extracts from SharePoint already security-trimmed for the signed-in user. If you need your own index, use the Azure AI Search SharePoint indexer with ACL ingestion (preview) and pass the user's Microsoft Entra token in the x-ms-query-source-authorization header on every query, so the search service filters out anything the user can't open before your model ever sees it.
Who this is for and what you will have at the end
This guide is for developers and architects building an internal assistant, chatbot or agent that answers questions from SharePoint libraries, and for the Microsoft 365 administrators who have to approve the app registration.
At the end you will:
- Understand why the common "crawl everything with an app registration, then answer" design leaks data.
- Be able to choose between retrieving in place and indexing with permissions.
- Have working request shapes for the Copilot Retrieval API, a remote SharePoint knowledge source, and an ACL-aware Azure AI Search index.
- Know the preview limitations that decide whether the indexed approach is acceptable for your content.
The failure mode to avoid
Many first RAG prototypes look like this:
App registration (Sites.Read.All, application)
|
v
Crawl every SharePoint library --> chunk --> embed --> one shared vector index
|
User question --> vector search (no user identity) -------+--> LLM --> answerEvery user now queries a corpus assembled with the app's permissions. A question from an intern can retrieve a chunk from a board paper because nothing in the query path knows who the intern is. Prompt instructions such as "only answer from documents the user may see" don't help; the model can't check permissions.
Permission trimming has to happen in retrieval, using the caller's identity, before any text reaches the model. Microsoft offers three supported ways to do that for SharePoint.
Choose a retrieval pattern
| Copilot Retrieval API | Remote SharePoint knowledge source (Azure AI Search) | Indexed SharePoint with ACL ingestion (Azure AI Search) | |
|---|---|---|---|
| Where content lives | Stays in Microsoft 365 | Stays in Microsoft 365 | Copied into your search index |
| How trimming works | Microsoft 365 trims results for the delegated user | Passes the user token to the Retrieval API | Stored ACL metadata compared with the user's Entra token at query time |
| Permission model | Full SharePoint permissions and sensitivity labels | Full SharePoint permissions and sensitivity labels | Basic ACLs; some link types and guests not supported |
| Freshness | Live | Live | As fresh as the last indexer run or resync |
| Status | Graph v1.0 endpoint; Copilot APIs terms of use apply | Preview | Preview |
| Licensing | Microsoft 365 Copilot licence per user, or pay-as-you-go (preview) for SharePoint | Same as the Retrieval API, plus Azure AI Search | Azure AI Search, Basic tier or higher |
| Main limits | 25 results per query, 200 requests per user per hour, 1,500-character query | Retrieval API limits plus Azure AI Search concurrency | 1,000 permission entries per file, preview restrictions |
Microsoft's own guidance for the indexed approach says that when you need the full SharePoint permissions model, sensitivity labels and out-of-the-box trimming, you should use the remote knowledge source, which calls SharePoint through the Retrieval API. Use your own index when you need control over chunking, embeddings, ranking or content from other sources in the same index, and when the preview limitations are acceptable.
Prerequisites
- An app registration in Microsoft Entra ID for your assistant, with a sign-in flow that produces a delegated token for the user.
- For the Retrieval API: the Files.Read.All and Sites.Read.All delegated Microsoft Graph permissions with admin consent, and either a Microsoft 365 Copilot add-on licence for each user or Retrieval API pay-as-you-go consumption enabled. Pay-as-you-go covers tenant-level sources such as SharePoint, not OneDrive.
- For Azure AI Search: a search service on the Basic tier or higher for the indexer, in a region that supports agentic retrieval if you use knowledge sources, and the latest preview REST API (
2026-08-01-previewin the examples here) or a preview SDK. - A model deployment for answer generation, such as an Azure OpenAI deployment.
- Permission hygiene in the source. RAG faithfully reproduces whatever access SharePoint grants, so clean up broad grants first; see Prepare a tenant for Microsoft 365 Copilot by fixing oversharing first.
Pattern 1: Retrieve in place with the Copilot Retrieval API
The Retrieval API returns relevant text extracts from the same hybrid index that powers Microsoft 365 Copilot. You don't build crawlers, parsers, chunking or a vector store, and the API "security trims content for the calling user and respects the defined access controls within the tenant".
Call it with the user's delegated Graph token:
POST https://graph.microsoft.com/v1.0/copilot/retrieval
Authorization: Bearer {user-delegated-token}
Content-Type: application/json
{
"queryString": "What is the approval limit for capital expenditure?",
"dataSource": "sharePoint",
"filterExpression": "path:\"https://contoso.sharepoint.com/sites/Finance/\"",
"resourceMetadata": [ "title", "author" ],
"maximumNumberOfResults": 10
}Each item in retrievalHits contains a webUrl, one or more extracts with text and relevanceScore, the requested resourceMetadata, and, where present, a sensitivityLabel object with the label ID, display name and priority. Use webUrl for citations and the label to decide how to display or handle the answer.
Scoping with filterExpression uses Keyword Query Language over a fixed set of SharePoint and OneDrive properties: Author, FileExtension, Filename, FileType, InformationProtectionLabelId, LastModifiedTime, ModifiedBy, Path, SiteID and Title. For example, to limit retrieval to recent Word and PDF files:
{
"filterExpression": "(FileExtension:\"docx\" OR FileExtension:\"pdf\") AND LastModifiedTime>= 2026-01-01"
}Things that catch teams out:
- Invalid KQL is silently ignored. If the filter syntax is wrong, the query runs with no scoping. Validate filters in testing by checking that every returned
webUrlmatches the scope you intended. - Results are unordered. Microsoft recommends not reducing
maximumNumberOfResultsunless you have strict token limits, and sending all returned extracts to the model. - One data source per call. Use
sharePoint,oneDriveBusinessorexternalItem(Copilot connectors). To combine them, batch requests with$batch; the API accepts up to 20 requests per batch. - Path filters must use the path from the item's Details pane in SharePoint, not a sharing link or the browser address bar.
- File coverage. Semantic and hybrid retrieval apply to
.doc,.docx,.pptx,.pdf,.aspxand.one; other types get lexical retrieval only. Images and charts aren't retrieved, and very large files (over 512 MB for .docx, .pptx and .pdf; over 150 MB for others) aren't supported. - Queries should be a single, context-rich sentence of up to 1,500 characters.
Then build the prompt for your model from the extracts, cite each webUrl, and instruct the model to answer only from the supplied text. Because the extracts were already trimmed for this user, the model can only repeat what the user could have opened.
Pattern 2: Use a remote SharePoint knowledge source
If you use Azure AI Search agentic retrieval or Foundry IQ knowledge bases, a remote SharePoint knowledge source gives you the same in-place retrieval without writing Graph calls. It uses the Copilot Retrieval API behind the scenes, needs no index or connection string, and honors SharePoint permissions and Purview sensitivity labels when you pass the user's token.
PUT {search-endpoint}/knowledgesources/finance-sharepoint-ks?api-version=2026-08-01-preview
Authorization: Bearer {search-access-token}
Content-Type: application/json
{
"name": "finance-sharepoint-ks",
"kind": "remoteSharePoint",
"description": "Finance policies in SharePoint",
"remoteSharePointParameters": {
"filterExpression": "Path:\"https://contoso.sharepoint.com/sites/Finance\"",
"resourceMetadata": [ "Author", "Title" ]
}
}Add the knowledge source to a knowledge base, then query it with the user's token in x-ms-query-source-authorization:
POST {search-endpoint}/knowledgebases/finance-kb/retrieve?api-version=2026-08-01-preview
Authorization: Bearer {search-access-token}
x-ms-query-source-authorization: {user-access-token}
Content-Type: application/json
{
"messages": [
{ "role": "user", "content": [ { "type": "text", "text": "capital expenditure approval limits" } ] }
],
"knowledgeSourceParams": [
{ "knowledgeSourceName": "finance-sharepoint-ks", "kind": "remoteSharePoint" }
]
}The SharePoint tenant must be in the same Microsoft Entra tenant as Azure, OneDrive and Copilot connectors aren't supported through this source, and the Retrieval API limits still apply. On a dedicated search service, each replica runs one remote SharePoint query at a time, so add replicas if retrieval latency rises under load.
Pattern 3: Index SharePoint with ACLs in Azure AI Search
When you need your own index, the SharePoint indexer can store each item's effective permissions alongside the content (preview). At query time the service compares the user's identity with those stored IDs.
Register the ingestion app
ACL ingestion requires application permissions; delegated permissions aren't supported for the indexer. Add permissions under API permissions > Add a permission and grant admin consent:
| What you index | Permissions | Credential |
|---|---|---|
| Library files shared to Entra users and groups | Microsoft Graph: Files.Read.All, Sites.FullControl.All (or Sites.Selected) | Client secret or federated credential |
| Library files where SharePoint site groups must be honored | As above, plus SharePoint: Sites.FullControl.All (or Sites.Selected) | Federated credential |
| List items or ASPX site pages | Graph and SharePoint permissions above, plus Graph User.Read.All | Federated credential |
If you use Sites.Selected, grant the app access to each target site explicitly. The app reads everything so that it can record who may see it; that is acceptable because trimming happens at query time, but restrict who can query the index without a user token.
Enable permission metadata
In the data source, ask for user and group IDs:
{
"name": "finance-sharepoint-acl-ds",
"type": "sharepoint",
"indexerPermissionOptions": ["userIds", "groupIds"],
"credentials": { "connectionString": "<connection-string>;" },
"container": { "name": "<library-name>" }
}In the index, add filterable permission fields and turn on permission filtering:
{
"fields": [
{ "name": "UserIds", "type": "Collection(Edm.String)", "permissionFilter": "userIds", "filterable": true, "retrievable": false },
{ "name": "GroupIds", "type": "Collection(Edm.String)", "permissionFilter": "groupIds", "filterable": true, "retrievable": false }
],
"permissionFilterOption": "enabled"
}Map metadata_user_ids to UserIds and metadata_group_ids to GroupIds in the indexer field mappings. If your skillset chunks documents for vectorization with projectionMode set to skipIndexingParentDocuments, field mappings are bypassed; project the ACL fields onto every chunk through indexProjections instead. A chunk without ACL fields can't be returned to the right user. To honor SharePoint site groups (Owners, Members, Visitors), also add the sharePointConnectorAppRegistration block and a SharePointSiteUrl field mapped from metadata_spo_site_url; those group IDs appear with an spg: prefix.
Query with the user's token
POST {search-endpoint}/indexes/finance-index/docs/search?api-version=2026-08-01-preview
Authorization: Bearer {app-or-user-token}
x-ms-query-source-authorization: {user-access-token}
Content-Type: application/json
{
"search": "capital expenditure approval limit",
"select": "title,content,url",
"top": 10
}The Authorization header proves your app may read the index (Search Index Data Reader or Search Index Data Contributor). The x-ms-query-source-authorization header carries the user; the service resolves the user's object ID and transitive Entra group memberships through Microsoft Graph and returns only documents where at least one of them matches. If the header is missing, only documents public to everyone come back. If Graph can't be reached, the query fails with a 5xx rather than returning partially filtered results.
Keep permissions in sync
- Item-level permission changes are picked up on the next successful indexer run (2026-05-01-preview and later).
- Changes at site, library, list or folder level aren't detected. Call
POST /indexers/{indexer}/resyncwith{"options": ["permissions"]}and then run the indexer, or use/resetdocsfor specific documents. - Schedule the resync after any known permission cleanup, such as a site access review.
Know what isn't honored
In the current preview, "Anyone" and "People in your organization" links, guest users and Information Management policies aren't supported, and a Microsoft Entra group nested inside a SharePoint group isn't expanded. Each file supports up to 1,000 permission entries; entries beyond that might not be enforced. If your content relies on any of these, use pattern 1 or 2.
Verify trimming before go-live
- Pick two test users with different access, for example a finance member and a non-member.
- Ask the same question as each user. The non-member must get no extracts or citations from restricted documents.
- For indexed content, temporarily set
retrievabletotrueon the ACL fields and run an elevated-read query (headerx-ms-enable-elevated-read: true, which needs Search Index Data Contributor or a custom role) to confirm every chunk has populatedUserIdsandGroupIds. Setretrievableback tofalseafterwards. - Remove the test user from the site, run the resync if the change was at site level, and confirm results disappear.
- Check that any answer cache is keyed per user. A shared semantic cache can replay one user's answer to another and undo all the trimming above; the caching guidance in production LLMOps and enterprise RAG architecture needs this adjustment for permissioned content.
Troubleshooting
| Symptom | Cause and fix |
|---|---|
| Retrieval API returns results from outside the intended site | The filterExpression has invalid KQL and was ignored. Fix quoting and property names. |
| Retrieval API calls are denied | Missing delegated Files.Read.All and Sites.Read.All, or the user has no Copilot licence and pay-as-you-go isn't enabled. |
| Throttling during load tests | The Retrieval API allows 200 requests per user per hour. Test with several users, cache retrieval per user, or batch. |
UserIds or GroupIds empty in the index | Your skillset chunks with skipIndexingParentDocuments, so field mappings were bypassed. Map ACLs in indexProjections. |
| Indexer returns 401 or 403 | Admin consent missing on Graph or SharePoint permissions, or a client secret used where a federated credential is required. |
| User still sees a document after losing access | The change was at a parent scope. Run /resync with the permissions option. |
| User can't find a document shared with them via an organization link | Organization-wide and Anyone links aren't honored by ACL ingestion. Grant access through a group or use pattern 1. |
| No results at all for every user | The app isn't sending x-ms-query-source-authorization, so only public documents return. |
Summary checklist
- Retrieval always runs with the end user's identity; no app-only path reaches the model.
- Pattern chosen: Retrieval API or remote knowledge source for full SharePoint semantics; ACL-aware index only where preview limits are acceptable.
- Delegated Graph permissions consented for the Retrieval API; ingestion app permissions and credential match the ACL scenario.
- ACL fields present on every chunk;
permissionFilterOptionenabled. - Resync scheduled for parent-scope permission changes.
- Two-user trimming test, ACL field check and per-user caching verified before go-live.
References
- https://learn.microsoft.com/en-us/microsoft-365/copilot/extensibility/api/ai-services/retrieval/overview
- https://learn.microsoft.com/en-us/microsoft-365/copilot/extensibility/api/ai-services/retrieval/copilotroot-retrieval
- https://learn.microsoft.com/en-us/azure/search/search-document-level-access-overview
- https://learn.microsoft.com/en-us/azure/search/search-indexer-sharepoint-access-control-lists
- https://learn.microsoft.com/en-us/azure/search/search-query-access-control-rbac-enforcement
- https://learn.microsoft.com/en-us/azure/search/agentic-knowledge-source-how-to-sharepoint-remote