AI engineering

Build RAG Over SharePoint Documents Without Breaking Permissions

Ground an internal AI assistant on SharePoint files so each user only gets answers from documents they can open, using the Copilot Retrieval API or Azure AI Search with ACL ingestion.

12 min read
On this page

To build retrieval-augmented generation (RAG) over SharePoint without leaking documents, retrieve with the end user's identity rather than an all-powerful app identity. The simplest option is the Microsoft 365 Copilot Retrieval API, which returns text extracts from SharePoint already security-trimmed for the signed-in user. If you need your own index, use the Azure AI Search SharePoint indexer with ACL ingestion (preview) and pass the user's Microsoft Entra token in the x-ms-query-source-authorization header on every query, so the search service filters out anything the user can't open before your model ever sees it.

Who this is for and what you will have at the end

This guide is for developers and architects building an internal assistant, chatbot or agent that answers questions from SharePoint libraries, and for the Microsoft 365 administrators who have to approve the app registration.

At the end you will:

  • Understand why the common "crawl everything with an app registration, then answer" design leaks data.
  • Be able to choose between retrieving in place and indexing with permissions.
  • Have working request shapes for the Copilot Retrieval API, a remote SharePoint knowledge source, and an ACL-aware Azure AI Search index.
  • Know the preview limitations that decide whether the indexed approach is acceptable for your content.

The failure mode to avoid

Many first RAG prototypes look like this:

App registration (Sites.Read.All, application)
        |
        v
Crawl every SharePoint library --> chunk --> embed --> one shared vector index
                                                          |
User question --> vector search (no user identity) -------+--> LLM --> answer

Every user now queries a corpus assembled with the app's permissions. A question from an intern can retrieve a chunk from a board paper because nothing in the query path knows who the intern is. Prompt instructions such as "only answer from documents the user may see" don't help; the model can't check permissions.

Permission trimming has to happen in retrieval, using the caller's identity, before any text reaches the model. Microsoft offers three supported ways to do that for SharePoint.

Choose a retrieval pattern

Copilot Retrieval APIRemote SharePoint knowledge source (Azure AI Search)Indexed SharePoint with ACL ingestion (Azure AI Search)
Where content livesStays in Microsoft 365Stays in Microsoft 365Copied into your search index
How trimming worksMicrosoft 365 trims results for the delegated userPasses the user token to the Retrieval APIStored ACL metadata compared with the user's Entra token at query time
Permission modelFull SharePoint permissions and sensitivity labelsFull SharePoint permissions and sensitivity labelsBasic ACLs; some link types and guests not supported
FreshnessLiveLiveAs fresh as the last indexer run or resync
StatusGraph v1.0 endpoint; Copilot APIs terms of use applyPreviewPreview
LicensingMicrosoft 365 Copilot licence per user, or pay-as-you-go (preview) for SharePointSame as the Retrieval API, plus Azure AI SearchAzure AI Search, Basic tier or higher
Main limits25 results per query, 200 requests per user per hour, 1,500-character queryRetrieval API limits plus Azure AI Search concurrency1,000 permission entries per file, preview restrictions

Microsoft's own guidance for the indexed approach says that when you need the full SharePoint permissions model, sensitivity labels and out-of-the-box trimming, you should use the remote knowledge source, which calls SharePoint through the Retrieval API. Use your own index when you need control over chunking, embeddings, ranking or content from other sources in the same index, and when the preview limitations are acceptable.

Prerequisites

  • An app registration in Microsoft Entra ID for your assistant, with a sign-in flow that produces a delegated token for the user.
  • For the Retrieval API: the Files.Read.All and Sites.Read.All delegated Microsoft Graph permissions with admin consent, and either a Microsoft 365 Copilot add-on licence for each user or Retrieval API pay-as-you-go consumption enabled. Pay-as-you-go covers tenant-level sources such as SharePoint, not OneDrive.
  • For Azure AI Search: a search service on the Basic tier or higher for the indexer, in a region that supports agentic retrieval if you use knowledge sources, and the latest preview REST API (2026-08-01-preview in the examples here) or a preview SDK.
  • A model deployment for answer generation, such as an Azure OpenAI deployment.
  • Permission hygiene in the source. RAG faithfully reproduces whatever access SharePoint grants, so clean up broad grants first; see Prepare a tenant for Microsoft 365 Copilot by fixing oversharing first.

Pattern 1: Retrieve in place with the Copilot Retrieval API

The Retrieval API returns relevant text extracts from the same hybrid index that powers Microsoft 365 Copilot. You don't build crawlers, parsers, chunking or a vector store, and the API "security trims content for the calling user and respects the defined access controls within the tenant".

Call it with the user's delegated Graph token:

POST https://graph.microsoft.com/v1.0/copilot/retrieval
Authorization: Bearer {user-delegated-token}
Content-Type: application/json
 
{
  "queryString": "What is the approval limit for capital expenditure?",
  "dataSource": "sharePoint",
  "filterExpression": "path:\"https://contoso.sharepoint.com/sites/Finance/\"",
  "resourceMetadata": [ "title", "author" ],
  "maximumNumberOfResults": 10
}

Each item in retrievalHits contains a webUrl, one or more extracts with text and relevanceScore, the requested resourceMetadata, and, where present, a sensitivityLabel object with the label ID, display name and priority. Use webUrl for citations and the label to decide how to display or handle the answer.

Scoping with filterExpression uses Keyword Query Language over a fixed set of SharePoint and OneDrive properties: Author, FileExtension, Filename, FileType, InformationProtectionLabelId, LastModifiedTime, ModifiedBy, Path, SiteID and Title. For example, to limit retrieval to recent Word and PDF files:

{
  "filterExpression": "(FileExtension:\"docx\" OR FileExtension:\"pdf\") AND LastModifiedTime>= 2026-01-01"
}

Things that catch teams out:

  • Invalid KQL is silently ignored. If the filter syntax is wrong, the query runs with no scoping. Validate filters in testing by checking that every returned webUrl matches the scope you intended.
  • Results are unordered. Microsoft recommends not reducing maximumNumberOfResults unless you have strict token limits, and sending all returned extracts to the model.
  • One data source per call. Use sharePoint, oneDriveBusiness or externalItem (Copilot connectors). To combine them, batch requests with $batch; the API accepts up to 20 requests per batch.
  • Path filters must use the path from the item's Details pane in SharePoint, not a sharing link or the browser address bar.
  • File coverage. Semantic and hybrid retrieval apply to .doc, .docx, .pptx, .pdf, .aspx and .one; other types get lexical retrieval only. Images and charts aren't retrieved, and very large files (over 512 MB for .docx, .pptx and .pdf; over 150 MB for others) aren't supported.
  • Queries should be a single, context-rich sentence of up to 1,500 characters.

Then build the prompt for your model from the extracts, cite each webUrl, and instruct the model to answer only from the supplied text. Because the extracts were already trimmed for this user, the model can only repeat what the user could have opened.

Pattern 2: Use a remote SharePoint knowledge source

If you use Azure AI Search agentic retrieval or Foundry IQ knowledge bases, a remote SharePoint knowledge source gives you the same in-place retrieval without writing Graph calls. It uses the Copilot Retrieval API behind the scenes, needs no index or connection string, and honors SharePoint permissions and Purview sensitivity labels when you pass the user's token.

PUT {search-endpoint}/knowledgesources/finance-sharepoint-ks?api-version=2026-08-01-preview
Authorization: Bearer {search-access-token}
Content-Type: application/json
 
{
  "name": "finance-sharepoint-ks",
  "kind": "remoteSharePoint",
  "description": "Finance policies in SharePoint",
  "remoteSharePointParameters": {
    "filterExpression": "Path:\"https://contoso.sharepoint.com/sites/Finance\"",
    "resourceMetadata": [ "Author", "Title" ]
  }
}

Add the knowledge source to a knowledge base, then query it with the user's token in x-ms-query-source-authorization:

POST {search-endpoint}/knowledgebases/finance-kb/retrieve?api-version=2026-08-01-preview
Authorization: Bearer {search-access-token}
x-ms-query-source-authorization: {user-access-token}
Content-Type: application/json
 
{
  "messages": [
    { "role": "user", "content": [ { "type": "text", "text": "capital expenditure approval limits" } ] }
  ],
  "knowledgeSourceParams": [
    { "knowledgeSourceName": "finance-sharepoint-ks", "kind": "remoteSharePoint" }
  ]
}

The SharePoint tenant must be in the same Microsoft Entra tenant as Azure, OneDrive and Copilot connectors aren't supported through this source, and the Retrieval API limits still apply. On a dedicated search service, each replica runs one remote SharePoint query at a time, so add replicas if retrieval latency rises under load.

When you need your own index, the SharePoint indexer can store each item's effective permissions alongside the content (preview). At query time the service compares the user's identity with those stored IDs.

Register the ingestion app

ACL ingestion requires application permissions; delegated permissions aren't supported for the indexer. Add permissions under API permissions > Add a permission and grant admin consent:

What you indexPermissionsCredential
Library files shared to Entra users and groupsMicrosoft Graph: Files.Read.All, Sites.FullControl.All (or Sites.Selected)Client secret or federated credential
Library files where SharePoint site groups must be honoredAs above, plus SharePoint: Sites.FullControl.All (or Sites.Selected)Federated credential
List items or ASPX site pagesGraph and SharePoint permissions above, plus Graph User.Read.AllFederated credential

If you use Sites.Selected, grant the app access to each target site explicitly. The app reads everything so that it can record who may see it; that is acceptable because trimming happens at query time, but restrict who can query the index without a user token.

Enable permission metadata

In the data source, ask for user and group IDs:

{
  "name": "finance-sharepoint-acl-ds",
  "type": "sharepoint",
  "indexerPermissionOptions": ["userIds", "groupIds"],
  "credentials": { "connectionString": "<connection-string>;" },
  "container": { "name": "<library-name>" }
}

In the index, add filterable permission fields and turn on permission filtering:

{
  "fields": [
    { "name": "UserIds",  "type": "Collection(Edm.String)", "permissionFilter": "userIds",  "filterable": true, "retrievable": false },
    { "name": "GroupIds", "type": "Collection(Edm.String)", "permissionFilter": "groupIds", "filterable": true, "retrievable": false }
  ],
  "permissionFilterOption": "enabled"
}

Map metadata_user_ids to UserIds and metadata_group_ids to GroupIds in the indexer field mappings. If your skillset chunks documents for vectorization with projectionMode set to skipIndexingParentDocuments, field mappings are bypassed; project the ACL fields onto every chunk through indexProjections instead. A chunk without ACL fields can't be returned to the right user. To honor SharePoint site groups (Owners, Members, Visitors), also add the sharePointConnectorAppRegistration block and a SharePointSiteUrl field mapped from metadata_spo_site_url; those group IDs appear with an spg: prefix.

Query with the user's token

POST {search-endpoint}/indexes/finance-index/docs/search?api-version=2026-08-01-preview
Authorization: Bearer {app-or-user-token}
x-ms-query-source-authorization: {user-access-token}
Content-Type: application/json
 
{
  "search": "capital expenditure approval limit",
  "select": "title,content,url",
  "top": 10
}

The Authorization header proves your app may read the index (Search Index Data Reader or Search Index Data Contributor). The x-ms-query-source-authorization header carries the user; the service resolves the user's object ID and transitive Entra group memberships through Microsoft Graph and returns only documents where at least one of them matches. If the header is missing, only documents public to everyone come back. If Graph can't be reached, the query fails with a 5xx rather than returning partially filtered results.

Keep permissions in sync

  • Item-level permission changes are picked up on the next successful indexer run (2026-05-01-preview and later).
  • Changes at site, library, list or folder level aren't detected. Call POST /indexers/{indexer}/resync with {"options": ["permissions"]} and then run the indexer, or use /resetdocs for specific documents.
  • Schedule the resync after any known permission cleanup, such as a site access review.

Know what isn't honored

In the current preview, "Anyone" and "People in your organization" links, guest users and Information Management policies aren't supported, and a Microsoft Entra group nested inside a SharePoint group isn't expanded. Each file supports up to 1,000 permission entries; entries beyond that might not be enforced. If your content relies on any of these, use pattern 1 or 2.

Verify trimming before go-live

  1. Pick two test users with different access, for example a finance member and a non-member.
  2. Ask the same question as each user. The non-member must get no extracts or citations from restricted documents.
  3. For indexed content, temporarily set retrievable to true on the ACL fields and run an elevated-read query (header x-ms-enable-elevated-read: true, which needs Search Index Data Contributor or a custom role) to confirm every chunk has populated UserIds and GroupIds. Set retrievable back to false afterwards.
  4. Remove the test user from the site, run the resync if the change was at site level, and confirm results disappear.
  5. Check that any answer cache is keyed per user. A shared semantic cache can replay one user's answer to another and undo all the trimming above; the caching guidance in production LLMOps and enterprise RAG architecture needs this adjustment for permissioned content.

Troubleshooting

SymptomCause and fix
Retrieval API returns results from outside the intended siteThe filterExpression has invalid KQL and was ignored. Fix quoting and property names.
Retrieval API calls are deniedMissing delegated Files.Read.All and Sites.Read.All, or the user has no Copilot licence and pay-as-you-go isn't enabled.
Throttling during load testsThe Retrieval API allows 200 requests per user per hour. Test with several users, cache retrieval per user, or batch.
UserIds or GroupIds empty in the indexYour skillset chunks with skipIndexingParentDocuments, so field mappings were bypassed. Map ACLs in indexProjections.
Indexer returns 401 or 403Admin consent missing on Graph or SharePoint permissions, or a client secret used where a federated credential is required.
User still sees a document after losing accessThe change was at a parent scope. Run /resync with the permissions option.
User can't find a document shared with them via an organization linkOrganization-wide and Anyone links aren't honored by ACL ingestion. Grant access through a group or use pattern 1.
No results at all for every userThe app isn't sending x-ms-query-source-authorization, so only public documents return.

Summary checklist

  • Retrieval always runs with the end user's identity; no app-only path reaches the model.
  • Pattern chosen: Retrieval API or remote knowledge source for full SharePoint semantics; ACL-aware index only where preview limits are acceptable.
  • Delegated Graph permissions consented for the Retrieval API; ingestion app permissions and credential match the ACL scenario.
  • ACL fields present on every chunk; permissionFilterOption enabled.
  • Resync scheduled for parent-scope permission changes.
  • Two-user trimming test, ACL field check and per-user caching verified before go-live.

References

Questions people ask

Can I call the Microsoft 365 Copilot Retrieval API with an app-only token?

No. The Retrieval API supports only delegated permissions for work or school accounts; application permissions aren't supported. Retrieving SharePoint or OneDrive content needs both Files.Read.All and Sites.Read.All, and results are security-trimmed for the signed-in user.

Does the Azure AI Search SharePoint indexer honor "Anyone" and "People in your organization" links?

No. In the current preview, only sharing links scoped to "Specific people" are supported. Organization-wide and Anyone links, guest users and SharePoint Information Management policies aren't honored, so content shared only through those mechanisms won't be returned to users who rely on them.

What happens if my app forgets to send the user token to Azure AI Search?

For indexes with permission filters, a query without the x-ms-query-source-authorization token returns only documents that are public to everyone, and ACL-protected content isn't returned. Since November 2025 this applies even when you authenticate with an admin API key.

How quickly do SharePoint permission changes reach an Azure AI Search index?

Changes to items with unique permissions are picked up on the next successful indexer run, starting with the 2026-05-01-preview REST API. Changes made at a parent scope such as a site, library or folder aren't detected automatically; you must call the resync operation with the permissions option or reset the affected documents.

SharePoint OnlineMicrosoft GraphAzure AI SearchAzure OpenAIMicrosoft Entra ID
  1. Azure AI Search Integrated Vectorization: Indexing PDFs from Blob Storage

    Index business PDFs from Azure Blob Storage for RAG without custom code: an indexer, a Text Split and Azure OpenAI embedding skillset, index projections, and a vectorizer for text-to-vector queries.

    AI engineering12 min read
  2. Azure OpenAI Keyless Access: Managed Identity, Entra ID and Private Endpoints

    Remove API keys from Azure OpenAI: call it with a managed identity and Entra ID RBAC, disable local auth, and reach it only through a private endpoint with public network access turned off.

    AI engineering12 min read
  3. Copilot Studio SharePoint Knowledge: Set It Up and Fix No-Answer Errors

    Add SharePoint sites and lists as Copilot Studio knowledge, choose the right authentication, and fix agents that answer "I'm not sure how to help with that."

    AI engineering12 min read