To index business PDFs from Blob Storage for RAG without writing a chunking pipeline, create four Azure AI Search objects: a blob data source that connects with the search service's managed identity, a skillset that splits each document's text with the Text Split skill and embeds every chunk with the Azure OpenAI Embedding skill, an index with a vector field and an Azure OpenAI vectorizer, and an indexer that runs the pipeline on a schedule. Index projections in the skillset write one search document per chunk, with the parent file's name and path repeated on each, so your RAG app can run hybrid queries in plain text and cite the source PDF.
Who this is for and what you will have at the end
This guide is for engineers building retrieval-augmented generation over internal documents, such as policies, contracts, manuals and reports, that already live in an Azure Storage container. It uses the Azure AI Search REST API so every setting is visible and repeatable, and it notes where the portal wizard can do the same work.
At the end you will have:
- A keyless pipeline: the search service reads blobs and calls the embedding model with its managed identity.
- A chunk-level index with text, vector, title and source URL fields, ready for hybrid and vector queries.
- A schedule that picks up new and changed PDFs automatically.
- A verification procedure and a troubleshooting table for the failures most people hit first.
How the pipeline works
Blob container (PDFs)
| data source: azureblob, ResourceId connection string (managed identity)
v
Indexer --> document cracking --> /document/content (extracted text)
|
v
Skillset
1. Text Split skill /document/content -> /document/pages/*
2. Azure OpenAI Embedding /document/pages/* -> /document/pages/*/text_vector
3. Index projections one search doc per page, parent fields repeated
v
Chunk index: chunk_id (key), parent_id, chunk, text_vector, title, source_url
+ vectorizer (same embedding deployment) for text-to-vector queriesIntegrated vectorization is available in all regions and tiers, and the Text Split skill is free. You pay for Azure OpenAI embedding calls at standard rates, and you need the Basic tier or higher if the search service is to use a managed identity for outbound connections to Storage and Azure OpenAI.
Wizard or REST
| Import data wizard (RAG) | REST API | |
|---|---|---|
| Effort | A few minutes in the portal | Four requests |
| Chunking | Fixed: 2,000-character pages, 500-character overlap | Fully configurable |
| Index schema | Generated fields you can extend but not modify | Any schema you design |
| Network | Requires public access to storage and Azure OpenAI while it runs | Works from inside a virtual network |
| Repeatability | Manual | Scriptable and reviewable |
The wizard (search service Overview > Import data, then your data source and RAG) is a good way to see a working configuration. Its generated objects are visible under Search management, and reviewing their JSON is a quick way to learn the structure used below.
Prerequisites
- An Azure AI Search service, Basic tier or higher. Services created before January 1, 2019 might not support vector fields; if adding one fails, create a new service.
- A general-purpose v2 storage account (standard performance) with a container of PDFs. Blobs in hot, cool and cold tiers are indexed.
- An Azure OpenAI resource created in the Azure portal with a custom subdomain (for example
https://contoso-aoai.openai.azure.com), and a deployment oftext-embedding-3-small,text-embedding-3-largeortext-embedding-ada-002. Microsoft notes that Azure OpenAI resources created in the Foundry portal aren't supported for this scenario. Use a deployment dedicated to indexing if you can, so indexing and query traffic don't compete for the same tokens-per-minute quota. - A REST client such as the REST Client extension for Visual Studio Code, and the Azure CLI.
Step 1: Set up identities and roles
- On the search service, enable role-based access for the data plane and turn on the system-assigned managed identity.
- On the storage account, assign Storage Blob Data Reader to the search service's managed identity.
- On the Azure OpenAI resource, assign Cognitive Services OpenAI User to the search service's managed identity. This is the only role the embedding skill needs; don't grant broader roles.
- Assign yourself Search Service Contributor and Search Index Data Contributor on the search service (add Search Index Data Reader for querying).
Then get a token for your REST calls:
az account get-access-token --scope https://search.azure.com/.default --query accessToken --output tsvSet up the variables in your .http file:
@baseUrl = https://contoso-search.search.windows.net
@token = <token-from-previous-command>
@aoaiEndpoint = https://contoso-aoai.openai.azure.com
@aoaiDeployment = text-embedding-3-small
@aoaiModel = text-embedding-3-smallIf you're also locking Azure OpenAI down with keyless access and private endpoints, the keyless access guide covers the network side; the search-specific options are in the security section below.
Step 2: Create the data source
With a managed identity, the connection string contains only the storage account's resource ID, with no key:
POST {{baseUrl}}/datasources?api-version=2026-04-01
Content-Type: application/json
Authorization: Bearer {{token}}
{
"name": "policies-blob-ds",
"type": "azureblob",
"credentials": {
"connectionString": "ResourceId=/subscriptions/<sub-id>/resourceGroups/rg-data/providers/Microsoft.Storage/storageAccounts/contosodocs/;"
},
"container": {
"name": "policies",
"query": null
}
}The query property can limit indexing to a virtual folder. If you need deleted blobs to disappear from the index, enable soft delete on the storage account and add a deletion detection policy to the data source; without one, chunks of deleted files stay in the index.
Step 3: Create the chunk index with a vectorizer
Index projections require a key field of type Edm.String with the keyword analyzer, and a separate filterable string field for the parent key. The vector field's dimensions must match the embedding output; text-embedding-3-small produces up to 1,536 dimensions.
POST {{baseUrl}}/indexes?api-version=2026-04-01
Content-Type: application/json
Authorization: Bearer {{token}}
{
"name": "policies-chunks",
"fields": [
{ "name": "chunk_id", "type": "Edm.String", "key": true, "filterable": true, "analyzer": "keyword" },
{ "name": "parent_id", "type": "Edm.String", "filterable": true },
{ "name": "title", "type": "Edm.String", "searchable": true, "filterable": true, "retrievable": true },
{ "name": "source_url", "type": "Edm.String", "filterable": true, "retrievable": true },
{ "name": "chunk", "type": "Edm.String", "searchable": true, "retrievable": true },
{
"name": "text_vector",
"type": "Collection(Edm.Single)",
"searchable": true,
"retrievable": false,
"stored": false,
"dimensions": 1536,
"vectorSearchProfile": "hnsw-aoai"
}
],
"vectorSearch": {
"algorithms": [
{ "name": "hnsw", "kind": "hnsw", "hnswParameters": { "metric": "cosine" } }
],
"profiles": [
{ "name": "hnsw-aoai", "algorithm": "hnsw", "vectorizer": "aoai-vectorizer" }
],
"vectorizers": [
{
"name": "aoai-vectorizer",
"kind": "azureOpenAI",
"azureOpenAIParameters": {
"resourceUri": "{{aoaiEndpoint}}",
"deploymentId": "{{aoaiDeployment}}",
"modelName": "{{aoaiModel}}"
}
}
]
}
}Vector fields must be searchable and can't be filterable, facetable or sortable. Setting stored to false saves space because the raw vectors are never returned to the app. The vectorizer must use the same embedding model as the skillset, or query vectors won't be comparable with document vectors.
Step 4: Create the skillset with index projections
The skillset does three things: split, embed and project.
POST {{baseUrl}}/skillsets?api-version=2026-04-01
Content-Type: application/json
Authorization: Bearer {{token}}
{
"name": "policies-skillset",
"skills": [
{
"@odata.type": "#Microsoft.Skills.Text.SplitSkill",
"name": "split-pages",
"context": "/document",
"textSplitMode": "pages",
"maximumPageLength": 2000,
"pageOverlapLength": 500,
"maximumPagesToTake": 0,
"defaultLanguageCode": "en",
"inputs": [ { "name": "text", "source": "/document/content" } ],
"outputs": [ { "name": "textItems", "targetName": "pages" } ]
},
{
"@odata.type": "#Microsoft.Skills.Text.AzureOpenAIEmbeddingSkill",
"name": "embed-pages",
"context": "/document/pages/*",
"resourceUri": "{{aoaiEndpoint}}",
"deploymentId": "{{aoaiDeployment}}",
"modelName": "{{aoaiModel}}",
"dimensions": 1536,
"inputs": [ { "name": "text", "source": "/document/pages/*" } ],
"outputs": [ { "name": "embedding", "targetName": "text_vector" } ]
}
],
"indexProjections": {
"selectors": [
{
"targetIndexName": "policies-chunks",
"parentKeyFieldName": "parent_id",
"sourceContext": "/document/pages/*",
"mappings": [
{ "name": "chunk", "source": "/document/pages/*" },
{ "name": "text_vector", "source": "/document/pages/*/text_vector" },
{ "name": "title", "source": "/document/metadata_storage_name" },
{ "name": "source_url", "source": "/document/metadata_storage_path" }
]
}
],
"parameters": { "projectionMode": "skipIndexingParentDocuments" }
}
}Points that matter:
- Chunk size.
maximumPageLengthis in characters by default (minimum 300, maximum 50,000, default 5,000), and the skill tries to break on sentence boundaries. The values here match the portal wizard. A previewunitofazureOpenAITokenslets you size by tokens instead; Microsoft's general recommendation for embedding models is 512 tokens, and the tokenizer doesn't supporto200k_base. - Embedding limit. The embedding skill accepts at most 8,000 tokens per input and fails above that, which is another reason to chunk.
- No key on the skill. With
apiKeyandauthIdentityboth omitted, the skill uses the search service's system-assigned identity. - Mappings. Every child field except the key and the parent ID must be mapped explicitly. Don't create a mapping for the parent key field; that breaks change tracking.
- Projection mode.
skipIndexingParentDocumentskeeps the index uniform: one document per chunk, with parent fields repeated. The default would add an extra, mostly empty document per PDF.
Step 5: Create and schedule the indexer
The indexer ties everything together. It runs as soon as it's created and then on the schedule.
POST {{baseUrl}}/indexers?api-version=2026-04-01
Content-Type: application/json
Authorization: Bearer {{token}}
{
"name": "policies-indexer",
"dataSourceName": "policies-blob-ds",
"targetIndexName": "policies-chunks",
"skillsetName": "policies-skillset",
"schedule": { "interval": "PT2H" },
"parameters": {
"maxFailedItems": 10,
"maxFailedItemsPerBatch": 10,
"configuration": {
"indexedFileNameExtensions": ".pdf",
"dataToExtract": "contentAndMetadata",
"parsingMode": "default",
"failOnUnsupportedContentType": false
}
}
}The blob indexer extracts each file's text into content and detects new and changed blobs by their last-modified timestamp, so scheduled runs only process what changed. Microsoft also recommends a schedule because Azure AI Search retries embedding calls that are throttled by Azure OpenAI, but if the quota stays exhausted those retries fail, and a later run picks up the documents that were missed. You don't need output field mappings for projected chunks.
Verify the index
- Indexer status. Check the last run, item counts and any errors or warnings:
GET {{baseUrl}}/indexers/policies-indexer/status?api-version=2026-04-01
Authorization: Bearer {{token}}- Hybrid query in plain text. Because the vector field has a vectorizer, you send text and the service vectorizes it. Combining
searchwithvectorQueriesruns keyword and vector retrieval together:
POST {{baseUrl}}/indexes('policies-chunks')/docs/search.post.search?api-version=2026-04-01
Content-Type: application/json
Authorization: Bearer {{token}}
{
"search": "How many days of annual leave can be carried over?",
"count": true,
"select": "title, source_url, chunk",
"vectorQueries": [
{
"kind": "text",
"text": "How many days of annual leave can be carried over?",
"fields": "text_vector",
"k": 5
}
]
}For integrated vectorization, kind must be text. Results should be chunks of the relevant PDF, each with its file name and blob URL for citation.
- One document's chunks. Filter on
parent_idwith a value from a result to see every chunk produced from that file, and confirm the overlap between consecutive chunks looks right.
You can run the same queries in the portal's Search explorer: open the index, switch to JSON view, and under Query options hide vector values to keep results readable.
Securing the pipeline
- Storage behind a firewall. If storage is network-protected and in the same region as the search service, use the system-assigned identity and either allow the search service as a trusted service or add a resource instance rule for it.
- Azure OpenAI behind a private endpoint. Create a shared private link from the search service with the
openai_accountgroup ID, or allow trusted Azure services on the Azure OpenAI resource. Either way the managed identity still needs its role. - Through a gateway. The embedding skill and vectorizer also accept Azure API Management endpoints (not custom domains), which lets you apply the same token governance as your apps; see Azure API Management AI gateway for OpenAI.
- Who can edit skillsets. Anyone who can change a skillset controls the
resourceUrithat receives the search service's token. Restrict skillset write access to trusted administrators and pointresourceUrionly at endpoints you own.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
Indexer fails to read blobs (403) | Search identity lacks Storage Blob Data Reader, or storage firewall blocks it | Assign the role; add the trusted service exception or a resource instance rule |
| Embedding skill errors on authentication | Search identity lacks Cognitive Services OpenAI User, or resourceUri isn't the custom subdomain | Assign the role; use https://<name>.openai.azure.com |
| Error that text is larger than 8,000 tokens | Chunks too large for the embedding model | Lower maximumPageLength |
| Many throttling warnings, documents missing | Azure OpenAI tokens-per-minute quota exhausted | Keep the schedule running, raise quota, or use a dedicated embedding deployment |
| Index creation fails on the vector field | Older service without vector support, or dimensions mismatch | Create a new service; match dimensions in skill and field |
| Vector query returns nothing | Vectorizer missing from the profile, or profile not assigned to the field | Add the vectorizer to the profile used by text_vector |
| Indexer stops on non-PDF files | Unsupported content types in the container | Use indexedFileNameExtensions or failOnUnsupportedContentType: false |
| Managed identity options unavailable | Free tier | Move to Basic or higher, or use keys |
| Wizard fails with private resources | Wizard needs public access | Use the REST steps from inside the network |
| Chunks of deleted PDFs still returned | No deletion detection on the data source | Enable soft delete and a deletion detection policy, or delete child documents manually |
Scanned PDFs that contain only images produce little or no content. The wizard can add OCR through Extract text from images; in a REST pipeline you would add an OCR skill before splitting.
Closing checklist
- Search service on Basic or higher, with RBAC enabled and a system-assigned identity.
- Identity has Storage Blob Data Reader and Cognitive Services OpenAI User, nothing broader.
- Data source uses a
ResourceIdconnection string; deletion detection decided. - Index has a
keyword-analyzed key, a filterableparent_id, a vector field whose dimensions match the model, and a vectorizer on its profile. - Skillset splits
/document/content, embeds/document/pages/*, and projects withskipIndexingParentDocuments. - Indexer limited to
.pdf, scheduled, and its status checked after the first run. - Hybrid queries return cited chunks. For evaluation, monitoring and the rest of the RAG lifecycle, see production LLMOps for enterprise RAG.
References
- https://learn.microsoft.com/en-us/azure/search/vector-search-integrated-vectorization
- https://learn.microsoft.com/en-us/azure/search/search-how-to-integrated-vectorization
- https://learn.microsoft.com/en-us/azure/search/search-how-to-define-index-projections
- https://learn.microsoft.com/en-us/azure/search/cognitive-search-skill-textsplit
- https://learn.microsoft.com/en-us/azure/search/cognitive-search-skill-azure-openai-embedding
- https://learn.microsoft.com/en-us/azure/search/search-how-to-index-azure-blob-storage
- https://learn.microsoft.com/en-us/azure/search/search-howto-managed-identities-storage
- https://learn.microsoft.com/en-us/azure/search/search-get-started-portal-import-vectors
- https://learn.microsoft.com/en-us/azure/search/search-indexer-howto-access-private