AI engineering

Azure AI Search Integrated Vectorization: Indexing PDFs from Blob Storage

Index business PDFs from Azure Blob Storage for RAG without custom code: an indexer, a Text Split and Azure OpenAI embedding skillset, index projections, and a vectorizer for text-to-vector queries.

12 min read
On this page

To index business PDFs from Blob Storage for RAG without writing a chunking pipeline, create four Azure AI Search objects: a blob data source that connects with the search service's managed identity, a skillset that splits each document's text with the Text Split skill and embeds every chunk with the Azure OpenAI Embedding skill, an index with a vector field and an Azure OpenAI vectorizer, and an indexer that runs the pipeline on a schedule. Index projections in the skillset write one search document per chunk, with the parent file's name and path repeated on each, so your RAG app can run hybrid queries in plain text and cite the source PDF.

Who this is for and what you will have at the end

This guide is for engineers building retrieval-augmented generation over internal documents, such as policies, contracts, manuals and reports, that already live in an Azure Storage container. It uses the Azure AI Search REST API so every setting is visible and repeatable, and it notes where the portal wizard can do the same work.

At the end you will have:

  • A keyless pipeline: the search service reads blobs and calls the embedding model with its managed identity.
  • A chunk-level index with text, vector, title and source URL fields, ready for hybrid and vector queries.
  • A schedule that picks up new and changed PDFs automatically.
  • A verification procedure and a troubleshooting table for the failures most people hit first.

How the pipeline works

Blob container (PDFs)
   | data source: azureblob, ResourceId connection string (managed identity)
   v
Indexer --> document cracking --> /document/content (extracted text)
   |
   v
Skillset
   1. Text Split skill        /document/content     -> /document/pages/*
   2. Azure OpenAI Embedding  /document/pages/*     -> /document/pages/*/text_vector
   3. Index projections       one search doc per page, parent fields repeated
   v
Chunk index: chunk_id (key), parent_id, chunk, text_vector, title, source_url
   + vectorizer (same embedding deployment) for text-to-vector queries

Integrated vectorization is available in all regions and tiers, and the Text Split skill is free. You pay for Azure OpenAI embedding calls at standard rates, and you need the Basic tier or higher if the search service is to use a managed identity for outbound connections to Storage and Azure OpenAI.

Wizard or REST

Import data wizard (RAG)REST API
EffortA few minutes in the portalFour requests
ChunkingFixed: 2,000-character pages, 500-character overlapFully configurable
Index schemaGenerated fields you can extend but not modifyAny schema you design
NetworkRequires public access to storage and Azure OpenAI while it runsWorks from inside a virtual network
RepeatabilityManualScriptable and reviewable

The wizard (search service Overview > Import data, then your data source and RAG) is a good way to see a working configuration. Its generated objects are visible under Search management, and reviewing their JSON is a quick way to learn the structure used below.

Prerequisites

  • An Azure AI Search service, Basic tier or higher. Services created before January 1, 2019 might not support vector fields; if adding one fails, create a new service.
  • A general-purpose v2 storage account (standard performance) with a container of PDFs. Blobs in hot, cool and cold tiers are indexed.
  • An Azure OpenAI resource created in the Azure portal with a custom subdomain (for example https://contoso-aoai.openai.azure.com), and a deployment of text-embedding-3-small, text-embedding-3-large or text-embedding-ada-002. Microsoft notes that Azure OpenAI resources created in the Foundry portal aren't supported for this scenario. Use a deployment dedicated to indexing if you can, so indexing and query traffic don't compete for the same tokens-per-minute quota.
  • A REST client such as the REST Client extension for Visual Studio Code, and the Azure CLI.

Step 1: Set up identities and roles

  1. On the search service, enable role-based access for the data plane and turn on the system-assigned managed identity.
  2. On the storage account, assign Storage Blob Data Reader to the search service's managed identity.
  3. On the Azure OpenAI resource, assign Cognitive Services OpenAI User to the search service's managed identity. This is the only role the embedding skill needs; don't grant broader roles.
  4. Assign yourself Search Service Contributor and Search Index Data Contributor on the search service (add Search Index Data Reader for querying).

Then get a token for your REST calls:

az account get-access-token --scope https://search.azure.com/.default --query accessToken --output tsv

Set up the variables in your .http file:

@baseUrl = https://contoso-search.search.windows.net
@token = <token-from-previous-command>
@aoaiEndpoint = https://contoso-aoai.openai.azure.com
@aoaiDeployment = text-embedding-3-small
@aoaiModel = text-embedding-3-small

If you're also locking Azure OpenAI down with keyless access and private endpoints, the keyless access guide covers the network side; the search-specific options are in the security section below.

Step 2: Create the data source

With a managed identity, the connection string contains only the storage account's resource ID, with no key:

POST {{baseUrl}}/datasources?api-version=2026-04-01
Content-Type: application/json
Authorization: Bearer {{token}}
 
{
  "name": "policies-blob-ds",
  "type": "azureblob",
  "credentials": {
    "connectionString": "ResourceId=/subscriptions/<sub-id>/resourceGroups/rg-data/providers/Microsoft.Storage/storageAccounts/contosodocs/;"
  },
  "container": {
    "name": "policies",
    "query": null
  }
}

The query property can limit indexing to a virtual folder. If you need deleted blobs to disappear from the index, enable soft delete on the storage account and add a deletion detection policy to the data source; without one, chunks of deleted files stay in the index.

Step 3: Create the chunk index with a vectorizer

Index projections require a key field of type Edm.String with the keyword analyzer, and a separate filterable string field for the parent key. The vector field's dimensions must match the embedding output; text-embedding-3-small produces up to 1,536 dimensions.

POST {{baseUrl}}/indexes?api-version=2026-04-01
Content-Type: application/json
Authorization: Bearer {{token}}
 
{
  "name": "policies-chunks",
  "fields": [
    { "name": "chunk_id", "type": "Edm.String", "key": true, "filterable": true, "analyzer": "keyword" },
    { "name": "parent_id", "type": "Edm.String", "filterable": true },
    { "name": "title", "type": "Edm.String", "searchable": true, "filterable": true, "retrievable": true },
    { "name": "source_url", "type": "Edm.String", "filterable": true, "retrievable": true },
    { "name": "chunk", "type": "Edm.String", "searchable": true, "retrievable": true },
    {
      "name": "text_vector",
      "type": "Collection(Edm.Single)",
      "searchable": true,
      "retrievable": false,
      "stored": false,
      "dimensions": 1536,
      "vectorSearchProfile": "hnsw-aoai"
    }
  ],
  "vectorSearch": {
    "algorithms": [
      { "name": "hnsw", "kind": "hnsw", "hnswParameters": { "metric": "cosine" } }
    ],
    "profiles": [
      { "name": "hnsw-aoai", "algorithm": "hnsw", "vectorizer": "aoai-vectorizer" }
    ],
    "vectorizers": [
      {
        "name": "aoai-vectorizer",
        "kind": "azureOpenAI",
        "azureOpenAIParameters": {
          "resourceUri": "{{aoaiEndpoint}}",
          "deploymentId": "{{aoaiDeployment}}",
          "modelName": "{{aoaiModel}}"
        }
      }
    ]
  }
}

Vector fields must be searchable and can't be filterable, facetable or sortable. Setting stored to false saves space because the raw vectors are never returned to the app. The vectorizer must use the same embedding model as the skillset, or query vectors won't be comparable with document vectors.

Step 4: Create the skillset with index projections

The skillset does three things: split, embed and project.

POST {{baseUrl}}/skillsets?api-version=2026-04-01
Content-Type: application/json
Authorization: Bearer {{token}}
 
{
  "name": "policies-skillset",
  "skills": [
    {
      "@odata.type": "#Microsoft.Skills.Text.SplitSkill",
      "name": "split-pages",
      "context": "/document",
      "textSplitMode": "pages",
      "maximumPageLength": 2000,
      "pageOverlapLength": 500,
      "maximumPagesToTake": 0,
      "defaultLanguageCode": "en",
      "inputs": [ { "name": "text", "source": "/document/content" } ],
      "outputs": [ { "name": "textItems", "targetName": "pages" } ]
    },
    {
      "@odata.type": "#Microsoft.Skills.Text.AzureOpenAIEmbeddingSkill",
      "name": "embed-pages",
      "context": "/document/pages/*",
      "resourceUri": "{{aoaiEndpoint}}",
      "deploymentId": "{{aoaiDeployment}}",
      "modelName": "{{aoaiModel}}",
      "dimensions": 1536,
      "inputs": [ { "name": "text", "source": "/document/pages/*" } ],
      "outputs": [ { "name": "embedding", "targetName": "text_vector" } ]
    }
  ],
  "indexProjections": {
    "selectors": [
      {
        "targetIndexName": "policies-chunks",
        "parentKeyFieldName": "parent_id",
        "sourceContext": "/document/pages/*",
        "mappings": [
          { "name": "chunk", "source": "/document/pages/*" },
          { "name": "text_vector", "source": "/document/pages/*/text_vector" },
          { "name": "title", "source": "/document/metadata_storage_name" },
          { "name": "source_url", "source": "/document/metadata_storage_path" }
        ]
      }
    ],
    "parameters": { "projectionMode": "skipIndexingParentDocuments" }
  }
}

Points that matter:

  • Chunk size. maximumPageLength is in characters by default (minimum 300, maximum 50,000, default 5,000), and the skill tries to break on sentence boundaries. The values here match the portal wizard. A preview unit of azureOpenAITokens lets you size by tokens instead; Microsoft's general recommendation for embedding models is 512 tokens, and the tokenizer doesn't support o200k_base.
  • Embedding limit. The embedding skill accepts at most 8,000 tokens per input and fails above that, which is another reason to chunk.
  • No key on the skill. With apiKey and authIdentity both omitted, the skill uses the search service's system-assigned identity.
  • Mappings. Every child field except the key and the parent ID must be mapped explicitly. Don't create a mapping for the parent key field; that breaks change tracking.
  • Projection mode. skipIndexingParentDocuments keeps the index uniform: one document per chunk, with parent fields repeated. The default would add an extra, mostly empty document per PDF.

Step 5: Create and schedule the indexer

The indexer ties everything together. It runs as soon as it's created and then on the schedule.

POST {{baseUrl}}/indexers?api-version=2026-04-01
Content-Type: application/json
Authorization: Bearer {{token}}
 
{
  "name": "policies-indexer",
  "dataSourceName": "policies-blob-ds",
  "targetIndexName": "policies-chunks",
  "skillsetName": "policies-skillset",
  "schedule": { "interval": "PT2H" },
  "parameters": {
    "maxFailedItems": 10,
    "maxFailedItemsPerBatch": 10,
    "configuration": {
      "indexedFileNameExtensions": ".pdf",
      "dataToExtract": "contentAndMetadata",
      "parsingMode": "default",
      "failOnUnsupportedContentType": false
    }
  }
}

The blob indexer extracts each file's text into content and detects new and changed blobs by their last-modified timestamp, so scheduled runs only process what changed. Microsoft also recommends a schedule because Azure AI Search retries embedding calls that are throttled by Azure OpenAI, but if the quota stays exhausted those retries fail, and a later run picks up the documents that were missed. You don't need output field mappings for projected chunks.

Verify the index

  1. Indexer status. Check the last run, item counts and any errors or warnings:
GET {{baseUrl}}/indexers/policies-indexer/status?api-version=2026-04-01
Authorization: Bearer {{token}}
  1. Hybrid query in plain text. Because the vector field has a vectorizer, you send text and the service vectorizes it. Combining search with vectorQueries runs keyword and vector retrieval together:
POST {{baseUrl}}/indexes('policies-chunks')/docs/search.post.search?api-version=2026-04-01
Content-Type: application/json
Authorization: Bearer {{token}}
 
{
  "search": "How many days of annual leave can be carried over?",
  "count": true,
  "select": "title, source_url, chunk",
  "vectorQueries": [
    {
      "kind": "text",
      "text": "How many days of annual leave can be carried over?",
      "fields": "text_vector",
      "k": 5
    }
  ]
}

For integrated vectorization, kind must be text. Results should be chunks of the relevant PDF, each with its file name and blob URL for citation.

  1. One document's chunks. Filter on parent_id with a value from a result to see every chunk produced from that file, and confirm the overlap between consecutive chunks looks right.

You can run the same queries in the portal's Search explorer: open the index, switch to JSON view, and under Query options hide vector values to keep results readable.

Securing the pipeline

  • Storage behind a firewall. If storage is network-protected and in the same region as the search service, use the system-assigned identity and either allow the search service as a trusted service or add a resource instance rule for it.
  • Azure OpenAI behind a private endpoint. Create a shared private link from the search service with the openai_account group ID, or allow trusted Azure services on the Azure OpenAI resource. Either way the managed identity still needs its role.
  • Through a gateway. The embedding skill and vectorizer also accept Azure API Management endpoints (not custom domains), which lets you apply the same token governance as your apps; see Azure API Management AI gateway for OpenAI.
  • Who can edit skillsets. Anyone who can change a skillset controls the resourceUri that receives the search service's token. Restrict skillset write access to trusted administrators and point resourceUri only at endpoints you own.

Troubleshooting

SymptomLikely causeFix
Indexer fails to read blobs (403)Search identity lacks Storage Blob Data Reader, or storage firewall blocks itAssign the role; add the trusted service exception or a resource instance rule
Embedding skill errors on authenticationSearch identity lacks Cognitive Services OpenAI User, or resourceUri isn't the custom subdomainAssign the role; use https://<name>.openai.azure.com
Error that text is larger than 8,000 tokensChunks too large for the embedding modelLower maximumPageLength
Many throttling warnings, documents missingAzure OpenAI tokens-per-minute quota exhaustedKeep the schedule running, raise quota, or use a dedicated embedding deployment
Index creation fails on the vector fieldOlder service without vector support, or dimensions mismatchCreate a new service; match dimensions in skill and field
Vector query returns nothingVectorizer missing from the profile, or profile not assigned to the fieldAdd the vectorizer to the profile used by text_vector
Indexer stops on non-PDF filesUnsupported content types in the containerUse indexedFileNameExtensions or failOnUnsupportedContentType: false
Managed identity options unavailableFree tierMove to Basic or higher, or use keys
Wizard fails with private resourcesWizard needs public accessUse the REST steps from inside the network
Chunks of deleted PDFs still returnedNo deletion detection on the data sourceEnable soft delete and a deletion detection policy, or delete child documents manually

Scanned PDFs that contain only images produce little or no content. The wizard can add OCR through Extract text from images; in a REST pipeline you would add an OCR skill before splitting.

Closing checklist

  • Search service on Basic or higher, with RBAC enabled and a system-assigned identity.
  • Identity has Storage Blob Data Reader and Cognitive Services OpenAI User, nothing broader.
  • Data source uses a ResourceId connection string; deletion detection decided.
  • Index has a keyword-analyzed key, a filterable parent_id, a vector field whose dimensions match the model, and a vectorizer on its profile.
  • Skillset splits /document/content, embeds /document/pages/*, and projects with skipIndexingParentDocuments.
  • Indexer limited to .pdf, scheduled, and its status checked after the first run.
  • Hybrid queries return cited chunks. For evaluation, monitoring and the rest of the RAG lifecycle, see production LLMOps for enterprise RAG.

References

Questions people ask

Do I need to write code to chunk and embed documents in Azure AI Search?

No. With integrated vectorization, an indexer cracks the documents, the Text Split skill chunks the text, the Azure OpenAI Embedding skill generates vectors, and index projections write one search document per chunk. A vectorizer on the index converts text queries to vectors at query time, so the app sends plain text.

Which roles does the search service need for keyless integrated vectorization?

Give the search service's managed identity Storage Blob Data Reader on the storage account and Cognitive Services OpenAI User on the Azure OpenAI resource. Managed identities for outbound connections require the Basic tier or higher; the Free tier must use keys.

How large can a chunk be for the Azure OpenAI Embedding skill?

The skill accepts up to 8,000 tokens of input; larger inputs fail with an error. Use the Text Split skill to keep chunks well below that. The Import data wizard uses 2,000-character pages with 500 characters of overlap, which is a reasonable starting point for business documents.

Why does the Import data wizard fail when my storage or Azure OpenAI is private?

The wizard requires public access to the data source and embedding model while it runs. Either run the wizard first and then enable firewalls and private endpoints, or create the objects with the REST API from a machine inside the virtual network.

Azure AI SearchAzure Blob StorageAzure OpenAIEmbeddings
  1. Build RAG Over SharePoint Documents Without Breaking Permissions

    Ground an internal AI assistant on SharePoint files so each user only gets answers from documents they can open, using the Copilot Retrieval API or Azure AI Search with ACL ingestion.

    AI engineering12 min read
  2. Azure API Management AI Gateway: Token Limits and Load Balancing for OpenAI

    Put Azure API Management in front of Azure OpenAI to give each app its own token rate limit and quota, emit token metrics to Application Insights, and balance traffic across regions with circuit breakers.

    AI engineering12 min read
  3. Azure OpenAI Deployment Types: Global Standard vs Data Zone vs Provisioned

    Choose between Global Standard, Data Zone, Standard and Provisioned deployments in Azure OpenAI based on where data is processed, latency variance, throughput guarantees and billing model.

    AI engineering12 min read