For most workloads, start with Global Standard: it gets new models first, has the lowest price and the highest default quota, and it can process prompts in any Azure region. Choose Data Zone Standard when inference must stay inside the US, EU or APAC data zone, Standard when it must stay inside one Azure geography, and a Provisioned type (Global, Data Zone or Regional) when you need reserved throughput and low latency variance at sustained volume, paid per PTU per hour or through a reservation.
Who this is for and what you will have at the end
This comparison is for architects and platform owners who deploy models in Azure OpenAI in Microsoft Foundry and must justify the choice to security, compliance or finance. It applies to the serverless deployment option for models sold directly by Azure; models running on managed compute don't use these types.
At the end you will have:
- A clear picture of the three axes behind every deployment type: data processing location, billing model and latency behavior.
- A decision table that maps common requirements to a type.
- Azure CLI commands to create each kind of deployment, including a provisioned deployment with spillover.
- An Azure Policy definition that blocks deployment types your compliance team hasn't approved.
The three decisions behind a deployment type
Every deployment type is a combination of three choices.
Where inference is processed. Data stored at rest stays in the designated Azure geography for all types. The difference is where prompts and completions are processed:
- Global types can process inference in any Azure region where the model is deployed.
- Data Zone types process inference only inside the Microsoft-defined data zone: United States, European Union or Asia Pacific. The EU zone follows the EU Data Boundary, which can include EFTA countries such as Norway and Switzerland, and Microsoft can add regions to a zone without prior notice.
- Standard and Regional Provisioned process inference within the Azure geography you choose, possibly moving between regions inside that geography for operational reasons.
How you pay. Standard types are pay-per-token. Provisioned types bill per provisioned throughput unit (PTU) per hour, whether or not you send traffic, with Azure reservations available as a discount. Batch types are pay-per-token at a 50% discount compared with Global Standard, with a 24-hour target turnaround.
How latency behaves. Standard types are best-effort shared capacity; under high, consistent volume, latency variance increases, with the threshold set per model. Provisioned types hold dedicated capacity for your deployment and come with a defined latency target per model.
Deployment types at a glance
The SKU code is what you pass as --sku-name on the CLI and what appears in templates and Azure Policy.
| Deployment type | SKU code | Inference processed in | Billing | Typical use |
|---|---|---|---|---|
| Global Standard | GlobalStandard | Any Azure region | Pay-per-token | Default for most workloads, highest quota |
| Global Provisioned | GlobalProvisionedManaged | Any Azure region | Per PTU per hour, or reservation | Predictable high throughput without residency limits |
| Global Batch | GlobalBatch | Any Azure region | Pay-per-token, 50% discount | Large asynchronous jobs |
| Data Zone Standard | DataZoneStandard | Within the data zone | Pay-per-token | EU, US or APAC zone compliance |
| Data Zone Provisioned | DataZoneProvisionedManaged | Within the data zone | Per PTU per hour, or reservation | Zone compliance plus reserved throughput |
| Data Zone Batch | DataZoneBatch | Within the data zone | Pay-per-token, 50% discount | Batch jobs with zone compliance |
| Standard | Standard | Within the Azure geography | Pay-per-token | Geography compliance, low to medium volume |
| Regional Provisioned | ProvisionedManaged | Within the Azure geography | Per PTU per hour, or reservation | Strict single-geography residency plus throughput |
| Developer | DeveloperTier | Any Azure region | Pay-per-token | Evaluating fine-tuned models only; 24-hour lifetime, no SLA |
Not every model supports every type in every region. New models typically arrive in a set order: Global first, then Data Zone, then geography-based types, and geography-based types have no guaranteed availability date. Always check the model's region availability table before committing to a design.
Data residency: global, data zone or geography
Residency requirements usually decide the first branch of the tree.
If your policy only constrains data at rest, Global types already satisfy it, because stored data stays in the resource's geography. Many compliance reviews stop here once this distinction is explained.
If your policy constrains processing to a jurisdiction such as the EU, use a Data Zone type with the resource in a region inside that zone. Data Zone Standard also has higher default quotas than geography-based types, so it's usually a better fit than Standard for EU or US workloads.
If processing must stay in one Azure geography, use Standard or Regional Provisioned, and accept narrower model availability and lower throughput.
One resilience point applies to Global Standard and Data Zone Standard: if the primary region of the resource has an interruption, all traffic initially routed to that region is affected. Plan a second resource in another region and a gateway in front if you need regional failover. The API Management AI gateway guide in this series shows how to load balance across them.
Throughput and latency: standard versus provisioned
Standard types
Standard deployments share capacity across customers. You allocate part of your subscription's tokens-per-minute quota to each deployment, and requests above the deployment's limit are throttled. Global Standard offers the highest default quota and, according to Microsoft, removes the need to load balance across multiple resources for throughput alone. Global Standard also supports request-level processing tiers: Flex processing for delay-tolerant work and Priority processing for faster responses on a pay-as-you-go basis. Data Zone Standard supports Priority processing.
Provisioned types
A provisioned deployment holds a fixed amount of processing capacity, measured in PTUs, for your deployment only. Key properties:
- PTUs are model-independent. Quota is granted per subscription, per region and per deployment type (Global, Data Zone and Regional are separate pools), and can be used for any supported model.
- Throughput per PTU depends on the model. Heavier models need more PTUs for the same tokens per minute, and each model has a minimum deployment size.
- Quota isn't capacity. Quota is a free policy limit; capacity is what's actually available in the region at deployment time. A deployment can fail for lack of capacity even with quota. Deleting or scaling down releases capacity, with no guarantee you get it back.
- Sizing inputs. Requests per minute, prompt and completion sizes, the model's output-to-input token ratio and your prompt cache rate. Cached tokens don't consume PTU capacity. Use the capacity calculator in the Foundry portal and then benchmark with your own traffic.
When a provisioned deployment reaches 100% utilization, the service returns 429 immediately with retry-after and retry-after-ms headers rather than queuing, so accepted requests keep predictable latency. Utilization is estimated per request from prompt tokens (minus cached tokens) plus max_tokens, so setting max_tokens close to the real generation size increases concurrency.
Spillover lets a provisioned deployment send requests it can't serve (429, some 400 long-context errors, 500 or 503) to a standard deployment of the same model and version in the same resource. Spilled requests are billed at the standard deployment's token rates.
Cost model
Provisioned deployments are billed per PTU per hour from creation until deletion, regardless of tokens processed. Hourly billing suits short benchmarks or events; for sustained production, Azure reservations give a discounted effective rate for a 1-month or 1-year commitment.
Points that commonly cause surprises:
- Reservations are bought per deployment type (Global, Data Zone or Regional) and apply to the billing meter, not to a specific deployment. A single Global reservation can cover Global PTU deployments across regions.
- Reservations don't guarantee capacity. Deploy first, confirm capacity, then purchase.
- Scaling provisioned deployments up and down with traffic to stay on hourly billing is risky: capacity might not be available when you scale back up.
- Deleting a deployment doesn't cancel a reservation, and charges for deployments on a deleted resource continue until the resource is purged. Delete deployments before deleting the resource.
For current prices, use the Azure OpenAI pricing page; this article doesn't quote figures because they change by model and region.
Decision guide
| Requirement | Recommended type |
|---|---|
| Newest models, lowest price, broadest regions | Global Standard |
| Inference must stay in the EU, US or APAC zone | Data Zone Standard |
| Inference must stay in one Azure geography | Standard (or Regional Provisioned) |
| Consistent high volume, low latency variance, no residency limit | Global Provisioned |
| Zone residency plus reserved throughput | Data Zone Provisioned |
| Geography residency plus reserved throughput | Regional Provisioned |
| Large offline jobs, results within about 24 hours acceptable | Global Batch or Data Zone Batch |
| Short-lived evaluation of a fine-tuned model | Developer |
A common production pattern combines types: a provisioned deployment sized for the steady baseline, spillover to a Global Standard or Data Zone Standard deployment for bursts, and Batch for overnight enrichment jobs.
Create deployments with the Azure CLI
Deployments are created on the Azure OpenAI or Foundry resource with az cognitiveservices account deployment create. The --sku-name value is the SKU code from the table. For provisioned types, --sku-capacity is the number of PTUs; for standard types it allocates part of your quota to the deployment.
# Global Standard
az cognitiveservices account deployment create \
--resource-group rg-ai --name contoso-aoai \
--deployment-name gpt41-global \
--model-name gpt-4.1 --model-version 2025-04-14 --model-format OpenAI \
--sku-name GlobalStandard --sku-capacity 10
# Data Zone Standard (resource must be in a region inside the zone)
az cognitiveservices account deployment create \
--resource-group rg-ai --name contoso-aoai-eu \
--deployment-name gpt41-dz \
--model-name gpt-4.1 --model-version 2025-04-14 --model-format OpenAI \
--sku-name DataZoneStandard --sku-capacity 10For a provisioned deployment with spillover, first create the standard deployment of the same model and version, then reference it:
az cognitiveservices account deployment create \
--resource-group rg-ai --name contoso-aoai-eu \
--deployment-name gpt41-dz-ptu \
--model-name gpt-4.1 --model-version 2025-04-14 --model-format OpenAI \
--sku-name DataZoneProvisionedManaged --sku-capacity <ptu-count> \
--spillover-deployment-name gpt41-dzCheck the model's minimum PTU count and region availability before running the provisioned command; if capacity is short, the deployment fails and the Foundry portal suggests alternative regions. Creating or modifying deployments requires Cognitive Services OpenAI Contributor, Cognitive Services Contributor or Contributor on the resource; viewing quota requires Cognitive Services Usages Reader at subscription scope.
Spillover can also be requested per call with the x-ms-spillover-deployment header. If the deployment property and the header are both set, the deployment property wins.
Enforce approved types with Azure Policy
If compliance approves only Data Zone processing for a subscription, deny every other SKU at deployment time. The alias Microsoft.CognitiveServices/accounts/deployments/sku.name exposes the deployment type:
{
"mode": "All",
"policyRule": {
"if": {
"allOf": [
{
"field": "type",
"equals": "Microsoft.CognitiveServices/accounts/deployments"
},
{
"field": "Microsoft.CognitiveServices/accounts/deployments/sku.name",
"notIn": [
"DataZoneStandard",
"DataZoneProvisionedManaged",
"DataZoneBatch"
]
}
]
},
"then": {
"effect": "deny"
}
}
}Create it as a custom policy definition, assign it at the subscription or management group, and test with a Global Standard deployment, which should now be rejected. Start with the audit effect if existing deployments need to be inventoried first.
Verify what you deployed
- Confirm the type of each deployment:
az cognitiveservices account deployment list \
--resource-group rg-ai --name contoso-aoai-eu \
--query "[].{name:name, sku:sku.name, capacity:sku.capacity}" --output table-
For provisioned deployments, open the resource in the Azure portal, select Metrics, and chart Provisioned-managed utilization V2, splitting by deployment if there are several.
-
For spillover, check response headers:
x-ms-spillover-from-deploymentappears on spilled requests,x-ms-deployment-nameshows the deployment that served the call, andx-ms-spillover-errorcontains the status code that triggered spillover. In metrics, split Azure OpenAI Requests byModelDeploymentName,StatusCodeandIsSpillover.
Troubleshooting
| Symptom | Cause | Resolution |
|---|---|---|
| Deployment type not offered for a model | The model doesn't support that type in the region | Check model availability by deployment type and region |
| Provisioned deployment fails with quota available | No PTU capacity in the region | Try fewer PTUs, another region, retry later, or use Global Provisioned |
429 from a provisioned deployment | Utilization reached 100% | Honor retry-after-ms, enable spillover, or add PTUs |
429 from a standard deployment | Deployment's share of quota exhausted | Raise capacity within quota, request more quota, or spread load |
| Lower concurrency than sized on PTU | max_tokens unset or much larger than real output | Set max_tokens close to expected generation length |
| Charges continue after deleting the resource | Deleted resource not purged, deployments still billed | Delete deployments first, then purge the resource |
| Policy blocks a deployment unexpectedly | SKU not in the allowed list | Confirm the SKU code with deployment list and update the policy |
Summary
- Default to Global Standard; move off it only for residency, reserved throughput or batch economics.
- Data at rest stays in your geography for every type; the choice is about where inference runs.
- Data Zone types are the practical answer to EU or US processing requirements, with higher quotas than geography-based types.
- Provisioned throughput buys dedicated capacity and predictable latency; deploy first, then reserve.
- Use spillover for bursts, Batch for offline work, and Azure Policy to keep deployments within approved types.
- Pair the deployment choice with keyless access; see Azure OpenAI keyless access with managed identity and private endpoints. For the operating model around it, see production LLMOps for enterprise RAG.
References
- https://learn.microsoft.com/en-us/azure/ai-foundry/foundry-models/concepts/deployment-types
- https://learn.microsoft.com/en-us/azure/ai-foundry/openai/concepts/provisioned-throughput
- https://learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/provisioned-get-started
- https://learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/spillover-traffic-management
- https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput-billing
- https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/role-based-access-control
- https://learn.microsoft.com/en-us/cli/azure/cognitiveservices/account/deployment
- https://learn.microsoft.com/en-us/azure/governance/policy/concepts/definition-structure-policy-rule
- https://azure.microsoft.com/pricing/details/cognitive-services/openai-service/