AI engineering

Azure OpenAI Deployment Types: Global Standard vs Data Zone vs Provisioned

Choose between Global Standard, Data Zone, Standard and Provisioned deployments in Azure OpenAI based on where data is processed, latency variance, throughput guarantees and billing model.

12 min read
On this page

For most workloads, start with Global Standard: it gets new models first, has the lowest price and the highest default quota, and it can process prompts in any Azure region. Choose Data Zone Standard when inference must stay inside the US, EU or APAC data zone, Standard when it must stay inside one Azure geography, and a Provisioned type (Global, Data Zone or Regional) when you need reserved throughput and low latency variance at sustained volume, paid per PTU per hour or through a reservation.

Who this is for and what you will have at the end

This comparison is for architects and platform owners who deploy models in Azure OpenAI in Microsoft Foundry and must justify the choice to security, compliance or finance. It applies to the serverless deployment option for models sold directly by Azure; models running on managed compute don't use these types.

At the end you will have:

  • A clear picture of the three axes behind every deployment type: data processing location, billing model and latency behavior.
  • A decision table that maps common requirements to a type.
  • Azure CLI commands to create each kind of deployment, including a provisioned deployment with spillover.
  • An Azure Policy definition that blocks deployment types your compliance team hasn't approved.

The three decisions behind a deployment type

Every deployment type is a combination of three choices.

Where inference is processed. Data stored at rest stays in the designated Azure geography for all types. The difference is where prompts and completions are processed:

  • Global types can process inference in any Azure region where the model is deployed.
  • Data Zone types process inference only inside the Microsoft-defined data zone: United States, European Union or Asia Pacific. The EU zone follows the EU Data Boundary, which can include EFTA countries such as Norway and Switzerland, and Microsoft can add regions to a zone without prior notice.
  • Standard and Regional Provisioned process inference within the Azure geography you choose, possibly moving between regions inside that geography for operational reasons.

How you pay. Standard types are pay-per-token. Provisioned types bill per provisioned throughput unit (PTU) per hour, whether or not you send traffic, with Azure reservations available as a discount. Batch types are pay-per-token at a 50% discount compared with Global Standard, with a 24-hour target turnaround.

How latency behaves. Standard types are best-effort shared capacity; under high, consistent volume, latency variance increases, with the threshold set per model. Provisioned types hold dedicated capacity for your deployment and come with a defined latency target per model.

Deployment types at a glance

The SKU code is what you pass as --sku-name on the CLI and what appears in templates and Azure Policy.

Deployment typeSKU codeInference processed inBillingTypical use
Global StandardGlobalStandardAny Azure regionPay-per-tokenDefault for most workloads, highest quota
Global ProvisionedGlobalProvisionedManagedAny Azure regionPer PTU per hour, or reservationPredictable high throughput without residency limits
Global BatchGlobalBatchAny Azure regionPay-per-token, 50% discountLarge asynchronous jobs
Data Zone StandardDataZoneStandardWithin the data zonePay-per-tokenEU, US or APAC zone compliance
Data Zone ProvisionedDataZoneProvisionedManagedWithin the data zonePer PTU per hour, or reservationZone compliance plus reserved throughput
Data Zone BatchDataZoneBatchWithin the data zonePay-per-token, 50% discountBatch jobs with zone compliance
StandardStandardWithin the Azure geographyPay-per-tokenGeography compliance, low to medium volume
Regional ProvisionedProvisionedManagedWithin the Azure geographyPer PTU per hour, or reservationStrict single-geography residency plus throughput
DeveloperDeveloperTierAny Azure regionPay-per-tokenEvaluating fine-tuned models only; 24-hour lifetime, no SLA

Not every model supports every type in every region. New models typically arrive in a set order: Global first, then Data Zone, then geography-based types, and geography-based types have no guaranteed availability date. Always check the model's region availability table before committing to a design.

Data residency: global, data zone or geography

Residency requirements usually decide the first branch of the tree.

If your policy only constrains data at rest, Global types already satisfy it, because stored data stays in the resource's geography. Many compliance reviews stop here once this distinction is explained.

If your policy constrains processing to a jurisdiction such as the EU, use a Data Zone type with the resource in a region inside that zone. Data Zone Standard also has higher default quotas than geography-based types, so it's usually a better fit than Standard for EU or US workloads.

If processing must stay in one Azure geography, use Standard or Regional Provisioned, and accept narrower model availability and lower throughput.

One resilience point applies to Global Standard and Data Zone Standard: if the primary region of the resource has an interruption, all traffic initially routed to that region is affected. Plan a second resource in another region and a gateway in front if you need regional failover. The API Management AI gateway guide in this series shows how to load balance across them.

Throughput and latency: standard versus provisioned

Standard types

Standard deployments share capacity across customers. You allocate part of your subscription's tokens-per-minute quota to each deployment, and requests above the deployment's limit are throttled. Global Standard offers the highest default quota and, according to Microsoft, removes the need to load balance across multiple resources for throughput alone. Global Standard also supports request-level processing tiers: Flex processing for delay-tolerant work and Priority processing for faster responses on a pay-as-you-go basis. Data Zone Standard supports Priority processing.

Provisioned types

A provisioned deployment holds a fixed amount of processing capacity, measured in PTUs, for your deployment only. Key properties:

  • PTUs are model-independent. Quota is granted per subscription, per region and per deployment type (Global, Data Zone and Regional are separate pools), and can be used for any supported model.
  • Throughput per PTU depends on the model. Heavier models need more PTUs for the same tokens per minute, and each model has a minimum deployment size.
  • Quota isn't capacity. Quota is a free policy limit; capacity is what's actually available in the region at deployment time. A deployment can fail for lack of capacity even with quota. Deleting or scaling down releases capacity, with no guarantee you get it back.
  • Sizing inputs. Requests per minute, prompt and completion sizes, the model's output-to-input token ratio and your prompt cache rate. Cached tokens don't consume PTU capacity. Use the capacity calculator in the Foundry portal and then benchmark with your own traffic.

When a provisioned deployment reaches 100% utilization, the service returns 429 immediately with retry-after and retry-after-ms headers rather than queuing, so accepted requests keep predictable latency. Utilization is estimated per request from prompt tokens (minus cached tokens) plus max_tokens, so setting max_tokens close to the real generation size increases concurrency.

Spillover lets a provisioned deployment send requests it can't serve (429, some 400 long-context errors, 500 or 503) to a standard deployment of the same model and version in the same resource. Spilled requests are billed at the standard deployment's token rates.

Cost model

Provisioned deployments are billed per PTU per hour from creation until deletion, regardless of tokens processed. Hourly billing suits short benchmarks or events; for sustained production, Azure reservations give a discounted effective rate for a 1-month or 1-year commitment.

Points that commonly cause surprises:

  • Reservations are bought per deployment type (Global, Data Zone or Regional) and apply to the billing meter, not to a specific deployment. A single Global reservation can cover Global PTU deployments across regions.
  • Reservations don't guarantee capacity. Deploy first, confirm capacity, then purchase.
  • Scaling provisioned deployments up and down with traffic to stay on hourly billing is risky: capacity might not be available when you scale back up.
  • Deleting a deployment doesn't cancel a reservation, and charges for deployments on a deleted resource continue until the resource is purged. Delete deployments before deleting the resource.

For current prices, use the Azure OpenAI pricing page; this article doesn't quote figures because they change by model and region.

Decision guide

RequirementRecommended type
Newest models, lowest price, broadest regionsGlobal Standard
Inference must stay in the EU, US or APAC zoneData Zone Standard
Inference must stay in one Azure geographyStandard (or Regional Provisioned)
Consistent high volume, low latency variance, no residency limitGlobal Provisioned
Zone residency plus reserved throughputData Zone Provisioned
Geography residency plus reserved throughputRegional Provisioned
Large offline jobs, results within about 24 hours acceptableGlobal Batch or Data Zone Batch
Short-lived evaluation of a fine-tuned modelDeveloper

A common production pattern combines types: a provisioned deployment sized for the steady baseline, spillover to a Global Standard or Data Zone Standard deployment for bursts, and Batch for overnight enrichment jobs.

Create deployments with the Azure CLI

Deployments are created on the Azure OpenAI or Foundry resource with az cognitiveservices account deployment create. The --sku-name value is the SKU code from the table. For provisioned types, --sku-capacity is the number of PTUs; for standard types it allocates part of your quota to the deployment.

# Global Standard
az cognitiveservices account deployment create \
  --resource-group rg-ai --name contoso-aoai \
  --deployment-name gpt41-global \
  --model-name gpt-4.1 --model-version 2025-04-14 --model-format OpenAI \
  --sku-name GlobalStandard --sku-capacity 10
 
# Data Zone Standard (resource must be in a region inside the zone)
az cognitiveservices account deployment create \
  --resource-group rg-ai --name contoso-aoai-eu \
  --deployment-name gpt41-dz \
  --model-name gpt-4.1 --model-version 2025-04-14 --model-format OpenAI \
  --sku-name DataZoneStandard --sku-capacity 10

For a provisioned deployment with spillover, first create the standard deployment of the same model and version, then reference it:

az cognitiveservices account deployment create \
  --resource-group rg-ai --name contoso-aoai-eu \
  --deployment-name gpt41-dz-ptu \
  --model-name gpt-4.1 --model-version 2025-04-14 --model-format OpenAI \
  --sku-name DataZoneProvisionedManaged --sku-capacity <ptu-count> \
  --spillover-deployment-name gpt41-dz

Check the model's minimum PTU count and region availability before running the provisioned command; if capacity is short, the deployment fails and the Foundry portal suggests alternative regions. Creating or modifying deployments requires Cognitive Services OpenAI Contributor, Cognitive Services Contributor or Contributor on the resource; viewing quota requires Cognitive Services Usages Reader at subscription scope.

Spillover can also be requested per call with the x-ms-spillover-deployment header. If the deployment property and the header are both set, the deployment property wins.

Enforce approved types with Azure Policy

If compliance approves only Data Zone processing for a subscription, deny every other SKU at deployment time. The alias Microsoft.CognitiveServices/accounts/deployments/sku.name exposes the deployment type:

{
  "mode": "All",
  "policyRule": {
    "if": {
      "allOf": [
        {
          "field": "type",
          "equals": "Microsoft.CognitiveServices/accounts/deployments"
        },
        {
          "field": "Microsoft.CognitiveServices/accounts/deployments/sku.name",
          "notIn": [
            "DataZoneStandard",
            "DataZoneProvisionedManaged",
            "DataZoneBatch"
          ]
        }
      ]
    },
    "then": {
      "effect": "deny"
    }
  }
}

Create it as a custom policy definition, assign it at the subscription or management group, and test with a Global Standard deployment, which should now be rejected. Start with the audit effect if existing deployments need to be inventoried first.

Verify what you deployed

  1. Confirm the type of each deployment:
az cognitiveservices account deployment list \
  --resource-group rg-ai --name contoso-aoai-eu \
  --query "[].{name:name, sku:sku.name, capacity:sku.capacity}" --output table
  1. For provisioned deployments, open the resource in the Azure portal, select Metrics, and chart Provisioned-managed utilization V2, splitting by deployment if there are several.

  2. For spillover, check response headers: x-ms-spillover-from-deployment appears on spilled requests, x-ms-deployment-name shows the deployment that served the call, and x-ms-spillover-error contains the status code that triggered spillover. In metrics, split Azure OpenAI Requests by ModelDeploymentName, StatusCode and IsSpillover.

Troubleshooting

SymptomCauseResolution
Deployment type not offered for a modelThe model doesn't support that type in the regionCheck model availability by deployment type and region
Provisioned deployment fails with quota availableNo PTU capacity in the regionTry fewer PTUs, another region, retry later, or use Global Provisioned
429 from a provisioned deploymentUtilization reached 100%Honor retry-after-ms, enable spillover, or add PTUs
429 from a standard deploymentDeployment's share of quota exhaustedRaise capacity within quota, request more quota, or spread load
Lower concurrency than sized on PTUmax_tokens unset or much larger than real outputSet max_tokens close to expected generation length
Charges continue after deleting the resourceDeleted resource not purged, deployments still billedDelete deployments first, then purge the resource
Policy blocks a deployment unexpectedlySKU not in the allowed listConfirm the SKU code with deployment list and update the policy

Summary

  • Default to Global Standard; move off it only for residency, reserved throughput or batch economics.
  • Data at rest stays in your geography for every type; the choice is about where inference runs.
  • Data Zone types are the practical answer to EU or US processing requirements, with higher quotas than geography-based types.
  • Provisioned throughput buys dedicated capacity and predictable latency; deploy first, then reserve.
  • Use spillover for bursts, Batch for offline work, and Azure Policy to keep deployments within approved types.
  • Pair the deployment choice with keyless access; see Azure OpenAI keyless access with managed identity and private endpoints. For the operating model around it, see production LLMOps for enterprise RAG.

References

Questions people ask

Does Global Standard store my data outside my Azure geography?

Data stored at rest stays in the designated Azure geography for every deployment type. What changes is where prompts and responses are processed: Global types can process inference in any Azure region, Data Zone types only inside the Microsoft-defined US, EU or APAC zone, and Standard or Regional Provisioned within the customer-specified geography.

What is the difference between PTU quota and PTU capacity?

Quota is a free policy limit on how many PTUs you may deploy per subscription, region and deployment type. Capacity is the actual model processing capacity available in the region when you deploy. Having quota doesn't guarantee capacity, so a provisioned deployment can fail even when quota is available.

Do Azure reservations guarantee provisioned throughput capacity?

No. A reservation is a billing discount on the PTU meter for a 1-month or 1-year term and is loosely coupled to deployments. Microsoft recommends creating the provisioned deployment first to confirm capacity, then buying a reservation that matches its deployment type and scope, and for Data Zone or Regional deployments its region as well.

What happens when a provisioned deployment is fully used?

The service returns HTTP 429 immediately with retry-after and retry-after-ms headers instead of queuing the request. You can retry after the indicated time, or configure spillover so overflow requests are routed to a standard deployment of the same model in the same resource.

Azure OpenAIMicrosoft FoundryProvisioned ThroughputData Residency
  1. Fix Azure OpenAI 429 Too Many Requests: TPM, RPM and PTU limits

    Diagnose Azure OpenAI 429 errors on Standard and provisioned deployments and fix them with quota changes, max_tokens tuning, correct retries and spillover.

    AI engineering14 min read
  2. Azure OpenAI model retirement runbook: avoid 410 errors, upgrade safely

    Track Azure OpenAI model retirement dates, set versionUpgradeOption, validate the replacement model and migrate provisioned deployments before the cutoff.

    AI engineering11 min read
  3. Fix Azure OpenAI content filter 400 errors: ResponsibleAIPolicyViolation

    Find which Azure OpenAI content filter category blocked a request, then fix false positives with custom content filters, Prompt Shields settings and blocklists.

    AI engineering12 min read