AI engineering

Fix Azure OpenAI 429 Too Many Requests: TPM, RPM and PTU limits

Diagnose Azure OpenAI 429 errors on Standard and provisioned deployments and fix them with quota changes, max_tokens tuning, correct retries and spillover.

14 min read
On this page

An Azure OpenAI 429 Too Many Requests means the deployment you called refused the request because a limit was reached: the tokens-per-minute (TPM) or requests-per-minute (RPM) limit of a Standard deployment, 100% utilization of a provisioned (PTU) deployment, or temporary capacity pressure in the shared Standard pool. The fix depends on which of these you hit, so read the error message and the x-ratelimit-* response headers first, then raise or rebalance quota, lower max_tokens, retry using retry-after-ms, smooth bursts, and for provisioned deployments add spillover to a standard deployment. Requesting more quota only helps when the cause really is the quota.

Who this is for and what you will have

This guide is for developers and platform engineers running chat, RAG or agent workloads on Azure OpenAI in Microsoft Foundry who see intermittent or sustained 429 responses. At the end you will have:

  • A way to tell a quota 429 from a capacity 429 and from a self-inflicted max_tokens 429.
  • Commands to read your per-region quota and change the TPM assigned to a deployment.
  • Retry settings that honour the service's wait hint instead of hammering the endpoint.
  • Spillover configured for a provisioned deployment, and metrics to prove it works.

If you are building the wider platform around these deployments, the gateway, caching and evaluation layers are covered in production LLMOps for enterprise RAG.

How Azure OpenAI decides to return 429

Standard deployments: TPM and RPM

Quota is assigned per subscription, per region, per model and per deployment type, in units of TPM. When you create a deployment you assign part of that quota to it, and the assigned TPM becomes the deployment's rate limit. An RPM limit is set in proportion to the TPM. You don't control TPM and RPM separately; quota is allocated in capacity units:

Model1 capacity unitRPMTPM
Older chat models1 unit61,000
o1 and o1-preview1 unit16,000
o31 unit11,000
o4-mini1 unit11,000
o3-mini, o1-mini, o3-pro1 unit110,000

Two details cause most surprises:

  • TPM is enforced on an estimate, not on billed tokens. As each request arrives, the service estimates the maximum tokens it could process from the prompt text, the max_tokens value and the best_of value. That estimate is added to a running count that resets every minute. Once the count reaches the TPM limit, further requests get 429 until the counter resets. The estimate is based partly on character count, so throttling can start earlier than an exact token count would suggest.
  • RPM is enforced over short windows. The service evaluates the request rate over a small period, typically 1 or 10 seconds. Microsoft's example: a 600 RPM deployment monitored on 1-second intervals is throttled if it receives more than 10 requests in a second, even if the minute total stays under 600.

Provisioned deployments: utilization

Provisioned deployments use a variation of the leaky bucket algorithm. Each request adds an estimate (prompt tokens less cached tokens, plus max_tokens) to utilization; utilization drains continuously at a rate proportional to the deployed PTUs, and the estimate is corrected with actual token counts when the request finishes. When utilization is at 100%, the service returns 429 immediately with retry-after-ms and retry-after headers. Microsoft describes this 429 as a traffic-management signal rather than a service error: accepted requests keep predictable latency because excess traffic is rejected instead of queued. If you don't specify max_tokens, the service estimates a value, which can lower the concurrency you get.

Prerequisites

  • Cognitive Services Usages Reader at subscription scope to read quota. Microsoft notes that a resource group or resource-level assignment doesn't authorize the subscription-scoped Usages API.
  • Cognitive Services Contributor on the resource to change deployments or configure spillover.
  • Azure CLI 2.51.0 or later for the quota commands (az upgrade updates an older install), signed in with az login to the subscription that holds the resource. The az rest examples below use that sign-in for the Azure Resource Manager calls.
  • The resource endpoint and an API key (or a Microsoft Entra ID token) for the inference calls. The examples use the api-key header with the key in AZURE_OPENAI_API_KEY and the endpoint in AZURE_OPENAI_ENDPOINT.
  • Access to the resource in the Azure portal for Azure Monitor metrics.

Step 1: Identify which 429 you have

Capture one failing response with its headers. Every response includes rate limit headers:

HeaderMeaning
x-ratelimit-limit-requestsRequests per minute permitted for the deployment
x-ratelimit-limit-tokensTokens per minute permitted for the deployment
x-ratelimit-remaining-requestsRequests left before the limit
x-ratelimit-remaining-tokensTokens left before the limit
x-ratelimit-reset-requestsTime until the request limit resets
x-ratelimit-reset-tokensTime until the token limit resets
retry-after-msOn 429 responses, the recommended wait in milliseconds

A quick way to see them is a raw call with curl -i:

curl -i "https://contoso-aoai.openai.azure.com/openai/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -H "api-key: $AZURE_OPENAI_API_KEY" \
  -d '{"model": "gpt-4o-prod", "max_tokens": 200,
       "messages": [{"role": "user", "content": "ping"}]}'

Then match the message to the cause. Microsoft documents four scenarios:

Message or signalRoot causeAction
"Requests to ... have been limited" or "Rate limit is exceeded"TPM or RPM limit of the deployment's quotaRaise or rebalance the deployment's TPM, or request more quota
"The service is temporarily unable to process your request" or "System is experiencing high demand"Backend capacity constrained, often transientRetry after retry-after-ms; consider provisioned throughput if persistent
429s while quota is unchanged and x-ratelimit-limit-tokens is lower than the configured TPMTemporary rate limit adjustment on the shared Standard poolRetry with backoff; Microsoft says it typically resolves within a few hours
Throttled while token metrics look lowmax_tokens and prompt estimate consuming the budgetLower max_tokens to the expected response size

Microsoft warns that capacity 429s are often misread as quota problems. Compare x-ratelimit-limit-tokens with the TPM you configured before filing a quota request.

Step 2: Check quota and resize the deployment

List the quota lines for a region. Each line shows currentValue (quota consumed by deployments) against limit:

az cognitiveservices usage list -l eastus -o table

The same data is available from the Usages API, which is useful in scripts and alerts:

SUBSCRIPTION_ID="00000000-0000-0000-0000-000000000000"
az rest --method get \
  --url "https://management.azure.com/subscriptions/$SUBSCRIPTION_ID/providers/Microsoft.CognitiveServices/locations/eastus/usages?api-version=2024-10-01"

Quota names follow {Provider}.{DeploymentType}.{Model}, for example OpenAI.Standard.gpt-4o, with the limit expressed in thousands of TPM.

A common finding is approved quota in the subscription that isn't assigned to the busy deployment. To change a deployment's TPM, send a create-or-update request with a new sku.capacity. Capacity is set in whole units; for the older chat models in the table above, a capacity of 1 equals 1,000 TPM, so a capacity of 150 is 150,000 TPM. Check the table before converting units for reasoning models, where one unit carries a different TPM value. Because this is a PUT, read the deployment first and copy its current sku.name (for example Standard or GlobalStandard), model name, model version and any other properties you set, such as raiPolicyName or versionUpgradeOption, into the body so the update changes only the capacity:

RG="rg-ai-prod"; ACCOUNT="contoso-aoai"; DEPLOYMENT="gpt-4o-prod"
BASE="https://management.azure.com/subscriptions/$SUBSCRIPTION_ID/resourceGroups/$RG/providers/Microsoft.CognitiveServices/accounts/$ACCOUNT/deployments/$DEPLOYMENT"
 
az rest --method get --url "$BASE?api-version=2024-10-01"
 
az rest --method put --url "$BASE?api-version=2024-10-01" \
  --body '{"sku": {"name": "Standard", "capacity": 150},
           "properties": {"model": {"format": "OpenAI", "name": "gpt-4o", "version": "2024-11-20"}}}'

In the Foundry portal, select Manage in the upper-right navigation, then Quota in the left pane, and stay on the Token per minute tab. Select a deployment to open its details pane, then use the pencil icon in the Affiliated deployments using shared quota section to change its allocation. Allow up to 15 minutes for an edited allocation to propagate, then refresh the page. If the region is out of quota, reduce TPM on underused deployments of the same model or select Request quota to submit an increase request. Microsoft notes that requests are prioritized for customers who actively use their existing allocation.

Before creating a new deployment elsewhere, the Model Capacities API shows where capacity exists for a model and version:

az rest --method get \
  --url "https://management.azure.com/subscriptions/$SUBSCRIPTION_ID/providers/Microsoft.CognitiveServices/modelCapacities?api-version=2024-10-01&modelFormat=OpenAI&modelName=gpt-4o&modelVersion=2024-08-06"

One more quota trap: if you delete a resource through the REST API or another programmatic method while it still has deployments, its quota stays unavailable for 48 hours until the resource is purged. Delete deployments first, or purge the deleted resource.

Step 3: Stop wasting budget with max_tokens

Because the TPM estimate includes max_tokens, an oversized value consumes rate limit budget even when the real answer is short. Microsoft's guidance:

  • Set max_tokens to the smallest value that serves the scenario. If responses are around 200 tokens, don't set 4,000.
  • Keep best_of at 1 unless you need multiple completions; each increment multiplies the estimate.
  • Trim prompts. Shorter system prompts and fewer retrieved chunks reduce the estimate directly.

The same applies to provisioned deployments: setting max_tokens close to the true generation size gives the highest concurrency. For RAG pipelines this usually means capping retrieved context and measuring actual output lengths before choosing a ceiling.

Reasoning models such as o3, o4-mini and the GPT-5 series don't accept max_tokens. With the Chat Completions API they only work with max_completion_tokens, and with the Responses API the cap is max_output_tokens. Both limits cover reasoning tokens as well as visible output, so leave room for the reasoning the model does before it answers.

Step 4: Retry correctly

Retries are expected for Standard deployments, but they must back off. Unsuccessful requests still count toward the per-minute limit.

The OpenAI Python library (v1 and later) retries 429 and transient errors automatically, respects the retry-after headers and uses exponential backoff with jitter. The default is two retries; raise it on the client or per call:

import os
from openai import OpenAI
 
client = OpenAI(
    api_key=os.getenv("AZURE_OPENAI_API_KEY"),
    base_url="https://contoso-aoai.openai.azure.com/openai/v1/",
    max_retries=5,  # default is 2
)
 
response = client.with_options(max_retries=8).chat.completions.create(
    model="gpt-4o-prod",  # deployment name
    max_tokens=300,
    messages=[{"role": "user", "content": "Summarize the incident report."}],
)

If you use your own retry library such as tenacity or Polly, set max_retries=0 on the client. Otherwise each outer attempt can trigger extra SDK retries and multiply the traffic you send to an already throttled deployment. A sound custom policy waits for retry-after-ms when present, otherwise doubles a randomized delay, and stops after a fixed number of attempts (Microsoft suggests something in the range of 5 to 10).

For .NET, AzureOpenAIClientOptions has built-in retry settings: set options.Retry.MaxRetries and options.Retry.Mode = RetryMode.Exponential. Microsoft suggests Polly only when you need more advanced patterns such as circuit breakers or bulkheads.

Step 5: Smooth bursts and spread load

Because RPM is checked in 1 to 10 second windows, a batch job that fires 200 requests at once can be throttled with a per-minute total far below the limit. Practical patterns:

  • Put work on a queue and drain it at a controlled rate instead of fanning out unbounded parallel calls. The trade-offs between queueing technologies are covered in Kafka vs RabbitMQ vs AWS SQS.
  • Ramp up new workloads gradually.
  • Read x-ratelimit-remaining-tokens and x-ratelimit-remaining-requests and slow down before you reach zero.
  • Spread traffic over several deployments or regions when one deployment can't supply the throughput you need.
  • Move work that doesn't need an immediate answer to asynchronous processing.

Step 6: Handle 429 on provisioned deployments

On a provisioned deployment you have two options when you receive 429: retry after the retry-after-ms wait if you need that deployment and can accept extra latency, or redirect the request to another deployment, which adds the least latency. Spillover automates the redirect.

Spillover requires a standard deployment of the same model and version in the same resource. It triggers when PTUs are fully used (429), when a request exceeds the context length the provisioned deployment supports (400), and on server errors (500 or 503). To enable it for every request, set spilloverDeploymentName on the provisioned deployment. When you add the property to an existing deployment, keep its current sku.name, sku.capacity and model version in the body, as in Step 2:

az rest --method put \
  --url "https://management.azure.com/subscriptions/$SUBSCRIPTION_ID/resourceGroups/$RG/providers/Microsoft.CognitiveServices/accounts/$ACCOUNT/deployments/gpt-4o-mini-ptu?api-version=2024-10-01" \
  --body '{"sku": {"name": "GlobalProvisionedManaged", "capacity": 100},
           "properties": {"spilloverDeploymentName": "gpt-4o-mini-standard",
                          "model": {"format": "OpenAI", "name": "gpt-4o-mini", "version": "2024-07-18"}}}'

To control it per request instead, leave the property unset and send the header on the calls that may spill over:

curl "$AZURE_OPENAI_ENDPOINT/openai/deployments/gpt-4o-mini-ptu/chat/completions?api-version=2024-10-21" \
  -H "Content-Type: application/json" \
  -H "x-ms-spillover-deployment: gpt-4o-mini-standard" \
  -H "api-key: $AZURE_OPENAI_API_KEY" \
  -d '{"messages": [{"role": "user", "content": "Classify this ticket."}]}'

If both are configured, the deployment property wins. Requests served by the provisioned deployment cost nothing beyond the hourly PTU charge; requests that spill over are billed at the standard deployment's token rates. The service still prioritizes the provisioned deployment, which can add some latency.

Verify the fix

Use Azure Monitor rather than application logs alone:

  1. In the Azure portal, open the resource and select Monitoring > Metrics.
  2. Add the Azure OpenAI Requests metric, select Apply splitting and split by ModelDeploymentName and StatusCode. The 429 series for the deployment you changed should drop.
  3. For provisioned deployments, chart Provisioned-managed Utilization V2, split by ModelDeploymentName. Calls are throttled with 429 whenever utilization reaches 100%, so long stretches at 100% mean the traffic needs more PTUs, smaller max_tokens values or spillover.
  4. For spillover, split the requests metric by IsSpillover on the standard deployment. Spilled requests appear there with their final status; they aren't double-counted as 429s on the provisioned deployment.

On individual responses, the header x-ms-spillover-from-deployment shows that a request spilled over, x-ms-deployment-name shows which deployment served it, and x-ms-spillover-error carries the status code that triggered the spillover.

Troubleshooting

SymptomLikely causeFix
429 with quota available in the regionTPM not assigned to the deployment receiving trafficIncrease that deployment's capacity or rebalance from idle deployments
429 at low request countsBursts inside the 1 to 10 second RPM windowQueue and pace requests
429 while billed tokens are lowLarge max_tokens, or oversized requests rejected with 400 that can still count toward the limitLower max_tokens; fix requests that fail with 400
x-ratelimit-limit-tokens lower than configured TPMTemporary adjustment on the shared Standard poolBack off; consider provisioned throughput for consistent capacity
Far more calls than expected during throttlingSDK retries stacked under a custom retry librarySet max_retries=0 on the client when using your own policy
Spillover never triggersHeader or property points to a deployment of a different model or version, or in another resourceUse a standard deployment of the same model and version in the same resource
Sustained 429s in production below approved quotaPossible service-side issueOpen a support request after confirming deployment-level allocation

Checklist

  • Read the 429 message and x-ratelimit-* headers before changing anything.
  • Confirm TPM is assigned to the deployment that receives the traffic.
  • Set max_tokens close to real output size and keep best_of at 1.
  • Use the SDK's retries, or your own policy with SDK retries disabled.
  • Pace bursts with a queue and ramp new workloads gradually.
  • On provisioned deployments, configure spillover to a matching standard deployment.
  • Watch Azure OpenAI Requests by StatusCode and Provisioned-managed Utilization V2 after each change.

When a 429 storm coincides with a model version change, check the deployment against the Azure OpenAI model retirement runbook, because a replacement model can have a different RPM-to-TPM ratio.

References

Questions people ask

Why do I get 429 errors when Azure Monitor shows token usage below my quota?

Rate limiting and usage metrics measure different things. The TPM limit is enforced on an estimate made when the request arrives, which includes the prompt and the max_tokens value, while token metrics show billed tokens from successful requests. Bursts inside a 1 or 10 second window, rejected requests and temporary rate limit adjustments on Standard deployments can also produce 429s while the metrics look low.

How long should I wait before retrying an Azure OpenAI 429?

Wait for the value in the retry-after-ms response header, then retry. If the header is missing, use exponential backoff with random jitter and a maximum number of retries. Failed requests still count toward the per-minute limit, so retrying immediately makes throttling worse.

How does Azure OpenAI convert TPM quota into an RPM limit?

Quota is allocated in capacity units and each unit carries a fixed amount of TPM and RPM that depends on the model. For older chat models one unit is 1,000 TPM and 6 RPM, while reasoning models such as o3 and o4-mini get 1 RPM per unit. You can't set TPM and RPM independently.

What does spillover do for a provisioned Azure OpenAI deployment?

Spillover sends requests that a provisioned deployment can't serve, such as 429s when PTUs are fully used or 500 and 503 errors, to a standard deployment in the same resource. Requests served by the standard deployment are billed at its token rates. You enable it with the spilloverDeploymentName deployment property or per request with the x-ms-spillover-deployment header.

Azure OpenAIMicrosoft FoundryProvisioned ThroughputAzure Monitor
  1. Azure OpenAI Deployment Types: Global Standard vs Data Zone vs Provisioned

    Choose between Global Standard, Data Zone, Standard and Provisioned deployments in Azure OpenAI based on where data is processed, latency variance, throughput guarantees and billing model.

    AI engineering12 min read
  2. Azure API Management AI Gateway: Token Limits and Load Balancing for OpenAI

    Put Azure API Management in front of Azure OpenAI to give each app its own token rate limit and quota, emit token metrics to Application Insights, and balance traffic across regions with circuit breakers.

    AI engineering12 min read
  3. Azure OpenAI model retirement runbook: avoid 410 errors, upgrade safely

    Track Azure OpenAI model retirement dates, set versionUpgradeOption, validate the replacement model and migrate provisioned deployments before the cutoff.

    AI engineering11 min read