An Azure OpenAI 429 Too Many Requests means the deployment you called refused the request because a limit was reached: the tokens-per-minute (TPM) or requests-per-minute (RPM) limit of a Standard deployment, 100% utilization of a provisioned (PTU) deployment, or temporary capacity pressure in the shared Standard pool. The fix depends on which of these you hit, so read the error message and the x-ratelimit-* response headers first, then raise or rebalance quota, lower max_tokens, retry using retry-after-ms, smooth bursts, and for provisioned deployments add spillover to a standard deployment. Requesting more quota only helps when the cause really is the quota.
Who this is for and what you will have
This guide is for developers and platform engineers running chat, RAG or agent workloads on Azure OpenAI in Microsoft Foundry who see intermittent or sustained 429 responses. At the end you will have:
- A way to tell a quota 429 from a capacity 429 and from a self-inflicted
max_tokens429. - Commands to read your per-region quota and change the TPM assigned to a deployment.
- Retry settings that honour the service's wait hint instead of hammering the endpoint.
- Spillover configured for a provisioned deployment, and metrics to prove it works.
If you are building the wider platform around these deployments, the gateway, caching and evaluation layers are covered in production LLMOps for enterprise RAG.
How Azure OpenAI decides to return 429
Standard deployments: TPM and RPM
Quota is assigned per subscription, per region, per model and per deployment type, in units of TPM. When you create a deployment you assign part of that quota to it, and the assigned TPM becomes the deployment's rate limit. An RPM limit is set in proportion to the TPM. You don't control TPM and RPM separately; quota is allocated in capacity units:
| Model | 1 capacity unit | RPM | TPM |
|---|---|---|---|
| Older chat models | 1 unit | 6 | 1,000 |
| o1 and o1-preview | 1 unit | 1 | 6,000 |
| o3 | 1 unit | 1 | 1,000 |
| o4-mini | 1 unit | 1 | 1,000 |
| o3-mini, o1-mini, o3-pro | 1 unit | 1 | 10,000 |
Two details cause most surprises:
- TPM is enforced on an estimate, not on billed tokens. As each request arrives, the service estimates the maximum tokens it could process from the prompt text, the
max_tokensvalue and thebest_ofvalue. That estimate is added to a running count that resets every minute. Once the count reaches the TPM limit, further requests get 429 until the counter resets. The estimate is based partly on character count, so throttling can start earlier than an exact token count would suggest. - RPM is enforced over short windows. The service evaluates the request rate over a small period, typically 1 or 10 seconds. Microsoft's example: a 600 RPM deployment monitored on 1-second intervals is throttled if it receives more than 10 requests in a second, even if the minute total stays under 600.
Provisioned deployments: utilization
Provisioned deployments use a variation of the leaky bucket algorithm. Each request adds an estimate (prompt tokens less cached tokens, plus max_tokens) to utilization; utilization drains continuously at a rate proportional to the deployed PTUs, and the estimate is corrected with actual token counts when the request finishes. When utilization is at 100%, the service returns 429 immediately with retry-after-ms and retry-after headers. Microsoft describes this 429 as a traffic-management signal rather than a service error: accepted requests keep predictable latency because excess traffic is rejected instead of queued. If you don't specify max_tokens, the service estimates a value, which can lower the concurrency you get.
Prerequisites
- Cognitive Services Usages Reader at subscription scope to read quota. Microsoft notes that a resource group or resource-level assignment doesn't authorize the subscription-scoped Usages API.
- Cognitive Services Contributor on the resource to change deployments or configure spillover.
- Azure CLI 2.51.0 or later for the quota commands (
az upgradeupdates an older install), signed in withaz loginto the subscription that holds the resource. Theaz restexamples below use that sign-in for the Azure Resource Manager calls. - The resource endpoint and an API key (or a Microsoft Entra ID token) for the inference calls. The examples use the
api-keyheader with the key inAZURE_OPENAI_API_KEYand the endpoint inAZURE_OPENAI_ENDPOINT. - Access to the resource in the Azure portal for Azure Monitor metrics.
Step 1: Identify which 429 you have
Capture one failing response with its headers. Every response includes rate limit headers:
| Header | Meaning |
|---|---|
x-ratelimit-limit-requests | Requests per minute permitted for the deployment |
x-ratelimit-limit-tokens | Tokens per minute permitted for the deployment |
x-ratelimit-remaining-requests | Requests left before the limit |
x-ratelimit-remaining-tokens | Tokens left before the limit |
x-ratelimit-reset-requests | Time until the request limit resets |
x-ratelimit-reset-tokens | Time until the token limit resets |
retry-after-ms | On 429 responses, the recommended wait in milliseconds |
A quick way to see them is a raw call with curl -i:
curl -i "https://contoso-aoai.openai.azure.com/openai/v1/chat/completions" \
-H "Content-Type: application/json" \
-H "api-key: $AZURE_OPENAI_API_KEY" \
-d '{"model": "gpt-4o-prod", "max_tokens": 200,
"messages": [{"role": "user", "content": "ping"}]}'Then match the message to the cause. Microsoft documents four scenarios:
| Message or signal | Root cause | Action |
|---|---|---|
| "Requests to ... have been limited" or "Rate limit is exceeded" | TPM or RPM limit of the deployment's quota | Raise or rebalance the deployment's TPM, or request more quota |
| "The service is temporarily unable to process your request" or "System is experiencing high demand" | Backend capacity constrained, often transient | Retry after retry-after-ms; consider provisioned throughput if persistent |
429s while quota is unchanged and x-ratelimit-limit-tokens is lower than the configured TPM | Temporary rate limit adjustment on the shared Standard pool | Retry with backoff; Microsoft says it typically resolves within a few hours |
| Throttled while token metrics look low | max_tokens and prompt estimate consuming the budget | Lower max_tokens to the expected response size |
Microsoft warns that capacity 429s are often misread as quota problems. Compare x-ratelimit-limit-tokens with the TPM you configured before filing a quota request.
Step 2: Check quota and resize the deployment
List the quota lines for a region. Each line shows currentValue (quota consumed by deployments) against limit:
az cognitiveservices usage list -l eastus -o tableThe same data is available from the Usages API, which is useful in scripts and alerts:
SUBSCRIPTION_ID="00000000-0000-0000-0000-000000000000"
az rest --method get \
--url "https://management.azure.com/subscriptions/$SUBSCRIPTION_ID/providers/Microsoft.CognitiveServices/locations/eastus/usages?api-version=2024-10-01"Quota names follow {Provider}.{DeploymentType}.{Model}, for example OpenAI.Standard.gpt-4o, with the limit expressed in thousands of TPM.
A common finding is approved quota in the subscription that isn't assigned to the busy deployment. To change a deployment's TPM, send a create-or-update request with a new sku.capacity. Capacity is set in whole units; for the older chat models in the table above, a capacity of 1 equals 1,000 TPM, so a capacity of 150 is 150,000 TPM. Check the table before converting units for reasoning models, where one unit carries a different TPM value. Because this is a PUT, read the deployment first and copy its current sku.name (for example Standard or GlobalStandard), model name, model version and any other properties you set, such as raiPolicyName or versionUpgradeOption, into the body so the update changes only the capacity:
RG="rg-ai-prod"; ACCOUNT="contoso-aoai"; DEPLOYMENT="gpt-4o-prod"
BASE="https://management.azure.com/subscriptions/$SUBSCRIPTION_ID/resourceGroups/$RG/providers/Microsoft.CognitiveServices/accounts/$ACCOUNT/deployments/$DEPLOYMENT"
az rest --method get --url "$BASE?api-version=2024-10-01"
az rest --method put --url "$BASE?api-version=2024-10-01" \
--body '{"sku": {"name": "Standard", "capacity": 150},
"properties": {"model": {"format": "OpenAI", "name": "gpt-4o", "version": "2024-11-20"}}}'In the Foundry portal, select Manage in the upper-right navigation, then Quota in the left pane, and stay on the Token per minute tab. Select a deployment to open its details pane, then use the pencil icon in the Affiliated deployments using shared quota section to change its allocation. Allow up to 15 minutes for an edited allocation to propagate, then refresh the page. If the region is out of quota, reduce TPM on underused deployments of the same model or select Request quota to submit an increase request. Microsoft notes that requests are prioritized for customers who actively use their existing allocation.
Before creating a new deployment elsewhere, the Model Capacities API shows where capacity exists for a model and version:
az rest --method get \
--url "https://management.azure.com/subscriptions/$SUBSCRIPTION_ID/providers/Microsoft.CognitiveServices/modelCapacities?api-version=2024-10-01&modelFormat=OpenAI&modelName=gpt-4o&modelVersion=2024-08-06"One more quota trap: if you delete a resource through the REST API or another programmatic method while it still has deployments, its quota stays unavailable for 48 hours until the resource is purged. Delete deployments first, or purge the deleted resource.
Step 3: Stop wasting budget with max_tokens
Because the TPM estimate includes max_tokens, an oversized value consumes rate limit budget even when the real answer is short. Microsoft's guidance:
- Set
max_tokensto the smallest value that serves the scenario. If responses are around 200 tokens, don't set 4,000. - Keep
best_ofat 1 unless you need multiple completions; each increment multiplies the estimate. - Trim prompts. Shorter system prompts and fewer retrieved chunks reduce the estimate directly.
The same applies to provisioned deployments: setting max_tokens close to the true generation size gives the highest concurrency. For RAG pipelines this usually means capping retrieved context and measuring actual output lengths before choosing a ceiling.
Reasoning models such as o3, o4-mini and the GPT-5 series don't accept max_tokens. With the Chat Completions API they only work with max_completion_tokens, and with the Responses API the cap is max_output_tokens. Both limits cover reasoning tokens as well as visible output, so leave room for the reasoning the model does before it answers.
Step 4: Retry correctly
Retries are expected for Standard deployments, but they must back off. Unsuccessful requests still count toward the per-minute limit.
The OpenAI Python library (v1 and later) retries 429 and transient errors automatically, respects the retry-after headers and uses exponential backoff with jitter. The default is two retries; raise it on the client or per call:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.getenv("AZURE_OPENAI_API_KEY"),
base_url="https://contoso-aoai.openai.azure.com/openai/v1/",
max_retries=5, # default is 2
)
response = client.with_options(max_retries=8).chat.completions.create(
model="gpt-4o-prod", # deployment name
max_tokens=300,
messages=[{"role": "user", "content": "Summarize the incident report."}],
)If you use your own retry library such as tenacity or Polly, set max_retries=0 on the client. Otherwise each outer attempt can trigger extra SDK retries and multiply the traffic you send to an already throttled deployment. A sound custom policy waits for retry-after-ms when present, otherwise doubles a randomized delay, and stops after a fixed number of attempts (Microsoft suggests something in the range of 5 to 10).
For .NET, AzureOpenAIClientOptions has built-in retry settings: set options.Retry.MaxRetries and options.Retry.Mode = RetryMode.Exponential. Microsoft suggests Polly only when you need more advanced patterns such as circuit breakers or bulkheads.
Step 5: Smooth bursts and spread load
Because RPM is checked in 1 to 10 second windows, a batch job that fires 200 requests at once can be throttled with a per-minute total far below the limit. Practical patterns:
- Put work on a queue and drain it at a controlled rate instead of fanning out unbounded parallel calls. The trade-offs between queueing technologies are covered in Kafka vs RabbitMQ vs AWS SQS.
- Ramp up new workloads gradually.
- Read
x-ratelimit-remaining-tokensandx-ratelimit-remaining-requestsand slow down before you reach zero. - Spread traffic over several deployments or regions when one deployment can't supply the throughput you need.
- Move work that doesn't need an immediate answer to asynchronous processing.
Step 6: Handle 429 on provisioned deployments
On a provisioned deployment you have two options when you receive 429: retry after the retry-after-ms wait if you need that deployment and can accept extra latency, or redirect the request to another deployment, which adds the least latency. Spillover automates the redirect.
Spillover requires a standard deployment of the same model and version in the same resource. It triggers when PTUs are fully used (429), when a request exceeds the context length the provisioned deployment supports (400), and on server errors (500 or 503). To enable it for every request, set spilloverDeploymentName on the provisioned deployment. When you add the property to an existing deployment, keep its current sku.name, sku.capacity and model version in the body, as in Step 2:
az rest --method put \
--url "https://management.azure.com/subscriptions/$SUBSCRIPTION_ID/resourceGroups/$RG/providers/Microsoft.CognitiveServices/accounts/$ACCOUNT/deployments/gpt-4o-mini-ptu?api-version=2024-10-01" \
--body '{"sku": {"name": "GlobalProvisionedManaged", "capacity": 100},
"properties": {"spilloverDeploymentName": "gpt-4o-mini-standard",
"model": {"format": "OpenAI", "name": "gpt-4o-mini", "version": "2024-07-18"}}}'To control it per request instead, leave the property unset and send the header on the calls that may spill over:
curl "$AZURE_OPENAI_ENDPOINT/openai/deployments/gpt-4o-mini-ptu/chat/completions?api-version=2024-10-21" \
-H "Content-Type: application/json" \
-H "x-ms-spillover-deployment: gpt-4o-mini-standard" \
-H "api-key: $AZURE_OPENAI_API_KEY" \
-d '{"messages": [{"role": "user", "content": "Classify this ticket."}]}'If both are configured, the deployment property wins. Requests served by the provisioned deployment cost nothing beyond the hourly PTU charge; requests that spill over are billed at the standard deployment's token rates. The service still prioritizes the provisioned deployment, which can add some latency.
Verify the fix
Use Azure Monitor rather than application logs alone:
- In the Azure portal, open the resource and select Monitoring > Metrics.
- Add the Azure OpenAI Requests metric, select Apply splitting and split by
ModelDeploymentNameandStatusCode. The 429 series for the deployment you changed should drop. - For provisioned deployments, chart Provisioned-managed Utilization V2, split by
ModelDeploymentName. Calls are throttled with 429 whenever utilization reaches 100%, so long stretches at 100% mean the traffic needs more PTUs, smallermax_tokensvalues or spillover. - For spillover, split the requests metric by
IsSpilloveron the standard deployment. Spilled requests appear there with their final status; they aren't double-counted as 429s on the provisioned deployment.
On individual responses, the header x-ms-spillover-from-deployment shows that a request spilled over, x-ms-deployment-name shows which deployment served it, and x-ms-spillover-error carries the status code that triggered the spillover.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| 429 with quota available in the region | TPM not assigned to the deployment receiving traffic | Increase that deployment's capacity or rebalance from idle deployments |
| 429 at low request counts | Bursts inside the 1 to 10 second RPM window | Queue and pace requests |
| 429 while billed tokens are low | Large max_tokens, or oversized requests rejected with 400 that can still count toward the limit | Lower max_tokens; fix requests that fail with 400 |
x-ratelimit-limit-tokens lower than configured TPM | Temporary adjustment on the shared Standard pool | Back off; consider provisioned throughput for consistent capacity |
| Far more calls than expected during throttling | SDK retries stacked under a custom retry library | Set max_retries=0 on the client when using your own policy |
| Spillover never triggers | Header or property points to a deployment of a different model or version, or in another resource | Use a standard deployment of the same model and version in the same resource |
| Sustained 429s in production below approved quota | Possible service-side issue | Open a support request after confirming deployment-level allocation |
Checklist
- Read the 429 message and
x-ratelimit-*headers before changing anything. - Confirm TPM is assigned to the deployment that receives the traffic.
- Set
max_tokensclose to real output size and keepbest_ofat 1. - Use the SDK's retries, or your own policy with SDK retries disabled.
- Pace bursts with a queue and ramp new workloads gradually.
- On provisioned deployments, configure spillover to a matching standard deployment.
- Watch Azure OpenAI Requests by
StatusCodeand Provisioned-managed Utilization V2 after each change.
When a 429 storm coincides with a model version change, check the deployment against the Azure OpenAI model retirement runbook, because a replacement model can have a different RPM-to-TPM ratio.
References
- Manage Azure OpenAI in Microsoft Foundry Models quota
- Automate Azure OpenAI deployments with quota
- Operate provisioned throughput deployments in production
- Provisioned throughput for Foundry Models
- Manage traffic with spillover for provisioned deployments
- Monitoring data reference for Azure OpenAI
- Azure OpenAI in Microsoft Foundry Models v1 API
- Azure OpenAI reasoning models
- Deployments - Create Or Update REST API