To govern Azure OpenAI with Azure API Management, import the model endpoint as an API, authenticate to every backend with the gateway's managed identity, and add the llm-token-limit policy keyed on the caller's subscription so each application gets its own tokens-per-minute rate and token quota. Add llm-emit-token-metric to send prompt and completion token counts to Application Insights with dimensions such as subscription and API, and route requests to a backend pool of Azure OpenAI resources in different regions, with priorities and circuit breakers that honor Retry-After when a backend returns 429.
Who this is for and what you will have at the end
This guide is for platform teams that run a shared Azure OpenAI capacity for several internal applications and need to stop one noisy app from consuming everybody's tokens, charge usage back, and survive a regional throttling event. It assumes you already have Azure OpenAI deployments and an API Management instance in the Developer, Basic, Standard or Premium tier, or one of their v2 equivalents.
At the end you will have:
- An Azure OpenAI API in API Management that authenticates to backends with a managed identity, so apps never see a model key.
- Per-application token rate limits and monthly token quotas.
- Token consumption metrics in Application Insights, split by application.
- A backend pool across two regions with priority routing and circuit breakers.
- A test procedure and a troubleshooting table.
Architecture
App A (subscription key A) --\
App B (subscription key B) ---+--> API Management gateway
App C (subscription key C) --/ inbound:
llm-token-limit (per subscription: TPM + monthly quota)
llm-emit-token-metric (Subscription ID, Product ID, API ID)
authentication-managed-identity (cognitiveservices.azure.com)
set-backend-service -> aoai-pool
backend pool "aoai-pool":
priority 1: aoai-swc (circuit breaker on 429, honors Retry-After)
priority 2: aoai-frc (circuit breaker on 429, honors Retry-After)
--> Azure OpenAI resources (same deployment names in both)| Capability | API Management feature | Tier notes |
|---|---|---|
| Per-app token rate and quota | llm-token-limit policy | Developer, Basic, Basic v2, Standard, Standard v2, Premium, Premium v2 |
| Token metrics | llm-emit-token-metric policy | All tiers |
| Load balancing | Backend pool (round-robin, weighted, priority, session-aware) | Up to 30 backends per pool |
| Failure isolation | Backend circuit breaker | Not supported in Consumption; one rule per backend |
| Keyless backend auth | authentication-managed-identity or backend credentials | All tiers |
The policies are named llm-token-limit and llm-emit-token-metric because they work with any API that follows the OpenAI Chat Completions or Responses schema, the Anthropic Messages API (v2 tiers) or Google Vertex AI. Older samples use the Azure OpenAI-specific names azure-openai-token-limit and azure-openai-emit-token-metric; Microsoft's documentation for those now points to the llm- policies.
Prerequisites
- An API Management instance in a tier that supports
llm-token-limit: Developer, Basic, Basic v2, Standard, Standard v2, Premium or Premium v2. Backend circuit breakers aren't supported in Consumption. - Two Azure OpenAI or Foundry resources in different regions, each with a deployment of the same model using the same deployment name. With the Azure OpenAI client compatibility option, the deployment name is in the request path, so every backend in the pool must accept the same path.
- An Application Insights resource.
- Rights to assign roles on the Azure OpenAI resources and to edit API Management policies.
- Ideally, keyless access already in place on the model resources; see Azure OpenAI keyless access with managed identity and private endpoints.
Step 1: Import the Azure OpenAI API
The fastest route is the import wizard, which creates operations, a backend, a set-backend-service policy, and managed identity authentication for you.
- In the Azure portal, open your API Management instance and select APIs > APIs > + Add API.
- Under Create from Azure resource, select Microsoft Foundry.
- On Select AI Service, choose the subscription and the resource in your primary region.
- On Configure API, enter a display name and a base path such as
aoai, optionally select products, and for Client compatibility select Azure OpenAI (clients call/openai/deployments/<deployment>/chat/completions) or Azure OpenAI v1 (clients call the v1 endpoint and pass the deployment name in the body). - On Manage token consumption, you can let the wizard add the token limit and token metric policies. This guide configures them by hand in later steps, so you can skip or accept the defaults and then edit them.
- Select Review, then Create.
If you import from an OpenAPI specification instead, set an API URL suffix that ends in /openai and configure authentication yourself, as described next.
Step 2: Grant the gateway access to every backend
Enable the API Management system-assigned managed identity if the wizard didn't, and assign it Cognitive Services OpenAI User on each Azure OpenAI resource in the pool. The wizard only grants access to the resource you selected, so the second region needs a manual assignment.
APIM_PRINCIPAL=$(az apim show --resource-group rg-gateway --name apim-contoso \
--query identity.principalId --output tsv)
for AOAI in aoai-swc aoai-frc; do
SCOPE=$(az cognitiveservices account show --resource-group rg-ai --name "$AOAI" --query id --output tsv)
az role assignment create \
--assignee-object-id "$APIM_PRINCIPAL" \
--assignee-principal-type ServicePrincipal \
--role "Cognitive Services OpenAI User" \
--scope "$SCOPE"
doneIn policy, authentication is two statements: request a token for the https://cognitiveservices.azure.com resource, then set it as the bearer token. Because the same identity and audience work for every backend, one policy covers the whole pool.
<authentication-managed-identity resource="https://cognitiveservices.azure.com" output-token-variable-name="managed-id-access-token" ignore-error="false" />
<set-header name="Authorization" exists-action="override">
<value>@("Bearer " + (string)context.Variables["managed-id-access-token"])</value>
</set-header>Step 3: Create backends with circuit breakers
Create one backend per region. A circuit breaker rule that trips on 429 and accepts the Retry-After header is important for Azure OpenAI: Microsoft notes that a throttled Azure OpenAI backend can return a Retry-After value that's very large, and with acceptRetryAfter the breaker waits exactly that long before sending traffic to that backend again.
resource backendSwc 'Microsoft.ApiManagement/service/backends@2023-09-01-preview' = {
name: 'apim-contoso/aoai-swc'
properties: {
url: 'https://aoai-swc.openai.azure.com/openai'
protocol: 'http'
circuitBreaker: {
rules: [
{
name: 'throttle-breaker'
failureCondition: {
count: 1
interval: 'PT1M'
statusCodeRanges: [
{
min: 429
max: 429
}
]
}
tripDuration: 'PT1M'
acceptRetryAfter: true
}
]
}
}
}Repeat for aoai-frc with its own URL. You can do the same in the portal under APIs > Backends > your backend > Circuit breaker settings > Add new, setting the failure status code range and selecting True (Accept) for the Retry-After check. When a breaker trips, API Management stops sending requests to that backend for the trip duration. Called directly, a tripped backend returns 503 Service Unavailable to the client; inside a pool, traffic moves to the backends that are still available.
Step 4: Create the load-balanced pool
A pool is a backend of type Pool that lists other backends with a priority and weight. API Management uses lower-priority backends only when every backend in a higher priority group is unavailable because its circuit breaker has tripped. Within a priority group, requests are spread evenly or by weight.
resource aoaiPool 'Microsoft.ApiManagement/service/backends@2023-09-01-preview' = {
name: 'apim-contoso/aoai-pool'
properties: {
description: 'Azure OpenAI regional pool'
type: 'Pool'
pool: {
services: [
{
id: '/subscriptions/<sub-id>/resourceGroups/rg-gateway/providers/Microsoft.ApiManagement/service/apim-contoso/backends/aoai-swc'
priority: 1
weight: 1
}
{
id: '/subscriptions/<sub-id>/resourceGroups/rg-gateway/providers/Microsoft.ApiManagement/service/apim-contoso/backends/aoai-frc'
priority: 2
weight: 1
}
]
}
}
}A typical design puts a provisioned (PTU) deployment at priority 1 so its prepaid capacity is used first, and pay-per-token deployments at priority 2. Check that the regions you combine meet your residency rules; the deployment types comparison covers which types keep inference inside a data zone. If you use stateful APIs where a conversation must stay on one backend, enable session awareness on the pool, which sets a session cookie the client must return.
Keep two caveats in mind. Load balancing and circuit breaker decisions are approximate, because gateway instances don't synchronize them. And the circuit breaker decides where later requests go; to re-send a request that already received a 429, add a retry policy, as in the full policy below.
Step 5: Give each application its own token budget
Issue each application its own API Management subscription (one per app, or one product per tier of service). Then key the token limit on the subscription:
tokens-per-minutecaps prompt plus completion tokens per minute. Exceeding it returns429.token-quotawithtoken-quota-period(Hourly,Daily,Weekly,MonthlyorYearly) caps the total over a fixed window that starts at the UTC boundary of the period. Exceeding it returns403.estimate-prompt-tokensis required. Withfalse, the policy uses actual usage from the model response, so a request can reach the backend after the limit is crossed and the next one is blocked. Withtrue, it estimates prompt tokens beforehand, which avoids unnecessary backend calls at some performance cost.
A single counter is kept per counter-key value across all scopes where the policy uses that key. To give products different budgets, put the policy at product scope and include the product in the key, for example @(context.Product.Id + "-" + context.Subscription.Id). In v2 tiers, which use a token bucket algorithm, keep tokens-per-minute identical wherever the same key is used.
Step 6: Emit token metrics to Application Insights
Token metrics need three settings outside the policy:
- Connect Application Insights to API Management (Monitoring > Application Insights > + Add), and enable Application Insights logging on the API (Settings tab > Diagnostics Logs). Microsoft recommends a connection string with managed identity credentials, which requires giving the API Management identity Monitoring Metrics Publisher on the Application Insights resource and creating the logger through the REST API, Bicep or ARM.
- In Application Insights, select Usage and estimated costs > Custom metrics (Preview) > With dimensions.
- Set
"metrics": trueon theapplicationinsightsdiagnostic entity through the Diagnostic - Create or Update REST API.
Each policy can have up to five custom dimensions, each dimension is limited to 100 unique values, and each metric namespace to 1,000 active time series; beyond those limits, new values are silently dropped. Default dimensions such as Subscription ID, Product ID and API ID can be used without a value.
The complete API policy
<policies>
<inbound>
<base />
<llm-token-limit
counter-key="@(context.Subscription.Id)"
tokens-per-minute="20000"
token-quota="5000000"
token-quota-period="Monthly"
estimate-prompt-tokens="false"
remaining-tokens-header-name="x-remaining-tokens"
remaining-quota-tokens-header-name="x-remaining-quota-tokens"
tokens-consumed-header-name="x-tokens-consumed" />
<llm-emit-token-metric namespace="aoai-gateway">
<dimension name="Subscription ID" />
<dimension name="Product ID" />
<dimension name="API ID" />
</llm-emit-token-metric>
<authentication-managed-identity resource="https://cognitiveservices.azure.com" output-token-variable-name="managed-id-access-token" ignore-error="false" />
<set-header name="Authorization" exists-action="override">
<value>@("Bearer " + (string)context.Variables["managed-id-access-token"])</value>
</set-header>
<set-backend-service backend-id="aoai-pool" />
</inbound>
<backend>
<retry condition="@(context.Response != null && context.Response.StatusCode == 429)" count="2" interval="1" first-fast-retry="true">
<forward-request buffer-request-body="true" />
</retry>
</backend>
<outbound>
<base />
</outbound>
<on-error>
<base />
</on-error>
</policies>The numbers are examples; size them from each application's expected load and the quota of the deployments behind the pool. The retry block re-sends a throttled request with the buffered body; test in your environment that, with the breaker open, retries land on the next backend in the pool. Remove the wizard-generated set-backend-service and duplicate token policies if they're still present, so each policy runs once.
Verify the gateway
- Call through the gateway with an application's subscription key:
curl -i "https://apim-contoso.azure-api.net/aoai/openai/deployments/gpt-4o/chat/completions?api-version=2024-10-21" \
-H "Ocp-Apim-Subscription-Key: <app-a-key>" \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "Hello"}], "max_tokens": 50}'Adjust the path to the base URL shown on the API's Settings tab. Expect 200, a usage section in the body, and the x-remaining-tokens, x-remaining-quota-tokens and x-tokens-consumed headers.
-
Trigger the rate limit by temporarily setting
tokens-per-minutevery low for a test subscription. Expect429withRetry-After. Set a tinytoken-quotato see403. -
Check metrics. In the Application Insights resource, open Metrics, choose the
aoai-gatewaycustom namespace, and split the token metrics bySubscription ID. Allow a few minutes for data to appear. -
Test failover. Point the priority-1 backend URL at a deployment you've throttled or temporarily removed, then confirm requests are served by the priority-2 backend and that the breaker resets after the trip duration.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
401 from the backend | Gateway identity lacks Cognitive Services OpenAI User on that resource, or the token audience is wrong | Assign the role on every pool member; use https://cognitiveservices.azure.com |
404 from one region only | Deployment name differs between backends | Use identical deployment names across the pool |
429 with headers from your gateway policy | App reached its tokens-per-minute | Expected; raise the limit or have the app back off |
403 after heavy use | App's token-quota for the period is exhausted | Raise the quota or wait for the next period |
503 Service Unavailable | Circuit breakers open on all backends | Check backend health and throttling; add capacity or a lower-priority backend |
| Traffic never fails over | Breaker rule doesn't match 429, or both backends share one priority and both are throttled | Add a 429 status range, acceptRetryAfter, and a distinct priority |
| No custom metrics | Custom metrics with dimensions not enabled, or "metrics": true missing on the diagnostic | Enable both and confirm logging is on for the API |
| Token counts missing for streamed calls | Model doesn't return usage when streaming | Send include_usage as true in the request |
| Limits looser than configured in multi-region | Counters are per gateway, not global | Divide budgets per region or route each app to one region |
Closing checklist
- Every app calls the gateway with its own subscription; no app holds an Azure OpenAI key.
- The gateway identity has Cognitive Services OpenAI User on every backend resource.
llm-token-limitenforces a per-app TPM and a periodic quota; limits are sized against deployment quota.llm-emit-token-metricsends tokens with application dimensions; custom metrics and"metrics": trueare enabled.- Backends have circuit breakers on
429withacceptRetryAfter, inside a priority-based pool. - Failover has been tested, and per-region counter behavior is understood. For the broader production picture, see production LLMOps for enterprise RAG.
References
- https://learn.microsoft.com/en-us/azure/api-management/genai-gateway-capabilities
- https://learn.microsoft.com/en-us/azure/api-management/llm-token-limit-policy
- https://learn.microsoft.com/en-us/azure/api-management/llm-emit-token-metric-policy
- https://learn.microsoft.com/en-us/azure/api-management/backends
- https://learn.microsoft.com/en-us/azure/api-management/azure-ai-foundry-api
- https://learn.microsoft.com/en-us/azure/api-management/azure-openai-api-from-specification
- https://learn.microsoft.com/en-us/azure/api-management/api-management-authenticate-authorize-ai-apis
- https://learn.microsoft.com/en-us/azure/api-management/api-management-howto-app-insights
- https://learn.microsoft.com/en-us/azure/api-management/retry-policy
- https://learn.microsoft.com/en-us/cli/azure/apim
- https://learn.microsoft.com/en-us/azure/role-based-access-control/role-assignments-cli