AI engineering

Azure API Management AI Gateway: Token Limits and Load Balancing for OpenAI

Put Azure API Management in front of Azure OpenAI to give each app its own token rate limit and quota, emit token metrics to Application Insights, and balance traffic across regions with circuit breakers.

12 min read
On this page

To govern Azure OpenAI with Azure API Management, import the model endpoint as an API, authenticate to every backend with the gateway's managed identity, and add the llm-token-limit policy keyed on the caller's subscription so each application gets its own tokens-per-minute rate and token quota. Add llm-emit-token-metric to send prompt and completion token counts to Application Insights with dimensions such as subscription and API, and route requests to a backend pool of Azure OpenAI resources in different regions, with priorities and circuit breakers that honor Retry-After when a backend returns 429.

Who this is for and what you will have at the end

This guide is for platform teams that run a shared Azure OpenAI capacity for several internal applications and need to stop one noisy app from consuming everybody's tokens, charge usage back, and survive a regional throttling event. It assumes you already have Azure OpenAI deployments and an API Management instance in the Developer, Basic, Standard or Premium tier, or one of their v2 equivalents.

At the end you will have:

  • An Azure OpenAI API in API Management that authenticates to backends with a managed identity, so apps never see a model key.
  • Per-application token rate limits and monthly token quotas.
  • Token consumption metrics in Application Insights, split by application.
  • A backend pool across two regions with priority routing and circuit breakers.
  • A test procedure and a troubleshooting table.

Architecture

App A (subscription key A) --\
App B (subscription key B) ---+--> API Management gateway
App C (subscription key C) --/      inbound:
                                      llm-token-limit      (per subscription: TPM + monthly quota)
                                      llm-emit-token-metric (Subscription ID, Product ID, API ID)
                                      authentication-managed-identity (cognitiveservices.azure.com)
                                      set-backend-service -> aoai-pool
                                    backend pool "aoai-pool":
                                      priority 1: aoai-swc  (circuit breaker on 429, honors Retry-After)
                                      priority 2: aoai-frc  (circuit breaker on 429, honors Retry-After)
                                    --> Azure OpenAI resources (same deployment names in both)
CapabilityAPI Management featureTier notes
Per-app token rate and quotallm-token-limit policyDeveloper, Basic, Basic v2, Standard, Standard v2, Premium, Premium v2
Token metricsllm-emit-token-metric policyAll tiers
Load balancingBackend pool (round-robin, weighted, priority, session-aware)Up to 30 backends per pool
Failure isolationBackend circuit breakerNot supported in Consumption; one rule per backend
Keyless backend authauthentication-managed-identity or backend credentialsAll tiers

The policies are named llm-token-limit and llm-emit-token-metric because they work with any API that follows the OpenAI Chat Completions or Responses schema, the Anthropic Messages API (v2 tiers) or Google Vertex AI. Older samples use the Azure OpenAI-specific names azure-openai-token-limit and azure-openai-emit-token-metric; Microsoft's documentation for those now points to the llm- policies.

Prerequisites

  • An API Management instance in a tier that supports llm-token-limit: Developer, Basic, Basic v2, Standard, Standard v2, Premium or Premium v2. Backend circuit breakers aren't supported in Consumption.
  • Two Azure OpenAI or Foundry resources in different regions, each with a deployment of the same model using the same deployment name. With the Azure OpenAI client compatibility option, the deployment name is in the request path, so every backend in the pool must accept the same path.
  • An Application Insights resource.
  • Rights to assign roles on the Azure OpenAI resources and to edit API Management policies.
  • Ideally, keyless access already in place on the model resources; see Azure OpenAI keyless access with managed identity and private endpoints.

Step 1: Import the Azure OpenAI API

The fastest route is the import wizard, which creates operations, a backend, a set-backend-service policy, and managed identity authentication for you.

  1. In the Azure portal, open your API Management instance and select APIs > APIs > + Add API.
  2. Under Create from Azure resource, select Microsoft Foundry.
  3. On Select AI Service, choose the subscription and the resource in your primary region.
  4. On Configure API, enter a display name and a base path such as aoai, optionally select products, and for Client compatibility select Azure OpenAI (clients call /openai/deployments/<deployment>/chat/completions) or Azure OpenAI v1 (clients call the v1 endpoint and pass the deployment name in the body).
  5. On Manage token consumption, you can let the wizard add the token limit and token metric policies. This guide configures them by hand in later steps, so you can skip or accept the defaults and then edit them.
  6. Select Review, then Create.

If you import from an OpenAPI specification instead, set an API URL suffix that ends in /openai and configure authentication yourself, as described next.

Step 2: Grant the gateway access to every backend

Enable the API Management system-assigned managed identity if the wizard didn't, and assign it Cognitive Services OpenAI User on each Azure OpenAI resource in the pool. The wizard only grants access to the resource you selected, so the second region needs a manual assignment.

APIM_PRINCIPAL=$(az apim show --resource-group rg-gateway --name apim-contoso \
  --query identity.principalId --output tsv)
 
for AOAI in aoai-swc aoai-frc; do
  SCOPE=$(az cognitiveservices account show --resource-group rg-ai --name "$AOAI" --query id --output tsv)
  az role assignment create \
    --assignee-object-id "$APIM_PRINCIPAL" \
    --assignee-principal-type ServicePrincipal \
    --role "Cognitive Services OpenAI User" \
    --scope "$SCOPE"
done

In policy, authentication is two statements: request a token for the https://cognitiveservices.azure.com resource, then set it as the bearer token. Because the same identity and audience work for every backend, one policy covers the whole pool.

<authentication-managed-identity resource="https://cognitiveservices.azure.com" output-token-variable-name="managed-id-access-token" ignore-error="false" />
<set-header name="Authorization" exists-action="override">
    <value>@("Bearer " + (string)context.Variables["managed-id-access-token"])</value>
</set-header>

Step 3: Create backends with circuit breakers

Create one backend per region. A circuit breaker rule that trips on 429 and accepts the Retry-After header is important for Azure OpenAI: Microsoft notes that a throttled Azure OpenAI backend can return a Retry-After value that's very large, and with acceptRetryAfter the breaker waits exactly that long before sending traffic to that backend again.

resource backendSwc 'Microsoft.ApiManagement/service/backends@2023-09-01-preview' = {
  name: 'apim-contoso/aoai-swc'
  properties: {
    url: 'https://aoai-swc.openai.azure.com/openai'
    protocol: 'http'
    circuitBreaker: {
      rules: [
        {
          name: 'throttle-breaker'
          failureCondition: {
            count: 1
            interval: 'PT1M'
            statusCodeRanges: [
              {
                min: 429
                max: 429
              }
            ]
          }
          tripDuration: 'PT1M'
          acceptRetryAfter: true
        }
      ]
    }
  }
}

Repeat for aoai-frc with its own URL. You can do the same in the portal under APIs > Backends > your backend > Circuit breaker settings > Add new, setting the failure status code range and selecting True (Accept) for the Retry-After check. When a breaker trips, API Management stops sending requests to that backend for the trip duration. Called directly, a tripped backend returns 503 Service Unavailable to the client; inside a pool, traffic moves to the backends that are still available.

Step 4: Create the load-balanced pool

A pool is a backend of type Pool that lists other backends with a priority and weight. API Management uses lower-priority backends only when every backend in a higher priority group is unavailable because its circuit breaker has tripped. Within a priority group, requests are spread evenly or by weight.

resource aoaiPool 'Microsoft.ApiManagement/service/backends@2023-09-01-preview' = {
  name: 'apim-contoso/aoai-pool'
  properties: {
    description: 'Azure OpenAI regional pool'
    type: 'Pool'
    pool: {
      services: [
        {
          id: '/subscriptions/<sub-id>/resourceGroups/rg-gateway/providers/Microsoft.ApiManagement/service/apim-contoso/backends/aoai-swc'
          priority: 1
          weight: 1
        }
        {
          id: '/subscriptions/<sub-id>/resourceGroups/rg-gateway/providers/Microsoft.ApiManagement/service/apim-contoso/backends/aoai-frc'
          priority: 2
          weight: 1
        }
      ]
    }
  }
}

A typical design puts a provisioned (PTU) deployment at priority 1 so its prepaid capacity is used first, and pay-per-token deployments at priority 2. Check that the regions you combine meet your residency rules; the deployment types comparison covers which types keep inference inside a data zone. If you use stateful APIs where a conversation must stay on one backend, enable session awareness on the pool, which sets a session cookie the client must return.

Keep two caveats in mind. Load balancing and circuit breaker decisions are approximate, because gateway instances don't synchronize them. And the circuit breaker decides where later requests go; to re-send a request that already received a 429, add a retry policy, as in the full policy below.

Step 5: Give each application its own token budget

Issue each application its own API Management subscription (one per app, or one product per tier of service). Then key the token limit on the subscription:

  • tokens-per-minute caps prompt plus completion tokens per minute. Exceeding it returns 429.
  • token-quota with token-quota-period (Hourly, Daily, Weekly, Monthly or Yearly) caps the total over a fixed window that starts at the UTC boundary of the period. Exceeding it returns 403.
  • estimate-prompt-tokens is required. With false, the policy uses actual usage from the model response, so a request can reach the backend after the limit is crossed and the next one is blocked. With true, it estimates prompt tokens beforehand, which avoids unnecessary backend calls at some performance cost.

A single counter is kept per counter-key value across all scopes where the policy uses that key. To give products different budgets, put the policy at product scope and include the product in the key, for example @(context.Product.Id + "-" + context.Subscription.Id). In v2 tiers, which use a token bucket algorithm, keep tokens-per-minute identical wherever the same key is used.

Step 6: Emit token metrics to Application Insights

Token metrics need three settings outside the policy:

  1. Connect Application Insights to API Management (Monitoring > Application Insights > + Add), and enable Application Insights logging on the API (Settings tab > Diagnostics Logs). Microsoft recommends a connection string with managed identity credentials, which requires giving the API Management identity Monitoring Metrics Publisher on the Application Insights resource and creating the logger through the REST API, Bicep or ARM.
  2. In Application Insights, select Usage and estimated costs > Custom metrics (Preview) > With dimensions.
  3. Set "metrics": true on the applicationinsights diagnostic entity through the Diagnostic - Create or Update REST API.

Each policy can have up to five custom dimensions, each dimension is limited to 100 unique values, and each metric namespace to 1,000 active time series; beyond those limits, new values are silently dropped. Default dimensions such as Subscription ID, Product ID and API ID can be used without a value.

The complete API policy

<policies>
    <inbound>
        <base />
        <llm-token-limit
            counter-key="@(context.Subscription.Id)"
            tokens-per-minute="20000"
            token-quota="5000000"
            token-quota-period="Monthly"
            estimate-prompt-tokens="false"
            remaining-tokens-header-name="x-remaining-tokens"
            remaining-quota-tokens-header-name="x-remaining-quota-tokens"
            tokens-consumed-header-name="x-tokens-consumed" />
        <llm-emit-token-metric namespace="aoai-gateway">
            <dimension name="Subscription ID" />
            <dimension name="Product ID" />
            <dimension name="API ID" />
        </llm-emit-token-metric>
        <authentication-managed-identity resource="https://cognitiveservices.azure.com" output-token-variable-name="managed-id-access-token" ignore-error="false" />
        <set-header name="Authorization" exists-action="override">
            <value>@("Bearer " + (string)context.Variables["managed-id-access-token"])</value>
        </set-header>
        <set-backend-service backend-id="aoai-pool" />
    </inbound>
    <backend>
        <retry condition="@(context.Response != null && context.Response.StatusCode == 429)" count="2" interval="1" first-fast-retry="true">
            <forward-request buffer-request-body="true" />
        </retry>
    </backend>
    <outbound>
        <base />
    </outbound>
    <on-error>
        <base />
    </on-error>
</policies>

The numbers are examples; size them from each application's expected load and the quota of the deployments behind the pool. The retry block re-sends a throttled request with the buffered body; test in your environment that, with the breaker open, retries land on the next backend in the pool. Remove the wizard-generated set-backend-service and duplicate token policies if they're still present, so each policy runs once.

Verify the gateway

  1. Call through the gateway with an application's subscription key:
curl -i "https://apim-contoso.azure-api.net/aoai/openai/deployments/gpt-4o/chat/completions?api-version=2024-10-21" \
  -H "Ocp-Apim-Subscription-Key: <app-a-key>" \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "Hello"}], "max_tokens": 50}'

Adjust the path to the base URL shown on the API's Settings tab. Expect 200, a usage section in the body, and the x-remaining-tokens, x-remaining-quota-tokens and x-tokens-consumed headers.

  1. Trigger the rate limit by temporarily setting tokens-per-minute very low for a test subscription. Expect 429 with Retry-After. Set a tiny token-quota to see 403.

  2. Check metrics. In the Application Insights resource, open Metrics, choose the aoai-gateway custom namespace, and split the token metrics by Subscription ID. Allow a few minutes for data to appear.

  3. Test failover. Point the priority-1 backend URL at a deployment you've throttled or temporarily removed, then confirm requests are served by the priority-2 backend and that the breaker resets after the trip duration.

Troubleshooting

SymptomLikely causeFix
401 from the backendGateway identity lacks Cognitive Services OpenAI User on that resource, or the token audience is wrongAssign the role on every pool member; use https://cognitiveservices.azure.com
404 from one region onlyDeployment name differs between backendsUse identical deployment names across the pool
429 with headers from your gateway policyApp reached its tokens-per-minuteExpected; raise the limit or have the app back off
403 after heavy useApp's token-quota for the period is exhaustedRaise the quota or wait for the next period
503 Service UnavailableCircuit breakers open on all backendsCheck backend health and throttling; add capacity or a lower-priority backend
Traffic never fails overBreaker rule doesn't match 429, or both backends share one priority and both are throttledAdd a 429 status range, acceptRetryAfter, and a distinct priority
No custom metricsCustom metrics with dimensions not enabled, or "metrics": true missing on the diagnosticEnable both and confirm logging is on for the API
Token counts missing for streamed callsModel doesn't return usage when streamingSend include_usage as true in the request
Limits looser than configured in multi-regionCounters are per gateway, not globalDivide budgets per region or route each app to one region

Closing checklist

  • Every app calls the gateway with its own subscription; no app holds an Azure OpenAI key.
  • The gateway identity has Cognitive Services OpenAI User on every backend resource.
  • llm-token-limit enforces a per-app TPM and a periodic quota; limits are sized against deployment quota.
  • llm-emit-token-metric sends tokens with application dimensions; custom metrics and "metrics": true are enabled.
  • Backends have circuit breakers on 429 with acceptRetryAfter, inside a priority-based pool.
  • Failover has been tested, and per-region counter behavior is understood. For the broader production picture, see production LLMOps for enterprise RAG.

References

Questions people ask

What status code does API Management return when an app exceeds its token limit?

With the llm-token-limit policy, exceeding the tokens-per-minute rate returns 429 Too Many Requests, and exceeding the token quota for the period returns 403 Forbidden. A Retry-After header carries the recommended wait in seconds, and you can rename it with retry-after-header-name.

Are token limits shared across API Management regions?

No. The llm-token-limit policy tracks usage independently at each gateway, including regional gateways in a multi-region deployment and workspace gateways. It doesn't aggregate token counts across the whole instance, so size per-region limits accordingly.

Why are my token metrics missing or wrong for streamed responses?

Some OpenAI models don't include token counts in streamed responses by default; send include_usage set to true in the request. Also confirm that custom metrics with dimensions are enabled in Application Insights and that the API Management diagnostic has metrics set to true. If a stream terminates unexpectedly, captured counts are inaccurate.

Which role does API Management need on Azure OpenAI?

Assign the API Management managed identity the Cognitive Services OpenAI User role on each Azure OpenAI or Foundry resource it calls. The authentication-managed-identity policy then requests a token for https://cognitiveservices.azure.com and sends it in the Authorization header.

Azure API ManagementAzure OpenAIAzure MonitorApplication Insights
  1. Fix Azure OpenAI 429 Too Many Requests: TPM, RPM and PTU limits

    Diagnose Azure OpenAI 429 errors on Standard and provisioned deployments and fix them with quota changes, max_tokens tuning, correct retries and spillover.

    AI engineering14 min read
  2. Azure AI Search Integrated Vectorization: Indexing PDFs from Blob Storage

    Index business PDFs from Azure Blob Storage for RAG without custom code: an indexer, a Text Split and Azure OpenAI embedding skillset, index projections, and a vectorizer for text-to-vector queries.

    AI engineering12 min read
  3. Azure OpenAI Deployment Types: Global Standard vs Data Zone vs Provisioned

    Choose between Global Standard, Data Zone, Standard and Provisioned deployments in Azure OpenAI based on where data is processed, latency variance, throughput guarantees and billing model.

    AI engineering12 min read