To protect a backend API in Azure API Management, add a validate-jwt (or validate-azure-ad-token) policy in the inbound section so every request must carry a valid Microsoft Entra access token with the right audience, issuer and app role, and then add rate-limit-by-key keyed on the calling application's client ID so one client can't exhaust capacity for the others. Requests without a valid token get 401 Unauthorized at the gateway, and clients that exceed their rate get 429 Too Many Requests with a Retry-After header, so neither reaches your backend.
Who this is for and what you will have at the end
This guide is for platform and integration engineers who publish internal or partner APIs through Azure API Management and want authentication and throttling enforced at the gateway instead of in each backend service. It assumes the callers are applications (daemons, services, partner integrations) using the OAuth 2.0 client credentials flow, though the same policies work for user tokens.
At the end you will have:
- An Entra app registration for the API with an app role, and a client app granted that role.
- A policy that validates v2.0 access tokens, requires the app role, and returns a clean
401otherwise. - Per-client rate limiting and an optional monthly quota keyed on the client application ID.
- A test procedure with
curl, and a troubleshooting table for the failures you are most likely to see.
How the policies fit together
The order of policies in the inbound section matters. A sensible order is: throttle anonymous traffic by IP, validate the token, then throttle by the authenticated client.
Client --> [APIM gateway inbound]
1. rate-limit-by-key (per IP, cheap flood protection)
2. validate-jwt (signature, exp, aud, iss, roles) --> 401 on failure
3. rate-limit-by-key (per client app, from token claim) --> 429 when exceeded
4. quota-by-key (per client app, long period) --> 403 when exceeded
--> backend API| Policy | Purpose | Response when it blocks | Notes |
|---|---|---|---|
validate-jwt | Validate any JWT from an OpenID Connect provider | 401 by default, configurable | Available in all tiers |
validate-azure-ad-token | Validate a token issued by Microsoft Entra | 401 by default, configurable | Available in all tiers |
rate-limit-by-key | Short-term call rate per arbitrary key | 429 Too Many Requests | Not available in the Consumption tier; renewal-period maximum is 300 seconds |
quota-by-key | Call count or bandwidth over a longer period per key | 403 Forbidden with Retry-After | The policy reference lists the Developer, Basic, Standard and Premium tiers; minimum renewal-period is 300 seconds |
Rate limits protect against short, intense bursts. Quotas control volume over a longer period such as a month, which suits tiered or monetized access.
Prerequisites
- An API Management instance in a tier that supports
rate-limit-by-key(Developer, Basic, Basic v2, Standard, Standard v2, Premium or Premium v2), with the API already imported. - Permission to create app registrations in Microsoft Entra ID, at least Cloud Application Administrator to define app roles and grant admin consent.
- Your directory (tenant) ID.
curland Python 3 on a workstation for testing.
If you are designing a multi-tenant product where the tenant identity in the token also drives data access, read multi-tenant SaaS data isolation on Azure alongside this guide.
Step 1: Register the API in Microsoft Entra ID
- In the Azure portal, open App registrations and select New registration. Name it, for example,
orders-api, choose the supported account type, leave Redirect URI empty and select Register. - On the Overview page, record the Application (client) ID.
- Under Manage, select Expose an API and set the Application ID URI to the default value, which takes the form
api://<client-id>. - Under Manage, select App roles and then Create app role. Set Display name to
Orders Reader, Allowed member types to Applications, Value toOrders.Read, add a description, make sure the role is enabled and select Apply. Create a second roleOrders.ReadWritethe same way if you need a write tier.
Request v2.0 access tokens
The format of an access token is controlled by the API's registration, not by the endpoint the client calls. The requestedAccessTokenVersion property in the app manifest accepts 1, 2 or null, and null defaults to 1. Open Manifest on the API registration and set it to 2:
"api": {
"requestedAccessTokenVersion": 2
}With v2.0 tokens the aud claim is always the API's client ID, the iss claim ends in /v2.0, and the calling application is identified by azp. In v1.0 tokens the caller is in appid and the audience can be either the client ID or the api:// URI used in the request. Choosing one version and writing the policy for it avoids most validation failures.
Step 2: Register the client application and grant the role
- Create a second app registration, for example
orders-sync-daemon, and record its client ID. - Add a client secret to the registration for testing. For production, prefer a certificate or a federated credential, which Microsoft describes as a higher level of assurance than a shared secret.
- Select API permissions > Add a permission > My APIs, choose
orders-api, select Application permissions, tickOrders.Readand select Add permissions. - Select Grant admin consent for your tenant and confirm. The Status column should show the consent as granted.
App roles granted this way appear in the roles claim of app-only tokens. Be aware that Entra can issue app-only tokens without a roles claim to any application unless you require assignment on the API's enterprise application, so the gateway policy must check the role rather than assume it.
Step 3: Store identifiers as named values
Keep the tenant ID and API client ID out of the policy text. In API Management, create two named values, for example entra-tenant-id and orders-api-client-id. Policies reference them with double braces such as {{entra-tenant-id}}.
Step 4: Validate the token at API scope
In the portal, open your API Management instance, select APIs, select the API, open the Design tab, select All operations, and in Inbound processing select the code editor icon. Add the policy after <base />.
<validate-jwt header-name="Authorization" require-scheme="Bearer"
failed-validation-httpcode="401"
failed-validation-error-message="Unauthorized. Access token is missing or invalid."
output-token-variable-name="jwt">
<openid-config url="https://login.microsoftonline.com/{{entra-tenant-id}}/v2.0/.well-known/openid-configuration" />
<audiences>
<audience>{{orders-api-client-id}}</audience>
</audiences>
<issuers>
<issuer>https://login.microsoftonline.com/{{entra-tenant-id}}/v2.0</issuer>
</issuers>
<required-claims>
<claim name="roles" match="any">
<value>Orders.Read</value>
<value>Orders.ReadWrite</value>
</claim>
</required-claims>
</validate-jwt>What each part does:
openid-configpoints at the tenant's v2.0 metadata. API Management downloads the signing keys and caches them. Microsoft documents a refresh every hour, plus a fetch at most once every five minutes when a token references a key ID that isn't cached (intervals that can change), so key rollover is handled for you.require-expiration-timeandrequire-signed-tokensdefault totrue, so unsigned tokens and tokens withoutexpare rejected without extra configuration.clock-skewdefaults to 0 seconds.output-token-variable-namestores the parsed token incontext.Variables["jwt"]so later policies can read its claims.required-claimswithmatch="any"accepts the token if at least one listed role is present.
If your instance is injected into a virtual network and the metadata URL can't be resolved through public DNS, add validate-connectivity="false" to the openid-config element.
The Entra-specific alternative
validate-azure-ad-token does the same job with less configuration for Entra-issued tokens. It takes the tenant ID directly and can restrict which client applications may call:
<validate-azure-ad-token tenant-id="{{entra-tenant-id}}" output-token-variable-name="jwt">
<client-application-ids>
<application-id>{{orders-sync-daemon-client-id}}</application-id>
</client-application-ids>
<audiences>
<audience>{{orders-api-client-id}}</audience>
</audiences>
<required-claims>
<claim name="roles" match="any">
<value>Orders.Read</value>
<value>Orders.ReadWrite</value>
</claim>
</required-claims>
</validate-azure-ad-token>Use one or the other, not both.
Step 5: Enforce write access per operation
A role check at API scope admits both readers and writers. To stop readers from calling write operations, add a check that reads the stored token and returns 403 when the role is missing. Place it at API scope after the validation policy, or at the scope of the write operations.
<choose>
<when condition="@(context.Request.Method != "GET" && !((Jwt)context.Variables["jwt"]).Claims.GetValueOrDefault("roles", "").Contains("Orders.ReadWrite"))">
<return-response>
<set-status code="403" reason="Forbidden" />
</return-response>
</when>
</choose>Claims.GetValueOrDefault returns the claim values as a comma-separated string, which is why Contains works for a multi-valued roles claim. Keep authorization checks in the backend as well; the gateway check is the first layer, not the only one.
Step 6: Throttle per client application
Rate limit by client ID
Key the counter on the azp claim from the validated token so each client application gets its own budget. Prefix the key with a label so it can't collide with other counters; API Management uses a single counter per key value across every scope where the policy appears.
<rate-limit-by-key calls="60" renewal-period="60"
counter-key="@("orders-client-" + ((Jwt)context.Variables["jwt"]).Claims.GetValueOrDefault("azp", "unknown"))"
remaining-calls-header-name="X-RateLimit-Remaining"
total-calls-header-name="X-RateLimit-Limit" />The remaining-calls-header-name and total-calls-header-name attributes add response headers so well-behaved clients can pace themselves. When the limit is exceeded the gateway returns 429 and a Retry-After header (you can rename it with retry-after-header-name).
Flood protection before validation
Token validation costs more than a counter lookup. A coarse per-IP limit placed before validate-jwt cuts off floods of junk requests. Keep it generous, because many callers can share one public IP behind NAT.
<rate-limit-by-key calls="300" renewal-period="60"
counter-key="@("orders-ip-" + context.Request.IpAddress)" />Monthly quota per client
For tiered access, add a quota with the same client key. The quota's renewal-period is a fixed window in seconds with a minimum of 300; the Microsoft example uses 2629800 seconds for a month. The quota-by-key policy reference lists only the classic Developer, Basic, Standard and Premium tiers, so confirm support before you rely on it in a v2 or Consumption tier instance.
<quota-by-key calls="1000000" renewal-period="2629800"
counter-key="@("orders-quota-" + ((Jwt)context.Variables["jwt"]).Claims.GetValueOrDefault("azp", "unknown"))" />Things to know about the counters
- Rate limiting is never completely accurate because the throttling architecture is distributed; the actual number allowed varies with request volume, rate and backend latency.
- Classic tiers use a sliding window; v2 tiers use a token bucket whose initial size equals
calls, which allows an initial burst. In v2 tiers every rate limit policy that shares a counter key must use the samecallsandrenewal-period, or behavior is unpredictable. - Rate limit counters are per gateway. In a multi-region deployment each regional gateway enforces the limit separately, and self-hosted gateways don't synchronize with the managed gateway. Quotas use one counter for the whole instance.
- If
increment-conditionorincrement-countuse an expression, the counter is updated at the end of the outbound pipeline, so the429arrives one call later than usual.
The complete inbound section
<policies>
<inbound>
<base />
<rate-limit-by-key calls="300" renewal-period="60"
counter-key="@("orders-ip-" + context.Request.IpAddress)" />
<validate-jwt header-name="Authorization" require-scheme="Bearer"
failed-validation-httpcode="401"
failed-validation-error-message="Unauthorized. Access token is missing or invalid."
output-token-variable-name="jwt">
<openid-config url="https://login.microsoftonline.com/{{entra-tenant-id}}/v2.0/.well-known/openid-configuration" />
<audiences>
<audience>{{orders-api-client-id}}</audience>
</audiences>
<issuers>
<issuer>https://login.microsoftonline.com/{{entra-tenant-id}}/v2.0</issuer>
</issuers>
<required-claims>
<claim name="roles" match="any">
<value>Orders.Read</value>
<value>Orders.ReadWrite</value>
</claim>
</required-claims>
</validate-jwt>
<rate-limit-by-key calls="60" renewal-period="60"
counter-key="@("orders-client-" + ((Jwt)context.Variables["jwt"]).Claims.GetValueOrDefault("azp", "unknown"))"
remaining-calls-header-name="X-RateLimit-Remaining"
total-calls-header-name="X-RateLimit-Limit" />
</inbound>
<backend>
<base />
</backend>
<outbound>
<base />
</outbound>
<on-error>
<base />
</on-error>
</policies>Select Calculate effective policy in the editor to see how this combines with any global or product policies.
Verification
Request a token with the client credentials flow. The scope is the API's Application ID URI followed by /.default.
TENANT_ID="<tenant-id>"
CLIENT_ID="<orders-sync-daemon-client-id>"
CLIENT_SECRET="<secret>"
API_APP_ID_URI="api://<orders-api-client-id>"
TOKEN=$(curl -s -X POST "https://login.microsoftonline.com/$TENANT_ID/oauth2/v2.0/token" \
-H "Content-Type: application/x-www-form-urlencoded" \
--data-urlencode "client_id=$CLIENT_ID" \
--data-urlencode "scope=$API_APP_ID_URI/.default" \
--data-urlencode "client_secret=$CLIENT_SECRET" \
--data-urlencode "grant_type=client_credentials" | python3 -c "import sys,json; print(json.load(sys.stdin)['access_token'])")Decode the payload and check ver, aud, iss, azp and roles before testing the gateway:
python3 -c "import sys,base64,json; p=sys.argv[1].split('.')[1]; p+='='*(-len(p)%4); print(json.dumps(json.loads(base64.urlsafe_b64decode(p)), indent=2))" "$TOKEN"Then run three tests against your gateway URL (add your subscription key header if the API requires a subscription):
# 1. No token: expect 401
curl -s -o /dev/null -w "%{http_code}\n" https://api.contoso.com/orders/v1/orders
# 2. Valid token: expect 200 and the rate limit headers
curl -s -D - -o /dev/null https://api.contoso.com/orders/v1/orders \
-H "Authorization: Bearer $TOKEN"
# 3. Burst past the limit: expect 429 responses near the end
for i in $(seq 1 80); do
curl -s -o /dev/null -w "%{http_code} " https://api.contoso.com/orders/v1/orders \
-H "Authorization: Bearer $TOKEN"
done; echoBecause counting is distributed and v2 tiers allow an initial burst, the exact call number where 429 starts can vary. What matters is that it appears and that Retry-After is present.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
401 with a token that works elsewhere; decoded token shows "ver": "1.0" and iss of https://sts.windows.net/<tenant-id>/ | The API registration still issues v1.0 tokens | Set requestedAccessTokenVersion to 2 and request a new token, or write the policy for v1.0 issuer, audience and the v1 metadata URL |
401; aud is api://<client-id> | v1.0 token requested with the URI as resource | Same as above, or add the api:// value as a second audience |
401; token has no roles claim | Application permission added but admin consent not granted, or role not assigned | Grant admin consent on the client registration and request a new token |
401 with the message "JWT not present." | Header missing, or sent in a custom header while relying on require-scheme | Send Authorization: Bearer <token>; require-scheme is ignored for custom headers |
401 in a VNet-injected instance for every request | Gateway can't reach the OpenID metadata endpoint | Allow outbound access to login.microsoftonline.com and set validate-connectivity="false" if DNS resolution is the issue |
429 arrives one call late | increment-condition or increment-count uses an expression | Expected behavior; counters are evaluated at the end of outbound |
| Limits look doubled across regions | Each regional gateway keeps its own counter | Divide the intended global limit by region count, or use a quota for a global ceiling |
| Unpredictable throttling on a v2 tier | Same counter key used with different calls or renewal-period | Use identical values, or different key prefixes per scope |
429 with no matching policy | Gateway capacity throttling, which is separate from rate limit policies | Review gateway capacity; this throttling protects the instance, not a specific API |
Checklist
- The API registration exposes an Application ID URI, defines app roles, and requests v2.0 tokens.
- Client applications are granted app roles with admin consent, using certificates or federated credentials in production.
- Tenant and client IDs are stored as named values.
validate-jwt(orvalidate-azure-ad-token) checks audience, issuer and required roles, and stores the token in a variable.- Write operations check for the write role and return
403otherwise; the backend also authorizes. rate-limit-by-keyis keyed on the client ID claim with a unique prefix; a coarse per-IP limit sits before validation.- Quotas are added where clients have contractual volumes.
- Tests confirm
401,200and429behavior, and alerts watch for spikes in each.
For the wider network and identity model that a gateway like this fits into, see the zero trust enterprise remote access architecture.
References
- validate-jwt policy
- validate-azure-ad-token policy
- rate-limit-by-key policy
- quota-by-key policy
- Advanced request throttling with Azure API Management
- Protect an API in API Management using OAuth 2.0 and Microsoft Entra ID
- Set or edit API Management policies
- API Management policy expressions
- Access token claims reference
- OAuth 2.0 client credentials flow
- Add app roles and get them from a token
- Microsoft Graph app manifest reference