To run Azure OpenAI without API keys, give the calling app a managed identity, assign that identity the Cognitive Services OpenAI User role on the resource, switch the client to Microsoft Entra ID tokens, and then set disableLocalAuth to true so keys stop working. To take the resource off the internet as well, create a private endpoint in your virtual network, link the privatelink.openai.azure.com private DNS zone, and set public network access to Disabled, so the only path in is an authenticated call from inside your network.
Who this is for and what you will have at the end
This guide is for platform, security and application engineers who run workloads on Azure OpenAI in Microsoft Foundry and need to meet a "no shared secrets, no public endpoint" requirement. It assumes the calling application runs in Azure (App Service is used in the examples, but the same pattern applies to Functions, Container Apps and virtual machines).
At the end you will have:
- An application that authenticates to Azure OpenAI with a managed identity and short-lived Entra tokens, with no key in configuration or Key Vault.
- Least-privilege Azure RBAC on the Azure OpenAI resource.
- Key-based (local) authentication disabled, enforced by Azure Policy for new resources.
- A private endpoint with working private DNS, and public network access disabled.
- A verification routine and a troubleshooting table for the errors you are most likely to hit.
How the controls fit together
Identity and network controls are independent layers. Entra ID decides who may call the model; the network settings decide from where. You want both, and you want to enable them in an order that never leaves the application broken.
App Service (managed identity)
| 1. token for https://ai.azure.com/.default from Microsoft Entra ID
| 2. outbound through VNet integration subnet
v
Private endpoint (subnet in your VNet, private IP)
| DNS: contoso-aoai.openai.azure.com -> CNAME contoso-aoai.privatelink.openai.azure.com -> 10.x.x.x
v
Azure OpenAI resource
- disableLocalAuth: true (keys rejected with 401)
- publicNetworkAccess: Disabled (internet callers rejected)
- RBAC: Cognitive Services OpenAI User for the app identity| Control | Setting | What it stops |
|---|---|---|
| Entra ID with RBAC | Cognitive Services OpenAI User on the resource | Callers without a role assignment get 403 |
| Disable local auth | disableLocalAuth: true | Any request using an api-key header |
| Private endpoint | Subresource account, zone privatelink.openai.azure.com | Nothing by itself; it adds a private path |
| Public network access | publicNetworkAccess: Disabled | Every request that doesn't arrive through a private endpoint |
| Trusted services exception | networkAcls.bypass: AzureServices | Lets Azure AI Search, Azure Machine Learning and Foundry Tools through with their managed identities |
The safe order is: identity first, then remove keys, then add the private path, then close the public path. Each step can be verified before the next one.
Prerequisites
- An Azure OpenAI (or Foundry) resource with a custom subdomain, for example
https://contoso-aoai.openai.azure.com. Entra ID authentication requires one; resources created in the Azure portal get it automatically. - A model deployment on the resource. Note the deployment name, which is what you pass as
model. - Permission to create role assignments on the resource (for example Role Based Access Control Administrator or User Access Administrator).
- A virtual network with one subnet for private endpoints and, for App Service, a separate empty subnet for virtual network integration. App Service integration requires a subnet of at least
/28(Microsoft recommends/26), delegated toMicrosoft.Web/serverFarms, and the subnet can't contain other resources such as private endpoints. - Azure CLI and the Az PowerShell module, including a current
Az.CognitiveServices.
Step 1: Give the application a managed identity
Use a system-assigned identity when one app needs access, or a user-assigned identity when several apps share the same access or you want to grant the role before the app exists.
# System-assigned identity on an App Service app
az webapp identity assign --resource-group rg-app --name app-contoso-chat
# Or: create a user-assigned identity and attach it
az identity create --resource-group rg-app --name id-contoso-chat
az webapp identity assign --resource-group rg-app --name app-contoso-chat \
--identities /subscriptions/<sub-id>/resourceGroups/rg-app/providers/Microsoft.ManagedIdentity/userAssignedIdentities/id-contoso-chatRecord the identity's object (principal) ID for the role assignment, and for a user-assigned identity also its client ID, which the code needs.
Step 2: Assign Cognitive Services OpenAI User
The inference role is Cognitive Services OpenAI User. It lets the identity call deployed models with Entra ID and see the endpoint, but not read or regenerate keys, create deployments or fine-tune. Choose it over broader roles:
| Role | Inference with Entra ID | Create or edit deployments | View or regenerate keys |
|---|---|---|---|
| Cognitive Services OpenAI User | Yes | No | No |
| Cognitive Services OpenAI Contributor | Yes | Yes | No |
| Cognitive Services Contributor | No | Yes | Yes |
That last row catches many teams: Cognitive Services Contributor can manage the resource and read keys, but it can't make inference calls with an Entra token. Microsoft also notes that subscription-level Owner and Contributor roles are inherited and take priority over Azure OpenAI roles applied at resource group level.
Assign the role at resource scope. Specifying the principal type avoids failures caused by replication delay right after an identity is created:
AOAI_ID=$(az cognitiveservices account show --resource-group rg-ai \
--name contoso-aoai --query id --output tsv)
az role assignment create \
--assignee-object-id "<principal-id-of-the-identity>" \
--assignee-principal-type "ServicePrincipal" \
--role "Cognitive Services OpenAI User" \
--scope "$AOAI_ID"Give developers the same role on non-production resources so they can test with az login instead of keys.
Step 3: Switch the client to Entra ID tokens
With the OpenAI Python library and Azure Identity, the client requests a token for the https://ai.azure.com/.default scope and sends it as a bearer token. get_bearer_token_provider caches and refreshes tokens for you.
from openai import OpenAI
from azure.identity import DefaultAzureCredential, get_bearer_token_provider
token_provider = get_bearer_token_provider(
DefaultAzureCredential(), "https://ai.azure.com/.default"
)
client = OpenAI(
base_url="https://contoso-aoai.openai.azure.com/openai/v1/",
api_key=token_provider, # a callable: a fresh Entra token is sent per request
)
response = client.chat.completions.create(
model="gpt-4o", # your deployment name
messages=[{"role": "user", "content": "Summarize our travel policy."}],
)
print(response.choices[0].message.content)The parameter is still called api_key, but it receives a callable, so no key is created, stored or sent. DefaultAzureCredential uses your az login session on a laptop and the managed identity in Azure. If the host has a user-assigned identity, set the AZURE_CLIENT_ID app setting to that identity's client ID, or pass ManagedIdentityCredential(client_id=...) directly.
Remove the old key from app settings and Key Vault references once the new code is deployed and healthy. Other Azure services that call Azure OpenAI with their own managed identity, such as API Management or Azure AI Search, need the same role assignment on the resource.
Step 4: Disable key-based authentication
Once every caller uses Entra ID, turn keys off. For an existing resource:
Connect-AzAccount
Set-AzCognitiveServicesAccount -ResourceGroupName "rg-ai" -Name "contoso-aoai" -DisableLocalAuth $true
# Confirm the setting
(Get-AzCognitiveServicesAccount -ResourceGroupName "rg-ai" -Name "contoso-aoai").DisableLocalAuthIf PowerShell reports that -DisableLocalAuth isn't recognized, run Update-Module -Name Az.CognitiveServices.
The change isn't instantaneous. The control plane updates at once, but the shared gateway can keep accepting previously valid keys until its cache refreshes, which Microsoft says usually takes a few minutes and can take several hours. Don't treat the cutoff as immediate for an incident response; verify it as shown later.
To stop new resources from being created with keys enabled, assign the built-in policy Azure AI Services resources should have key access disabled (disable local authentication) at subscription or resource group scope. In Bicep or ARM templates, set disableLocalAuth: true on the account.
Step 5: Add a private endpoint and private DNS
Create the private endpoint in the endpoint subnet. For Azure OpenAI the subresource (group ID) is account, and the recommended private DNS zone is privatelink.openai.azure.com.
az network private-endpoint create \
--resource-group rg-network \
--name pe-contoso-aoai \
--vnet-name vnet-app \
--subnet snet-private-endpoints \
--private-connection-resource-id "$AOAI_ID" \
--group-id account \
--connection-name contoso-aoai-connection
az network private-dns zone create \
--resource-group rg-network \
--name "privatelink.openai.azure.com"
az network private-dns link vnet create \
--resource-group rg-network \
--zone-name "privatelink.openai.azure.com" \
--name link-vnet-app \
--virtual-network vnet-app \
--registration-enabled false
az network private-endpoint dns-zone-group create \
--resource-group rg-network \
--endpoint-name pe-contoso-aoai \
--name default \
--private-dns-zone "privatelink.openai.azure.com" \
--zone-name openaiA Foundry resource can expose three endpoints: openai.azure.com, cognitiveservices.azure.com and services.ai.azure.com. Microsoft's guidance is to configure the matching private zone (privatelink.openai.azure.com, privatelink.cognitiveservices.azure.com, privatelink.services.ai.azure.com) for each endpoint your workload actually uses. If you run your own DNS servers, Microsoft lists openai.azure.com as the public DNS zone to forward for this resource type; your DNS servers must resolve the resource name to the private endpoint IP, either by forwarding to the Azure private DNS zone or by hosting the privatelink zone and its A records yourself. The same pattern applies to every PaaS service you reach through private endpoints in a landing zone; see the enterprise Azure migration playbook for how that fits into a Corp landing zone design.
Keep calling the custom subdomain. Clients must use https://contoso-aoai.openai.azure.com; DNS resolves it to the private IP through a CNAME. Never put the *.privatelink.openai.azure.com name in configuration.
For App Service, route outbound traffic into the virtual network so the app reaches the private IP:
az webapp vnet-integration add \
--resource-group rg-app \
--name app-contoso-chat \
--vnet vnet-app \
--subnet snet-appservice-integration
az resource update \
--resource-group rg-app \
--name app-contoso-chat \
--resource-type "Microsoft.Web/sites" \
--set properties.outboundVnetRouting.allTraffic=trueStep 6: Close the public endpoint
When the app resolves the private IP and calls succeed, disable public network access:
Set-AzCognitiveServicesAccount -ResourceGroupName "rg-ai" -Name "contoso-aoai" -PublicNetworkAccess "Disabled"If Azure AI Search, Azure Machine Learning or other Foundry Tools must call the model (for example an indexer's embedding skill), allow trusted services. In the portal, open the resource, select Networking, and under Exceptions select Allow Azure services on the trusted services list to access this cognitive services account. The equivalent property is networkAcls.bypass set to AzureServices; set it to None to revoke. Those services still authenticate with their managed identities, so they need a role assignment too. The integrated vectorization guide in this series covers that case.
Microsoft notes that network restrictions also block requests that come from the Azure portal, so plan for portal and playground access to the resource only from networks that can reach the private endpoint.
Infrastructure as code
The same end state in Bicep, using the GA API version 2026-07-01:
resource aoai 'Microsoft.CognitiveServices/accounts@2026-07-01' = {
name: 'contoso-aoai'
location: location
kind: 'OpenAI'
sku: { name: 'S0' }
properties: {
customSubDomainName: 'contoso-aoai'
disableLocalAuth: true
publicNetworkAccess: 'Disabled'
networkAcls: {
defaultAction: 'Deny'
bypass: 'AzureServices'
}
}
}Deploy the private endpoint, DNS zone, virtual network link and role assignment in the same template so a fresh environment never exists in a half-open state.
Verify the configuration
Run these checks from a VM or Cloud Shell inside the linked virtual network, and once from outside it.
- DNS. From inside the network, the resource name must resolve to the private IP through the
privatelinkalias:
Resolve-DnsName contoso-aoai.openai.azure.comFrom the internet, the same name resolves to a public address, which is expected. DNS answers don't grant access; the service rejects the connection when public access is disabled.
- Entra token works. From inside the network, with an identity that has the role:
TOKEN=$(az account get-access-token --resource https://ai.azure.com --query accessToken --output tsv)
curl https://contoso-aoai.openai.azure.com/openai/v1/chat/completions \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-4o", "messages": [{"role": "user", "content": "Hello"}]}'-
Keys are rejected. Repeat the call with
-H "api-key: <old key>". Expect HTTP 401 withAccess denied due to invalid subscription key or wrong API endpoint. Until you see that, treat keys as still valid. -
Public path is closed. The same token-authenticated call from outside the virtual network must fail.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
401 Unauthorized with a valid token | Client uses a regional endpoint, or the resource has no custom subdomain | Configure a custom subdomain and use https://<name>.openai.azure.com |
403 Forbidden or PermissionDenied right after granting the role | Role assignment hasn't propagated | Wait up to five minutes and retry before changing anything else |
| Works locally, fails in Azure | Role was assigned to your user, not the app identity, or the identity isn't enabled | Assign Cognitive Services OpenAI User to the app's identity |
| Fails in Azure only with a user-assigned identity | Several identities on the host and none selected | Set AZURE_CLIENT_ID or pass the client ID to ManagedIdentityCredential |
| Access still wrong hours after a group membership change | Group and role memberships are claims in the token, and managed identity back ends cache tokens per resource URI for around 24 hours | Assign the role directly to the identity instead of through a group, and grant it before rollout |
DefaultAzureCredential failed to retrieve a token on a laptop | Not signed in, or signed in to the wrong tenant | az login --tenant <tenant-id> |
| Old keys still work after disabling local auth | Gateway cache hasn't refreshed | Wait and re-test for the 401; propagation can take hours |
| App times out after public access is disabled | App traffic isn't entering the VNet, or the DNS zone isn't linked to the VNet it uses | Check VNet integration, route-all, the zone link and the A record |
404 Not Found on a correct endpoint | model doesn't match a deployment name | Use the deployment name, which can differ from the model name |
Closing checklist
- Every caller uses a managed identity or developer sign-in; no key in app settings, pipelines or Key Vault.
- App identities hold Cognitive Services OpenAI User at resource scope, nothing broader.
disableLocalAuthistrue, the old key returns401, and the deny policy is assigned for new resources.- The private endpoint uses subresource
account, theprivatelink.openai.azure.comzone is linked to every virtual network that calls the model, and clients use the custom subdomain. publicNetworkAccessisDisabled, with the trusted services exception only if Azure AI Search or similar services need it.- If several applications share the resource, put a gateway in front for per-app limits; see Azure API Management AI Gateway for OpenAI. For the wider production picture, see production LLMOps for enterprise RAG.
References
- https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/managed-identity
- https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/role-based-access-control
- https://learn.microsoft.com/en-us/entra/identity/managed-identities-azure-resources/managed-identity-best-practice-recommendations
- https://learn.microsoft.com/en-us/azure/ai-services/disable-local-auth
- https://learn.microsoft.com/en-us/azure/ai-services/cognitive-services-virtual-networks
- https://learn.microsoft.com/en-us/azure/private-link/private-endpoint-dns
- https://learn.microsoft.com/en-us/azure/private-link/create-private-endpoint-cli
- https://learn.microsoft.com/en-us/azure/role-based-access-control/role-assignments-cli
- https://learn.microsoft.com/en-us/azure/app-service/overview-managed-identity
- https://learn.microsoft.com/en-us/azure/app-service/configure-vnet-integration-enable
- https://learn.microsoft.com/en-us/powershell/module/az.cognitiveservices/set-azcognitiveservicesaccount
- https://learn.microsoft.com/en-us/azure/templates/microsoft.cognitiveservices/accounts