When an Azure OpenAI model version reaches its retirement date, every request to a deployment still running that version returns 410 Gone. Standard, Global Standard and Data Zone Standard deployments are upgraded automatically on a rolling, region-by-region schedule, unless versionUpgradeOption is set to NoAutoUpgrade, in which case they simply stop working. Provisioned deployments are never auto-upgraded and must be migrated by you. A safe runbook therefore has four parts: inventory every deployment and its retirement date, set versionUpgradeOption deliberately, validate the replacement model against your own data, and migrate provisioned deployments well before the cutoff.
Who this is for and what you will have
This runbook is for platform teams and application owners who run Azure OpenAI deployments in production, especially those with several resources, regions or provisioned throughput. At the end you will have:
- An inventory of deployments with model, version, SKU and upgrade policy.
- Retirement dates pulled from the Models API rather than from memory.
- A chosen
versionUpgradeOptionfor each deployment, set with PowerShell or REST. - A validation step that compares the current and replacement model on the same frozen dataset.
- A tested procedure for in-place and side-by-side provisioned migrations.
How the model lifecycle works
Every model in the Foundry catalog is in one of five stages:
| Stage | New deployments | Existing deployments |
|---|---|---|
| Preview | Yes | Yes |
| Generally Available | Yes | Yes |
| Legacy (optional) | Yes, until deprecation | Yes |
| Deprecated | Existing customers only | Yes |
| Retired | No | No, requests return 410 Gone |
"Existing customer" is decided per subscription: whether that Azure subscription has ever deployed the specific model version. A new subscription in the same tenant doesn't inherit access, which matters when you build a new environment late in a model's life.
For GA models on the standard lifecycle, the timeline is set at launch:
| Point in time | What happens |
|---|---|
| Launch | Retirement date set 18 months out and exposed in the Models API |
| 12 months | Deprecated: new customers can't deploy it |
| About 90 days before retirement | Replacement available to test in Global Standard |
| About 30 days before retirement | Replacement available in provisioned regions where the predecessor is retiring |
| 18 months | Retired: inference returns 410 Gone |
Microsoft notes exceptions: GA models from Anthropic, DeepSeek, Fireworks and Mistral AI follow a 12-month lifecycle, preview models usually start with a "not sooner than" date about 90 days out and are force-upgraded or retired with at least 30 days notice, and a model with security or compliance issues can be retired early with shortened notice. Fine-tuned models have separate training and deployment retirement dates. Retirement dates can't be extended.
Prerequisites
- Reader access to the subscriptions you want to inventory, and Cognitive Services Contributor (or Contributor) on resources you will change.
- Azure CLI for
az rest, and the Az PowerShell module for theAz.CognitiveServicescmdlets. - A Log Analytics workspace or at least portal access to Azure Monitor metrics for each resource.
- A representative test dataset for each workload, described in Step 5.
Step 1: Inventory deployments and upgrade policy
List deployments per resource with the Deployments List API. If versionUpgradeOption is missing from a deployment, its value is null:
SUBSCRIPTION_ID="00000000-0000-0000-0000-000000000000"
RG="rg-ai-prod"; ACCOUNT="contoso-aoai"
az rest --method get \
--url "https://management.azure.com/subscriptions/$SUBSCRIPTION_ID/resourceGroups/$RG/providers/Microsoft.CognitiveServices/accounts/$ACCOUNT/deployments?api-version=2023-05-01" \
--query "value[].{deployment:name, sku:sku.name, capacity:sku.capacity, model:properties.model.name, version:properties.model.version, upgrade:properties.versionUpgradeOption}" \
-o tableThe PowerShell equivalent returns the same properties:
Get-AzCognitiveServicesAccountDeployment -ResourceGroupName 'rg-ai-prod' -AccountName 'contoso-aoai' |
Select-Object Name,
@{n='Sku'; e={$_.Sku.Name}},
@{n='Model'; e={$_.Properties.Model.Name}},
@{n='Version'; e={$_.Properties.Model.Version}},
@{n='Upgrade'; e={$_.Properties.VersionUpgradeOption}}Record the deployment type for each row. Anything with a provisioned SKU (ProvisionedManaged, DataZoneProvisionedManaged or GlobalProvisionedManaged) needs a manual migration plan.
Step 2: Pull retirement dates from the Models API
The Models API returns every model available in a region with its lifecycle status and deprecation dates:
az rest --method get \
--url "https://management.azure.com/subscriptions/$SUBSCRIPTION_ID/providers/Microsoft.CognitiveServices/locations/eastus/models?api-version=2024-10-01" \
--query "value[?model.name=='gpt-4o'].{version:model.version, status:model.lifecycleStatus, inferenceRetires:model.deprecation.inference, isDefault:model.isDefaultVersion}" \
-o tableEach SKU in model.skus also carries its own deprecationDate, which can differ between Standard and provisioned SKUs.
The API uses different words than the documentation and portal, and this is the most common source of confusion:
| Docs and portal stage | lifecycleStatus | Meaning |
|---|---|---|
| Preview | Preview | Experimental |
| Generally Available | GenerallyAvailable | Production-ready |
| Deprecated | Deprecating | Still serves inference, blocked for new customers |
| Retired | Deprecated | Inference returns 410 Gone |
Microsoft's guidance is to check both fields: treat a deprecation.inference date in the past as retired even if lifecycleStatus hasn't caught up. Join this output to the Step 1 inventory on model name and version, and sort by retirement date. That list is your migration backlog.
Step 3: Turn on notifications
Microsoft emails subscription owners with active deployments at least 60 days before a GA model retires (30 days for preview models), and posts Azure Service Health advisories. Owners are often not the people running the workload, so route the advisories:
- In the Azure portal, open Service Health > Health advisories.
- Filter the service to Azure OpenAI Service. There is no separate Microsoft Foundry entry in Service Health.
- Create an alert rule that sends email, SMS or a webhook to the team that owns the deployments.
Combine this with a scheduled run of the Step 2 query so a new retirement date appears in your backlog even if an email is missed.
Step 4: Set versionUpgradeOption deliberately
Automatic upgrades only apply to Standard deployment types. Three options exist:
| Value | Behaviour |
|---|---|
OnceNewDefaultVersionAvailable | Upgrades within two weeks after a new version becomes the default |
OnceCurrentVersionExpired | Upgrades to the current default version at the retirement date |
NoAutoUpgrade | Never upgrades; the deployment stops working at retirement |
A null value behaves like OnceCurrentVersionExpired. A practical policy: use OnceNewDefaultVersionAvailable for development and early testing, OnceCurrentVersionExpired for production workloads you pin and migrate on your own schedule (it keeps a safety net), and NoAutoUpgrade only when an unvalidated model change is worse for you than an outage, and your monitoring will catch a 410 immediately.
The Azure CLI can read the property but can't currently update it. Use PowerShell:
$deployment = Get-AzCognitiveServicesAccountDeployment -ResourceGroupName 'rg-ai-prod' -AccountName 'contoso-aoai' -Name 'gpt-4o-prod'
$deployment.Properties.VersionUpgradeOption = 'OnceCurrentVersionExpired'
New-AzCognitiveServicesAccountDeployment -ResourceGroupName 'rg-ai-prod' -AccountName 'contoso-aoai' -Name 'gpt-4o-prod' `
-Properties $deployment.Properties -Sku $deployment.SkuOr REST, sending the existing SKU and model with the new value:
az rest --method put \
--url "https://management.azure.com/subscriptions/$SUBSCRIPTION_ID/resourceGroups/$RG/providers/Microsoft.CognitiveServices/accounts/$ACCOUNT/deployments/gpt-4o-prod?api-version=2025-06-01" \
--body '{"sku": {"name": "Standard", "capacity": 120},
"properties": {"model": {"format": "OpenAI", "name": "gpt-4o", "version": "2024-11-20"},
"versionUpgradeOption": "OnceCurrentVersionExpired"}}'If you manage deployments with Bicep or Terraform, set the property in code too, otherwise the next deployment run can reset it.
Step 5: Validate the replacement model
An endpoint that still answers after an upgrade doesn't mean the application still behaves correctly. Microsoft's model migration guidance describes six phases: Discover, Assess, Adapt, Validate, Roll out and Retire. The parts that prevent incidents:
- Freeze a test dataset first. Collect representative inputs, expected outputs and success criteria as CSV or JSONL from captured production traffic or curated examples. Keep them fixed for the whole migration; if they change, results stop being comparable. Content capture is opt-in and never retroactive, so start logging prompts, responses, latency and token counts before you need them.
- Confirm the target is deployable. Check region, deployment type, quota and capacity, and that the replacement can run next to the current model so rollback stays possible. The RPM-to-TPM ratio can differ between model families; see fixing Azure OpenAI 429 rate limit errors for how quota maps to limits.
- Replay unchanged, then adapt. Run the current prompts on the new model with no changes to get a baseline, then adjust prompts, parameters, tool definitions, output schemas and calling code. Parameters such as
temperature,top_p,max_tokensand reasoning-effort controls don't map one-to-one across generations; for example, reasoning models don't acceptmax_tokensand only work withmax_completion_tokensin the Chat Completions API (max_output_tokensin the Responses API). - Score source and target with the same evaluators. The Azure AI Evaluation SDK provides more than 30 built-in evaluators, including groundedness, relevance, coherence, fluency, F1, BLEU, ROUGE, safety and agent evaluators, plus custom LLM-as-judge graders. Run the current model first to set the baseline, then the target, and compare quality, latency and cost per request together.
Approach the evaluation the same way as the rest of your LLM operations; production LLMOps for enterprise RAG covers where the evaluation gate sits in a release pipeline.
Step 6: Roll out on Standard deployments
For Standard-family deployments you have two options:
- Side by side: create a new deployment of the replacement version, route a small share of traffic to it from your gateway or application, watch errors, latency and quality, then move the rest. Keep the old deployment reachable until you are confident.
- In place: update the model version on the existing deployment with the same PUT request shown in Step 4, changing
properties.model.version. This is faster but gives you no parallel rollback path.
Quota carries over automatically when Microsoft auto-upgrades a Standard deployment. When you create a new deployment yourself, it needs quota for the target model in that region.
Step 7: Migrate provisioned deployments
Provisioned deployments have no safety net, so schedule them first. Before you start, confirm that the target model and version support your deployment type (migrations only happen between provisioned deployments of the same type), that capacity is available in the region, and for side-by-side moves that you have quota for both deployments at once.
In-place migration keeps the deployment name and PTU count. Azure moves traffic to the new model over a 20 to 30 minute window, during which the provisioning state shows Updating and the deployment keeps serving requests. To change only the version, edit the deployment in the Foundry portal (Deployments, select the deployment, Edit, choose the new model version). To change model family, use REST or the Azure CLI, for example:
az rest --method put \
--url "https://management.azure.com/subscriptions/$SUBSCRIPTION_ID/resourceGroups/$RG/providers/Microsoft.CognitiveServices/accounts/$ACCOUNT/deployments/gpt-4o-ptu-deployment?api-version=2024-10-01" \
--body '{"sku": {"name": "GlobalProvisionedManaged", "capacity": 100},
"properties": {"model": {"format": "OpenAI", "name": "gpt-4o-mini", "version": "2024-07-18"}}}'Side-by-side migration gives you control over pace:
- Create a new provisioned deployment with the target model.
- Shift traffic from the old deployment to the new one at the rate you choose.
- Confirm the old deployment has received no calls: the Azure OpenAI Requests metric, split by
ModelDeploymentName, should show no API calls for 5 to 10 minutes after traffic moved. - Delete the old deployment.
Deleting a deployment doesn't cancel or change a PTU reservation. Review reservations afterwards so they still match the deployment type and region you run.
Verify
- Rerun the Step 1 inventory: every deployment shows the expected model, version and
versionUpgradeOption, withprovisioningStateback toSucceeded. - Chart Azure OpenAI Requests split by
ModelDeploymentNameandStatusCode: traffic is on the new deployments and there are no410responses. - Compare live quality signals against the pre-migration baseline for at least the first days after rollout.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
410 Gone on every request | Deployment still on a retired version: provisioned, or NoAutoUpgrade | Create or update a deployment on a supported version and repoint the application |
Model shows Deprecated in the API but works in another region | API Deprecated means retired; regions can differ | Check deprecation.inference and per-SKU deprecationDate for each region |
| Can't deploy a model that other subscriptions use | Model is deprecated and the subscription never deployed that version | Deploy a supported model, or use the subscription that already has it |
az can't change the upgrade option | Not supported in Azure CLI | Use PowerShell or the REST API |
| 401 or 403 from Azure Resource Manager | Expired token or missing permissions | Refresh the token and check deployment write access |
| Provisioned in-place migration fails | Different deployment type or no capacity | Check supported deployment types and capacity, or migrate side by side |
| Fine-tuned deployment not upgraded | Fine-tuned models aren't auto-upgraded | Re-tune on the replacement base model before the deployment retirement date |
Checklist
- Inventory every deployment with model, version, SKU and
versionUpgradeOption. - Pull retirement dates from the Models API and read
DeprecatingversusDeprecatedcorrectly. - Route Service Health advisories for Azure OpenAI Service to the owning team.
- Set
versionUpgradeOptionexplicitly, in code and in infrastructure templates. - Freeze a test dataset and score current and replacement models with the same evaluators.
- Migrate provisioned deployments before the replacement window closes.
- Confirm zero traffic on old deployments, delete them and review PTU reservations.
References
- Foundry Models lifecycle and support policy
- Working with Azure OpenAI models
- Model migration: upgrade or switch models in Microsoft Foundry
- Models - List REST API
- Deployments - Create Or Update REST API
- Operate provisioned throughput deployments in production
- Azure OpenAI v1 API and changelog
- Azure OpenAI reasoning models
- Provisioned throughput billing and cost management