To protect Azure VMs against a regional outage, create a Recovery Services vault outside the source region, open Site Recovery > Azure virtual machines > Enable replication, select the VMs and accept or customize the target region resources Site Recovery proposes. Site Recovery installs the Mobility service extension, copies disk writes through a cache storage account to replica managed disks in the target region, and creates crash-consistent recovery points every five minutes. Once a VM reaches the Protected state, run Test Failover into an isolated virtual network, validate the application, then select Cleanup test failover; replication and production are unaffected throughout.
Who this is for and what you will have at the end
This guide is for administrators who run production workloads on Azure VMs and need a documented, tested way to bring them up in another region. It focuses on the decisions that make a drill meaningful, not only on getting replication green.
At the end you will have:
- Outbound connectivity that lets the Mobility service replicate without exceptions.
- VMs replicating to a target region with a replication policy you chose.
- A recovery plan that starts tiers in order.
- A completed test failover, validated and cleaned up, with notes recorded.
How Azure-to-Azure replication works
When you enable replication, the Site Recovery Mobility service extension installs on the VM and registers it. Disk writes go immediately to a cache storage account in the source region; Site Recovery processes them and writes them to replica managed disks in the target region, named with an -ASRReplica suffix. Crash-consistent recovery points are created every five minutes, and app-consistent points at the frequency in the replication policy.
Site Recovery can create the target resources for you:
| Target resource | Default |
|---|---|
| Subscription | Same as the source |
| Resource group | New group in the target region with an asr suffix |
| Virtual network | New VNet and subnet with an asr suffix, plus a network mapping in both directions |
| Availability set | New set with an asr suffix if the source VM is in one |
| Replication policy | Recovery point retention one day, app-consistent snapshots disabled |
You can change most target settings after replication starts, including the VM size, but not the availability type (single instance, availability set or zone); changing that means disabling and re-enabling replication.
Prerequisites
Permissions and capacity
- Owner or administrator rights on the subscription to create the vault.
- Site Recovery Contributor to manage replication in the vault.
- Virtual Machine Contributor (or equivalent rights to create VMs in the chosen network and write to storage and managed disks) in the target region.
- Enough quota in the target region for VM sizes matching the source. Site Recovery picks the same size or the closest available one.
The VMs must use managed disks. Unmanaged Azure VM disks retired on 31 March 2026, and the current procedures support managed disks only.
Outbound connectivity
Replicated VMs need outbound HTTPS only; Site Recovery never needs inbound connectivity to the VM. If you control egress with network security groups, use service tags. IP address allowlists are no longer supported.
| Service tag (HTTPS 443) | Purpose |
|---|---|
Storage.<region> | Write to the cache storage account (source region rule) and target storage |
AzureActiveDirectory | Authentication to Site Recovery |
EventHub.<region> | Site Recovery monitoring |
AzureSiteRecovery | The Site Recovery service |
GuestAndHybridManagement | Only if you let Site Recovery upgrade the Mobility agent automatically |
AzureKeyVault | Only for enabling replication of Azure Disk Encryption VMs in the portal |
With a URL-based firewall, allow *.blob.core.windows.net, login.microsoftonline.com, *.hypervrecoverymanager.windowsazure.com and *.servicebus.windows.net. Site Recovery doesn't support authentication proxies for this scenario. The Mobility agent also queries the Azure Instance Metadata Service, so make sure any proxy bypasses 169.254.169.254.
Certificates and DNS
Install current Windows updates (or your Linux distribution's CA bundle) so the trusted root certificates are present; without them registration fails. If the VMs use custom DNS servers, make sure those servers are reachable from the target region, otherwise reprotection after a failover can fail.
Step 1: Create the vault
- In the portal, open Recovery Services vaults and select + Create.
- Choose the subscription and resource group, name the vault, and pick a region that is not the source region. The target region is a sensible choice.
- Select Review + create, then Create, and open the vault when deployment finishes.
Step 2: Enable replication
- In the vault overview select Enable Site Recovery, then under Azure virtual machines select Enable replication.
- On Source, pick the region, subscription and resource group of the VMs, keep Resource Manager and leave Disaster recovery between availability zones at No for cross-region protection.
- On Virtual machines, select up to 10 VMs per run.
- On Replication settings, review the target region, subscription, resource group, network and storage. For VMs with a high data change rate, open Storage > View/edit storage configuration and set Churn for the VM to High Churn, which uses a Premium Block Blob cache account. VMs with Premium SSD v2 disks require High Churn.
- On Manage, choose the replication policy. Create a Replication group only for VMs that run the same workload and need shared, multi-VM-consistent recovery points; this costs performance and requires the VMs to talk to each other on port 20004. Under Extension settings, choose whether Site Recovery updates the Mobility agent and which automation account it uses.
- On Review, select Enable replication.
The VMs appear under Replicated items. Initial replication copies all data first; when it completes, the state becomes Protected and you can test.
Policy and monitoring with PowerShell
A custom policy with 24 hours of retention and app-consistent snapshots every 4 hours, then a health check of everything in the protection container:
# $vault from Get-AzRecoveryServicesVault; $PrimaryProtContainer from
# Get-AzRecoveryServicesAsrFabric and Get-AzRecoveryServicesAsrProtectionContainer
Set-AzRecoveryServicesAsrVaultContext -Vault $vault
$job = New-AzRecoveryServicesAsrPolicy -AzureToAzure -Name "A2A-24h-AppCon4h" `
-RecoveryPointRetentionInHours 24 -ApplicationConsistentSnapshotFrequencyInHours 4
Get-AzRecoveryServicesAsrJob -Job $job
Get-AzRecoveryServicesAsrReplicationProtectedItem -ProtectionContainer $PrimaryProtContainer |
Select-Object FriendlyName, ProtectionState, ReplicationHealthKeep the app-consistent frequency shorter than the retention period. App-consistent snapshots use VSS (copy-only, so SQL Server log backup chains aren't changed) and cost some performance, so enable them only where the application needs them. Longer retention also increases storage cost.
Step 3: Prepare the test failover network
A drill is only as good as the network it runs in. Microsoft recommends a test network isolated from production that mirrors it:
- The same number of subnets with the same names.
- The same address ranges.
- DNS pointing at a DNS server VM that you fail over with the test, or a copy of your Active Directory environment if the application depends on it.
When you test a single VM, choose a non-production network, and not the network Site Recovery created when you enabled replication. During a recovery plan test, Site Recovery tries to place each VM in a subnet with the same name and the same IP address as its Compute and Network settings; if no subnet with that name exists it uses the first subnet alphabetically, and if the IP is taken it assigns another one.
For finer control, open Replicated items, select the VM, then Network > Edit. You can choose a test failover virtual network and, per NIC, pre-created resources such as subnet, public IP, internal load balancer and network security group for both test failover and failover. The resources must be in the same subscription and region as the target VM, and the public IP and load balancer SKUs must match.
Step 4: Build a recovery plan
Single-VM failover works, but applications need ordering.
- In the vault, select Recovery Plans (Site Recovery) > +Recovery Plan.
- Name it, select the source and target Azure regions, and choose Resource Manager.
- In Select items virtual machines, add the replicated VMs or a replication group. All VMs in one plan must replicate into a single subscription.
- Right-click the plan, select Customize, and use +Group to split tiers, for example database, application, web. A plan can have up to seven groups; machines in the same group start together.
- Add a pre- or post-action to a group: Manual action pauses the plan with instructions for the operator, and Script attaches an Azure Automation runbook. Azure-to-Azure supports runbooks for both failover and failback. Specify whether a manual action applies to test failover as well.
Step 5: Run the test failover
For a single VM:
- Open Replicated items, select the VM, and confirm it is protected and healthy on Overview.
- Select Test Failover.
- Choose a recovery point:
- Latest processed: the newest point already processed; lowest recovery time.
- Latest: processes all pending data first; lowest recovery point objective.
- Latest app-consistent: the newest application-consistent point.
- Custom: a specific point (single-VM failover only).
- Select the isolated Azure virtual network and select OK.
- Watch the notifications. The job runs a prerequisites check, prepares the data, optionally creates the latest point, then starts the VM.
For a recovery plan, open Recovery Plans > plan name > Test Failover. The plan offers the same options plus Latest multi-VM processed and Latest multi-VM app-consistent for replication groups. Track progress on the Jobs tab.
The same test in PowerShell, including cleanup:
$TFOVnet = New-AzVirtualNetwork -Name "vnet-dr-test" -ResourceGroupName "rg-dr-weu" -Location "West Europe" -AddressPrefix "10.3.0.0/16"
Add-AzVirtualNetworkSubnetConfig -Name "default" -VirtualNetwork $TFOVnet -AddressPrefix "10.3.0.0/20" | Set-AzVirtualNetwork
$rpi = Get-AzRecoveryServicesAsrReplicationProtectedItem -FriendlyName "app01" -ProtectionContainer $PrimaryProtContainer
$tfo = Start-AzRecoveryServicesAsrTestFailoverJob -ReplicationProtectedItem $rpi -AzureVMNetworkId $TFOVnet.Id -Direction PrimaryToRecovery
Get-AzRecoveryServicesAsrJob -Job $tfo
# After validation
$cleanup = Start-AzRecoveryServicesAsrTestFailoverCleanupJob -ReplicationProtectedItem $rpi
Get-AzRecoveryServicesAsrJob -Job $cleanup | Select-Object StateValidate, then clean up
After the job succeeds, the test VM appears under Virtual machines in the target region. Check that it is running, correctly sized and attached to the network you chose. Use Boot diagnostics to see a screenshot of the VM if you can't connect. To connect over RDP or SSH, the test VM's network security group must allow the port, and you either add a public IP or connect from a host inside the test network. Then test the application itself: service start-up, database connectivity, name resolution and any dependency on resources outside the plan.
When finished, select Cleanup test failover, write your observations in Notes, and select Testing is complete. Site Recovery deletes the test VMs. Any changes made inside them are discarded.
If you must test inside the production network in the target region, shut down the primary VM first so two machines with the same identity don't run on the same network, and accept that the application is down for the duration.
Verification
- Every replicated item shows Protected with normal replication health.
- The latest recovery points are minutes old, with app-consistent points at the expected interval.
- The recovery plan starts groups in the right order and pauses at manual actions.
- The test failover job completed, the application worked in the test network, and cleanup finished.
- Drill notes record recovery time, issues and owners for follow-up.
For point-in-time restores of individual VMs and files, pair this with Azure VM backup in a Recovery Services vault. If you also replicate on-premises VMware machines, see moving VMware DR from classic to the modernized appliance.
Troubleshooting
150097: Replication couldn't be enabled for the virtual machine. The subscription lacks quota or isn't enabled for the required VM size in the target region. Request quota, or replicate to a region with capacity.
151066: Site Recovery configuration failed. Trusted root certificates are missing on the VM. Install current updates, then check that the VM can reach login.microsoftonline.com.
151037 or 151072. A Site Recovery endpoint can't be reached, often because a custom DNS server isn't reachable from the region in use. Make DNS available in that region. For an on-premises proxy, set it for the Mobility service in ProxyInfo.conf under C:\ProgramData\Microsoft Azure Site Recovery\Config or /usr/local/InMage/config/; only unauthenticated proxies work.
151196 or 151197: Site Recovery configuration failed. Outbound access to Microsoft Entra ID or the Site Recovery service is blocked. Replace IP-based NSG rules with the AzureActiveDirectory and AzureSiteRecovery service tags.
151025: Site Recovery extension failed to install. The COM+ System Application or Volume Shadow Copy service is disabled. Set both to manual or automatic start.
150172: Protection couldn't be enabled ... lesser than the minimum supported size 1024 MB. A disk is smaller than 1,024 MB; resize it into the supported range and retry.
Protection fails because a replica disk exists. A previous protection left an -ASRReplica disk without the expected tags in the target resource group. Delete the disk named in the error and retry.
151083: Mobility service update finished with warnings. The filter driver update needs a restart to load. Replication keeps working; restart at a convenient time.
Checklist
- Vault in a region other than the source.
- NSG or firewall rules using service tags, IMDS reachable, root certificates current.
- Target quota confirmed and target network, resource group and availability type chosen.
- Replication policy set with retention and app-consistent frequency the application needs.
- All VMs Protected with healthy replication.
- Isolated test network mirroring production subnets, address ranges and DNS.
- Recovery plan with ordered groups and manual actions.
- Test failover run, validated, cleaned up and documented, and scheduled to repeat.
References
- Set up Azure VM disaster recovery with Azure Site Recovery
- Azure to Azure disaster recovery architecture
- Run an Azure VM disaster recovery drill
- Run a test failover to Azure
- Customize networking configurations for a failover VM
- Create and customize recovery plans
- Disaster recovery for Azure VMs using Azure PowerShell
- Troubleshoot Azure VM replication errors