Cloud & infrastructure

Migrating Enterprise VMs to Azure: A Zero-Downtime Cutover Playbook and Runbook

An enterprise migration architecture and minute-by-minute cutover runbook for moving VMware virtual machines and multi-tier business applications to Microsoft Azure with minimal, planned downtime.

Updated 6 min read
On this page

Executive Summary & Architecture Takeaways

  • Example Scope: This playbook assumes a representative estate of roughly two hundred VMware vSphere virtual machines, around fifteen business-critical multi-tier applications, and tens of terabytes of transactional data in a regulated industry. Scale the numbers to your own estate.
  • The Execution Constraint: A maximum cutover maintenance window of 4 hours on a Sunday morning, with zero business transaction loss permitted.
  • Key Architectural Decisions:
    • Build Azure landing zones (CAF) with Terraform before touching a single workload.
    • Provision a dedicated Azure ExpressRoute circuit sized for replication traffic, with a site-to-site VPN as fallback.
    • Use continuous block-level replication (Azure Migrate server migration or Azure Site Recovery) for application tiers and SQL Server Always On / PostgreSQL streaming replication for database tiers.
    • Define rollback gates and a hard go/no-go time before the window starts.

1. Enterprise-Scale Landing Zone Architecture

Before replicating any servers, establish a governance, network, and security foundation.

                         +-----------------------------------+
                         |    Enterprise Management Group    |
                         +-----------------+-----------------+
                                           |
                   +-----------------------+-----------------------+
                   |                                               |
         +---------v---------+                           +---------v---------+
         |   Platform MG     |                           |   Workloads MG    |
         +----+----+----+----+                           +----+----+----+----+
              |    |    |                                     |    |    |
      +-------+    |    +-------+                             |    |    |
      |            |            |                             |    |    |
+-----v----+ +-----v----+ +-----v----+                  +-----v----+ +--v-------+
|Management| |Connectiv.| | Identity |                  |Corp Spoke| |Online Spk|
|Subscript.| |Subscript.| |Subscript.|                  |Subscript.| |Subscript.|
+----------+ +-----+----+ +----------+                  +-----+----+ +----------+
                   |                                          |
                   |       Peer: Spoke-to-Hub Peering         |
                   +==========================================+
                   |
     +-------------v-------------+
     | ExpressRoute Gateway      | <==== Dedicated Circuit ====> On-Premises
     | Azure Firewall Premium    |                               Datacenter
     +---------------------------+

1.1 Network Segmentation Rules

  • Hub-Spoke VNet Topology: The Connectivity subscription hosts the central hub VNet. All on-premises traffic terminates on the ExpressRoute gateway.
  • Routing via Azure Firewall: Inter-spoke and internet-bound traffic is sent through the central Azure Firewall using user-defined routes (0.0.0.0/0 -> next hop: the firewall's private IP, for example 10.0.0.4).
  • No Public IP Rule: No compute workload in the Corp landing zone gets a public IP address (enforce it with Azure Policy). Administrative access goes through Azure Bastion, and PaaS services are reached via Private Endpoints.

2. Data Replication Architecture: Minimizing Cutover Delta

The critical mistake in cloud migrations is attempting to copy disk images over the network during the cutover window. As a rough calculation, 45 TB over a fully utilized 1 Gbps link takes more than four days, which no maintenance window can absorb.

2.1 The Two-Tier Continuous Sync Model

[On-Premises VMware Datacenter]                   [Azure Target Region]
  |                                                  |
  +-- Application VMs                                +-- Replica Managed Disks
  |   Continuous replication (Azure Migrate / ASR) ----> Block-level deltas applied
  |                                                  |   continuously
  |                                                  |
  +-- Database Tier                                  +-- Azure IaaS DB / Managed Instance
      SQL Always On / Postgres Replication ------------> Continuous log streaming
      (Transaction Log / WAL)                        |   (lag monitored to zero at cutover)
  1. Application Servers (Stateless / File Stores):
    • Replicated with the Azure Migrate server migration tool (Microsoft's recommended path for VMware to Azure) or Azure Site Recovery (ASR).
    • Initial replication runs weeks before cutover while systems are live.
    • Changed blocks are replicated continuously, so the replica disks stay within minutes of the source.
  2. Database Servers (Stateful Transaction Engines):
    • Exclude active database volumes from block-level replication, because a crash-consistent copy of a busy database is not something you want to promote.
    • Configure native replication instead:
      • SQL Server: Always On availability groups (or distributed availability groups) extending across ExpressRoute to SQL Server on Azure VMs, or the Managed Instance link for Azure SQL Managed Instance.
      • PostgreSQL / MySQL: streaming replication to target standby instances on Azure.

3. A Minute-by-Minute Cutover Runbook

The sample schedule below fits a Sunday 01:00 to 05:00 window. Rehearse it at least once end to end against a non-production copy before the real weekend.

Time (UTC)Action ItemOwnerVerification Gate
01:00Commence Cutover Window: Place on-premises frontends in maintenance modeRelease LeadPublic status page updated; maintenance banner active.
01:10Stop background queue processors (message consumers, cron jobs)App OpsMessage queues drained to 0 pending tasks.
01:25Validate final database replication syncLead DBASQL Server log send/redo queue = 0; Postgres replay lag = 0 bytes.
01:30Stop replication; promote Azure database instances to primaryLead DBAWrite test verified on the Azure primary.
01:45Migrate / fail over application VMs with source VMs shut down firstCloud ArchitectTarget VMs booted; expected private IPs and DNS registration confirmed.
02:15Run automated smoke-test suites (Selenium / Cypress)QA LeadCore login, checkout and ledger balance tests pass.
02:30Go/no-go decisionMigration LeadAll gates green, or rollback starts now.
02:40Update public and internal DNS recordsNetwork LeadDNS records point to Azure Application Gateway / load balancer front ends.
03:00End-to-end business integration validationBusiness LeadTest transactions executed in production and confirmed by finance.
03:20Maintenance Mode Disabled: Service live on AzureRelease LeadPublic traffic entering Azure; monitoring green.

4. Rollback Triggers & Safety Gates

A successful migration requires pre-defined, non-negotiable abort criteria. If any of the following occur before the 02:30 go/no-go point, the cutover is aborted and traffic remains on-premises:

  1. Database Promotion Failure: The Azure database does not accept writes within 15 minutes of the promotion attempt.
  2. Smoke Test Failure: More than 2 core business flows fail automated regression tests.
  3. Network Packet Loss: Sustained packet loss above 0.5% on the ExpressRoute private peering during application start-up.

Rollback Procedure: Re-enable on-premises frontend services, point DNS back (TTL lowered to 60 seconds at least 48 hours in advance), and resume background workers. Because no user writes have reached Azure before the go/no-go point, the on-premises databases are still authoritative. After go-live, rolling back requires reverse replication or accepting data loss, which is why the gate sits before DNS is switched. Aim to restore the on-premises baseline in well under 30 minutes, and time it during rehearsal.


5. Post-Migration Optimization: Reducing Cloud Waste

Lift-and-shift sizing is usually conservative. Plan a right-sizing and optimization pass in the first 60 days after migration:

  1. Azure Hybrid Benefit (AHB): Apply existing Windows Server and SQL Server licenses with Software Assurance (or qualifying subscriptions) to Azure VMs to remove the license component from compute pricing.
  2. Savings Plans & Reserved Instances: Once utilization data is stable, buy 1- or 3-year commitments for baseline workloads. Model the discount for your SKUs and region with the Azure pricing calculator rather than relying on headline percentages.
  3. Auto-Shutdown on Dev/Test: Use Azure Automation runbooks or the built-in VM auto-shutdown to stop non-production environments outside business hours.
  4. Right-Sizing: Use Azure Advisor and Azure Monitor metrics to downsize over-provisioned VMs and move suitable disks to cheaper tiers.
  5. Track the Business Case: Compare actual run cost against the on-premises baseline (hardware leases, power and cooling, virtualization licensing) quarterly, so savings are measured rather than assumed.

Questions people ask

How do you minimize downtime when migrating monolithic databases to Azure?

Do not rely on cold VM snapshot exports. Use continuous replication for application tiers (Azure Migrate or Azure Site Recovery) and native database replication for database tiers (SQL Server Always On availability groups, the Managed Instance link, or PostgreSQL streaming replication) over ExpressRoute or a site-to-site VPN. During the cutover window, drain application queues, confirm replication lag is zero, promote the Azure databases, and switch DNS records whose TTL was lowered in advance.

What is the recommended Azure Landing Zone topology for regulated enterprises?

Adopt the Microsoft Cloud Adoption Framework (CAF) Azure landing zone architecture. It separates platform subscriptions for Management, Connectivity (hub VNet with Azure Firewall and ExpressRoute gateways) and Identity (domain controllers) from Corp and Online landing zone subscriptions, with user-defined routes sending egress and inter-spoke traffic through a central firewall.

How do you resolve IP address overlap between on-premises subnets and cloud landing zones?

The cleanest fix is to allocate non-overlapping address space for Azure and re-IP where possible. When hardcoded addresses make that infeasible, options include NAT rules on Azure VPN Gateway for site-to-site connections, a NAT-capable network virtual appliance in a transit VNet, or exposing specific services through Private Link private endpoints placed in a non-conflicting address range. Azure NAT Gateway only provides outbound internet SNAT and does not solve overlap.

AzureCloud MigrationEnterprise ArchitectureTerraformDisaster Recovery
  1. Production LLMOps & Enterprise RAG: Low-Latency, Privacy-Preserving Architecture at Scale

    A blueprint for building accurate, privacy-preserving Retrieval-Augmented Generation (RAG) systems, covering hybrid dense-sparse search, cross-encoder re-ranking, semantic caching and latency tuning.

    AI engineering6 min read
  2. Automate blog and social media posting with Claude, GitHub Actions and Make

    An architecture for publishing one researched article a day and turning it into a narrated vertical video for YouTube Shorts, Instagram, Facebook and LinkedIn, with the platform limits that shape it.

    AI engineering9 min read
  3. Google Workspace to Microsoft 365 Migration: A Complete Technical Guide for Mail, Calendar, Contacts and Drive

    A step-by-step guide to moving from Google Workspace to Microsoft 365 with Exchange Online's native Gmail migration and Migration Manager, from routing subdomains and service accounts to MX cutover.

    Microsoft 36515 min read