Executive Summary & Architecture Takeaways
- Example Scope: This playbook assumes a representative estate of roughly two hundred VMware vSphere virtual machines, around fifteen business-critical multi-tier applications, and tens of terabytes of transactional data in a regulated industry. Scale the numbers to your own estate.
- The Execution Constraint: A maximum cutover maintenance window of 4 hours on a Sunday morning, with zero business transaction loss permitted.
- Key Architectural Decisions:
- Build Azure landing zones (CAF) with Terraform before touching a single workload.
- Provision a dedicated Azure ExpressRoute circuit sized for replication traffic, with a site-to-site VPN as fallback.
- Use continuous block-level replication (Azure Migrate server migration or Azure Site Recovery) for application tiers and SQL Server Always On / PostgreSQL streaming replication for database tiers.
- Define rollback gates and a hard go/no-go time before the window starts.
1. Enterprise-Scale Landing Zone Architecture
Before replicating any servers, establish a governance, network, and security foundation.
+-----------------------------------+
| Enterprise Management Group |
+-----------------+-----------------+
|
+-----------------------+-----------------------+
| |
+---------v---------+ +---------v---------+
| Platform MG | | Workloads MG |
+----+----+----+----+ +----+----+----+----+
| | | | | |
+-------+ | +-------+ | | |
| | | | | |
+-----v----+ +-----v----+ +-----v----+ +-----v----+ +--v-------+
|Management| |Connectiv.| | Identity | |Corp Spoke| |Online Spk|
|Subscript.| |Subscript.| |Subscript.| |Subscript.| |Subscript.|
+----------+ +-----+----+ +----------+ +-----+----+ +----------+
| |
| Peer: Spoke-to-Hub Peering |
+==========================================+
|
+-------------v-------------+
| ExpressRoute Gateway | <==== Dedicated Circuit ====> On-Premises
| Azure Firewall Premium | Datacenter
+---------------------------+1.1 Network Segmentation Rules
- Hub-Spoke VNet Topology: The Connectivity subscription hosts the central hub VNet. All on-premises traffic terminates on the ExpressRoute gateway.
- Routing via Azure Firewall: Inter-spoke and internet-bound traffic is sent through the central Azure Firewall using user-defined routes (
0.0.0.0/0-> next hop: the firewall's private IP, for example10.0.0.4). - No Public IP Rule: No compute workload in the Corp landing zone gets a public IP address (enforce it with Azure Policy). Administrative access goes through Azure Bastion, and PaaS services are reached via Private Endpoints.
2. Data Replication Architecture: Minimizing Cutover Delta
The critical mistake in cloud migrations is attempting to copy disk images over the network during the cutover window. As a rough calculation, 45 TB over a fully utilized 1 Gbps link takes more than four days, which no maintenance window can absorb.
2.1 The Two-Tier Continuous Sync Model
[On-Premises VMware Datacenter] [Azure Target Region]
| |
+-- Application VMs +-- Replica Managed Disks
| Continuous replication (Azure Migrate / ASR) ----> Block-level deltas applied
| | continuously
| |
+-- Database Tier +-- Azure IaaS DB / Managed Instance
SQL Always On / Postgres Replication ------------> Continuous log streaming
(Transaction Log / WAL) | (lag monitored to zero at cutover)- Application Servers (Stateless / File Stores):
- Replicated with the Azure Migrate server migration tool (Microsoft's recommended path for VMware to Azure) or Azure Site Recovery (ASR).
- Initial replication runs weeks before cutover while systems are live.
- Changed blocks are replicated continuously, so the replica disks stay within minutes of the source.
- Database Servers (Stateful Transaction Engines):
- Exclude active database volumes from block-level replication, because a crash-consistent copy of a busy database is not something you want to promote.
- Configure native replication instead:
- SQL Server: Always On availability groups (or distributed availability groups) extending across ExpressRoute to SQL Server on Azure VMs, or the Managed Instance link for Azure SQL Managed Instance.
- PostgreSQL / MySQL: streaming replication to target standby instances on Azure.
3. A Minute-by-Minute Cutover Runbook
The sample schedule below fits a Sunday 01:00 to 05:00 window. Rehearse it at least once end to end against a non-production copy before the real weekend.
| Time (UTC) | Action Item | Owner | Verification Gate |
|---|---|---|---|
| 01:00 | Commence Cutover Window: Place on-premises frontends in maintenance mode | Release Lead | Public status page updated; maintenance banner active. |
| 01:10 | Stop background queue processors (message consumers, cron jobs) | App Ops | Message queues drained to 0 pending tasks. |
| 01:25 | Validate final database replication sync | Lead DBA | SQL Server log send/redo queue = 0; Postgres replay lag = 0 bytes. |
| 01:30 | Stop replication; promote Azure database instances to primary | Lead DBA | Write test verified on the Azure primary. |
| 01:45 | Migrate / fail over application VMs with source VMs shut down first | Cloud Architect | Target VMs booted; expected private IPs and DNS registration confirmed. |
| 02:15 | Run automated smoke-test suites (Selenium / Cypress) | QA Lead | Core login, checkout and ledger balance tests pass. |
| 02:30 | Go/no-go decision | Migration Lead | All gates green, or rollback starts now. |
| 02:40 | Update public and internal DNS records | Network Lead | DNS records point to Azure Application Gateway / load balancer front ends. |
| 03:00 | End-to-end business integration validation | Business Lead | Test transactions executed in production and confirmed by finance. |
| 03:20 | Maintenance Mode Disabled: Service live on Azure | Release Lead | Public traffic entering Azure; monitoring green. |
4. Rollback Triggers & Safety Gates
A successful migration requires pre-defined, non-negotiable abort criteria. If any of the following occur before the 02:30 go/no-go point, the cutover is aborted and traffic remains on-premises:
- Database Promotion Failure: The Azure database does not accept writes within 15 minutes of the promotion attempt.
- Smoke Test Failure: More than 2 core business flows fail automated regression tests.
- Network Packet Loss: Sustained packet loss above 0.5% on the ExpressRoute private peering during application start-up.
Rollback Procedure: Re-enable on-premises frontend services, point DNS back (TTL lowered to 60 seconds at least 48 hours in advance), and resume background workers. Because no user writes have reached Azure before the go/no-go point, the on-premises databases are still authoritative. After go-live, rolling back requires reverse replication or accepting data loss, which is why the gate sits before DNS is switched. Aim to restore the on-premises baseline in well under 30 minutes, and time it during rehearsal.
5. Post-Migration Optimization: Reducing Cloud Waste
Lift-and-shift sizing is usually conservative. Plan a right-sizing and optimization pass in the first 60 days after migration:
- Azure Hybrid Benefit (AHB): Apply existing Windows Server and SQL Server licenses with Software Assurance (or qualifying subscriptions) to Azure VMs to remove the license component from compute pricing.
- Savings Plans & Reserved Instances: Once utilization data is stable, buy 1- or 3-year commitments for baseline workloads. Model the discount for your SKUs and region with the Azure pricing calculator rather than relying on headline percentages.
- Auto-Shutdown on Dev/Test: Use Azure Automation runbooks or the built-in VM auto-shutdown to stop non-production environments outside business hours.
- Right-Sizing: Use Azure Advisor and Azure Monitor metrics to downsize over-provisioned VMs and move suitable disks to cheaper tiers.
- Track the Business Case: Compare actual run cost against the on-premises baseline (hardware leases, power and cooling, virtualization licensing) quarterly, so savings are measured rather than assumed.