Building blocks
Design goal
Recover applications—not just virtual machines—within agreed RPO and RTO by planning identity, DNS, networking, data consistency, recovery order, and failback together.
Success criteria
- Define RPO/RTO by workload tier
- Map recovery networks
- Sequence interdependent servers
- Test without disrupting production
Architecture
Replicate supported VM workloads to a recovery region, maintain network/recovery configuration, and use recovery plans to coordinate application recovery steps.
Architecture flow
- Replicate protected VMs
- Monitor replication health
- Maintain recovery network mappings
- Run test failovers
- Initiate recovery during an event
- Validate application and redirect traffic
Architecture decisions
Application consistency must be workload-aware
WhyVM replication does not replace native application/database protection where transaction consistency or application-specific recovery is required.
Trade-offNative application protection can improve consistency beyond VM replication, but introduces workload-specific tooling and runbooks alongside ASR.
Identity, DNS, and network recovery are first-class dependencies
WhyRecovered VMs are useful only if authentication, name resolution, routing, and application dependencies are available.
Trade-offPreparing shared dependencies improves recoverability, but increases DR scope beyond the replicated virtual machines themselves.
Networking
Make the traffic path explicit. Predefine recovery VNets/subnets, routing, DNS, firewall rules, and traffic redirection. Avoid public IP dependency unless explicitly required.
Security
Protect the control and data paths deliberately. Protect the vault and recovery operations with least-privilege access, MFA, monitoring, and change control.
Availability
Design for the failure domain that must be survived. Use high availability within the primary region separately from regional DR; HA and DR solve different failure scopes.
Disaster recovery
Treat regional recovery as a separate operating state. Define recovery plans, test failover cadence, application validation steps, traffic redirection, and failback procedures.
Cost drivers
- Replication storage
- ASR protected instances
- Secondary-region network/security services
- Test failover compute
- Backup retention if used alongside DR
Design assumptions
- Workloads are supported by Azure Site Recovery
- Application owners define validation steps
- Secondary-region networking is preplanned
Implementation plan
- Classify applications by RPO/RTO and map recovery order, data consistency, identity, DNS, and network dependencies before enabling replication.
- Prepare the recovery Region with address space, routing, security controls, DNS, quotas, and any services that cannot be created during an outage.
- Configure replication policies, cache/storage, target disks, and recovery settings according to workload churn and supported ASR limits.
- Enable replication in dependency-aware groups and build recovery plans for the sequence and automation required to restore the application.
- Run test failovers in isolated networks, measure achieved RPO/RTO, and document production failover, validation, failback, and re-protection.
Validate the design
- Run an isolated test failover and measure achieved RPO and RTO for the complete application sequence.
- Verify recovery-region DNS, routing, security rules, and dependent services before application validation.
- Confirm application and database consistency after failover rather than validating VM boot alone.
- Review and rerun the recovery plan after major architecture or dependency changes, including failback steps.