Building blocks
Design goal
Meet server-recovery objectives without running a full duplicate environment continuously, and prove that applications can actually start and reconnect during a recovery event.
Success criteria
- Continuously replicate source servers
- Use low-footprint staging resources
- Run non-disruptive recovery drills
- Document traffic redirection and failback
Architecture
Replication agents send block-level changes into a staging area; during drills or recovery, DRS launches converted EC2 recovery instances in the target VPC.
Architecture flow
- Install replication agent
- Replicate data into staging area
- Monitor replication health
- Launch drill/recovery instances
- Validate applications and redirect traffic
- Fail back when primary environment is ready
Architecture decisions
Recovery testing is part of the service design
WhyA configured replication pipeline is not enough; periodic drills validate launch settings, networking, dependencies, and operational runbooks.
Trade-offRegular drills consume time and temporary recovery capacity, but without them launch settings and runbooks remain assumptions.
Traffic redirection remains an application/network responsibility
WhyRecovery instances must be integrated with DNS, load balancing, identity, and client connectivity before the service is restored.
Trade-offDRS can launch recovery compute, but DNS, load balancing, identity, certificates, and client routing still need separate recovery procedures.
Networking
Make the traffic path explicit. Predefine target subnets, security groups, routes, DNS, and connectivity required by recovered applications.
Security
Protect the control and data paths deliberately. Protect replication agents, DRS permissions, staging resources, and recovery actions with least privilege and monitoring.
Availability
Design for the failure domain that must be survived. DR protects against broader outages; application-level multi-AZ design should still be used where high availability is required during normal operation.
Disaster recovery
Treat regional recovery as a separate operating state. Use drills and recovery plans to coordinate interdependent server groups, then perform explicit traffic redirection and failback.
Cost drivers
- Replication staging storage
- Replication servers
- Drill/recovery EC2 runtime
- Network transfer
- Recovery-region shared services
Design assumptions
- Source operating systems/workloads are supported
- Target VPC and dependencies are prepared before an event
Implementation plan
- Prepare the recovery Region with the staging subnet, security controls, replication connectivity, and quotas required by Elastic Disaster Recovery.
- Install replication agents on supported source servers and confirm replication lag is inside the required recovery-point objective.
- Define launch settings for instance type, disks, networking, security groups, IAM, and post-launch actions rather than accepting defaults blindly.
- Create recovery sequencing around application dependencies such as identity, DNS, databases, middleware, and front-end services.
- Run non-disruptive recovery drills, measure actual RPO/RTO, and document the production failover and failback procedure.
Validate the design
- Confirm replication lag remains inside the target RPO during representative write activity.
- Launch a recovery drill into the isolated recovery network and verify every required application dependency starts in sequence.
- Measure achieved RPO and RTO from the drill rather than relying on configured targets.
- Validate application data, DNS, routing, identity, and external connectivity before declaring the recovered service usable.