Building blocks
Design goal
Define recovery outcomes before selecting technology so data-loss tolerance and recovery-time tolerance drive the design instead of being inferred afterward.
Architecture
Business services are classified by impact and assigned recovery objectives. Architecture choices are then mapped to those objectives and validated through recovery testing.
Architecture decisions
Start with business impact
WhyTechnology should follow required recovery outcomes, not the other way around.
Trade-offOutcome-first recovery prevents overengineering, but requires business owners to quantify tolerable downtime and data loss before technology is selected.
Treat RPO and RTO separately
WhyData-loss tolerance and recovery-time tolerance drive different architecture decisions.
Trade-offSeparating data-loss and recovery-time objectives produces better designs, but may require different technologies and costs to meet each target.
Security
Protect the control and data paths deliberately. Recovery access, privileged accounts, encryption keys, and emergency procedures must remain available during a primary-environment outage.
Availability
Design for the failure domain that must be survived. High availability reduces some local failures but does not replace backup or disaster recovery.
Disaster recovery
Treat regional recovery as a separate operating state. Define failover authority, recovery order, dependencies, communication, and failback as part of the design.
Cost drivers
- Secondary capacity
- replication
- backup storage/retention
- cross-region transfer
- recovery testing
Design assumptions
- Business owners can classify workload criticality
Implementation plan
- Run a business-impact assessment to identify the services, data, and business processes whose loss creates material impact.
- Define recovery tiers with explicit RPO and RTO targets, and document who is authorized to accept exceptions or additional cost.
- Map upstream and downstream dependencies so recovery order reflects the complete service rather than individual components.
- Select backup, replication, standby, redeployment, and traffic-recovery patterns that meet each tier without overengineering lower-criticality workloads.
- Document ownership and runbooks, then run tabletop and technical recovery exercises that measure actual RPO/RTO.
Validate the design
- Run a tabletop exercise and confirm decision owners, dependencies, and escalation paths are understood.
- Perform a technical recovery test for each recovery tier rather than validating documentation only.
- Measure actual recovery point and recovery time against the approved RPO/RTO targets.
- Track every gap found during testing to an owner and closure date, then retest material fixes.