Building blocks
Design goal
Protect data from deletion, corruption, ransomware, and infrastructure failure with backup policies that reflect workload criticality instead of one global retention rule.
Success criteria
- Define protection by business criticality
- Separate backup administration from workload administration
- Maintain recoverable copies outside the primary failure domain
- Test recovery regularly
Architecture
Classify workloads, protect them with tiered policies, maintain isolated recovery copies, and continuously validate restorability.
Architecture flow
- Classify data and workloads
- Assign RPO/RTO and retention
- Create primary backups
- Maintain isolated or immutable recovery copy
- Run scheduled restore tests
Architecture decisions
Recovery objectives drive policy
WhyFrequency and retention should be derived from business RPO/RTO rather than applying one global policy.
Trade-offTiered policies align protection with business value, but require stakeholders to agree on recovery objectives instead of defaulting to one retention policy.
Backups are not DR by themselves
WhyBackup protects data; business continuity additionally requires compute, networking, identity, runbooks, and recovery testing.
Trade-offSeparating data protection from service recovery prevents false confidence, but broadens the recovery design to include compute, identity, networking, and runbooks.
Security
Protect the control and data paths deliberately. Separate backup privileges, protect backup configuration with strong identity controls, and use immutability or deletion protection where available.
- Least-privilege backup administration
- MFA for privileged operations
- Encryption
- Isolated recovery copies
- Restore audit trail
Cost drivers
- Protected data volume
- Backup frequency
- Retention duration
- Cross-region or cross-account copy
- Restore/test frequency
Design assumptions
- Workloads have documented owners and criticality
- Recovery targets are approved by the business
Implementation plan
- Classify workloads and data by business impact, then assign RPO, RTO, retention, and regulatory requirements to each recovery tier.
- Choose backup methods that are application-consistent where the workload requires it rather than relying on crash-consistent snapshots everywhere.
- Keep at least one recovery copy isolated from the primary administrative and failure boundary when ransomware or account compromise is in scope.
- Protect backup configuration and deletion with separate roles, immutability/retention controls, encryption, and monitored access.
- Schedule restore tests by recovery tier and record actual restore time, data usability, dependencies, and remediation actions.
Validate the design
- Restore representative workloads from each recovery tier on the documented schedule.
- Record achieved restore point and recovery time from each test and compare them with the agreed RPO/RTO.
- Verify the isolated/immutable copy remains accessible when primary workload credentials are assumed compromised.
- Track failed backups, failed restores, and remediation actions to closure.