Building blocks
Design goal
Remove single points of failure for incidents that should be handled by normal availability mechanisms rather than invoking a disaster-recovery process.
Architecture
Workloads use redundant components, health-based traffic distribution, stateless scaling where possible, and resilient data and services aligned to failure domains.
Architecture decisions
Design for failure domains
WhyRedundancy must cross the failure boundary you are trying to survive.
Trade-offSpreading components across the intended failure boundary improves availability, but adds duplicate capacity and more distributed operating paths.
Remove hidden single points of failure
WhyIdentity, DNS, secrets, storage, monitoring, and administration can be as critical as application compute.
Trade-offAddressing non-compute dependencies improves real availability, but expands the design and testing scope beyond the obvious application tier.
Security
Protect the control and data paths deliberately. HA components should preserve the same security controls on all active and standby paths.
Availability
Design for the failure domain that must be survived. Use redundant instances and managed zone-resilient capabilities where available; test failure rather than assuming redundancy works.
Disaster recovery
Treat regional recovery as a separate operating state. HA addresses local or component failures; use separate DR design for region-scale or catastrophic events.
Cost drivers
- duplicate capacity
- load balancing
- zone-aware data services
- observability
Design assumptions
- Failure scenarios are identified
Implementation plan
- Map application dependencies and identify the actual failure domains—process, instance, host, zone, network, data service, and region.
- Define which failures the workload must survive without recovery intervention; do not design every component for the maximum possible availability.
- Apply redundancy to each critical tier and remove shared dependencies that would defeat the intended failure-domain separation.
- Configure health checks and traffic failover around end-to-end application health rather than host reachability alone.
- Test component and failure-domain loss, measure user impact and recovery behavior, and feed the results back into the design.
Validate the design
- Stop or disable one component in each intended failure domain and observe whether the service remains available.
- Verify traffic is redistributed only to healthy capacity and that health checks detect application failure, not just host reachability.
- Confirm alternate paths do not bypass authentication, inspection, logging, or other security controls.
- Measure user impact and recovery behavior, then compare the result with the availability objective.