← Cloud projects

Cloud · Architecture Concept

High Availability Design Patterns

Design workloads to survive component and zone failures by making failure domains visible across compute, network, data, identity, and operational dependencies.

PlatformCloud
DomainResilience
LevelFoundation
Last reviewed2026-09-19
Use whenDesign applications to tolerate component or zone failures without requiring disaster-recovery procedures for every incident.
Key decisionDesign for failure domains
Key conceptsFailure domains · Redundancy · Health checks · Traffic distribution
Cost focusduplicate capacity
On this page

Building blocks

Failure domainsRedundancyHealth checksTraffic distributionStateless scalingResilient stateFailover testing

Design goal

Remove single points of failure for incidents that should be handled by normal availability mechanisms rather than invoking a disaster-recovery process.

Architecture

Workloads use redundant components, health-based traffic distribution, stateless scaling where possible, and resilient data and services aligned to failure domains.

Architecture decisions

Decision

Design for failure domains

Why

Redundancy must cross the failure boundary you are trying to survive.

Trade-off

Spreading components across the intended failure boundary improves availability, but adds duplicate capacity and more distributed operating paths.

Decision

Remove hidden single points of failure

Why

Identity, DNS, secrets, storage, monitoring, and administration can be as critical as application compute.

Trade-off

Addressing non-compute dependencies improves real availability, but expands the design and testing scope beyond the obvious application tier.

Security

Protect the control and data paths deliberately. HA components should preserve the same security controls on all active and standby paths.

Availability

Design for the failure domain that must be survived. Use redundant instances and managed zone-resilient capabilities where available; test failure rather than assuming redundancy works.

Disaster recovery

Treat regional recovery as a separate operating state. HA addresses local or component failures; use separate DR design for region-scale or catastrophic events.

Cost drivers

  • duplicate capacity
  • load balancing
  • zone-aware data services
  • observability

Design assumptions

  • Failure scenarios are identified

Implementation plan

  1. Map application dependencies and identify the actual failure domains—process, instance, host, zone, network, data service, and region.
  2. Define which failures the workload must survive without recovery intervention; do not design every component for the maximum possible availability.
  3. Apply redundancy to each critical tier and remove shared dependencies that would defeat the intended failure-domain separation.
  4. Configure health checks and traffic failover around end-to-end application health rather than host reachability alone.
  5. Test component and failure-domain loss, measure user impact and recovery behavior, and feed the results back into the design.

Validate the design

  • Stop or disable one component in each intended failure domain and observe whether the service remains available.
  • Verify traffic is redistributed only to healthy capacity and that health checks detect application failure, not just host reachability.
  • Confirm alternate paths do not bypass authentication, inspection, logging, or other security controls.
  • Measure user impact and recovery behavior, then compare the result with the availability objective.

Architecture basis

Continue a learning path