An organization is designing a cloud application that must remain available even if an entire AWS availability zone fails. Which architecture pattern should they implement?
Running active-active across multiple availability zones within one region lets traffic fail over instantly when an entire AZ becomes unavailable, because each AZ has independent power, cooling and networking. This satisfies the requirement to survive a full AZ failure without regional latency penalties.
Why this answer
The correct architecture is a single region with multiple Availability Zones (AZs) in an active-active configuration. This ensures that if one AZ fails, the application continues to serve traffic from the remaining AZs without any manual intervention, as all AZs are actively handling requests. AWS Availability Zones are physically separate data centers within a region, and an active-active pattern distributes the workload across them to achieve high availability and fault tolerance.
Exam trap
ISC2 often tests the distinction between surviving an AZ failure versus a region failure, and the trap here is that candidates may overcomplicate the solution by choosing multi-region active-active, not realizing that a single region with multiple AZs is sufficient and more cost-effective for the given requirement.
How to eliminate wrong answers
Option B (Single region with multiple AZs active-standby) is wrong because it introduces a standby component that is not actively serving traffic, leading to potential downtime during failover and resource underutilization; the question requires continuous availability even during an AZ failure, which active-standby does not guarantee without a failover delay. Option C (Active-passive in a single region) is wrong because it typically relies on a single AZ for the active component, making it vulnerable to AZ failure, and the passive component requires manual or automated failover, which introduces downtime. Option D (Multi-region active-active) is wrong because while it provides high availability, it is over-engineered for the requirement of surviving a single AZ failure; it adds unnecessary complexity, latency, and cost, and the question specifically asks for an architecture that remains available if an entire AWS availability zone fails, not a full region failure.