SOA-C02 Reliability and Business Continuity Practice Question
A company runs a stateless web application on EC2 instances in an Auto Scaling group. The application is deployed across multiple Availability Zones. The SysOps administrator wants to ensure that the application remains available even if an entire Availability Zone fails. What is the MOST effective way to achieve this?
⚠ Common exam trap
A common mix-up: candidates confuse instance-level recovery mechanisms (like rebooting or replacing a single failed instance) with AZ-level fault tolerance, leading them to choose options that address individual instance health rather than the architectural redundancy required for AZ failure scenarios.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure the Auto Scaling group to launch instances in at least two Availability Zones.
By configuring the Auto Scaling group to launch instances in at least two Availability Zones, the application can survive the failure of an entire AZ because the Auto Scaling group will automatically replace failed instances in the remaining healthy AZs. This design ensures that the stateless web application remains available as long as at least one AZ is operational, leveraging the fault isolation that AWS Availability Zones provide. The Auto Scaling group distributes instances across the specified AZs and will maintain the desired capacity even if one AZ becomes completely unavailable.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Configure the Auto Scaling group to launch instances in at least two Availability Zones.
Why this is correct
Configuring the Auto Scaling group to span at least two Availability Zones distributes your stateless web application's instances across independent failure domains. If an entire AZ becomes unavailable—due to power loss, cooling failure, or networking issues—the remaining healthy AZs continue to serve traffic, and the Auto Scaling group automatically replaces failed instances in the healthy zones. This is the core mechanism for high availability, because it removes any single data-center-level dependency and lets a load balancer route only to surviving instances.
- ✗
Create a CloudWatch alarm to reboot instances when they become unhealthy.
Why it's wrong here
A CloudWatch alarm configured to reboot an instance when it becomes unhealthy can only address problems tied to that specific instance's operating system or software state, such as memory leaks or a hung kernel. It cannot help if the underlying physical host fails or if the entire Availability Zone becomes unavailable, because a reboot in a failed zone will not restore connectivity or service. Furthermore, a reboot is a disruptive, slow recovery action compared to launching a fresh replacement instance, and it provides no protection against correlated failures that affect many instances at once.
- ✗
Use a single Availability Zone to reduce complexity.
Why it's wrong here
Using a single Availability Zone may simplify network topology and reduce the number of subnets, but it introduces a catastrophic single point of failure at the data-center level. If that one AZ suffers an outage, every instance hosting the web application becomes unreachable simultaneously, leaving the stateless app completely offline. For stateless workloads, the added operational complexity of a second AZ is trivial compared to the risk of losing the entire application, so this approach directly contradicts the goal of fault tolerance.
- ✗
Use a larger instance type to handle more traffic.
Why it's wrong here
Choosing a larger instance type, known as vertical scaling, increases the compute capacity of a single server but does nothing to protect against the failure of that instance or its Availability Zone. If the instance crashes or the AZ it resides in goes down, all traffic is lost regardless of instance size—there is no automatic failover or redundancy. Larger instances also increase the blast radius because more users are impacted by a single failure, and they typically cost more per hour without improving availability, unlike horizontal scaling across multiple smaller instances in multiple AZs.
Go deeper
Related to this question
About these practice questions
This SOA-C02 question is part of Courseiva's 1,169-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This SOA-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SOA-C02 exam.