SOA-C02 Reliability and Business Continuity Practice Question
A SysOps administrator needs to ensure that an EC2 instance automatically recovers from an underlying hardware failure. Which action should be taken?
⚠ Common exam trap
A common mix-up: candidates confuse Auto Scaling recovery (which replaces the instance) with CloudWatch alarm recovery (which recovers the same instance), leading them to choose Option D despite the requirement to preserve the original instance.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create a CloudWatch alarm on the StatusCheckFailed metric and configure the recovery action.
A CloudWatch alarm on the StatusCheckFailed metric can be configured with the 'recover' action to automatically restart the EC2 instance on a new underlying host if a hardware failure is detected. This recovery action preserves the instance ID, private IP, Elastic IP, and instance metadata, ensuring minimal disruption. The StatusCheckFailed metric specifically monitors the instance's system status checks, which detect AWS hardware issues.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Launch a second instance in a different Availability Zone.
Why it's wrong here
Launching a second instance in a different Availability Zone provides failover capacity, but it does not recover the original EC2 instance. The failed instance remains in a failed state, and the replacement instance receives a new instance ID, private IP, and Elastic IP attachment, so applications referencing the original instance's identity or IP address will break. This approach is a manual or orchestrated disaster-recovery pattern, not the automated system-status-check recovery that CloudWatch can trigger on the existing instance.
- ✗
Assign an Elastic IP address to the instance.
Why it's wrong here
Assigning an Elastic IP address to the instance only guarantees a static public IP; it has no effect on the recovery or health of the underlying EC2 instance. When an instance fails a system status check due to hardware or host issues, the Elastic IP remains mapped to the unhealthy instance, and there is no automatic action to move it to a healthy host or restart the instance. Manual remapping of the Elastic IP could redirect traffic, but it does not repair or restart the original failed instance.
- ✓
Create a CloudWatch alarm on the StatusCheckFailed metric and configure the recovery action.
Why this is correct
Creating a CloudWatch alarm on the StatusCheckFailed metric and configuring the recovery action is the correct method because EC2 instance recovery automatically restarts the instance on new hardware when the underlying host fails. This recovery action preserves the instance ID, private IP address, Elastic IP address, and all EBS volumes, so the instance's identity and configuration are maintained. The alarm must monitor the StatusCheckFailed_System metric (or the aggregate StatusCheckFailed metric) and invoke the 'recover' action to trigger the automatic recovery.
- ✗
Place the instance in an Auto Scaling group with a min size of 1.
Why it's wrong here
An Auto Scaling group with a minimum size of 1 will detect an unhealthy instance and replace it, but replacement means terminating the old instance and launching a brand new one. The new instance will have a different instance ID, a different private IP address (unless a separate Elastic IP is attached), and any instance-store data will be lost, so the original instance is not recovered. Auto Scaling is designed for elasticity and capacity management, not for preserving the state or identity of a specific failed EC2 instance.
Go deeper
Related to this question
About these practice questions
This SOA-C02 question is part of Courseiva's 1,169-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on SOA-C02
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A SysOps administrator needs to ensure that an EC2 instance automatically recovers from an underlying hardware failure. Which configuration should be used?
easy- A.Use AWS Lambda to periodically check instance health and reboot if necessary.
- B.Enable termination protection on the instance.
- ✓ C.Create a CloudWatch alarm on the StatusCheckFailed metric and configure the recovery action.
- D.Place the instance in an Auto Scaling group with a minimum size of 1.
Why C: A CloudWatch alarm on the StatusCheckFailed metric can be configured with an EC2 recovery action. When the alarm triggers (e.g., due to an underlying hardware failure), the recovery action automatically stops and starts the instance on healthy hardware, preserving the instance ID, private IP, Elastic IP, and EBS attachments. This is the native AWS mechanism for automatic instance recovery from hardware failures.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This SOA-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SOA-C02 exam.