SOA-C02 Monitoring, Logging, and Remediation Practice Question
A SysOps administrator receives an alarm that an EC2 instance's status check has failed. The instance is part of an Auto Scaling group behind an Application Load Balancer. The administrator needs to ensure that the instance is automatically replaced and that the root cause is investigated. What is the MOST efficient combination of actions to achieve this?
⚠ Common exam trap
It's easy for candidates to think manual actions (reboot, stop/start) are sufficient for recovery, but the question explicitly requires automatic replacement and root cause investigation, which only a lifecycle hook with data capture provides.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure an Auto Scaling lifecycle hook to terminate the unhealthy instance and send the instance system log to an S3 bucket for analysis.
It combines automatic instance replacement via the Auto Scaling group's health check (which marks the instance unhealthy and terminates it) with a lifecycle hook that captures the instance's system log before termination and sends it to S3 for root cause analysis. This is the most efficient approach as it requires no manual intervention and preserves diagnostic data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Configure an Auto Scaling lifecycle hook to terminate the unhealthy instance and send the instance system log to an S3 bucket for analysis.
Why this is correct
A lifecycle hook on the terminating state of an Auto Scaling group intercepts the EC2 instance termination process, enabling a Lambda function to capture the instance's system log (console output) and upload it to an S3 bucket before the instance is destroyed. Because the Auto Scaling group has already marked the instance unhealthy via its health checks, the group will automatically launch a replacement instance after the lifecycle action completes, ensuring fully automated recovery. The hook's timeout and the ability to call complete-lifecycle-action guarantee that the log is securely stored, and the system log provides the diagnostic evidence needed for root cause analysis.
- ✗
Create a CloudWatch alarm that triggers an SNS notification to the administrator to manually replace the instance.
Why it's wrong here
This solution is notification-only: a CloudWatch alarm publishes to an SNS topic, but the actual recovery depends on a human receiving the message, logging into the console, and manually terminating or replacing the failing instance. There is no automated step to initiate a new instance, so the system remains impaired during the entire manual turnaround. Moreover, CloudWatch alarms cannot directly execute EC2 instance replacement actions, so this approach neither scales nor satisfies the requirement for automatic replacement.
- ✗
Reboot the instance from the AWS Management Console and then review CloudTrail logs.
Why it's wrong here
Rebooting from the console is a manual, one-off action that does not provide a durable fix if the instance is trapped in a failing state, and it is not integrated with any automated health-check replacement workflow. Additionally, CloudTrail only records control-plane API calls made to AWS services; it does not capture the instance's Operating System logs, kernel messages, or console output. To retrieve those diagnostics, you would need to use the EC2 console output (get-console-output) or an agent, and the reboot itself might destroy the very evidence you need if the instance fails to boot.
- ✗
Manually stop and start the instance to recover it, then check the system logs.
Why it's wrong here
A manual stop and start is performed by an operator, and while it may migrate the instance to a new underlying host, it does not automate the recovery process or scale to a fleet of instances. Capturing system logs after a stop/start is complicated because the instance's console output is cleared once it reboots, and you may lose the logs from the failed boot. The entire workflow requires human intervention, so it is not an acceptable automated solution for replacing an unhealthy instance.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every SOA-C02 question from scratch — 1,169 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on SOA-C02
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A SysOps admin notices that an EC2 instance's status check fails intermittently. The instance is part of an Auto Scaling group. What is the most appropriate first step to diagnose the issue?
medium- A.Stop and start the instance
- B.Terminate the instance and let Auto Scaling replace it
- C.Reboot the instance
- ✓ D.Review the instance's status check history in the EC2 console
Why D: The most appropriate first step is to review the instance's status check history in the EC2 console (Option D). This allows the SysOps admin to determine whether the failures are due to system status checks (e.g., underlying hardware issues) or instance status checks (e.g., OS-level problems). Since the instance is part of an Auto Scaling group, understanding the root cause is critical before taking any corrective action, as premature termination or reboot could mask the issue or lead to unnecessary replacements.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This SOA-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SOA-C02 exam.