Courseiva

SOA-C02 Reliability and Business Continuity Practice Question

A SysOps administrator needs to ensure that an EC2 instance automatically recovers from an underlying hardware failure. Which configuration should be used?

⚠ Common exam trap

Many exam-takers confuse termination protection (which only prevents deletion) or Auto Scaling replacement (which creates a new instance) with the native recovery action that preserves the instance's identity and state.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Create a CloudWatch alarm on the StatusCheckFailed metric and configure the recovery action.

A CloudWatch alarm on the StatusCheckFailed metric can be configured with an EC2 recovery action. When the alarm triggers (e.g., due to an underlying hardware failure), the recovery action automatically stops and starts the instance on healthy hardware, preserving the instance ID, private IP, Elastic IP, and EBS attachments. This is the native AWS mechanism for automatic instance recovery from hardware failures.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use AWS Lambda to periodically check instance health and reboot if necessary.

    Why it's wrong here

    Lambda’s periodic health checks can detect an unhealthy instance and trigger a reboot, but the scenario demands recovery from underlying hardware failure, not a software-level restart. A reboot does not migrate the instance to new physical hardware; it merely restarts the operating system on the same host. This option is tempting because Lambda is commonly used for automated instance monitoring and remediation, and would be correct for scenarios requiring a simple restart after an application crash or OS hang.

  • ✗

    Enable termination protection on the instance.

    Why it's wrong here

    Termination protection is an EC2 attribute that blocks accidental API, console, or CLI termination of the instance, but it does not monitor the health of the underlying physical host or trigger any automatic remediation. When hardware fails, the instance becomes impaired and AWS may eventually stop or terminate it, even if termination protection is enabled, because the protected flag only stops user-initiated requests. It therefore guards against human error, not against hardware failure, and cannot keep the workload running or migrate it to new infrastructure.

  • ✓

    Create a CloudWatch alarm on the StatusCheckFailed metric and configure the recovery action.

    Why this is correct

    A CloudWatch alarm on the StatusCheckFailed metric can be configured with the EC2 recovery action, which automatically stops and starts the instance on a different physical host when the underlying hardware or network is impaired. This recovery process preserves the instance ID, private IP address, Elastic IP, and instance metadata, and the EBS root volume remains attached with its existing data. Because the action moves the instance to healthy hardware while maintaining its identity and configuration, it is the correct mechanism to recover from a physical host failure without requiring manual intervention.

  • ✗

    Place the instance in an Auto Scaling group with a minimum size of 1.

    Why it's wrong here

    An Auto Scaling group with a minimum size of 1 treats an unhealthy instance as a candidate for replacement: it terminates the old instance and launches a brand-new instance from the launch template or configuration. This new instance receives a different instance ID and typically a new private IP; even if an Elastic IP is associated, it is not automatically transferred to the replacement. Additionally, the original instance's local instance-store volumes are lost and the new instance starts from a fresh EBS root volume, so it cannot preserve the original instance's identity, attached data, or custom metadata, making ASG replacement unsuitable for a scenario that requires recovering the same instance.

About these practice questions

This SOA-C02 question is part of Courseiva's 1,169-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This SOA-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SOA-C02 exam.