Courseiva

SAA-C03 Design Resilient Architectures Practice Question

A startup runs a nightly batch job on a single EC2 instance that reads a large dataset from Amazon S3, performs transformations, and writes results back to S3. The job takes about two hours, and the team wants the job to restart automatically if the instance fails or is terminated by AWS. The job is idempotent and can safely resume from the beginning. What is the MOST operationally efficient way to meet this requirement?

⚠ Common exam trap

The trap here is reaching for CloudWatch alarms and Lambda automation when the built-in Auto Scaling group replacement behavior already satisfies the restart requirement.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Place the instance in an Auto Scaling group with a minimum and desired capacity of one, use a launch template that installs and starts the job at boot, and configure a lifecycle hook to keep the instance in service until the job completes.

A single-instance Auto Scaling group is the managed way to keep exactly one healthy instance running and to replace it automatically when it is terminated or fails. Supplying the job through a launch template's user data means every replacement instance starts the job again, and because the job is idempotent, restarting from the beginning is acceptable. No custom alarm or recovery scripting is required.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use EC2 Auto Recovery to recover the instance onto new hardware and attach an Amazon EBS volume that persists the job's progress.

    Why it's wrong here

    EC2 Auto Recovery applies to instances in an Auto Scaling group and recovers them when the underlying host is impaired, but it is not a general replacement mechanism and does not restart application processes automatically. Persisting progress to EBS does not cause the batch job to resume by itself, so this still needs custom logic to detect and restart the work.

  • ✓

    Place the instance in an Auto Scaling group with a minimum and desired capacity of one, use a launch template that installs and starts the job at boot, and configure a lifecycle hook to keep the instance in service until the job completes.

    Why this is correct

    An Auto Scaling group with a desired capacity of one continuously maintains a single healthy instance, so if AWS terminates or the instance fails, the group replaces it and the launch template's user data reruns the idempotent job. A lifecycle hook can delay termination long enough for a clean shutdown, and this is fully managed with no custom monitoring code.

  • ✗

    Create an Amazon CloudWatch alarm on the StatusCheckFailed metric that invokes an AWS Lambda function to call the EC2 RebootInstances API.

    Why it's wrong here

    Rebooting an instance does not help when the underlying host has failed, and AWS-initiated termination or host retirement requires the instance to be replaced rather than rebooted. RebootInstances also does not restart the batch process unless the operating system is configured to do so, so the job would not reliably resume after a hardware failure.

  • ✗

    Convert the job into an AWS Lambda function with a 15-minute timeout and trigger it on a nightly Amazon EventBridge schedule.

    Why it's wrong here

    Lambda functions have a maximum invocation timeout of 15 minutes, and this job runs for roughly two hours, so it would be terminated before finishing. Increasing memory does not extend the timeout limit, and splitting the job into many chained invocations would require substantial re-engineering that the scenario does not ask for.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This SAA-C03 question is part of Courseiva's 935-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.