Courseiva

SAA-C03 Design Resilient Architectures Practice Question

A media company runs a video-transcoding fleet on Amazon EC2 instances that read source files from an Amazon S3 bucket and write output to a second bucket. The fleet is spread across three Availability Zones in one Region, and instances are launched by an Auto Scaling group. The company needs the architecture to survive the loss of an entire Availability Zone without losing in-flight transcoding work or requiring manual intervention. Which combination of design elements should a solutions architect implement to meet these requirements?

⚠ Common exam trap

The trap here is assuming that storing source and output objects in Amazon S3 is sufficient for workload resilience, when S3 durability says nothing about resuming in-flight compute work after a zone failure.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Deploy the Auto Scaling group across three Availability Zones, make transcoding jobs idempotent and store progress in Amazon DynamoDB, and have instances poll an Amazon SQS queue for work so that unfinished jobs are retried by healthy instances.

Resilience across an Availability Zone failure requires both compute capacity in the surviving zones and durable, externalized job state. Distributing the Auto Scaling group across three zones keeps instances running, while an SQS queue with visibility timeouts and DynamoDB progress tracking lets interrupted jobs be reclaimed and retried automatically. This removes any dependency on the failed zone's instances or storage.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Deploy the Auto Scaling group across three Availability Zones, make transcoding jobs idempotent and store progress in Amazon DynamoDB, and have instances poll an Amazon SQS queue for work so that unfinished jobs are retried by healthy instances.

    Why this is correct

    Spreading the Auto Scaling group across three Availability Zones means instances in surviving zones continue running when one zone fails. Because work is pulled from an SQS queue and progress is tracked in DynamoDB, a job interrupted in the failed zone becomes visible again after its visibility timeout and is retried by a healthy instance, so no manual intervention is required.

  • ✗

    Deploy the Auto Scaling group in a single Availability Zone with a spot fleet, use an Amazon EBS volume attached to each instance to persist transcoding progress, and enable EBS snapshots every five minutes to another zone.

    Why it's wrong here

    A single-zone Auto Scaling group is itself the failure domain, so losing that Availability Zone stops all transcoding capacity. EBS volumes are tied to one Availability Zone and snapshots only let you create new volumes there; they cannot transparently move an instance or resume an interrupted job in another zone without manual recovery steps.

  • ✗

    Deploy the Auto Scaling group across three Availability Zones, place a Network Load Balancer in front of the instances, and configure the load balancer to retry failed transcoding requests against the same instance until the zone recovers.

    Why it's wrong here

    A Network Load Balancer distributes incoming connections but has no visibility into which Availability Zone an instance sits in or whether a transcoding job is complete. Retrying against an instance in the failed zone cannot succeed, and the fleet still needs a durable queue and state store to resume work without operator intervention.

  • ✗

    Deploy the Auto Scaling group across three Availability Zones, store transcoding state in an Amazon S3 bucket configured with S3 Cross-Region Replication, and rely on the S3 Standard storage class for automatic recovery.

    Why it's wrong here

    Cross-Region Replication copies objects to a different Region but does not provide any mechanism to detect a failed Availability Zone or move the transcoding fleet. S3 Standard already replicates data within a Region across multiple Availability Zones, so replication adds latency and cost without restoring compute capacity in the surviving zones.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This SAA-C03 question is part of Courseiva's 935-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.