SAA-C03 Design Resilient Architectures Practice Question
An internal worker consumes messages from an Amazon SQS Standard queue. Recently, some messages fail validation in the worker (for example, missing required fields), causing the worker to crash before it can successfully process those messages. Those messages keep getting retried repeatedly, slowing down processing of valid messages. The team wants a resilient mechanism to quarantine bad messages after a limited number of receive attempts. What should they implement?
⚠ Common exam trap
A common mix-up: candidates think increasing the visibility timeout (Option A) solves the retry problem, but it only delays retries without eliminating the root cause, while the DLQ mechanism (Option B) provides a proper quarantine by moving messages after a configurable number of receive attempts.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure a redrive policy with a Dead-Letter Queue (DLQ) and set maxReceiveCount so poison messages are moved to the DLQ after repeated failures.
Amazon SQS supports configuring a redrive policy with a Dead-Letter Queue (DLQ) that automatically moves messages after a specified number of receive attempts (maxReceiveCount). This isolates poison messages that fail validation and cause crashes, preventing them from being retried indefinitely and slowing down valid message processing. The worker can then focus on valid messages while the DLQ stores the problematic ones for later analysis or manual intervention.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the SQS visibility timeout to several hours so the worker does not retry too quickly.
Why it's wrong here
Increasing the SQS visibility timeout to several hours only changes how long a message stays hidden from other consumers after a worker receives it. When the worker crashes or fails to delete the message, it becomes visible again after the timeout expires, so the same poison message is repeatedly redelivered and reprocessed, just at a slower rate. The visibility timeout has no counter or threshold to permanently quarantine a message after repeated failures; it merely delays the inevitable retry loop.
- ✓
Configure a redrive policy with a Dead-Letter Queue (DLQ) and set maxReceiveCount so poison messages are moved to the DLQ after repeated failures.
Why this is correct
An SQS DLQ with a redrive policy is specifically designed for poison-message handling. When a message exceeds maxReceiveCount without successful processing (for example, the worker crashes before deletion), SQS moves the message to the DLQ. This quarantines bad messages and protects throughput for valid messages.
- ✗
Switch the queue to an SNS topic and subscribe the worker directly, eliminating message retries.
Why it's wrong here
SNS-to-SQS delivery still needs application-level failure handling, and SNS does not provide the same poison-message quarantine mechanism as SQS DLQs with maxReceiveCount. The underlying issue (failed processing and retry/redelivery) would still need a strategy to isolate bad messages.
- ✗
Enable KMS encryption with a new CMK to ensure validation errors stop occurring.
Why it's wrong here
Enabling KMS encryption with a new customer master key (CMK) adds server-side encryption for message bodies at rest, which protects data confidentiality and controls access to the encryption key. However, encryption does not modify the message content or the worker's business logic, so any validation error caused by malformed data or a bug in the consumer will still occur exactly the same way. KMS is about protecting data in transit and at rest, not about fixing application-level processing failures or stopping retry loops.
Go deeper
Related to this question
About these practice questions
Courseiva writes every SAA-C03 question from scratch — 935 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.