Courseiva

SAA-C03 Design Resilient Architectures Practice Question

An order processing workflow uses Amazon SQS as the decoupling layer between a producer and a consumer Lambda function. The consumer intermittently fails due to a downstream dependency. The team has observed that certain “poison” messages keep being retried repeatedly and prevent other messages from being processed efficiently. Which SQS configuration most directly addresses this issue?

⚠ Common exam trap

Many candidates think increasing visibility timeout or switching to FIFO alone will handle failed messages, but without a DLQ, poison messages remain in the queue and continue to block other messages, which is the core issue described.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure a redrive policy with a dead-letter queue (DLQ) and set an appropriate visibility timeout greater than the maximum processing time.

Configuring a redrive policy with a dead-letter queue (DLQ) allows messages that repeatedly fail processing to be moved out of the main queue after a specified number of receive attempts. Setting an appropriate visibility timeout greater than the maximum processing time ensures that messages are not made visible again before the consumer finishes processing, preventing premature retries. This directly isolates poison messages so they no longer block the processing of other messages in the queue.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set the SQS queue’s retention period to 10 years and rely on application retries to eventually succeed.

    Why it's wrong here

    SQS retention period defines the maximum length of time a message can remain in the queue, from 1 minute to 14 days; a value of 10 years becomes invalid as SQS will reject it, and even the maximum 14 days merely keeps messages available rather than protecting the workflow from repeat failures. Relying on application retries alone means the same poison message will be attempted again and again, wasting compute and potentially growing fetch/processing costs without advancing past the fault. Retention controls storage duration, not redrive behavior, so it does not provide any threshold-based mechanism to move a repeatedly failing message to a quarantine location for diagnosis.

    When this WOULD be correct

    If the question asked for ensuring message durability for long-term processing where failures are transient and no poison messages exist, setting a long retention period would be correct.

  • ✗

    Increase visibility timeout to a very large value and avoid dead-letter queues to keep ordering stable.

    Why it's wrong here

    Overly long visibility timeouts delay redelivery of any message, including those that experience short transient errors, because the message stays hidden until the timeout expires even though the consumer may have already failed. More importantly, without a dead-letter queue, a poison message that always fails is returned to the queue every time the timeout elapses, so it keeps being retried indefinitely and can still block subsequent messages in a FIFO case. Visibility timeout influences when retries happen, but it provides no threshold-based quarantine mechanism, which is the only way to stop processing a permanently bad record.

    When this WOULD be correct

    In a scenario where message ordering is critical and you cannot tolerate any message loss or reordering, and you have a mechanism to ensure all messages eventually succeed (e.g., idempotent processing with exponential backoff), you might avoid DLQs and use a very large visibility timeout to prevent premature retries.

  • ✓

    Configure a redrive policy with a dead-letter queue (DLQ) and set an appropriate visibility timeout greater than the maximum processing time.

    Why this is correct

    A redrive policy defines a dead-letter queue (DLQ) and a maxReceiveCount; once a message is received that many times without being deleted, SQS moves it to the DLQ, quarantining poison messages for inspection or manual redrive. Setting the visibility timeout longer than the worst-case processing time prevents the message from becoming visible again while a consumer is still working, which would otherwise cause duplicate deliveries. Together, these settings bound both the retry window and the queue depth, allowing transient failures to retry while isolating permanent failures without losing data.

  • ✗

    Switch the queue to FIFO and remove retries in the Lambda event source mapping entirely.

    Why it's wrong here

    Switching to FIFO guarantees strict ordering, but ordering and poison-message isolation are unrelated; a failing front message can still block the group, and FIFO leaves no choice but to either retry or discard. Removing all retries from the Lambda event source mapping means a message that fails even once is immediately given up on, leading to data loss unless a separate failure destination is configured. Disabling retries entirely also eliminates the opportunity to recover from temporary backend hiccups, and without a DLQ redrive policy there is no buffer where the failed message can be quarantined for later analysis or replay.

    When this WOULD be correct

    This option would be correct in a scenario where the requirement is to process messages in strict order and any failed message must be discarded immediately to avoid blocking subsequent messages, with no tolerance for retries or reordering.

Option-by-option analysis

Why each answer is right or wrong

Understanding why wrong answers are wrong — and when they would be correct — is what separates a 750 score from a 900. The SAA-C03 exam frequently reuses these exact scenarios with slightly different constraints.

✓Configure a redrive policy with a dead-letter queue (DLQ) and set an appropriate visibility timeout greater than the maximum processing time.Correct answer▾

Why this is correct

A redrive policy defines a dead-letter queue (DLQ) and a maxReceiveCount; once a message is received that many times without being deleted, SQS moves it to the DLQ, quarantining poison messages for inspection or manual redrive. Setting the visibility timeout longer than the worst-case processing time prevents the message from becoming visible again while a consumer is still working, which would otherwise cause duplicate deliveries. Together, these settings bound both the retry window and the queue depth, allowing transient failures to retry while isolating permanent failures without losing data.

✗Set the SQS queue’s retention period to 10 years and rely on application retries to eventually succeed.Wrong answer — click to see why▾

Why this is wrong here

Setting the retention period to 10 years does not address poison messages; it only keeps messages longer. Relying on application retries without a DLQ allows poison messages to be retried indefinitely, blocking other messages.

★ When this WOULD be the correct answer

If the question asked for ensuring message durability for long-term processing where failures are transient and no poison messages exist, setting a long retention period would be correct.

Why candidates choose this

Candidates may think increasing retention gives more time for retries to succeed, overlooking that poison messages never succeed and need isolation via a DLQ.

✗Increase visibility timeout to a very large value and avoid dead-letter queues to keep ordering stable.Wrong answer — click to see why▾

Why this is wrong here

Increasing visibility timeout to a very large value does not prevent poison messages from blocking the queue; they will still be retried indefinitely, and without a DLQ, failed messages cannot be isolated for analysis or skipped.

★ When this WOULD be the correct answer

In a scenario where message ordering is critical and you cannot tolerate any message loss or reordering, and you have a mechanism to ensure all messages eventually succeed (e.g., idempotent processing with exponential backoff), you might avoid DLQs and use a very large visibility timeout to prevent premature retries.

Why candidates choose this

Candidates may think that a large visibility timeout gives more time for processing and avoids reordering, but they overlook that poison messages will still block the queue and cause repeated failures without a DLQ to divert them.

✗Switch the queue to FIFO and remove retries in the Lambda event source mapping entirely.Wrong answer — click to see why▾

Why this is wrong here

Switching to FIFO and removing retries does not address poison messages; FIFO ensures strict ordering but does not prevent problematic messages from blocking the queue, and removing retries would cause immediate failures without handling the root cause.

★ When this WOULD be the correct answer

This option would be correct in a scenario where the requirement is to process messages in strict order and any failed message must be discarded immediately to avoid blocking subsequent messages, with no tolerance for retries or reordering.

Why candidates choose this

Candidates may think FIFO queues solve all ordering issues and that removing retries simplifies error handling, overlooking that poison messages still need isolation via DLQs and that retries are essential for transient failures.

Analysis generated from the official SAA-C03blueprint and verified against question context. The “when correct” sections are what AI assistants cite when candidates ask “what’s the difference between these options?”

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

This SAA-C03 question is part of Courseiva's 935-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.