DOP-C02 Resilient Cloud Solutions Practice Question
A company runs a critical application on AWS Lambda that processes messages from an Amazon SQS queue. The application must be resilient to downstream service failures. The team notices that when the downstream service is unhealthy, messages are repeatedly retried and eventually sent to the dead-letter queue (DLQ) before the service recovers. What design change would improve resilience by allowing automatic retries after the downstream service recovers?
⚠ Common exam trap
Many exam-takers think increasing the DLQ retention or reducing retries (maxReceiveCount) is the solution, but the real key is controlling the retry timing via the visibility timeout to allow the downstream service to recover before messages are exhausted.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure the SQS queue with a large visibility timeout (e.g., 6 hours) and use a redrive policy only after a high number of receives. Keep the messages in the queue and retry when the downstream service becomes healthy.
Increasing the visibility timeout to a long duration (e.g., 6 hours) prevents messages from being repeatedly retried and sent to the DLQ while the downstream service is unhealthy. Instead, messages remain in the SQS queue and become visible again only after the visibility timeout expires, allowing automatic retries once the downstream service recovers. This approach avoids premature DLQ delivery and leverages SQS's built-in redrive policy based on maxReceiveCount.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Configure the SQS queue with a large visibility timeout (e.g., 6 hours) and use a redrive policy only after a high number of receives. Keep the messages in the queue and retry when the downstream service becomes healthy.
Why this is correct
This approach leverages the SQS visibility timeout as a built-in retry buffer: setting it to 6 hours prevents messages from being redelivered to consumers while the downstream service is unhealthy, effectively holding them in the queue. Combined with a high maxReceiveCount (e.g., 100) in the redrive policy, messages are not prematurely moved to the DLQ, allowing the Lambda consumer to keep retrying for hours until the service recovers. This ensures no data loss and avoids manual intervention, making it the correct pattern for intermittent downstream outages.
- ✗
Reduce the maxReceiveCount to 1 so that messages are sent to DLQ immediately, then reprocess them from DLQ later.
Why it's wrong here
Setting maxReceiveCount to 1 means any message that fails processing—even a transient failure—goes straight to the DLQ after a single attempt. This is wrong because it treats all failures as permanent, bypassing the retry logic that SQS provides. Messages would need to be manually reprocessed from the DLQ after the downstream service recovers, which is operationally heavy and risks message expiration if not handled promptly. It also defeats the purpose of using SQS as a durable, automatic retry buffer.
- ✗
Increase the message retention period to 14 days and use a DLQ with high retention.
Why it's wrong here
Increasing the message retention period (to 14 days) or using a high-retention DLQ only controls how long messages remain in the system; it does not affect whether or when they are retried. Once a message reaches maxReceiveCount, it is moved to the DLQ and becomes passive—nothing automatically retries it from there. This option merely delays eventual data loss or forces a custom reprocessing solution, so it fails to provide the continuous, automatic retry mechanism that the application requires.
- ✗
Use Amazon SNS to fan out messages to multiple SQS queues, each with different retry policies.
Why it's wrong here
Amazon SNS is a pub/sub service that simply fans out messages to subscribed SQS queues; it has no retry logic, no visibility timeout, and no awareness of downstream health. Configuring multiple SQS queues with different retry policies adds complexity and duplicates messages across queues, but each queue's consumer still needs to handle its own retries independently. It does not solve the core problem of holding messages while waiting for the downstream service to recover, and it introduces unnecessary operational overhead.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
This DOP-C02 question is part of Courseiva's 1,298-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.