DOP-C02 Incident and Event Response Practice Question
A company is using AWS Lambda to process events from an Amazon SQS queue. The Lambda function is configured with a batch size of 10 and a maximum concurrency of 5. Recently, the function started experiencing high error rates and the SQS queue's ApproximateNumberOfMessagesVisible metric is increasing. The CloudWatch logs show that the function is timing out after 30 seconds. The function makes calls to an external API that sometimes takes more than 30 seconds to respond. The DevOps engineer needs to reduce the backlog and prevent message loss. The engineer is considering the following actions: A) Increase the Lambda function timeout to 60 seconds and increase the SQS visibility timeout to 90 seconds. B) Decrease the batch size to 1 to avoid processing multiple messages at once. C) Increase the Lambda function reserved concurrency to 100 to allow more concurrent executions. D) Use a dead-letter queue to capture messages that fail processing after all retries. Which combination of actions should the engineer take?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase the Lambda function timeout to 60 seconds and increase the SQS visibility timeout to 90 seconds.
The correct action because increasing the Lambda function timeout to 60 seconds allows the function to wait longer for the external API, and increasing the SQS visibility timeout to 90 seconds prevents messages from becoming visible again before the function completes. This reduces unnecessary retries and helps clear the backlog. Option A (DLQ) is useful for capturing failed messages but does not address the timeout issue. Option B (decrease batch size) reduces throughput and worsens the backlog. Option D (increase concurrency) may lead to more timeouts if the function still cannot complete within the existing timeout.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a dead-letter queue to capture messages that fail processing after all retries.
Why it's wrong here
A dead-letter queue does help by capturing messages that exhaust retries and preventing data loss, but it does nothing to clear the messages still sitting in the source queue or to stop new failures from piling up. If the Lambda function is timing out on every attempt, the DLQ will simply accumulate those failed messages while the backlog remains unchanged. The proper fix is to reduce the per-invocation workload or extend timeout and visibility timeout, not divert the failures.
- ✗
Decrease the batch size to 1 to avoid processing multiple messages at once.
Why it's wrong here
Lowering the batch size from the default to 1 cuts the number of messages processed per invocation, which ironically increases the total number of Lambda invocations and reduces overall throughput. It also does not address the root cause: if a single message still takes longer to process than the function timeout, that one message will still time out and be retried. With a large backlog, this change will only make the pile-up worse while adding more SQS polling and invocation overhead.
- ✓
Increase the Lambda function timeout to 60 seconds and increase the SQS visibility timeout to 90 seconds.
Why this is correct
This is correct because an SQS-triggered Lambda invocation has a maximum execution window set by the function timeout, and the SQS visibility timeout controls when unacknowledged messages become visible again for redelivery. If the function timeout is too short, valid work gets aborted, and if the visibility timeout is shorter than the processing time, the message is re-delivered before the first attempt finishes, causing duplicate work and retries that inflate the backlog. Setting the visibility timeout to 90 seconds (longer than the 60-second function timeout) ensures the message stays hidden until the Lambda function either succeeds or itself times out, giving the function the full time it needs.
- ✗
Increase the Lambda function reserved concurrency to 100 to allow more concurrent executions.
Why it's wrong here
Adding reserved concurrency up to 100 only permits more executions to run in parallel; it does not shorten the time each execution spends doing work. If every execution is already timing out, raising concurrency means 100 executions may all fail simultaneously, creating a burst of retries and potentially overwhelming downstream dependencies. Also, for SQS event source mappings, high concurrency can cause the mapping to throttle or invoke with less optimal batch behavior, so the backlog problem remains—or gets worse.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DOP-C02 question from scratch — 1,298 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.