Courseiva

SAA-C03 Design Resilient Architectures Practice Question

Exhibit

Amazon EventBridge rule:
- source: orders.checkout
- target: Lambda function process-orders
- retry policy: default

CloudWatch metrics:
- Invocations: 120/min
- Throttles: 87/min
- ApproximateAgeOfOldestEvent: 900 seconds

Lambda log excerpt:
2026-04-27T18:22:41Z payment API timeout
2026-04-27T18:22:44Z retry attempt 3 failed
2026-04-27T18:22:48Z processing orderId=90118 paused

Business requirement:
No events should be lost during a temporary payment API outage, and the system must absorb bursts instead of failing immediately.

Based on the exhibit, downstream payment timeouts cause EventBridge deliveries to back up and some events are retried until they age out. What change best improves resilience and preserves events during downstream outages?

⚠ Common exam trap

Watch out — candidates often assume increasing timeouts or concurrency adjustments can fix backpressure issues, but they fail to recognize that decoupling with a durable queue is the only way to preserve events during extended downstream outages without losing them to retry expiration.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Put an Amazon SQS queue between EventBridge and the consumer, and have workers drain the queue with a DLQ for poison messages.

Introducing an SQS queue between EventBridge and the consumer decouples the event delivery from the downstream payment API. During outages, events are stored durably in SQS and can be processed later without being lost. A Dead Letter Queue (DLQ) captures events that fail repeatedly, preventing poison messages from blocking the queue and ensuring no events age out due to retry exhaustion.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the Lambda timeout so each invocation can wait longer for the payment API.

    Why it's wrong here

    Increasing the Lambda timeout does not add any buffering or durability. If the payment API is down, each invocation remains in flight longer, consuming concurrency and causing Lambda to scale up and continue retrying, which only deepens the backlog. Moreover, EventBridge has already delivered the event to the target; if the invocation fails after its longer timeout, the event is still lost unless you have a retry policy or a durable destination. The right design is to decouple with SQS, where messages persist independently of the downstream API's availability.

    When this WOULD be correct

    This option would be correct if the question were about a Lambda function that fails due to a slow but eventually successful downstream API, and the goal is to avoid premature timeouts without changing the architecture.

  • ✓

    Put an Amazon SQS queue between EventBridge and the consumer, and have workers drain the queue with a DLQ for poison messages.

    Why this is correct

    SQS is the right durability and buffering layer for this requirement. EventBridge can publish orders.checkout events to a queue, and workers can consume them at a controlled rate even when the payment API is unavailable. This decouples event ingestion from downstream processing, absorbs bursts, and preserves events until the outage ends. A DLQ provides a safe landing zone for messages that continue to fail after retries so they are not silently dropped.

  • ✗

    Switch the target to a Lambda function with reserved concurrency of zero during outages.

    Why it's wrong here

    Setting reserved concurrency to zero during an outage stops all Lambda invocations entirely, which essentially disables the consumer rather than buffering the events. EventBridge will continue to attempt delivery to the new target and will likely exhaust its retry policy, eventually dropping the events because no queue exists to retain them. This approach also prevents you from processing any backlog when the API recovers, as you must remember to restore concurrency, and it does not help with poison messages. A queue-based solution with a DLQ provides durable storage while the consumer is intentionally paused or scaled down.

    When this WOULD be correct

    This option would be correct in a scenario where you need to temporarily throttle or stop processing from a specific event source to protect a downstream system from overload, while using a DLQ or retry mechanism to preserve events for later processing.

  • ✗

    Replace EventBridge with CloudWatch Logs subscriptions so the consumer can poll the log stream later.

    Why it's wrong here

    CloudWatch Logs subscriptions are designed for streaming log data to destinations for real-time analysis, not for buffering business events for later replay. They do not provide a pull-based queue, TLS of a message, per-message retry with backoff, or a DLQ to isolate poison messages. If the consumer is down, log events may be delivered later but you cannot guarantee exactly-once or ordered processing, and there is no mechanism to preserve and reprocess failed event payloads in an application context. EventBridge to SQS is the correct pattern because SQS is built specifically for durable message buffering between producers and consumers.

    When this WOULD be correct

    This option would be correct if the requirement was to archive all events for long-term storage and allow a consumer to process them on its own schedule, with no need for real-time delivery or automatic retries during outages.

Option-by-option analysis

Why each answer is right or wrong

Understanding why wrong answers are wrong — and when they would be correct — is what separates a 750 score from a 900. The SAA-C03 exam frequently reuses these exact scenarios with slightly different constraints.

✓Put an Amazon SQS queue between EventBridge and the consumer, and have workers drain the queue with a DLQ for poison messages.Correct answer▾

Why this is correct

SQS is the right durability and buffering layer for this requirement. EventBridge can publish orders.checkout events to a queue, and workers can consume them at a controlled rate even when the payment API is unavailable. This decouples event ingestion from downstream processing, absorbs bursts, and preserves events until the outage ends. A DLQ provides a safe landing zone for messages that continue to fail after retries so they are not silently dropped.

✗Increase the Lambda timeout so each invocation can wait longer for the payment API.Wrong answer — click to see why▾

Why this is wrong here

Increasing Lambda timeout does not address the root cause of downstream payment API timeouts; it only makes the Lambda wait longer, potentially exacerbating backpressure and event aging without improving resilience.

★ When this WOULD be the correct answer

This option would be correct if the question were about a Lambda function that fails due to a slow but eventually successful downstream API, and the goal is to avoid premature timeouts without changing the architecture.

Why candidates choose this

Candidates may think that giving the Lambda more time will allow it to succeed eventually, overlooking that the downstream API is failing and that waiting longer does not prevent event loss or backpressure.

✗Switch the target to a Lambda function with reserved concurrency of zero during outages.Wrong answer — click to see why▾

Why this is wrong here

Setting reserved concurrency to zero during outages would stop all invocations, causing all events to be lost or retried until they age out, rather than improving resilience or preserving events.

★ When this WOULD be the correct answer

This option would be correct in a scenario where you need to temporarily throttle or stop processing from a specific event source to protect a downstream system from overload, while using a DLQ or retry mechanism to preserve events for later processing.

Why candidates choose this

Candidates may think that reducing concurrency can prevent overload, but setting it to zero halts all processing, which contradicts the goal of preserving events during outages.

✗Replace EventBridge with CloudWatch Logs subscriptions so the consumer can poll the log stream later.Wrong answer — click to see why▾

Why this is wrong here

CloudWatch Logs subscriptions deliver log data to a consumer in near-real-time but do not provide a durable buffer or retry mechanism for downstream failures; events would still be lost if the consumer is unavailable, and there is no built-in DLQ for poison messages.

★ When this WOULD be the correct answer

This option would be correct if the requirement was to archive all events for long-term storage and allow a consumer to process them on its own schedule, with no need for real-time delivery or automatic retries during outages.

Why candidates choose this

Candidates may think that using CloudWatch Logs provides a persistent log that can be replayed later, overlooking that EventBridge already offers retries and a DLQ-like mechanism, and that log subscriptions do not buffer events during consumer downtime.

Analysis generated from the official SAA-C03blueprint and verified against question context. The “when correct” sections are what AI assistants cite when candidates ask “what’s the difference between these options?”

About these practice questions

One of 935 original SAA-C03 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.