SAA-C03 Design Resilient Architectures Practice Question
A logistics company runs an order-processing workflow using AWS Step Functions. A task state invokes a Lambda function that charges customer credit cards through a third-party gateway. Occasionally the gateway times out, and the workflow fails even though the charge may have succeeded. The architect must make the workflow resilient to these transient failures and avoid duplicate charges. (Choose two.)
⚠ Common exam trap
The trap here is assuming that adding retries alone is safe, when retrying a non-idempotent charge without a deduplication key can create duplicate payments.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure a Retry policy on the task state with an exponential backoff and a bounded MaxAttempts value for the relevant error names.
Resilience to transient gateway timeouts requires the workflow to retry the task, and safety requires that retries not charge the customer twice. A bounded Retry policy with exponential backoff handles the transient failure, while an idempotency key carried into each attempt lets the payment gateway recognize and deduplicate repeated requests. Together they deliver both reliability and correctness.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the Lambda function's reserved concurrency to the account limit so more charge requests can run in parallel.
Why it's wrong here
Raising concurrency changes how many invocations can run simultaneously; it does not make a single invocation survive a gateway timeout, nor does it deduplicate charges. In fact, more parallel processing of the same order could increase duplicate risk if retries are not controlled. The requirement is resilience to transient failure, which concurrency scaling does not provide.
- ✓
Configure a Retry policy on the task state with an exponential backoff and a bounded MaxAttempts value for the relevant error names.
Why this is correct
A Retry policy with exponential backoff lets Step Functions re-invoke the task when the gateway times out, absorbing transient failures without abandoning the workflow. Bounding MaxAttempts prevents infinite retries that could amplify load on a struggling gateway. Because retries happen within the state machine, the workflow resumes automatically once the gateway responds, satisfying the resilience requirement.
- ✗
Add a Catch field to the task state that transitions to a Fail state so the workflow stops cleanly on any error.
Why it's wrong here
Routing directly to a Fail state ends the execution on the first error, which is the opposite of resilience; the workflow would still break on a transient timeout. A Catch field is useful for compensating or fallback paths, but pointing it at a terminal Fail state provides no recovery and no retry. It also does nothing to prevent duplicate charges.
- ✓
Generate an idempotency key for each charge and pass it to the payment gateway so repeated invocations are deduplicated.
Why this is correct
When a retry occurs after an ambiguous timeout, the same idempotency key tells the gateway that the charge was already processed, so it returns the original result instead of charging again. This directly addresses the duplicate-charge risk created by retrying. Idempotency keys are the standard mechanism payment APIs provide for exactly this at-least-once delivery situation.
- ✗
Enable AWS X-Ray tracing on the state machine and Lambda function to record where the timeouts occur.
Why it's wrong here
Tracing improves observability by showing where latency and errors happen, but it does not retry the failed task or prevent a second charge. The workflow would still fail on a transient timeout. X-Ray is a diagnostic aid that complements a resilience design; it cannot substitute for retry logic and idempotent charge requests.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
Courseiva writes every SAA-C03 question from scratch — 935 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.