Courseiva

SAA-C03 · domain

Design Resilient Architectures

Use this page to practise high availability and resilience questions. The SAA-C03 exam tests your ability to match an architecture pattern to an RTO/RPO requirement — know the cost and recovery time of each pattern.

257 questions55 easy141 medium61 hard

Focused practice

Practice Design Resilient Architectures questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Design Resilient Architectures

High availability and resilience questions test multi-AZ vs multi-Region patterns, Auto Scaling, load balancing and the right service for a given recovery time objective.

Multi-AZ vs multi-Region deployment trade-offs.

Auto Scaling policies and when to scale horizontally vs vertically.

Elastic Load Balancing: ALB, NLB, CLB and their use cases.

RTO and RPO targets matched to the correct AWS architecture.

Watch out for

Common Design Resilient Architectures exam traps

  • ▸Multi-AZ protects against AZ failure; multi-Region protects against Region failure.
  • ▸Auto Scaling does not guarantee zero downtime without a load balancer.
  • ▸ALB operates at Layer 7; NLB operates at Layer 4.
  • ▸Pilot light is cheaper than warm standby but has longer recovery time.

Question index

All Design Resilient Architectures questions (257)

Click any question to see the full explanation, or start a practice session above.

1

A production Amazon RDS database already has automated backups enabled. At 10:45 UTC, the team discovers that a faulty migration corrupted rows in a table at 10:30 UTC. The business wants the database restored to exactly the state it had at 10:30 UTC with minimal risk. Which two actions should the team take? Select two.

Medium
2

Based on the exhibit, the web team wants the application to continue serving traffic if one Availability Zone fails. Which change best meets the requirement with the least operational overhead?

Easy
3

A trading dashboard uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The architecture review board prefers a managed AWS-native control.

Medium
4

An internal worker consumes messages from an Amazon SQS queue. Occasionally, a message fails validation in the worker (for example, missing required fields). Reprocessing the same bad message repeatedly wastes processing time and delays healthy messages. What is the best AWS approach to handle these poison messages without blocking the rest of the queue?

Easy
5

A production application uses an Amazon RDS Multi-AZ DB instance. During an unplanned failover, the database endpoint remains the same. What change should the application team make to handle the failover reliably?

Easy
6

Based on the exhibit, a web application must stay available if one Availability Zone fails. What is the best change to improve resilience?

Easy
7

An order-processing service consumes messages from an Amazon SQS Standard queue using a custom worker. During traffic spikes, the worker occasionally times out after performing some work but before acknowledging the message, so SQS redelivers it and it may be processed again. You also observe that a small set of “poison” messages always fail validation. What change most directly improves resilience by (1) preventing poison messages from retrying indefinitely and (2) avoiding duplicate side effects caused by legitimate retries?

Medium
8

A healthcare provider hosts a patient-records API on Amazon EC2 instances in a single Availability Zone behind an Application Load Balancer. An audit finds the architecture cannot tolerate the loss of that Availability Zone. Budget is limited, and the API reads from an Amazon Aurora MySQL cluster that currently has one writer instance and no replicas. Which change most effectively addresses the audit finding?

Medium
9

A company runs a stateless API on Amazon EC2 instances in a single Availability Zone behind an Application Load Balancer. The ALB currently has a listener on port 80 only. The company wants the API to remain available if the single Availability Zone fails. What should the solutions architect do to meet this requirement with the LEAST operational overhead?

Easy
10

A web application runs on an EC2 Auto Scaling group (ASG) behind an Application Load Balancer (ALB). The ASG spans three Availability Zones. After a deployment, new instances frequently fail the ALB target group health checks with HTTP 5xx responses and are quickly terminated by the ASG. What change most improves resiliency during deployments with minimal downtime by preventing premature removal of instances that are still starting?

Medium
11

Based on the exhibit, the application sees several minutes of connection errors during an Aurora failover. What is the best change to reduce failover impact?

Medium
12

A patient portal must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The architecture review board prefers a managed AWS-native control.

Hard
13

A healthcare analytics platform stores derived datasets in an Amazon S3 bucket. Regulatory rules require that every object remain recoverable for 90 days after creation even if an application bug issues a delete, and that no object version be permanently destroyed during that window. The team wants the strongest protection with the least custom code. Which S3 feature should the solutions architect enable?

Hard
14

A company runs a critical two-tier web application on AWS. The web tier consists of Amazon EC2 instances behind an Application Load Balancer (ALB) in a single Availability Zone. The database tier is an Amazon RDS for MySQL DB instance in the same Availability Zone. A recent power outage in that Availability Zone caused a full application outage. The company wants to redesign the architecture to survive an Availability Zone failure with minimal operational overhead. Which solution meets these requirements?

Medium
15

An orders service publishes payment instructions to an Amazon SQS Standard queue. The downstream processor sometimes times out after it has already applied the payment, but before it can delete the message from the queue. As a result, the same payment instruction can be processed more than once. The team wants the strongest way to prevent duplicate side effects while keeping the system decoupled. What should they implement?

Medium
16

Your order-processing system uses EventBridge rules to send events to a Lambda function that updates order status. Over the last week, some events fail with a transient database timeout, and the Lambda retries intermittently but then the events are lost (no alerts after failures). You want at-least-once processing, bounded retries, and a way to inspect unprocessable events for later reprocessing. Which architecture change best meets these requirements?

Medium
17

A inventory service exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most? The architecture review board prefers a managed AWS-native control.

Easy
18

A ticket booking system stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured?

Medium
19

A retail API runs on Amazon EC2 instances behind an Application Load Balancer and stores orders in an Amazon RDS for PostgreSQL database. A test that stopped one Availability Zone caused the API to return errors because all application servers were in the same AZ and the database was single-AZ. Which two changes should the architect make to continue serving traffic during a single-AZ failure? Select two.

Medium
20

Match the disaster recovery strategy to the recovery posture it best fits for a Regional outage.

Medium
21

A payments service receives payment orders by consuming messages from an Amazon SQS Standard queue. The downstream processor occasionally exceeds its processing timeout. As a result, some messages reappear in the queue and may be processed more than once. The team wants to prevent duplicate side effects (for example, double-charging) and also ensure poison messages do not repeatedly consume processing capacity. What approach best satisfies both goals?

Medium
22

A financial services firm runs a batch settlement job on a fleet of Amazon EC2 instances that pull work from an Amazon SQS queue. The job must not lose messages if an instance is terminated mid-processing, and duplicate processing must be minimized because each settlement charge is expensive. The team also wants to avoid indefinite reprocessing of a message that repeatedly fails. Which two changes should the solutions architect make to meet these requirements? (Choose two.)

Hard
23

A company runs a stateful workload on Amazon EC2 instances in an Auto Scaling group. The workload writes session data to the instance store and to an Amazon EBS volume attached at launch. The company wants the workload to survive an Availability Zone failure without losing session data. What should the solutions architect do?

Hard
24

A fintech company has a two-Region DR requirement: RPO must be within 15 minutes and RTO must be under 2 hours. To control cost, they do not want to run full production infrastructure in the secondary Region continuously. They plan to continuously replicate the database and keep the application infrastructure in the secondary Region prepared, but at reduced capacity. Which DR strategy best matches this requirement and accurately describes their plan?

Medium
25

A media company stores generated video thumbnails in an Amazon S3 bucket. The bucket currently uses the S3 Standard storage class, and the objects are accessed frequently for the first 30 days and then almost never. The company wants to reduce storage costs automatically without changing the application and must retain the objects for at least one year. Which action should a solutions architect take?

Easy
26

A logistics company runs an order-processing workflow using AWS Step Functions. A task state invokes a Lambda function that charges customer credit cards through a third-party gateway. Occasionally the gateway times out, and the workflow fails even though the charge may have succeeded. The architect must make the workflow resilient to these transient failures and avoid duplicate charges. (Choose two.)

Medium
27

A healthcare company runs a stateless patient-intake API on a fleet of Amazon EC2 instances in a single VPC. The compliance team requires the workload to survive the complete loss of one Availability Zone with no manual intervention, and the instances must be replaced automatically if they fail health checks. The application stores no local state and writes all data to Amazon RDS. Which approach meets these requirements with the LEAST operational effort?

Medium
28

A payments API uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure?

Hard
29

A content publishing system exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most?

Easy
30

A company runs an application behind an Application Load Balancer (ALB). An Auto Scaling group (ASG) is configured with desired capacity 2, but it is attached only to subnets in a single Availability Zone. The ALB is healthy because it is configured across multiple Availability Zones. When the Availability Zone that contains the ASG subnets experiences an outage, what change most directly improves resilience and allows capacity to be restored automatically?

Medium
31

A logistics company runs an order-processing workload that reads messages from an Amazon SQS queue and writes results to an Amazon DynamoDB table. Occasionally the same order is processed twice and produces duplicate shipments. The architects must ensure each order is processed exactly once end to end, while keeping throughput as high as possible. What should they do?

Hard
32

A global application experiences frequent writes and must survive a full Regional outage with near-zero data loss. The product team also requires that users can continue to write during the incident using the closest Region. Which approach is most aligned with these requirements?

Medium
33

A patient portal receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers?

Medium
34

An engineering team deploys a stateless web API on EC2 using an Auto Scaling group and an Application Load Balancer (ALB). During a recent test, they noticed that when one Availability Zone was unavailable, traffic failed until new instances were manually launched. Which change most directly improves automatic failover for the compute layer within a single Region?

Easy
35

A payments API requires point-in-time recovery and accidental-delete protection for a DynamoDB table. Which two settings should the architect enable? The team wants the control to be enforceable during normal operations.

Hard
36

A healthcare company runs a containerized claims-processing service on Amazon ECS with the Fargate launch type in a single AWS Region. The service must survive the loss of an entire Availability Zone with no manual intervention, and the architecture must keep the same service endpoint for callers. The service is fronted by an Application Load Balancer. Which combination of actions should a solutions architect take to meet these requirements with the LEAST operational overhead?

Medium
37

A claims workflow uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure?

Hard
38

An orders service publishes payment instructions to an Amazon SQS queue. After occasional processing timeouts, the downstream consumer sometimes processes the same instruction twice, resulting in duplicate payment attempts. The team currently uses an SQS Standard queue with a visibility timeout of 2 minutes and relies on the consumer to finish before the timeout expires. What approach best improves resilience against duplicate processing?

Medium
39

An orders system sends payment instructions to an Amazon SQS queue. The consumer sometimes times out after it has already created the payment record but before it deletes the SQS message. As a result, the same instruction can be processed more than once. Which design best ensures the consumer remains resilient and does not create duplicate payments when the same instruction is delivered multiple times?

Medium
40

A logistics company runs an order processing system on Amazon EC2 instances that read and write to an Amazon RDS for MySQL database. The database is currently a Single-AZ deployment. The company needs the database to survive an Availability Zone failure with automatic failover and minimal downtime. The application connects using a hardcoded DNS name. Which change should a solutions architect make?

Hard
41

A customer portal must recover from a regional outage within a few hours. The business wants lower ongoing cost than a fully active second Region and does not want to rebuild everything from scratch during the outage. Which two DR patterns best fit that goal? Select two.

Medium
42

A media company stores master video files in an Amazon S3 bucket in the us-east-1 Region. A compliance policy requires that the data remain readable even if the entire us-east-1 Region becomes unavailable, and the recovery point objective is 15 minutes. The team wants the lowest operational overhead and does not want to modify application code. Which solution should the architect implement?

Hard
43

A financial analytics platform runs an Amazon Aurora MySQL cluster with one writer and two readers. During month-end reporting, read traffic spikes and the application sometimes receives TooManyConnections errors on the reader endpoint. The architect wants to absorb bursts without changing application code and must keep failover behaviour intact. Which change meets these requirements?

Hard
44

A logistics company runs a shipment-tracking service on a single Amazon EC2 instance in one Availability Zone. The instance stores tracking state in an attached Amazon EBS volume and writes nightly backups to Amazon S3. The company needs the service to survive the loss of an entire Availability Zone with minimal data loss and automatic recovery, while keeping changes minimal. Which design change should a solutions architect recommend?

Medium
45

A production Amazon RDS database has automated backups enabled. At 10:45 UTC, an issue is discovered. The team needs to restore the database to its state as of 10:30 UTC. Which capability should they use?

Easy
46

A solutions architect is designing a highly available and resilient architecture for a critical internal application that processes financial transactions. The application runs on Amazon EC2 instances inside an Auto Scaling group. The database layer uses an Amazon Aurora MySQL cluster. The company requires that if an entire AWS Availability Zone (AZ) fails, the application must remain operational with minimal impact and automatically recover without manual intervention. Which combination of architectural decisions will meet these requirements? (Choose four.)

Medium
47

A startup runs a single Amazon EC2 instance hosting both a web application and its MySQL database. The founders want the application to survive the failure of the underlying hardware without changing the database engine, and they want the smallest possible operational change. What should the architect recommend?

Easy
48

Your public API is hosted in two regions. You want Route 53 to automatically send traffic to the secondary region when the primary region’s endpoint fails. The primary API health check is returning failure codes, but clients still reach the primary region for several minutes. Which Route 53 configuration most directly addresses this behavior?

Medium
49

Based on the exhibit, some SQS messages fail validation repeatedly and continue consuming worker time. What change best prevents the bad messages from being retried forever?

Easy
50

A patient portal receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers? The architecture review board prefers a managed AWS-native control.

Medium
51

A team runs an Amazon RDS for MySQL database in a single Availability Zone. They want automatic failover with minimal downtime if the primary database instance becomes unavailable. Automated backups are already enabled. Which configuration change best meets the requirement?

Easy
52

A startup runs a nightly batch job on a single Amazon EC2 instance that stores results in an Amazon EBS volume. The job takes six hours, and the team wants to resume from the last completed step if the instance is terminated unexpectedly. Which approach provides the required durability with the least operational effort?

Easy
53

A company needs an Amazon RDS database that automatically fails over to a standby when the primary DB instance becomes unavailable. Which approach best meets the requirement with minimal operational effort?

Easy
54

A SaaS platform serves an API using two regional deployments: us-east-1 (primary) and us-west-2 (secondary). Each region has its own ALB. The business requires automated DNS-based failover when the primary region becomes unhealthy, and they do not want manual DNS changes during incidents. Which Route 53 configuration is the best match?

Medium
55

Based on the exhibit, DNS still sends traffic to the primary Region even though Route 53 health checks show the primary endpoint is unhealthy. What is the best change to make failover work as intended?

Hard
56

A patient portal must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used?

Hard
57

A claims workflow uses an RDS MySQL database and must remain available during an Availability Zone failure with minimal application changes. What should the architect enable?

Medium
58

A ticket booking system stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured? The architecture review board prefers a managed AWS-native control.

Medium
59

A developer accidentally corrupts part of a production Amazon RDS database, and the issue is discovered 45 minutes later. The team needs to restore the database to the state immediately before the change. Which two actions should be part of the recovery plan? Select two.

Easy
60

Your media processing pipeline writes original uploads to an S3 bucket and later generates derivative files. An operator accidentally deletes a subset of original uploads in production. You need to (1) restore the deleted objects with minimal data loss and (2) protect against both regional disasters and future operator mistakes. The company requires recovery even if objects are deleted and later overwritten. What is the most effective change to meet these requirements?

Medium
61

A patient portal must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The design must avoid adding custom operational scripts.

Hard
62

A healthcare provider runs a patient-record API on Amazon EC2 instances behind an Application Load Balancer in one AWS Region. The API reads from an Amazon RDS for MySQL DB instance. The provider must be able to continue serving read traffic if the primary database instance fails, and must minimize the time the application is unavailable. Which change should a solutions architect make?

Medium
63

An organization hosts the same public API in two AWS Regions. Normal traffic should go to the primary Region. If the primary endpoint becomes unhealthy, Route 53 should automatically route users to the secondary Region. What is the best Route 53 configuration approach?

Easy
64

Based on the exhibit, the web application must remain available even if one Availability Zone fails. What is the best change to improve resilience with the least redesign?

Medium
65

A logistics firm runs an order-processing service that reads from an Amazon SQS queue and writes results to an Amazon DynamoDB table. During a marketing event, the consumer fleet scaled out aggressively and DynamoDB began returning ProvisionedThroughputExceededException errors, causing messages to be retried and some orders to be processed twice. The architects want to absorb traffic spikes without overprovisioning capacity and without duplicate processing. Which combination of changes should they make?

Hard
66

A warehouse integration service must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable?

Hard
67

A team needs a relational database solution that can automatically fail over to a standby instance if the primary database becomes unavailable. They want the standby to be located in a different Availability Zone. Which RDS/Aurora configuration best satisfies this requirement?

Easy
68

A healthcare analytics platform processes streaming records with an AWS Lambda function that writes results to an Amazon DynamoDB table. The pipeline must not lose records if the function throws an error, and the operations team wants to inspect and reprocess failed records without writing custom retry code. Which approach should the solutions architect use?

Hard
69

A financial services firm runs a critical API on Amazon EC2 instances behind a Network Load Balancer. The API must handle a sudden loss of one Availability Zone and continue serving traffic with no manual failover. The instances are in an Auto Scaling group that currently uses a single subnet in one Availability Zone. Which change should the architect make?

Medium
70

A trading dashboard runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include? The architecture review board prefers a managed AWS-native control.

Medium
71

An event-driven order processing service consumes messages from an Amazon SQS Standard queue. After a deployment, about 1% of messages start failing validation because a required field is missing. The consumer catches the exception and returns control, so the messages are retried. However, those poison messages keep reappearing and repeatedly consuming processing time for hours, delaying handling of valid messages. What is the most resilient way to handle the poison messages while keeping the system available?

Medium
72

A claims workflow uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure? The architecture review board prefers a managed AWS-native control.

Hard
73

An application uses an Amazon RDS Multi-AZ DB instance. During a failover test, connections fail until the application is restarted, even though the database comes back online. Which two changes should the team make to improve resilience during failover? Select two.

Medium
74

An Auto Scaling group behind an Application Load Balancer frequently replaces new EC2 instances. The application needs ~6 minutes to warm up after instance launch. However, the ALB target group health checks start immediately and mark the targets unhealthy until the application is ready. Because the targets become unhealthy early, the Auto Scaling group then terminates the instances and launches replacements, creating a repeated unhealthy/termination loop. What configuration change will most directly improve recovery by preventing premature ASG termination while the application is warming up?

Medium
75

A company runs an internet-facing API in two AWS Regions. Route 53 currently uses simple routing to a primary Application Load Balancer (ALB) DNS name. When the primary Region experiences an outage, customers wait a long time because the DNS entry is not changed automatically. The team wants automatic failover: if the primary Region ALB health check fails for a sustained period, Route 53 should route users to the secondary Region ALB. Which Route 53 approach best meets this requirement?

Medium
76

A healthcare company needs to store patient records in Amazon DynamoDB. The records must be highly available and durable across multiple Availability Zones. The company also requires the ability to recover the table to any point in time within the last 35 days in case of accidental writes or deletions. Which solution meets these requirements?

Medium
77

A company stores critical documents in an Amazon S3 bucket in the us-east-1 Region. The documents must survive an unlikely loss of the entire us-east-1 Region. The company wants a recovery point objective (RPO) of 15 minutes and a recovery time objective (RTO) of 1 hour. What should the solutions architect recommend?

Medium
78

A healthcare company runs a patient portal on Amazon EC2 instances behind an Application Load Balancer across two Availability Zones. A new compliance rule requires that if an entire Availability Zone fails, the portal must remain available with no manual intervention. The EC2 instances are stateless and store no session data. Which design change should the architect implement to meet this requirement?

Medium
79

A patient portal must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable? The team wants the control to be enforceable during normal operations.

Hard
80

A team accidentally updates critical rows in an Amazon RDS for PostgreSQL database. Automated backups are enabled. They need to recover the data to the exact state as of 90 minutes ago. They also cannot risk interrupting the current production database instance while investigators validate the restored data. Which recovery strategy best meets these constraints?

Medium
81

Order the steps to create a static website using Amazon S3 and CloudFront.

Medium
82

A logistics company runs a stateless order-tracking API on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer. The architect must ensure the API survives the loss of an entire Availability Zone and that unhealthy instances are replaced automatically. (Choose two.)

Medium
83

A company runs a critical API on Amazon EC2 behind an Application Load Balancer in a single AWS Region. The business requires the API to keep serving traffic if an entire Availability Zone becomes unavailable, and the recovery must not depend on any manual step. The database is Amazon RDS for PostgreSQL configured as a Single-AZ instance. Which combination of changes should a solutions architect implement to meet these requirements?

Hard
84

A SaaS provider runs a multi-tenant application on Amazon EC2 instances behind an Application Load Balancer. Tenants are identified by a subdomain, and each tenant's data is stored in a separate Amazon S3 bucket. The provider wants HTTPS with a single certificate, automatic renewal, and the ability to add new tenant subdomains without redeploying or replacing the certificate. Which solution meets these requirements?

Hard
85

Based on the exhibit, the database must fail over automatically if the primary Availability Zone goes down. Which solution should the architect choose?

Easy
86

A ticket booking system uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered?

Medium
87

A warehouse integration service must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The architecture review board prefers a managed AWS-native control.

Hard
88

A worker service consumes messages from an Amazon SQS queue. Some messages are malformed and always fail validation. The worker retries, but it keeps reprocessing the same bad messages and consumes processing capacity that should be used for valid work. What is the best solution to prevent “poison messages” from blocking progress?

Easy
89

A warehouse integration service receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers? The design must avoid adding custom operational scripts.

Medium
90

A media company stores video files in an Amazon S3 bucket in the us-east-1 Region. The company wants to ensure that the files are automatically replicated to us-west-2 for disaster recovery, and that replication occurs within 15 minutes of upload. Which solution meets these requirements with the LEAST operational overhead?

Medium
91

A company runs its customer-facing web app on EC2 behind an Application Load Balancer. The database is Amazon RDS for PostgreSQL. The requirement is that if a single Availability Zone fails, the database must automatically fail over within the same AWS Region with minimal application changes. Which database setup best meets this requirement?

Easy
92

A company is deploying a stateless web application on Amazon ECS with Fargate. The application must be resilient to individual task failures and Availability Zone failures. Which three steps should the company take to achieve this resilience? (Choose three.)

Medium
93

Your company hosts an internal API in two AWS Regions. You want Amazon Route 53 to automatically send traffic to the secondary Region if the primary Region’s endpoint becomes unhealthy. Which Route 53 configuration best meets this requirement?

Easy
94

A ticket booking system uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The architecture review board prefers a managed AWS-native control.

Medium
95

A trading dashboard stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured? The architecture review board prefers a managed AWS-native control.

Medium
96

A financial analytics platform runs a stateless containerized service on Amazon ECS with AWS Fargate tasks spread across three Availability Zones. The service reads from an Amazon Aurora MySQL cluster and must continue serving read traffic if one Availability Zone fails. The team wants the read capacity to remain available with the least operational overhead and no changes to application connection strings during a zone failure. Which approach meets these requirements?

Hard
97

A web application runs on an Auto Scaling group (ASG) behind an Application Load Balancer (ALB). The ASG uses the ALB target group health checks to decide when instances are healthy (for example, by using the ELB/target-group health check integration). During a deployment, the ASG performs instance replacement. Shortly after the deployment starts and while new instances are still bootstrapping, CloudWatch shows the ALB target group briefly has zero healthy targets, and users intermittently receive 502 responses. Which ASG deployment configuration best reduces the chance that there will be a period with zero healthy ALB targets, while still keeping failover behavior resilient?

Medium
98

A claims workflow uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure? The team wants the control to be enforceable during normal operations.

Hard
99

A startup runs a small internal tool on a single Amazon EC2 instance that uses an instance store volume for its database files. After a routine host maintenance event, the instance rebooted and the database was empty. The team wants the data to persist independently of the instance lifecycle and to survive a stop-and-start of the instance. What should they change?

Easy
100

A ticket booking system runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include?

Medium
101

A media company stores daily financial exports in Amazon S3. The files must be protected against accidental overwrite or deletion, and the business also wants a second copy in another Region for recovery after a regional outage. Which two actions should the architect take? Select two.

Medium
102

A serverless order-ingestion API writes directly to a database. During traffic spikes, the database occasionally throttles, Lambda retries create duplicate order records, and some requests time out. Which two changes best improve buffering and safe retry behavior? Select two.

Medium
103

A company runs a production MySQL database on Amazon RDS in us-east-1. A read replica exists in us-west-2 for disaster recovery. The primary region experiences a complete outage. Which of the following describes the correct procedure to restore database service using the cross-region read replica?

Hard
104

An orders service consumes payment instructions from an Amazon SQS queue. Sometimes the consumer times out after applying the payment but before deleting the SQS message. As a result, the same payment instruction is processed again. Which design change most directly prevents duplicate side effects caused by message retries?

Easy
105

A public API is deployed in two AWS Regions: us-east-1 (primary) and us-west-2 (secondary). The team wants Route 53 to automatically route users to the secondary region if the primary API becomes unhealthy. They will use Route 53 health checks that monitor the API’s /status endpoint over HTTPS. Which Route 53 configuration most directly implements this failover behavior?

Medium
106

A regional web application for a inventory service must fail over automatically to a secondary Region if the primary endpoint becomes unhealthy. Which two services or features are required? The team wants the control to be enforceable during normal operations.

Hard
107

A financial services company runs a critical application on Amazon EC2 instances in an Auto Scaling group. The application writes to an Amazon RDS for MySQL database. The company needs a recovery point objective (RPO) of 1 second and a recovery time objective (RTO) of 1 minute for the database in the event of a Regional disaster. Which solution meets these requirements?

Hard
108

A payments platform requires disaster recovery across Regions. Requirements: RPO of 15 minutes and RTO of about 1 hour. The business cannot afford full duplicate capacity in both Regions all the time, but the team wants automated readiness so failover is mostly operationally guided rather than a slow rebuild. Which DR strategy is the best fit?

Medium
109

A financial services firm runs a latency-sensitive trading application on Amazon EC2 instances distributed across three Availability Zones behind a Network Load Balancer. The application must continue serving traffic with no manual intervention if an entire Availability Zone becomes impaired, and each instance must receive a fair share of connections. Which combination of features meets these requirements?

Hard
110

A company runs a stateful web application on a fleet of Amazon EC2 instances in an Auto Scaling group. The application stores session state locally on each instance. During an Availability Zone failure, the Auto Scaling group replaces the unhealthy instances in a different AZ, but users lose their sessions and must log in again. The company wants to make the application resilient to AZ failures without requiring users to re-authenticate. Which solution should a solutions architect recommend?

Hard
111

Based on the exhibit, the database must continue serving if the current Availability Zone fails. What should you change?

Easy
112

A inventory service exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most?

Easy
113

Based on the exhibit, a faulty deployment corrupted production data at 10:30 UTC and the issue was discovered at 10:55 UTC. The team needs to recover the database to the last good state before the corruption. Which action should they take?

Medium
114

An order system receives events and uses a Lambda function to write each order into a database. During traffic spikes, the database sometimes throttles, and Lambda retries lead to occasional message loss in the event flow. The team wants buffering, automatic retries, and a way to isolate messages that repeatedly fail so they can be inspected later. What design change best meets this need?

Easy
115

A company uses Amazon RDS for a PostgreSQL database powering a customer-facing application. The application’s availability depends on fast database failover with minimal manual intervention. The RDS instance currently runs as a single-AZ deployment in one DB subnet group. Which change most directly meets the goal?

Medium
116

A trading dashboard runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include?

Medium
117

A trading dashboard stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured?

Medium
118

An internal service is hosted behind an Application Load Balancer (ALB) with targets spread across two Availability Zones. If the targets in one Availability Zone become unhealthy, the service must continue serving traffic from the healthy AZ. What change most directly improves resilience at the load-balancing layer?

Easy
119

A regional web application for a inventory service must fail over automatically to a secondary Region if the primary endpoint becomes unhealthy. Which two services or features are required? The design must avoid adding custom operational scripts.

Hard
120

A media company runs a stateless transcoding fleet on Amazon EC2 instances spread across three Availability Zones behind a Network Load Balancer. The fleet must keep processing jobs even if an entire Availability Zone becomes unavailable, and the architect wants to minimize manual intervention. Which combination of actions should the architect take to meet these requirements?

Medium
121

An event consumer sometimes processes the same SQS message more than once due to timeouts and retries. The consumer must ensure the payment is not charged twice. What design choice best addresses this requirement?

Easy
122

A inventory service uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The architecture review board prefers a managed AWS-native control.

Medium
123

A media company runs a transcoding fleet on Amazon EC2 instances behind an Application Load Balancer in a single Availability Zone. The business requires the workload to survive the loss of that Availability Zone with no manual intervention and minimal downtime. The instances store intermediate files on instance store volumes and the fleet is managed by an Auto Scaling group. Which change should a solutions architect make to meet the requirement?

Medium
124

A media company runs a video transcoding pipeline on Amazon EC2 instances in a single Availability Zone. The pipeline writes intermediate files to an Amazon EBS volume attached to each instance. The company needs the pipeline to survive the failure of any single Availability Zone and to recover automatically with minimal data loss. Which change should a solutions architect make?

Medium
125

A inventory service uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured?

Medium
126

Based on the exhibit, the application team wants the database to keep the same connection endpoint during failover and to reconnect automatically after the primary instance becomes unavailable. Which change best meets the requirement?

Medium
127

A content publishing system uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The architecture review board prefers a managed AWS-native control.

Medium
128

Based on the exhibit, the application tier is not replacing unhealthy instances even though the Auto Scaling group spans two Availability Zones. What change most directly improves automatic recovery when the application process fails?

Hard
129

A payments API requires point-in-time recovery and accidental-delete protection for a DynamoDB table. Which two settings should the architect enable? The architecture review board prefers a managed AWS-native control.

Hard
130

A trading dashboard uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered?

Medium
131

A startup runs a stateless web application on a single Amazon EC2 instance in one Availability Zone. The application has become popular, and the startup wants to ensure that the application can survive the failure of an Availability Zone and can handle increased traffic. Which architecture change should the startup make FIRST?

Easy
132

An events service publishes critical notifications using Amazon SNS. Three independent downstream systems (A, B, and C) subscribe to the topic. Downstream system B sometimes fails to process certain messages (for example, it times out or returns an error while handling the message), and you want: 1) failures in B to be isolated so A and C keep processing unaffected, and 2) messages that B cannot successfully process after retries to be sent to a DLQ for B. Which design best meets these requirements?

Medium
133

A stateless web API runs on EC2 instances behind an Application Load Balancer (ALB). The Auto Scaling group (ASG) currently uses subnets from only one Availability Zone, even though the ALB spans two Availability Zones. During maintenance of that single AZ, the ALB remains up but clients see timeouts because there are no healthy targets. Which change most directly improves resilience against an AZ failure?

Medium
134

An orders service publishes payment instructions to an Amazon SQS Standard queue. A downstream consumer sometimes times out or crashes after it has partially completed processing, causing the same instruction to be processed more than once. You must keep the design resilient without attempting to guarantee exactly-once processing. Which approach best handles duplicates safely?

Medium
135

A claims workflow uses an RDS MySQL database and must remain available during an Availability Zone failure with minimal application changes. What should the architect enable? The design must avoid adding custom operational scripts.

Medium
136

Based on the exhibit, the team must restore an Amazon RDS for PostgreSQL database to the exact state just before a bad delete happened. What is the best recovery approach?

Hard
137

A patient portal must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable?

Hard
138

A worker consumes messages from an Amazon SQS queue. Some messages consistently fail validation and are retried until the worker can no longer process them. What is the most appropriate AWS mechanism to handle these poison messages while keeping the queue usable?

Easy
139

A trading dashboard runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include? The team wants the control to be enforceable during normal operations.

Medium
140

Based on the exhibit, the company wants DNS traffic to fail over automatically from the primary Region to a secondary Region when the primary endpoint is unhealthy. Which Route 53 change is best?

Medium
141

A financial analytics platform stores results in an Amazon S3 bucket. Compliance requires that objects be recoverable for 30 days after deletion and that no user, including administrators, be able to permanently erase them during that window. Objects must also remain readable throughout the retention period. Which approach should the architect implement?

Hard
142

A warehouse integration service must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used?

Hard
143

A web application runs on an Auto Scaling group (ASG) behind an Application Load Balancer (ALB). The ASG is currently attached to subnets in only two Availability Zones (AZs). During a planned maintenance window, one AZ becomes unavailable for about 25 minutes. Monitoring shows that targets in the remaining AZ go healthy, and the ALB/target group health checks report normal. However, users still experience intermittent connection failures and slower responses during the AZ outage. What change will most directly improve resilience against an AZ loss while keeping the same ALB-based design?

Medium
144

A warehouse integration service receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers? The architecture review board prefers a managed AWS-native control.

Medium
145

A patient portal must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable? The design must avoid adding custom operational scripts.

Hard
146

A payments API requires point-in-time recovery and accidental-delete protection for a DynamoDB table. Which two settings should the architect enable? The design must avoid adding custom operational scripts.

Hard
147

A company hosts an internal API behind an Application Load Balancer (ALB) in two AWS Regions. They want Amazon Route 53 to automatically fail over to the secondary Region when the primary Region’s ALB is unhealthy. Health checks for the primary ALB are already configured, but the DNS record currently uses a latency-based routing policy. Which Route 53 configuration most directly provides automatic failover based on health status?

Medium
148

An orders service publishes payment instructions to an Amazon SQS Standard queue. A downstream consumer sometimes times out and retries the work, causing the consumer to process the same instruction more than once. Operationally, the team must ensure that duplicate processing does not create duplicate charges. The queue type cannot be changed. What is the most resilient application-side approach?

Medium
149

A ticket booking system runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include? The architecture review board prefers a managed AWS-native control.

Medium
150

A service processes messages from an Amazon SQS queue. Sometimes the worker finishes the business logic but does not delete the message before the visibility timeout expires, so the message is delivered again. Which two changes improve resilience and reduce the impact of duplicate processing? Select two.

Easy
151

Based on the exhibit, the team wants to stop poison messages from consuming worker capacity and also prevent duplicate side effects if the same message is delivered more than once. Which design change best meets the requirement?

Medium
152

A caching layer uses Amazon ElastiCache for Redis in front of a stateless web service. The service must continue to read cached responses during maintenance events and should automatically fail over to another node if one AZ becomes impaired. Which design change best satisfies this requirement?

Medium
153

An ECS service runs on EC2 instances and is fronted by an ALB. The ALB spans two Availability Zones, and the ECS service desired count is 2 tasks. The underlying EC2 capacity uses an Auto Scaling group (ASG) with min size set to 1, and the ASG also spans only one subnet in practice. What is the most effective change to meet the requirement that the service continues during a single-AZ instance loss?

Medium
154

A small e-commerce company runs a web application on a single Amazon EC2 instance in one Availability Zone. The instance stores session state locally and the database runs on the same instance. The company wants the application to survive an Availability Zone failure with minimal changes and no data loss for committed orders. Which combination of changes should the architect recommend?

Easy
155

A warehouse integration service must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable? The architecture review board prefers a managed AWS-native control.

Hard
156

An order-processing worker consumes messages from Amazon SQS. Occasionally, the worker times out after successfully creating a payment record but before deleting the message, which causes duplicate charges during retries. Some messages also fail validation repeatedly because required fields are missing. Which two changes should the team make? Select two.

Medium
157

A payments API uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure? The design must avoid adding custom operational scripts.

Hard
158

A content publishing system exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most? The design must avoid adding custom operational scripts.

Easy
159

A startup runs a nightly batch job on a single EC2 instance that reads a large dataset from Amazon S3, performs transformations, and writes results back to S3. The job takes about two hours, and the team wants the job to restart automatically if the instance fails or is terminated by AWS. The job is idempotent and can safely resume from the beginning. What is the MOST operationally efficient way to meet this requirement?

Easy
160

Based on the exhibit, downstream payment timeouts cause EventBridge deliveries to back up and some events are retried until they age out. What change best improves resilience and preserves events during downstream outages?

Hard
161

A company runs a stateful web application on a single Amazon EC2 instance in a public subnet. The application stores session data on the instance's root volume. The company wants to make the application highly available across two Availability Zones and ensure that session data is preserved if an instance fails. Which solution should a solutions architect recommend?

Hard
162

A healthcare company runs a critical patient-records API on Amazon EC2 instances behind an Application Load Balancer in a single AWS Region. The compliance team mandates that the API remain available even if an entire AWS Region becomes unavailable. The company wants a cost-effective solution that avoids running full production capacity in a second Region at all times. Which approach BEST meets these requirements?

Medium
163

A company runs a stateful analytics workload on EC2 instances that use EBS volumes. The data must be restorable in another Region after a major outage, with frequent point-in-time recovery. Which approach provides the most suitable replication mechanism for the EBS-backed data?

Medium
164

A company hosts a web application on EC2 instances behind an Application Load Balancer (ALB) in us-east-1. A static failover site is hosted in an S3 bucket with static website hosting enabled. The company needs automatic DNS failover to the S3 bucket if the primary ALB becomes unhealthy. Which Route 53 configuration achieves this?

Medium
165

A company runs a critical application on Amazon EC2 instances in a single Availability Zone. The application writes data to an Amazon RDS for MySQL DB instance that is not Multi-AZ. The company wants to improve the resilience of the database tier so that it can survive an Availability Zone failure with minimal downtime and no data loss. The application uses the database endpoint from the RDS console. Which solution meets these requirements?

Hard
166

A web application runs on an Amazon EC2 Auto Scaling group (ASG) behind an Application Load Balancer (ALB). The ALB is configured to use at least two Availability Zones (AZs), but the ASG currently uses subnets in only one AZ. If that AZ becomes unavailable, the application stops serving requests. Which change most directly improves resilience to an AZ outage?

Easy
167

A healthcare analytics platform ingests records into an Amazon Aurora MySQL cluster. Compliance rules require that the cluster remain writable even if an entire Availability Zone is lost, and that recovery happen without operator action. The team also wants read traffic to scale independently of the writer. Which configuration should a solutions architect choose?

Hard
168

An order-processing system publishes an event whenever a payment succeeds. Three downstream services (inventory, shipping, and analytics) must react independently. Analytics sometimes has high latency, but order processing must not be blocked. What is the best AWS approach to decouple these consumers?

Easy
169

An internal worker consumes messages from an Amazon SQS Standard queue. Recently, some messages fail validation in the worker (for example, missing required fields), causing the worker to crash before it can successfully process those messages. Those messages keep getting retried repeatedly, slowing down processing of valid messages. The team wants a resilient mechanism to quarantine bad messages after a limited number of receive attempts. What should they implement?

Medium
170

A logistics company stores shipment events in an Amazon S3 bucket. An analytics team must be able to recover any object version that is accidentally overwritten or deleted for at least 90 days, and objects must be protected from permanent deletion by any user, including the root user, during that window. Which combination of S3 features meets these requirements with the LEAST operational overhead?

Hard
171

A ticket booking system runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include? The design must avoid adding custom operational scripts.

Medium
172

A company uses an Amazon Aurora DB cluster in a Multi-AZ configuration. During a planned failover of the writer instance, the database endpoints in the application are updated incorrectly. After failover, reads work but writes fail with connection errors and timeouts for several minutes. The team currently uses the instance endpoint for the writer. What should they change to improve write resilience during failovers?

Medium
173

A claims workflow requires point-in-time recovery and accidental-delete protection for a DynamoDB table. Which two settings should the architect enable? The design must avoid adding custom operational scripts.

Hard
174

A trading dashboard stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured? The design must avoid adding custom operational scripts.

Medium
175

An order processing workflow uses Amazon SQS as the decoupling layer between a producer and a consumer Lambda function. The consumer intermittently fails due to a downstream dependency. The team has observed that certain “poison” messages keep being retried repeatedly and prevent other messages from being processed efficiently. Which SQS configuration most directly addresses this issue?

Medium
176

A startup runs a stateless web application on a single Amazon EC2 instance in one Availability Zone. The application must remain available if the instance fails or if its Availability Zone becomes unavailable. The startup wants a managed solution that requires minimal operational overhead. Which solution should a solutions architect recommend?

Easy
177

A system processes events from Amazon SQS and sometimes sees duplicate messages due to retries. The business requirement is that each payment must be charged at most once. What design choice best addresses this resiliency requirement?

Easy
178

A company hosts a web application on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer (ALB). The ALB and the Auto Scaling group are currently deployed in only one Availability Zone (AZ). The business wants the application to keep running if that AZ has an outage. What is the best change?

Easy
179

A web application runs on an Amazon EC2 Auto Scaling group behind an Application Load Balancer (ALB). After each deployment, new instances take about 2 minutes to download artifacts and become ready to accept requests on the target port. In the last deployment, the ALB started marking targets unhealthy before the app was ready, and the Auto Scaling group then replaced those instances repeatedly, causing a prolonged outage. Which change best improves resilience during instance start-up without reducing actual availability once the application is healthy?

Medium
180

An internal API is deployed in two AWS Regions behind separate Application Load Balancers. The company wants clients to use the primary Region when it is healthy and automatically switch to the secondary Region if the primary health check fails. Which two Route 53 record configurations are required? Select two.

Medium
181

A SaaS platform plans to run in two AWS Regions for lower latency. The team wants to enable active-active writes (both regions accept updates) to avoid failover downtime. However, the business requires strong consistency for order status transitions (for example, only one transition from “Paid” to “Shipped” must be allowed). Which statement is the best architectural choice to meet the consistency requirement?

Medium
182

A consumer application reads from an Amazon SQS queue. Some messages have an invalid format and always fail processing. They are retried repeatedly and consume consumer capacity. What is the best way to prevent these "poison pill" messages from blocking normal processing?

Easy
183

A healthcare company stores patient records in an Amazon DynamoDB table. The table must be recoverable to any point within the last 35 days, and the data must remain available if an entire AWS Region becomes unavailable. Which two actions should a solutions architect take to meet these requirements? (Choose two.)

Medium
184

A company is building a serverless application that processes messages from an Amazon SQS queue using AWS Lambda. The application must not lose messages and must handle occasional downstream failures gracefully. The Lambda function sometimes fails due to a transient error in a downstream service. The company wants to ensure that failed messages are retried and eventually processed, but also wants to avoid infinite retries that could block the queue. What should the company do?

Hard
185

A healthcare analytics team runs a containerized reporting service on Amazon ECS with the Fargate launch type in a single Availability Zone. The service must remain available if one Availability Zone fails, and it must scale automatically based on CPU utilization. The tasks are stateless and write output to Amazon S3. Which configuration should a solutions architect implement?

Medium
186

Based on the exhibit, the web tier becomes unavailable if us-west-2a has an outage. What is the best change to improve resilience with the least redesign?

Easy
187

A warehouse integration service must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The design must avoid adding custom operational scripts.

Hard
188

A media company stores finalized video masters in an Amazon S3 bucket in the us-east-1 Region. Compliance requires that the objects be recoverable if they are accidentally deleted or overwritten for at least 90 days, and that no user, including administrators, be able to permanently erase them during that period. Which S3 feature should the solutions architect enable?

Easy
189

Based on the exhibit, the database is manually promoted during an Availability Zone failure and the application outage lasts longer than the target. What change best improves resilience with the least operational intervention?

Hard
190

An application writes to an Amazon Aurora DB cluster. After a planned Aurora failover, the application experiences several minutes of connection errors. The logs show the application continues connecting to the specific DB instance endpoint that was the primary before the failover. What change most directly improves resilience during Aurora failovers?

Medium
191

A company uses Amazon RDS with automated backups enabled (retention period: 7 days). At 10:30 UTC, a bad release corrupts specific rows in a production table. The team detects the issue at 11:10 UTC. They need to revert the database state to what it was from 10:00–10:30 UTC, recover quickly, and minimize risk to the currently running workload. What is the best option?

Medium
192

A service processes customer payments from a message queue. Because the queue provides at-least-once delivery, the same payment message can be delivered more than once if the consumer times out before committing its state. Currently, the service sometimes charges the customer twice. Which design change most directly prevents duplicate charges while still allowing safe retries?

Medium
193

A media company stores original uploads in an S3 bucket. They must recover from accidental overwrites/deletes and also recover quickly from a full Region outage. The required RPO is about 1 hour. Which configuration best meets these requirements?

Medium
194

A warehouse integration service must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable? The design must avoid adding custom operational scripts.

Hard
195

A content publishing system uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The team wants the control to be enforceable during normal operations.

Medium
196

A SaaS application is deployed in us-east-1 and us-west-2 behind separate ALBs. The business wants DNS to send new clients to the primary Region when it is healthy and automatically fail over to the secondary Region when the primary endpoint is unhealthy. Which two Route 53 settings are required? Select two.

Medium
197

A logistics company runs an order-tracking API on a fleet of EC2 instances in a single Availability Zone behind a Network Load Balancer. The architecture team must make the API resilient to the loss of that Availability Zone without changing the API endpoint that clients already use. The instances are stateless and store session data in a shared Amazon ElastiCache cluster. Which change should the solutions architect make to meet these requirements?

Medium
198

A payments API uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure? The architecture review board prefers a managed AWS-native control.

Hard
199

A patient portal receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers? The design must avoid adding custom operational scripts.

Medium
200

A media company stores original video masters in an Amazon S3 bucket in the us-east-1 Region. Compliance requires that a readable copy of every object exist in the eu-west-1 Region within 15 minutes of upload, and that the objects in eu-west-1 be usable directly by an application there. No transformations are required. Which S3 feature should the solutions architect enable?

Medium
201

A regional web application for a content publishing system must fail over automatically to a secondary Region if the primary endpoint becomes unhealthy. Which two services or features are required? The design must avoid adding custom operational scripts.

Hard
202

A logistics company runs an order-tracking service that exposes a REST API. The service must remain available during a single Availability Zone failure and must keep read latency low for a globally distributed user base. The data store must support automatic multi-AZ replication without the team managing database servers. Which solution meets these requirements?

Medium
203

Your web application is deployed in two AWS Regions (Region A and Region B). You want Route 53 to automatically fail over DNS traffic from Region A to Region B when Region A is unhealthy. The failover decision must be based on health checks that verify whether the application in Region A is reachable. Which Route 53 routing configuration best meets these requirements?

Medium
204

A media company stores original video masters in an Amazon S3 bucket in the us-east-1 Region. Compliance requires that a readable copy of every object exists in eu-west-1 within 15 minutes of upload, and that objects deleted in the source bucket do not automatically disappear from the destination. Which S3 feature should the solutions architect enable?

Medium
205

A media processing company runs a stateless thumbnail-generation fleet on Amazon EC2 instances behind an Application Load Balancer. The instances store no local state, and the team wants the fleet to survive the loss of an entire Availability Zone without manual intervention. The fleet must also scale out automatically based on CPU. Which combination of AWS services should the solutions architect use to meet these requirements with the LEAST operational overhead?

Medium
206

A trading dashboard uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The design must avoid adding custom operational scripts.

Medium
207

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer. The application must be highly available and able to withstand the failure of an entire AWS Region. The company wants to minimize operational overhead and ensure that failover is automatic. Which solution should a solutions architect recommend?

Medium
208

A financial services firm runs a stateful trading application on EC2 instances in an Auto Scaling group. Each instance maintains an in-memory cache that takes several minutes to rebuild after a restart, and the team wants the application to survive the loss of an Availability Zone with minimal disruption. The application cannot be made stateless in the near term. Which approach should a solutions architect recommend?

Hard
209

An application uses an Amazon Aurora DB cluster. The cluster performs an automatic failover from the writer instance to a standby instance. After failover completes, reads succeed, but all new writes fail with errors indicating the application is connecting to the old writer endpoint. Which change best fixes the resiliency issue after failover?

Medium
210

A company needs to store application logs in a durable and highly available manner. The logs are written continuously by multiple EC2 instances and are accessed infrequently for compliance audits. The company wants a solution that provides 99.999999999% (11 9's) durability and automatically replicates data across multiple Availability Zones. Which AWS service should the company use?

Easy
211

Based on the exhibit, the application should continue serving requests if one Availability Zone fails. Which change best improves resilience with the least operational complexity?

Medium
212

A startup runs a stateless web tier on Amazon EC2 instances in an Auto Scaling group that spans three Availability Zones. The team wants the application to keep serving requests even if one instance becomes unresponsive, without operator involvement. What should the solutions architect configure?

Easy
213

A startup runs a customer-facing web application on a single Amazon EC2 instance in one Availability Zone, with the database on the same instance. The founders want the application to survive the failure of that Availability Zone with minimal changes and no server management for the database tier. Which action should the solutions architect take first?

Easy
214

Based on the exhibit, the payment worker sometimes processes the same SQS Standard message more than once after a timeout. What change best prevents duplicate charges while keeping the queue architecture?

Medium
215

A production Amazon RDS database has automated backups enabled. At 10:00 UTC, an application deploy accidentally overwrote a subset of rows due to a faulty migration. The issue is detected at 10:45 UTC. The team confirms that the required retention window is still available. Which approach offers the most resilient and least disruptive way to recover the affected data close to the time of the event?

Medium
216

A company runs an Amazon Aurora DB cluster with a Multi-AZ deployment. The application is configured with a hard-coded endpoint that points to the current writer *DB instance* (an instance-specific endpoint), rather than the Aurora cluster writer endpoint. During an unexpected AZ failure, Aurora promotes the standby to become the new writer. However, the application continues to fail to connect until an operator updates the hard-coded endpoint. What change most directly improves resiliency so the application automatically reconnects after failover?

Medium
217

Based on the exhibit, duplicate payment charges occasionally occur when the worker times out after the charge is submitted but before the message is deleted. What change best prevents duplicate charges while keeping retry behavior?

Hard
218

A team wants a web application to keep serving traffic if one Availability Zone fails. Match each architecture element to the resilience behavior it provides.

Medium
219

A ticket booking system uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The team wants the control to be enforceable during normal operations.

Medium
220

A logistics company runs an order-tracking service on Amazon EC2 instances that write state to an Amazon DynamoDB table. A recent incident showed that a single Availability Zone failure caused the service to lose capacity, and the team also discovered that a developer accidentally deleted a production table. The architect must improve both Availability Zone resilience and protection against accidental table deletion. (Choose two.)

Medium
221

A warehouse integration service must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The team wants the control to be enforceable during normal operations.

Hard
222

A team uses an S3 bucket to store important customer-generated exports. They need protection against accidental overwrites and also want copies of the data in another AWS Region for disaster recovery. Which S3 configuration best satisfies both requirements?

Easy
223

A media company runs a video-transcoding fleet on Amazon EC2 instances that read source files from an Amazon S3 bucket and write output to a second bucket. The fleet is spread across three Availability Zones in one Region, and instances are launched by an Auto Scaling group. The company needs the architecture to survive the loss of an entire Availability Zone without losing in-flight transcoding work or requiring manual intervention. Which combination of design elements should a solutions architect implement to meet these requirements?

Medium
224

A company is designing a disaster recovery plan for a critical application hosted on AWS. The application runs on EC2 instances with data stored in Amazon EBS volumes and Amazon S3. The recovery time objective (RTO) is 15 minutes, and the recovery point objective (RPO) is 1 hour. Which three strategies would help meet these objectives? (Choose three.)

Medium
225

A service consumes messages from an SQS queue. Recently, a new message format started failing validation in the consumer. The consumer catches the exception but cannot successfully process those messages without code changes. The team wants failed messages to be isolated for later investigation instead of being retried indefinitely. What should they configure?

Medium
226

A production Amazon RDS database has automated backups enabled with sufficient retention. At 10:30 UTC, a release corrupts specific rows. The issue is detected at 10:45 UTC. The team wants to restore the database state to before the corruption with minimal complexity. What should they do?

Easy
227

A Multi-AZ Amazon RDS database experiences incorrect writes at 10:15 UTC due to a buggy release. The team detects the problem at 10:25 UTC. They want to restore the data to a known-good point around 10:15 UTC, and validate the recovered data, without taking the current production instance offline during the recovery process. What is the most appropriate AWS action?

Medium
228

A financial analytics platform ingests events into an Amazon Kinesis Data Stream with four shards. During month-end peaks, producers receive ProvisionedThroughputExceededException errors and consumers fall behind. The architects want to increase capacity without changing producer code and must preserve the order of records that share the same partition key. What should they do?

Hard
229

A production team accidentally deletes critical rows in an Amazon RDS for PostgreSQL database. The deletion occurred about 6 hours ago. The team wants to recover to a specific point in time with minimal disruption. Assuming automated backups are enabled, which approach provides the best resilience outcome?

Medium
230

A company wants a disaster recovery setup for a web application. They want to keep costs low but still recover within a couple of hours after a regional disruption. They are willing to run only minimal infrastructure in the secondary location and scale it up during the outage. Which DR approach best matches this requirement?

Easy
231

A financial services company is designing a new payment processing platform. The platform must continue to accept and process transactions even if an entire AWS Region becomes unavailable, and it must not lose any accepted transaction. The architects have decided to run active-active deployments in two Regions and use Amazon Route 53 for traffic management. Which two additional design elements are required to meet the durability and availability goals? (Choose two.)

Hard
232

A inventory service exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most? The design must avoid adding custom operational scripts.

Easy
233

A financial services company runs a critical application on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer. The application must be able to survive the failure of an entire AWS Region. The company wants a cost-effective solution that minimizes operational overhead. Which approach should the architect recommend?

Hard
234

Based on the exhibit, which Route 53 configuration should be used so traffic automatically returns to the secondary Region only when the primary Region becomes unhealthy?

Medium
235

A web application runs on an Auto Scaling group (ASG) behind an Application Load Balancer (ALB). After a new release, instances begin failing ALB health checks with errors like 502 while the application is still starting up. CloudWatch shows that the ASG replaces the instances before they finish initializing, so traffic never reaches healthy targets. Which change most directly prevents premature replacement during startup so traffic can resume as soon as the instances are actually healthy?

Medium
236

A company is designing a multi-Region disaster recovery (DR) strategy for a stateless web application running on Amazon EC2 instances behind an Application Load Balancer (ALB). The application uses an Amazon RDS for MySQL database as its data store. The architecture must provide rapid failover with the lowest possible Recovery Point Objective (RPO) and Recovery Time Objective (RTO). Which of the following design choices will help achieve these objectives? (Choose four.)

Medium
237

A startup runs a stateless image-resizing API on a fleet of EC2 instances behind an Application Load Balancer. The instances store uploaded source images on their own instance store volumes before processing. During a routine scale-in event, an instance was terminated and several in-flight uploads were lost. The architect must make the design resilient to instance loss without changing the API code. What should the architect do?

Easy
238

A inventory service uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The design must avoid adding custom operational scripts.

Medium
239

A company runs a customer portal on an Amazon Aurora PostgreSQL cluster. The application currently connects directly to the writer instance endpoint and keeps long-lived connections open. During a maintenance failover, writes fail until clients are restarted. The team wants the application to reconnect to the correct Aurora endpoint automatically and reduce user-visible write interruptions. Which change is most likely to achieve this?

Medium
240

A developer accidentally deletes important rows in an RDS database. The mistake is discovered 45 minutes later. The database has automated backups enabled with a retention period of 7 days. What is the best way to restore the database to a point just before the deletion?

Medium
241

A payments API uses an RDS MySQL database and must remain available during an Availability Zone failure with minimal application changes. What should the architect enable?

Medium
242

You host a public API using Amazon API Gateway in two AWS Regions: us-east-1 (primary) and us-west-2 (secondary). You want Route 53 to send client traffic to the secondary region only when the primary API is unhealthy. Which Route 53 setup best meets this requirement?

Medium
243

An ECS service runs on EC2 instances and is fronted by an ALB. The ALB spans two Availability Zones, and the ECS service desired count is 2 tasks. The underlying EC2 capacity uses an Auto Scaling group (ASG) with min size set to 1, and the ASG also spans only one subnet in practice. What is the most effective change to meet the requirement that the service continues during a single-AZ instance loss?

Medium
244

A solutions architect is designing a highly available relational database tier for a customer-facing order system that must survive the loss of an entire Availability Zone with minimal administrative effort and no application connection-string changes during failover. (Choose two.)

Medium
245

A ticket booking system uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The design must avoid adding custom operational scripts.

Medium
246

A healthcare data platform stores patient documents in an Amazon S3 bucket in us-east-1. Regulations require that the data remain readable with low latency even if the entire us-east-1 Region becomes unavailable, and that writes continue in a secondary Region. The team wants object-level replication with minimal operational overhead and must preserve version history. Which solution BEST meets these requirements?

Hard
247

Based on the exhibit, an administrator accidentally deleted data from Amazon RDS for PostgreSQL about 90 minutes ago. Which recovery approach best restores the database to the exact required point in time?

Medium
248

A warehouse integration service receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers?

Medium
249

A financial analytics platform runs a stateless API on Amazon EC2 instances in an Auto Scaling group behind a Network Load Balancer. The API reads from an Amazon Aurora MySQL cluster that has a single writer instance and one reader instance in a different Availability Zone. During a recent Availability Zone event, the writer instance failed and the API saw several minutes of failed writes. The team wants writes to resume automatically with minimal downtime and no application code changes. What should the solutions architect do?

Hard
250

An internal API is hosted in two AWS Regions behind Route 53. Under normal conditions, clients should use the primary region. If the primary endpoint becomes unhealthy, traffic must automatically switch to the secondary region. Which Route 53 setup best meets this requirement?

Easy
251

An orders service currently sends HTTP requests directly to two downstream services (inventory and shipping). During peak load, inventory slows down, causing the orders service to slow as well. The team wants the orders service to remain responsive even when a downstream service is temporarily slow or restarted. Which design change best achieves this resiliency goal?

Easy
252

A content publishing system uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The design must avoid adding custom operational scripts.

Medium
253

A healthcare company runs a batch ingestion pipeline on Amazon EC2 instances that read messages from an Amazon SQS queue and write results to Amazon DynamoDB. The pipeline must be resilient so that a single instance failure does not stop processing and no messages are lost. Which two architectural changes should a solutions architect make to meet these requirements? (Choose two.)

Medium
254

A media startup stores user-uploaded video files in an Amazon S3 bucket in the us-east-1 Region. The compliance team requires that the data remain recoverable if an entire AWS Region becomes unavailable, and that recovery can be performed by pointing applications at a different endpoint. Cost should be minimized while still meeting the requirement. Which solution should a solutions architect recommend?

Easy
255

Based on the exhibit, the current disaster recovery design misses the RTO target even though the database replica is current. Which deployment model best meets the requirements with the least always-on cost?

Hard
256

A retail platform needs disaster recovery across AWS Regions. The business requirement is: RTO up to 6 hours, RPO up to 1 hour, and they want the ability to start serving quickly during a Region outage but do not want to run full production capacity continuously. Which DR strategy best fits these requirements?

Easy
257

A media company stores original video assets in an Amazon S3 bucket in the us-east-1 Region. Editors in Europe report slow downloads, and the legal team requires that a copy of every asset exist in eu-west-1 within 15 minutes of upload, with the ability to fail over reads to the European copy during a Regional impairment. Which S3 feature should the architects enable?

Medium

Frequently asked questions

What does the Design Resilient Architectures domain cover on the SAA-C03 exam?
High availability and resilience questions test multi-AZ vs multi-Region patterns, Auto Scaling, load balancing and the right service for a given recovery time objective.
How many questions are in this domain?
This page lists all 257 Design Resilient Architectures questions in the SAA-C03 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Design Resilient Architectures questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
aws-saa AWS-SAA design resilient Practice Questions