SAA-C03 · domain
Design Resilient Architectures
Use this page to practise high availability and resilience questions. The SAA-C03 exam tests your ability to match an architecture pattern to an RTO/RPO requirement — know the cost and recovery time of each pattern.
Focused practice
Practice Design Resilient Architectures questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Design Resilient Architectures
High availability and resilience questions test multi-AZ vs multi-Region patterns, Auto Scaling, load balancing and the right service for a given recovery time objective.
Multi-AZ vs multi-Region deployment trade-offs.
Auto Scaling policies and when to scale horizontally vs vertically.
Elastic Load Balancing: ALB, NLB, CLB and their use cases.
RTO and RPO targets matched to the correct AWS architecture.
Watch out for
Common Design Resilient Architectures exam traps
- ▸Multi-AZ protects against AZ failure; multi-Region protects against Region failure.
- ▸Auto Scaling does not guarantee zero downtime without a load balancer.
- ▸ALB operates at Layer 7; NLB operates at Layer 4.
- ▸Pilot light is cheaper than warm standby but has longer recovery time.
Question index
All Design Resilient Architectures questions (257)
Click any question to see the full explanation, or start a practice session above.
A production Amazon RDS database already has automated backups enabled. At 10:45 UTC, the team discovers that a faulty migration corrupted rows in a table at 10:30 UTC. The business wants the database restored to exactly the state it had at 10:30 UTC with minimal risk. Which two actions should the team take? Select two.
Medium2Based on the exhibit, the web team wants the application to continue serving traffic if one Availability Zone fails. Which change best meets the requirement with the least operational overhead?
Easy3A trading dashboard uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The architecture review board prefers a managed AWS-native control.
Medium4An internal worker consumes messages from an Amazon SQS queue. Occasionally, a message fails validation in the worker (for example, missing required fields). Reprocessing the same bad message repeatedly wastes processing time and delays healthy messages. What is the best AWS approach to handle these poison messages without blocking the rest of the queue?
Easy5A production application uses an Amazon RDS Multi-AZ DB instance. During an unplanned failover, the database endpoint remains the same. What change should the application team make to handle the failover reliably?
Easy6Based on the exhibit, a web application must stay available if one Availability Zone fails. What is the best change to improve resilience?
Easy7An order-processing service consumes messages from an Amazon SQS Standard queue using a custom worker. During traffic spikes, the worker occasionally times out after performing some work but before acknowledging the message, so SQS redelivers it and it may be processed again. You also observe that a small set of “poison” messages always fail validation. What change most directly improves resilience by (1) preventing poison messages from retrying indefinitely and (2) avoiding duplicate side effects caused by legitimate retries?
Medium8A healthcare provider hosts a patient-records API on Amazon EC2 instances in a single Availability Zone behind an Application Load Balancer. An audit finds the architecture cannot tolerate the loss of that Availability Zone. Budget is limited, and the API reads from an Amazon Aurora MySQL cluster that currently has one writer instance and no replicas. Which change most effectively addresses the audit finding?
Medium9A company runs a stateless API on Amazon EC2 instances in a single Availability Zone behind an Application Load Balancer. The ALB currently has a listener on port 80 only. The company wants the API to remain available if the single Availability Zone fails. What should the solutions architect do to meet this requirement with the LEAST operational overhead?
Easy10A web application runs on an EC2 Auto Scaling group (ASG) behind an Application Load Balancer (ALB). The ASG spans three Availability Zones. After a deployment, new instances frequently fail the ALB target group health checks with HTTP 5xx responses and are quickly terminated by the ASG. What change most improves resiliency during deployments with minimal downtime by preventing premature removal of instances that are still starting?
Medium11Based on the exhibit, the application sees several minutes of connection errors during an Aurora failover. What is the best change to reduce failover impact?
Medium12A patient portal must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The architecture review board prefers a managed AWS-native control.
Hard13A healthcare analytics platform stores derived datasets in an Amazon S3 bucket. Regulatory rules require that every object remain recoverable for 90 days after creation even if an application bug issues a delete, and that no object version be permanently destroyed during that window. The team wants the strongest protection with the least custom code. Which S3 feature should the solutions architect enable?
Hard14A company runs a critical two-tier web application on AWS. The web tier consists of Amazon EC2 instances behind an Application Load Balancer (ALB) in a single Availability Zone. The database tier is an Amazon RDS for MySQL DB instance in the same Availability Zone. A recent power outage in that Availability Zone caused a full application outage. The company wants to redesign the architecture to survive an Availability Zone failure with minimal operational overhead. Which solution meets these requirements?
Medium15An orders service publishes payment instructions to an Amazon SQS Standard queue. The downstream processor sometimes times out after it has already applied the payment, but before it can delete the message from the queue. As a result, the same payment instruction can be processed more than once. The team wants the strongest way to prevent duplicate side effects while keeping the system decoupled. What should they implement?
Medium16Your order-processing system uses EventBridge rules to send events to a Lambda function that updates order status. Over the last week, some events fail with a transient database timeout, and the Lambda retries intermittently but then the events are lost (no alerts after failures). You want at-least-once processing, bounded retries, and a way to inspect unprocessable events for later reprocessing. Which architecture change best meets these requirements?
Medium17A inventory service exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most? The architecture review board prefers a managed AWS-native control.
Easy18A ticket booking system stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured?
Medium19A retail API runs on Amazon EC2 instances behind an Application Load Balancer and stores orders in an Amazon RDS for PostgreSQL database. A test that stopped one Availability Zone caused the API to return errors because all application servers were in the same AZ and the database was single-AZ. Which two changes should the architect make to continue serving traffic during a single-AZ failure? Select two.
Medium20Match the disaster recovery strategy to the recovery posture it best fits for a Regional outage.
Medium21A payments service receives payment orders by consuming messages from an Amazon SQS Standard queue. The downstream processor occasionally exceeds its processing timeout. As a result, some messages reappear in the queue and may be processed more than once. The team wants to prevent duplicate side effects (for example, double-charging) and also ensure poison messages do not repeatedly consume processing capacity. What approach best satisfies both goals?
Medium22A financial services firm runs a batch settlement job on a fleet of Amazon EC2 instances that pull work from an Amazon SQS queue. The job must not lose messages if an instance is terminated mid-processing, and duplicate processing must be minimized because each settlement charge is expensive. The team also wants to avoid indefinite reprocessing of a message that repeatedly fails. Which two changes should the solutions architect make to meet these requirements? (Choose two.)
Hard23A company runs a stateful workload on Amazon EC2 instances in an Auto Scaling group. The workload writes session data to the instance store and to an Amazon EBS volume attached at launch. The company wants the workload to survive an Availability Zone failure without losing session data. What should the solutions architect do?
Hard24A fintech company has a two-Region DR requirement: RPO must be within 15 minutes and RTO must be under 2 hours. To control cost, they do not want to run full production infrastructure in the secondary Region continuously. They plan to continuously replicate the database and keep the application infrastructure in the secondary Region prepared, but at reduced capacity. Which DR strategy best matches this requirement and accurately describes their plan?
Medium25A media company stores generated video thumbnails in an Amazon S3 bucket. The bucket currently uses the S3 Standard storage class, and the objects are accessed frequently for the first 30 days and then almost never. The company wants to reduce storage costs automatically without changing the application and must retain the objects for at least one year. Which action should a solutions architect take?
Easy26A logistics company runs an order-processing workflow using AWS Step Functions. A task state invokes a Lambda function that charges customer credit cards through a third-party gateway. Occasionally the gateway times out, and the workflow fails even though the charge may have succeeded. The architect must make the workflow resilient to these transient failures and avoid duplicate charges. (Choose two.)
Medium27A healthcare company runs a stateless patient-intake API on a fleet of Amazon EC2 instances in a single VPC. The compliance team requires the workload to survive the complete loss of one Availability Zone with no manual intervention, and the instances must be replaced automatically if they fail health checks. The application stores no local state and writes all data to Amazon RDS. Which approach meets these requirements with the LEAST operational effort?
Medium28A payments API uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure?
Hard29A content publishing system exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most?
Easy30A company runs an application behind an Application Load Balancer (ALB). An Auto Scaling group (ASG) is configured with desired capacity 2, but it is attached only to subnets in a single Availability Zone. The ALB is healthy because it is configured across multiple Availability Zones. When the Availability Zone that contains the ASG subnets experiences an outage, what change most directly improves resilience and allows capacity to be restored automatically?
Medium31A logistics company runs an order-processing workload that reads messages from an Amazon SQS queue and writes results to an Amazon DynamoDB table. Occasionally the same order is processed twice and produces duplicate shipments. The architects must ensure each order is processed exactly once end to end, while keeping throughput as high as possible. What should they do?
Hard32A global application experiences frequent writes and must survive a full Regional outage with near-zero data loss. The product team also requires that users can continue to write during the incident using the closest Region. Which approach is most aligned with these requirements?
Medium33A patient portal receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers?
Medium34An engineering team deploys a stateless web API on EC2 using an Auto Scaling group and an Application Load Balancer (ALB). During a recent test, they noticed that when one Availability Zone was unavailable, traffic failed until new instances were manually launched. Which change most directly improves automatic failover for the compute layer within a single Region?
Easy35A payments API requires point-in-time recovery and accidental-delete protection for a DynamoDB table. Which two settings should the architect enable? The team wants the control to be enforceable during normal operations.
Hard36A healthcare company runs a containerized claims-processing service on Amazon ECS with the Fargate launch type in a single AWS Region. The service must survive the loss of an entire Availability Zone with no manual intervention, and the architecture must keep the same service endpoint for callers. The service is fronted by an Application Load Balancer. Which combination of actions should a solutions architect take to meet these requirements with the LEAST operational overhead?
Medium37A claims workflow uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure?
Hard38An orders service publishes payment instructions to an Amazon SQS queue. After occasional processing timeouts, the downstream consumer sometimes processes the same instruction twice, resulting in duplicate payment attempts. The team currently uses an SQS Standard queue with a visibility timeout of 2 minutes and relies on the consumer to finish before the timeout expires. What approach best improves resilience against duplicate processing?
Medium39An orders system sends payment instructions to an Amazon SQS queue. The consumer sometimes times out after it has already created the payment record but before it deletes the SQS message. As a result, the same instruction can be processed more than once. Which design best ensures the consumer remains resilient and does not create duplicate payments when the same instruction is delivered multiple times?
Medium40A logistics company runs an order processing system on Amazon EC2 instances that read and write to an Amazon RDS for MySQL database. The database is currently a Single-AZ deployment. The company needs the database to survive an Availability Zone failure with automatic failover and minimal downtime. The application connects using a hardcoded DNS name. Which change should a solutions architect make?
Hard41A customer portal must recover from a regional outage within a few hours. The business wants lower ongoing cost than a fully active second Region and does not want to rebuild everything from scratch during the outage. Which two DR patterns best fit that goal? Select two.
Medium42A media company stores master video files in an Amazon S3 bucket in the us-east-1 Region. A compliance policy requires that the data remain readable even if the entire us-east-1 Region becomes unavailable, and the recovery point objective is 15 minutes. The team wants the lowest operational overhead and does not want to modify application code. Which solution should the architect implement?
Hard43A financial analytics platform runs an Amazon Aurora MySQL cluster with one writer and two readers. During month-end reporting, read traffic spikes and the application sometimes receives TooManyConnections errors on the reader endpoint. The architect wants to absorb bursts without changing application code and must keep failover behaviour intact. Which change meets these requirements?
Hard44A logistics company runs a shipment-tracking service on a single Amazon EC2 instance in one Availability Zone. The instance stores tracking state in an attached Amazon EBS volume and writes nightly backups to Amazon S3. The company needs the service to survive the loss of an entire Availability Zone with minimal data loss and automatic recovery, while keeping changes minimal. Which design change should a solutions architect recommend?
Medium45A production Amazon RDS database has automated backups enabled. At 10:45 UTC, an issue is discovered. The team needs to restore the database to its state as of 10:30 UTC. Which capability should they use?
Easy46A solutions architect is designing a highly available and resilient architecture for a critical internal application that processes financial transactions. The application runs on Amazon EC2 instances inside an Auto Scaling group. The database layer uses an Amazon Aurora MySQL cluster. The company requires that if an entire AWS Availability Zone (AZ) fails, the application must remain operational with minimal impact and automatically recover without manual intervention. Which combination of architectural decisions will meet these requirements? (Choose four.)
Medium47A startup runs a single Amazon EC2 instance hosting both a web application and its MySQL database. The founders want the application to survive the failure of the underlying hardware without changing the database engine, and they want the smallest possible operational change. What should the architect recommend?
Easy48Your public API is hosted in two regions. You want Route 53 to automatically send traffic to the secondary region when the primary region’s endpoint fails. The primary API health check is returning failure codes, but clients still reach the primary region for several minutes. Which Route 53 configuration most directly addresses this behavior?
Medium49Based on the exhibit, some SQS messages fail validation repeatedly and continue consuming worker time. What change best prevents the bad messages from being retried forever?
Easy50A patient portal receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers? The architecture review board prefers a managed AWS-native control.
Medium51A team runs an Amazon RDS for MySQL database in a single Availability Zone. They want automatic failover with minimal downtime if the primary database instance becomes unavailable. Automated backups are already enabled. Which configuration change best meets the requirement?
Easy52A startup runs a nightly batch job on a single Amazon EC2 instance that stores results in an Amazon EBS volume. The job takes six hours, and the team wants to resume from the last completed step if the instance is terminated unexpectedly. Which approach provides the required durability with the least operational effort?
Easy53A company needs an Amazon RDS database that automatically fails over to a standby when the primary DB instance becomes unavailable. Which approach best meets the requirement with minimal operational effort?
Easy54A SaaS platform serves an API using two regional deployments: us-east-1 (primary) and us-west-2 (secondary). Each region has its own ALB. The business requires automated DNS-based failover when the primary region becomes unhealthy, and they do not want manual DNS changes during incidents. Which Route 53 configuration is the best match?
Medium55Based on the exhibit, DNS still sends traffic to the primary Region even though Route 53 health checks show the primary endpoint is unhealthy. What is the best change to make failover work as intended?
Hard56A patient portal must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used?
Hard57A claims workflow uses an RDS MySQL database and must remain available during an Availability Zone failure with minimal application changes. What should the architect enable?
Medium58A ticket booking system stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured? The architecture review board prefers a managed AWS-native control.
Medium59A developer accidentally corrupts part of a production Amazon RDS database, and the issue is discovered 45 minutes later. The team needs to restore the database to the state immediately before the change. Which two actions should be part of the recovery plan? Select two.
Easy60Your media processing pipeline writes original uploads to an S3 bucket and later generates derivative files. An operator accidentally deletes a subset of original uploads in production. You need to (1) restore the deleted objects with minimal data loss and (2) protect against both regional disasters and future operator mistakes. The company requires recovery even if objects are deleted and later overwritten. What is the most effective change to meet these requirements?
Medium61A patient portal must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The design must avoid adding custom operational scripts.
Hard62A healthcare provider runs a patient-record API on Amazon EC2 instances behind an Application Load Balancer in one AWS Region. The API reads from an Amazon RDS for MySQL DB instance. The provider must be able to continue serving read traffic if the primary database instance fails, and must minimize the time the application is unavailable. Which change should a solutions architect make?
Medium63An organization hosts the same public API in two AWS Regions. Normal traffic should go to the primary Region. If the primary endpoint becomes unhealthy, Route 53 should automatically route users to the secondary Region. What is the best Route 53 configuration approach?
Easy64Based on the exhibit, the web application must remain available even if one Availability Zone fails. What is the best change to improve resilience with the least redesign?
Medium65A logistics firm runs an order-processing service that reads from an Amazon SQS queue and writes results to an Amazon DynamoDB table. During a marketing event, the consumer fleet scaled out aggressively and DynamoDB began returning ProvisionedThroughputExceededException errors, causing messages to be retried and some orders to be processed twice. The architects want to absorb traffic spikes without overprovisioning capacity and without duplicate processing. Which combination of changes should they make?
Hard66A warehouse integration service must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable?
Hard67A team needs a relational database solution that can automatically fail over to a standby instance if the primary database becomes unavailable. They want the standby to be located in a different Availability Zone. Which RDS/Aurora configuration best satisfies this requirement?
Easy68A healthcare analytics platform processes streaming records with an AWS Lambda function that writes results to an Amazon DynamoDB table. The pipeline must not lose records if the function throws an error, and the operations team wants to inspect and reprocess failed records without writing custom retry code. Which approach should the solutions architect use?
Hard69A financial services firm runs a critical API on Amazon EC2 instances behind a Network Load Balancer. The API must handle a sudden loss of one Availability Zone and continue serving traffic with no manual failover. The instances are in an Auto Scaling group that currently uses a single subnet in one Availability Zone. Which change should the architect make?
Medium70A trading dashboard runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include? The architecture review board prefers a managed AWS-native control.
Medium71An event-driven order processing service consumes messages from an Amazon SQS Standard queue. After a deployment, about 1% of messages start failing validation because a required field is missing. The consumer catches the exception and returns control, so the messages are retried. However, those poison messages keep reappearing and repeatedly consuming processing time for hours, delaying handling of valid messages. What is the most resilient way to handle the poison messages while keeping the system available?
Medium72A claims workflow uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure? The architecture review board prefers a managed AWS-native control.
Hard73An application uses an Amazon RDS Multi-AZ DB instance. During a failover test, connections fail until the application is restarted, even though the database comes back online. Which two changes should the team make to improve resilience during failover? Select two.
Medium74An Auto Scaling group behind an Application Load Balancer frequently replaces new EC2 instances. The application needs ~6 minutes to warm up after instance launch. However, the ALB target group health checks start immediately and mark the targets unhealthy until the application is ready. Because the targets become unhealthy early, the Auto Scaling group then terminates the instances and launches replacements, creating a repeated unhealthy/termination loop. What configuration change will most directly improve recovery by preventing premature ASG termination while the application is warming up?
Medium75A company runs an internet-facing API in two AWS Regions. Route 53 currently uses simple routing to a primary Application Load Balancer (ALB) DNS name. When the primary Region experiences an outage, customers wait a long time because the DNS entry is not changed automatically. The team wants automatic failover: if the primary Region ALB health check fails for a sustained period, Route 53 should route users to the secondary Region ALB. Which Route 53 approach best meets this requirement?
Medium76A healthcare company needs to store patient records in Amazon DynamoDB. The records must be highly available and durable across multiple Availability Zones. The company also requires the ability to recover the table to any point in time within the last 35 days in case of accidental writes or deletions. Which solution meets these requirements?
Medium77A company stores critical documents in an Amazon S3 bucket in the us-east-1 Region. The documents must survive an unlikely loss of the entire us-east-1 Region. The company wants a recovery point objective (RPO) of 15 minutes and a recovery time objective (RTO) of 1 hour. What should the solutions architect recommend?
Medium78A healthcare company runs a patient portal on Amazon EC2 instances behind an Application Load Balancer across two Availability Zones. A new compliance rule requires that if an entire Availability Zone fails, the portal must remain available with no manual intervention. The EC2 instances are stateless and store no session data. Which design change should the architect implement to meet this requirement?
Medium79A patient portal must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable? The team wants the control to be enforceable during normal operations.
Hard80A team accidentally updates critical rows in an Amazon RDS for PostgreSQL database. Automated backups are enabled. They need to recover the data to the exact state as of 90 minutes ago. They also cannot risk interrupting the current production database instance while investigators validate the restored data. Which recovery strategy best meets these constraints?
Medium81Order the steps to create a static website using Amazon S3 and CloudFront.
Medium82A logistics company runs a stateless order-tracking API on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer. The architect must ensure the API survives the loss of an entire Availability Zone and that unhealthy instances are replaced automatically. (Choose two.)
Medium83A company runs a critical API on Amazon EC2 behind an Application Load Balancer in a single AWS Region. The business requires the API to keep serving traffic if an entire Availability Zone becomes unavailable, and the recovery must not depend on any manual step. The database is Amazon RDS for PostgreSQL configured as a Single-AZ instance. Which combination of changes should a solutions architect implement to meet these requirements?
Hard84A SaaS provider runs a multi-tenant application on Amazon EC2 instances behind an Application Load Balancer. Tenants are identified by a subdomain, and each tenant's data is stored in a separate Amazon S3 bucket. The provider wants HTTPS with a single certificate, automatic renewal, and the ability to add new tenant subdomains without redeploying or replacing the certificate. Which solution meets these requirements?
Hard85Based on the exhibit, the database must fail over automatically if the primary Availability Zone goes down. Which solution should the architect choose?
Easy86A ticket booking system uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered?
Medium87A warehouse integration service must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The architecture review board prefers a managed AWS-native control.
Hard88A worker service consumes messages from an Amazon SQS queue. Some messages are malformed and always fail validation. The worker retries, but it keeps reprocessing the same bad messages and consumes processing capacity that should be used for valid work. What is the best solution to prevent “poison messages” from blocking progress?
Easy89A warehouse integration service receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers? The design must avoid adding custom operational scripts.
Medium90A media company stores video files in an Amazon S3 bucket in the us-east-1 Region. The company wants to ensure that the files are automatically replicated to us-west-2 for disaster recovery, and that replication occurs within 15 minutes of upload. Which solution meets these requirements with the LEAST operational overhead?
Medium91A company runs its customer-facing web app on EC2 behind an Application Load Balancer. The database is Amazon RDS for PostgreSQL. The requirement is that if a single Availability Zone fails, the database must automatically fail over within the same AWS Region with minimal application changes. Which database setup best meets this requirement?
Easy92A company is deploying a stateless web application on Amazon ECS with Fargate. The application must be resilient to individual task failures and Availability Zone failures. Which three steps should the company take to achieve this resilience? (Choose three.)
Medium93Your company hosts an internal API in two AWS Regions. You want Amazon Route 53 to automatically send traffic to the secondary Region if the primary Region’s endpoint becomes unhealthy. Which Route 53 configuration best meets this requirement?
Easy94A ticket booking system uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The architecture review board prefers a managed AWS-native control.
Medium95A trading dashboard stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured? The architecture review board prefers a managed AWS-native control.
Medium96A financial analytics platform runs a stateless containerized service on Amazon ECS with AWS Fargate tasks spread across three Availability Zones. The service reads from an Amazon Aurora MySQL cluster and must continue serving read traffic if one Availability Zone fails. The team wants the read capacity to remain available with the least operational overhead and no changes to application connection strings during a zone failure. Which approach meets these requirements?
Hard97A web application runs on an Auto Scaling group (ASG) behind an Application Load Balancer (ALB). The ASG uses the ALB target group health checks to decide when instances are healthy (for example, by using the ELB/target-group health check integration). During a deployment, the ASG performs instance replacement. Shortly after the deployment starts and while new instances are still bootstrapping, CloudWatch shows the ALB target group briefly has zero healthy targets, and users intermittently receive 502 responses. Which ASG deployment configuration best reduces the chance that there will be a period with zero healthy ALB targets, while still keeping failover behavior resilient?
Medium98A claims workflow uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure? The team wants the control to be enforceable during normal operations.
Hard99A startup runs a small internal tool on a single Amazon EC2 instance that uses an instance store volume for its database files. After a routine host maintenance event, the instance rebooted and the database was empty. The team wants the data to persist independently of the instance lifecycle and to survive a stop-and-start of the instance. What should they change?
Easy100A ticket booking system runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include?
Medium101A media company stores daily financial exports in Amazon S3. The files must be protected against accidental overwrite or deletion, and the business also wants a second copy in another Region for recovery after a regional outage. Which two actions should the architect take? Select two.
Medium102A serverless order-ingestion API writes directly to a database. During traffic spikes, the database occasionally throttles, Lambda retries create duplicate order records, and some requests time out. Which two changes best improve buffering and safe retry behavior? Select two.
Medium103A company runs a production MySQL database on Amazon RDS in us-east-1. A read replica exists in us-west-2 for disaster recovery. The primary region experiences a complete outage. Which of the following describes the correct procedure to restore database service using the cross-region read replica?
Hard104An orders service consumes payment instructions from an Amazon SQS queue. Sometimes the consumer times out after applying the payment but before deleting the SQS message. As a result, the same payment instruction is processed again. Which design change most directly prevents duplicate side effects caused by message retries?
Easy105A public API is deployed in two AWS Regions: us-east-1 (primary) and us-west-2 (secondary). The team wants Route 53 to automatically route users to the secondary region if the primary API becomes unhealthy. They will use Route 53 health checks that monitor the API’s /status endpoint over HTTPS. Which Route 53 configuration most directly implements this failover behavior?
Medium106A regional web application for a inventory service must fail over automatically to a secondary Region if the primary endpoint becomes unhealthy. Which two services or features are required? The team wants the control to be enforceable during normal operations.
Hard107A financial services company runs a critical application on Amazon EC2 instances in an Auto Scaling group. The application writes to an Amazon RDS for MySQL database. The company needs a recovery point objective (RPO) of 1 second and a recovery time objective (RTO) of 1 minute for the database in the event of a Regional disaster. Which solution meets these requirements?
Hard108A payments platform requires disaster recovery across Regions. Requirements: RPO of 15 minutes and RTO of about 1 hour. The business cannot afford full duplicate capacity in both Regions all the time, but the team wants automated readiness so failover is mostly operationally guided rather than a slow rebuild. Which DR strategy is the best fit?
Medium109A financial services firm runs a latency-sensitive trading application on Amazon EC2 instances distributed across three Availability Zones behind a Network Load Balancer. The application must continue serving traffic with no manual intervention if an entire Availability Zone becomes impaired, and each instance must receive a fair share of connections. Which combination of features meets these requirements?
Hard110A company runs a stateful web application on a fleet of Amazon EC2 instances in an Auto Scaling group. The application stores session state locally on each instance. During an Availability Zone failure, the Auto Scaling group replaces the unhealthy instances in a different AZ, but users lose their sessions and must log in again. The company wants to make the application resilient to AZ failures without requiring users to re-authenticate. Which solution should a solutions architect recommend?
Hard111Based on the exhibit, the database must continue serving if the current Availability Zone fails. What should you change?
Easy112A inventory service exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most?
Easy113Based on the exhibit, a faulty deployment corrupted production data at 10:30 UTC and the issue was discovered at 10:55 UTC. The team needs to recover the database to the last good state before the corruption. Which action should they take?
Medium114An order system receives events and uses a Lambda function to write each order into a database. During traffic spikes, the database sometimes throttles, and Lambda retries lead to occasional message loss in the event flow. The team wants buffering, automatic retries, and a way to isolate messages that repeatedly fail so they can be inspected later. What design change best meets this need?
Easy115A company uses Amazon RDS for a PostgreSQL database powering a customer-facing application. The application’s availability depends on fast database failover with minimal manual intervention. The RDS instance currently runs as a single-AZ deployment in one DB subnet group. Which change most directly meets the goal?
Medium116A trading dashboard runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include?
Medium117A trading dashboard stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured?
Medium118An internal service is hosted behind an Application Load Balancer (ALB) with targets spread across two Availability Zones. If the targets in one Availability Zone become unhealthy, the service must continue serving traffic from the healthy AZ. What change most directly improves resilience at the load-balancing layer?
Easy119A regional web application for a inventory service must fail over automatically to a secondary Region if the primary endpoint becomes unhealthy. Which two services or features are required? The design must avoid adding custom operational scripts.
Hard120A media company runs a stateless transcoding fleet on Amazon EC2 instances spread across three Availability Zones behind a Network Load Balancer. The fleet must keep processing jobs even if an entire Availability Zone becomes unavailable, and the architect wants to minimize manual intervention. Which combination of actions should the architect take to meet these requirements?
Medium121An event consumer sometimes processes the same SQS message more than once due to timeouts and retries. The consumer must ensure the payment is not charged twice. What design choice best addresses this requirement?
Easy122A inventory service uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The architecture review board prefers a managed AWS-native control.
Medium123A media company runs a transcoding fleet on Amazon EC2 instances behind an Application Load Balancer in a single Availability Zone. The business requires the workload to survive the loss of that Availability Zone with no manual intervention and minimal downtime. The instances store intermediate files on instance store volumes and the fleet is managed by an Auto Scaling group. Which change should a solutions architect make to meet the requirement?
Medium124A media company runs a video transcoding pipeline on Amazon EC2 instances in a single Availability Zone. The pipeline writes intermediate files to an Amazon EBS volume attached to each instance. The company needs the pipeline to survive the failure of any single Availability Zone and to recover automatically with minimal data loss. Which change should a solutions architect make?
Medium125A inventory service uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured?
Medium126Based on the exhibit, the application team wants the database to keep the same connection endpoint during failover and to reconnect automatically after the primary instance becomes unavailable. Which change best meets the requirement?
Medium127A content publishing system uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The architecture review board prefers a managed AWS-native control.
Medium128Based on the exhibit, the application tier is not replacing unhealthy instances even though the Auto Scaling group spans two Availability Zones. What change most directly improves automatic recovery when the application process fails?
Hard129A payments API requires point-in-time recovery and accidental-delete protection for a DynamoDB table. Which two settings should the architect enable? The architecture review board prefers a managed AWS-native control.
Hard130A trading dashboard uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered?
Medium131A startup runs a stateless web application on a single Amazon EC2 instance in one Availability Zone. The application has become popular, and the startup wants to ensure that the application can survive the failure of an Availability Zone and can handle increased traffic. Which architecture change should the startup make FIRST?
Easy132An events service publishes critical notifications using Amazon SNS. Three independent downstream systems (A, B, and C) subscribe to the topic. Downstream system B sometimes fails to process certain messages (for example, it times out or returns an error while handling the message), and you want: 1) failures in B to be isolated so A and C keep processing unaffected, and 2) messages that B cannot successfully process after retries to be sent to a DLQ for B. Which design best meets these requirements?
Medium133A stateless web API runs on EC2 instances behind an Application Load Balancer (ALB). The Auto Scaling group (ASG) currently uses subnets from only one Availability Zone, even though the ALB spans two Availability Zones. During maintenance of that single AZ, the ALB remains up but clients see timeouts because there are no healthy targets. Which change most directly improves resilience against an AZ failure?
Medium134An orders service publishes payment instructions to an Amazon SQS Standard queue. A downstream consumer sometimes times out or crashes after it has partially completed processing, causing the same instruction to be processed more than once. You must keep the design resilient without attempting to guarantee exactly-once processing. Which approach best handles duplicates safely?
Medium135A claims workflow uses an RDS MySQL database and must remain available during an Availability Zone failure with minimal application changes. What should the architect enable? The design must avoid adding custom operational scripts.
Medium136Based on the exhibit, the team must restore an Amazon RDS for PostgreSQL database to the exact state just before a bad delete happened. What is the best recovery approach?
Hard137A patient portal must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable?
Hard138A worker consumes messages from an Amazon SQS queue. Some messages consistently fail validation and are retried until the worker can no longer process them. What is the most appropriate AWS mechanism to handle these poison messages while keeping the queue usable?
Easy139A trading dashboard runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include? The team wants the control to be enforceable during normal operations.
Medium140Based on the exhibit, the company wants DNS traffic to fail over automatically from the primary Region to a secondary Region when the primary endpoint is unhealthy. Which Route 53 change is best?
Medium141A financial analytics platform stores results in an Amazon S3 bucket. Compliance requires that objects be recoverable for 30 days after deletion and that no user, including administrators, be able to permanently erase them during that window. Objects must also remain readable throughout the retention period. Which approach should the architect implement?
Hard142A warehouse integration service must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used?
Hard143A web application runs on an Auto Scaling group (ASG) behind an Application Load Balancer (ALB). The ASG is currently attached to subnets in only two Availability Zones (AZs). During a planned maintenance window, one AZ becomes unavailable for about 25 minutes. Monitoring shows that targets in the remaining AZ go healthy, and the ALB/target group health checks report normal. However, users still experience intermittent connection failures and slower responses during the AZ outage. What change will most directly improve resilience against an AZ loss while keeping the same ALB-based design?
Medium144A warehouse integration service receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers? The architecture review board prefers a managed AWS-native control.
Medium145A patient portal must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable? The design must avoid adding custom operational scripts.
Hard146A payments API requires point-in-time recovery and accidental-delete protection for a DynamoDB table. Which two settings should the architect enable? The design must avoid adding custom operational scripts.
Hard147A company hosts an internal API behind an Application Load Balancer (ALB) in two AWS Regions. They want Amazon Route 53 to automatically fail over to the secondary Region when the primary Region’s ALB is unhealthy. Health checks for the primary ALB are already configured, but the DNS record currently uses a latency-based routing policy. Which Route 53 configuration most directly provides automatic failover based on health status?
Medium148An orders service publishes payment instructions to an Amazon SQS Standard queue. A downstream consumer sometimes times out and retries the work, causing the consumer to process the same instruction more than once. Operationally, the team must ensure that duplicate processing does not create duplicate charges. The queue type cannot be changed. What is the most resilient application-side approach?
Medium149A ticket booking system runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include? The architecture review board prefers a managed AWS-native control.
Medium150A service processes messages from an Amazon SQS queue. Sometimes the worker finishes the business logic but does not delete the message before the visibility timeout expires, so the message is delivered again. Which two changes improve resilience and reduce the impact of duplicate processing? Select two.
Easy151Based on the exhibit, the team wants to stop poison messages from consuming worker capacity and also prevent duplicate side effects if the same message is delivered more than once. Which design change best meets the requirement?
Medium152A caching layer uses Amazon ElastiCache for Redis in front of a stateless web service. The service must continue to read cached responses during maintenance events and should automatically fail over to another node if one AZ becomes impaired. Which design change best satisfies this requirement?
Medium153An ECS service runs on EC2 instances and is fronted by an ALB. The ALB spans two Availability Zones, and the ECS service desired count is 2 tasks. The underlying EC2 capacity uses an Auto Scaling group (ASG) with min size set to 1, and the ASG also spans only one subnet in practice. What is the most effective change to meet the requirement that the service continues during a single-AZ instance loss?
Medium154A small e-commerce company runs a web application on a single Amazon EC2 instance in one Availability Zone. The instance stores session state locally and the database runs on the same instance. The company wants the application to survive an Availability Zone failure with minimal changes and no data loss for committed orders. Which combination of changes should the architect recommend?
Easy155A warehouse integration service must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable? The architecture review board prefers a managed AWS-native control.
Hard156An order-processing worker consumes messages from Amazon SQS. Occasionally, the worker times out after successfully creating a payment record but before deleting the message, which causes duplicate charges during retries. Some messages also fail validation repeatedly because required fields are missing. Which two changes should the team make? Select two.
Medium157A payments API uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure? The design must avoid adding custom operational scripts.
Hard158A content publishing system exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most? The design must avoid adding custom operational scripts.
Easy159A startup runs a nightly batch job on a single EC2 instance that reads a large dataset from Amazon S3, performs transformations, and writes results back to S3. The job takes about two hours, and the team wants the job to restart automatically if the instance fails or is terminated by AWS. The job is idempotent and can safely resume from the beginning. What is the MOST operationally efficient way to meet this requirement?
Easy160Based on the exhibit, downstream payment timeouts cause EventBridge deliveries to back up and some events are retried until they age out. What change best improves resilience and preserves events during downstream outages?
Hard161A company runs a stateful web application on a single Amazon EC2 instance in a public subnet. The application stores session data on the instance's root volume. The company wants to make the application highly available across two Availability Zones and ensure that session data is preserved if an instance fails. Which solution should a solutions architect recommend?
Hard162A healthcare company runs a critical patient-records API on Amazon EC2 instances behind an Application Load Balancer in a single AWS Region. The compliance team mandates that the API remain available even if an entire AWS Region becomes unavailable. The company wants a cost-effective solution that avoids running full production capacity in a second Region at all times. Which approach BEST meets these requirements?
Medium163A company runs a stateful analytics workload on EC2 instances that use EBS volumes. The data must be restorable in another Region after a major outage, with frequent point-in-time recovery. Which approach provides the most suitable replication mechanism for the EBS-backed data?
Medium164A company hosts a web application on EC2 instances behind an Application Load Balancer (ALB) in us-east-1. A static failover site is hosted in an S3 bucket with static website hosting enabled. The company needs automatic DNS failover to the S3 bucket if the primary ALB becomes unhealthy. Which Route 53 configuration achieves this?
Medium165A company runs a critical application on Amazon EC2 instances in a single Availability Zone. The application writes data to an Amazon RDS for MySQL DB instance that is not Multi-AZ. The company wants to improve the resilience of the database tier so that it can survive an Availability Zone failure with minimal downtime and no data loss. The application uses the database endpoint from the RDS console. Which solution meets these requirements?
Hard166A web application runs on an Amazon EC2 Auto Scaling group (ASG) behind an Application Load Balancer (ALB). The ALB is configured to use at least two Availability Zones (AZs), but the ASG currently uses subnets in only one AZ. If that AZ becomes unavailable, the application stops serving requests. Which change most directly improves resilience to an AZ outage?
Easy167A healthcare analytics platform ingests records into an Amazon Aurora MySQL cluster. Compliance rules require that the cluster remain writable even if an entire Availability Zone is lost, and that recovery happen without operator action. The team also wants read traffic to scale independently of the writer. Which configuration should a solutions architect choose?
Hard168An order-processing system publishes an event whenever a payment succeeds. Three downstream services (inventory, shipping, and analytics) must react independently. Analytics sometimes has high latency, but order processing must not be blocked. What is the best AWS approach to decouple these consumers?
Easy169An internal worker consumes messages from an Amazon SQS Standard queue. Recently, some messages fail validation in the worker (for example, missing required fields), causing the worker to crash before it can successfully process those messages. Those messages keep getting retried repeatedly, slowing down processing of valid messages. The team wants a resilient mechanism to quarantine bad messages after a limited number of receive attempts. What should they implement?
Medium170A logistics company stores shipment events in an Amazon S3 bucket. An analytics team must be able to recover any object version that is accidentally overwritten or deleted for at least 90 days, and objects must be protected from permanent deletion by any user, including the root user, during that window. Which combination of S3 features meets these requirements with the LEAST operational overhead?
Hard171A ticket booking system runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include? The design must avoid adding custom operational scripts.
Medium172A company uses an Amazon Aurora DB cluster in a Multi-AZ configuration. During a planned failover of the writer instance, the database endpoints in the application are updated incorrectly. After failover, reads work but writes fail with connection errors and timeouts for several minutes. The team currently uses the instance endpoint for the writer. What should they change to improve write resilience during failovers?
Medium173A claims workflow requires point-in-time recovery and accidental-delete protection for a DynamoDB table. Which two settings should the architect enable? The design must avoid adding custom operational scripts.
Hard174A trading dashboard stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured? The design must avoid adding custom operational scripts.
Medium175An order processing workflow uses Amazon SQS as the decoupling layer between a producer and a consumer Lambda function. The consumer intermittently fails due to a downstream dependency. The team has observed that certain “poison” messages keep being retried repeatedly and prevent other messages from being processed efficiently. Which SQS configuration most directly addresses this issue?
Medium176A startup runs a stateless web application on a single Amazon EC2 instance in one Availability Zone. The application must remain available if the instance fails or if its Availability Zone becomes unavailable. The startup wants a managed solution that requires minimal operational overhead. Which solution should a solutions architect recommend?
Easy177A system processes events from Amazon SQS and sometimes sees duplicate messages due to retries. The business requirement is that each payment must be charged at most once. What design choice best addresses this resiliency requirement?
Easy178A company hosts a web application on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer (ALB). The ALB and the Auto Scaling group are currently deployed in only one Availability Zone (AZ). The business wants the application to keep running if that AZ has an outage. What is the best change?
Easy179A web application runs on an Amazon EC2 Auto Scaling group behind an Application Load Balancer (ALB). After each deployment, new instances take about 2 minutes to download artifacts and become ready to accept requests on the target port. In the last deployment, the ALB started marking targets unhealthy before the app was ready, and the Auto Scaling group then replaced those instances repeatedly, causing a prolonged outage. Which change best improves resilience during instance start-up without reducing actual availability once the application is healthy?
Medium180An internal API is deployed in two AWS Regions behind separate Application Load Balancers. The company wants clients to use the primary Region when it is healthy and automatically switch to the secondary Region if the primary health check fails. Which two Route 53 record configurations are required? Select two.
Medium181A SaaS platform plans to run in two AWS Regions for lower latency. The team wants to enable active-active writes (both regions accept updates) to avoid failover downtime. However, the business requires strong consistency for order status transitions (for example, only one transition from “Paid” to “Shipped” must be allowed). Which statement is the best architectural choice to meet the consistency requirement?
Medium182A consumer application reads from an Amazon SQS queue. Some messages have an invalid format and always fail processing. They are retried repeatedly and consume consumer capacity. What is the best way to prevent these "poison pill" messages from blocking normal processing?
Easy183A healthcare company stores patient records in an Amazon DynamoDB table. The table must be recoverable to any point within the last 35 days, and the data must remain available if an entire AWS Region becomes unavailable. Which two actions should a solutions architect take to meet these requirements? (Choose two.)
Medium184A company is building a serverless application that processes messages from an Amazon SQS queue using AWS Lambda. The application must not lose messages and must handle occasional downstream failures gracefully. The Lambda function sometimes fails due to a transient error in a downstream service. The company wants to ensure that failed messages are retried and eventually processed, but also wants to avoid infinite retries that could block the queue. What should the company do?
Hard185A healthcare analytics team runs a containerized reporting service on Amazon ECS with the Fargate launch type in a single Availability Zone. The service must remain available if one Availability Zone fails, and it must scale automatically based on CPU utilization. The tasks are stateless and write output to Amazon S3. Which configuration should a solutions architect implement?
Medium186Based on the exhibit, the web tier becomes unavailable if us-west-2a has an outage. What is the best change to improve resilience with the least redesign?
Easy187A warehouse integration service must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The design must avoid adding custom operational scripts.
Hard188A media company stores finalized video masters in an Amazon S3 bucket in the us-east-1 Region. Compliance requires that the objects be recoverable if they are accidentally deleted or overwritten for at least 90 days, and that no user, including administrators, be able to permanently erase them during that period. Which S3 feature should the solutions architect enable?
Easy189Based on the exhibit, the database is manually promoted during an Availability Zone failure and the application outage lasts longer than the target. What change best improves resilience with the least operational intervention?
Hard190An application writes to an Amazon Aurora DB cluster. After a planned Aurora failover, the application experiences several minutes of connection errors. The logs show the application continues connecting to the specific DB instance endpoint that was the primary before the failover. What change most directly improves resilience during Aurora failovers?
Medium191A company uses Amazon RDS with automated backups enabled (retention period: 7 days). At 10:30 UTC, a bad release corrupts specific rows in a production table. The team detects the issue at 11:10 UTC. They need to revert the database state to what it was from 10:00–10:30 UTC, recover quickly, and minimize risk to the currently running workload. What is the best option?
Medium192A service processes customer payments from a message queue. Because the queue provides at-least-once delivery, the same payment message can be delivered more than once if the consumer times out before committing its state. Currently, the service sometimes charges the customer twice. Which design change most directly prevents duplicate charges while still allowing safe retries?
Medium193A media company stores original uploads in an S3 bucket. They must recover from accidental overwrites/deletes and also recover quickly from a full Region outage. The required RPO is about 1 hour. Which configuration best meets these requirements?
Medium194A warehouse integration service must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable? The design must avoid adding custom operational scripts.
Hard195A content publishing system uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The team wants the control to be enforceable during normal operations.
Medium196A SaaS application is deployed in us-east-1 and us-west-2 behind separate ALBs. The business wants DNS to send new clients to the primary Region when it is healthy and automatically fail over to the secondary Region when the primary endpoint is unhealthy. Which two Route 53 settings are required? Select two.
Medium197A logistics company runs an order-tracking API on a fleet of EC2 instances in a single Availability Zone behind a Network Load Balancer. The architecture team must make the API resilient to the loss of that Availability Zone without changing the API endpoint that clients already use. The instances are stateless and store session data in a shared Amazon ElastiCache cluster. Which change should the solutions architect make to meet these requirements?
Medium198A payments API uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure? The architecture review board prefers a managed AWS-native control.
Hard199A patient portal receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers? The design must avoid adding custom operational scripts.
Medium200A media company stores original video masters in an Amazon S3 bucket in the us-east-1 Region. Compliance requires that a readable copy of every object exist in the eu-west-1 Region within 15 minutes of upload, and that the objects in eu-west-1 be usable directly by an application there. No transformations are required. Which S3 feature should the solutions architect enable?
Medium201A regional web application for a content publishing system must fail over automatically to a secondary Region if the primary endpoint becomes unhealthy. Which two services or features are required? The design must avoid adding custom operational scripts.
Hard202A logistics company runs an order-tracking service that exposes a REST API. The service must remain available during a single Availability Zone failure and must keep read latency low for a globally distributed user base. The data store must support automatic multi-AZ replication without the team managing database servers. Which solution meets these requirements?
Medium203Your web application is deployed in two AWS Regions (Region A and Region B). You want Route 53 to automatically fail over DNS traffic from Region A to Region B when Region A is unhealthy. The failover decision must be based on health checks that verify whether the application in Region A is reachable. Which Route 53 routing configuration best meets these requirements?
Medium204A media company stores original video masters in an Amazon S3 bucket in the us-east-1 Region. Compliance requires that a readable copy of every object exists in eu-west-1 within 15 minutes of upload, and that objects deleted in the source bucket do not automatically disappear from the destination. Which S3 feature should the solutions architect enable?
Medium205A media processing company runs a stateless thumbnail-generation fleet on Amazon EC2 instances behind an Application Load Balancer. The instances store no local state, and the team wants the fleet to survive the loss of an entire Availability Zone without manual intervention. The fleet must also scale out automatically based on CPU. Which combination of AWS services should the solutions architect use to meet these requirements with the LEAST operational overhead?
Medium206A trading dashboard uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The design must avoid adding custom operational scripts.
Medium207A company runs a web application on Amazon EC2 instances behind an Application Load Balancer. The application must be highly available and able to withstand the failure of an entire AWS Region. The company wants to minimize operational overhead and ensure that failover is automatic. Which solution should a solutions architect recommend?
Medium208A financial services firm runs a stateful trading application on EC2 instances in an Auto Scaling group. Each instance maintains an in-memory cache that takes several minutes to rebuild after a restart, and the team wants the application to survive the loss of an Availability Zone with minimal disruption. The application cannot be made stateless in the near term. Which approach should a solutions architect recommend?
Hard209An application uses an Amazon Aurora DB cluster. The cluster performs an automatic failover from the writer instance to a standby instance. After failover completes, reads succeed, but all new writes fail with errors indicating the application is connecting to the old writer endpoint. Which change best fixes the resiliency issue after failover?
Medium210A company needs to store application logs in a durable and highly available manner. The logs are written continuously by multiple EC2 instances and are accessed infrequently for compliance audits. The company wants a solution that provides 99.999999999% (11 9's) durability and automatically replicates data across multiple Availability Zones. Which AWS service should the company use?
Easy211Based on the exhibit, the application should continue serving requests if one Availability Zone fails. Which change best improves resilience with the least operational complexity?
Medium212A startup runs a stateless web tier on Amazon EC2 instances in an Auto Scaling group that spans three Availability Zones. The team wants the application to keep serving requests even if one instance becomes unresponsive, without operator involvement. What should the solutions architect configure?
Easy213A startup runs a customer-facing web application on a single Amazon EC2 instance in one Availability Zone, with the database on the same instance. The founders want the application to survive the failure of that Availability Zone with minimal changes and no server management for the database tier. Which action should the solutions architect take first?
Easy214Based on the exhibit, the payment worker sometimes processes the same SQS Standard message more than once after a timeout. What change best prevents duplicate charges while keeping the queue architecture?
Medium215A production Amazon RDS database has automated backups enabled. At 10:00 UTC, an application deploy accidentally overwrote a subset of rows due to a faulty migration. The issue is detected at 10:45 UTC. The team confirms that the required retention window is still available. Which approach offers the most resilient and least disruptive way to recover the affected data close to the time of the event?
Medium216A company runs an Amazon Aurora DB cluster with a Multi-AZ deployment. The application is configured with a hard-coded endpoint that points to the current writer *DB instance* (an instance-specific endpoint), rather than the Aurora cluster writer endpoint. During an unexpected AZ failure, Aurora promotes the standby to become the new writer. However, the application continues to fail to connect until an operator updates the hard-coded endpoint. What change most directly improves resiliency so the application automatically reconnects after failover?
Medium217Based on the exhibit, duplicate payment charges occasionally occur when the worker times out after the charge is submitted but before the message is deleted. What change best prevents duplicate charges while keeping retry behavior?
Hard218A team wants a web application to keep serving traffic if one Availability Zone fails. Match each architecture element to the resilience behavior it provides.
Medium219A ticket booking system uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The team wants the control to be enforceable during normal operations.
Medium220A logistics company runs an order-tracking service on Amazon EC2 instances that write state to an Amazon DynamoDB table. A recent incident showed that a single Availability Zone failure caused the service to lose capacity, and the team also discovered that a developer accidentally deleted a production table. The architect must improve both Availability Zone resilience and protection against accidental table deletion. (Choose two.)
Medium221A warehouse integration service must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The team wants the control to be enforceable during normal operations.
Hard222A team uses an S3 bucket to store important customer-generated exports. They need protection against accidental overwrites and also want copies of the data in another AWS Region for disaster recovery. Which S3 configuration best satisfies both requirements?
Easy223A media company runs a video-transcoding fleet on Amazon EC2 instances that read source files from an Amazon S3 bucket and write output to a second bucket. The fleet is spread across three Availability Zones in one Region, and instances are launched by an Auto Scaling group. The company needs the architecture to survive the loss of an entire Availability Zone without losing in-flight transcoding work or requiring manual intervention. Which combination of design elements should a solutions architect implement to meet these requirements?
Medium224A company is designing a disaster recovery plan for a critical application hosted on AWS. The application runs on EC2 instances with data stored in Amazon EBS volumes and Amazon S3. The recovery time objective (RTO) is 15 minutes, and the recovery point objective (RPO) is 1 hour. Which three strategies would help meet these objectives? (Choose three.)
Medium225A service consumes messages from an SQS queue. Recently, a new message format started failing validation in the consumer. The consumer catches the exception but cannot successfully process those messages without code changes. The team wants failed messages to be isolated for later investigation instead of being retried indefinitely. What should they configure?
Medium226A production Amazon RDS database has automated backups enabled with sufficient retention. At 10:30 UTC, a release corrupts specific rows. The issue is detected at 10:45 UTC. The team wants to restore the database state to before the corruption with minimal complexity. What should they do?
Easy227A Multi-AZ Amazon RDS database experiences incorrect writes at 10:15 UTC due to a buggy release. The team detects the problem at 10:25 UTC. They want to restore the data to a known-good point around 10:15 UTC, and validate the recovered data, without taking the current production instance offline during the recovery process. What is the most appropriate AWS action?
Medium228A financial analytics platform ingests events into an Amazon Kinesis Data Stream with four shards. During month-end peaks, producers receive ProvisionedThroughputExceededException errors and consumers fall behind. The architects want to increase capacity without changing producer code and must preserve the order of records that share the same partition key. What should they do?
Hard229A production team accidentally deletes critical rows in an Amazon RDS for PostgreSQL database. The deletion occurred about 6 hours ago. The team wants to recover to a specific point in time with minimal disruption. Assuming automated backups are enabled, which approach provides the best resilience outcome?
Medium230A company wants a disaster recovery setup for a web application. They want to keep costs low but still recover within a couple of hours after a regional disruption. They are willing to run only minimal infrastructure in the secondary location and scale it up during the outage. Which DR approach best matches this requirement?
Easy231A financial services company is designing a new payment processing platform. The platform must continue to accept and process transactions even if an entire AWS Region becomes unavailable, and it must not lose any accepted transaction. The architects have decided to run active-active deployments in two Regions and use Amazon Route 53 for traffic management. Which two additional design elements are required to meet the durability and availability goals? (Choose two.)
Hard232A inventory service exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most? The design must avoid adding custom operational scripts.
Easy233A financial services company runs a critical application on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer. The application must be able to survive the failure of an entire AWS Region. The company wants a cost-effective solution that minimizes operational overhead. Which approach should the architect recommend?
Hard234Based on the exhibit, which Route 53 configuration should be used so traffic automatically returns to the secondary Region only when the primary Region becomes unhealthy?
Medium235A web application runs on an Auto Scaling group (ASG) behind an Application Load Balancer (ALB). After a new release, instances begin failing ALB health checks with errors like 502 while the application is still starting up. CloudWatch shows that the ASG replaces the instances before they finish initializing, so traffic never reaches healthy targets. Which change most directly prevents premature replacement during startup so traffic can resume as soon as the instances are actually healthy?
Medium236A company is designing a multi-Region disaster recovery (DR) strategy for a stateless web application running on Amazon EC2 instances behind an Application Load Balancer (ALB). The application uses an Amazon RDS for MySQL database as its data store. The architecture must provide rapid failover with the lowest possible Recovery Point Objective (RPO) and Recovery Time Objective (RTO). Which of the following design choices will help achieve these objectives? (Choose four.)
Medium237A startup runs a stateless image-resizing API on a fleet of EC2 instances behind an Application Load Balancer. The instances store uploaded source images on their own instance store volumes before processing. During a routine scale-in event, an instance was terminated and several in-flight uploads were lost. The architect must make the design resilient to instance loss without changing the API code. What should the architect do?
Easy238A inventory service uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The design must avoid adding custom operational scripts.
Medium239A company runs a customer portal on an Amazon Aurora PostgreSQL cluster. The application currently connects directly to the writer instance endpoint and keeps long-lived connections open. During a maintenance failover, writes fail until clients are restarted. The team wants the application to reconnect to the correct Aurora endpoint automatically and reduce user-visible write interruptions. Which change is most likely to achieve this?
Medium240A developer accidentally deletes important rows in an RDS database. The mistake is discovered 45 minutes later. The database has automated backups enabled with a retention period of 7 days. What is the best way to restore the database to a point just before the deletion?
Medium241A payments API uses an RDS MySQL database and must remain available during an Availability Zone failure with minimal application changes. What should the architect enable?
Medium242You host a public API using Amazon API Gateway in two AWS Regions: us-east-1 (primary) and us-west-2 (secondary). You want Route 53 to send client traffic to the secondary region only when the primary API is unhealthy. Which Route 53 setup best meets this requirement?
Medium243An ECS service runs on EC2 instances and is fronted by an ALB. The ALB spans two Availability Zones, and the ECS service desired count is 2 tasks. The underlying EC2 capacity uses an Auto Scaling group (ASG) with min size set to 1, and the ASG also spans only one subnet in practice. What is the most effective change to meet the requirement that the service continues during a single-AZ instance loss?
Medium244A solutions architect is designing a highly available relational database tier for a customer-facing order system that must survive the loss of an entire Availability Zone with minimal administrative effort and no application connection-string changes during failover. (Choose two.)
Medium245A ticket booking system uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The design must avoid adding custom operational scripts.
Medium246A healthcare data platform stores patient documents in an Amazon S3 bucket in us-east-1. Regulations require that the data remain readable with low latency even if the entire us-east-1 Region becomes unavailable, and that writes continue in a secondary Region. The team wants object-level replication with minimal operational overhead and must preserve version history. Which solution BEST meets these requirements?
Hard247Based on the exhibit, an administrator accidentally deleted data from Amazon RDS for PostgreSQL about 90 minutes ago. Which recovery approach best restores the database to the exact required point in time?
Medium248A warehouse integration service receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers?
Medium249A financial analytics platform runs a stateless API on Amazon EC2 instances in an Auto Scaling group behind a Network Load Balancer. The API reads from an Amazon Aurora MySQL cluster that has a single writer instance and one reader instance in a different Availability Zone. During a recent Availability Zone event, the writer instance failed and the API saw several minutes of failed writes. The team wants writes to resume automatically with minimal downtime and no application code changes. What should the solutions architect do?
Hard250An internal API is hosted in two AWS Regions behind Route 53. Under normal conditions, clients should use the primary region. If the primary endpoint becomes unhealthy, traffic must automatically switch to the secondary region. Which Route 53 setup best meets this requirement?
Easy251An orders service currently sends HTTP requests directly to two downstream services (inventory and shipping). During peak load, inventory slows down, causing the orders service to slow as well. The team wants the orders service to remain responsive even when a downstream service is temporarily slow or restarted. Which design change best achieves this resiliency goal?
Easy252A content publishing system uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The design must avoid adding custom operational scripts.
Medium253A healthcare company runs a batch ingestion pipeline on Amazon EC2 instances that read messages from an Amazon SQS queue and write results to Amazon DynamoDB. The pipeline must be resilient so that a single instance failure does not stop processing and no messages are lost. Which two architectural changes should a solutions architect make to meet these requirements? (Choose two.)
Medium254A media startup stores user-uploaded video files in an Amazon S3 bucket in the us-east-1 Region. The compliance team requires that the data remain recoverable if an entire AWS Region becomes unavailable, and that recovery can be performed by pointing applications at a different endpoint. Cost should be minimized while still meeting the requirement. Which solution should a solutions architect recommend?
Easy255Based on the exhibit, the current disaster recovery design misses the RTO target even though the database replica is current. Which deployment model best meets the requirements with the least always-on cost?
Hard256A retail platform needs disaster recovery across AWS Regions. The business requirement is: RTO up to 6 hours, RPO up to 1 hour, and they want the ability to start serving quickly during a Region outage but do not want to run full production capacity continuously. Which DR strategy best fits these requirements?
Easy257A media company stores original video assets in an Amazon S3 bucket in the us-east-1 Region. Editors in Europe report slow downloads, and the legal team requires that a copy of every asset exist in eu-west-1 within 15 minutes of upload, with the ability to fail over reads to the European copy during a Regional impairment. Which S3 feature should the architects enable?
MediumOther domains
All SAA-C03 exam domains
Frequently asked questions
- What does the Design Resilient Architectures domain cover on the SAA-C03 exam?
- High availability and resilience questions test multi-AZ vs multi-Region patterns, Auto Scaling, load balancing and the right service for a given recovery time objective.
- How many questions are in this domain?
- This page lists all 257 Design Resilient Architectures questions in the SAA-C03 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Design Resilient Architectures questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.