Courseiva

CCNA Design Resilient Architectures Questions

75 of 257 questions · Page 3/4 · Design Resilient Architectures · Answers revealed

151
MCQmedium

Based on the exhibit, the team wants to stop poison messages from consuming worker capacity and also prevent duplicate side effects if the same message is delivered more than once. Which design change best meets the requirement?

A.Increase the SQS queue batch size so each worker processes more messages per request.
B.Replace SQS with Amazon SNS and let each worker subscribe directly to the topic.
C.Configure a dead-letter queue and make the handler idempotent by storing a durable processed-message key.
D.Disable retries and shorten the visibility timeout so failed messages disappear sooner.
AnswerC

A dead-letter queue isolates messages that repeatedly fail so they stop wasting worker capacity. Idempotency ensures a message processed more than once does not create duplicate side effects, which is essential when visibility timeouts expire or retries occur. Together, these controls address both poison-message handling and at-least-once delivery behavior.

Why this answer

A dead-letter queue isolates poison messages that repeatedly fail processing, preventing them from consuming worker capacity. Making the handler idempotent by storing a durable processed-message key (e.g., using DynamoDB or a database) ensures that even if the same message is delivered more than once, duplicate side effects are avoided. This combination directly addresses both requirements: stopping poison messages from wasting resources and preventing duplicate processing.

Exam trap

The trap here is that candidates often think disabling retries or increasing batch size solves poison messages, but they fail to realize that only a dead-letter queue isolates problematic messages, and idempotency is required to handle duplicate deliveries inherent in SQS's at-least-once delivery model.

How to eliminate wrong answers

Option A is wrong because increasing the SQS batch size does not prevent poison messages from consuming worker capacity; it only makes each worker process more messages per request, which could actually increase the impact of poison messages. Option B is wrong because replacing SQS with Amazon SNS and having workers subscribe directly to the topic removes the ability to decouple producers and consumers, and SNS does not provide message retention, retries, or a dead-letter queue mechanism, so poison messages would still be delivered and could cause duplicate side effects. Option D is wrong because disabling retries and shortening the visibility timeout would cause failed messages to disappear sooner, but this does not prevent duplicate side effects (messages could still be redelivered before being deleted) and does not isolate poison messages—they would simply be lost, not handled.

152
MCQmedium

A caching layer uses Amazon ElastiCache for Redis in front of a stateless web service. The service must continue to read cached responses during maintenance events and should automatically fail over to another node if one AZ becomes impaired. Which design change best satisfies this requirement?

A.Deploy a single-node Redis cluster and rely on application-level retries when cache misses occur.
B.Configure an ElastiCache Redis replication group with automatic failover across multiple Availability Zones.
C.Move the cache into the VPC but keep it in one Availability Zone to reduce network latency.
D.Use a Memcached cluster and configure only client-side connection pooling without failover support.
AnswerB

A Redis replication group with automatic failover maintains a primary and replicas across Availability Zones; if the primary's AZ is impaired, ElastiCache promotes a replica automatically, so cached reads continue during maintenance events as the stem requires.

Why this answer

An ElastiCache Redis replication group with automatic failover across multiple Availability Zones ensures that if the primary node or its AZ becomes impaired, a read-replica in another AZ is automatically promoted to primary. This allows the stateless web service to continue reading cached responses without interruption, satisfying both the maintenance and AZ impairment requirements.

Exam trap

The trap here is that candidates often confuse Memcached's simplicity with Redis's replication capabilities, assuming that client-side connection pooling alone can handle failover, when in fact Memcached lacks any built-in replication or automatic failover mechanism.

Why the other options are wrong

A

A single-node Redis cluster lacks automatic failover; if the node or its AZ becomes impaired, the service cannot read cached responses until the node is restored, violating the requirement for continued reads during maintenance and AZ impairment.

C

Keeping the cache in one Availability Zone does not provide automatic failover to another node if that AZ becomes impaired, failing the requirement for high availability during maintenance events.

D

Memcached does not support automatic failover or multi-AZ replication; if an AZ becomes impaired, the cache becomes unavailable, violating the requirement for continued reads during maintenance.

153
MCQmedium

An ECS service runs on EC2 instances and is fronted by an ALB. The ALB spans two Availability Zones, and the ECS service desired count is 2 tasks. The underlying EC2 capacity uses an Auto Scaling group (ASG) with min size set to 1, and the ASG also spans only one subnet in practice. What is the most effective change to meet the requirement that the service continues during a single-AZ instance loss?

A.Set the ECS deployment configuration to maximum percent 100 so tasks replace instances faster during rollouts.
B.Increase ASG min size to at least 2 and ensure the ASG uses subnets in at least two Availability Zones.
C.Enable ALB connection draining longer than expected so existing connections survive longer during an AZ event.
D.Reduce task memory reservations to pack both tasks onto a single EC2 instance.
AnswerB

Multi-AZ instance capacity ensures tasks have eligible compute in another AZ when one AZ loses instances.

Why this answer

The current architecture has a single point of failure: the ASG spans only one subnet (one AZ), so if that AZ fails, all EC2 instances are lost, and the ECS service cannot run any tasks. By increasing the ASG min size to at least 2 and ensuring it uses subnets in at least two AZs, the ASG will maintain at least one healthy instance in each AZ, allowing the ECS service to survive a single-AZ outage. This aligns with the AWS Well-Architected Framework's principle of deploying across multiple AZs for high availability.

Exam trap

The trap here is that candidates may focus on the ALB's multi-AZ configuration and overlook that the compute layer (ASG/EC2) is the actual bottleneck, leading them to choose connection draining or deployment settings that do not address the fundamental lack of cross-AZ capacity.

How to eliminate wrong answers

Option A is wrong because setting the ECS deployment configuration maximum percent to 100 controls how many tasks are replaced during a rolling update, not the ability to survive an AZ failure; it does not address the underlying lack of EC2 capacity across AZs. Option C is wrong because ALB connection draining only gracefully terminates existing connections during deregistration or health check failures; it does not prevent service interruption when all EC2 instances in the single AZ become unavailable. Option D is wrong because reducing task memory reservations to pack both tasks onto a single EC2 instance actually increases risk—if that single instance or its AZ fails, both tasks are lost, violating the resilience requirement.

154
MCQeasy

A small e-commerce company runs a web application on a single Amazon EC2 instance in one Availability Zone. The instance stores session state locally and the database runs on the same instance. The company wants the application to survive an Availability Zone failure with minimal changes and no data loss for committed orders. Which combination of changes should the architect recommend?

A.Create an Amazon Machine Image of the instance and configure a weekly cron job to launch a replacement instance in a second Availability Zone if the first fails.
B.Attach an additional Amazon EBS volume to the instance and take hourly snapshots to Amazon S3 for disaster recovery.
C.Move the database to Amazon RDS with a Multi-AZ deployment, place the web tier in an Auto Scaling group across multiple Availability Zones, and store session state in Amazon ElastiCache for Redis.
D.Enable detailed monitoring on the EC2 instance and create an Amazon CloudWatch alarm that reboots the instance when the status check fails.
AnswerC

Amazon RDS Multi-AZ provides a synchronous standby in another Availability Zone with automatic failover, protecting committed orders. An Auto Scaling group across AZs keeps the web tier available, and ElastiCache for Redis externalizes session state so users are not tied to one instance, delivering resilience with minimal application redesign.

Why this answer

RDS Multi-AZ keeps a synchronous standby in a second Availability Zone and fails over automatically, protecting committed orders. An Auto Scaling group across AZs maintains the web tier, and externalizing sessions to ElastiCache for Redis removes dependence on a single instance, so the application survives an AZ failure with minimal redesign.

Exam trap

The trap here is treating EBS snapshots or instance reboots as high availability, when they only provide backup or recover the same failed Availability Zone.

155
MCQhard

A warehouse integration service must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable? The architecture review board prefers a managed AWS-native control.

A.Use CloudFront signed URLs
B.Use Amazon SQS standard queue and design consumers to be idempotent
C.Use UDP messages sent directly to workers
D.Use an in-memory queue on one EC2 instance
AnswerB

An SQS standard queue is designed for decoupling components and provides at-least-once delivery with high throughput, ensuring that every message placed in the queue is eventually delivered to a consumer. Because at-least-once delivery can produce duplicate messages, consumers must be built to process each event idempotently so that repeated handling does not corrupt state or create duplicate outcomes. This combination satisfies the requirement that every event is processed.

Why this answer

Amazon SQS standard queues provide at-least-once delivery, which guarantees that every message is processed at least once, meeting the requirement that every event must be processed. Duplicate processing is acceptable because the consumer can be designed to handle idempotency. This is a managed, AWS-native service that aligns with the architecture review board's preference.

Exam trap

The trap here is that candidates may confuse 'at-least-once' delivery with 'exactly-once' delivery and incorrectly choose a solution like FIFO queues (not listed) or dismiss SQS standard queues due to the duplicate processing allowance, but the question explicitly states duplicates are acceptable if idempotency is handled, making SQS standard the correct choice.

How to eliminate wrong answers

Option A is wrong because CloudFront signed URLs are used for securing content delivery, not for event processing or messaging. Option C is wrong because UDP is a connectionless, unreliable protocol that does not guarantee delivery, so it cannot ensure at-least-once processing. Option D is wrong because an in-memory queue on a single EC2 instance is not managed, not AWS-native, and introduces a single point of failure, violating the requirement for a resilient, managed service.

156
Multi-Selectmedium

An order-processing worker consumes messages from Amazon SQS. Occasionally, the worker times out after successfully creating a payment record but before deleting the message, which causes duplicate charges during retries. Some messages also fail validation repeatedly because required fields are missing. Which two changes should the team make? Select two.

Select 2 answers
A.Make the payment step idempotent using a unique transaction identifier.
B.Configure an SQS dead-letter queue with a redrive policy.
C.Reduce the visibility timeout so failed messages return to the queue faster.
D.Run only one long-lived worker instance so the queue can never be processed twice.
E.Switch from a standard queue to a FIFO queue and remove all other changes.
AnswersA, B

Correct. SQS provides at-least-once delivery, so the same message can be processed more than once if the worker times out, retries, or crashes after partially completing the work. An idempotency key lets the application recognize that the payment was already created and prevents duplicate charges.

Why this answer

Making the payment step idempotent using a unique transaction identifier ensures that if the same message is processed multiple times due to a timeout, the payment is only charged once. This is a common pattern for handling at-least-once delivery semantics in Amazon SQS, where the worker must be designed to handle duplicate messages safely.

Exam trap

The trap here is that candidates often think reducing the visibility timeout will speed up recovery, but it actually increases the chance of duplicate processing, and they may also overlook that a FIFO queue alone does not fix the worker's failure to delete the message after processing.

157
MCQhard

A payments API uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure? The design must avoid adding custom operational scripts.

A.A FIFO queue without a redrive policy
B.A dead-letter queue with an appropriate maxReceiveCount
C.A larger message retention period only
D.Short polling instead of long polling
AnswerB

A dead-letter queue with a configured maxReceiveCount is the standard SQS mechanism for poison messages. When a message is received that many times without being deleted, SQS automatically moves it to a separate DLQ, isolating the problematic message from the main queue. This lets you inspect and debug the payload while healthy traffic continues unaffected. This also protects downstream consumers from continuous failure loops.

Why this answer

A dead-letter queue (DLQ) with an appropriate maxReceiveCount allows messages that repeatedly fail processing to be moved out of the source queue after a specified number of receive attempts. This prevents poison messages from blocking retries and consuming processing resources, without requiring custom operational scripts.

Exam trap

The trap here is that candidates may think increasing retention or switching polling modes solves poison messages, but only a DLQ with maxReceiveCount directly addresses repeated failures without custom scripts.

How to eliminate wrong answers

Option A is wrong because a FIFO queue without a redrive policy does not automatically handle poison messages; it still requires a DLQ configuration to move failing messages out. Option C is wrong because increasing the message retention period only keeps messages longer but does not prevent poison messages from repeatedly failing and blocking retries. Option D is wrong because short polling (immediate return with fewer messages) does not address poison message handling; it only affects message availability and latency, not failure management.

158
MCQeasy

A content publishing system exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most? The design must avoid adding custom operational scripts.

A.IAM Access Analyzer
B.AWS Backup Vault Lock
C.CloudFront caching with appropriate TTLs
D.S3 Select
AnswerC

CloudFront caching with appropriate TTL values lets the CDN store static objects at edge locations and continue serving those cached copies even when the S3 origin is temporarily unreachable, as long as the cached object has not expired. Setting longer TTLs for immutable content (e.g., versioned images, scripts, and stylesheets) reduces the frequency of origin fetches and widens the window of resilience during an S3 outage. This is the only proposed solution that keeps content available to end users during an origin failure, directly addressing the requirement to tolerate an S3 outage.

Why this answer

CloudFront caches responses from the S3 origin based on configured TTLs (Cache-Control or Expires headers). If the S3 origin experiences a short outage, CloudFront can still serve cached content to users as long as the TTL has not expired, ensuring availability without custom scripts. This is the most direct and resilient feature for this use case.

Exam trap

The trap here is that candidates may confuse backup or access control features (like Backup Vault Lock or IAM Access Analyzer) with availability mechanisms, or think S3 Select provides caching, when the correct answer is simply leveraging CloudFront's built-in caching TTLs to serve stale content during origin outages.

How to eliminate wrong answers

Option A is wrong because IAM Access Analyzer helps identify unintended access to resources, not caching or origin resilience. Option B is wrong because AWS Backup Vault Lock prevents deletion of backups, not caching or serving stale content during origin outages. Option D is wrong because S3 Select is a feature to retrieve subsets of data from objects using SQL queries, not related to caching or origin resilience.

159
MCQeasy

A startup runs a nightly batch job on a single EC2 instance that reads a large dataset from Amazon S3, performs transformations, and writes results back to S3. The job takes about two hours, and the team wants the job to restart automatically if the instance fails or is terminated by AWS. The job is idempotent and can safely resume from the beginning. What is the MOST operationally efficient way to meet this requirement?

A.Use EC2 Auto Recovery to recover the instance onto new hardware and attach an Amazon EBS volume that persists the job's progress.
B.Place the instance in an Auto Scaling group with a minimum and desired capacity of one, use a launch template that installs and starts the job at boot, and configure a lifecycle hook to keep the instance in service until the job completes.
C.Create an Amazon CloudWatch alarm on the StatusCheckFailed metric that invokes an AWS Lambda function to call the EC2 RebootInstances API.
D.Convert the job into an AWS Lambda function with a 15-minute timeout and trigger it on a nightly Amazon EventBridge schedule.
AnswerB

An Auto Scaling group with a desired capacity of one continuously maintains a single healthy instance, so if AWS terminates or the instance fails, the group replaces it and the launch template's user data reruns the idempotent job. A lifecycle hook can delay termination long enough for a clean shutdown, and this is fully managed with no custom monitoring code.

Why this answer

A single-instance Auto Scaling group is the managed way to keep exactly one healthy instance running and to replace it automatically when it is terminated or fails. Supplying the job through a launch template's user data means every replacement instance starts the job again, and because the job is idempotent, restarting from the beginning is acceptable. No custom alarm or recovery scripting is required.

Exam trap

The trap here is reaching for CloudWatch alarms and Lambda automation when the built-in Auto Scaling group replacement behavior already satisfies the restart requirement.

160
MCQhard

Based on the exhibit, downstream payment timeouts cause EventBridge deliveries to back up and some events are retried until they age out. What change best improves resilience and preserves events during downstream outages?

A.Increase the Lambda timeout so each invocation can wait longer for the payment API.
B.Put an Amazon SQS queue between EventBridge and the consumer, and have workers drain the queue with a DLQ for poison messages.
C.Switch the target to a Lambda function with reserved concurrency of zero during outages.
D.Replace EventBridge with CloudWatch Logs subscriptions so the consumer can poll the log stream later.
AnswerB

SQS is the right durability and buffering layer for this requirement. EventBridge can publish orders.checkout events to a queue, and workers can consume them at a controlled rate even when the payment API is unavailable. This decouples event ingestion from downstream processing, absorbs bursts, and preserves events until the outage ends. A DLQ provides a safe landing zone for messages that continue to fail after retries so they are not silently dropped.

Why this answer

Introducing an SQS queue between EventBridge and the consumer decouples the event delivery from the downstream payment API. During outages, events are stored durably in SQS and can be processed later without being lost. A Dead Letter Queue (DLQ) captures events that fail repeatedly, preventing poison messages from blocking the queue and ensuring no events age out due to retry exhaustion.

Exam trap

The trap here is that candidates often assume increasing timeouts or concurrency adjustments can fix backpressure issues, but they fail to recognize that decoupling with a durable queue is the only way to preserve events during extended downstream outages without losing them to retry expiration.

Why the other options are wrong

A

Increasing Lambda timeout does not address the root cause of downstream payment API timeouts; it only makes the Lambda wait longer, potentially exacerbating backpressure and event aging without improving resilience.

C

Setting reserved concurrency to zero during outages would stop all invocations, causing all events to be lost or retried until they age out, rather than improving resilience or preserving events.

D

CloudWatch Logs subscriptions deliver log data to a consumer in near-real-time but do not provide a durable buffer or retry mechanism for downstream failures; events would still be lost if the consumer is unavailable, and there is no built-in DLQ for poison messages.

161
MCQhard

A company runs a stateful web application on a single Amazon EC2 instance in a public subnet. The application stores session data on the instance's root volume. The company wants to make the application highly available across two Availability Zones and ensure that session data is preserved if an instance fails. Which solution should a solutions architect recommend?

A.Create an Amazon Machine Image (AMI) of the instance, launch two new instances in two Availability Zones, and configure an Application Load Balancer with sticky sessions.
B.Move the session data to an Amazon ElastiCache for Redis cluster with Multi-AZ enabled, and configure an Auto Scaling group across two Availability Zones behind an Application Load Balancer.
C.Attach an Amazon Elastic Block Store (EBS) volume to the instance and create a snapshot schedule. Launch a second instance in another Availability Zone and attach the same EBS volume.
D.Enable termination protection on the instance and create a scheduled AWS Lambda function to reboot the instance if it becomes unhealthy.
AnswerB

Storing session data in ElastiCache for Redis with Multi-AZ enabled externalizes the state, allowing any instance in the Auto Scaling group to access it. The Auto Scaling group across two Availability Zones ensures high availability, and the Application Load Balancer distributes traffic. If an instance fails, a new one launches and retrieves session data from ElastiCache.

Why this answer

The recommended solution is to externalize session data to an ElastiCache for Redis cluster with Multi-AZ enabled and run the application in an Auto Scaling group across two Availability Zones behind an Application Load Balancer. This design ensures that session data survives instance failures and the application remains available if one AZ fails. The load balancer distributes traffic to healthy instances.

Exam trap

The trap here is assuming that sticky sessions or EBS snapshots can preserve session state across instance failures, when session data stored on an instance's root volume is lost if the instance fails.

162
MCQmedium

A healthcare company runs a critical patient-records API on Amazon EC2 instances behind an Application Load Balancer in a single AWS Region. The compliance team mandates that the API remain available even if an entire AWS Region becomes unavailable. The company wants a cost-effective solution that avoids running full production capacity in a second Region at all times. Which approach BEST meets these requirements?

A.Deploy the API in a second Region with a minimal warm standby environment, and use Amazon Route 53 failover routing with health checks to shift traffic when the primary Region fails.
B.Create an Amazon CloudFront distribution with the existing ALB as the origin, and enable origin failover to a second ALB in another Region.
C.Take regular Amazon EBS snapshots of the EC2 instances and copy them to a second Region, then restore the instances manually if the primary Region fails.
D.Enable Multi-AZ deployment for the EC2 instances and the Application Load Balancer, and rely on the existing single-Region architecture to survive a Regional failure.
AnswerA

A warm standby keeps a scaled-down but functional stack in the secondary Region, which can be scaled up during a failover. Route 53 failover routing with health checks automatically redirects traffic when the primary endpoint becomes unhealthy. This balances cost and recovery time, meeting the cross-Region availability requirement without paying for full duplicate capacity.

Why this answer

A warm standby in a second Region with Route 53 failover routing provides cross-Region resilience while controlling cost. The standby runs at reduced capacity and can be scaled up during a Regional failure. Health checks trigger automatic DNS failover, so clients reach the surviving Region without manual intervention.

This design directly addresses the need to survive a full Region outage without duplicating full production capacity at all times.

Exam trap

The trap here is assuming that Multi-AZ deployment provides protection against a complete AWS Region failure, when it only protects against Availability Zone failures within a single Region.

163
MCQmedium

A company runs a stateful analytics workload on EC2 instances that use EBS volumes. The data must be restorable in another Region after a major outage, with frequent point-in-time recovery. Which approach provides the most suitable replication mechanism for the EBS-backed data?

A.Create scheduled EBS snapshots and copy them to another Region, then restore the volumes from those snapshots during recovery.
B.Enable EBS multi-attach to spread the workload across AZs and replicate snapshots automatically between Regions.
C.Use RDS read replicas in another Region and keep the analytics dataset in an RDS instance only.
D.Rely on instance store for durability and copy only AMIs across Regions.
AnswerA

Scheduled EBS snapshots copied cross-Region provide point-in-time recovery and durability in a second Region, satisfying the restore-elsewhere requirement. Snapshots capture only changed blocks after the first, keeping replication efficient, and restoration creates fresh volumes from any snapshot in the target Region.

Why this answer

Scheduled EBS snapshots provide point-in-time backups of EBS volumes, which can be copied to another Region using the cross-Region snapshot copy feature. During recovery, you restore volumes from those snapshots in the target Region, ensuring the data is restorable after a major outage. This approach meets the requirements for frequent point-in-time recovery and cross-Region durability.

Exam trap

The trap here is that candidates may confuse EBS multi-attach (which is for high availability within a single AZ) with cross-Region replication, or mistakenly think instance store provides durability for long-term data recovery.

Why the other options are wrong

B

EBS multi-attach allows attaching a volume to multiple EC2 instances in the same AZ, but it does not replicate snapshots across Regions or provide cross-Region disaster recovery. It is designed for clustered applications within a single AZ, not for multi-Region replication.

C

RDS read replicas are for relational databases, not for analytics workloads on EC2 with EBS volumes. The question specifies EBS-backed data, not RDS-managed data, so using RDS would require migrating the dataset and does not replicate EBS snapshots.

D

Instance store volumes are ephemeral and lose data on instance stop/termination, making them unsuitable for durable, restorable data. Copying AMIs does not replicate the analytics data itself.

164
MCQmedium

A company hosts a web application on EC2 instances behind an Application Load Balancer (ALB) in us-east-1. A static failover site is hosted in an S3 bucket with static website hosting enabled. The company needs automatic DNS failover to the S3 bucket if the primary ALB becomes unhealthy. Which Route 53 configuration achieves this?

A.Configure Route 53 Failover routing with a health check on the ALB as PRIMARY and the S3 bucket website endpoint as SECONDARY
B.Configure Route 53 Weighted routing with 100% weight on the ALB and 0% on the S3 bucket
C.Configure Route 53 Latency routing with records in both regions to route to the healthiest endpoint
D.Configure Route 53 Geolocation routing with North American users directed to the ALB and all others to S3
AnswerA

Failover routing is the correct Route 53 policy for active-passive architecture, where the ALB is the primary endpoint and the S3 static website endpoint is the secondary. You configure a primary record pointing to the ALB, associate it with a Route 53 health check that monitors the ALB, and create a secondary record pointing to the S3 website endpoint. When the health check fails, Route 53 automatically returns the S3 endpoint, providing DNS-level failover. The S3 bucket must be configured for static website hosting with a publicly accessible endpoint, and this setup requires no manual intervention to switch traffic.

Why this answer

Route 53 Failover routing uses health checks to route traffic to a primary resource and automatically switch to a secondary when the primary health check fails.

Configuration: Create a Route 53 health check targeting the ALB endpoint. Create a PRIMARY alias A record pointing to the ALB with the health check associated. Create a SECONDARY alias A record pointing to the S3 static website endpoint. When the ALB health check fails, Route 53 returns the S3 endpoint automatically.

Exam trap

Route 53 offers multiple routing policies. Failover routing is active-passive — one primary resource, one standby. Weighted routing splits traffic percentages (active-active).

Latency routing picks the lowest-latency endpoint. Geolocation routes by user geography. Only Failover routing provides automatic primary/secondary switchover based on health checks.

Weighted routing at 100%/0% does NOT failover when the 100% target fails.

Why the other options are wrong

B

Weighted routing at 100%/0% does not failover. When the 100% target (ALB) is unhealthy, Route 53 does not automatically redirect to the 0% target (S3). Weighted routing splits traffic by percentage without health-check-based switching.

C

Latency routing routes to the lowest-latency endpoint for each client. It does not implement primary/secondary logic. If us-east-1 is unhealthy, some clients may still be routed there unless combined with health checks (but even then this is not a defined primary/secondary failover).

D

Geolocation routing directs traffic by user geography — North American users always go to the ALB even when it fails. S3 only receives other-region traffic. This is not a failover configuration.

165
MCQhard

A company runs a critical application on Amazon EC2 instances in a single Availability Zone. The application writes data to an Amazon RDS for MySQL DB instance that is not Multi-AZ. The company wants to improve the resilience of the database tier so that it can survive an Availability Zone failure with minimal downtime and no data loss. The application uses the database endpoint from the RDS console. Which solution meets these requirements?

A.Create a read replica in another Availability Zone and manually promote it if the primary fails.
B.Convert the DB instance to a Multi-AZ deployment and update the application to use the same endpoint.
C.Enable automated backups with a retention period of 35 days and copy the backups to another Region.
D.Migrate the database to Amazon DynamoDB with global tables and update the application to use the DynamoDB endpoint.
AnswerB

Converting to Multi-AZ creates a synchronous standby in a different AZ. RDS automatically fails over to the standby in the event of an AZ failure, and the application can continue using the same endpoint. This provides minimal downtime and no data loss because replication is synchronous.

Why this answer

Multi-AZ deployments for Amazon RDS provide a synchronous standby replica in a different Availability Zone. In the event of an AZ failure, RDS automatically fails over to the standby, and the application continues to use the same endpoint. Because replication is synchronous, there is no data loss.

This meets the requirements of minimal downtime and no data loss without application changes.

Exam trap

The trap here is confusing read replicas with Multi-AZ standbys. Read replicas are for offloading reads and can be promoted manually, but they use asynchronous replication and do not provide automatic failover.

166
MCQeasy

A web application runs on an Amazon EC2 Auto Scaling group (ASG) behind an Application Load Balancer (ALB). The ALB is configured to use at least two Availability Zones (AZs), but the ASG currently uses subnets in only one AZ. If that AZ becomes unavailable, the application stops serving requests. Which change most directly improves resilience to an AZ outage?

A.Keep the ASG in one Availability Zone, but reduce ALB health check intervals.
B.Place the ASG across multiple Availability Zones by configuring it with subnets in at least two AZs.
C.Switch the load balancer from an ALB to an NLB to remove HTTP health check dependency.
D.Add an Amazon SQS queue to buffer requests during failures.
AnswerB

An ASG launches instances into the AZs of the subnets you specify. By placing the ASG in at least two AZs, the ALB can route traffic to healthy targets in the remaining AZ(s) if one AZ fails, enabling recovery as new instances maintain desired capacity.

Why this answer

Distributing an Auto Scaling group across multiple Availability Zones (AZs) ensures that if one AZ fails, the remaining AZs continue to serve traffic. The Application Load Balancer (ALB) is already configured for at least two AZs, but the ASG’s single-AZ subnet placement creates a single point of failure. By adding subnets in at least two AZs to the ASG, the application becomes resilient to an AZ outage without any other architectural changes.

Exam trap

The trap here is that candidates assume the ALB’s multi-AZ configuration automatically protects the application, overlooking that the ASG must also span multiple AZs to provide compute redundancy.

How to eliminate wrong answers

Option A is wrong because reducing health check intervals only detects failures faster but does not eliminate the single point of failure; if the sole AZ becomes unavailable, no healthy instances exist to serve traffic. Option C is wrong because switching from an ALB to an NLB does not address the root cause—the ASG is still in one AZ—and HTTP health checks are not the issue; the ALB can already perform health checks across AZs. Option D is wrong because adding an SQS queue buffers requests but does not provide compute capacity in another AZ; without instances in a second AZ, the queue cannot process requests during an AZ outage.

167
MCQhard

A healthcare analytics platform ingests records into an Amazon Aurora MySQL cluster. Compliance rules require that the cluster remain writable even if an entire Availability Zone is lost, and that recovery happen without operator action. The team also wants read traffic to scale independently of the writer. Which configuration should a solutions architect choose?

A.Create an Aurora MySQL cluster with one writer instance and two Aurora Replicas distributed across three Availability Zones, and connect the application to the cluster reader endpoint for read queries and the cluster endpoint for writes.
B.Create an Aurora MySQL cluster with one writer instance and enable Aurora Global Database with a secondary cluster in a different Region, directing all reads to the secondary Region's reader endpoint.
C.Create an Aurora MySQL cluster with one writer instance, and configure an Amazon Route 53 failover record that points to a standby cluster in the same Region when health checks fail.
D.Create an Aurora MySQL cluster with a single writer instance and one reader instance in the same Availability Zone, and enable automated backups with a 35-day retention period.
AnswerA

Aurora stores six copies of data across three Availability Zones, and distributing the writer plus two Aurora Replicas across those zones means a zone failure leaves surviving replicas available. Aurora automatically promotes a replica to writer and updates the cluster endpoint, so the application reconnects without manual steps, while the reader endpoint load-balances read queries across the replicas.

Why this answer

Aurora's storage layer replicates each data volume six ways across three Availability Zones, and placing the writer plus two Aurora Replicas in separate zones means a zone failure does not remove all compute endpoints. Aurora automatically promotes a surviving replica and repoints the cluster endpoint, so writes resume with no operator involvement, while the reader endpoint spreads read queries across the replicas for independent read scaling.

Exam trap

The trap here is treating automated backups or a Route 53 failover record as Availability Zone resilience, when both depend on a manual or external recovery path rather than Aurora's own automatic replica promotion.

168
MCQeasy

An order-processing system publishes an event whenever a payment succeeds. Three downstream services (inventory, shipping, and analytics) must react independently. Analytics sometimes has high latency, but order processing must not be blocked. What is the best AWS approach to decouple these consumers?

A.Have order processing call each service synchronously via HTTPS and retry on failures.
B.Publish payment events to SNS (or EventBridge) and let each downstream service consume independently (for example, via SQS queues or other async targets).
C.Store events in a single relational database table and let consumers poll continuously for new rows.
D.Send events directly from the producer to each consumer EC2 instance using SSH tunnels.
AnswerB

Using pub/sub decouples the producer from consumers. Order processing publishes once and can complete without waiting for each downstream service. Each consumer receives events independently, so analytics latency does not directly block inventory or shipping processing.

Why this answer

Amazon SNS (or EventBridge) enables asynchronous, fan-out messaging where a single payment-success event is published once and delivered independently to multiple downstream services (inventory, shipping, analytics) via SQS queues or other targets. This decouples the producer from consumer latency—analytics can take its time without blocking order processing—and ensures each consumer processes the event at its own pace, meeting the requirement for independent, non-blocking reactions.

Exam trap

The trap here is that candidates may choose synchronous integration (Option A) because it seems simpler, failing to recognize that the requirement 'must not be blocked' explicitly demands asynchronous decoupling, not just retries.

How to eliminate wrong answers

Option A is wrong because synchronous HTTPS calls with retries tightly couple the producer to all consumers; if analytics has high latency, order processing is blocked waiting for responses, violating the requirement that it must not be blocked. Option C is wrong because storing events in a single relational database table introduces a single point of failure, creates a polling bottleneck, and tightly couples consumers to a shared schema and table, which is not a decoupled, scalable architecture. Option D is wrong because sending events directly via SSH tunnels requires direct network connectivity to each EC2 instance, introduces security risks, and tightly couples the producer to consumer instances, making it brittle and unscalable.

169
MCQmedium

An internal worker consumes messages from an Amazon SQS Standard queue. Recently, some messages fail validation in the worker (for example, missing required fields), causing the worker to crash before it can successfully process those messages. Those messages keep getting retried repeatedly, slowing down processing of valid messages. The team wants a resilient mechanism to quarantine bad messages after a limited number of receive attempts. What should they implement?

A.Increase the SQS visibility timeout to several hours so the worker does not retry too quickly.
B.Configure a redrive policy with a Dead-Letter Queue (DLQ) and set maxReceiveCount so poison messages are moved to the DLQ after repeated failures.
C.Switch the queue to an SNS topic and subscribe the worker directly, eliminating message retries.
D.Enable KMS encryption with a new CMK to ensure validation errors stop occurring.
AnswerB

An SQS DLQ with a redrive policy is specifically designed for poison-message handling. When a message exceeds maxReceiveCount without successful processing (for example, the worker crashes before deletion), SQS moves the message to the DLQ. This quarantines bad messages and protects throughput for valid messages.

Why this answer

Amazon SQS supports configuring a redrive policy with a Dead-Letter Queue (DLQ) that automatically moves messages after a specified number of receive attempts (maxReceiveCount). This isolates poison messages that fail validation and cause crashes, preventing them from being retried indefinitely and slowing down valid message processing. The worker can then focus on valid messages while the DLQ stores the problematic ones for later analysis or manual intervention.

Exam trap

The trap here is that candidates may think increasing the visibility timeout (Option A) solves the retry problem, but it only delays retries without eliminating the root cause, while the DLQ mechanism (Option B) provides a proper quarantine by moving messages after a configurable number of receive attempts.

How to eliminate wrong answers

Option A is wrong because increasing the visibility timeout to several hours would only delay retries, not prevent them; the worker would still crash repeatedly on the same invalid messages after each timeout expires, and valid messages would be blocked for hours. Option C is wrong because switching to an SNS topic eliminates message retries entirely, but the worker would still crash on invalid messages without any retry mechanism or quarantine, and SNS does not provide a built-in DLQ for consumer-side failures. Option D is wrong because enabling KMS encryption with a new CMK addresses data encryption at rest and in transit, but has no effect on message content validation errors or crash handling; encryption does not fix missing required fields or prevent retries.

170
MCQhard

A logistics company stores shipment events in an Amazon S3 bucket. An analytics team must be able to recover any object version that is accidentally overwritten or deleted for at least 90 days, and objects must be protected from permanent deletion by any user, including the root user, during that window. Which combination of S3 features meets these requirements with the LEAST operational overhead?

A.Enable S3 Versioning and configure S3 Cross-Region Replication to a second bucket in another Region with a 90-day lifecycle expiration.
B.Enable S3 Versioning and configure an S3 Lifecycle rule to transition noncurrent versions to S3 Glacier Deep Archive after 90 days.
C.Enable S3 Versioning and apply an S3 Object Lock retention period of 90 days in compliance mode on the bucket.
D.Enable S3 Versioning and add a bucket policy that denies s3:DeleteObject for all principals for 90 days.
AnswerC

S3 Object Lock with a 90-day compliance-mode retention prevents any principal, including the root user, from deleting or overwriting the protected object versions until the retention expires. Versioning is required for Object Lock, and compliance mode provides the strongest immutability, directly satisfying both recovery and permanent-deletion-prevention requirements with minimal ongoing effort.

Why this answer

S3 Object Lock requires versioning and enforces a retention period during which object versions cannot be deleted or overwritten by anyone, including the root user. Compliance mode cannot be shortened or bypassed, so it uniquely satisfies the requirement to prevent permanent deletion for at least 90 days while still allowing recovery of prior versions.

Exam trap

The trap here is thinking that versioning plus a restrictive bucket policy or replication equals immutability, when only Object Lock retention actually blocks deletion by the root user.

171
MCQmedium

A ticket booking system runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include? The design must avoid adding custom operational scripts.

A.Subnets in at least two Availability Zones with health checks enabled
B.All instances in one larger subnet
C.A Network Load Balancer in one subnet
D.A single EC2 instance with detailed monitoring
AnswerA

An Auto Scaling group configured with subnets in at least two Availability Zones ensures that if an entire AZ becomes unavailable, the ASG can launch replacement instances in another AZ, preserving capacity. Health checks on the load balancer and EC2 allow the ASG to detect failed instances and replace them automatically. This architecture eliminates the single point of failure that comes with a single AZ and provides fault tolerance for both planned maintenance and unexpected outages.

Why this answer

Placing subnets in at least two Availability Zones ensures that if one AZ fails, the Auto Scaling group can launch instances in the remaining healthy AZ, maintaining application availability. Health checks integrated with the Application Load Balancer allow the Auto Scaling group to automatically replace unhealthy instances without custom scripts, aligning with the requirement to avoid operational overhead.

Exam trap

The trap here is that candidates often assume a single larger subnet or a Network Load Balancer provides AZ resilience, but they fail to recognize that without multiple subnets in distinct AZs, the architecture cannot survive an AZ failure, and custom scripts would be needed for health checks without ELB integration.

How to eliminate wrong answers

Option B is wrong because placing all instances in one larger subnet confines them to a single Availability Zone, violating the requirement to tolerate the failure of one AZ. Option C is wrong because a Network Load Balancer operates at Layer 4 and does not provide the health check integration needed for Auto Scaling group instance replacement; additionally, placing it in one subnet creates a single point of failure. Option D is wrong because a single EC2 instance, even with detailed monitoring, cannot survive an AZ failure and does not leverage Auto Scaling for automatic recovery.

172
MCQmedium

A company uses an Amazon Aurora DB cluster in a Multi-AZ configuration. During a planned failover of the writer instance, the database endpoints in the application are updated incorrectly. After failover, reads work but writes fail with connection errors and timeouts for several minutes. The team currently uses the instance endpoint for the writer. What should they change to improve write resilience during failovers?

A.Continue using the instance endpoint, but increase application retry count so the writer changes are handled more quickly.
B.Use the Aurora cluster writer endpoint for all write operations.
C.Use a read replica endpoint for writes because it is typically stable across failovers.
D.Disable Multi-AZ failover so the writer instance never changes and writes remain consistent.
AnswerB

Aurora provides a writer endpoint designed specifically for write traffic. During failover, Aurora updates where the writer endpoint points, so the same DNS name continues to resolve to the current writer instance without requiring manual endpoint changes in the application.

Why this answer

The Aurora cluster writer endpoint always points to the current primary (writer) instance, even after a failover. By using this endpoint instead of a static instance endpoint, the application automatically resolves to the new writer without manual updates, eliminating connection errors and timeouts during failover transitions.

Exam trap

The trap here is that candidates confuse the instance endpoint (which is static and tied to a specific instance) with the cluster endpoint (which is dynamic and always points to the current writer), assuming any endpoint will automatically follow failover.

How to eliminate wrong answers

Option A is wrong because increasing the retry count does not fix the root cause—the application is still pointing to the old (now read-only) instance endpoint, so writes will continue to fail until the endpoint is manually corrected. Option C is wrong because read replica endpoints point to read-only instances; writes to a read replica will always fail with an error, regardless of failover state. Option D is wrong because disabling Multi-AZ failover removes high availability entirely, making the database vulnerable to a single point of failure, which contradicts the goal of improving write resilience.

173
Multi-Selecthard

A claims workflow requires point-in-time recovery and accidental-delete protection for a DynamoDB table. Which two settings should the architect enable? The design must avoid adding custom operational scripts.

Select 2 answers
A.Point-in-time recovery
B.DAX
C.Deletion protection or tightly controlled delete permissions
D.Global secondary indexes
AnswersA, C

PITR allows restoration to a specific second within the supported recovery window.

Why this answer

Point-in-time recovery (PITR) for DynamoDB enables continuous backups with 35-day granularity, allowing restoration to any second within that window. This directly satisfies the point-in-time recovery requirement without custom scripts, as it is a native AWS feature.

Exam trap

The trap here is that candidates often confuse DAX with a data protection feature, but DAX only accelerates reads and has no role in backup or deletion prevention.

174
MCQmedium

A trading dashboard stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured? The design must avoid adding custom operational scripts.

A.An EBS snapshot schedule
B.S3 Cross-Region Replication with versioning enabled
C.S3 lifecycle transition to Glacier Flexible Retrieval
D.A CloudFront distribution
AnswerB

Cross-Region Replication automatically copies objects to a destination bucket in another Region, and versioning is a prerequisite for it. This meets the disaster-recovery copy requirement without custom operational scripts, since replication is handled natively by S3.

Why this answer

S3 Cross-Region Replication (CRR) automatically replicates objects to a destination bucket in a different AWS Region, meeting the disaster recovery requirement without custom scripts. Versioning must be enabled on both source and destination buckets for CRR to function, as it tracks object versions and ensures consistency during replication.

Exam trap

The trap here is that candidates may confuse S3 Cross-Region Replication with S3 lifecycle policies or CloudFront, thinking they provide cross-region replication, but only CRR with versioning enabled meets the DR requirement without custom scripts.

How to eliminate wrong answers

Option A is wrong because EBS snapshots are for Amazon Elastic Block Store volumes attached to EC2 instances, not for S3 objects; they cannot replicate S3 data across regions. Option C is wrong because S3 lifecycle transitions to Glacier Flexible Retrieval only change storage class within the same region for cost optimization, not replicate data to another region. Option D is wrong because CloudFront is a content delivery network (CDN) that caches content at edge locations for low-latency access, not a replication mechanism for disaster recovery across regions.

175
MCQmedium

An order processing workflow uses Amazon SQS as the decoupling layer between a producer and a consumer Lambda function. The consumer intermittently fails due to a downstream dependency. The team has observed that certain “poison” messages keep being retried repeatedly and prevent other messages from being processed efficiently. Which SQS configuration most directly addresses this issue?

A.Set the SQS queue’s retention period to 10 years and rely on application retries to eventually succeed.
B.Increase visibility timeout to a very large value and avoid dead-letter queues to keep ordering stable.
C.Configure a redrive policy with a dead-letter queue (DLQ) and set an appropriate visibility timeout greater than the maximum processing time.
D.Switch the queue to FIFO and remove retries in the Lambda event source mapping entirely.
AnswerC

A redrive policy defines a dead-letter queue (DLQ) and a maxReceiveCount; once a message is received that many times without being deleted, SQS moves it to the DLQ, quarantining poison messages for inspection or manual redrive. Setting the visibility timeout longer than the worst-case processing time prevents the message from becoming visible again while a consumer is still working, which would otherwise cause duplicate deliveries. Together, these settings bound both the retry window and the queue depth, allowing transient failures to retry while isolating permanent failures without losing data.

Why this answer

Configuring a redrive policy with a dead-letter queue (DLQ) allows messages that repeatedly fail processing to be moved out of the main queue after a specified number of receive attempts. Setting an appropriate visibility timeout greater than the maximum processing time ensures that messages are not made visible again before the consumer finishes processing, preventing premature retries. This directly isolates poison messages so they no longer block the processing of other messages in the queue.

Exam trap

The trap here is that candidates may think increasing visibility timeout or switching to FIFO alone will handle failed messages, but without a DLQ, poison messages remain in the queue and continue to block other messages, which is the core issue described.

Why the other options are wrong

A

Setting the retention period to 10 years does not address poison messages; it only keeps messages longer. Relying on application retries without a DLQ allows poison messages to be retried indefinitely, blocking other messages.

B

Increasing visibility timeout to a very large value does not prevent poison messages from blocking the queue; they will still be retried indefinitely, and without a DLQ, failed messages cannot be isolated for analysis or skipped.

D

Switching to FIFO and removing retries does not address poison messages; FIFO ensures strict ordering but does not prevent problematic messages from blocking the queue, and removing retries would cause immediate failures without handling the root cause.

176
MCQeasy

A startup runs a stateless web application on a single Amazon EC2 instance in one Availability Zone. The application must remain available if the instance fails or if its Availability Zone becomes unavailable. The startup wants a managed solution that requires minimal operational overhead. Which solution should a solutions architect recommend?

A.Launch a second EC2 instance in the same Availability Zone and configure an Elastic IP address for failover.
B.Create an Auto Scaling group spanning at least two Availability Zones and attach it to an Application Load Balancer.
C.Migrate the application to AWS Lambda and expose it through an API Gateway REST API.
D.Create an Auto Scaling group with a desired capacity of one instance and place it in a single Availability Zone.
AnswerB

An Auto Scaling group spanning multiple Availability Zones automatically replaces failed instances and can launch replacements in healthy zones when one zone fails. Attaching an Application Load Balancer distributes traffic across healthy instances and provides a single stable endpoint. This is a managed, low-overhead solution that meets both instance and zone failure requirements.

Why this answer

An Auto Scaling group that spans at least two Availability Zones can replace a failed instance and can also launch instances in a healthy zone when one zone becomes unavailable. Attaching an Application Load Balancer provides a stable endpoint and health-based routing, delivering a managed, resilient architecture with minimal operational effort.

Exam trap

The trap here is believing that a single-zone Auto Scaling group provides Availability Zone resilience, when it can only replace instances within the same zone.

177
MCQeasy

A system processes events from Amazon SQS and sometimes sees duplicate messages due to retries. The business requirement is that each payment must be charged at most once. What design choice best addresses this resiliency requirement?

A.Assume duplicates never occur because the consumer deletes messages immediately after receiving them.
B.Implement idempotent processing using a deduplication key (for example, paymentId) and record completed charges so duplicates are safely ignored.
C.Increase the SQS visibility timeout until duplicates never happen.
D.Use SNS topics instead of SQS so retries are disabled by default.
AnswerB

Idempotency ensures at-most-once side effects even when duplicates are delivered. Persist a record keyed by paymentId (e.g., a unique constraint/conditional write). If the record indicates the payment was already charged, skip the charge for any subsequent duplicate message.

Why this answer

Implementing idempotent processing with a deduplication key (e.g., paymentId) ensures that even if duplicate messages arrive from SQS (due to retries or at-least-once delivery), the consumer can check a record of completed charges and safely ignore duplicates. This satisfies the business requirement of charging each payment at most once without relying on SQS’s best-effort deduplication or message ordering.

Exam trap

The trap here is that candidates assume SQS guarantees exactly-once delivery or that increasing visibility timeouts can prevent duplicates, but SQS is designed for at-least-once delivery, and the only reliable way to handle duplicates is to make the consumer idempotent.

How to eliminate wrong answers

Option A is wrong because SQS provides at-least-once delivery, and deleting a message immediately after receiving it does not prevent duplicates that may arrive before the delete is processed or due to visibility timeout expiration; assuming duplicates never occur violates the fundamental reliability guarantee of SQS. Option C is wrong because increasing the visibility timeout cannot eliminate duplicates; it only delays the redelivery of unacknowledged messages, and duplicates can still occur due to network retries, consumer crashes, or SQS’s internal replication. Option D is wrong because SNS topics do not disable retries by default; SNS uses at-least-once delivery and can retry HTTP/S endpoints, and switching to SNS does not solve the duplicate problem—it may even introduce additional delivery attempts without built-in deduplication.

178
MCQeasy

A company hosts a web application on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer (ALB). The ALB and the Auto Scaling group are currently deployed in only one Availability Zone (AZ). The business wants the application to keep running if that AZ has an outage. What is the best change?

A.Increase the desired capacity in the existing Availability Zone to handle all traffic during an outage.
B.Deploy the ALB and the Auto Scaling group across at least two Availability Zones so healthy targets remain.
C.Enable longer ALB health check intervals so failing instances are detected more slowly.
D.Switch from the ALB to an Internet Gateway so instances can fail over to the public internet.
AnswerB

To tolerate an AZ outage, both the load-balancing entry point (the ALB) and the compute capacity (the Auto Scaling instances) must be available in more than one AZ. With the ALB in multiple AZs and the Auto Scaling group using multiple subnets/AZs, requests can be routed to healthy targets in a remaining AZ while Auto Scaling replaces unhealthy instances.

Why this answer

Deploying the ALB and Auto Scaling group across at least two Availability Zones (AZs) ensures that if one AZ fails, the ALB can route traffic to healthy EC2 instances in the remaining AZ(s). This is the fundamental AWS best practice for high availability: an ALB is a regional service that requires targets in multiple AZs to survive an AZ outage, and the Auto Scaling group must also span those AZs to maintain capacity. Without multi-AZ deployment, a single AZ failure makes the entire application unavailable regardless of instance health checks.

Exam trap

The trap here is that candidates think increasing capacity or adjusting health check intervals can compensate for a single-AZ deployment, but AWS high availability fundamentally requires distributing resources across multiple isolated failure domains (AZs).

How to eliminate wrong answers

Option A is wrong because increasing the desired capacity in a single AZ does not protect against an AZ outage; all instances are in the same failure domain, so they all become unreachable simultaneously. Option C is wrong because enabling longer health check intervals would delay detection of failing instances, making the application less responsive to failures and increasing downtime, not improving availability. Option D is wrong because an Internet Gateway (IGW) is a VPC component that enables outbound internet access for instances, not a load balancer; it cannot perform health checks, distribute traffic, or fail over traffic between instances, and it does not replace the ALB's role in high availability.

179
MCQmedium

A web application runs on an Amazon EC2 Auto Scaling group behind an Application Load Balancer (ALB). After each deployment, new instances take about 2 minutes to download artifacts and become ready to accept requests on the target port. In the last deployment, the ALB started marking targets unhealthy before the app was ready, and the Auto Scaling group then replaced those instances repeatedly, causing a prolonged outage. Which change best improves resilience during instance start-up without reducing actual availability once the application is healthy?

A.Increase the Auto Scaling group’s health check grace period so it exceeds the ~2-minute initialization time.
B.Add more subnets across additional Availability Zones to distribute the same instances more widely.
C.Switch the load balancer target type from instance targets to IP targets to avoid health check failures.
D.Reduce the ALB health check interval so unhealthy targets are removed faster.
AnswerA

A health check grace period prevents the Auto Scaling group from treating early health check failures as instance health problems. This avoids terminating instances before the application finishes initializing, which stops the restart/replace loop during deployments while still allowing normal health checks to apply once the app is ready.

Why this answer

The Auto Scaling group's health check grace period allows instances to initialize without being marked unhealthy by the ELB health checks. By setting this grace period to exceed the ~2-minute artifact download time, the ASG will not replace instances that are still starting up, preventing the cascade of terminations and redeployments that caused the outage. This directly addresses the root cause—premature health check failures—without changing the health check configuration or reducing availability once the app is ready.

Exam trap

The trap here is that candidates confuse the ALB health check interval or target type with the Auto Scaling group's lifecycle management, mistakenly thinking that changing how the ALB checks health (interval or target type) will fix the premature replacement, when the correct solution is to adjust the ASG's grace period to align with the application's startup time.

How to eliminate wrong answers

Option B is wrong because adding more subnets across additional Availability Zones distributes instances more widely for fault tolerance but does not prevent the ALB from marking starting instances as unhealthy, so it does not solve the premature replacement issue. Option C is wrong because switching from instance targets to IP targets changes how the ALB routes traffic but does not alter the health check logic or timing; the ALB will still mark the target as unhealthy if the health check fails during the initialization window. Option D is wrong because reducing the ALB health check interval causes unhealthy targets to be detected and removed faster, which would worsen the problem by accelerating the replacement cycle, not improving resilience during start-up.

180
Multi-Selectmedium

An internal API is deployed in two AWS Regions behind separate Application Load Balancers. The company wants clients to use the primary Region when it is healthy and automatically switch to the secondary Region if the primary health check fails. Which two Route 53 record configurations are required? Select two.

Select 2 answers
A.Create a primary failover record that points to the primary ALB and associates a Route 53 health check.
B.Create a weighted record set that sends 50 percent of traffic to each Region.
C.Create a secondary failover record that points to the secondary ALB.
D.Create a latency-based record set so Route 53 always prefers the fastest Region.
E.Create a multivalue answer record to return both ALB addresses on each lookup.
AnswersA, C

A primary failover record designates the active endpoint in an active-passive configuration. Route 53 associates a health check with this record, and as long as the check returns healthy, every DNS response returns the primary ALB's IP address. When that health check fails, Route 53 stops serving the primary record and instead returns the secondary failover record, enabling automatic regional failover.

Why this answer

A primary failover record in Amazon Route 53 directs traffic to the primary ALB and is associated with a Route 53 health check. If the health check fails, Route 53 automatically fails over to the secondary failover record, ensuring high availability across Regions.

Exam trap

The trap here is that candidates often confuse failover routing with weighted or latency routing, assuming any health-aware routing provides automatic primary/secondary failover, but only failover records enforce a strict active-passive pattern.

181
MCQmedium

A SaaS platform plans to run in two AWS Regions for lower latency. The team wants to enable active-active writes (both regions accept updates) to avoid failover downtime. However, the business requires strong consistency for order status transitions (for example, only one transition from “Paid” to “Shipped” must be allowed). Which statement is the best architectural choice to meet the consistency requirement?

A.Use active-active writes only when the workload tolerates eventual consistency; for strongly consistent transitions, use a single-writer pattern with failover (active-passive/pilot light).
B.Active-active writes always provide strong consistency because AWS replicates data across Regions automatically and immediately.
C.Active-active writes can be used safely by simply enabling retries and expecting the application to resolve conflicts without coordination.
D.To ensure strong consistency, run both Regions with different IAM roles and block cross-Region writes at the API layer only.
AnswerA

Correct. Active-active multi-Region writes rely on asynchronous cross-Region replication (e.g., DynamoDB Global Tables or Aurora Global Database), so a write in one Region is not immediately visible in the other; concurrent updates to the same item can conflict and require last-write-wins or custom conflict resolution, which is inherently eventually consistent. Strongly consistent transitions, such as a transaction that must read its own write or maintain a linearizable order, cannot be guaranteed with multiple concurrent writers. A single-writer pattern (active-passive or pilot light) ensures only one Region accepts writes at a time; failover transfers write authority to the secondary Region, preserving strong consistency because there is never more than one authoritative writer.

Why this answer

Active-active writes across AWS Regions cannot guarantee strong consistency due to the inherent latency and lack of synchronous replication between Regions. For order status transitions that require exactly-once semantics (e.g., only one transition from 'Paid' to 'Shipped'), a single-writer pattern (active-passive or pilot light) ensures that only one Region accepts writes at a time, avoiding conflicts and maintaining a single source of truth. AWS services like DynamoDB global tables offer eventual consistency for multi-region writes, while Aurora Global Database provides read replicas with failover but not active-active writes for strong consistency.

Exam trap

The trap here is that candidates assume AWS's global services (like DynamoDB global tables or Aurora Global Database) inherently provide strong consistency for multi-region writes, when in fact they are designed for eventual consistency and require careful trade-offs for strict ordering requirements.

Why the other options are wrong

B

AWS does not replicate data across Regions automatically or immediately; cross-Region replication is asynchronous, so active-active writes cannot guarantee strong consistency.

C

Active-active writes without coordination cannot guarantee strong consistency for order status transitions; retries and application-level conflict resolution are insufficient to prevent two regions from simultaneously accepting conflicting transitions (e.g., 'Paid' to 'Shipped' in both regions).

D

Blocking cross-Region writes at the API layer with different IAM roles does not prevent concurrent writes from being accepted in both regions before the API layer can reject them, so it cannot guarantee strong consistency for order status transitions.

182
MCQeasy

A consumer application reads from an Amazon SQS queue. Some messages have an invalid format and always fail processing. They are retried repeatedly and consume consumer capacity. What is the best way to prevent these "poison pill" messages from blocking normal processing?

A.Enable long polling and increase the maximum message retention to 30 days.
B.Configure a dead-letter queue (DLQ) with a redrive policy and a maxReceiveCount.
C.Switch the queue to FIFO and disable retries in the consumer code.
D.Delete the main queue and recreate it after every failure.
AnswerB

A DLQ with a redrive policy isolates poison-pill messages. After a message fails processing and is received more than maxReceiveCount times, SQS stops returning it to the main queue and moves it to the DLQ. Normal messages continue to be processed without repeatedly consuming consumer capacity.

Why this answer

A dead-letter queue (DLQ) with a redrive policy and a maxReceiveCount allows messages that repeatedly fail processing to be moved to a separate queue after a specified number of receive attempts. This prevents poison pill messages from being retried indefinitely, freeing consumer capacity for valid messages. Amazon SQS automatically redirects messages to the DLQ once the maxReceiveCount threshold is exceeded, ensuring normal processing is not blocked.

Exam trap

The trap here is that candidates may think increasing retention or polling settings will solve the problem, but they fail to recognize that only a DLQ with a redrive policy isolates repeatedly failing messages from consuming consumer capacity.

How to eliminate wrong answers

Option A is wrong because enabling long polling and increasing maximum message retention does not address the root cause of invalid messages; it only reduces empty responses and keeps messages longer, but poison pills will still be retried. Option C is wrong because switching to a FIFO queue does not prevent poison pills; FIFO ensures exactly-once processing but still retries failed messages, and disabling retries in consumer code would cause message loss without moving them to a DLQ. Option D is wrong because deleting and recreating the main queue after every failure is disruptive, loses all messages, and does not provide a systematic way to isolate or inspect poison pills.

183
Multi-Selectmedium

A healthcare company stores patient records in an Amazon DynamoDB table. The table must be recoverable to any point within the last 35 days, and the data must remain available if an entire AWS Region becomes unavailable. Which two actions should a solutions architect take to meet these requirements? (Choose two.)

Select 2 answers
A.Enable point-in-time recovery on the DynamoDB table.
B.Enable DynamoDB Streams and archive the stream to an Amazon S3 bucket.
C.Configure on-demand backup and restore with a daily backup schedule.
D.Enable DynamoDB Accelerator (DAX) for the table.
E.Create a DynamoDB global table with replica tables in additional AWS Regions.
AnswersA, E

Point-in-time recovery continuously backs up the DynamoDB table and allows restoration to any second within the last 35 days, up to the configured retention period. This directly satisfies the requirement to recover patient records to any point within 35 days. It protects against accidental writes or deletes but does not provide cross-Region availability, so it must be combined with another action.

Why this answer

Point-in-time recovery provides continuous backups that allow restoration to any second within the last 35 days, meeting the recovery requirement. DynamoDB global tables replicate the table across multiple AWS Regions and support automatic failover, meeting the Region-level availability requirement. Together, these two features address both dimensions of the scenario.

Exam trap

The trap here is treating DynamoDB Streams or on-demand backups as substitutes for point-in-time recovery, when only PITR provides second-level restore within a 35-day window.

184
MCQhard

A company is building a serverless application that processes messages from an Amazon SQS queue using AWS Lambda. The application must not lose messages and must handle occasional downstream failures gracefully. The Lambda function sometimes fails due to a transient error in a downstream service. The company wants to ensure that failed messages are retried and eventually processed, but also wants to avoid infinite retries that could block the queue. What should the company do?

A.Increase the Lambda function's timeout and memory allocation to reduce the chance of transient errors.
B.Configure a dead-letter queue (DLQ) on the source SQS queue and set the maximumReceiveCount to a reasonable value.
C.Set the SQS queue's visibility timeout to a very high value so that messages are not retried until the downstream service recovers.
D.Configure the Lambda function to write failed messages to an Amazon S3 bucket and delete them from the queue.
AnswerB

Setting a DLQ on the source queue and a maximumReceiveCount allows messages that repeatedly fail to be moved to a separate queue after a set number of attempts. This prevents infinite retries and preserves failed messages for later analysis or reprocessing. Lambda automatically returns messages to the queue if the function fails, so the receive count increments with each attempt.

Why this answer

Using a dead-letter queue on the source SQS queue with a maximumReceiveCount ensures that messages are retried a limited number of times and then moved to a DLQ for separate handling. This prevents poison messages from blocking the queue and allows for later analysis or reprocessing, meeting the requirement to avoid infinite retries while not losing messages.

Exam trap

The trap here is assuming that increasing Lambda timeout or visibility timeout solves the problem of persistent downstream failures, when in fact a DLQ is needed to prevent infinite retries.

185
MCQmedium

A healthcare analytics team runs a containerized reporting service on Amazon ECS with the Fargate launch type in a single Availability Zone. The service must remain available if one Availability Zone fails, and it must scale automatically based on CPU utilization. The tasks are stateless and write output to Amazon S3. Which configuration should a solutions architect implement?

A.Create an ECS service with a task count of two, place the tasks in subnets in two Availability Zones, register them with a target group behind an Application Load Balancer, and attach a target-tracking scaling policy based on ECSServiceAverageCPUUtilization.
B.Configure the ECS service with two tasks in the same subnet of one Availability Zone and add a Network Load Balancer with cross-zone load balancing enabled.
C.Use the EC2 launch type with an Auto Scaling group spanning two Availability Zones and a scheduled scaling policy that adds instances at peak hours.
D.Deploy the ECS service with a task count of one and enable deployment circuit breaker with rollback so that a failed task is automatically replaced.
AnswerA

Running at least two tasks in subnets across two Availability Zones removes the single-zone dependency, and the Application Load Balancer health checks route around a task in a failed zone. Target-tracking on ECSServiceAverageCPUUtilization is the native Application Auto Scaling mechanism for ECS services, so the service adds and removes tasks automatically as CPU load changes.

Why this answer

Resilience against an Availability Zone failure requires task capacity in more than one zone, and an Application Load Balancer with health checks directs traffic only to healthy tasks. Because the tasks are stateless and store output in S3, no shared state must be replicated. Application Auto Scaling with a target-tracking policy on ECSServiceAverageCPUUtilization then adjusts task count to match demand without manual intervention.

Exam trap

The trap here is confusing the deployment circuit breaker, which only handles failed rollouts, with genuine multi-Availability-Zone redundancy and demand-based scaling.

186
MCQeasy

Based on the exhibit, the web tier becomes unavailable if us-west-2a has an outage. What is the best change to improve resilience with the least redesign?

A.Increase the Auto Scaling group desired capacity from 2 to 3 in the same subnet.
B.Attach the Application Load Balancer and Auto Scaling group to subnets in a second Availability Zone.
C.Replace the Application Load Balancer with a Network Load Balancer.
D.Increase the health check grace period so instances stay registered longer.
AnswerB

Spanning the load balancer and Auto Scaling group across at least two Availability Zones removes the single-AZ dependency shown in the exhibit. If us-west-2a fails, the remaining AZ can continue serving traffic and Auto Scaling can replace unhealthy instances there. This is the smallest architectural change that directly improves availability.

Why this answer

The web tier is currently deployed in a single Availability Zone (us-west-2a), so an outage of that AZ makes the entire tier unavailable. By attaching the Application Load Balancer and Auto Scaling group to subnets in a second Availability Zone, the application can continue serving traffic from the healthy AZ, achieving high availability with minimal architectural changes. This is the standard AWS best practice for multi-AZ resilience.

Exam trap

The trap here is that candidates may think increasing instance count or changing load balancer type improves resilience, but the core issue is the single-AZ deployment, which only multi-AZ subnets can fix.

Why the other options are wrong

A

Increasing desired capacity in the same subnet does not add fault tolerance across Availability Zones; the web tier remains vulnerable to a single AZ outage.

C

Replacing the Application Load Balancer with a Network Load Balancer does not address the single-Availability Zone failure; the web tier would still be in us-west-2a only, so an outage of that AZ would still cause unavailability.

D

Increasing the health check grace period only delays instance deregistration during an outage, but does not address the root cause: the web tier is in a single Availability Zone. Instances in us-west-2a will still become unhealthy and eventually be terminated, causing unavailability.

187
MCQhard

A warehouse integration service must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The design must avoid adding custom operational scripts.

A.Amazon EFS with mount targets in multiple Availability Zones
B.S3 mounted as a POSIX file system without a file gateway
C.Instance store volumes
D.An EBS volume attached to all instances
AnswerA

EFS mount targets in multiple Availability Zones give each Linux instance a local endpoint, so file storage stays reachable if one AZ fails. EBS cannot attach across AZs, and the multi-AZ design needs no custom scripts.

Why this answer

Amazon EFS provides a fully managed, POSIX-compliant NFSv4.1 shared file system that can be mounted concurrently across multiple Linux EC2 instances. By deploying mount targets in multiple Availability Zones, the file system remains accessible even if one AZ fails, satisfying the high-availability requirement without any custom scripts.

Exam trap

The trap here is that candidates may confuse EBS Multi-Attach (which has strict limitations and requires cluster-aware file systems) with a true shared file system, or assume that S3 with a FUSE mount is a viable POSIX alternative without considering the operational overhead and lack of native consistency.

How to eliminate wrong answers

Option B is wrong because mounting S3 as a POSIX file system (e.g., via s3fs-fuse) requires custom operational scripts and does not provide native POSIX semantics or strong consistency, making it unsuitable for shared file storage across AZs. Option C is wrong because instance store volumes are ephemeral, tied to a single EC2 instance, and cannot be shared across instances or survive AZ failures. Option D is wrong because a single EBS volume cannot be attached to multiple EC2 instances; it can only be attached to one instance at a time, and while Multi-Attach EBS exists, it is limited to specific instance types and does not provide a shared file system without additional cluster-aware software.

188
MCQeasy

A media company stores finalized video masters in an Amazon S3 bucket in the us-east-1 Region. Compliance requires that the objects be recoverable if they are accidentally deleted or overwritten for at least 90 days, and that no user, including administrators, be able to permanently erase them during that period. Which S3 feature should the solutions architect enable?

A.S3 Versioning with a lifecycle rule that transitions noncurrent versions to S3 Glacier Deep Archive after 30 days.
B.S3 Object Lock in governance mode with a retention period of 90 days on the bucket.
C.S3 Object Lock in compliance mode with a retention period of 90 days on the bucket.
D.S3 Cross-Region Replication to a bucket in another Region with versioning enabled on both buckets.
AnswerC

S3 Object Lock in compliance mode prevents any user, including the root account, from deleting or overwriting a protected object version until the retention period expires. A 90-day retention period matches the compliance window exactly, and the protection cannot be shortened or bypassed, which is what the requirement demands.

Why this answer

The requirement is legal-hold-style immutability that even administrators cannot override. S3 Object Lock in compliance mode enforces exactly that: protected object versions cannot be deleted or overwritten by any identity until the retention period lapses, and the retention cannot be shortened. Governance mode and replication both leave a path for privileged users to destroy the data.

Exam trap

The trap here is treating governance mode and compliance mode as interchangeable, when only compliance mode removes the ability of privileged users to bypass retention.

189
MCQhard

Based on the exhibit, the database is manually promoted during an Availability Zone failure and the application outage lasts longer than the target. What change best improves resilience with the least operational intervention?

A.Keep the read replica and automate promotion with a runbook after CloudWatch alarms fire.
B.Convert the database to an RDS Multi-AZ deployment so a synchronous standby can fail over automatically.
C.Use a cross-Region read replica so promotion happens faster during an AZ failure.
D.Increase the application retry count and keep the current database design.
AnswerB

Multi-AZ is designed for automatic failover within the same Region and maintains a synchronous standby for high availability. The exhibit shows that the current read replica requires manual promotion and produces an outage longer than the target. Switching to Multi-AZ removes the manual step and aligns the database layer with the desired recovery time.

Why this answer

B is correct because RDS Multi-AZ automatically synchronously replicates data to a standby in a different Availability Zone and triggers an automatic failover with zero manual intervention when an AZ failure occurs. This directly addresses the requirement to improve resilience while minimizing operational effort, as the failover is handled by AWS without any runbook execution or manual promotion.

Exam trap

The trap here is that candidates often confuse read replicas (designed for read scaling and manual promotion) with Multi-AZ deployments (designed for automatic failover), and incorrectly assume that automating a runbook for read replica promotion is equivalent to the native automatic failover of Multi-AZ.

Why the other options are wrong

A

Manual promotion via runbook after CloudWatch alarms still requires human intervention, which does not meet the goal of 'least operational intervention' and will not reduce outage duration as effectively as automatic failover.

C

Cross-Region read replicas are asynchronous and do not support automatic failover; promoting them requires manual intervention and takes longer than Multi-AZ failover, so they do not improve resilience with the least operational intervention during an AZ failure.

D

Increasing the application retry count does not address the root cause of the outage (AZ failure) and does not improve resilience; it only masks the symptom and may lead to degraded user experience or timeouts.

190
MCQmedium

An application writes to an Amazon Aurora DB cluster. After a planned Aurora failover, the application experiences several minutes of connection errors. The logs show the application continues connecting to the specific DB instance endpoint that was the primary before the failover. What change most directly improves resilience during Aurora failovers?

A.Update the application to use the Aurora cluster writer endpoint for write traffic so it always resolves to the current writer instance.
B.Increase Aurora storage autoscaling so failovers are unnecessary.
C.Point both reads and writes to the Aurora reader endpoint to keep the DNS name the same.
D.Disable Aurora failover capability so the cluster never switches writer instances.
AnswerA

During failover, Aurora changes which underlying DB instance is the writer. The cluster writer endpoint (for the cluster) always resolves to the current writer. Using the writer endpoint prevents the application from being pinned to an old instance endpoint that may stop accepting writes after failover.

Why this answer

The Aurora cluster writer endpoint always resolves to the current primary DB instance, even after a failover. By using this endpoint instead of a specific instance endpoint, the application automatically reconnects to the new writer without manual intervention or connection errors.

Exam trap

The trap here is that candidates may think using any Aurora endpoint (like the reader endpoint) is sufficient, but they must understand that only the cluster writer endpoint guarantees write availability after a failover, while the reader endpoint is strictly for read traffic.

Why the other options are wrong

B

Increasing Aurora storage autoscaling does not prevent failovers; it only adjusts storage capacity. Failovers occur due to instance-level issues, not storage limits, so this change does not address the application's connection errors after failover.

C

The reader endpoint is intended for read-only traffic and does not handle write operations; pointing writes to it would cause failures. Moreover, the reader endpoint resolves to multiple reader instances, not the current writer, so it does not ensure connectivity to the writer after failover.

D

Disabling failover capability prevents the cluster from switching to a healthy instance during a failure, which would cause prolonged downtime instead of resolving the connection errors. The question asks for improving resilience, and disabling failover directly undermines that goal.

191
MCQmedium

A company uses Amazon RDS with automated backups enabled (retention period: 7 days). At 10:30 UTC, a bad release corrupts specific rows in a production table. The team detects the issue at 11:10 UTC. They need to revert the database state to what it was from 10:00–10:30 UTC, recover quickly, and minimize risk to the currently running workload. What is the best option?

A.Reboot the DB instance and rely on the corrupted data being overwritten by storage-level changes.
B.Perform a point-in-time restore to a new DB instance using a timestamp before the corruption (for example, a time within 10:00–10:30 UTC).
C.Restore only the most recent automated backup snapshot, even if it is after the corruption timestamp.
D.Create a read replica of the current DB instance and overwrite the corrupted table using SELECT queries from the replica.
AnswerB

With automated backups enabled, RDS supports point-in-time recovery (PITR) within the retention window. Restoring to a timestamp before the corruption creates a consistent copy from that moment. The team can validate the restored DB and then cut over application traffic, reducing risk to the currently running workload.

Why this answer

Amazon RDS automated backups enable point-in-time recovery (PITR) to any second within the retention window. By restoring to a timestamp between 10:00 and 10:30 UTC, you recover the database to a state before the corruption occurred, without affecting the current production instance. This minimizes risk to the running workload because the restore creates a new DB instance, leaving the original untouched until you are ready to switch.

Exam trap

The trap here is that candidates may confuse automated backup snapshots (which are full backups taken once per day) with point-in-time recovery (which uses transaction logs to restore to any point within the retention window), leading them to choose Option C instead of B.

How to eliminate wrong answers

Option A is wrong because rebooting a DB instance does not revert data; it only restarts the database engine and does not undo committed transactions or storage-level changes. Option C is wrong because restoring the most recent automated backup snapshot includes the corrupted data, so it does not achieve the goal of reverting to a pre-corruption state. Option D is wrong because a read replica mirrors the current (corrupted) data; using SELECT queries from it cannot overwrite the corrupted table with clean data, and it does not provide a mechanism to roll back changes.

192
MCQmedium

A service processes customer payments from a message queue. Because the queue provides at-least-once delivery, the same payment message can be delivered more than once if the consumer times out before committing its state. Currently, the service sometimes charges the customer twice. Which design change most directly prevents duplicate charges while still allowing safe retries?

A.Delete the message from the queue immediately after receive to prevent redelivery.
B.Make the payment processing idempotent by recording an idempotency key for each payment and ensuring repeated deliveries do not apply the charge twice.
C.Increase the queue visibility timeout to a very large value so messages rarely reappear.
D.Switch to a single-threaded consumer with one worker so messages are processed in order.
AnswerB

Because message queues provide at-least-once delivery, a payment message may be delivered multiple times if a consumer times out or fails after processing. By writing the payment's idempotency key (for example, a payment reference) to a durable store with a unique constraint before applying the charge, the consumer can detect and ignore repeated deliveries. This ensures the charge happens exactly once even when a redelivery occurs, making the system safe against at-least-once semantics.

Why this answer

Making payment processing idempotent using an idempotency key ensures that even if the same message is delivered multiple times due to at-least-once delivery semantics, the charge is applied only once. The consumer records a unique key (e.g., payment ID) in a durable store (like DynamoDB or Redis) and checks it before processing; if the key already exists, the charge is skipped. This directly prevents duplicate charges while still allowing safe retries, as the consumer can safely reprocess messages without side effects.

Exam trap

The trap here is that candidates often confuse at-least-once delivery with exactly-once delivery and assume that increasing visibility timeouts or using single-threaded consumers will prevent duplicates, when in fact only idempotency guarantees safe retries without duplicate charges.

Why the other options are wrong

A

Deleting the message immediately after receive prevents redelivery but also eliminates the ability to retry if processing fails, which violates the requirement of allowing safe retries.

C

Increasing the visibility timeout to a very large value does not guarantee that a consumer won't crash or timeout, and it can delay processing of other messages, leading to potential bottlenecks and still allowing duplicate charges if the consumer fails after processing but before deleting the message.

D

Single-threading does not prevent duplicate charges because the same message can still be redelivered after a timeout, even with one worker; the core issue is at-least-once delivery, not concurrency.

193
MCQmedium

A media company stores original uploads in an S3 bucket. They must recover from accidental overwrites/deletes and also recover quickly from a full Region outage. The required RPO is about 1 hour. Which configuration best meets these requirements?

A.Enable an S3 lifecycle policy to transition objects to Glacier after 7 days without enabling versioning.
B.Enable S3 cross-Region replication (CRR) but leave the bucket without versioning enabled.
C.Enable S3 versioning and configure cross-Region replication to a bucket in another Region.
D.Rely on frequent EBS snapshots of a temporary cache used during uploads.
AnswerC

Enabling S3 versioning preserves every version of an object, allowing retrieval of prior versions if an object is accidentally overwritten or deleted. Cross-Region Replication (CRR) asynchronously copies objects to a bucket in another Region, protecting against a Regional disaster. Because CRR requires versioning on both source and destination buckets, this combination satisfies both logical protection and geographic redundancy.

Why this answer

Enabling S3 versioning protects against accidental overwrites and deletes by preserving all object versions, while cross-Region replication (CRR) asynchronously replicates objects to a bucket in another Region, providing recovery from a full Region outage. With versioning enabled, CRR replicates both current and previous object versions, meeting the ~1-hour RPO (typically within minutes for new objects) and ensuring data durability across Regions.

Exam trap

AWS often tests the misconception that CRR can work without versioning, but the S3 API explicitly requires versioning on the source bucket for replication to function, and candidates may overlook that versioning is also the mechanism that protects against accidental overwrites and deletes.

Why the other options are wrong

A

Without versioning, the lifecycle policy cannot protect against accidental overwrites or deletes, and Glacier transition does not provide quick recovery from a full Region outage (RPO ~1 hour).

B

Without versioning, S3 cross-Region replication cannot protect against accidental overwrites or deletes because replication only copies the current version; deleted or overwritten objects are not recoverable.

D

EBS snapshots are for EC2 instance volumes, not for S3 data. They cannot protect against accidental overwrites/deletes in S3, nor do they provide cross-Region recovery for S3 objects.

194
MCQhard

A warehouse integration service must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable? The design must avoid adding custom operational scripts.

A.Use CloudFront signed URLs
B.Use Amazon SQS standard queue and design consumers to be idempotent
C.Use UDP messages sent directly to workers
D.Use an in-memory queue on one EC2 instance
AnswerB

An Amazon SQS standard queue provides high-throughput, distributed, and durable messaging with at-least-once delivery semantics. Because duplicates can occur due to the distributed architecture, consumers must be idempotent to correctly handle repeated messages without side effects. The queue also decouples event producers from workers, automatically buffering events and allowing independent scaling of processing capacity, which satisfies the requirement to process every event reliably.

Why this answer

Amazon SQS standard queues provide at-least-once delivery, ensuring every event is processed at least once, with the possibility of duplicates. Designing consumers to be idempotent handles duplicates without requiring custom scripts, aligning with the requirement to avoid operational overhead. This approach is serverless, scalable, and fits the warehouse integration use case.

Exam trap

The trap here is that candidates may choose UDP (Option C) thinking it is lightweight and fast, but they overlook its lack of delivery guarantees, which fails the 'process every event at least once' requirement.

How to eliminate wrong answers

Option A is wrong because CloudFront signed URLs are for controlling access to content, not for event processing or messaging; they do not provide at-least-once delivery guarantees. Option C is wrong because UDP is a connectionless, unreliable protocol that does not guarantee message delivery, making it unsuitable for processing every event at least once. Option D is wrong because an in-memory queue on a single EC2 instance introduces a single point of failure and requires custom scripts for management, violating the 'avoid adding custom operational scripts' constraint.

195
MCQmedium

A content publishing system uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The team wants the control to be enforceable during normal operations.

A.Lambda reserved concurrency set to zero
B.A larger deployment package
C.CloudFront error pages
D.A Lambda dead-letter queue or failure destination
AnswerD

A Lambda dead-letter queue (an SQS queue or SNS topic) or an asynchronous failure destination is the AWS-recommended way to retain events that exhaust Lambda's two built-in retries. After the retry attempts fail, Lambda writes the event payload to the configured SQS/SNS or invokes a destination like another Lambda, EventBridge, or SNS, enabling a separate process to analyze or replay the failed event. To use it, the function must have a resource-based policy granting Lambda permissions to send to the DLQ/destination, and the configuration must be set on the function's event-invoke settings.

Why this answer

A Lambda dead-letter queue (DLQ) or failure destination allows you to capture events that have exhausted all retry attempts from an asynchronous invocation. This ensures failed events are retained in Amazon SQS or SNS for later investigation, providing enforceable control during normal operations without impacting the function's ability to process successful events.

Exam trap

The trap here is that candidates may confuse Lambda's synchronous invocation error handling (where DLQs are not supported) with asynchronous invocation, or mistakenly think that increasing function resources (like deployment package size) can improve reliability against external API failures.

How to eliminate wrong answers

Option A is wrong because setting Lambda reserved concurrency to zero would completely disable the function, preventing any invocations and thus failing to process events at all, which does not address the need to retain failed events after retries. Option B is wrong because a larger deployment package does not affect error handling or retention of failed events; it only increases the function's size, potentially impacting cold start times and deployment limits. Option C is wrong because CloudFront error pages are used for customizing HTTP error responses for web distributions, not for capturing or retaining Lambda invocation failures from asynchronous API calls.

196
Multi-Selectmedium

A SaaS application is deployed in us-east-1 and us-west-2 behind separate ALBs. The business wants DNS to send new clients to the primary Region when it is healthy and automatically fail over to the secondary Region when the primary endpoint is unhealthy. Which two Route 53 settings are required? Select two.

Select 2 answers
A.Use a failover routing policy with a primary and secondary record.
B.Create a health check and associate it with the primary endpoint.
C.Use weighted routing with a 50/50 traffic split between both Regions.
D.Use latency-based routing so clients always choose the fastest Region.
E.Use a geolocation policy without health checks.
AnswersA, B

Failover routing is Route 53's active-passive policy: you create a primary record pointing to the us-east-1 endpoint and a secondary record pointing to us-west-2. When a health check attached to the primary fails, Route 53 automatically returns the secondary record in DNS responses. This is the exact mechanism that provides automatic region-level failover for your SaaS application, and it also lets you designate a definitive secondary failover target.

Why this answer

A failover routing policy is correct because it allows you to designate one record as primary and another as secondary. Route 53 will route traffic to the primary record as long as it is healthy, and automatically fail over to the secondary record when the primary is unhealthy. This directly meets the requirement to send new clients to the primary region when healthy and fail over to the secondary region.

Exam trap

The trap here is that candidates often confuse failover routing with weighted or latency-based routing, thinking any multi-region setup with health checks will automatically fail over, but only failover routing policy provides the explicit primary/secondary failover behavior required.

Why the other options are wrong

C

Weighted routing with a 50/50 split distributes traffic evenly regardless of health, failing to automatically fail over to the secondary region when the primary endpoint is unhealthy.

D

Latency-based routing directs traffic to the region with the lowest latency for each user, not to a primary region with failover to a secondary region when the primary is unhealthy.

E

Geolocation routing directs traffic based on client location, not health. Without health checks, it cannot automatically fail over when the primary endpoint is unhealthy, which is the core requirement.

197
MCQmedium

A logistics company runs an order-tracking API on a fleet of EC2 instances in a single Availability Zone behind a Network Load Balancer. The architecture team must make the API resilient to the loss of that Availability Zone without changing the API endpoint that clients already use. The instances are stateless and store session data in a shared Amazon ElastiCache cluster. Which change should the solutions architect make to meet these requirements?

A.Enable termination protection on the EC2 instances and configure an Elastic IP address per instance so that clients can reconnect to the same address after a zone failure.
B.Replace the Network Load Balancer with an Application Load Balancer deployed in one Availability Zone and attach an Auto Scaling group with a desired capacity of two instances.
C.Create a second Network Load Balancer in a different Availability Zone and use Amazon Route 53 weighted routing to send half the client traffic to each load balancer.
D.Create an Auto Scaling group that spans at least two Availability Zones, register the instances with a target group attached to the existing Network Load Balancer, and enable cross-zone load balancing.
AnswerD

A Network Load Balancer is a Regional resource with nodes in each enabled Availability Zone, so extending the Auto Scaling group across multiple Availability Zones lets healthy instances in a surviving zone serve traffic without any DNS or endpoint change. Cross-zone load balancing distributes traffic evenly so the remaining zone absorbs the full load, and because sessions live in ElastiCache, no instance holds state that would be lost.

Why this answer

The resilient pattern is to spread stateless compute across multiple Availability Zones and let the existing Regional Network Load Balancer route to whichever targets are healthy. Because the load balancer already has nodes in every enabled zone, adding an Auto Scaling group that spans zones gives the surviving zone capacity to absorb traffic, and cross-zone load balancing keeps distribution even. Session state already lives in ElastiCache, so no failover logic is needed in the application.

Exam trap

The trap here is assuming the load balancer itself needs replacing or duplicating, when the real single point of failure is the compute capacity confined to one Availability Zone.

198
MCQhard

A payments API uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure? The architecture review board prefers a managed AWS-native control.

A.A FIFO queue without a redrive policy
B.A dead-letter queue with an appropriate maxReceiveCount
C.A larger message retention period only
D.Short polling instead of long polling
AnswerB

A dead-letter queue with maxReceiveCount moves messages aside after a set number of failed receives, so poison messages stop blocking the main queue and useful retries continue. This is the managed, AWS-native control the review board requires, needing no custom code.

Why this answer

A dead-letter queue (DLQ) with an appropriate maxReceiveCount is the correct AWS-native solution for handling poison messages. When a message is repeatedly received from an SQS queue but fails processing, it is considered a poison message. By configuring a DLQ and setting a maxReceiveCount (e.g., 3 or 5), the message is automatically moved to the DLQ after exceeding that threshold, preventing it from blocking further retries and allowing the main queue to process valid messages.

Exam trap

The trap here is that candidates may confuse poison message handling with ordering or polling optimizations, and incorrectly choose FIFO queues or short polling, not realizing that only a DLQ with a redrive policy isolates repeatedly failing messages.

How to eliminate wrong answers

Option A is wrong because a FIFO queue without a redrive policy does not automatically handle poison messages; it only ensures strict ordering and exactly-once processing, but failed messages remain in the queue and continue to block retries. Option C is wrong because increasing the message retention period only keeps messages longer in the queue, but does nothing to isolate or remove poison messages that are repeatedly failing. Option D is wrong because short polling (returning immediately even if no messages are available) versus long polling (waiting for messages) affects latency and cost, but does not address the poison message problem; poison messages are a content/processing issue, not a polling mechanism issue.

199
MCQmedium

A patient portal receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers? The design must avoid adding custom operational scripts.

A.AWS WAF
B.Amazon CloudFront
C.Amazon SQS queue
D.Amazon Route 53 weighted routing
AnswerC

Amazon SQS decouples the web tier from fulfilment workers, buffering burst order volumes so spikes are absorbed rather than dropped. Its native at-least-once delivery and visibility-timeout retry mechanism reprocess failed messages without custom operational scripts, satisfying the no-scripts constraint.

Why this answer

Amazon SQS is the correct choice because it acts as a durable, fully managed message buffer that decouples the web tier from the fulfilment workers. When bursts of orders arrive, SQS queues the messages and allows workers to poll at their own pace, absorbing spikes without data loss. The built-in retry logic (visibility timeout and dead-letter queue) ensures failed processing attempts are automatically retried, and no custom operational scripts are needed.

Exam trap

The trap here is that candidates often confuse decoupling with caching or DNS-level distribution, picking CloudFront or Route 53 because they think 'absorbing spikes' means scaling web servers, but the question specifically requires buffering and retry without custom scripts, which only a queue service like SQS provides.

How to eliminate wrong answers

Option A is wrong because AWS WAF is a web application firewall that filters HTTP/S traffic based on rules (e.g., SQL injection, XSS); it does not buffer or retry messages between tiers. Option B is wrong because Amazon CloudFront is a content delivery network (CDN) that caches and accelerates static/dynamic content at edge locations; it cannot queue or retry asynchronous order processing. Option D is wrong because Amazon Route 53 weighted routing distributes DNS traffic across multiple endpoints based on weights; it provides load balancing at the DNS level but does not absorb spikes or provide retry mechanisms for message processing.

200
MCQmedium

A media company stores original video masters in an Amazon S3 bucket in the us-east-1 Region. Compliance requires that a readable copy of every object exist in the eu-west-1 Region within 15 minutes of upload, and that the objects in eu-west-1 be usable directly by an application there. No transformations are required. Which S3 feature should the solutions architect enable?

A.S3 Multi-Region Access Points with an active-passive routing configuration
B.S3 Same-Region Replication (SRR) between the us-east-1 bucket and a second bucket in us-east-1
C.S3 Lifecycle policies that transition objects to S3 Glacier Instant Retrieval in eu-west-1 after 15 minutes
D.S3 Cross-Region Replication (CRR) with S3 Versioning enabled on both the source and destination buckets
AnswerD

CRR asynchronously replicates new and updated objects to a destination bucket in another Region, and most objects replicate within minutes, satisfying the 15-minute requirement. Versioning is mandatory on both source and destination buckets for replication rules to work. Because the replicated objects are full readable copies, the application in eu-west-1 can read them directly without any restore step.

Why this answer

Cross-Region Replication is the purpose-built S3 capability for maintaining an automatically updated, readable object copy in a different Region, and its asynchronous replication latency generally falls well inside a 15-minute window. Versioning on both buckets is a prerequisite, so the architect must enable it as part of the design. The other choices either keep data in one Region, move data only between storage classes, or route requests without creating a second copy.

Exam trap

The trap here is confusing request-routing features such as Multi-Region Access Points with data-replication features, and forgetting that S3 Versioning must be enabled on both buckets before a replication rule can be created.

201
Multi-Selecthard

A regional web application for a content publishing system must fail over automatically to a secondary Region if the primary endpoint becomes unhealthy. Which two services or features are required? The design must avoid adding custom operational scripts.

Select 2 answers
A.AWS Organizations service control policies
B.Route 53 failover routing with health checks
C.S3 Transfer Acceleration
D.A deployed standby application stack in the secondary Region
AnswersB, D

Route 53 failover routing with health checks monitors the primary endpoint and automatically redirects DNS to the secondary when it turns unhealthy. This satisfies the automatic failover constraint without custom operational scripts, since health evaluation and DNS switching are managed by Route 53 itself.

Why this answer

Option B is correct because Route 53 failover routing with health checks automatically redirects DNS queries to the secondary endpoint when the primary health check is deemed unhealthy, providing the required automatic failover without custom scripts. Option D is correct because a deployed standby application stack in the secondary Region is necessary to actually serve traffic after failover; DNS redirection alone cannot run the application. Option A is incorrect because AWS Organizations service control policies govern permissions and guardrails, not traffic failover or health-based routing.

Option C is incorrect because S3 Transfer Acceleration only speeds up uploads to S3 buckets and does not provide regional failover or health checking.

Exam trap

The trap here is that candidates often assume Route 53 alone is sufficient, forgetting that the secondary Region must have a fully deployed and running application stack to receive traffic after failover.

202
MCQmedium

A logistics company runs an order-tracking service that exposes a REST API. The service must remain available during a single Availability Zone failure and must keep read latency low for a globally distributed user base. The data store must support automatic multi-AZ replication without the team managing database servers. Which solution meets these requirements?

A.Amazon DynamoDB with global tables and on-demand capacity mode.
B.Amazon RDS for MySQL with a read replica in a second Availability Zone.
C.Amazon ElastiCache for Redis with a cluster mode disabled replication group in one Availability Zone.
D.Amazon RDS for PostgreSQL with Multi-AZ DB instance deployment and a cross-Region read replica.
AnswerA

DynamoDB is a fully managed, multi-AZ service by default, and global tables replicate data across Regions for low-latency reads near users. On-demand capacity removes provisioning concerns. The team does not manage servers, and the design remains available during a single Availability Zone failure without manual failover steps.

Why this answer

DynamoDB is a fully managed, multi-AZ key-value store, so a single Availability Zone failure does not interrupt service and no database servers are managed. Global tables replicate data across Regions so users read from a nearby replica with low latency, and on-demand capacity removes the need to provision throughput, meeting both resilience and performance goals.

Exam trap

The trap here is choosing an RDS read replica for failover, when read replicas are not automatic failover targets and require manual promotion.

203
MCQmedium

Your web application is deployed in two AWS Regions (Region A and Region B). You want Route 53 to automatically fail over DNS traffic from Region A to Region B when Region A is unhealthy. The failover decision must be based on health checks that verify whether the application in Region A is reachable. Which Route 53 routing configuration best meets these requirements?

A.Latency-based routing with regional aliases to split traffic based on measured latency.
B.Geolocation routing using country-based routing policies.
C.Failover routing using a primary record with an associated health check for Region A and a secondary record for Region B.
D.Weighted routing with weights set to 100 for Region A and 0 for Region B.
AnswerC

Route 53 failover routing is designed for active/standby patterns. You configure the Region A record as primary with a health check. When that health check fails, Route 53 automatically returns the Region B (secondary) record, enabling health-check-driven regional failover.

Why this answer

Route 53 failover routing allows you to create a primary record with an associated health check for Region A and a secondary record for Region B. When the health check for Region A fails, Route 53 automatically returns the secondary record's IP address, directing traffic to Region B. This directly meets the requirement for automatic failover based on application reachability.

Exam trap

The trap here is that candidates often confuse failover routing with weighted routing, mistakenly thinking that setting weights to 100/0 will achieve failover, but weighted routing does not automatically adjust weights based on health checks.

Why the other options are wrong

A

Latency-based routing directs traffic based on lowest latency, not health status. It cannot automatically fail over to Region B when Region A is unhealthy because it lacks health check integration.

B

Geolocation routing directs traffic based on the geographic location of the user, not on the health of the endpoint. It cannot automatically failover from Region A to Region B when Region A becomes unhealthy.

D

Weighted routing with 100/0 weights does not provide automatic failover; it simply sends all traffic to Region A until you manually change weights. It lacks health checks to trigger failover when Region A becomes unhealthy.

204
MCQmedium

A media company stores original video masters in an Amazon S3 bucket in the us-east-1 Region. Compliance requires that a readable copy of every object exists in eu-west-1 within 15 minutes of upload, and that objects deleted in the source bucket do not automatically disappear from the destination. Which S3 feature should the solutions architect enable?

A.S3 Same-Region Replication into a bucket in us-east-1 with S3 Object Lock in compliance mode
B.AWS Backup with a cross-Region backup vault and a daily scheduled backup plan for the bucket
C.S3 Cross-Region Replication with S3 Replication Time Control and DeleteMarkerReplication disabled
D.S3 Transfer Acceleration on the source bucket with a lifecycle rule transitioning objects to S3 Glacier Deep Archive
AnswerC

S3 Cross-Region Replication copies objects to a bucket in another Region, and S3 Replication Time Control provides an SLA that 99.99% of objects replicate within 15 minutes. Disabling DeleteMarkerReplication prevents delete markers created in the source bucket from being replicated, so deletions in us-east-1 will not remove the copy in eu-west-1. This satisfies both the timing and the retention requirements.

Why this answer

S3 Cross-Region Replication is the AWS feature that asynchronously copies new and updated objects to a bucket in another Region. S3 Replication Time Control adds an SLA that 99.99% of objects replicate within 15 minutes, matching the compliance window. Disabling DeleteMarkerReplication keeps the destination copy intact when the source object is deleted, which is exactly the retention behaviour the compliance team requires.

Exam trap

The trap here is assuming that any replication option guarantees a 15-minute window, when only S3 Replication Time Control provides that SLA.

205
MCQmedium

A media processing company runs a stateless thumbnail-generation fleet on Amazon EC2 instances behind an Application Load Balancer. The instances store no local state, and the team wants the fleet to survive the loss of an entire Availability Zone without manual intervention. The fleet must also scale out automatically based on CPU. Which combination of AWS services should the solutions architect use to meet these requirements with the LEAST operational overhead?

A.Two independent Auto Scaling groups in separate Availability Zones, each attached to its own Application Load Balancer, with Route 53 failover routing between them.
B.A single Auto Scaling group pinned to one Availability Zone, with a Network Load Balancer in front and an Amazon Route 53 latency record.
C.An Amazon ECS cluster on AWS Fargate with tasks placed in one Availability Zone and an Application Load Balancer using sticky sessions.
D.An Auto Scaling group configured across multiple Availability Zones with a health check against the load balancer, plus an Application Load Balancer with cross-zone load balancing enabled.
AnswerD

Spreading the Auto Scaling group across multiple Availability Zones means capacity remains if one AZ fails, and the ELB health check replaces unhealthy instances automatically. Cross-zone load balancing distributes requests evenly to remaining healthy targets, so the stateless fleet keeps serving traffic without manual action and still scales on CPU.

Why this answer

Resilience against an Availability Zone failure for a stateless fleet is achieved by distributing capacity across multiple Availability Zones and letting the Auto Scaling group replace unhealthy instances using load balancer health checks. A single multi-AZ Auto Scaling group behind one Application Load Balancer with cross-zone load balancing delivers this with minimal operational effort, while still supporting CPU-based scaling.

Exam trap

The trap here is assuming a load balancer alone provides Availability Zone resilience, when the compute capacity must also be spread across multiple zones by the Auto Scaling group.

206
MCQmedium

A trading dashboard uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The design must avoid adding custom operational scripts.

A.A single-AZ Aurora cluster
B.Aurora Global Database
C.Manual snapshots copied monthly
D.An ElastiCache Redis replica
AnswerB

Aurora Global Database is a feature specifically designed for cross-region disaster recovery and low-latency global reads. It replicates data from a primary Region to up to five secondary Regions with typical latency of under a second, using dedicated storage-based replication rather than binlog-based replication. In a regional failure, you can promote one of the secondary regions to become the new primary in as little as one minute, which gives a low RTO, while snapshot-based approaches would take much longer.

Why this answer

Aurora Global Database is the correct choice because it provides a fully managed cross-Region disaster recovery solution with a typical RPO of 1 second or less, using storage-based replication that does not require custom scripts. This meets the low RPO requirement while avoiding operational overhead, as replication is handled automatically by the Aurora storage layer.

Exam trap

The trap here is that candidates may confuse cross-Region read replicas (which require manual promotion and scripting) with Aurora Global Database, which provides automated, low-latency replication without custom operational scripts.

How to eliminate wrong answers

Option A is wrong because a single-AZ Aurora cluster lacks any cross-Region replication or failover capability, offering no disaster recovery across Regions. Option C is wrong because manual snapshots copied monthly result in an RPO of up to one month, which is far too high for a trading dashboard requiring low RPO. Option D is wrong because an ElastiCache Redis replica is an in-memory cache, not a database with persistent cross-Region replication, and it does not provide the required disaster recovery for Aurora MySQL data.

207
MCQmedium

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer. The application must be highly available and able to withstand the failure of an entire AWS Region. The company wants to minimize operational overhead and ensure that failover is automatic. Which solution should a solutions architect recommend?

A.Deploy the application in two AWS Regions. Use Amazon Route 53 with a weighted routing policy and equal weights for both Regions.
B.Deploy the application in two Availability Zones within a single Region. Use an Application Load Balancer to distribute traffic across both AZs.
C.Deploy the application in two AWS Regions. Use Amazon Route 53 with a failover routing policy and health checks to route traffic to the secondary Region if the primary becomes unhealthy.
D.Deploy the application in a single Region using AWS Global Accelerator to route traffic to the nearest edge location.
AnswerC

A multi-Region active-passive setup with Route 53 failover routing and health checks provides automatic failover at the DNS level. The secondary Region must have the same application stack and data replication. This meets the requirements for Region-level resilience and minimal operational overhead because failover is automatic.

Why this answer

To withstand a Region failure with automatic failover, the application must be deployed in at least two Regions. Route 53 failover routing with health checks automatically redirects traffic to the secondary Region when the primary is unhealthy. This minimizes operational overhead because the failover is managed by Route 53 and does not require manual intervention.

Exam trap

The trap here is confusing weighted routing with failover routing. Weighted routing distributes traffic based on weights but does not automatically remove an unhealthy Region unless health checks are explicitly configured to do so, and even then, it may not fully redirect all traffic.

208
MCQhard

A financial services firm runs a stateful trading application on EC2 instances in an Auto Scaling group. Each instance maintains an in-memory cache that takes several minutes to rebuild after a restart, and the team wants the application to survive the loss of an Availability Zone with minimal disruption. The application cannot be made stateless in the near term. Which approach should a solutions architect recommend?

A.Move the in-memory cache to an Amazon ElastiCache for Redis cluster with Multi-AZ enabled and keep the Auto Scaling group in one Availability Zone.
B.Create a second Auto Scaling group in another Availability Zone and use Amazon Route 53 failover routing with health checks to switch traffic when the primary zone fails.
C.Enable EC2 Auto Recovery on all instances and configure the Auto Scaling group to use a single Availability Zone with a larger instance type.
D.Configure the Auto Scaling group to span three Availability Zones with a capacity that leaves headroom in each zone, and enable instance warm-up and health check grace periods so that replacement instances are fully initialized before receiving traffic.
AnswerD

Spreading capacity across three zones means the loss of one zone leaves two zones with running, already-warmed instances that can absorb the load immediately. Instance warm-up and health check grace periods prevent the Auto Scaling group from treating a still-initializing instance as unhealthy, which matters because the in-memory cache takes several minutes to rebuild and premature termination would cause a restart loop.

Why this answer

For a stateful application that cannot be re-architected immediately, the practical resilience pattern is to run warmed instances in more than two Availability Zones so that a single zone loss leaves sufficient capacity already serving traffic. Warm-up and grace period settings are essential because they stop the Auto Scaling group from killing instances that are still rebuilding their in-memory caches, which would otherwise produce repeated restarts and prolonged unavailability.

Exam trap

The trap here is assuming that Auto Recovery or DNS failover replaces the need for live, warmed capacity in multiple Availability Zones for a stateful workload.

209
MCQmedium

An application uses an Amazon Aurora DB cluster. The cluster performs an automatic failover from the writer instance to a standby instance. After failover completes, reads succeed, but all new writes fail with errors indicating the application is connecting to the old writer endpoint. Which change best fixes the resiliency issue after failover?

A.Update the application to use the Aurora cluster writer endpoint (or the cluster endpoint intended for writes) rather than an instance-specific endpoint.
B.Enable Multi-AZ on the individual writer instance settings so it can automatically create a new instance during failover.
C.Increase the failover timeout for Aurora to 60 minutes to ensure the app finishes reconnecting.
D.Switch the cluster to a single-AZ configuration to reduce connection retries after failover.
AnswerA

During Aurora failover, the writer role moves to a different underlying DB instance. The cluster writer endpoint is stable and always resolves to the current writer, even after failover. An instance-specific endpoint continues to point to the original (now non-writer) instance, so write operations fail if the application keeps using that stale endpoint.

Why this answer

The application is failing writes because it is connecting to the old writer instance's endpoint, which is no longer the writer after failover. The Aurora cluster writer endpoint is a DNS name that always points to the current primary (writer) instance, regardless of failovers. By using the cluster writer endpoint, the application automatically connects to the new writer after failover, eliminating the need to manually update connection strings.

Exam trap

The trap here is that candidates often confuse instance-specific endpoints with cluster endpoints, assuming that failover automatically updates all DNS records, but only the cluster endpoint is dynamically updated to reflect the new writer.

How to eliminate wrong answers

Option B is wrong because Multi-AZ is already inherent in Aurora clusters (by default, Aurora stores data across three Availability Zones) and enabling it on an individual instance does not change the failover behavior or fix the endpoint issue. Option C is wrong because increasing the failover timeout to 60 minutes does not address the root cause; the application will still connect to the old writer endpoint and fail writes indefinitely. Option D is wrong because switching to a single-AZ configuration would actually reduce resiliency and increase the risk of data loss, and it does not solve the problem of the application using the wrong endpoint.

210
MCQeasy

A company needs to store application logs in a durable and highly available manner. The logs are written continuously by multiple EC2 instances and are accessed infrequently for compliance audits. The company wants a solution that provides 99.999999999% (11 9's) durability and automatically replicates data across multiple Availability Zones. Which AWS service should the company use?

A.Amazon EC2 instance store volumes.
B.Amazon Elastic Block Store (Amazon EBS) volumes attached to each EC2 instance.
C.Amazon RDS for MySQL with a Multi-AZ deployment.
D.Amazon S3 with the STANDARD storage class.
AnswerD

Amazon S3 Standard provides 99.999999999% durability and automatically replicates data across multiple Availability Zones within a Region. It is highly available and suitable for infrequently accessed compliance logs. Multiple EC2 instances can write to the same bucket concurrently.

Why this answer

Amazon S3 Standard is designed for 99.999999999% durability and automatically replicates data across multiple Availability Zones within a Region. It supports concurrent writes from multiple EC2 instances and is a cost-effective solution for storing and retrieving logs. This meets the durability and availability requirements.

Exam trap

The trap here is assuming that EBS volumes or instance store provide similar durability to S3. EBS volumes are replicated within a single AZ, and instance store is ephemeral, so neither meets the cross-AZ durability requirement.

211
MCQmedium

Based on the exhibit, the application should continue serving requests if one Availability Zone fails. Which change best improves resilience with the least operational complexity?

A.Increase the desired capacity in AZ-a so more instances can absorb the failure of that same Availability Zone.
B.Add at least one subnet from a second Availability Zone to both the ALB and the Auto Scaling group.
C.Disable health checks so the ALB stops removing targets during brief infrastructure issues.
D.Move the application to a single larger instance type so the fleet has fewer moving parts.
AnswerB

A resilient design needs the load balancer and the Auto Scaling group to span multiple Availability Zones. If one AZ fails, the ALB can still route to healthy targets in the remaining AZs and the Auto Scaling group can replenish capacity there. This is the simplest and most common way to achieve AZ-level fault tolerance.

Why this answer

Adding subnets from a second Availability Zone to both the ALB and the Auto Scaling group distributes the application across multiple AZs. This ensures that if one AZ fails, the ALB can route traffic to healthy targets in the remaining AZ, and the Auto Scaling group can maintain capacity by launching instances in the surviving AZ. This approach directly addresses the requirement to continue serving requests during an AZ failure with minimal operational complexity.

Exam trap

The trap here is that candidates often think increasing capacity in a single AZ (Option A) provides resilience, but it actually concentrates risk in that AZ, while the correct answer requires distributing resources across multiple AZs to achieve true fault tolerance.

How to eliminate wrong answers

Option A is wrong because increasing the desired capacity in a single AZ does not provide resilience against the failure of that same AZ; all instances would be lost if the AZ fails. Option C is wrong because disabling health checks would prevent the ALB from detecting and removing unhealthy targets, causing traffic to be routed to failed instances and degrading application availability. Option D is wrong because moving to a single larger instance type creates a single point of failure; if that instance fails, the entire application becomes unavailable, and it does not address AZ-level failures.

212
MCQeasy

A startup runs a stateless web tier on Amazon EC2 instances in an Auto Scaling group that spans three Availability Zones. The team wants the application to keep serving requests even if one instance becomes unresponsive, without operator involvement. What should the solutions architect configure?

A.An Application Load Balancer with health checks that route traffic only to healthy targets.
B.A Network Load Balancer configured with a single target group and no health checks.
C.An Elastic IP address associated with each EC2 instance in the Auto Scaling group.
D.An Auto Scaling group with a scheduled scaling policy that adds instances every morning.
AnswerA

The load balancer performs health checks against each target and stops sending traffic to targets that fail, so an unresponsive instance is removed from rotation automatically. Combined with the Auto Scaling group, replacement capacity is launched, keeping the stateless web tier available without human action.

Why this answer

Health-checked load balancing is the mechanism that detects an unresponsive instance and removes it from the traffic path automatically. Because the tier is stateless and already spans three Availability Zones, the Application Load Balancer can shift requests to remaining healthy instances while the Auto Scaling group replaces the failed one.

Exam trap

The trap here is expecting an Elastic IP or scheduled scaling to provide automatic failover for an unhealthy instance, when only health-checked load balancing removes it from service.

213
MCQeasy

A startup runs a customer-facing web application on a single Amazon EC2 instance in one Availability Zone, with the database on the same instance. The founders want the application to survive the failure of that Availability Zone with minimal changes and no server management for the database tier. Which action should the solutions architect take first?

A.Create an Amazon Machine Image of the instance and copy it to a second Region for disaster recovery.
B.Enable termination protection on the EC2 instance and take regular Amazon EBS snapshots.
C.Place an Application Load Balancer in front of the single instance and enable sticky sessions.
D.Move the database to Amazon RDS with a Multi-AZ DB instance deployment and place the application instances in an Auto Scaling group spanning multiple Availability Zones.
AnswerD

Moving the database to RDS Multi-AZ provides automatic failover to a standby in another Availability Zone and removes server management for the database. Placing the application in a multi-AZ Auto Scaling group ensures compute capacity survives a zonal failure, directly addressing the single-AZ risk with a well-understood pattern.

Why this answer

The core problem is that both the application and database reside in a single Availability Zone. Relocating the database to Amazon RDS with Multi-AZ deployment gives automatic failover and removes database server management, while running the application in a multi-AZ Auto Scaling group provides compute redundancy. Together these changes let the application survive a zonal failure with minimal rework.

Exam trap

The trap here is believing that snapshots or termination protection provide high availability, when they only aid recovery and do not keep the workload running through an Availability Zone failure.

214
MCQmedium

Based on the exhibit, the payment worker sometimes processes the same SQS Standard message more than once after a timeout. What change best prevents duplicate charges while keeping the queue architecture?

A.Increase the SQS visibility timeout to 15 minutes and leave the worker unchanged.
B.Replace the Standard queue with a FIFO queue and rely only on message ordering.
C.Make the payment workflow idempotent by recording a unique order key before charging.
D.Add a second consumer so duplicate messages are processed faster.
AnswerC

SQS Standard queues are at-least-once delivery, so duplicate messages are always possible. The correct safeguard is idempotency: store a unique order or payment request key, check whether that key has already been processed, and only perform the charge the first time it is seen. Any later delivery is safely ignored.

Why this answer

Making the payment workflow idempotent ensures that even if the same SQS Standard message is processed more than once (due to a visibility timeout), the duplicate charge is prevented by checking a unique order key before processing. This is the most robust solution for handling at-least-once delivery semantics of Standard queues without changing the queue architecture.

Exam trap

The trap here is that candidates often think increasing the visibility timeout (Option A) or switching to a FIFO queue (Option B) will solve duplicate processing, but they overlook that the root cause is the worker's timeout behavior, which requires application-level idempotency to prevent duplicate charges.

Why the other options are wrong

A

Increasing the visibility timeout to 15 minutes does not prevent duplicate processing; it only reduces the likelihood of timeouts causing duplicates. The worker can still process the same message twice if the timeout expires after 15 minutes, and the change does not address the root cause of duplicate charges.

B

FIFO queues guarantee exactly-once processing and ordering, but the question asks to prevent duplicate charges while keeping the queue architecture. Replacing Standard with FIFO changes the queue type, which may not be desired, and FIFO alone does not prevent duplicate charges if the worker is not idempotent.

D

Adding a second consumer does not prevent duplicate processing; it may even increase the chance of duplicates if both consumers process the same message after a timeout.

215
MCQmedium

A production Amazon RDS database has automated backups enabled. At 10:00 UTC, an application deploy accidentally overwrote a subset of rows due to a faulty migration. The issue is detected at 10:45 UTC. The team confirms that the required retention window is still available. Which approach offers the most resilient and least disruptive way to recover the affected data close to the time of the event?

A.Perform a snapshot restore and attach the restored instance, then manually copy only the affected rows back into the current database.
B.Use point-in-time recovery to restore the database to a timestamp just before 10:00 UTC, then swap application connectivity to the recovered instance.
C.Rely on automated backups to roll forward automatically until the data becomes correct.
D.Disable automated backups going forward to prevent future corruption, then reindex the corrupted table.
AnswerB

Point-in-time recovery (PITR) for Amazon RDS uses automated backups and transaction logs to restore the database to any second within the backup retention period, allowing you to target a timestamp just before 10:00 UTC when the corruption occurred. This minimizes data loss to only the changes made in the seconds immediately preceding the incident, far more precise than a full snapshot. After restoring to a new RDS instance, you swap the application connection string (or use Route 53/RDS Proxy) to point to the recovered instance, enabling a clean recovery with minimal disruption and no manual row copying.

Why this answer

Point-in-time recovery (PITR) allows you to restore the RDS instance to any second within the backup retention window, such as just before the faulty migration at 10:00 UTC. This restores a complete, consistent database state, minimizing data loss and avoiding manual row-by-row recovery. Swapping application connectivity to the restored instance is the least disruptive approach, as it avoids complex manual data merging and reduces downtime.

Exam trap

The trap here is that candidates may choose snapshot restore (Option A) thinking it is faster or simpler, but they overlook that PITR provides a more precise, consistent recovery point without manual data extraction and reinsertion.

How to eliminate wrong answers

Option A is wrong because performing a snapshot restore and manually copying affected rows is error-prone, time-consuming, and risks data inconsistency, especially if the affected rows have dependencies. Option C is wrong because automated backups do not 'roll forward' to correct data corruption; they are used for restore operations, not automatic healing. Option D is wrong because disabling automated backups does not recover lost data and actually increases future risk; reindexing does not restore overwritten rows.

216
MCQmedium

A company runs an Amazon Aurora DB cluster with a Multi-AZ deployment. The application is configured with a hard-coded endpoint that points to the current writer *DB instance* (an instance-specific endpoint), rather than the Aurora cluster writer endpoint. During an unexpected AZ failure, Aurora promotes the standby to become the new writer. However, the application continues to fail to connect until an operator updates the hard-coded endpoint. What change most directly improves resiliency so the application automatically reconnects after failover?

A.Keep using the writer DB instance endpoint, but increase the client connection timeout.
B.Connect using the Aurora cluster writer endpoint so DNS resolves to the current writer after failover.
C.Disable Multi-AZ failover and rely on manual snapshot restore to bring the database back online.
D.Enable cross-Region read replicas and route application traffic to the replica during the outage.
AnswerB

Aurora cluster endpoints are designed to provide continuity across failovers. The Aurora cluster writer endpoint (writer endpoint for the cluster) updates so DNS resolves to the promoted writer. The application can reconnect without manual endpoint changes.

Why this answer

The Aurora cluster writer endpoint is a DNS name that always resolves to the current writer instance in the cluster, even after a failover. By using this endpoint instead of a hard-coded instance-specific endpoint, the application automatically reconnects to the new writer without manual intervention, directly improving resiliency.

Exam trap

The trap here is that candidates may confuse the instance-specific endpoint with the cluster writer endpoint, or think that increasing timeouts or using read replicas can solve a writer failover issue, when the core problem is the hard-coded reference to a specific instance that no longer exists.

How to eliminate wrong answers

Option A is wrong because increasing the client connection timeout does not change the fact that the hard-coded endpoint points to a failed instance; the connection will still fail after the timeout expires. Option C is wrong because disabling Multi-AZ failover and relying on manual snapshot restore would cause significant downtime and data loss, directly contradicting the goal of improving resiliency. Option D is wrong because cross-Region read replicas are read-only and cannot accept writes; routing application traffic to a read replica during an outage would not allow the application to write data, and it does not address the failover of the writer instance.

217
MCQhard

Based on the exhibit, duplicate payment charges occasionally occur when the worker times out after the charge is submitted but before the message is deleted. What change best prevents duplicate charges while keeping retry behavior?

A.Switch the queue to FIFO and rely on content-based deduplication to guarantee exactly-once processing.
B.Make the consumer idempotent by storing a processed payment key and rejecting repeat charges.
C.Reduce the visibility timeout so the message becomes available again sooner after a timeout.
D.Add a dead-letter queue and disable retries so the message is never processed twice.
AnswerB

The worker can still receive the same message more than once because SQS Standard is at-least-once delivery and the delete happened after the charge. Idempotency is the correct safety control because it prevents the payment from being applied twice even when the message is retried. A processed-payment record or conditional write lets retries remain possible without creating duplicate charges.

Why this answer

Making the consumer idempotent ensures that even if the same message is processed more than once (due to a timeout after the charge is submitted but before the message is deleted), the duplicate charge will be rejected. By storing a processed payment key (e.g., a unique transaction ID) and checking it before processing, the system can safely retry without causing duplicate payments. This approach preserves retry behavior while preventing duplicates, which is the core requirement.

Exam trap

The trap here is that candidates often assume FIFO queues with deduplication guarantee exactly-once processing, but they fail to recognize that deduplication only prevents duplicate message delivery, not duplicate processing when the consumer times out after processing but before acknowledging the message.

Why the other options are wrong

A

FIFO queues with content-based deduplication provide exactly-once delivery, but the issue here is a timeout after submission but before deletion, which can still cause duplicate processing if the worker retries. FIFO deduplication does not prevent duplicate charges if the same message is sent again after a timeout, as deduplication is based on message content within a 5-minute window, not on processing state.

C

Reducing the visibility timeout would cause the message to reappear sooner after a timeout, increasing the likelihood of duplicate processing rather than preventing it. It does not address the root cause of duplicate charges when the worker times out after submitting the charge.

D

Adding a dead-letter queue and disabling retries prevents duplicate charges by eliminating retries, but the question explicitly requires keeping retry behavior. This option removes retries, which violates the requirement.

218
Matchingmedium

A team wants a web application to keep serving traffic if one Availability Zone fails. Match each architecture element to the resilience behavior it provides.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Stop sending requests to unhealthy targets and keep only healthy instances in rotation.

Launch replacement instances in healthy AZs when capacity is lost.

Maintain a synchronous standby in another AZ and fail over automatically.

Allow instances to be replaced without losing user sessions that are stored elsewhere.

Why these pairings

These pairs match architecture elements with their resilience behaviors for surviving an Availability Zone failure, focusing on AWS services that provide high availability and fault tolerance.

219
MCQmedium

A ticket booking system uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The team wants the control to be enforceable during normal operations.

A.Aurora Global Database
B.A single-AZ Aurora cluster
C.An ElastiCache Redis replica
D.Manual snapshots copied monthly
AnswerA

Aurora Global Database replicates data at the storage layer to secondary Regions with typical latency under one second, using an asynchronous but dedicated replication channel. It supports both planned switchover and unplanned failover promotion, achieving RPO of seconds and RTO of minutes, which far exceeds the ticket system's need for fast failover. Secondary Regions can also serve local reads, improving both availability and recovery performance.

Why this answer

Aurora Global Database is designed for cross-Region disaster recovery with a typical RPO of 1 second or less, using storage-based replication that does not impact database performance. It provides fast failover to a secondary Region and allows the primary Region to enforce write control during normal operations, meeting the low RPO and enforceable control requirements.

Exam trap

The trap here is that candidates may confuse Aurora Global Database with cross-Region read replicas or manual snapshot copy strategies, underestimating the RPO and failover speed requirements for disaster recovery.

How to eliminate wrong answers

Option B is wrong because a single-AZ Aurora cluster lacks any cross-Region replication or failover capability, resulting in no DR protection and an RPO that depends on manual backups. Option C is wrong because ElastiCache Redis is an in-memory cache, not a persistent database, and its cross-Region replication (Global Datastore) does not provide the same transactional consistency or DR guarantees as Aurora Global Database for a ticket booking system. Option D is wrong because manual snapshots copied monthly would yield an RPO of up to 30 days, far exceeding the low RPO requirement, and they require manual intervention for recovery, which is not fast.

220
Multi-Selectmedium

A logistics company runs an order-tracking service on Amazon EC2 instances that write state to an Amazon DynamoDB table. A recent incident showed that a single Availability Zone failure caused the service to lose capacity, and the team also discovered that a developer accidentally deleted a production table. The architect must improve both Availability Zone resilience and protection against accidental table deletion. (Choose two.)

Select 2 answers
A.Enable DynamoDB point-in-time recovery on the table and attach a resource-based policy that denies the dynamodb:DeleteTable action to non-administrative principals.
B.Create a DynamoDB global secondary index on the partition key used by the tracking queries and project all attributes into the index.
C.Enable DynamoDB Streams on the table and write a consumer that copies every change into an Amazon S3 bucket for long-term retention.
D.Deploy the EC2 instances in an Auto Scaling group that spans multiple Availability Zones and attach the instances to an Application Load Balancer.
E.Convert the DynamoDB table to use provisioned capacity mode with auto scaling so that read and write capacity automatically adjusts during traffic spikes.
AnswersA, D

Point-in-time recovery allows restoration of the table to any second within the previous 35 days, which recovers data after an accidental deletion. Adding an IAM policy that denies dynamodb:DeleteTable to ordinary principals reduces the chance of the same mistake recurring, so together they address the accidental-deletion risk.

Why this answer

Zone resilience for the compute tier comes from running instances across multiple Availability Zones behind a load balancer, so a single zone loss does not remove all capacity. Accidental table deletion is mitigated by enabling point-in-time recovery, which allows restore to a recent point in time, and by restricting who holds the delete-table permission so the mistake is far less likely to recur.

Exam trap

The trap here is assuming DynamoDB needs multi-AZ configuration like a relational database, when DynamoDB already replicates data across zones and the real gaps are compute placement and deletion protection.

221
MCQhard

A warehouse integration service must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The team wants the control to be enforceable during normal operations.

A.Amazon EFS with mount targets in multiple Availability Zones
B.S3 mounted as a POSIX file system without a file gateway
C.Instance store volumes
D.An EBS volume attached to all instances
AnswerA

EFS is regional file storage and supports mount targets across AZs.

Why this answer

Amazon EFS provides a fully managed, POSIX-compliant NFS file system that can be mounted concurrently on multiple Linux EC2 instances across different Availability Zones. By creating mount targets in each AZ, the file system remains accessible even if one AZ fails, because the other mount targets continue to serve traffic. EFS also supports lifecycle policies and IAM enforcement to control access during normal operations, meeting the requirement for enforceable control.

Exam trap

The trap here is that candidates often confuse EBS multi-attach (which is limited to specific instance types and a single AZ) with the cross-AZ shared file system capability that only EFS provides, or they mistakenly think S3 with a FUSE mount is a reliable POSIX file system for production workloads.

How to eliminate wrong answers

Option B is wrong because mounting S3 as a POSIX file system (e.g., using s3fs-fuse) does not provide true POSIX semantics (e.g., no file locking, eventual consistency) and is not designed for shared file storage across AZs with high availability during an AZ failure. Option C is wrong because instance store volumes are ephemeral, tied to a single EC2 instance, and data is lost if the instance stops or fails; they cannot be shared across instances or survive an AZ failure. Option D is wrong because an EBS volume can only be attached to a single EC2 instance at a time (except for multi-attach EBS, which is limited to specific instance types and is not designed for shared file storage across AZs); attaching the same EBS volume to multiple instances is not supported.

222
MCQeasy

A team uses an S3 bucket to store important customer-generated exports. They need protection against accidental overwrites and also want copies of the data in another AWS Region for disaster recovery. Which S3 configuration best satisfies both requirements?

A.Enable S3 lifecycle policies to automatically move objects to Glacier after 30 days only.
B.Enable S3 versioning and configure Cross-Region Replication to a destination bucket in another Region.
C.Disable all versioning and rely on AWS Backup to restore objects from a scheduled backup window.
D.Enable S3 Block Public Access and SSE-S3 encryption, without using versioning or replication.
AnswerB

Enabling S3 versioning preserves every version of an object, so accidental overwrites or deletes can be undone by restoring a prior version or removing a delete marker. Cross-Region Replication then asynchronously copies new and updated objects to a bucket in another Region, providing a geographically separate copy for disaster recovery. Together these features directly address both object-level corruption and Region-level failures, making them the correct solution.

Why this answer

Enabling S3 versioning protects against accidental overwrites by preserving all object versions, allowing recovery of previous versions. Configuring Cross-Region Replication (CRR) automatically replicates objects to a destination bucket in another AWS Region, providing disaster recovery by maintaining a copy of the data in a separate geographic location.

Exam trap

The trap here is that candidates may think lifecycle policies or AWS Backup alone can handle both accidental overwrites and disaster recovery, but they fail to address the real-time protection and cross-region copy requirements that versioning and CRR specifically provide.

Why the other options are wrong

A

Lifecycle policies to Glacier only address storage cost optimization, not protection against accidental overwrites or cross-region disaster recovery.

C

AWS Backup does not prevent accidental overwrites; it only provides scheduled backups. Without versioning, overwritten objects are permanently lost until the next backup, and recovery point objectives may not align with real-time protection.

D

Block Public Access and SSE-S3 encryption protect against unauthorized access and encrypt data at rest, but they do not prevent accidental overwrites or provide cross-region disaster recovery copies.

223
MCQmedium

A media company runs a video-transcoding fleet on Amazon EC2 instances that read source files from an Amazon S3 bucket and write output to a second bucket. The fleet is spread across three Availability Zones in one Region, and instances are launched by an Auto Scaling group. The company needs the architecture to survive the loss of an entire Availability Zone without losing in-flight transcoding work or requiring manual intervention. Which combination of design elements should a solutions architect implement to meet these requirements?

A.Deploy the Auto Scaling group across three Availability Zones, make transcoding jobs idempotent and store progress in Amazon DynamoDB, and have instances poll an Amazon SQS queue for work so that unfinished jobs are retried by healthy instances.
B.Deploy the Auto Scaling group in a single Availability Zone with a spot fleet, use an Amazon EBS volume attached to each instance to persist transcoding progress, and enable EBS snapshots every five minutes to another zone.
C.Deploy the Auto Scaling group across three Availability Zones, place a Network Load Balancer in front of the instances, and configure the load balancer to retry failed transcoding requests against the same instance until the zone recovers.
D.Deploy the Auto Scaling group across three Availability Zones, store transcoding state in an Amazon S3 bucket configured with S3 Cross-Region Replication, and rely on the S3 Standard storage class for automatic recovery.
AnswerA

Spreading the Auto Scaling group across three Availability Zones means instances in surviving zones continue running when one zone fails. Because work is pulled from an SQS queue and progress is tracked in DynamoDB, a job interrupted in the failed zone becomes visible again after its visibility timeout and is retried by a healthy instance, so no manual intervention is required.

Why this answer

Resilience across an Availability Zone failure requires both compute capacity in the surviving zones and durable, externalized job state. Distributing the Auto Scaling group across three zones keeps instances running, while an SQS queue with visibility timeouts and DynamoDB progress tracking lets interrupted jobs be reclaimed and retried automatically. This removes any dependency on the failed zone's instances or storage.

Exam trap

The trap here is assuming that storing source and output objects in Amazon S3 is sufficient for workload resilience, when S3 durability says nothing about resuming in-flight compute work after a zone failure.

224
Multi-Selectmedium

A company is designing a disaster recovery plan for a critical application hosted on AWS. The application runs on EC2 instances with data stored in Amazon EBS volumes and Amazon S3. The recovery time objective (RTO) is 15 minutes, and the recovery point objective (RPO) is 1 hour. Which three strategies would help meet these objectives? (Choose three.)

Select 3 answers
.Use AWS Backup to create hourly snapshots of EBS volumes and copy them to a different AWS Region.
.Pre-provision EC2 instances in the disaster recovery region and keep them running 24/7.
.Replicate critical data to S3 in the disaster recovery region using S3 Cross-Region Replication (CRR).
.Store Amazon Machine Images (AMIs) in the source region and use AWS Lambda to copy them after a disaster.
.Configure Amazon Route 53 with a failover routing policy and health checks to redirect traffic to the DR region.
.Set up an AWS Direct Connect link between the primary and DR regions for faster data transfer.

Why this answer

AWS Backup can create hourly snapshots of EBS volumes and copy them to a different AWS Region, meeting the 1-hour RPO by ensuring backups are taken every hour. S3 Cross-Region Replication (CRR) asynchronously replicates objects to a bucket in another region, keeping data synchronized within minutes and supporting the RPO. Amazon Route 53 with a failover routing policy and health checks can automatically redirect traffic to the DR region within seconds to minutes, enabling the 15-minute RTO by quickly failing over to pre-prepared infrastructure.

Exam trap

The trap here is that candidates may confuse operational readiness (like pre-provisioning instances) with a specific strategy that directly contributes to meeting RTO/RPO, or they may think Direct Connect is a disaster recovery strategy when it is merely a connectivity option that does not automate failover or data replication.

225
MCQmedium

A service consumes messages from an SQS queue. Recently, a new message format started failing validation in the consumer. The consumer catches the exception but cannot successfully process those messages without code changes. The team wants failed messages to be isolated for later investigation instead of being retried indefinitely. What should they configure?

A.Set the queue’s retention period to 1 minute and rely on messages expiring naturally.
B.Configure a dead-letter queue (DLQ) with a redrive policy and set maxReceiveCount so messages move after repeated failed receives.
C.Increase the visibility timeout to 7 days so failed messages cannot be retried.
D.Publish the same message again to SNS on every failure so a different subscriber might succeed.
AnswerB

A DLQ isolates “poison messages” that repeatedly fail processing. With a redrive policy, SQS tracks receives; once a message exceeds maxReceiveCount without successful processing, SQS moves it to the DLQ. This prevents infinite retries on the bad format while preserving the failed messages for debugging and code fixes.

Why this answer

A dead-letter queue (DLQ) with a redrive policy is the correct solution because it allows messages that repeatedly fail processing to be moved to a separate queue after exceeding the maxReceiveCount. This isolates problematic messages for later investigation without blocking the main queue or causing infinite retries. The consumer catches the exception, so the message is not deleted and is returned to the queue for redelivery; the DLQ ensures that after a configurable number of attempts, the message is redirected instead of being retried indefinitely.

Exam trap

The trap here is that candidates may think increasing the visibility timeout or relying on message expiration is sufficient, but they fail to understand that those approaches either affect all messages or only temporarily hide the message, whereas a DLQ provides a permanent, targeted isolation mechanism for repeatedly failing messages.

How to eliminate wrong answers

Option A is wrong because setting the retention period to 1 minute would cause all messages (including valid ones) to expire quickly, leading to data loss and not isolating only the failed messages. Option C is wrong because increasing the visibility timeout to 7 days would simply hide the message from consumers for that period, but after the timeout expires the message would become visible again and be retried, failing to isolate it permanently. Option D is wrong because publishing the same message to SNS on every failure would create an infinite loop of republishing, and SNS subscribers would also fail if they use the same validation logic, not solving the isolation requirement.

← PreviousPage 3 of 4 · 257 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Design Resilient Architectures questions.