Courseiva

CCNA Design Resilient Questions

75 of 257 questions · Page 1/4 · Design Resilient topic · Answers revealed

1
Multi-Selectmedium

A production Amazon RDS database already has automated backups enabled. At 10:45 UTC, the team discovers that a faulty migration corrupted rows in a table at 10:30 UTC. The business wants the database restored to exactly the state it had at 10:30 UTC with minimal risk. Which two actions should the team take? Select two.

Select 2 answers
A.Restore the database to a new instance using point-in-time restore for 10:30 UTC.
B.Validate the restored database, then switch the application endpoint to the restored database.
C.Restore the most recent manual snapshot because it will include the 10:30 UTC state.
D.Overwrite the existing database instance in place so the application keeps the same storage volume.
E.Wait for automated backups to complete again, then replay the migration to restore the missing rows.
AnswersA, B

Correct. Point-in-time restore is the RDS recovery method for returning to a specific moment before the corruption occurred. Restoring to a new instance gives the team a clean database copy at the desired timestamp without risking the current production instance.

Why this answer

Amazon RDS Point-in-Time Restore (PITR) allows you to restore a DB instance to any second within the backup retention period, including 10:30 UTC. This uses automated backups and transaction logs to reconstruct the exact database state at that specific time, providing a precise recovery point with minimal data loss.

Exam trap

The trap here is that candidates may think manual snapshots can be used for point-in-time recovery, but they only capture a single moment and cannot roll forward to a specific time like automated backups can.

2
MCQeasy

Based on the exhibit, the web team wants the application to continue serving traffic if one Availability Zone fails. Which change best meets the requirement with the least operational overhead?

A.Increase desired capacity to 3 in the same Availability Zone so one extra instance is always available.
B.Add the unused subnet in us-east-1b to the Auto Scaling group so instances can launch in both AZs.
C.Replace the Application Load Balancer with a Network Load Balancer because it will automatically keep the app online.
D.Move the application to a larger EC2 instance type so a single server can handle the full workload.
AnswerB

Placing the Auto Scaling group in at least two Availability Zones allows AWS to distribute and replace instances across zones. Because the Application Load Balancer can route only to healthy targets, adding the second subnet is the lowest-complexity change that gives the application resilience to a full AZ outage.

Why this answer

It adds the unused subnet in us-east-1b to the Auto Scaling group, enabling EC2 instances to launch across two Availability Zones. This provides fault isolation: if one AZ fails, the ALB can route traffic to healthy instances in the other AZ. The change requires only a configuration update to the Auto Scaling group, minimizing operational overhead while meeting the high-availability requirement.

Exam trap

The trap here is that candidates often assume increasing instance count in a single AZ or using a different load balancer type alone provides high availability, but true resilience requires distributing instances across multiple Availability Zones.

How to eliminate wrong answers

Option A is wrong because increasing desired capacity to 3 in the same Availability Zone does not protect against an AZ failure; all instances remain in a single AZ, so if that AZ fails, all traffic is lost. Option C is wrong because replacing the Application Load Balancer with a Network Load Balancer does not inherently provide cross-AZ failover; the NLB still requires instances in multiple AZs to maintain availability, and the change introduces unnecessary operational overhead. Option D is wrong because moving to a larger EC2 instance type does not eliminate the single point of failure; if the AZ hosting that single instance fails, the application goes down regardless of instance size.

3
MCQmedium

A trading dashboard uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The architecture review board prefers a managed AWS-native control.

A.A single-AZ Aurora cluster
B.Aurora Global Database
C.Manual snapshots copied monthly
D.An ElastiCache Redis replica
AnswerB

Aurora Global Database replicates with low latency to secondary Regions and supports faster disaster recovery than snapshot-only approaches.

Why this answer

Aurora Global Database is the correct choice because it provides a managed, cross-Region disaster recovery solution with a Recovery Point Objective (RPO) of less than 1 second and a Recovery Time Objective (RTO) of typically less than 1 minute. It uses storage-based replication to keep a secondary cluster in another AWS Region up to date with minimal latency, meeting the low RPO requirement without manual intervention.

Exam trap

The trap here is that candidates may confuse cross-Region replication with multi-AZ deployments, or assume that manual snapshots or caching solutions can meet low RPO requirements, when only a managed global database service like Aurora Global Database provides the necessary sub-second RPO and automated failover.

How to eliminate wrong answers

Option A is wrong because a single-AZ Aurora cluster lacks any cross-Region replication or failover capability, offering no disaster recovery across Regions and resulting in an unacceptably high RPO if the primary Region fails. Option C is wrong because manual snapshots copied monthly provide an RPO of up to one month, which is far too high for the low RPO requirement, and the process is not automated or managed natively for rapid recovery. Option D is wrong because ElastiCache Redis is an in-memory cache, not a persistent database, and cannot serve as a primary data store for the trading dashboard's transactional data; it also lacks cross-Region replication for disaster recovery.

4
MCQeasy

An internal worker consumes messages from an Amazon SQS queue. Occasionally, a message fails validation in the worker (for example, missing required fields). Reprocessing the same bad message repeatedly wastes processing time and delays healthy messages. What is the best AWS approach to handle these poison messages without blocking the rest of the queue?

A.Configure an SQS dead-letter queue (DLQ) using a redrive policy with a maxReceiveCount.
B.Delete the SQS queue and recreate it daily to clear invalid messages.
C.Increase the consumer timeout/processing time so validation failures take longer to occur.
D.Use SNS fan-out without any DLQ and rely only on application retries.
AnswerA

With a redrive policy, SQS continues delivering the message to consumers until it has been received unsuccessfully maxReceiveCount times. After that threshold, SQS moves the poison message to a DLQ, isolating it from the main processing flow so healthy messages can continue being processed.

Why this answer

An SQS dead-letter queue (DLQ) with a redrive policy that sets a maxReceiveCount allows the worker to process a message up to a specified number of times. After that threshold is exceeded, the message is automatically moved to the DLQ, isolating the poison message and preventing it from blocking or delaying the processing of healthy messages in the main queue.

Exam trap

The trap here is that candidates may think increasing timeouts or relying on application retries alone can solve the problem, but they fail to recognize that only a DLQ with a redrive policy provides automatic, queue-level isolation of poison messages without blocking healthy message processing.

How to eliminate wrong answers

Option B is wrong because deleting and recreating the queue daily is disruptive, causes data loss of all messages (including valid ones), and does not provide a targeted mechanism to isolate only the poison messages. Option C is wrong because increasing the consumer timeout or processing time does not prevent validation failures; it only delays the retry cycle and does not remove the bad message from the queue, so it will still be reprocessed and waste resources. Option D is wrong because SNS fan-out without a DLQ and relying only on application retries means the poison message will be repeatedly delivered to all subscribers, causing infinite retries and blocking the processing of healthy messages; there is no automatic isolation mechanism.

5
MCQeasy

A production application uses an Amazon RDS Multi-AZ DB instance. During an unplanned failover, the database endpoint remains the same. What change should the application team make to handle the failover reliably?

A.Hard-code the new writer instance IP address after failover completes.
B.Keep using the same RDS endpoint and implement connection retry logic on failures.
C.Disable Multi-AZ and rely on manual intervention to switch endpoints.
D.Move reads to application-side caching only, and avoid reopening DB connections.
AnswerB

The RDS endpoint is DNS-based and remains constant across a Multi-AZ failover, so clients should continue using that same hostname. When failover occurs, existing active connections are dropped and in-flight transactions may fail; the application must treat those errors as transient and reconnect using retry logic with backoff or jitter. This allows the app to automatically resume writing to the new primary once DNS and the instance are ready. Do not assume an individual request succeeded; ensure retries are idempotent.

Why this answer

The RDS Multi-AZ DNS endpoint remains unchanged during a failover, automatically pointing to the new writer instance. Implementing connection retry logic with exponential backoff allows the application to handle the brief DNS propagation delay and connection interruption, ensuring reliable recovery without manual intervention.

Exam trap

The trap here is that candidates assume the endpoint changes or that Multi-AZ provides seamless failover without any application-side changes, but in reality the application must implement retry logic to handle the brief connection disruption during DNS propagation.

How to eliminate wrong answers

Option A is wrong because hard-coding the new writer instance IP address is impractical and error-prone; the IP address can change after failover, and this approach bypasses the automatic DNS update provided by Multi-AZ. Option C is wrong because disabling Multi-AZ removes high availability entirely, forcing manual endpoint switching which increases downtime and violates the goal of reliable failover handling. Option D is wrong because moving reads to application-side caching does not address the need to re-establish the database connection after failover; the application must still handle connection failures and retries for writes.

6
MCQeasy

Based on the exhibit, a web application must stay available if one Availability Zone fails. What is the best change to improve resilience?

A.Increase the desired capacity to 8 instances in the same subnet.
B.Add a subnet in another Availability Zone to the Auto Scaling group and keep the ALB spanning both AZs.
C.Replace the Application Load Balancer with a Network Load Balancer.
D.Move the instances to a larger instance type with more CPU and memory.
AnswerB

This places application instances across multiple Availability Zones, which protects the stateless tier from a single-AZ failure. The ALB already spans two AZs, so the missing piece is the Auto Scaling group using subnets in more than one AZ. That allows AWS to replace unhealthy instances and continue serving traffic from the surviving Zone.

Why this answer

Adding a subnet in another Availability Zone (AZ) to the Auto Scaling group and keeping the ALB spanning both AZs ensures that if one AZ fails, the ALB can route traffic to healthy instances in the other AZ. This is the standard pattern for building multi-AZ resilient architectures with Auto Scaling and ALB, as it eliminates the single point of failure at the AZ level.

Exam trap

The trap here is that candidates often think increasing instance count or size improves resilience, but without multi-AZ distribution, all instances remain vulnerable to a single AZ failure.

Why the other options are wrong

A

Increasing desired capacity within the same subnet does not protect against an Availability Zone failure; all instances remain in a single AZ, so if that AZ fails, the application becomes unavailable.

C

Replacing the ALB with a Network Load Balancer does not improve resilience across Availability Zones; it operates at a different layer and does not inherently provide cross-AZ fault tolerance for the application.

D

Increasing instance size (CPU/memory) does not address the requirement for availability zone failure resilience; it only improves performance, not fault tolerance across AZs.

7
MCQmedium

An order-processing service consumes messages from an Amazon SQS Standard queue using a custom worker. During traffic spikes, the worker occasionally times out after performing some work but before acknowledging the message, so SQS redelivers it and it may be processed again. You also observe that a small set of “poison” messages always fail validation. What change most directly improves resilience by (1) preventing poison messages from retrying indefinitely and (2) avoiding duplicate side effects caused by legitimate retries?

A.Increase the SQS visibility timeout and, when validation fails, call DeleteMessage in the consumer to remove the message immediately.
B.Move to SNS topics with subscriptions and rely on SNS to provide exactly-once delivery to eliminate duplicates automatically.
C.Configure a dead-letter queue (DLQ) with a redrive policy that moves messages after maxReceiveCount, and implement idempotent processing in the consumer using an idempotency key.
D.Change the queue to FIFO and enable content-based deduplication, leaving the consumer logic unchanged.
AnswerC

SQS Standard is at-least-once delivery, so timeouts can cause redelivery and duplicates. A DLQ with a redrive policy prevents poison messages from retrying forever by moving them after repeated failures. Idempotent processing (for example, storing a processed marker in a database with conditional logic keyed by an idempotency key) prevents duplicate side effects when retries occur for valid messages.

Why this answer

A dead-letter queue (DLQ) with a maxReceiveCount redrive policy directly addresses the poison message problem by moving messages that repeatedly fail validation out of the main queue after a set number of retries, preventing indefinite retries. Implementing idempotent processing using an idempotency key ensures that even if a legitimate message is redelivered due to a visibility timeout, the consumer can detect and skip duplicate side effects, thus solving both requirements most directly.

Exam trap

The trap here is that candidates often confuse FIFO queues as a universal solution for both deduplication and poison message handling, but FIFO only provides exactly-once processing within a deduplication window and does not automatically handle poison messages without a DLQ, nor does it address idempotency for retries outside that window.

Why the other options are wrong

A

Increasing visibility timeout does not prevent poison messages from retrying indefinitely; they would still be redelivered until deleted manually. Also, deleting on validation failure only removes poison messages but does not address duplicate side effects from legitimate retries, as the worker may still process the same message multiple times before the timeout expires.

B

SNS does not provide exactly-once delivery; it delivers messages at least once, so duplicates can still occur. Additionally, SNS does not handle poison messages or retries, so it fails to address both requirements.

D

FIFO queues guarantee exactly-once processing but do not prevent duplicate side effects from legitimate retries (e.g., after timeout) because the same message can be redelivered with a different deduplication ID. Also, poison messages would still retry indefinitely unless a DLQ is configured, which is not mentioned.

8
MCQmedium

A healthcare provider hosts a patient-records API on Amazon EC2 instances in a single Availability Zone behind an Application Load Balancer. An audit finds the architecture cannot tolerate the loss of that Availability Zone. Budget is limited, and the API reads from an Amazon Aurora MySQL cluster that currently has one writer instance and no replicas. Which change most effectively addresses the audit finding?

A.Add an Aurora Replica in a second Availability Zone and register EC2 instances in that same second Availability Zone with the load balancer's target group.
B.Enable Aurora backtrack on the cluster and increase the EC2 Auto Scaling group maximum size within the existing Availability Zone.
C.Convert the cluster to Aurora Global Database and add a secondary AWS Region for the EC2 fleet.
D.Replace Aurora with a single-instance Amazon RDS for MySQL deployment and place all EC2 instances in an Auto Scaling group with a desired capacity of one.
AnswerA

Placing an Aurora Replica in a second Availability Zone gives the database a failover target, and Aurora promotes a replica automatically if the writer fails. Adding EC2 instances in that same zone removes the single-zone compute risk and lets the Application Load Balancer route to healthy targets in either zone. This directly resolves the audit finding at modest cost.

Why this answer

The audit finding is about surviving the loss of one Availability Zone, so the fix must add capacity in a second zone for both tiers. An Aurora Replica in another zone gives the database an automatic failover target, and EC2 instances in that same zone let the load balancer continue serving requests. Cross-Region designs exceed the requirement and the budget.

Exam trap

The trap here is confusing multi-Region disaster recovery with multi-Availability Zone high availability, which leads to an over-engineered and over-budget answer.

9
MCQeasy

A company runs a stateless API on Amazon EC2 instances in a single Availability Zone behind an Application Load Balancer. The ALB currently has a listener on port 80 only. The company wants the API to remain available if the single Availability Zone fails. What should the solutions architect do to meet this requirement with the LEAST operational overhead?

A.Enable cross-zone load balancing on the ALB and increase the desired capacity of the existing instances in the single Availability Zone.
B.Create a second ALB in another Availability Zone and use Amazon Route 53 failover routing with health checks to switch traffic when the primary zone fails.
C.Create an Amazon Machine Image (AMI) of the instances and configure an Auto Scaling group across at least two Availability Zones, registering the instances with the existing ALB target group.
D.Move the API to an Amazon S3 static website endpoint and serve the traffic through Amazon CloudFront with origin failover.
AnswerC

An Auto Scaling group spanning multiple Availability Zones automatically launches replacement instances in a healthy zone if one zone fails, and registers them with the existing target group. This provides zone-level resilience with minimal ongoing management because scaling and health replacement are handled by the service.

Why this answer

Running the stateless API in an Auto Scaling group that spans at least two Availability Zones lets the group replace instances in a surviving zone when one zone becomes unavailable. The existing ALB and target group can be reused, so the change adds zone redundancy with the least operational overhead.

Exam trap

The trap here is assuming that adding more instances inside the same Availability Zone or enabling cross-zone load balancing protects against a zone-level failure.

10
MCQmedium

A web application runs on an EC2 Auto Scaling group (ASG) behind an Application Load Balancer (ALB). The ASG spans three Availability Zones. After a deployment, new instances frequently fail the ALB target group health checks with HTTP 5xx responses and are quickly terminated by the ASG. What change most improves resiliency during deployments with minimal downtime by preventing premature removal of instances that are still starting?

A.Reduce the ASG health check grace period to 0 seconds so issues are detected faster.
B.Use a longer ASG health check grace period and deploy new instances using controlled replacement (for example, rolling instance refresh) so existing healthy instances continue serving while new ones warm up.
C.Restrict the ASG to a single Availability Zone so health check evaluation is simpler.
D.Disable ALB health checks so the ASG does not terminate instances on HTTP 5xx responses.
AnswerB

A longer ASG health check grace period prevents instances from being evaluated too early during normal startup time. Controlled replacement or rolling instance refresh ensures capacity is maintained while new instances warm up, so the ALB continues routing requests only to healthy targets.

Why this answer

Increasing the ASG health check grace period gives new instances more time to complete their startup and pass the ALB health checks before the ASG marks them unhealthy. A rolling instance refresh replaces instances in a controlled manner, ensuring that existing healthy instances continue serving traffic while new instances warm up, minimizing downtime and preventing premature termination.

Exam trap

The trap here is that candidates think reducing the grace period or disabling health checks will speed up recovery, when in fact it causes premature termination or serves traffic to unhealthy instances, increasing downtime.

How to eliminate wrong answers

Option A is wrong because reducing the grace period to 0 seconds would cause the ASG to terminate instances even faster when they return HTTP 5xx during startup, worsening the problem. Option C is wrong because restricting to a single Availability Zone reduces fault tolerance and does not address the root cause of premature termination during startup. Option D is wrong because disabling ALB health checks would prevent the ASG from detecting actual instance failures, leading to serving traffic to unhealthy instances and increasing downtime.

11
MCQmedium

Based on the exhibit, the application sees several minutes of connection errors during an Aurora failover. What is the best change to reduce failover impact?

A.Change the application to use the Aurora cluster writer endpoint and retry transient connections.
B.Add an Aurora read replica and keep using the same JDBC URL.
C.Increase the EC2 instance size of the application servers.
D.Switch to a single-AZ RDS PostgreSQL instance for simpler connectivity.
AnswerA

The current configuration targets a specific instance endpoint, which becomes stale after failover. The Aurora cluster writer endpoint always resolves to the current writer, so the application can reconnect without manual endpoint changes. Adding retries with backoff helps the application survive the short DNS and connection transition during failover.

Why this answer

The Aurora cluster writer endpoint always points to the current primary instance, even after a failover. By using this endpoint and implementing retry logic for transient connection errors, the application can automatically reconnect to the new writer without manual intervention, reducing the impact of the failover from several minutes to seconds.

Exam trap

The trap here is that candidates often think adding read replicas or scaling application servers will fix failover connectivity, but the real issue is that the application must use the correct endpoint and handle transient disconnections gracefully.

Why the other options are wrong

B

Adding a read replica does not reduce failover impact because the application still uses the same JDBC URL, which points to the writer endpoint. During failover, the writer endpoint may be unavailable, and read replicas cannot handle write operations, so connection errors persist.

C

Increasing EC2 instance size addresses application-side compute capacity, not database failover connectivity issues. The connection errors stem from DNS propagation delays and endpoint changes during Aurora failover, which are unaffected by application server size.

D

Switching to a single-AZ RDS PostgreSQL instance removes high availability and does not address failover impact; it actually increases downtime risk during a failure.

12
MCQhard

A patient portal must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The architecture review board prefers a managed AWS-native control.

A.Instance store volumes
B.Amazon EFS with mount targets in multiple Availability Zones
C.An EBS volume attached to all instances
D.S3 mounted as a POSIX file system without a file gateway
AnswerB

Amazon EFS is a fully managed, regional NFS file system that provides shared, elastic file storage for Linux EC2 instances. By creating mount targets in each Availability Zone where instances reside, every instance can mount the same file system and read/write concurrently with POSIX semantics, automatic failover, and scalable throughput. This directly satisfies the requirement for shared file storage across multiple Linux EC2 instances in a patient portal.

Why this answer

Amazon EFS provides a fully managed, shared POSIX-compliant file system that can be mounted concurrently across multiple Linux EC2 instances. By creating mount targets in multiple Availability Zones, the file system remains accessible even if one AZ fails, meeting the high-availability requirement. This aligns with the architecture review board's preference for a managed AWS-native control.

Exam trap

The trap here is that candidates may confuse EBS multi-attach (which is limited to a single AZ and specific volume types) with the cross-AZ shared file system capability of EFS, or incorrectly assume that mounting S3 as a POSIX file system is a reliable, managed solution for shared storage.

How to eliminate wrong answers

Option A is wrong because instance store volumes are ephemeral, tied to a single EC2 instance, and data is lost on instance stop or termination, so they cannot provide shared, durable storage across AZs. Option C is wrong because a single EBS volume can only be attached to one EC2 instance at a time (except for multi-attach io1/io2 volumes, which are limited to a few Nitro-based instances and still not designed for cross-AZ shared file storage). Option D is wrong because mounting an S3 bucket as a POSIX file system (e.g., via s3fs-fuse) does not provide native POSIX locking or consistency semantics, and it introduces performance and reliability issues; it is not a managed AWS-native file system service.

13
MCQhard

A healthcare analytics platform stores derived datasets in an Amazon S3 bucket. Regulatory rules require that every object remain recoverable for 90 days after creation even if an application bug issues a delete, and that no object version be permanently destroyed during that window. The team wants the strongest protection with the least custom code. Which S3 feature should the solutions architect enable?

A.S3 Object Lock in compliance mode with a 90-day retention period.
B.S3 Lifecycle rules that transition objects to S3 Glacier Deep Archive after 1 day and expire them after 90 days.
C.S3 Object Lock in governance mode with a 90-day retention period.
D.S3 Versioning combined with a bucket policy that denies s3:DeleteObject.
AnswerA

Compliance mode prevents any principal, including the AWS account root user, from deleting or overwriting a protected object version until the retention period expires. Setting a 90-day retention satisfies the recoverability window with no custom code, giving the strongest immutability guarantee available for S3 objects.

Why this answer

S3 Object Lock in compliance mode enforces write-once-read-many protection for a defined retention period and cannot be bypassed by any user, including the root user. A 90-day retention period meets the recoverability requirement without custom code, making it the strongest and least-effort control for preventing permanent destruction of object versions.

Exam trap

The trap here is treating governance mode as equivalent to compliance mode, when governance mode can be bypassed by users holding the s3:BypassGovernanceRetention permission.

14
MCQmedium

A company runs a critical two-tier web application on AWS. The web tier consists of Amazon EC2 instances behind an Application Load Balancer (ALB) in a single Availability Zone. The database tier is an Amazon RDS for MySQL DB instance in the same Availability Zone. A recent power outage in that Availability Zone caused a full application outage. The company wants to redesign the architecture to survive an Availability Zone failure with minimal operational overhead. Which solution meets these requirements?

A.Use an Application Load Balancer with cross-zone load balancing enabled and enable Multi-AZ on the RDS instance.
B.Enable RDS automated backups with a 30-day retention period and create a read replica in a second Availability Zone.
C.Place the EC2 instances in an Auto Scaling group spanning two Availability Zones, and take nightly snapshots of the RDS instance to restore in the second AZ.
D.Deploy the web tier across two Availability Zones with an ALB, and convert the database to an RDS Multi-AZ DB instance.
AnswerD

Spreading EC2 instances across two AZs behind an ALB removes the single point of failure for the web tier, and an RDS Multi-AZ DB instance maintains a synchronous standby in a different AZ. On failure, RDS automatically fails over to the standby, and the ALB routes traffic to healthy instances. This design survives an AZ outage with minimal operational overhead.

Why this answer

The correct design distributes the stateless web tier across multiple Availability Zones and uses an RDS Multi-AZ DB instance for automatic database failover. This eliminates single points of failure at both tiers and requires little ongoing operational effort. Other options either leave the web tier in one AZ or rely on manual recovery for the database, failing the resilience and minimal-overhead goals.

Exam trap

The trap here is assuming that enabling Multi-AZ on RDS alone provides full application resilience while ignoring the single-AZ web tier.

15
MCQmedium

An orders service publishes payment instructions to an Amazon SQS Standard queue. The downstream processor sometimes times out after it has already applied the payment, but before it can delete the message from the queue. As a result, the same payment instruction can be processed more than once. The team wants the strongest way to prevent duplicate side effects while keeping the system decoupled. What should they implement?

A.Keep the queue as SQS Standard but increase the visibility timeout so duplicates are less likely to reappear during timeouts.
B.Change the queue to an SQS FIFO queue and use a stable deduplication ID derived from the payment instruction ID.
C.Make the downstream processor idempotent by recording processed payment instruction IDs in a durable datastore and ignoring repeats.
D.Use an ALB health check to restart the downstream processor when timeouts occur.
AnswerC

SQS Standard is at-least-once delivery, so the same message can be delivered more than once if the consumer times out before deleting it. Idempotent processing is the strongest protection against duplicate side effects because it prevents repeat application of the payment even when the message is redelivered.

Why this answer

Making the downstream processor idempotent ensures that duplicate payment instructions are safely ignored, even if the same message is delivered more than once. This approach provides the strongest guarantee against duplicate side effects without requiring changes to the queue type or increasing visibility timeouts, and it keeps the system fully decoupled.

Exam trap

The trap here is that candidates often assume that switching to a FIFO queue or increasing visibility timeout fully solves duplicate processing, but they overlook that the downstream processor's timeout after applying the payment is the root cause, which idempotency directly addresses.

How to eliminate wrong answers

Option A is wrong because increasing the visibility timeout only reduces the likelihood of duplicates but does not eliminate them; a timeout can still occur after processing, leading to the same duplicate issue. Option B is wrong because switching to an SQS FIFO queue with a deduplication ID prevents duplicate messages from being delivered, but it does not prevent the downstream processor from timing out after applying the payment and before deleting the message, so the same message could be redelivered and processed again. Option D is wrong because an ALB health check only restarts the downstream processor when timeouts occur, but it does not prevent duplicate processing of the same payment instruction.

16
MCQmedium

Your order-processing system uses EventBridge rules to send events to a Lambda function that updates order status. Over the last week, some events fail with a transient database timeout, and the Lambda retries intermittently but then the events are lost (no alerts after failures). You want at-least-once processing, bounded retries, and a way to inspect unprocessable events for later reprocessing. Which architecture change best meets these requirements?

A.Send EventBridge events to an SQS queue, configure a redrive policy to move messages to a dead-letter queue (DLQ) after a defined receive count, and make the Lambda processing idempotent.
B.Invoke Lambda directly from EventBridge in asynchronous mode, and increase the Lambda timeout to reduce failures.
C.Use SNS topics with Lambda subscriptions, but remove all retry and DLQ configuration to minimize duplicate events.
D.Store failed events only in CloudWatch logs, and have operators manually copy log entries back into the database for reprocessing.
AnswerA

Routing events through SQS decouples delivery from Lambda, so the redrive policy bounds retries and diverts unprocessable messages to a DLQ for later inspection and reprocessing, while idempotent handlers preserve at-least-once semantics across duplicate deliveries.

Why this answer

It introduces an SQS queue between EventBridge and Lambda, which provides a durable buffer for events. The redrive policy moves events to a dead-letter queue (DLQ) after a defined number of failed processing attempts, ensuring bounded retries and preserving unprocessable events for later inspection and reprocessing. Making the Lambda idempotent guarantees at-least-once processing even if duplicate events occur.

Exam trap

The trap here is that candidates may think increasing Lambda timeout or relying on asynchronous invocation retries alone is sufficient, but they overlook the need for a DLQ to capture and inspect events that persistently fail, which is a key requirement for operational visibility and reprocessing.

Why the other options are wrong

B

Asynchronous Lambda invocation from EventBridge has limited retry (0-2 attempts) and no DLQ support, so events lost after transient failures cannot be inspected or reprocessed, failing the requirement for bounded retries and inspectability.

C

Removing retry and DLQ configuration prevents at-least-once processing and makes it impossible to inspect unprocessable events, directly contradicting the requirements.

D

Storing failed events only in CloudWatch logs and manually reprocessing them does not provide automated retries, bounded retries, or a systematic way to inspect and reprocess unprocessable events, violating the requirements for at-least-once processing and automated reprocessing.

17
MCQeasy

A inventory service exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most? The architecture review board prefers a managed AWS-native control.

A.CloudFront caching with appropriate TTLs
B.AWS Backup Vault Lock
C.IAM Access Analyzer
D.S3 Select
AnswerA

CloudFront can serve cached content from edge locations when the origin is temporarily unavailable.

Why this answer

CloudFront caching with appropriate TTLs allows cached responses to be served to users even when the S3 origin is temporarily unavailable. By setting a minimum TTL (e.g., 0 seconds for fresh content, but a higher default or maximum TTL for stale content), CloudFront can continue delivering previously cached pages from edge locations during an S3 outage, ensuring high availability and resilience. This is a managed AWS-native feature that aligns with the architecture review board's preference.

Exam trap

The trap here is that candidates may confuse data protection features (like Backup Vault Lock) or data retrieval tools (like S3 Select) with caching and origin resilience, overlooking that CloudFront's TTL-based caching is the direct AWS-managed solution for serving content during origin outages.

How to eliminate wrong answers

Option B (AWS Backup Vault Lock) is wrong because it is a data protection feature for backup vaults that prevents deletion of backups, not a mechanism to serve cached content during an origin outage. Option C (IAM Access Analyzer) is wrong because it analyzes resource-based policies to identify unintended public access, not to cache or serve static content. Option D (S3 Select) is wrong because it is a query-in-place feature that retrieves subsets of data from objects using SQL expressions, and it does not provide caching or resilience against origin outages.

18
MCQmedium

A ticket booking system stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured?

A.S3 lifecycle transition to Glacier Flexible Retrieval
B.An EBS snapshot schedule
C.S3 Cross-Region Replication with versioning enabled
D.A CloudFront distribution
AnswerC

S3 Cross-Region Replication (CRR) asynchronously replicates every new object to a destination bucket in a different AWS Region, and it requires versioning to be enabled on both the source and destination buckets. This maintains a separate, durable copy of the uploaded documents that can be promoted to production during a regional outage, satisfying disaster recovery and compliance requirements. Replication is automatic, can be filtered by prefix or tags, and preserves object metadata and versions.

Why this answer

S3 Cross-Region Replication (CRR) with versioning enabled automatically copies objects from a source bucket in one AWS Region to a destination bucket in another Region, meeting the disaster recovery requirement for a geographically separate copy. Versioning must be enabled on both buckets to support replication of all object versions, ensuring consistency and recoverability. This is the native S3 feature designed for cross-region data redundancy without custom scripting or third-party tools.

Exam trap

The trap here is that candidates confuse S3 Cross-Region Replication with S3 lifecycle policies or other storage services like EBS snapshots, failing to recognize that CRR is the only option that directly creates a second copy of S3 objects in a different AWS Region for disaster recovery.

How to eliminate wrong answers

Option A is wrong because S3 lifecycle transition to Glacier Flexible Retrieval only moves objects to a lower-cost storage class within the same bucket and region; it does not create a copy in another AWS Region. Option B is wrong because EBS snapshots are for Amazon Elastic Block Store volumes attached to EC2 instances, not for S3 objects, and they cannot replicate data across regions automatically without additional configuration like copying snapshots manually. Option D is wrong because CloudFront is a content delivery network (CDN) that caches content at edge locations for low-latency delivery; it does not provide persistent cross-region storage replication for disaster recovery.

19
Multi-Selectmedium

A retail API runs on Amazon EC2 instances behind an Application Load Balancer and stores orders in an Amazon RDS for PostgreSQL database. A test that stopped one Availability Zone caused the API to return errors because all application servers were in the same AZ and the database was single-AZ. Which two changes should the architect make to continue serving traffic during a single-AZ failure? Select two.

Select 2 answers
A.Increase the EC2 instance size and keep all application servers in the same subnet.
B.Configure the Auto Scaling group to launch instances across private subnets in at least two Availability Zones.
C.Replace the Application Load Balancer with a Network Load Balancer in a single Availability Zone.
D.Convert the RDS for PostgreSQL database to a Multi-AZ deployment.
E.Add an Amazon RDS read replica and point the application to the replica endpoint.
AnswersB, D

Spreading instances across private subnets in at least two Availability Zones lets the Auto Scaling group replace capacity in a surviving AZ, satisfying the requirement to keep serving traffic when one AZ fails. The load balancer then routes only to healthy targets.

Why this answer

Option B is correct because an Auto Scaling group that spans private subnets in at least two Availability Zones ensures application instances remain available if one AZ fails, and the ALB can route to healthy targets in the surviving AZ. Option D is correct because converting RDS for PostgreSQL to a Multi-AZ deployment creates a synchronous standby in a different AZ and automatically fails over the database endpoint, eliminating the single-AZ database as a point of failure. Option A is wrong because increasing instance size does not address the single-AZ placement of the application servers.

Option C is wrong because a single-AZ Network Load Balancer still fails when that AZ fails and does not improve database resilience. Option E is wrong because a read replica is asynchronous, is not an automatic failover target for writes, and pointing the application at the replica endpoint does not provide a highly available primary database.

Exam trap

The trap here is that candidates often think a read replica can serve as a high-availability solution for writes, but read replicas are asynchronous and do not support automatic failover for the primary database.

Why the other options are wrong

A

Increasing EC2 instance size and keeping all servers in one subnet does not provide fault tolerance across Availability Zones; a single AZ failure would still take down all application servers.

C

A Network Load Balancer (NLB) in a single AZ cannot provide cross-AZ failover; the question requires serving traffic during a single-AZ failure, which demands multi-AZ architecture. An NLB alone does not address the lack of application server redundancy.

E

A read replica does not provide automatic failover; the application would need to manually switch to the replica endpoint, which does not address the single-AZ failure of the primary database. The question requires continued serving traffic during a single-AZ failure, which Multi-AZ provides by automatic failover.

20
Matchingmedium

Match the disaster recovery strategy to the recovery posture it best fits for a Regional outage.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Lowest cost option where the environment is rebuilt from backups and hours of downtime are acceptable.

Keep only the critical core running in the secondary Region, then scale out after failover.

Run a scaled-down but functional environment in another Region for faster cutover.

Serve production traffic from more than one Region at the same time for the fastest recovery.

Why these pairings

These pairs match disaster recovery strategies to their recovery postures, aligning with AWS DR strategies where RTO and RPO define the recovery objectives.

21
MCQmedium

A payments service receives payment orders by consuming messages from an Amazon SQS Standard queue. The downstream processor occasionally exceeds its processing timeout. As a result, some messages reappear in the queue and may be processed more than once. The team wants to prevent duplicate side effects (for example, double-charging) and also ensure poison messages do not repeatedly consume processing capacity. What approach best satisfies both goals?

A.Implement idempotent processing (for example, store processed payment IDs in DynamoDB) and configure an SQS dead-letter queue (DLQ) using a redrive policy with an appropriate maxReceiveCount.
B.Rely only on increasing the SQS visibility timeout so duplicates rarely occur, without adding idempotency checks or a DLQ.
C.Switch to a FIFO queue and delete messages immediately upon receipt to avoid duplicates.
D.Move the workload to SNS and use synchronous HTTP endpoints so the sender retries until the receiver confirms success.
AnswerA

With SQS Standard’s at-least-once delivery, duplicates can occur. Idempotency ensures repeated processing of the same payment ID does not create duplicate side effects. A DLQ with redrive policy isolates poison messages: after a message is received and fails processing more than maxReceiveCount times, SQS moves it to the DLQ instead of cycling it back to the main queue indefinitely.

Why this answer

It addresses both requirements: idempotent processing (e.g., storing processed payment IDs in DynamoDB) ensures that even if a message is processed more than once, duplicate side effects like double-charging are prevented. Configuring an SQS dead-letter queue (DLQ) with a redrive policy and an appropriate maxReceiveCount (e.g., 3 or 5) automatically moves messages that exceed the maximum number of receives to the DLQ, preventing poison messages from repeatedly consuming processing capacity.

Exam trap

The trap here is that candidates often confuse 'exactly-once delivery' (FIFO queues) with 'exactly-once processing,' failing to realize that idempotency is still required to handle failures after message receipt, and that a DLQ is necessary to manage poison messages regardless of queue type.

Why the other options are wrong

B

Increasing visibility timeout reduces duplicates but does not guarantee idempotency; messages can still be processed multiple times if the timeout is exceeded. It also fails to handle poison messages that repeatedly fail processing.

C

FIFO queues guarantee exactly-once processing, but the question states messages reappear due to processing timeout; deleting immediately upon receipt would lose messages that fail processing, and FIFO does not prevent duplicate side effects if processing is not idempotent.

D

SNS with synchronous HTTP endpoints does not guarantee exactly-once processing; the sender may still retry, and the receiver could process duplicates. It also lacks a mechanism to handle poison messages that repeatedly fail, as there is no dead-letter queue.

22
Multi-Selecthard

A financial services firm runs a batch settlement job on a fleet of Amazon EC2 instances that pull work from an Amazon SQS queue. The job must not lose messages if an instance is terminated mid-processing, and duplicate processing must be minimized because each settlement charge is expensive. The team also wants to avoid indefinite reprocessing of a message that repeatedly fails. Which two changes should the solutions architect make to meet these requirements? (Choose two.)

Select 2 answers
A.Switch the queue to an Amazon SQS FIFO queue with content-based deduplication.
B.Increase the queue's visibility timeout so it exceeds the maximum expected processing time.
C.Store a processing flag in an Amazon ElastiCache for Redis cluster and check it before charging.
D.Enable long polling on the queue by setting ReceiveMessageWaitTimeSeconds to 20.
E.Configure a dead-letter queue on the source queue with a maximumReceiveCount.
AnswersB, E

If the visibility timeout is shorter than the processing time, the message becomes visible again and a second consumer can pick it up, causing duplicate settlement charges. Setting the timeout above the worst-case processing duration prevents this overlap and is essential for minimizing duplicates in a pull-based worker fleet.

Why this answer

Raising the visibility timeout above the maximum processing time prevents a still-running message from being redelivered to another consumer, which directly reduces duplicate settlement charges. Attaching a dead-letter queue with a maximumReceiveCount moves repeatedly failing messages aside after a set number of receives, so they are not reprocessed indefinitely while healthy traffic continues.

Exam trap

The trap here is assuming that switching to a FIFO queue alone guarantees no duplicates, when the visibility timeout and a dead-letter queue are what actually control redelivery and poison messages.

23
MCQhard

A company runs a stateful workload on Amazon EC2 instances in an Auto Scaling group. The workload writes session data to the instance store and to an Amazon EBS volume attached at launch. The company wants the workload to survive an Availability Zone failure without losing session data. What should the solutions architect do?

A.Create an Amazon EBS snapshot schedule for the attached volumes and restore the snapshots in another Availability Zone after a failure.
B.Move session data to Amazon ElastiCache for Redis with Multi-AZ enabled, and configure the Auto Scaling group to span multiple Availability Zones.
C.Attach an additional EBS volume to each instance and configure RAID 0 across the two volumes for higher durability.
D.Enable termination protection on the instances and configure the Auto Scaling group to use a single Availability Zone.
AnswerB

Storing session data in ElastiCache for Redis with Multi-AZ provides automatic failover to a replica in another Availability Zone. Spreading the Auto Scaling group across zones lets replacement instances launch in a healthy zone and reconnect to the session store, preserving sessions through a zone failure.

Why this answer

Session data must be moved off instance-bound storage to a multi-AZ data store. ElastiCache for Redis with Multi-AZ automatically fails over to a replica in another zone, and an Auto Scaling group spanning multiple zones can launch replacement instances that reconnect to the session store.

Exam trap

The trap here is assuming that EBS snapshots, RAID, or termination protection can preserve live session data when an entire Availability Zone fails.

24
MCQmedium

A fintech company has a two-Region DR requirement: RPO must be within 15 minutes and RTO must be under 2 hours. To control cost, they do not want to run full production infrastructure in the secondary Region continuously. They plan to continuously replicate the database and keep the application infrastructure in the secondary Region prepared, but at reduced capacity. Which DR strategy best matches this requirement and accurately describes their plan?

A.Pilot light: keep only minimal components (for example, replicated storage and a small amount of core services), so the app scales up during a disaster.
B.Warm standby: keep the essential parts of the application running in the secondary Region at reduced capacity, while using database replication to meet the RPO.
C.Active-active: run the application fully in both Regions with synchronized writes and share traffic continuously.
D.Cold standby: store backups in the secondary Region and provision all infrastructure only during a disaster.
AnswerB

Warm standby runs a fully functional but reduced-capacity version of the application stack in the secondary Region, including application servers and a database that is continuously replicated from the primary Region (for example, via Amazon RDS cross-Region read replicas or Aurora Global Database). Because the environment is already running, failover only requires scaling up the existing infrastructure and promoting the replica, which keeps RTO well under 2 hours. Continuous replication also keeps data lag within the 15-minute RPO, making this the correct choice for the given constraints.

Why this answer

Warm standby is the correct strategy because it runs a scaled-down version of the production application in the secondary Region continuously, with database replication (e.g., Amazon RDS Multi-Region or Aurora Global Database) meeting the 15-minute RPO. The reduced-capacity infrastructure can be scaled up within the 2-hour RTO during a disaster, balancing cost and recovery requirements.

Exam trap

The trap here is confusing pilot light with warm standby: candidates often think any pre-provisioned infrastructure qualifies as pilot light, but warm standby explicitly runs the application at reduced capacity, whereas pilot light keeps only core services and storage without running the application stack.

How to eliminate wrong answers

Option A is wrong because pilot light keeps only minimal core services and storage, not a running application at reduced capacity, and requires provisioning and scaling up compute resources during a disaster, which may not meet the 2-hour RTO if scaling takes significant time. Option C is wrong because active-active runs the application fully in both Regions with synchronized writes and continuous traffic sharing, which violates the cost control requirement of not running full production infrastructure continuously. Option D is wrong because cold standby stores only backups and provisions all infrastructure during a disaster, leading to RTOs that typically exceed 2 hours due to provisioning and data restoration delays.

25
MCQeasy

A media company stores generated video thumbnails in an Amazon S3 bucket. The bucket currently uses the S3 Standard storage class, and the objects are accessed frequently for the first 30 days and then almost never. The company wants to reduce storage costs automatically without changing the application and must retain the objects for at least one year. Which action should a solutions architect take?

A.Enable S3 Intelligent-Tiering on the bucket and let AWS move the objects between access tiers automatically.
B.Configure an S3 Lifecycle rule to transition objects to S3 One Zone-IA after 30 days and expire them after 365 days.
C.Configure an S3 Lifecycle rule to transition objects to S3 Standard-IA after 30 days and to S3 Glacier Instant Retrieval after 90 days.
D.Configure an S3 Lifecycle rule to transition objects to S3 Standard-IA after 30 days and retain them for the required period.
AnswerD

S3 Standard-IA is designed for data accessed infrequently but requiring rapid access when needed, and its lower storage price with a 30-day minimum duration matches the access pattern described. A lifecycle transition applies automatically without application changes and preserves the objects for the required retention period.

Why this answer

The access pattern is known and stable: frequent for about 30 days, then rarely. A lifecycle rule that transitions objects to S3 Standard-IA after 30 days lowers storage cost automatically, requires no application change, and keeps the objects available for the required retention period.

Exam trap

The trap here is reaching for S3 Intelligent-Tiering when the access pattern is already known and predictable rather than unknown.

26
Multi-Selectmedium

A logistics company runs an order-processing workflow using AWS Step Functions. A task state invokes a Lambda function that charges customer credit cards through a third-party gateway. Occasionally the gateway times out, and the workflow fails even though the charge may have succeeded. The architect must make the workflow resilient to these transient failures and avoid duplicate charges. (Choose two.)

Select 2 answers
A.Increase the Lambda function's reserved concurrency to the account limit so more charge requests can run in parallel.
B.Configure a Retry policy on the task state with an exponential backoff and a bounded MaxAttempts value for the relevant error names.
C.Add a Catch field to the task state that transitions to a Fail state so the workflow stops cleanly on any error.
D.Generate an idempotency key for each charge and pass it to the payment gateway so repeated invocations are deduplicated.
E.Enable AWS X-Ray tracing on the state machine and Lambda function to record where the timeouts occur.
AnswersB, D

A Retry policy with exponential backoff lets Step Functions re-invoke the task when the gateway times out, absorbing transient failures without abandoning the workflow. Bounding MaxAttempts prevents infinite retries that could amplify load on a struggling gateway. Because retries happen within the state machine, the workflow resumes automatically once the gateway responds, satisfying the resilience requirement.

Why this answer

Resilience to transient gateway timeouts requires the workflow to retry the task, and safety requires that retries not charge the customer twice. A bounded Retry policy with exponential backoff handles the transient failure, while an idempotency key carried into each attempt lets the payment gateway recognize and deduplicate repeated requests. Together they deliver both reliability and correctness.

Exam trap

The trap here is assuming that adding retries alone is safe, when retrying a non-idempotent charge without a deduplication key can create duplicate payments.

27
MCQmedium

A healthcare company runs a stateless patient-intake API on a fleet of Amazon EC2 instances in a single VPC. The compliance team requires the workload to survive the complete loss of one Availability Zone with no manual intervention, and the instances must be replaced automatically if they fail health checks. The application stores no local state and writes all data to Amazon RDS. Which approach meets these requirements with the LEAST operational effort?

A.Launch the instances directly with a launch template into two subnets in different Availability Zones, and attach each instance to a Network Load Balancer with cross-zone load balancing enabled.
B.Place the instances behind an Application Load Balancer in one Availability Zone and configure an Amazon Route 53 health check with a failover record pointing to a static Elastic IP address.
C.Deploy the instances as an Auto Scaling group in a single Availability Zone and enable detailed CloudWatch monitoring with a scaling policy based on CPU utilization.
D.Create an Auto Scaling group spanning two Availability Zones with a launch template, and attach it to an Application Load Balancer target group with health checks enabled.
AnswerD

An Auto Scaling group across multiple Availability Zones redistributes capacity when a zone is lost, and its health checks replace instances that fail. Attaching the group to an Application Load Balancer target group means ELB health checks drive replacement of unhealthy instances. This satisfies both zone-failure survival and automatic instance replacement with minimal operational overhead.

Why this answer

Resilience to a full Availability Zone loss requires capacity spread across at least two zones, and automatic instance replacement requires an Auto Scaling group whose health checks are tied to the load balancer. Combining a multi-AZ Auto Scaling group with an Application Load Balancer target group delivers both properties without manual intervention, which is exactly what the compliance requirement demands.

Exam trap

The trap here is assuming that a load balancer alone provides high availability, when it only distributes traffic and does not replace failed compute capacity.

28
MCQhard

A payments API uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure?

A.A FIFO queue without a redrive policy
B.A dead-letter queue with an appropriate maxReceiveCount
C.A larger message retention period only
D.Short polling instead of long polling
AnswerB

A dead-letter queue combined with a redrive policy's maxReceiveCount is the correct solution. Each time a message is received, SQS increments its receive count; when the count exceeds maxReceiveCount, SQS moves the message from the source queue into the configured DLQ, isolating it from normal traffic. This lets the payments consumer continue processing healthy messages while the poison message is held for investigation. An appropriate threshold (for example, 3 to 5) balances retrying transient errors against letting a permanently bad message consume all subsequent visibility timeouts and retries.

Why this answer

B is correct because a dead-letter queue (DLQ) with an appropriate maxReceiveCount allows the payments API to isolate poison messages after a specified number of failed processing attempts. This prevents repeated failures from blocking useful retries, as the problematic messages are moved to the DLQ for manual inspection or separate handling, while the main queue continues processing valid messages.

Exam trap

The trap here is that candidates often confuse increasing the retention period or switching polling methods as solutions for poison messages, when the correct mechanism is a dead-letter queue with a maxReceiveCount to limit retries.

How to eliminate wrong answers

Option A is wrong because a FIFO queue without a redrive policy does not automatically handle poison messages; without a DLQ, failed messages will continue to be retried indefinitely, blocking the queue. Option C is wrong because increasing the message retention period only extends how long messages stay in the queue, but does not address the repeated failure and blocking caused by poison messages. Option D is wrong because short polling (vs. long polling) affects how often the queue is polled for messages, not the handling of poison messages or retry behavior.

29
MCQeasy

A content publishing system exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most?

A.IAM Access Analyzer
B.AWS Backup Vault Lock
C.CloudFront caching with appropriate TTLs
D.S3 Select
AnswerC

CloudFront caching with appropriate TTLs is correct because CloudFront edge locations store copies of S3 objects and serve them directly to viewers based on the configured cache TTL. If the S3 origin becomes temporarily unavailable, requests can still be fulfilled from cached content at the edge as long as the object is still in the cache and its TTL has not expired—this effectively masks origin downtime. Choosing appropriate TTLs (e.g., long TTL for immutable static assets) maximizes the chance that cached copies remain available during an outage while balancing freshness, and CloudFront can even serve stale content when it fails to reach the origin.

Why this answer

CloudFront caches responses from the S3 origin based on configured TTLs (Cache-Control or Expires headers). If the S3 origin experiences a short outage, CloudFront can still serve cached content to users until the TTL expires, maintaining availability. This is the most direct way to ensure users receive pages during transient origin failures.

Exam trap

The trap here is confusing data protection features (like Backup Vault Lock) or data retrieval features (like S3 Select) with caching mechanisms that directly improve availability during origin outages.

How to eliminate wrong answers

Option A is wrong because IAM Access Analyzer helps identify unintended access to resources but does not provide caching or origin failover capabilities. Option B is wrong because AWS Backup Vault Lock prevents deletion of backups but does not affect content delivery or caching behavior. Option D is wrong because S3 Select is a feature to retrieve subsets of object data using SQL queries, not a mechanism for caching or serving static content during outages.

30
MCQmedium

A company runs an application behind an Application Load Balancer (ALB). An Auto Scaling group (ASG) is configured with desired capacity 2, but it is attached only to subnets in a single Availability Zone. The ALB is healthy because it is configured across multiple Availability Zones. When the Availability Zone that contains the ASG subnets experiences an outage, what change most directly improves resilience and allows capacity to be restored automatically?

A.Update the ASG to use subnet IDs that span at least two Availability Zones so it can launch replacement instances after an AZ outage.
B.Reduce the ALB health check interval to speed up detection of unhealthy targets.
C.Enable connection draining on the ALB so existing requests complete before targets are terminated.
D.Increase the ASG desired capacity from 2 to 6 to compensate for the missing subnets.
AnswerA

If the ASG is attached to subnets in multiple Availability Zones, when instances in the failed AZ become unhealthy/terminate, Auto Scaling can launch new instances in the remaining AZs to restore the desired capacity. This directly addresses the root cause: the ASG cannot create capacity outside the AZs it is configured for.

Why this answer

An Auto Scaling group (ASG) can only launch instances into the subnets explicitly assigned to it. If those subnets reside in a single Availability Zone (AZ) and that AZ fails, the ASG has no capacity to launch replacement instances, even though the ALB is multi-AZ. By configuring the ASG with subnet IDs spanning at least two AZs, the ASG can automatically launch instances in a healthy AZ, restoring capacity and resilience.

Exam trap

The trap here is that candidates assume a multi-AZ ALB automatically makes the entire architecture resilient, overlooking that the ASG must also be configured with subnets in multiple AZs to launch replacement instances after an AZ failure.

Why the other options are wrong

B

Reducing the ALB health check interval speeds up detection of unhealthy targets but does not address the root cause: the ASG is confined to a single AZ. Without instances in other AZs, the ASG cannot launch replacements during an AZ outage.

C

Connection draining helps complete in-flight requests before terminating instances, but it does not improve resilience or automatically restore capacity after an AZ outage. The issue is the ASG's lack of multi-AZ subnets, not request handling during termination.

D

Increasing desired capacity does not address the root cause: the ASG is confined to a single AZ. During an AZ outage, all instances in that AZ become unavailable, and the ASG cannot launch replacements because no subnets exist in other AZs. Higher desired capacity does not help if there are no subnets to launch into.

31
MCQhard

A logistics company runs an order-processing workload that reads messages from an Amazon SQS queue and writes results to an Amazon DynamoDB table. Occasionally the same order is processed twice and produces duplicate shipments. The architects must ensure each order is processed exactly once end to end, while keeping throughput as high as possible. What should they do?

A.Keep the SQS Standard queue and increase the visibility timeout to several hours.
B.Use an SQS FIFO queue with a message group ID per order and a DynamoDB conditional write on the order ID.
C.Enable DynamoDB Streams on the table and process changes with a second Lambda function.
D.Move the queue to Amazon SNS and subscribe a Lambda function to fan out the orders.
AnswerB

A FIFO queue with content-based deduplication and per-order message group IDs delivers each message once and preserves order within a group, while a DynamoDB conditional write using an order ID attribute rejects a second attempt even if a retry slips through. Together they provide exactly-once processing with high parallelism across order groups.

Why this answer

Exactly-once processing requires both a transport that deduplicates and orders messages per key, and an idempotent write at the destination. An SQS FIFO queue with per-order message group IDs provides the first, and a DynamoDB conditional write on the order ID provides the second, so retries cannot create duplicate shipments.

Exam trap

The trap here is assuming that a longer visibility timeout on a Standard queue eliminates duplicates, when Standard queues only guarantee at-least-once delivery.

32
MCQmedium

A global application experiences frequent writes and must survive a full Regional outage with near-zero data loss. The product team also requires that users can continue to write during the incident using the closest Region. Which approach is most aligned with these requirements?

A.Use an active/active design with multi-Region data replication (for example, global tables for the write-heavy datastore) and route traffic to multiple Regions based on health and latency.
B.Use warm standby with periodic backups of the primary write datastore every 24 hours.
C.Use pilot light where the secondary Region runs only infrastructure templates and starts data replication only after detecting failure.
D.Use a single-writer model in one Region and deploy read-only replicas in the other Region for continuity.
AnswerA

Active/active with multi-Region replication (such as DynamoDB global tables) allows writes to succeed in multiple AWS Regions simultaneously, ensuring near-zero RPO and immediate write availability during a Regional failure. Routing traffic by health and latency distributes load intelligently and automatically shifts users to the nearest healthy Region, which directly satisfies the global, write-heavy workload's requirement for continuous writes with minimal data loss.

Why this answer

An active/active design with multi-Region data replication, such as Amazon DynamoDB global tables, allows writes to occur in any Region and replicates them to all other Regions with near-real-time latency (typically sub-second). This meets the requirement for near-zero data loss during a full Regional outage, as data is asynchronously replicated to multiple Regions, and users can continue writing to the closest healthy Region via Route 53 latency-based or geolocation routing.

Exam trap

The trap here is that candidates often confuse 'multi-Region replication' with 'read replicas only' (Option D) or assume that periodic backups (Option B) provide sufficient durability, failing to recognize that near-zero data loss requires continuous asynchronous replication, not batch-based or on-demand replication.

Why the other options are wrong

B

This option does not meet the requirement for near-zero data loss during a full Regional outage, as backups every 24 hours could lose up to a day's worth of writes. It also fails to allow users to continue writing during the incident.

C

Pilot light does not support near-zero data loss because data replication starts only after failure detection, leading to potential data loss during the gap. It also does not allow writes during the incident as the secondary region is not active.

D

A single-writer model cannot survive a full Regional outage with near-zero data loss because writes are only accepted in the primary Region; if that Region fails, writes must stop until failover occurs, violating the requirement for continuous writes during the incident.

33
MCQmedium

A patient portal receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers?

A.AWS WAF
B.Amazon CloudFront
C.Amazon SQS queue
D.Amazon Route 53 weighted routing
AnswerC

Amazon SQS is a distributed message queue that decouples the patient portal from the order-processing backend: the portal sends each order as a message, and SQS durably stores it across multiple Availability Zones until a consumer polls and processes it. During a burst, SQS absorbs the unprocessed messages, allowing the backend to scale out or catch up at its own pace rather than being overwhelmed. The visibility timeout prevents the same message from being processed by multiple workers, while a dead-letter queue captures any messages that fail repeatedly after a configured number of attempts. This buffering plus retry capability directly addresses the bursting workload described in the scenario.

Why this answer

Amazon SQS is the correct choice because it acts as a durable, highly available message buffer between the web tier and the fulfilment workers. It decouples the components, allowing the web tier to enqueue requests immediately without waiting for the downstream service, and the workers can poll and process messages at their own pace. SQS automatically retains messages for up to 14 days and supports retries via a dead-letter queue, ensuring no requests are lost even during spikes.

Exam trap

The trap here is that candidates may confuse a load-balancing or caching service (like CloudFront or Route 53) with a message queue, failing to recognize that only a queue provides durable, asynchronous decoupling and retry capability for request processing.

How to eliminate wrong answers

Option A is wrong because AWS WAF is a web application firewall that filters HTTP/S traffic based on rules (e.g., SQL injection, XSS) and does not provide message buffering, retry logic, or decoupling for asynchronous processing. Option B is wrong because Amazon CloudFront is a content delivery network (CDN) that caches and accelerates static and dynamic content at edge locations; it cannot buffer or persist requests for downstream workers to process asynchronously. Option D is wrong because Route 53 weighted routing distributes DNS traffic across multiple endpoints based on weights, but it operates at the DNS level and cannot absorb spikes or retry failed requests; it provides no queueing or persistence.

34
MCQeasy

An engineering team deploys a stateless web API on EC2 using an Auto Scaling group and an Application Load Balancer (ALB). During a recent test, they noticed that when one Availability Zone was unavailable, traffic failed until new instances were manually launched. Which change most directly improves automatic failover for the compute layer within a single Region?

A.Place the Auto Scaling group in only one subnet so instance launches are simpler.
B.Ensure the ALB and Auto Scaling group span multiple subnets in at least two Availability Zones.
C.Increase the target group deregistration delay to allow old instances to stay longer.
D.Use a Network Load Balancer, but keep all subnets in a single Availability Zone.
AnswerB

An Application Load Balancer is a regional service that routes traffic to healthy targets across the Availability Zones it is enabled in. By spanning the ALB and Auto Scaling group across at least two AZs, an AZ failure leaves the remaining instances serving traffic while health checks automatically redirect requests away from the failed zone, preserving availability for the stateless web API. This aligns with AWS best practices for fault-tolerant multi-AZ architectures.

Why this answer

An Application Load Balancer (ALB) and Auto Scaling group must span multiple subnets in at least two Availability Zones (AZs) to provide automatic failover. When one AZ becomes unavailable, the ALB automatically reroutes traffic to healthy targets in the remaining AZs, and the Auto Scaling group can launch replacement instances in the surviving AZs. This architecture ensures that the compute layer remains available without manual intervention.

Exam trap

The trap here is that candidates often think a single-AZ deployment with a load balancer provides failover, but without multiple AZs, the load balancer itself becomes a single point of failure and cannot reroute traffic when the AZ goes down.

Why the other options are wrong

A

Placing the Auto Scaling group in only one subnet (single AZ) defeats the purpose of high availability; if that AZ fails, all instances are lost, and the ALB has no healthy targets in other AZs to route traffic to, causing complete failure.

C

Increasing the deregistration delay keeps old instances longer, but does not help automatically launch new instances in a healthy AZ when one AZ fails. It only delays connection draining, not failover.

D

Using a single Availability Zone for all subnets does not provide automatic failover; if that zone fails, the NLB and instances become unavailable, which does not solve the problem described.

35
Multi-Selecthard

A payments API requires point-in-time recovery and accidental-delete protection for a DynamoDB table. Which two settings should the architect enable? The team wants the control to be enforceable during normal operations.

Select 2 answers
A.Deletion protection or tightly controlled delete permissions
B.Point-in-time recovery
C.Global secondary indexes
D.DAX
AnswersA, B

Deletion protection on DynamoDB tables is a table-level setting that blocks delete table operations via the API, console, or SDK, while tightly controlled IAM permissions ensure only authorized principals can execute destructive actions. Together they create a robust defense-in-depth mechanism that directly mitigates the risk of accidental table removal. This is a required control for a payments API because losing the entire table would be catastrophic, even if backups exist.

Why this answer

Deletion protection (Option A) prevents accidental table deletion by blocking drop-table operations, which is enforceable during normal operations. Point-in-time recovery (Option B) enables continuous backups with 35-day granularity, allowing restoration to any second within that window. Together, they satisfy the requirements for accidental-delete protection and point-in-time recovery.

Exam trap

The trap here is that candidates often confuse point-in-time recovery with backup solutions like AWS Backup or assume that GSIs or DAX provide data protection, when in fact they serve entirely different purposes (performance optimization and caching).

36
MCQmedium

A healthcare company runs a containerized claims-processing service on Amazon ECS with the Fargate launch type in a single AWS Region. The service must survive the loss of an entire Availability Zone with no manual intervention, and the architecture must keep the same service endpoint for callers. The service is fronted by an Application Load Balancer. Which combination of actions should a solutions architect take to meet these requirements with the LEAST operational overhead?

A.Deploy the ECS tasks on Amazon EC2 launch type instances in an Auto Scaling group that spans one Availability Zone, and attach an Elastic Load Balancer health check to the group.
B.Create a second ECS service in a different AWS Region and use an Amazon Route 53 latency-based routing policy with health checks to send traffic to the Regional Application Load Balancers.
C.Increase the ECS task CPU and memory reservation so each task can absorb the load of a failed Availability Zone, and enable container health checks in the task definition.
D.Configure the ECS service with a desired count of at least two tasks and use the spread placement strategy across the subnets of multiple Availability Zones, then register those subnets with the Application Load Balancer.
AnswerD

Spreading ECS tasks across subnets in multiple Availability Zones with a desired count of two or more means the loss of one Availability Zone leaves running tasks in the remaining zones, and the Application Load Balancer health checks route traffic only to healthy targets. The ALB DNS name stays constant, so callers need no changes and no manual failover step is required.

Why this answer

Resilience to an Availability Zone failure requires capacity to exist in more than one zone simultaneously. An ECS service with a desired count of two or more and the spread placement strategy across multiple Availability Zone subnets keeps tasks running elsewhere when one zone fails, and the Application Load Balancer automatically stops routing to unhealthy targets while preserving a stable DNS endpoint.

Exam trap

The trap here is assuming that scaling up a single task's CPU and memory provides high availability, when resilience against an Availability Zone failure requires additional task replicas placed in separate zones.

37
MCQhard

A claims workflow uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure?

A.A FIFO queue without a redrive policy
B.Short polling instead of long polling
C.A dead-letter queue with an appropriate maxReceiveCount
D.A larger message retention period only
AnswerC

A dead-letter queue (DLQ) paired with a redrive policy that sets a specific maxReceiveCount (for example, 5) is the correct solution. When a message is received from the source queue more times than the configured limit without being deleted, SQS automatically moves the message to the DLQ. This quarantines unfixable messages so they can be analyzed or replayed later, preventing them from consuming worker instances and blocking normal workflow processing.

Why this answer

A dead-letter queue (DLQ) with an appropriate maxReceiveCount allows messages that repeatedly fail processing to be moved out of the source queue after a specified number of receive attempts. This prevents poison messages from blocking the queue and consuming retry capacity, enabling the workflow to continue processing valid messages without interruption.

Exam trap

The trap here is that candidates often confuse increasing the retention period or changing polling behavior with solving poison message issues, when the correct solution is to use a dead-letter queue with a maxReceiveCount to isolate failing messages.

How to eliminate wrong answers

Option A is wrong because a FIFO queue without a redrive policy does not automatically handle poison messages; without a DLQ, failed messages remain in the queue and continue to block retries. Option B is wrong because short polling reduces latency but does not address the issue of poison messages; it returns fewer messages per request and can increase costs, but it does not prevent repeated failures. Option D is wrong because increasing the message retention period only keeps messages in the queue longer; it does not remove or isolate poison messages, so they will continue to fail and block useful retries.

38
MCQmedium

An orders service publishes payment instructions to an Amazon SQS queue. After occasional processing timeouts, the downstream consumer sometimes processes the same instruction twice, resulting in duplicate payment attempts. The team currently uses an SQS Standard queue with a visibility timeout of 2 minutes and relies on the consumer to finish before the timeout expires. What approach best improves resilience against duplicate processing?

A.Decrease visibility timeout to 10 seconds so duplicates are less likely to occur.
B.Make the consumer idempotent using the order ID as a deduplication key, and set the visibility timeout longer than the worst-case processing time.
C.Use an EventBridge rule with a fixed retry policy that only retries when the payload matches exactly.
D.Enable a dead-letter queue (DLQ) only, without changing the queue type or consumer logic.
AnswerB

SQS Standard provides at-least-once delivery, so duplicates can still occur. The most resilient design is to make the payment handler idempotent so repeated deliveries do not create duplicate side effects, and to set the visibility timeout long enough to cover the worst-case processing time to reduce unnecessary re-delivery.

Why this answer

Making the consumer idempotent using the order ID as a deduplication key ensures that even if the same message is processed multiple times, the downstream system will only apply the payment once. Setting the visibility timeout longer than the worst-case processing time prevents the message from becoming visible again before the consumer finishes, eliminating the root cause of duplicate processing in a Standard queue.

Exam trap

The trap here is that candidates often think reducing the visibility timeout or adding a DLQ alone solves duplicates, but they overlook that Standard queues inherently allow at-least-once delivery, so idempotency is the only reliable solution.

How to eliminate wrong answers

Option A is wrong because decreasing the visibility timeout to 10 seconds would increase the likelihood of duplicates by making the message reappear sooner if the consumer takes longer than 10 seconds, exacerbating the timeout issue. Option C is wrong because an EventBridge rule with a fixed retry policy does not address duplicate processing; EventBridge is a event bus service, not a queue, and its retry policy cannot prevent duplicate delivery from SQS. Option D is wrong because enabling only a DLQ without changing the queue type or consumer logic does not prevent duplicates; a DLQ captures failed messages but does not make the consumer idempotent or adjust visibility timeout to avoid reprocessing.

39
MCQmedium

An orders system sends payment instructions to an Amazon SQS queue. The consumer sometimes times out after it has already created the payment record but before it deletes the SQS message. As a result, the same instruction can be processed more than once. Which design best ensures the consumer remains resilient and does not create duplicate payments when the same instruction is delivered multiple times?

A.Assume the consumer will always delete the SQS message in the same execution path, and ignore the timeout case.
B.Use idempotency: store a deterministic payment request identifier in a DynamoDB table and only create a payment when a conditional write indicates it was not processed before.
C.Switch to SQS Standard because it provides exactly-once delivery, so duplicates cannot happen.
D.Increase the consumer timeout and reduce the number of retries so that duplicates rarely occur.
AnswerB

Implementing idempotency with a DynamoDB table keyed by a deterministic payment request ID (for example, an MD5/SHA-256 hash of the order ID, amount, and currency) lets a consumer use a conditional PutItem with ConditionExpression 'attribute_not_exists(id)'. A successful conditional write claims the ID for the first request; a failed write on retry tells the consumer the payment was already handled, so no charge is created again. This turns at-least-once SQS delivery into effectively exactly-once processing because duplicate messages either see the existing record or lose the race to create it, and the payment action is triggered only on the first successful claim.

Why this answer

It implements idempotency using a DynamoDB table with a conditional write. By storing a deterministic payment request identifier (e.g., a hash of the message body) and only creating the payment if the conditional write succeeds (i.e., the identifier does not already exist), the consumer can safely process the same SQS message multiple times without creating duplicate payments. This pattern ensures resilience against the at-least-once delivery semantics of SQS and consumer timeouts that prevent message deletion.

Exam trap

The trap here is that candidates assume SQS FIFO queues provide exactly-once delivery, but the question specifies an SQS queue (likely Standard), and even FIFO queues only guarantee exactly-once processing within a limited deduplication window, not absolute idempotency; the correct solution is to make the consumer itself idempotent.

How to eliminate wrong answers

Option A is wrong because ignoring the timeout case violates the principle of designing for failure; SQS guarantees at-least-once delivery, and timeouts are a real-world occurrence that must be handled explicitly. Option C is wrong because SQS Standard does not provide exactly-once delivery; it offers at-least-once delivery, and duplicates can still occur due to network retries or consumer failures. Option D is wrong because increasing the consumer timeout and reducing retries only reduces the probability of duplicates but does not eliminate them, and it does not address the fundamental issue of at-least-once delivery semantics.

40
MCQhard

A logistics company runs an order processing system on Amazon EC2 instances that read and write to an Amazon RDS for MySQL database. The database is currently a Single-AZ deployment. The company needs the database to survive an Availability Zone failure with automatic failover and minimal downtime. The application connects using a hardcoded DNS name. Which change should a solutions architect make?

A.Convert the RDS for MySQL database to a Multi-AZ DB instance deployment and continue using the existing endpoint.
B.Enable automated backups with a longer retention period and restore the database to a new instance during a failure.
C.Create a read replica in another Availability Zone and update the application to use the read replica endpoint for writes.
D.Migrate the database to Amazon DynamoDB with global tables enabled.
AnswerA

Converting to a Multi-AZ DB instance deployment creates a synchronous standby in a different Availability Zone and provides automatic failover. The DNS endpoint remains the same, so the hardcoded connection string continues to work after failover. This meets the requirements for zone resilience and minimal downtime with the least application change.

Why this answer

Converting the RDS for MySQL database to a Multi-AZ DB instance deployment creates a synchronous standby in a second Availability Zone and enables automatic failover. The database endpoint remains unchanged, so the application's hardcoded DNS name continues to work, satisfying the requirements for zone resilience and minimal downtime.

Exam trap

The trap here is confusing a read replica with a Multi-AZ standby; a read replica is asynchronous and does not provide automatic failover for writes.

41
Multi-Selectmedium

A customer portal must recover from a regional outage within a few hours. The business wants lower ongoing cost than a fully active second Region and does not want to rebuild everything from scratch during the outage. Which two DR patterns best fit that goal? Select two.

Select 2 answers
A.Backup and restore
B.Pilot light
C.Warm standby
D.Multi-site active-active
E.Single-AZ deployment
AnswersB, C

Pilot light keeps only core components such as the database replication running in the second Region, so recovery within hours is achievable without rebuilding everything, while ongoing cost stays well below an active-active deployment. It matches the lower-cost, few-hours RTO constraint.

Why this answer

Pilot light (B) is correct because it keeps a minimal, always-on core (for example replicated data and pre-provisioned but scaled-down compute) in the DR Region, so recovery only requires scaling up and starting the remaining resources rather than rebuilding everything from scratch, and its low baseline footprint keeps ongoing cost well below a fully active second Region. Warm standby (C) is correct because it runs a scaled-down but fully functional copy of the workload in the second Region, allowing it to be scaled up to production capacity within hours while still costing less than a fully active second Region. Backup and restore (A) is not the best fit because recovery means redeploying and rebuilding the entire environment from backups, which typically exceeds a few hours and requires substantial manual effort.

Multi-site active-active (D) is excluded because running full production capacity in two Regions simultaneously is the most expensive option and contradicts the goal of lower ongoing cost. Single-AZ deployment (E) is not a cross-Region DR pattern at all and provides no protection against a regional outage.

Exam trap

AWS often tests the distinction between pilot light and warm standby—the trap here is that candidates may confuse pilot light with backup and restore, not realizing that pilot light maintains a live, minimal environment (e.g., database replicas) rather than just backup files, enabling faster recovery without full rebuild.

Why the other options are wrong

A

Backup and restore typically has a Recovery Time Objective (RTO) of hours to days, which may not meet the 'within a few hours' requirement, and it often involves rebuilding infrastructure from backups, which the question explicitly wants to avoid.

D

Multi-site active-active requires fully active resources in two regions simultaneously, which incurs higher ongoing costs than a fully active second Region, contradicting the requirement for lower cost.

E

Single-AZ deployment does not provide any cross-region recovery capability; a regional outage would cause complete downtime, contradicting the requirement to recover within hours.

42
MCQhard

A media company stores master video files in an Amazon S3 bucket in the us-east-1 Region. A compliance policy requires that the data remain readable even if the entire us-east-1 Region becomes unavailable, and the recovery point objective is 15 minutes. The team wants the lowest operational overhead and does not want to modify application code. Which solution should the architect implement?

A.Use AWS DataSync to copy the bucket to a second Region every hour and store the copy in an S3 bucket with versioning enabled.
B.Configure an S3 Lifecycle policy to transition objects to S3 Glacier Deep Archive in a second Region and restore them during a disaster.
C.Enable S3 Cross-Region Replication to a bucket in us-west-2 with S3 Replication Time Control set to 15 minutes and enable versioning on both buckets.
D.Enable S3 Transfer Acceleration on the bucket and use an Amazon CloudFront distribution with origin failover to serve the video files.
AnswerC

S3 Cross-Region Replication with Replication Time Control provides a predictable 15-minute replication SLA and replicates objects, versions, and metadata to the destination bucket. The application can be pointed to the destination bucket during a Regional outage, and no code changes are needed because S3 APIs remain the same, meeting both RPO and low-overhead requirements.

Why this answer

Cross-Region Replication with S3 Replication Time Control replicates new and updated objects to a second Region within a defined 15-minute window and supports versioning and metadata. The destination bucket can serve reads during a Regional outage without application changes, giving the required RPO with minimal operational effort.

Exam trap

The trap here is confusing S3 Lifecycle storage-class transitions with cross-Region replication, when lifecycle actions keep data in the same Region and archival restores are far too slow for a 15-minute RPO.

43
MCQhard

A financial analytics platform runs an Amazon Aurora MySQL cluster with one writer and two readers. During month-end reporting, read traffic spikes and the application sometimes receives TooManyConnections errors on the reader endpoint. The architect wants to absorb bursts without changing application code and must keep failover behaviour intact. Which change meets these requirements?

A.Increase the aurora_max_connections parameter on the cluster parameter group and reboot the readers
B.Replace the reader endpoint with a Network Load Balancer that distributes TCP connections across the Aurora readers
C.Add an Amazon RDS Proxy in front of the Aurora cluster and point the application at the proxy endpoint
D.Convert the readers to Aurora Replicas in a second Region and use a global database secondary cluster
AnswerC

RDS Proxy pools and shares database connections, so bursts of application connections are multiplexed onto a smaller set of database connections, eliminating TooManyConnections errors without code changes beyond the endpoint. It preserves Aurora failover by automatically redirecting to the new writer, and it can be associated with read-only endpoints for reader traffic. This directly addresses connection exhaustion while keeping resilience.

Why this answer

Amazon RDS Proxy sits between the application and Aurora, maintaining a warm pool of database connections that are shared across many client connections. This absorbs connection bursts and prevents TooManyConnections errors without application changes beyond updating the endpoint. Because RDS Proxy is failover-aware, it automatically routes to the new writer during a failover, preserving the resilience behaviour the architect requires.

Exam trap

The trap here is treating a load balancer as a connection pooler, when only RDS Proxy multiplexes client connections onto a smaller backend set.

44
MCQmedium

A logistics company runs a shipment-tracking service on a single Amazon EC2 instance in one Availability Zone. The instance stores tracking state in an attached Amazon EBS volume and writes nightly backups to Amazon S3. The company needs the service to survive the loss of an entire Availability Zone with minimal data loss and automatic recovery, while keeping changes minimal. Which design change should a solutions architect recommend?

A.Enable termination protection and detailed monitoring on the existing instance, and increase the EBS volume size to improve durability.
B.Create an Amazon Machine Image of the instance and place an Auto Scaling group behind an Application Load Balancer across at least two Availability Zones, storing state in Amazon RDS Multi-AZ.
C.Take more frequent EBS snapshots and copy them to another AWS Region so the volume can be restored after an outage.
D.Attach an additional Elastic Network Interface to the instance and assign an Elastic IP address so clients reconnect quickly after a failure.
AnswerB

Distributing stateless compute across multiple Availability Zones with an Auto Scaling group and an Application Load Balancer removes the single-AZ failure domain, and moving state to Amazon RDS Multi-AZ provides a synchronously replicated standby with automatic failover, so the service survives an AZ loss with minimal data loss and no manual intervention.

Why this answer

Surviving an Availability Zone failure requires eliminating single-AZ dependencies. Running stateless application instances in an Auto Scaling group behind an Application Load Balancer spreads compute across multiple zones and replaces failed capacity automatically, while Amazon RDS Multi-AZ keeps a synchronously updated standby in a different zone that takes over with the same endpoint, satisfying both automatic recovery and minimal data loss.

Exam trap

The trap here is assuming that making a single instance more durable, through snapshots, monitoring, or network addressing, converts it into a highly available system.

45
MCQeasy

A production Amazon RDS database has automated backups enabled. At 10:45 UTC, an issue is discovered. The team needs to restore the database to its state as of 10:30 UTC. Which capability should they use?

A.Point-in-time restore (PITR) using automated backups to a specific timestamp.
B.Perform a Multi-AZ manual failover of the standby to recover to the earlier timestamp.
C.Promote a cross-region replication target to replace the current database with the last-known good copy.
D.Switch to a read replica to access an older view of data without restoring.
AnswerA

Point-in-time restore (PITR) leverages RDS's automated backups and transaction logs to create a new DB instance reflecting the exact state at 10:30 UTC, precisely within your retention window. The restore process spins up a separate instance, so you can validate the data and then redirect traffic or promote it without altering your current production database. This is the only option that explicitly rolls back committed transactions to a specific timestamp, satisfying the team's requirement.

Why this answer

Amazon RDS automated backups enable point-in-time recovery (PITR) to any second within the backup retention period, restoring to a new DB instance. Since the issue was discovered at 10:45 UTC and the desired recovery point is 10:30 UTC, PITR can restore the database to that exact timestamp, provided it falls within the automated backup window and retention period.

Exam trap

The trap here is confusing Multi-AZ failover or read replicas with point-in-time recovery capabilities, leading candidates to think failover or replica promotion can roll back to a specific past state when they only provide high availability or read scaling.

How to eliminate wrong answers

Option B is wrong because Multi-AZ failover switches to a standby replica that is kept synchronously in sync with the primary; it does not provide a way to roll back to an earlier point in time, only to the current state of the primary. Option C is wrong because cross-region replication (e.g., using a read replica in another region) replicates data asynchronously and cannot be used to restore to a specific past timestamp; promoting it would give you a copy from a lagged point, not necessarily 10:30 UTC. Option D is wrong because a read replica provides a live, near-real-time copy of the primary database and does not retain historical snapshots or allow accessing an older view of data without a full restore.

46
Multi-Selectmedium

A solutions architect is designing a highly available and resilient architecture for a critical internal application that processes financial transactions. The application runs on Amazon EC2 instances inside an Auto Scaling group. The database layer uses an Amazon Aurora MySQL cluster. The company requires that if an entire AWS Availability Zone (AZ) fails, the application must remain operational with minimal impact and automatically recover without manual intervention. Which combination of architectural decisions will meet these requirements? (Choose four.)

Select 4 answers
.Configure the Auto Scaling group to span at least three Availability Zones in the same AWS Region.
.Deploy the Aurora cluster with a single DB instance to reduce complexity and cost.
.Configure the Aurora cluster to include at least one Aurora Replica in a different Availability Zone than the primary instance.
.Use an Application Load Balancer (ALB) to distribute traffic across EC2 instances in multiple Availability Zones.
.Place the EC2 instances in a single Availability Zone to ensure data locality with the primary database.
.Set up an Amazon RDS Proxy to manage database connections and provide connection pooling for improved resilience.

Why this answer

Configuring the Auto Scaling group to span at least three Availability Zones ensures that if one AZ fails, the remaining AZs have sufficient capacity to handle the load, and the Auto Scaling group can automatically launch new instances in the healthy AZs. Deploying the Aurora cluster with at least one Aurora Replica in a different AZ than the primary instance provides automatic failover to a replica in under 30 seconds, ensuring database resilience without manual intervention. Using an Application Load Balancer (ALB) to distribute traffic across EC2 instances in multiple AZs allows the ALB to automatically route traffic away from failed AZs and only to healthy targets, maintaining application availability.

Setting up an Amazon RDS Proxy manages database connections by pooling and reusing them, which reduces the load on the database during failover and improves resilience by providing seamless connection handling across AZ failures.

Exam trap

The trap here is that candidates often think a single Aurora instance with multi-AZ storage is sufficient, but without an Aurora Replica in a different AZ, automatic failover is not possible; similarly, they may assume that placing all EC2 instances in one AZ simplifies data locality, but this sacrifices availability for a false sense of performance optimization.

47
MCQeasy

A startup runs a single Amazon EC2 instance hosting both a web application and its MySQL database. The founders want the application to survive the failure of the underlying hardware without changing the database engine, and they want the smallest possible operational change. What should the architect recommend?

A.Move the database to an Amazon RDS for MySQL Multi-AZ DB instance and place the web application on an EC2 instance in an Auto Scaling group spanning two Availability Zones.
B.Keep both tiers on one EC2 instance but enable detailed monitoring and create an Amazon CloudWatch alarm that emails the founders when CPU is high.
C.Create an Amazon EBS snapshot schedule for the instance's volume and restore the snapshot to a new instance if the hardware fails.
D.Migrate the database to Amazon DynamoDB and rewrite the application to use the DynamoDB API.
AnswerA

RDS Multi-AZ maintains a synchronous standby replica in a different Availability Zone and fails over automatically if the primary fails, while an Auto Scaling group across two zones can replace a failed web instance. The engine stays MySQL, matching the no-engine-change constraint, and both tiers gain automatic recovery with minimal ongoing operational effort from the founders.

Why this answer

The goal is automatic recovery for both tiers without changing the database engine. RDS for MySQL Multi-AZ provides a managed standby that fails over automatically, and an Auto Scaling group across two Availability Zones can replace a failed web server. Together they remove the single point of failure with far less operational burden than self-managing replication or restoring snapshots.

Exam trap

The trap here is believing that monitoring, snapshots, or backups constitute high availability, when they only improve detection or recovery time rather than preventing downtime.

48
MCQmedium

Your public API is hosted in two regions. You want Route 53 to automatically send traffic to the secondary region when the primary region’s endpoint fails. The primary API health check is returning failure codes, but clients still reach the primary region for several minutes. Which Route 53 configuration most directly addresses this behavior?

A.Use a single Alias A record with simple routing and a short TTL so Route 53 quickly changes the IP address.
B.Use Route 53 failover routing with a primary record and a secondary record, each associated with its own health check, so Route 53 answers with the healthy region.
C.Use weighted routing to send a small percentage of traffic to the secondary region, increasing it manually when the primary fails.
D.Use latency routing only, letting Route 53 choose the lowest-latency region at query time, without health checks.
AnswerB

Failover routing is designed for this: Route 53 evaluates health checks and returns the primary record while it is healthy. When the primary health check fails, Route 53 automatically returns the secondary record. Note that clients may still see traffic for a few minutes due to DNS caching, but failover routing is the configuration that enables automatic region switching.

Why this answer

Route 53 failover routing with health checks on both primary and secondary records ensures that when the primary health check fails, Route 53 stops returning the primary record's IP and instead returns the secondary record's IP. This directly addresses the observed behavior where clients still reach the primary region for several minutes—likely because the primary record's health check was not configured or associated, or a simple routing policy was used without health check integration, causing stale DNS responses to be served until TTL expires.

Exam trap

The trap here is that candidates assume a short TTL alone (Option A) is sufficient for fast failover, but without health checks, Route 53 has no mechanism to detect endpoint failure and will continue returning the primary record until the TTL expires and the record is manually updated, causing the observed delay.

How to eliminate wrong answers

Option A is wrong because simple routing with a short TTL does not incorporate health checks; Route 53 will continue to return the primary record's IP even if the endpoint is unhealthy, and clients will still reach the failing region until the TTL expires and the record is manually updated. Option C is wrong because weighted routing requires manual intervention to adjust weights when the primary fails, which does not provide automatic failover and can still result in clients reaching the unhealthy primary region. Option D is wrong because latency routing without health checks will continue to return the primary region's IP if it has the lowest latency, even when the primary endpoint is returning failure codes, so clients will still be directed to the failing region.

49
MCQeasy

Based on the exhibit, some SQS messages fail validation repeatedly and continue consuming worker time. What change best prevents the bad messages from being retried forever?

A.Increase the visibility timeout so each message has more time to finish processing.
B.Configure a dead-letter queue and a redrive policy for messages that exceed the retry limit.
C.Replace the queue with an Amazon SNS topic so failed messages will not be retried.
D.Increase the number of workers so the queue drains faster during peak load.
AnswerB

Configuring a dead-letter queue (DLQ) with a redrive policy is the canonical AWS pattern for catching poison messages. The redrive policy specifies a maxReceiveCount (e.g., 5); after a message is received that many times without being deleted, SQS automatically moves it to the DLQ. This isolates the repeatedly failing message from the main queue so downstream workers can continue processing healthy messages without hitting the same bad payload repeatedly. Once in the DLQ, you can inspect, correct, or manually reprocess the message after fixing the underlying data issue.

Why this answer

A dead-letter queue (DLQ) with a redrive policy allows messages that have been received a maximum number of times (e.g., after the configured retry limit) to be moved to a separate queue for analysis or manual handling. This prevents the same invalid message from being repeatedly processed by workers, freeing up compute resources and avoiding infinite retry loops.

Exam trap

The trap here is that candidates may think increasing the visibility timeout or adding more workers will solve the retry problem, but neither addresses the root cause of a message that will always fail validation.

How to eliminate wrong answers

Option A is wrong because increasing the visibility timeout only gives workers more time to process a message before it becomes visible again; it does not stop a failing message from being retried indefinitely. Option C is wrong because Amazon SNS is a pub/sub messaging service that does not provide built-in retry logic or a mechanism to move failed messages out of the processing pipeline; it would still deliver the same bad message to subscribers repeatedly. Option D is wrong because adding more workers only increases throughput for valid messages but does not prevent the same invalid message from being retried forever; the bad message will still consume worker time on every retry.

50
MCQmedium

A patient portal receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers? The architecture review board prefers a managed AWS-native control.

A.AWS WAF
B.Amazon CloudFront
C.Amazon SQS queue
D.Amazon Route 53 weighted routing
AnswerC

Amazon SQS is a fully managed message queue that decouples order producers from the processing service, providing a durable buffer for bursty traffic so the backend only pulls messages at a sustainable pace. When a consumer fails or times out, the visibility timeout makes the message available again, and a dead-letter queue can isolate messages that repeatedly fail processing. This lets the patient portal safely absorb order spikes without dropping data, making SQS the right architectural fit.

Why this answer

Amazon SQS is the correct choice because it acts as a decoupling buffer between the web tier and the fulfilment workers. It can absorb sudden bursts of orders by storing messages durably, and workers can poll the queue at their own pace, retrying failed processing without losing any requests. This aligns with the requirement for a managed AWS-native service that handles spikes and retries.

Exam trap

The trap here is that candidates may confuse buffering and decoupling with services like CloudFront (caching) or Route 53 (traffic routing), failing to recognize that SQS is the only AWS-native service designed specifically for asynchronous message queuing and retry logic.

How to eliminate wrong answers

Option A is wrong because AWS WAF is a web application firewall that filters HTTP/S traffic based on rules, not a message queue for buffering and retrying requests. Option B is wrong because Amazon CloudFront is a content delivery network (CDN) that caches and accelerates static/dynamic content delivery, not a service for decoupling and buffering asynchronous workloads. Option D is wrong because Amazon Route 53 weighted routing is a DNS routing policy for distributing traffic across endpoints, not a message queuing or buffering service.

51
MCQeasy

A team runs an Amazon RDS for MySQL database in a single Availability Zone. They want automatic failover with minimal downtime if the primary database instance becomes unavailable. Automated backups are already enabled. Which configuration change best meets the requirement?

A.Keep the deployment as single-AZ, but increase automated backup retention to 35 days.
B.Create a read replica in another Availability Zone, but keep Multi-AZ disabled.
C.Enable RDS Multi-AZ so AWS maintains a standby in another Availability Zone for automatic failover.
D.Rely on restoring from the most recent manual snapshot after an outage.
AnswerC

RDS Multi-AZ creates a standby instance in a different AZ and replicates data to it. If the primary becomes unavailable, AWS performs an automatic failover, promoting the standby and maintaining high availability with minimal application disruption.

Why this answer

Enabling Multi-AZ on Amazon RDS for MySQL automatically provisions and maintains a synchronous standby replica in a different Availability Zone. If the primary instance fails, Amazon RDS automatically fails over to the standby, typically within 60–120 seconds, minimizing downtime without manual intervention. This meets the requirement for automatic failover with minimal downtime.

Exam trap

The trap here is that candidates often confuse read replicas (which are for read scaling and manual promotion) with Multi-AZ (which is for high availability and automatic failover), leading them to incorrectly choose Option B.

How to eliminate wrong answers

Option A is wrong because increasing automated backup retention to 35 days only extends the point-in-time recovery window; it does not provide automatic failover or a standby instance. Option B is wrong because a read replica in another AZ is asynchronous and does not support automatic failover; it requires manual promotion to become the primary, which incurs downtime. Option D is wrong because restoring from a manual snapshot is a manual process that can take significant time (minutes to hours depending on size) and does not provide automatic failover.

52
MCQeasy

A startup runs a nightly batch job on a single Amazon EC2 instance that stores results in an Amazon EBS volume. The job takes six hours, and the team wants to resume from the last completed step if the instance is terminated unexpectedly. Which approach provides the required durability with the least operational effort?

A.Enable EBS Multi-Attach on the volume and attach it to a standby instance in another Availability Zone
B.Store intermediate results in Amazon S3 and write a checkpoint file after each completed step
C.Move the EBS volume to a RAID 1 configuration across two Availability Zones using mdadm
D.Create a snapshot schedule for the EBS volume and restore from the latest snapshot after termination
AnswerB

Amazon S3 provides durable, highly available object storage that survives instance termination, and a checkpoint file lets the job resume from the last completed step. This requires minimal operational effort because S3 is fully managed and needs no volume attachment or replication configuration. It directly satisfies the durability and resume requirements.

Why this answer

Storing intermediate results in Amazon S3 with checkpoint files gives durable, externally persisted state that survives instance termination and lets the job resume from the last completed step. S3 is managed, so there is no replication or volume management to operate, which matches the low-effort requirement. Snapshots and RAID do not provide automatic step-level resume without custom logic.

Exam trap

The trap here is assuming that EBS snapshots alone let a job resume mid-step, when resume logic requires application-level checkpoints.

53
MCQeasy

A company needs an Amazon RDS database that automatically fails over to a standby when the primary DB instance becomes unavailable. Which approach best meets the requirement with minimal operational effort?

A.Keep the DB as a single-AZ instance and implement a manual process to promote a standby when needed.
B.Deploy the DB as a Multi-AZ DB instance so AWS maintains a synchronous standby in another Availability Zone and performs automated failover.
C.Enable versioned backups only, and restore the database each time the primary instance becomes unavailable.
D.Replicate the database to another region and switch clients to the secondary region using manual DNS changes.
AnswerB

RDS Multi-AZ provisions a synchronous standby in a different Availability Zone within the same AWS Region. When the primary DB instance is unavailable, AWS performs automated failover to the standby, reducing downtime without custom scripts.

Why this answer

Amazon RDS Multi-AZ automatically provisions and maintains a synchronous standby replica in a different Availability Zone. When the primary DB instance fails, AWS handles the automatic failover to the standby with zero manual intervention, meeting the requirement with minimal operational effort.

Exam trap

The trap here is that candidates often confuse Multi-AZ (synchronous replication, automatic failover) with Multi-Region (asynchronous replication, manual or automated cross-region failover) or assume that backups alone can provide high availability, but backups do not offer automatic failover or minimal downtime.

How to eliminate wrong answers

Option A is wrong because a single-AZ instance has no standby, and a manual process to promote a standby would require creating a new instance from a snapshot or read replica, which incurs significant downtime and operational overhead. Option C is wrong because versioned backups alone do not provide a standby; restoring from a backup can take minutes to hours, resulting in unacceptable downtime and data loss. Option D is wrong because cross-region replication requires manual DNS changes to redirect traffic, introduces higher latency, and involves more operational complexity than a Multi-AZ deployment within a single region.

54
MCQmedium

A SaaS platform serves an API using two regional deployments: us-east-1 (primary) and us-west-2 (secondary). Each region has its own ALB. The business requires automated DNS-based failover when the primary region becomes unhealthy, and they do not want manual DNS changes during incidents. Which Route 53 configuration is the best match?

A.Create a single Route 53 record using weighted routing across both ALBs with weights adjusted manually during an incident.
B.Use Route 53 failover routing with a primary record pointing to the us-east-1 ALB and a secondary record pointing to the us-west-2 ALB, each using health checks.
C.Use latency-based routing so Route 53 always selects the fastest region; health checks are unnecessary because client latency reflects availability.
D.Use a single A record with a static IP address that points to a NAT gateway, and update that IP during failure events.
AnswerB

Failover routing is the correct active–passive DNS pattern for regional redundancy. You create a primary A record for the us-east-1 ALB with an associated Route 53 health check, and a secondary A record for the us-west-2 ALB; when the health check fails for the primary, Route 53 automatically responds to DNS queries with the secondary record. This shifts client traffic to the healthy secondary region without manual intervention. Route 53 can also use health checks on both records to detect secondary failure, though the primary health check is the key trigger for failover.

Why this answer

Route 53 failover routing is designed specifically for active-passive failover scenarios where you have a primary and secondary resource. By associating health checks with each record, Route 53 automatically detects when the primary ALB in us-east-1 becomes unhealthy and routes traffic to the secondary ALB in us-west-2 without manual intervention. This meets the requirement for automated DNS-based failover without manual DNS changes.

Exam trap

The trap here is that candidates may confuse latency-based routing with failover routing, assuming that lowest latency implies health, but latency routing does not consider endpoint health and will continue sending traffic to an unhealthy region if it is still the fastest.

How to eliminate wrong answers

Option A is wrong because weighted routing requires manual adjustment of weights during an incident, which violates the requirement for automated failover without manual DNS changes. Option C is wrong because latency-based routing selects the region with the lowest latency for each user, not based on health; it does not provide failover when a region becomes unhealthy, and health checks are not used to determine routing decisions. Option D is wrong because using a static IP pointing to a NAT gateway is not a scalable or resilient approach for an API served by ALBs, and updating the IP during failure events requires manual intervention, which contradicts the automation requirement.

55
MCQhard

Based on the exhibit, DNS still sends traffic to the primary Region even though Route 53 health checks show the primary endpoint is unhealthy. What is the best change to make failover work as intended?

A.Change both records to weighted routing with a 50/50 split so Route 53 can shift traffic gradually.
B.Use a failover routing policy with a primary record and a secondary record, and attach the health check to the primary record.
C.Switch to latency-based routing so users are always directed to the lowest-latency Region.
D.Use geolocation routing so clients in one Region are sent to the healthier endpoint.
AnswerB

Failover routing is designed for active-passive DNS behavior. With a primary and secondary record, Route 53 answers with the primary record when it is healthy and returns the secondary record when the primary health check fails. The exhibit shows simple routing, which does not express the failover intent. Switching to failover routing aligns the DNS policy with the stated requirement.

Why this answer

A failover routing policy with a health check attached to the primary record is the only configuration that allows Route 53 to automatically stop sending traffic to an unhealthy primary endpoint and redirect it to the secondary endpoint. Without the health check attached to the primary record, Route 53 has no mechanism to detect the failure and will continue routing traffic to the primary Region, even if the health check status shows unhealthy.

Exam trap

The trap here is that candidates assume Route 53 automatically uses health check status to influence routing regardless of the routing policy, but in reality, health checks only affect routing when explicitly attached to a record in a failover or weighted routing policy.

Why the other options are wrong

A

Weighted routing distributes traffic based on weights, not health; it does not automatically failover when a health check fails, so unhealthy primary would still receive traffic.

C

Latency-based routing directs users to the region with the lowest latency, not based on health. Even if the primary endpoint is unhealthy, it may still receive traffic if it has lower latency, failing to achieve the desired failover.

D

Geolocation routing directs traffic based on the client's geographic location, not health status. Even if the primary endpoint is unhealthy, clients in the primary region would still be routed to it, failing to achieve failover.

56
MCQhard

A patient portal must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used?

A.Instance store volumes
B.Amazon EFS with mount targets in multiple Availability Zones
C.An EBS volume attached to all instances
D.S3 mounted as a POSIX file system without a file gateway
AnswerB

Amazon EFS is a fully managed, regional file storage service that provides a POSIX-compliant NFS file system automatically scaled to meet demand. By creating mount targets in multiple Availability Zones, you enable Linux EC2 instances in those AZs to access the same shared file data simultaneously, with high availability and low latency. This makes EFS the appropriate choice for shared file storage for a patient portal across multiple Linux EC2 instances.

Why this answer

Amazon EFS provides a scalable, fully managed NFS file system that can be mounted concurrently on multiple Linux EC2 instances. By creating mount targets in multiple Availability Zones, the file system remains accessible even if one AZ fails, ensuring high availability and shared file storage across instances.

Exam trap

The trap here is that candidates may confuse EBS multi-attach (which is limited to specific instance types and does not span AZs) with the true multi-AZ shared file system capability of EFS.

How to eliminate wrong answers

Option A is wrong because instance store volumes are ephemeral and tied to a single EC2 instance; they cannot be shared across instances or survive an AZ failure. Option C is wrong because an EBS volume can only be attached to one EC2 instance at a time (unless using multi-attach, which is limited to specific instance types and not designed for shared file storage across AZs). Option D is wrong because mounting S3 as a POSIX file system without a file gateway (e.g., using s3fs-fuse) does not provide consistent POSIX semantics, lacks strong read-after-write consistency, and is not designed for high-availability shared file storage across AZs.

57
MCQmedium

A claims workflow uses an RDS MySQL database and must remain available during an Availability Zone failure with minimal application changes. What should the architect enable?

A.S3 Cross-Region Replication
B.Multi-AZ deployment for the RDS DB instance
C.EBS snapshots every hour
D.Read replicas only
AnswerB

Multi-AZ deployment maintains a synchronous standby replica in a second Availability Zone and automatically fails over the DB instance during an AZ outage, requiring no application connection-string changes. This satisfies the availability and minimal-change constraints.

Why this answer

Multi-AZ deployment for RDS MySQL automatically provisions and maintains a synchronous standby replica in a different Availability Zone. In the event of an AZ failure, Amazon RDS automatically fails over to the standby, providing high availability with minimal application changes (the application simply reconnects to the same endpoint). This meets the requirement for availability during an AZ outage without requiring code modifications.

Exam trap

The trap here is that candidates often confuse read replicas (which are for read scaling and manual promotion) with Multi-AZ (which provides automatic failover and high availability), leading them to select 'Read replicas only' as a cheaper but incorrect alternative.

How to eliminate wrong answers

Option A is wrong because S3 Cross-Region Replication is designed for object-level replication across AWS regions, not for database high availability within a region, and it does not provide automatic failover for an RDS MySQL database. Option C is wrong because EBS snapshots every hour provide point-in-time backup and recovery, not automatic failover; restoring from a snapshot requires manual intervention and results in data loss for transactions after the last snapshot. Option D is wrong because read replicas only provide read scaling and asynchronous replication; they do not support automatic failover for write operations, and promoting a read replica to a primary requires manual action and potential data loss.

58
MCQmedium

A ticket booking system stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured? The architecture review board prefers a managed AWS-native control.

A.S3 lifecycle transition to Glacier Flexible Retrieval
B.An EBS snapshot schedule
C.S3 Cross-Region Replication with versioning enabled
D.A CloudFront distribution
AnswerC

S3 Cross-Region Replication (CRR) asynchronously replicates newly uploaded objects to a destination bucket in a different AWS Region, creating a durable secondary copy for disaster recovery. CRR requires versioning to be enabled on both the source and destination buckets, and it also replicates delete markers and metadata by default. Because the destination bucket is in another Region, applications can fail over and still access the ticket documents if the source Region becomes unavailable.

Why this answer

S3 Cross-Region Replication (CRR) is a fully managed AWS-native feature that automatically replicates objects from a source S3 bucket in one AWS Region to a destination bucket in another Region, meeting the disaster recovery requirement for a geographically separate copy. Enabling versioning on both buckets is mandatory for CRR to function, as it tracks object versions and ensures consistency during replication.

Exam trap

The trap here is that candidates often confuse S3 Cross-Region Replication with S3 lifecycle policies or Glacier transitions, mistakenly thinking that moving data to a cheaper storage class in the same region satisfies a disaster recovery requirement for geographic separation.

How to eliminate wrong answers

Option A is wrong because S3 lifecycle transition to Glacier Flexible Retrieval only moves data within the same bucket and region to a colder storage class for cost optimization, not to another AWS Region for disaster recovery. Option B is wrong because EBS snapshot schedules are used for backing up Amazon EBS volumes attached to EC2 instances, not for S3 objects, and they do not provide cross-region replication for S3 data. Option D is wrong because CloudFront is a content delivery network (CDN) that caches data at edge locations for low-latency access, not a replication mechanism to copy data to another AWS Region for disaster recovery.

59
Multi-Selecteasy

A developer accidentally corrupts part of a production Amazon RDS database, and the issue is discovered 45 minutes later. The team needs to restore the database to the state immediately before the change. Which two actions should be part of the recovery plan? Select two.

Select 2 answers
A.Enable automated backups with a retention period that covers the recovery window.
B.Perform a point-in-time restore to a new database instance.
C.Convert the database to a single-AZ deployment for faster restores.
D.Delete the corrupted rows manually and continue without restoring.
E.Use a read replica as the only recovery source for all deletions.
AnswersA, B

AWS RDS automated backups are foundational for Point-in-Time Recovery (PITR), capturing daily snapshots and continuously archiving transaction logs. A sufficient retention period is crucial, as it dictates how far back in time a database can be restored. If the corruption occurred outside the defined retention window, the specific recovery point before the incident would be unavailable, rendering PITR ineffective for that event.

Why this answer

Automated backups must be enabled to allow point-in-time recovery (PITR) within the retention window. Since the corruption occurred 45 minutes ago, the retention period must cover at least that duration to restore to the state immediately before the change. Option B is correct because PITR restores the database to a specified time (down to the second) within the backup retention period, creating a new DB instance that reflects the state just before the corruption.

Exam trap

The trap here is that candidates may think a read replica can be used for point-in-time recovery, but it only provides read scaling and asynchronous replication, not a restore point before the corruption occurred.

60
MCQmedium

Your media processing pipeline writes original uploads to an S3 bucket and later generates derivative files. An operator accidentally deletes a subset of original uploads in production. You need to (1) restore the deleted objects with minimal data loss and (2) protect against both regional disasters and future operator mistakes. The company requires recovery even if objects are deleted and later overwritten. What is the most effective change to meet these requirements?

A.Enable S3 versioning on the bucket and configure cross-Region replication so previous versions are available after regional loss and accidental deletion.
B.Move all objects to S3 Glacier Instant Retrieval and apply a lifecycle policy to keep only the latest object copy.
C.Use S3 server-side encryption with KMS keys and rely on access logs to manually recover the deleted objects.
D.Enable S3 bucket policies that deny DeleteObject, but do not enable versioning or replication.
AnswerA

Enabling S3 versioning preserves every PUT as a new version, so an accidental DELETE only adds a delete marker and the original object can be restored instantly. Cross-Region replication (CRR) asynchronously copies every version — including previous ones — to a secondary Region, providing recoverability if the entire source Region becomes unavailable. Together, versioning addresses user error and CRR addresses regional loss, satisfying the full recovery requirement.

Why this answer

Enabling S3 Versioning preserves all object versions, including overwrites and deletions (which become delete markers), allowing you to restore deleted objects by removing the delete marker. Cross-Region Replication (CRR) replicates both current and previous versions to a secondary Region, protecting against regional disasters. Together, they ensure recovery even if objects are deleted and later overwritten, meeting all requirements.

Exam trap

The trap here is that candidates may think a bucket policy denying DeleteObject is sufficient to prevent data loss, but it does not protect against overwrites, authorized user mistakes, or regional disasters, and without versioning, deleted objects are permanently lost.

Why the other options are wrong

B

Glacier Instant Retrieval does not provide versioning, so deleted objects cannot be restored. Additionally, a lifecycle policy keeping only the latest copy would not protect against accidental deletion or overwrites, and Glacier does not offer cross-region replication for disaster recovery.

C

Access logs only record requests; they do not restore deleted objects. Manual recovery from logs is impractical and cannot restore objects that were overwritten, failing the requirement for recovery after overwrite.

D

Denying DeleteObject via bucket policy does not protect against regional disasters, and if an operator deletes objects, they are still permanently lost because versioning is not enabled. Overwritten objects are also unrecoverable.

61
MCQhard

A patient portal must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The design must avoid adding custom operational scripts.

A.Instance store volumes
B.Amazon EFS with mount targets in multiple Availability Zones
C.An EBS volume attached to all instances
D.S3 mounted as a POSIX file system without a file gateway
AnswerB

Amazon EFS is a fully managed, regional NFS-based file system that supports the POSIX standard, making it ideal for Linux workloads. By creating mount targets in multiple Availability Zones, all EC2 instances can concurrently read and write the same files with low latency and high availability. EFS automatically scales capacity and is designed to be accessed by thousands of instances simultaneously, satisfying the shared file storage requirement across AZs.

Why this answer

Amazon EFS provides a fully managed, scalable, and shared file system that can be mounted concurrently on multiple Linux EC2 instances across different Availability Zones. By creating mount targets in each AZ, the file system remains accessible even if one AZ fails, meeting the high availability requirement without custom scripts.

Exam trap

The trap here is that candidates may confuse EBS multi-attach (which is limited to specific instance types and a single AZ) with a true cross-AZ shared file system, or assume S3 with a FUSE mount provides POSIX compliance without operational overhead.

How to eliminate wrong answers

Option A is wrong because instance store volumes are ephemeral, tied to a single EC2 instance, and data is lost on instance stop or termination, making them unsuitable for shared, durable storage across AZs. Option C is wrong because a single EBS volume can only be attached to one EC2 instance at a time (except for multi-attach EBS, which is limited to specific instance types and not designed for cross-AZ shared file storage). Option D is wrong because S3 mounted as a POSIX file system (e.g., via s3fs) requires custom scripts and does not provide native POSIX consistency or locking, and using it without a file gateway introduces performance and reliability issues for shared file storage.

62
MCQmedium

A healthcare provider runs a patient-record API on Amazon EC2 instances behind an Application Load Balancer in one AWS Region. The API reads from an Amazon RDS for MySQL DB instance. The provider must be able to continue serving read traffic if the primary database instance fails, and must minimize the time the application is unavailable. Which change should a solutions architect make?

A.Store the database credentials in AWS Secrets Manager and rotate them automatically on a schedule.
B.Convert the DB instance to an Amazon RDS Multi-AZ DB instance deployment so a standby is maintained in another Availability Zone.
C.Create a read replica in the same Availability Zone and point read traffic to it using a custom application load-balancing layer.
D.Enable automated backups with a longer retention period and increase the DB instance class size.
AnswerB

Amazon RDS Multi-AZ maintains a synchronously replicated standby in a different Availability Zone and automatically fails over by promoting the standby and updating the DNS endpoint, so the application reconnects without operator action. This directly addresses database failure and minimizes application downtime.

Why this answer

Database high availability with minimal application downtime is achieved by running Amazon RDS Multi-AZ, which keeps a synchronously replicated standby in another Availability Zone and performs automatic failover to it. Because the failover updates the same endpoint, applications reconnect without manual reconfiguration.

Exam trap

The trap here is confusing backups and read replicas with automatic failover, when only a Multi-AZ standby provides it.

63
MCQeasy

An organization hosts the same public API in two AWS Regions. Normal traffic should go to the primary Region. If the primary endpoint becomes unhealthy, Route 53 should automatically route users to the secondary Region. What is the best Route 53 configuration approach?

A.Use simple routing with one record that contains both regions as weighted targets.
B.Use weighted routing and set the secondary Region weight to 0 until needed.
C.Use Route 53 failover routing with health checks that mark the primary as unhealthy and fail over to the secondary.
D.Use latency-based routing so requests go to the region with the lowest latency, regardless of health.
AnswerC

Failover routing is designed for active/passive disaster recovery. You configure a primary record and a secondary record, each associated with health checks. When the primary fails its health checks, Route 53 automatically resolves the name to the secondary target.

Why this answer

Route 53 failover routing is designed for active-passive configurations where traffic is directed to a primary resource unless a health check marks it as unhealthy, at which point all traffic automatically shifts to the secondary resource. This directly matches the requirement of routing normal traffic to the primary Region and failing over to the secondary Region only when the primary endpoint becomes unhealthy.

Exam trap

The trap here is that candidates often confuse weighted routing with failover routing, mistakenly thinking that setting a weight of 0 on the secondary is a valid way to keep it inactive until needed, but Route 53 does not automatically adjust weights based on health checks.

How to eliminate wrong answers

Option A is wrong because simple routing does not support health checks or automatic failover; it simply returns all IP addresses in a random order, which cannot enforce a primary-secondary failover pattern. Option B is wrong because setting the secondary Region weight to 0 would prevent any traffic from reaching it even during a failure, and manually changing weights defeats the purpose of automatic failover. Option D is wrong because latency-based routing selects the Region with the lowest latency for each user, which does not guarantee that the primary Region handles normal traffic and does not automatically fail over based on endpoint health.

64
MCQmedium

Based on the exhibit, the web application must remain available even if one Availability Zone fails. What is the best change to improve resilience with the least redesign?

A.Increase DesiredCapacity to 4 while keeping all instances in subnet-a1.
B.Add subnet-b1 in a different Availability Zone to the Auto Scaling group.
C.Replace the Application Load Balancer with a Network Load Balancer.
D.Enable EBS encryption on the launch template volumes.
AnswerB

This spreads EC2 instances across two Availability Zones, so the Auto Scaling group can continue serving traffic if one AZ becomes unavailable. Because the ALB is already deployed in both subnets, this is the smallest change that adds true zonal resilience to the compute tier.

Why this answer

Adding subnet-b1 in a different Availability Zone to the Auto Scaling group ensures that EC2 instances are launched across two Availability Zones. If one zone fails, the ALB can route traffic to healthy instances in the other zone, maintaining application availability. This change requires minimal redesign because it only modifies the Auto Scaling group's subnet configuration without altering the load balancer or compute architecture.

Exam trap

The trap here is that candidates may think increasing instance count or changing load balancer type improves resilience, but without multi-AZ distribution, a single AZ failure still causes a total outage.

Why the other options are wrong

A

Increasing DesiredCapacity to 4 in a single subnet (subnet-a1) does not add resilience across Availability Zones; all instances remain in one AZ, so a failure of that AZ still causes total outage.

C

Replacing the Application Load Balancer with a Network Load Balancer does not improve resilience across Availability Zones; it only changes the load balancer type, which operates at a different layer and does not address the single-AZ failure risk.

D

Enabling EBS encryption does not improve availability or resilience across Availability Zones; it only protects data at rest. The question requires resilience against an AZ failure, which encryption does not address.

65
MCQhard

A logistics firm runs an order-processing service that reads from an Amazon SQS queue and writes results to an Amazon DynamoDB table. During a marketing event, the consumer fleet scaled out aggressively and DynamoDB began returning ProvisionedThroughputExceededException errors, causing messages to be retried and some orders to be processed twice. The architects want to absorb traffic spikes without overprovisioning capacity and without duplicate processing. Which combination of changes should they make?

A.Add a global secondary index on the message ID and query it before each write to check whether the order has already been stored.
B.Switch the DynamoDB table to on-demand capacity mode and have the consumer use a conditional write with an idempotency key derived from the message ID.
C.Increase the table's provisioned write capacity units substantially and enable DynamoDB Streams so that duplicate items can be detected after the fact.
D.Move the table to a different AWS Region and enable DynamoDB Accelerator (DAX) in front of it to cache the writes during the spike.
AnswerB

On-demand capacity mode removes the need to forecast or preprovision read and write capacity, so the table scales with the traffic spike automatically. Using a conditional write keyed on a stable message identifier makes repeated processing of the same message a no-op, eliminating duplicates. Together these address both the throttling and the duplicate-order symptoms described.

Why this answer

The symptom set has two distinct causes: a capacity model that cannot absorb spikes and a consumer that is not idempotent. On-demand capacity mode lets the table follow the workload without manual provisioning, while a conditional write on a stable idempotency key makes repeated delivery harmless. Addressing both together is what stops the throttling and the duplicate orders.

Exam trap

The trap here is treating throttling and duplicate processing as one problem, when they require two independent fixes.

66
MCQhard

A warehouse integration service must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable?

A.Use CloudFront signed URLs
B.Use Amazon SQS standard queue and design consumers to be idempotent
C.Use UDP messages sent directly to workers
D.Use an in-memory queue on one EC2 instance
AnswerB

Amazon SQS standard queues provide at-least-once delivery with high throughput, meaning every message is delivered but occasional duplicates can occur. Designing consumers to be idempotent ensures that processing the same event multiple times yields the same result, which satisfies the requirement to process every event. This is the recommended pattern for reliable, scalable event processing in AWS.

Why this answer

Amazon SQS standard queues provide at-least-once delivery, meaning each message is delivered at least once but can occasionally be delivered more than once. This matches the requirement to process every event at least once, and since duplicate processing is acceptable when consumers are idempotent, the standard queue is the most suitable and cost-effective choice. SQS also decouples the warehouse integration service from its consumers, improving resilience and scalability.

Exam trap

The trap here is that candidates may confuse 'at-least-once' with 'exactly-once' and incorrectly choose FIFO queues or other options, but the question explicitly accepts duplicates if idempotency is handled, making the standard queue the correct and simpler choice.

How to eliminate wrong answers

Option A is wrong because CloudFront signed URLs are used to control access to content delivered via CloudFront, not for event processing or message queuing; they provide no delivery guarantee mechanism. Option C is wrong because UDP is a connectionless, unreliable transport protocol that does not guarantee message delivery, order, or duplicate prevention, making it unsuitable for at-least-once processing. Option D is wrong because an in-memory queue on a single EC2 instance creates a single point of failure and lacks durability; if the instance fails, all queued events are lost, violating the requirement to process every event at least once.

67
MCQeasy

A team needs a relational database solution that can automatically fail over to a standby instance if the primary database becomes unavailable. They want the standby to be located in a different Availability Zone. Which RDS/Aurora configuration best satisfies this requirement?

A.Single-AZ DB deployment and rely on manual snapshot restore during failures.
B.Multi-AZ deployment with an automatically managed standby in a different Availability Zone and automatic failover.
C.Enable read replicas only, and promote a replica manually when the primary fails.
D.Enable point-in-time recovery (PITR) without configuring any Multi-AZ standby.
AnswerB

In a Multi-AZ deployment, Amazon RDS or Aurora provisions a physically separate standby instance in a different Availability Zone and synchronously replicates data from the primary to that standby. When the primary fails or the AZ is degraded, RDS automatically initiates failover to the standby, which becomes the new primary; because the DNS endpoint remains unchanged, the application continues to work without manual intervention. This provides both the required cross-AZ redundancy and automatic failover, meeting the availability requirement directly.

Why this answer

A Multi-AZ RDS deployment automatically provisions and maintains a standby instance in a different Availability Zone, and the failover is handled automatically by AWS without manual intervention. This meets the requirement for automatic failover to a standby in a different AZ, which is the core purpose of Multi-AZ deployments.

Exam trap

The trap here is that candidates often confuse read replicas with Multi-AZ standby, thinking that promoting a read replica provides automatic failover, but read replicas require manual promotion and do not serve as a synchronous standby.

How to eliminate wrong answers

Option A is wrong because a Single-AZ deployment has no standby instance, and manual snapshot restore requires significant downtime and manual steps, failing the automatic failover requirement. Option C is wrong because read replicas are designed for read scaling, not automatic failover; promoting a read replica manually introduces downtime and does not provide automatic failover to a standby. Option D is wrong because point-in-time recovery (PITR) only enables restoring to a specific time from backups, not automatic failover to a standby instance in a different AZ.

68
MCQhard

A healthcare analytics platform processes streaming records with an AWS Lambda function that writes results to an Amazon DynamoDB table. The pipeline must not lose records if the function throws an error, and the operations team wants to inspect and reprocess failed records without writing custom retry code. Which approach should the solutions architect use?

A.Write a wrapper inside the function that catches exceptions and re-sends the batch to the stream before returning success
B.Configure the event source mapping with a maximum retry count and a destination on failure set to an Amazon SQS queue configured as a dead-letter queue
C.Set the function's timeout to the maximum value and rely on Lambda's built-in retry of the entire batch until it succeeds
D.Increase the Lambda function's reserved concurrency so that retries happen faster and failures are less likely
AnswerB

For Lambda event source mappings, the maximum retry count controls how many times a failing batch is retried, and the on-failure destination sends the batch metadata to an SQS queue or SNS topic after retries are exhausted. That preserves failed records for later inspection and reprocessing without custom code. This directly meets both the no-loss and no-custom-retry requirements.

Why this answer

Lambda event source mappings support a maximum retry count plus an on-failure destination that captures the failed batch after retries are exhausted. Pointing that destination at an SQS queue gives the team a durable holding area where failed records can be examined and replayed, satisfying both no-loss and no-custom-code goals. The other options either only tune performance, rely on nonexistent indefinite retries, or reintroduce the custom logic the team wants to eliminate.

Exam trap

The trap here is believing Lambda retries a failing batch indefinitely; event source mappings retry a limited number of times, and without an on-failure destination the batch is discarded once retries are exhausted.

69
MCQmedium

A financial services firm runs a critical API on Amazon EC2 instances behind a Network Load Balancer. The API must handle a sudden loss of one Availability Zone and continue serving traffic with no manual failover. The instances are in an Auto Scaling group that currently uses a single subnet in one Availability Zone. Which change should the architect make?

A.Replace the Network Load Balancer with an Application Load Balancer and enable sticky sessions with a long cookie duration.
B.Create an Amazon Route 53 latency-based routing record that points to the Network Load Balancer and set a failover routing policy with a health check.
C.Enable cross-zone load balancing on the Network Load Balancer and increase the Auto Scaling group desired capacity to four instances in the existing subnet.
D.Recreate the Auto Scaling group with subnets in at least two Availability Zones, enable the Network Load Balancer across those subnets, and attach a target group with health checks.
AnswerD

A Network Load Balancer requires subnets in each Availability Zone where it should accept traffic, and an Auto Scaling group spanning multiple AZs keeps instances running if one AZ fails. Health checks remove failed targets, so the API continues serving with no manual intervention, which is exactly the resilience the firm needs.

Why this answer

The Network Load Balancer must be enabled in subnets across multiple Availability Zones, and the Auto Scaling group must launch instances in those same subnets. With health checks on the target group, failed instances and the failed AZ are removed from rotation automatically, so the API remains available without manual failover.

Exam trap

The trap here is believing that cross-zone load balancing or a larger desired capacity creates Availability Zone redundancy, when both still depend on having subnets and instances in more than one AZ.

70
MCQmedium

A trading dashboard runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include? The architecture review board prefers a managed AWS-native control.

A.A single EC2 instance with detailed monitoring
B.Subnets in at least two Availability Zones with health checks enabled
C.All instances in one larger subnet
D.A Network Load Balancer in one subnet
AnswerB

Placing the Auto Scaling group's subnets in at least two Availability Zones ensures that if one AZ becomes unavailable, the remaining AZs still have healthy instances to serve traffic. Health checks (either EC2 status checks or Elastic Load Balancing target health checks) allow the ASG to detect failed instances and replace them while maintaining the desired capacity. This architecture survives both individual instance crashes and full AZ outages, providing high availability for the trading dashboard.

Why this answer

Distributing EC2 instances across at least two Availability Zones (AZs) ensures that the application remains available if one AZ fails. The Auto Scaling group must include subnets in multiple AZs and use health checks (e.g., ELB health checks) to automatically replace unhealthy instances. This configuration meets the requirement for fault tolerance and aligns with AWS-managed best practices for high availability.

Exam trap

The trap here is that candidates often confuse 'scaling' with 'resilience' and think that a single large subnet or a different load balancer type (NLB) provides AZ fault tolerance, but only multi-AZ subnet configuration with health checks ensures automatic recovery from an AZ failure.

How to eliminate wrong answers

Option A is wrong because a single EC2 instance, even with detailed monitoring, cannot tolerate the failure of an Availability Zone; it represents a single point of failure. Option C is wrong because placing all instances in one larger subnet confines them to a single Availability Zone, which does not provide AZ-level fault tolerance. Option D is wrong because a Network Load Balancer in one subnet does not address the need for multi-AZ instance distribution; it also lacks the health-check-based auto-scaling capabilities required for instance replacement.

71
MCQmedium

An event-driven order processing service consumes messages from an Amazon SQS Standard queue. After a deployment, about 1% of messages start failing validation because a required field is missing. The consumer catches the exception and returns control, so the messages are retried. However, those poison messages keep reappearing and repeatedly consuming processing time for hours, delaying handling of valid messages. What is the most resilient way to handle the poison messages while keeping the system available?

A.Set the consumer visibility timeout to a very large value so failing messages are hidden for hours.
B.Configure an SQS redrive policy to send messages to a dead-letter queue (DLQ) after a limited number of receives (maxReceiveCount).
C.Switch the SQS queue from Standard to FIFO so poison messages do not retry.
D.Increase the consumer concurrency indefinitely so the system processes all messages even if some fail validation.
AnswerB

A DLQ redrive policy creates a deterministic stop condition for poison messages. After maxReceiveCount, the messages are moved to the DLQ instead of cycling in the main queue, preventing repeated failed deliveries from degrading capacity and availability for valid messages.

Why this answer

Configuring an SQS redrive policy with a maxReceiveCount (e.g., 3–5) automatically moves messages that repeatedly fail processing to a dead-letter queue (DLQ) after the specified number of receives. This isolates the poison messages, preventing them from consuming visibility timeout and processing resources, while allowing valid messages to be handled without delay. The DLQ can then be analyzed or reprocessed offline, maintaining system availability.

Exam trap

The trap here is that candidates may think increasing visibility timeout or concurrency solves the problem, but they fail to recognize that only a dead-letter queue permanently isolates poison messages from the processing pipeline.

How to eliminate wrong answers

Option A is wrong because setting the consumer visibility timeout to a very large value would hide failing messages for hours, but they would still reappear after the timeout expires, continuing the cycle of retries and delays without resolving the issue. Option C is wrong because switching from Standard to FIFO does not prevent poison messages from retrying; FIFO queues still retry messages on failure and require a DLQ for poison handling, and they also sacrifice throughput and ordering flexibility. Option D is wrong because increasing consumer concurrency indefinitely does not address the root cause—poison messages will still be retried and consume processing slots, potentially overwhelming the system and delaying valid messages further.

72
MCQhard

A claims workflow uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure? The architecture review board prefers a managed AWS-native control.

A.A FIFO queue without a redrive policy
B.Short polling instead of long polling
C.A dead-letter queue with an appropriate maxReceiveCount
D.A larger message retention period only
AnswerC

A dead-letter queue with maxReceiveCount moves messages aside once they exceed the receive threshold, so poison messages stop blocking useful retries. This satisfies the review board's managed AWS-native constraint, since DLQs are a native SQS feature requiring no custom code or third-party tooling.

Why this answer

A dead-letter queue (DLQ) with an appropriate maxReceiveCount allows messages that repeatedly fail processing to be moved out of the source queue after a specified number of receive attempts. This prevents poison messages from blocking useful retries and is a fully managed AWS-native pattern. The architecture review board's preference for a managed solution is satisfied because SQS DLQs are a built-in feature requiring no custom code.

Exam trap

The trap here is that candidates may confuse a DLQ with simply increasing retention or changing polling behavior, not realizing that poison messages require explicit isolation via a separate queue and a maxReceiveCount threshold to stop infinite retries.

How to eliminate wrong answers

Option A is wrong because a FIFO queue without a redrive policy does not automatically handle poison messages; without a DLQ, failed messages remain in the queue and continue to block retries. Option B is wrong because short polling reduces latency but does not address poison messages; it returns only a subset of servers' messages and can increase empty responses, but it has no effect on message failure handling. Option D is wrong because increasing the message retention period only keeps messages longer without removing failing ones; poison messages would still be retried until they expire, continuing to block useful retries.

73
Multi-Selectmedium

An application uses an Amazon RDS Multi-AZ DB instance. During a failover test, connections fail until the application is restarted, even though the database comes back online. Which two changes should the team make to improve resilience during failover? Select two.

Select 2 answers
A.Cache and reconnect to the current writer IP address to avoid DNS lookups during failover.
B.Use the RDS endpoint name instead of hard-coding the current instance IP or hostname in the application.
C.Switch to a read replica and let it promote manually after every outage.
D.Add retry logic with exponential backoff for transient connection and DNS resolution errors.
E.Disable connection pooling so each request opens a fresh socket during normal operation.
AnswersB, D

The RDS endpoint abstracts the underlying writer instance. When failover occurs, AWS updates the endpoint to point at the new writer, so the application should reconnect by using the managed name rather than a fixed IP or hostname.

Why this answer

The RDS endpoint is a DNS name that automatically resolves to the current writer instance's IP address. During a failover, the DNS record is updated to point to the new primary, so using the endpoint instead of a hard-coded IP or hostname allows the application to reconnect without manual intervention. Option D is correct because adding retry logic with exponential backoff handles transient failures during DNS resolution and connection establishment, which are common during the brief period when the DNS TTL has not yet expired after a failover.

Exam trap

The trap here is that candidates often think caching the IP (Option A) improves performance, but it actually breaks failover resilience because the application never learns the new writer's address after a failover.

74
MCQmedium

An Auto Scaling group behind an Application Load Balancer frequently replaces new EC2 instances. The application needs ~6 minutes to warm up after instance launch. However, the ALB target group health checks start immediately and mark the targets unhealthy until the application is ready. Because the targets become unhealthy early, the Auto Scaling group then terminates the instances and launches replacements, creating a repeated unhealthy/termination loop. What configuration change will most directly improve recovery by preventing premature ASG termination while the application is warming up?

A.Set a health check grace period on the Auto Scaling group that exceeds the application startup/warm-up time.
B.Increase the Auto Scaling group's desired capacity to a higher number than required.
C.Disable ALB target group health checks so instances are considered healthy as soon as they register.
D.Change the Auto Scaling health check type from ELB to EC2 so the ALB will no longer determine instance health.
AnswerA

A health check grace period delays when the Auto Scaling group starts evaluating instance health. This prevents the ASG from terminating instances due to ALB/target health being unhealthy during the initial warm-up window, breaking the unhealthy/termination loop.

Why this answer

The health check grace period on an Auto Scaling group (ASG) allows a newly launched EC2 instance to bypass health check failures for a specified duration. By setting this grace period to exceed the application's ~6-minute warm-up time, the ASG will not prematurely terminate the instance based on ALB health check results. This directly breaks the unhealthy/termination loop while the application initializes.

Exam trap

The trap here is that candidates may think disabling health checks or changing the health check type is a valid fix, but the correct solution is to use the ASG's built-in grace period to decouple early health check failures from termination decisions.

Why the other options are wrong

B

Increasing desired capacity does not prevent the Auto Scaling group from terminating instances that fail health checks; it only adds more instances, which may also fail and be terminated, perpetuating the loop.

C

Disabling ALB target group health checks would prevent the ALB from routing traffic to healthy instances, causing service disruption. The issue is premature termination by ASG, not health check failure; the grace period directly addresses this.

D

Changing the health check type to EC2 would make the Auto Scaling group ignore ALB health check results, but the ALB would still route traffic to unhealthy instances, causing application errors. The question requires preventing premature termination during warm-up, not ignoring health checks entirely.

75
MCQmedium

A company runs an internet-facing API in two AWS Regions. Route 53 currently uses simple routing to a primary Application Load Balancer (ALB) DNS name. When the primary Region experiences an outage, customers wait a long time because the DNS entry is not changed automatically. The team wants automatic failover: if the primary Region ALB health check fails for a sustained period, Route 53 should route users to the secondary Region ALB. Which Route 53 approach best meets this requirement?

A.Use Route 53 failover routing with a PRIMARY and SECONDARY record set for the same name, and attach health checks to the ALBs.
B.Use latency-based routing so Route 53 automatically spreads traffic to both Regions based on measured latency.
C.Use weighted routing and configure the secondary ALB to receive 100% traffic when the primary returns HTTP 5xx responses.
D.Use geolocation routing and restrict the primary Region record to specific countries only.
AnswerA

Route 53 failover routing is specifically designed for active-passive DNS failover. You create two records with the same name, designate one as PRIMARY and one as SECONDARY, and attach a Route 53 health check to each ALB endpoint. Route 53 continuously evaluates the PRIMARY health check; when it fails for the configured evaluation period, Route 53 responds with the SECONDARY record's IP or alias. This gives a deterministic, health-driven failover where the healthy secondary ALB starts receiving traffic once the primary is marked unhealthy, respecting the record TTL for propagation.

Why this answer

Route 53 failover routing is designed specifically for active-passive failover scenarios. By creating PRIMARY and SECONDARY record sets with the same DNS name and attaching health checks to the ALBs, Route 53 will automatically route traffic to the secondary ALB when the primary ALB health check fails for a sustained period. This meets the requirement for automatic failover without manual intervention.

Exam trap

The trap here is that candidates often confuse failover routing with latency-based or weighted routing, assuming that latency-based routing inherently provides failover, but it does not—it only optimizes for performance, not availability.

Why the other options are wrong

B

Latency-based routing distributes traffic based on lowest latency, not health. It does not provide automatic failover when a region is completely down; users may still be routed to the unhealthy primary if it has lower latency.

C

Weighted routing distributes traffic based on weights, not health. It cannot automatically shift 100% traffic to the secondary ALB based on HTTP 5xx responses; health checks are not integrated with weighted routing for automatic failover.

D

Geolocation routing directs traffic based on the geographic location of the user, not on the health or availability of the endpoint. It cannot automatically failover to a secondary Region when the primary ALB becomes unhealthy.

Page 1 of 4 · 257 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Design Resilient questions.