Courseiva

CCNA Design Resilient Questions

75 of 257 questions · Page 2/4 · Design Resilient topic · Answers revealed

76
MCQmedium

A healthcare company needs to store patient records in Amazon DynamoDB. The records must be highly available and durable across multiple Availability Zones. The company also requires the ability to recover the table to any point in time within the last 35 days in case of accidental writes or deletions. Which solution meets these requirements?

A.Create a DynamoDB table and enable point-in-time recovery (PITR).
B.Create a DynamoDB table with a read replica in another AWS Region and enable point-in-time recovery (PITR).
C.Create a DynamoDB table with a global secondary index and enable point-in-time recovery (PITR).
D.Create a DynamoDB table with on-demand capacity mode and enable point-in-time recovery (PITR).
AnswerA

DynamoDB automatically replicates data across multiple Availability Zones within an AWS Region, providing high availability and durability. Enabling point-in-time recovery (PITR) allows restoration to any point in time within the last 35 days, protecting against accidental writes or deletions. This solution meets both the durability and recovery requirements without additional complexity.

Why this answer

DynamoDB tables are automatically replicated across multiple Availability Zones within a Region, ensuring high availability and durability. Enabling point-in-time recovery (PITR) provides continuous backups and allows restoration to any point in time within the last 35 days. This combination meets the requirements without additional configuration such as global tables or secondary indexes.

Exam trap

The trap here is thinking that additional features like global secondary indexes or cross-Region replicas are needed for multi-AZ durability, when DynamoDB already provides it by default.

77
MCQmedium

A company stores critical documents in an Amazon S3 bucket in the us-east-1 Region. The documents must survive an unlikely loss of the entire us-east-1 Region. The company wants a recovery point objective (RPO) of 15 minutes and a recovery time objective (RTO) of 1 hour. What should the solutions architect recommend?

A.Enable S3 server access logging and store the logs in a separate bucket in us-east-1.
B.Configure an S3 Lifecycle rule to transition objects to S3 Glacier Deep Archive after 30 days.
C.Enable S3 Transfer Acceleration on the us-east-1 bucket to speed up uploads from global clients.
D.Enable S3 Cross-Region Replication (CRR) from the us-east-1 bucket to a bucket in us-west-2 and enable S3 Versioning on both buckets.
AnswerD

S3 Cross-Region Replication asynchronously copies objects to a bucket in another Region, and versioning is required for replication. With typical replication times well under 15 minutes, this meets the RPO and provides a target bucket that can be used within the 1-hour RTO.

Why this answer

S3 Cross-Region Replication copies objects to a bucket in a different Region, and versioning is a prerequisite. Replication is asynchronous but typically completes well within a 15-minute RPO, and the destination bucket can serve as the recovery source within the 1-hour RTO.

Exam trap

The trap here is confusing performance features such as Transfer Acceleration or storage-class transitions with actual cross-Region data replication.

78
MCQmedium

A healthcare company runs a patient portal on Amazon EC2 instances behind an Application Load Balancer across two Availability Zones. A new compliance rule requires that if an entire Availability Zone fails, the portal must remain available with no manual intervention. The EC2 instances are stateless and store no session data. Which design change should the architect implement to meet this requirement?

A.Enable an Auto Scaling group that spans both Availability Zones and configure the load balancer health checks to deregister unhealthy instances.
B.Convert the EC2 instances to a single larger instance type and enable detailed monitoring in Amazon CloudWatch.
C.Configure the Application Load Balancer to use cross-zone load balancing and attach an Elastic IP address to each EC2 instance.
D.Place the EC2 instances in a single Availability Zone and create an Amazon EBS snapshot schedule every hour.
AnswerA

An Auto Scaling group spanning both Availability Zones automatically replaces instances in the surviving AZ when one AZ fails, and ALB health checks remove failed targets. This provides resilience without manual intervention, directly satisfying the compliance requirement for AZ failure tolerance.

Why this answer

High availability across Availability Zones requires compute capacity that can be automatically replaced when an AZ fails. An Auto Scaling group spanning multiple AZs, combined with load balancer health checks, ensures that unhealthy instances are removed and new instances are launched in a surviving AZ without human action.

Exam trap

The trap here is assuming that cross-zone load balancing or a larger instance type provides AZ-level resilience, when only an Auto Scaling group spanning multiple AZs can automatically replace lost capacity.

79
MCQhard

A patient portal must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable? The team wants the control to be enforceable during normal operations.

A.Use an in-memory queue on one EC2 instance
B.Use UDP messages sent directly to workers
C.Use Amazon SQS standard queue and design consumers to be idempotent
D.Use CloudFront signed URLs
AnswerC

Amazon SQS standard queues guarantee at-least-once delivery, so no event is lost, and duplicates are possible; consumers must therefore be idempotent. This satisfies the stated requirement that duplicate processing is acceptable if consumers handle idempotency.

Why this answer

Amazon SQS standard queues provide at-least-once delivery, ensuring every event is processed at least once, which matches the requirement. Duplicate processing is acceptable because the team can design consumers to be idempotent, handling duplicates without side effects. SQS is a fully managed, scalable, and durable service that enforces this behavior during normal operations without requiring custom infrastructure.

Exam trap

The trap here is that candidates may confuse 'at-least-once' delivery with 'exactly-once' delivery, or incorrectly assume that UDP or in-memory queues can provide reliable event processing, when in fact only a managed queue service like SQS with idempotent consumers meets the stated requirement for enforceability during normal operations.

How to eliminate wrong answers

Option A is wrong because an in-memory queue on a single EC2 instance is not durable, cannot survive instance failures, and does not provide at-least-once delivery guarantees across restarts or scaling events. Option B is wrong because UDP is a connectionless, unreliable protocol that does not guarantee message delivery, order, or duplicate detection, making it unsuitable for at-least-once processing. Option D is wrong because CloudFront signed URLs are used for access control to content delivery, not for event processing or messaging, and they do not provide any delivery guarantee or queue semantics.

80
MCQmedium

A team accidentally updates critical rows in an Amazon RDS for PostgreSQL database. Automated backups are enabled. They need to recover the data to the exact state as of 90 minutes ago. They also cannot risk interrupting the current production database instance while investigators validate the restored data. Which recovery strategy best meets these constraints?

A.Use point-in-time recovery (PITR) to restore to a new RDS DB instance as of 90 minutes ago, then validate and cut over after approval.
B.Restore a manual snapshot and overwrite the existing production DB instance so the data matches exactly 90 minutes ago.
C.Wait for the next automated backup window and then restart the current DB instance to roll back changes automatically.
D.Use cross-region read replicas to rewind changes and promote the replica to become the writer immediately.
AnswerA

Amazon RDS point-in-time recovery (PITR) restores a new DB instance to any fraction of a second within the backup retention window by replaying transaction logs from the last automated snapshot. Launching a separate instance preserves the production database untouched, allowing you to validate the restored data at 90 minutes ago before promoting it and updating the application connection string. After approval, you can cut over by renaming the instances or changing the DNS endpoint, minimizing downtime and risk.

Why this answer

Point-in-time recovery (PITR) for Amazon RDS allows you to restore a DB instance to any second within the backup retention period, using automated backups and transaction logs. By restoring to a new RDS instance as of 90 minutes ago, you create an isolated copy for validation without affecting the production database. This meets both the recovery point objective (RPO) of 90 minutes and the constraint of no interruption to the current production instance.

Exam trap

The trap here is that candidates confuse point-in-time recovery with snapshot restoration or assume that read replicas can be used for time-based rollbacks, but only PITR provides the exact time-targeted restore without affecting the production instance.

Why the other options are wrong

B

Restoring a manual snapshot and overwriting the production DB instance would cause downtime and data loss, as it replaces the current database entirely, violating the constraint of not interrupting production while validating.

C

Waiting for the next automated backup window does not allow recovery to a specific point 90 minutes ago; automated backups are typically taken once per day and do not support rollback to an arbitrary time.

D

Cross-region read replicas do not support rewinding changes; they replicate data asynchronously and cannot restore to a specific past point in time. Promoting a replica does not roll back the database to a previous state.

81
Drag & Dropmedium

Order the steps to create a static website using Amazon S3 and CloudFront.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

S3 bucket with hosting, upload files, CloudFront distribution, configure CloudFront, then DNS.

82
Multi-Selectmedium

A logistics company runs a stateless order-tracking API on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer. The architect must ensure the API survives the loss of an entire Availability Zone and that unhealthy instances are replaced automatically. (Choose two.)

Select 2 answers
A.Enable termination protection on all EC2 instances so the Auto Scaling group cannot remove them during a zone failure
B.Create a second Auto Scaling group in a different Region and use Amazon Route 53 latency-based routing to distribute traffic
C.Attach an Application Load Balancer target group health check and enable ELB health checks on the Auto Scaling group so unhealthy instances are terminated and replaced
D.Place the instances in a single Availability Zone and enable detailed CloudWatch monitoring with a 1-minute granularity
E.Configure the Auto Scaling group to span at least two Availability Zones and set the desired capacity to a number that keeps instances running in each zone
AnswersC, E

Enabling Elastic Load Balancing health checks on the Auto Scaling group means the group uses the load balancer's health status rather than only EC2 status checks, so instances that fail application-level checks are replaced. This ensures automatic recovery from unhealthy instances, which is the second requirement. The target group health check defines what the load balancer considers healthy.

Why this answer

Zone-level resilience with EC2 Auto Scaling requires the group to span multiple Availability Zones with enough desired capacity to run in each, so a single zone failure leaves serving capacity. Automatic replacement of unhealthy instances requires ELB health checks on the group, which makes the group act on the load balancer's health determination rather than only EC2 status checks. Together these two settings deliver the required survival and self-healing behaviour.

Exam trap

The trap here is assuming that enabling monitoring or termination protection contributes to resilience, when only multi-AZ capacity and ELB health checks do.

83
MCQhard

A company runs a critical API on Amazon EC2 behind an Application Load Balancer in a single AWS Region. The business requires the API to keep serving traffic if an entire Availability Zone becomes unavailable, and the recovery must not depend on any manual step. The database is Amazon RDS for PostgreSQL configured as a Single-AZ instance. Which combination of changes should a solutions architect implement to meet these requirements?

A.Configure the Auto Scaling group to span at least two Availability Zones with a health check grace period, and convert the RDS instance to a Multi-AZ DB instance deployment.
B.Enable Multi-AZ on the Application Load Balancer by adding a second listener in a different Availability Zone.
C.Take hourly RDS snapshots and configure an Auto Scaling lifecycle hook to restore the snapshot into a new Availability Zone during a failure.
D.Create an RDS read replica in a second Availability Zone and update the application to write to the replica when the primary fails.
AnswerA

Spreading the Auto Scaling group across multiple Availability Zones lets the load balancer route to healthy instances when one zone fails, and Multi-AZ RDS maintains a synchronous standby in another zone with automatic failover to the same endpoint. Together these provide automatic, hands-off recovery for both compute and database tiers.

Why this answer

Automatic recovery across an Availability Zone failure needs redundant compute and a database that fails over on its own. A multi-AZ Auto Scaling group keeps the API serving through the load balancer, and Multi-AZ RDS maintains a synchronous standby that is promoted automatically to the same endpoint, so no human action is required.

Exam trap

The trap here is reaching for an RDS read replica for high availability, when replicas are asynchronous and read-only and require manual promotion.

84
MCQhard

A SaaS provider runs a multi-tenant application on Amazon EC2 instances behind an Application Load Balancer. Tenants are identified by a subdomain, and each tenant's data is stored in a separate Amazon S3 bucket. The provider wants HTTPS with a single certificate, automatic renewal, and the ability to add new tenant subdomains without redeploying or replacing the certificate. Which solution meets these requirements?

A.Request a public certificate in AWS Certificate Manager for the apex domain and a wildcard for its subdomains, validate it with DNS, and attach it to the Application Load Balancer HTTPS listener.
B.Import a self-signed certificate covering all current tenant subdomains into AWS Certificate Manager and attach it to the Application Load Balancer listener.
C.Terminate TLS on the EC2 instances using certificates issued by AWS Private Certificate Authority, and pass traffic from the Application Load Balancer to the instances over HTTP.
D.Store the private key and certificate in AWS Secrets Manager, and configure the Application Load Balancer to retrieve and rotate the certificate at each renewal.
AnswerA

A public ACM certificate that includes the apex domain and a wildcard covers all current and future tenant subdomains, and DNS validation allows ACM to renew the certificate automatically as long as the validation records remain in place. Attaching it to the Application Load Balancer HTTPS listener provides TLS termination without redeploying the application when tenants are added.

Why this answer

A public ACM certificate that includes the apex domain and a wildcard for its subdomains covers every current and future tenant hostname. DNS validation lets ACM renew the certificate automatically while the validation records persist, and attaching the certificate to the Application Load Balancer listener provides HTTPS termination without touching the application when new tenants are onboarded.

Exam trap

The trap here is importing a certificate that covers only today's subdomains, when a wildcard certificate requested through ACM with DNS validation covers future tenants and renews automatically.

85
MCQeasy

Based on the exhibit, the database must fail over automatically if the primary Availability Zone goes down. Which solution should the architect choose?

A.Create a read replica in the same Availability Zone as the primary database.
B.Convert the database to a Multi-AZ RDS deployment.
C.Increase the backup retention period to 35 days.
D.Move the database to an EC2 instance with an attached EBS volume.
AnswerB

A Multi-AZ RDS deployment keeps a synchronous standby in another Availability Zone and automatically fails over when the primary fails. This matches the requirement for minimal manual intervention and preserves the same database endpoint, so the application does not need connection string changes. It is the standard AWS choice for resilient relational databases.

Why this answer

A Multi-AZ RDS deployment automatically provisions and maintains a synchronous standby replica in a different Availability Zone. If the primary AZ fails, Amazon RDS automatically fails over to the standby, typically within 60–120 seconds, without requiring manual intervention or changes to the application connection string.

Exam trap

The trap here is that candidates confuse read replicas (which are asynchronous and require manual promotion) with Multi-AZ deployments (which provide automatic synchronous failover), often selecting a read replica in the same AZ because they think it offers high availability without understanding the fundamental replication mode difference.

How to eliminate wrong answers

Option A is wrong because a read replica in the same AZ does not provide automatic failover; it is designed for read scaling, not high availability, and requires manual promotion. Option C is wrong because increasing the backup retention period to 35 days only affects point-in-time recovery and automated backups, not failover capability. Option D is wrong because moving the database to an EC2 instance with an attached EBS volume requires custom scripting or third-party tools to implement automatic failover, and EBS volumes are AZ-specific, so they cannot survive an AZ outage without manual intervention.

86
MCQmedium

A ticket booking system uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered?

A.Aurora Global Database
B.A single-AZ Aurora cluster
C.An ElastiCache Redis replica
D.Manual snapshots copied monthly
AnswerA

Aurora Global Database is the correct choice because it uses storage-based replication to synchronize data across multiple AWS Regions with a typical latency of under one second. It supports up to five secondary Regions, and when a regional failure occurs, you can promote a secondary Region to primary in about a minute, giving you a very low RPO and RTO for fast disaster recovery. This managed feature also enables low-latency global reads and automatic failover, making it vastly superior to any snapshot-based approach.

Why this answer

Aurora Global Database is designed for cross-Region disaster recovery with a typical RPO of 1 second or less, using storage-based replication that does not impact database performance. This meets the requirement for fast failover and low data loss, unlike manual snapshot-based approaches which have higher RPO and slower recovery.

Exam trap

The trap here is that candidates may confuse cross-Region read replicas (which have higher lag and manual promotion) with Aurora Global Database, or assume that ElastiCache or single-AZ deployments can provide adequate DR, when only Aurora Global Database meets the low RPO and fast cross-Region recovery requirements.

How to eliminate wrong answers

Option B is wrong because a single-AZ Aurora cluster lacks any cross-Region replication or failover capability, providing no disaster recovery across Regions. Option C is wrong because ElastiCache Redis is an in-memory cache, not a persistent database, and cannot serve as the primary data store for ticket booking transactions or provide cross-Region DR for the Aurora MySQL data. Option D is wrong because manual snapshots copied monthly result in an RPO of up to one month, which is far too high for the low RPO requirement, and recovery would require provisioning a new cluster from the snapshot, leading to significant downtime.

87
MCQhard

A warehouse integration service must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used? The architecture review board prefers a managed AWS-native control.

A.Amazon EFS with mount targets in multiple Availability Zones
B.S3 mounted as a POSIX file system without a file gateway
C.Instance store volumes
D.An EBS volume attached to all instances
AnswerA

Amazon EFS is a regional, managed NFS file system that provides standard POSIX file semantics, including file locking and consistent reads/writes, making it ideal for concurrent access from many EC2 instances. By configuring mount targets in multiple Availability Zones, the service achieves high availability and fault tolerance while instances in any AZ can access the same shared file data.

Why this answer

Amazon EFS provides a fully managed, NFS-based shared file system that can be mounted concurrently by multiple Linux EC2 instances across different Availability Zones. By creating mount targets in each AZ, the file system remains accessible even if one AZ fails, as traffic is automatically routed to the surviving mount targets. This meets the requirement for shared, resilient storage with a managed AWS-native control plane.

Exam trap

The trap here is that candidates often confuse EBS Multi-Attach (which is limited to a single AZ and specific instance types) with the cross-AZ shared file system capability of EFS, or incorrectly assume S3 with a FUSE mount can replace a POSIX-compliant file system.

How to eliminate wrong answers

Option B is wrong because mounting S3 as a POSIX file system (e.g., using s3fs-fuse) does not provide true POSIX compliance, lacks strong consistency guarantees, and introduces performance and locking issues unsuitable for shared file workloads; it also requires a third-party tool, not a fully managed AWS-native service. Option C is wrong because instance store volumes are ephemeral, tied to the lifecycle of a single EC2 instance, and cannot be shared across instances or survive an AZ failure. Option D is wrong because an EBS volume can only be attached to a single EC2 instance at a time (unless using multi-attach, which is limited to specific EBS types and still not designed for cross-AZ shared file systems), and it cannot be simultaneously mounted by instances in multiple Availability Zones.

88
MCQeasy

A worker service consumes messages from an Amazon SQS queue. Some messages are malformed and always fail validation. The worker retries, but it keeps reprocessing the same bad messages and consumes processing capacity that should be used for valid work. What is the best solution to prevent “poison messages” from blocking progress?

A.Configure a Dead-Letter Queue (DLQ) and set a redrive policy so messages move to the DLQ after a maximum number of receives.
B.Increase the visibility timeout so the worker gets fewer retries per hour.
C.Disable SQS retries by deleting messages immediately on any processing error.
D.Create a second worker that polls the queue less frequently until the malformed message is processed successfully.
AnswerA

Configuring a Dead-Letter Queue (DLQ) with a redrive policy is the most effective solution. This mechanism automatically moves messages that fail processing a specified number of times (maxReceiveCount) from the source queue to the DLQ. This prevents 'poison pill' messages from continuously consuming worker resources and allows for their isolation, analysis, and eventual reprocessing or discarding without impacting the main message flow.

Why this answer

A Dead-Letter Queue (DLQ) with a redrive policy is the standard AWS mechanism for handling poison messages. By setting a maximum receive count (e.g., 5), the SQS queue automatically moves messages that fail processing repeatedly to the DLQ, isolating them from the main queue. This prevents the worker from wasting capacity on invalid messages and allows the main queue to continue processing valid work without interruption.

Exam trap

The trap here is that candidates may think increasing the visibility timeout or deleting messages on error is a valid solution, but AWS specifically designed the DLQ pattern to isolate poison messages without losing data or impacting throughput.

How to eliminate wrong answers

Option B is wrong because increasing the visibility timeout only delays the retry, it does not prevent the worker from eventually reprocessing the same bad message, so the poison message still consumes processing capacity. Option C is wrong because SQS does not support disabling retries; deleting messages immediately on error would lose the message entirely without any chance for recovery or analysis, which is not a best practice. Option D is wrong because creating a second worker that polls less frequently does not solve the problem—the malformed message will still be retried and block progress, and a slower poll rate only reduces throughput without addressing the root cause.

89
MCQmedium

A warehouse integration service receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers? The design must avoid adding custom operational scripts.

A.AWS WAF
B.Amazon Route 53 weighted routing
C.Amazon SQS queue
D.Amazon CloudFront
AnswerC

Amazon SQS is a fully managed message queue that decouples order-producing services from warehouse processing consumers. It durably stores messages, allowing bursts of orders to be buffered while consumers poll and process at a controlled rate, and it supports retries via visibility timeout and dead-letter queues for failed processing. This makes SQS the correct choice for absorbing spikes and ensuring no order is lost during high demand.

Why this answer

Amazon SQS is the correct choice because it acts as a durable, scalable message buffer that decouples the web tier from the fulfilment workers. When order bursts arrive, messages are stored reliably in the queue, and workers can poll at their own pace, retrying failed messages automatically without any custom scripts. This pattern absorbs spikes and ensures no requests are lost, meeting the requirement for a fully managed, serverless integration.

Exam trap

The trap here is that candidates often confuse load-balancing or traffic-routing services (like Route 53 or CloudFront) with message queuing, mistakenly thinking they can absorb processing spikes, whereas only a queue like SQS provides durable storage and asynchronous decoupling for request bursts.

How to eliminate wrong answers

Option A is wrong because AWS WAF is a web application firewall that filters HTTP/S traffic based on rules (e.g., SQL injection, XSS) and does not provide message buffering, decoupling, or retry capabilities for downstream services. Option B is wrong because Amazon Route 53 weighted routing distributes DNS traffic across multiple endpoints based on weights, but it operates at the DNS level and cannot absorb processing spikes or retry failed requests; it simply routes new connections. Option D is wrong because Amazon CloudFront is a content delivery network (CDN) that caches static and dynamic content at edge locations to reduce latency, but it does not offer message queuing, buffering, or retry logic for backend processing workloads.

90
MCQmedium

A media company stores video files in an Amazon S3 bucket in the us-east-1 Region. The company wants to ensure that the files are automatically replicated to us-west-2 for disaster recovery, and that replication occurs within 15 minutes of upload. Which solution meets these requirements with the LEAST operational overhead?

A.Configure an AWS Lambda function triggered by S3 event notifications to copy each object to the destination bucket.
B.Use AWS DataSync to schedule a recurring task that synchronizes the source bucket to the destination bucket every 15 minutes.
C.Enable S3 Cross-Region Replication (CRR) on the bucket with a replication rule that applies to all objects.
D.Enable S3 Same-Region Replication (SRR) to a bucket in us-east-1, then use S3 Batch Operations to copy objects to us-west-2 nightly.
AnswerC

S3 Cross-Region Replication automatically replicates objects to a bucket in another Region asynchronously, typically within minutes. It requires versioning on both buckets and an IAM role. Once configured, it operates with no additional infrastructure, providing low operational overhead and meeting the 15-minute replication goal for most objects.

Why this answer

S3 Cross-Region Replication is the managed solution for automatically replicating objects to a bucket in another Region. It is asynchronous but typically completes within minutes, satisfying the 15-minute requirement. It requires versioning and an IAM role, but after setup it runs without ongoing intervention.

This provides the lowest operational overhead compared to custom code or scheduled data transfer tasks.

Exam trap

The trap here is confusing S3 Same-Region Replication with Cross-Region Replication, or assuming that a scheduled copy job can meet a near-real-time replication objective.

91
MCQeasy

A company runs its customer-facing web app on EC2 behind an Application Load Balancer. The database is Amazon RDS for PostgreSQL. The requirement is that if a single Availability Zone fails, the database must automatically fail over within the same AWS Region with minimal application changes. Which database setup best meets this requirement?

A.Use an RDS single-AZ instance and periodically restore from automated backups if needed.
B.Deploy the RDS PostgreSQL instance as Multi-AZ with automatic failover enabled.
C.Create a read replica in a different AZ and use it only when the primary fails.
D.Use RDS with Multi-AZ disabled, but increase storage IOPS to prevent failover.
AnswerB

Multi-AZ maintains a synchronous standby replica in a second Availability Zone; on AZ failure, RDS automatically promotes it and repoints the DNS endpoint, typically within 60–120 seconds. Because the application reconnects via the same endpoint, no code changes are needed, satisfying the in-Region automatic failover and minimal-change constraints.

Why this answer

RDS Multi-AZ for PostgreSQL automatically provisions and maintains a synchronous standby replica in a different Availability Zone. If the primary AZ fails, Amazon RDS automatically fails over to the standby, typically within 60–120 seconds, with no changes required to the application's connection string (the DNS name remains the same). This meets the requirement for minimal application changes and automatic failover within the same Region.

Exam trap

The trap here is that candidates often confuse a read replica (which requires manual promotion and DNS changes) with a Multi-AZ standby (which provides automatic, transparent failover), leading them to incorrectly select Option C.

Why the other options are wrong

A

Single-AZ RDS with manual backup restoration does not provide automatic failover; it requires manual intervention and incurs significant downtime, failing the requirement for automatic failover within the same Region.

C

A read replica is not designed for automatic failover; promoting it requires manual intervention or additional scripting, which does not meet the 'automatically fail over' requirement with minimal application changes.

D

Multi-AZ disabled means no automatic failover; increasing IOPS improves performance but does not provide high availability across AZs, so it fails the requirement of automatic failover during an AZ failure.

92
Multi-Selectmedium

A company is deploying a stateless web application on Amazon ECS with Fargate. The application must be resilient to individual task failures and Availability Zone failures. Which three steps should the company take to achieve this resilience? (Choose three.)

Select 3 answers
.Configure the ECS service to use a spread placement strategy across Availability Zones.
.Set a minimum healthy percent of 50 and a maximum percent of 200 in the ECS service deployment configuration.
.Place all ECS tasks in a single subnet to minimize network latency.
.Use an Application Load Balancer (ALB) in front of the ECS service to distribute traffic across tasks.
.Store application session data in an attached EFS file system shared across all tasks.
.Disable automatic task replacement to avoid unnecessary task churn during failures.

Why this answer

Configuring the ECS service with a spread placement strategy across Availability Zones ensures tasks are distributed across multiple AZs, providing resilience against AZ failures. Setting a minimum healthy percent of 50 and a maximum percent of 200 allows the service to maintain at least half of the desired tasks during deployments or failures while scaling up to replace failed tasks without downtime. Using an Application Load Balancer (ALB) in front of the ECS service distributes incoming traffic across healthy tasks in different AZs, automatically rerouting traffic if a task or AZ fails.

Exam trap

The trap here is that candidates may confuse stateless applications with stateful ones and incorrectly choose to store session data in EFS, or they may think placing tasks in a single subnet improves performance without considering the single point of failure risk.

93
MCQeasy

Your company hosts an internal API in two AWS Regions. You want Amazon Route 53 to automatically send traffic to the secondary Region if the primary Region’s endpoint becomes unhealthy. Which Route 53 configuration best meets this requirement?

A.Latency-based routing with health checks for both Regions.
B.Failover routing with a primary record associated with a health check, and a secondary (failover) record associated with its own health check settings.
C.Weighted routing to distribute traffic evenly across both Regions.
D.Geolocation routing based on the client’s country to choose a Region.
AnswerB

Route 53 failover routing implements an active-passive pattern by defining a primary record (the desired region) paired with a health check that continuously verifies endpoint availability. When that health check returns an unhealthy status, Route 53 automatically removes the primary record from the response and answers DNS queries with the secondary (failover) record, which also has its own health check settings to ensure the backup is truly reachable. This gives you deterministic, health-driven failover where all traffic shifts to the secondary region after the primary's health check fails, exactly matching the requirement.

Why this answer

Failover routing in Route 53 is specifically designed for active-passive configurations where traffic is directed to a primary resource unless a health check indicates it is unhealthy, at which point traffic is automatically routed to the secondary (failover) record. By associating a health check with the primary record, Route 53 can monitor the endpoint's health and perform the failover seamlessly. This directly meets the requirement to send traffic to the secondary Region when the primary endpoint becomes unhealthy.

Exam trap

The trap here is that candidates often confuse failover routing with latency-based routing, assuming that latency-based routing with health checks will automatically redirect traffic to the next best Region when one is unhealthy, but in reality, latency-based routing only selects the lowest-latency healthy endpoint and does not enforce a strict primary-secondary failover order.

How to eliminate wrong answers

Option A is wrong because latency-based routing directs traffic to the Region with the lowest latency for the client, not based on health status; while health checks can be associated, they only mark records as unhealthy without automatically failing over to a specific secondary Region. Option C is wrong because weighted routing distributes traffic based on assigned weights, not health; if the primary endpoint is unhealthy, traffic would still be sent to it according to the weight, unless the record is marked unhealthy, but there is no automatic failover to a designated secondary. Option D is wrong because geolocation routing directs traffic based on the client's geographic location, not on endpoint health; it does not provide automatic failover to a secondary Region when the primary is unhealthy.

94
MCQmedium

A ticket booking system uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The architecture review board prefers a managed AWS-native control.

A.Aurora Global Database
B.A single-AZ Aurora cluster
C.An ElastiCache Redis replica
D.Manual snapshots copied monthly
AnswerA

Aurora Global Database is the correct answer because it replicates MySQL data at the storage layer to up to five secondary AWS Regions with typical replication latency under one second. This enables a promoted secondary cluster to become the primary in minutes, yielding a Recovery Time Objective (RTO) of roughly 1–2 minutes and a Recovery Point Objective (RPO) measured in seconds. Unlike snapshot-based backups, this is continuous, automatic replication that keeps the ticket booking system’s data nearly current in the disaster recovery Region, making it the only option that satisfies fast DR while preserving data consistency at scale.

Why this answer

Aurora Global Database is the correct choice because it provides a managed, cross-Region disaster recovery solution with a Recovery Point Objective (RPO) of typically less than 1 second, using storage-based replication that does not impact database performance. This meets the requirement for fast failover and low data loss, while being fully AWS-native and controlled by the architecture review board.

Exam trap

The trap here is that candidates may confuse cross-Region read replicas (which have higher RPO and require manual promotion) with Aurora Global Database, or assume that any caching layer like ElastiCache can substitute for database DR, when in fact only Aurora Global Database provides the required low RPO and managed failover.

How to eliminate wrong answers

Option B is wrong because a single-AZ Aurora cluster lacks any cross-Region replication or failover capability, offering no disaster recovery across Regions. Option C is wrong because ElastiCache Redis is an in-memory cache, not a persistent database, and cannot serve as a primary data store for ticket bookings or provide cross-Region DR with low RPO. Option D is wrong because manual snapshots copied monthly result in an RPO of up to a month, which is far too high for fast disaster recovery requirements.

95
MCQmedium

A trading dashboard stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured? The architecture review board prefers a managed AWS-native control.

A.An EBS snapshot schedule
B.S3 Cross-Region Replication with versioning enabled
C.S3 lifecycle transition to Glacier Flexible Retrieval
D.A CloudFront distribution
AnswerB

S3 Cross-Region Replication requires versioning enabled on both the source and destination buckets and asynchronously replicates every new object upload to a chosen destination Region. This creates a geographically separate, durable copy of the trading dashboard's documents, satisfying disaster-recovery requirements without manual intervention. Versioning is also essential because CRR relies on object version IDs to track replication state and to replicate delete markers or overwrites consistently.

Why this answer

S3 Cross-Region Replication (CRR) is the correct AWS-native managed solution for automatically replicating objects from a source S3 bucket in one region to a destination bucket in another region, meeting the disaster recovery requirement. Versioning must be enabled on both source and destination buckets for CRR to function, as replication relies on version IDs to track and copy objects. This provides asynchronous, automatic replication without custom scripting or third-party tools.

Exam trap

The trap here is that candidates may confuse S3 lifecycle policies (which only manage storage tiers within a region) with cross-region replication, or incorrectly assume CloudFront's global edge caching provides durable DR storage in another region.

How to eliminate wrong answers

Option A is wrong because EBS snapshots are for Amazon Elastic Block Store volumes attached to EC2 instances, not for S3 objects; they cannot replicate data across regions for S3-based storage. Option C is wrong because S3 lifecycle transitions to Glacier Flexible Retrieval only change the storage class within the same region for cost optimization, not replicate data to another region for disaster recovery. Option D is wrong because CloudFront is a content delivery network (CDN) that caches content at edge locations for low-latency access, but it does not provide cross-region replication or persistent storage in a secondary region for DR.

96
MCQhard

A financial analytics platform runs a stateless containerized service on Amazon ECS with AWS Fargate tasks spread across three Availability Zones. The service reads from an Amazon Aurora MySQL cluster and must continue serving read traffic if one Availability Zone fails. The team wants the read capacity to remain available with the least operational overhead and no changes to application connection strings during a zone failure. Which approach meets these requirements?

A.Place a Network Load Balancer in front of each Aurora Replica and configure the service to connect through the load balancer DNS name.
B.Add Aurora Replicas in multiple Availability Zones and connect the service to the cluster reader endpoint so reads are load-balanced and fail over automatically.
C.Increase the size of the Aurora writer instance so it can absorb read traffic if a replica becomes unavailable.
D.Create a separate Aurora cluster in each Availability Zone and have the service choose a cluster endpoint based on a health check.
AnswerB

Aurora Replicas in different Availability Zones provide redundant read capacity, and the cluster reader endpoint automatically distributes connections across healthy replicas. If a zone fails, Aurora removes the affected replica from the endpoint, so the application keeps reading without changing connection strings, satisfying the low-overhead and no-reconfiguration requirements.

Why this answer

The Aurora cluster reader endpoint is purpose-built to distribute read connections across replicas and to remove unhealthy replicas, including those in a failed Availability Zone. Deploying replicas in multiple zones and using that endpoint delivers automatic read failover with no application changes and minimal operational effort.

Exam trap

The trap here is adding external load balancing or per-zone clusters when Aurora's reader endpoint already provides managed read distribution and failover.

97
MCQmedium

A web application runs on an Auto Scaling group (ASG) behind an Application Load Balancer (ALB). The ASG uses the ALB target group health checks to decide when instances are healthy (for example, by using the ELB/target-group health check integration). During a deployment, the ASG performs instance replacement. Shortly after the deployment starts and while new instances are still bootstrapping, CloudWatch shows the ALB target group briefly has zero healthy targets, and users intermittently receive 502 responses. Which ASG deployment configuration best reduces the chance that there will be a period with zero healthy ALB targets, while still keeping failover behavior resilient?

A.Set the target group HealthCheckGracePeriod to a very short value so the ALB quickly declares instances healthy or unhealthy.
B.Use an ASG rolling update approach that launches replacement instances first, ensures the new instances pass the ALB target group health checks, and only then terminates the old instances (for example, by configuring sufficient minimum healthy capacity and waiting on ALB health).
C.Disable ALB target group health checks and route traffic to any registered targets so replacements do not depend on health check status.
D.Reduce the ASG desired capacity by one instance during deployments so the replacement happens faster.
AnswerB

This sequencing avoids a “no healthy targets” window. By keeping capacity stable (or maintaining a minimum healthy percentage) and waiting for the new instances to be marked healthy by the ALB, traffic is only sent to healthy targets during replacement.

Why this answer

It describes a rolling update strategy that launches new instances first, waits for them to pass ALB target group health checks, and only then terminates old instances. This ensures that at all times during the deployment, there is a sufficient number of healthy instances to serve traffic, preventing the ALB target group from ever having zero healthy targets. The ASG's minimum healthy capacity setting and the wait for ALB health check integration guarantee that failover remains resilient because the old instances continue to handle requests until the new ones are fully ready.

Exam trap

The trap here is that candidates often think reducing the health check grace period or disabling health checks will speed up recovery, but in reality, these actions either cause premature removal of healthy instances or allow traffic to unhealthy instances, both of which increase the likelihood of 502 errors and reduce resilience.

How to eliminate wrong answers

Option A is wrong because setting the HealthCheckGracePeriod to a very short value does not prevent zero healthy targets; it merely reduces the delay before the ALB marks instances as unhealthy, which could actually cause the ALB to prematurely remove instances and exacerbate the problem. Option C is wrong because disabling ALB target group health checks would cause the ALB to route traffic to any registered targets regardless of their actual health, leading to increased 502 errors and no failover resilience. Option D is wrong because reducing the ASG desired capacity by one instance during deployments does not address the root cause of zero healthy targets; it only reduces the number of instances being replaced, but the replacement process still creates a gap where old instances are terminated before new ones are healthy.

98
MCQhard

A claims workflow uses Amazon SQS. Poison messages are repeatedly failing and blocking useful retries. What should the architect configure? The team wants the control to be enforceable during normal operations.

A.A FIFO queue without a redrive policy
B.Short polling instead of long polling
C.A dead-letter queue with an appropriate maxReceiveCount
D.A larger message retention period only
AnswerC

A dead-letter queue with an appropriate maxReceiveCount is the formal SQS mechanism for poison messages. When a message is received more times than the configured threshold (e.g., 5), SQS automatically moves it to the DLQ using the source queue's redrive policy, isolating it from normal traffic. This lets the main queue continue processing healthy messages, while engineers can inspect and debug the quarantined messages later without disrupting the workflow.

Why this answer

A dead-letter queue (DLQ) with an appropriate maxReceiveCount is the correct solution because it automatically moves messages that have failed processing a specified number of times to a separate queue, preventing them from blocking subsequent retries. This enforces control during normal operations by isolating poison messages without manual intervention, allowing the main queue to continue processing valid messages.

Exam trap

The trap here is that candidates may confuse a dead-letter queue with simply increasing retention or changing polling methods, failing to recognize that only a DLQ with a maxReceiveCount enforces automatic removal of poison messages during normal operations.

How to eliminate wrong answers

Option A is wrong because a FIFO queue without a redrive policy does not handle poison messages; it only preserves message order and exactly-once processing, but without a DLQ, failed messages will continue to be retried indefinitely. Option B is wrong because short polling (returning immediately even if the queue is empty) does not address poison messages; it affects message availability timing, not retry behavior or failure handling. Option D is wrong because increasing the message retention period only extends how long messages stay in the queue; it does not limit retries or remove failing messages, so poison messages would still block retries until the retention period expires.

99
MCQeasy

A startup runs a small internal tool on a single Amazon EC2 instance that uses an instance store volume for its database files. After a routine host maintenance event, the instance rebooted and the database was empty. The team wants the data to persist independently of the instance lifecycle and to survive a stop-and-start of the instance. What should they change?

A.Enable termination protection on the instance and take an AMI of the instance before each maintenance window.
B.Increase the size of the instance store volume and enable detailed monitoring on the instance.
C.Move the database files to an Amazon EBS volume attached to the instance.
D.Place the instance in an Auto Scaling group with a minimum capacity of one so it is relaunched after maintenance.
AnswerC

Amazon EBS volumes are network-attached block storage that persists independently of the instance and remains available across stop and start operations within the same Availability Zone. Data is retained when the host is replaced, which directly fixes the loss the team observed. This is the standard persistent storage choice for EC2 databases.

Why this answer

Instance store volumes are physically attached to the host and their contents are lost when the host stops, fails, or is replaced during maintenance. Amazon EBS volumes live separately from the instance and retain data across stop and start, making them the correct storage layer for a database that must persist. The other options address compute lifecycle or monitoring rather than data durability.

Exam trap

The trap here is assuming that any storage attached to an instance survives a host event, when instance store is explicitly ephemeral.

100
MCQmedium

A ticket booking system runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include?

A.Subnets in at least two Availability Zones with health checks enabled
B.All instances in one larger subnet
C.A Network Load Balancer in one subnet
D.A single EC2 instance with detailed monitoring
AnswerA

Deploying an Auto Scaling group across subnets in at least two Availability Zones (AZs) is the core of high availability. Each AZ is an isolated failure domain, so if one AZ fails, the ASG continues to run instances in the other AZ(es). Coupled with ELB or ASG health checks, the group automatically detects unhealthy instances—whether due to EC2 failure or AZ impairment—and replaces them, maintaining desired capacity and absorbing request traffic. This design avoids any single point of failure at both the compute and network layers.

Why this answer

An Auto Scaling group configured with subnets in at least two Availability Zones and health checks enabled ensures that if one AZ fails, EC2 instances in the remaining AZs continue to serve traffic. The Application Load Balancer distributes requests across healthy instances in multiple AZs, and the Auto Scaling group replaces failed instances in the affected AZ, maintaining capacity. This design meets the requirement to tolerate the failure of one Availability Zone.

Exam trap

The trap here is that candidates often think a single larger subnet or a different load balancer type provides resilience, but only distributing subnets across multiple Availability Zones with health checks ensures the system can survive an AZ failure.

How to eliminate wrong answers

Option B is wrong because placing all instances in one larger subnet within a single Availability Zone creates a single point of failure; if that AZ goes down, all instances become unavailable. Option C is wrong because a Network Load Balancer operates at Layer 4 and does not provide the HTTP/HTTPS health checks or path-based routing needed for a ticket booking system, and placing it in one subnet does not address multi-AZ resilience. Option D is wrong because a single EC2 instance, even with detailed monitoring, cannot survive an AZ failure; there is no redundancy or automatic failover.

101
Multi-Selectmedium

A media company stores daily financial exports in Amazon S3. The files must be protected against accidental overwrite or deletion, and the business also wants a second copy in another Region for recovery after a regional outage. Which two actions should the architect take? Select two.

Select 2 answers
A.Enable bucket versioning on the S3 bucket.
B.Turn on S3 Transfer Acceleration for the bucket.
C.Use only lifecycle policies to move objects to Glacier.
D.Configure replication to a bucket in a second AWS Region.
E.Enable S3 Block Public Access on the bucket.
AnswersA, D

Bucket versioning preserves every object revision, so an accidental overwrite creates a new version while the prior data remains intact, and a delete merely adds a delete marker rather than erasing the object. This directly satisfies the stem's requirement to protect the daily financial exports against accidental overwrite or deletion.

Why this answer

Option A is correct because enabling S3 bucket versioning preserves every prior version of an object, so an accidental overwrite creates a new version and an accidental delete only adds a delete marker, allowing the original data to be restored rather than lost. Option D is correct because S3 Cross-Region Replication (CRR) automatically copies objects to a bucket in a second AWS Region, providing the required second copy for recovery from a regional outage; versioning must be enabled on both source and destination buckets for replication to work. Option B is not relevant because S3 Transfer Acceleration only speeds up uploads over long distances using edge locations, and does not protect data or create a second copy.

Option C is not appropriate because lifecycle policies to Glacier change storage class and cost, not immutability, and do not provide cross-Region recovery. Option E is not relevant because S3 Block Public Access only prevents public exposure and does not protect against overwrite, deletion, or regional failure.

Exam trap

The trap here is that candidates may confuse S3 Transfer Acceleration or Block Public Access with data protection features, when in fact only versioning and replication directly address the requirements for preventing accidental deletion and providing cross-region recovery.

Why the other options are wrong

B

S3 Transfer Acceleration speeds up uploads over long distances but does not protect against accidental deletion or overwrite, nor does it create a cross-region copy for disaster recovery.

C

Lifecycle policies to move objects to Glacier provide cost optimization for long-term storage, but do not protect against accidental overwrite/deletion or provide cross-region recovery.

E

Block Public Access prevents public access to S3 objects but does not protect against accidental overwrite or deletion by authorized users, nor does it provide cross-region replication for disaster recovery.

102
Multi-Selectmedium

A serverless order-ingestion API writes directly to a database. During traffic spikes, the database occasionally throttles, Lambda retries create duplicate order records, and some requests time out. Which two changes best improve buffering and safe retry behavior? Select two.

Select 2 answers
A.Increase the Lambda timeout and keep writing directly to the database.
B.Put an Amazon SQS queue between the API and the database-processing function.
C.Replace SQS with SNS so every request is delivered immediately to all subscribers.
D.Make the database write idempotent by using a unique request token or order ID.
E.Disable retries so failed writes are never duplicated.
AnswersB, D

Amazon SQS decouples ingestion from database writes, absorbing traffic spikes as a buffer while the processing function consumes at a controlled rate. It also enables safe retries via visibility timeouts and dead-letter queues, preventing duplicate order records and timeouts.

Why this answer

Option B is correct because inserting an Amazon SQS queue between the API and the database-processing function decouples ingestion from database writes, allowing the queue to absorb traffic spikes and buffer requests while the consumer processes them at a rate the database can sustain, which directly addresses throttling and timeouts. Option D is correct because making the database write idempotent using a unique request token or order ID ensures that Lambda retries of the same message do not create duplicate order records, which is the standard safe-retry pattern for at-least-once delivery systems like SQS. Option A is not appropriate because increasing the Lambda timeout while continuing direct database writes does nothing to buffer spikes or prevent duplicate records, and longer timeouts can worsen throttling pressure.

Option C is not appropriate because SNS is a pub/sub fan-out service, not a durable buffering queue, and immediate delivery to all subscribers does not provide the backpressure or retry buffering needed here. Option E is not appropriate because disabling retries sacrifices reliability and does not solve the underlying throttling or duplicate-write problem.

Exam trap

The trap here is that candidates often think SNS (Option C) is a suitable replacement for SQS because both are messaging services, but SNS lacks buffering and retry mechanics, making it inappropriate for smoothing traffic spikes and handling failures gracefully.

Why the other options are wrong

A

Increasing Lambda timeout does not address database throttling or duplicate records; it only allows the function to wait longer, but the database will still throttle under load, and retries will still create duplicates.

C

SNS pushes messages to all subscribers immediately without buffering or throttling, so it does not help with database throttling or retry management; it would still overwhelm the database and cause duplicate processing.

E

Disabling retries entirely would cause order writes to fail permanently during throttling, losing data and increasing timeouts, which contradicts the need for safe retry behavior.

103
MCQhard

A company runs a production MySQL database on Amazon RDS in us-east-1. A read replica exists in us-west-2 for disaster recovery. The primary region experiences a complete outage. Which of the following describes the correct procedure to restore database service using the cross-region read replica?

A.Wait for AWS to automatically fail over the read replica to become the new primary
B.Restore the primary database from the most recent automated snapshot in us-west-2
C.Manually promote the us-west-2 read replica to a standalone DB instance and update application endpoints
D.Create a new RDS instance in us-west-2 and manually restore data from application logs
AnswerC

Promoting the us-west-2 read replica is the correct DR procedure. Promotion converts the read replica into a standalone writable DB instance, and because the replica continuously replays transactions from the primary, it typically has only seconds of lag. After promotion, applications must update their connection strings to the new endpoint in us-west-2. This is the standard, lowest-RPO approach for cross-region failover.

Why this answer

Cross-region RDS read replicas support manual promotion to a standalone database instance. When the primary region fails, the replica must be manually promoted — this makes it an independent writable instance in us-west-2.

Key points: Promotion is NOT automatic (unlike RDS Multi-AZ failover). Promotion breaks the replication link — the replica becomes autonomous. After promotion, application connection strings must be updated to the new endpoint. Any replication lag at the time of the outage represents potential data loss (RPO > 0).

Exam trap

RDS Multi-AZ provides automatic failover — no manual action required. Cross-region read replicas do NOT failover automatically — promotion must be manually triggered. This distinction appears frequently.

For automatic cross-region failover with near-zero RPO, use Amazon Aurora Global Database.

Why the other options are wrong

A

RDS cross-region read replicas do NOT automatically failover. Only RDS Multi-AZ provides automatic same-region failover. Manual promotion is required for cross-region replicas.

B

Restoring from a snapshot creates a new instance from an older state. The read replica contains more recent data (continuously synchronized). Promotion is faster and yields less data loss than snapshot restoration when the replica is available.

D

Creating an empty new instance and manually re-entering data is not a valid DR procedure. The read replica already contains synchronized production data. Never manually re-enter data as part of a DR plan.

104
MCQeasy

An orders service consumes payment instructions from an Amazon SQS queue. Sometimes the consumer times out after applying the payment but before deleting the SQS message. As a result, the same payment instruction is processed again. Which design change most directly prevents duplicate side effects caused by message retries?

A.Delete the SQS message immediately after it is received, before processing, to ensure it is not retried.
B.Implement idempotency by recording a processed marker keyed by the instruction ID and ignoring duplicates.
C.Increase the SQS visibility timeout to a maximum value to avoid retries entirely.
D.Convert the queue to FIFO and enable content-based deduplication.
AnswerB

Idempotency ensures that repeated deliveries of the same instruction do not cause repeated side effects. By persisting a record keyed by instruction ID (or enforcing a unique constraint in a transactional store), the service can detect duplicates and safely skip or reconcile them even if SQS redelivers the message.

Why this answer

Implementing idempotency ensures that even if the same payment instruction is processed multiple times due to a timeout and retry, the side effect (e.g., applying the payment) occurs only once. By recording a processed marker keyed by the instruction ID (e.g., using a DynamoDB table or Redis), the consumer can check the marker before processing and ignore duplicates. This directly addresses the root cause—duplicate processing—without altering the queue's retry behavior.

Exam trap

The trap here is that candidates confuse message deduplication (preventing duplicate deliveries) with idempotent processing (preventing duplicate side effects), leading them to choose Option D, which only prevents redelivery but does not handle the case where the same message is processed twice due to a consumer timeout before deletion.

How to eliminate wrong answers

Option A is wrong because deleting the SQS message immediately after receipt, before processing, defeats the purpose of at-least-once delivery; if the consumer crashes after deletion but before processing, the payment instruction is lost permanently, leading to data loss. Option C is wrong because increasing the visibility timeout to a maximum value (e.g., 12 hours) does not prevent retries entirely; the message will still be retried if the consumer fails to delete it within the timeout, and it can also delay processing of other messages. Option D is wrong because converting to a FIFO queue with content-based deduplication deduplicates based on the message body, not the processing outcome; if the same message is received again due to a consumer timeout, the deduplication ID (derived from the body) remains the same, so the message is not redelivered—but this does not prevent the duplicate side effect from the first retry that already occurred, and it also requires the queue to be FIFO, which may not suit the existing architecture.

105
MCQmedium

A public API is deployed in two AWS Regions: us-east-1 (primary) and us-west-2 (secondary). The team wants Route 53 to automatically route users to the secondary region if the primary API becomes unhealthy. They will use Route 53 health checks that monitor the API’s /status endpoint over HTTPS. Which Route 53 configuration most directly implements this failover behavior?

A.Create two latency-based alias records for the same name, each with different health checks; Route 53 will automatically shift to the secondary when primary is unhealthy.
B.Create a primary alias record and a failover alias record (secondary), configure failover routing policy, and attach health checks to both records.
C.Use geolocation routing with a health check; when the primary is unhealthy, Route 53 will automatically change the region mapping globally.
D.Use simple routing with weighted records and a low health check threshold so traffic quickly moves to the secondary region.
AnswerB

Route 53 failover routing (primary/secondary) is designed for active-passive regional DR. When the primary health check fails, Route 53 automatically stops returning the primary alias and returns the secondary alias target; attaching health checks ensures the change is driven by the /status endpoint health.

Why this answer

B is correct because the failover routing policy in Route 53 is specifically designed for active-passive failover. By creating a primary alias record and a secondary failover alias record, each with an associated health check, Route 53 will automatically route traffic to the secondary region when the health check for the primary fails. This directly implements the required behavior without relying on latency or geographic proximity.

Exam trap

The trap here is that candidates often confuse failover routing with latency-based or geolocation routing, assuming that health checks automatically trigger failover in those policies, but only failover routing provides the explicit active-passive failover behavior described in the question.

How to eliminate wrong answers

Option A is wrong because latency-based routing does not support automatic failover based on health checks; it routes based on lowest latency, and while health checks can be associated, Route 53 does not automatically shift traffic to the secondary when the primary is unhealthy—it continues to return the primary record if it is still considered healthy, and if both are healthy, latency determines the response. Option C is wrong because geolocation routing routes based on the user's geographic location, not health; even with a health check, Route 53 does not automatically change region mappings globally—it would only return no answer for the unhealthy location, not redirect to another region. Option D is wrong because simple routing with weighted records does not support health checks for automatic failover; weighted routing distributes traffic based on weights and does not automatically shift all traffic to the secondary when the primary is unhealthy—it would require manual intervention or complex scripting.

106
Multi-Selecthard

A regional web application for a inventory service must fail over automatically to a secondary Region if the primary endpoint becomes unhealthy. Which two services or features are required? The team wants the control to be enforceable during normal operations.

Select 2 answers
A.Route 53 failover routing with health checks
B.S3 Transfer Acceleration
C.A deployed standby application stack in the secondary Region
D.AWS Organizations service control policies
AnswersA, C

Route 53 failover routing with health checks monitors the primary endpoint and automatically redirects DNS queries to the secondary record when it becomes unhealthy. This satisfies the automatic failover requirement, and the routing policy remains enforceable during normal operations.

Why this answer

Option A (Route 53 failover routing with health checks) is correct because Route 53 failover routing policies use health checks on the primary endpoint; when the health check reports the primary as unhealthy, Route 53 automatically returns the secondary Region's record, providing the required automatic DNS-level failover. Option C (a deployed standby application stack in the secondary Region) is correct because failover routing can only redirect traffic to a target that actually exists and can serve requests, so a running standby stack (for example, on EC2, ECS, or Elastic Beanstalk) must already be deployed in the secondary Region for the failover to succeed. Option B (S3 Transfer Acceleration) is incorrect because it only speeds up uploads/downloads to S3 buckets using edge locations and does not provide health-check-based regional failover for a web application.

Option D (AWS Organizations service control policies) is incorrect because SCPs only set permission guardrails for accounts in an organization and cannot detect an unhealthy endpoint or redirect traffic between Regions.

Exam trap

The trap here is that candidates may think DNS-level failover alone is sufficient, forgetting that a fully deployed standby stack in the secondary Region is required to actually serve traffic after failover.

107
MCQhard

A financial services company runs a critical application on Amazon EC2 instances in an Auto Scaling group. The application writes to an Amazon RDS for MySQL database. The company needs a recovery point objective (RPO) of 1 second and a recovery time objective (RTO) of 1 minute for the database in the event of a Regional disaster. Which solution meets these requirements?

A.Configure an RDS for MySQL read replica in a second Region and promote it during a disaster.
B.Enable automated backups on the RDS for MySQL instance and copy the backups to a second Region using a cross-Region snapshot copy.
C.Use AWS Database Migration Service (DMS) to continuously replicate data from the primary RDS instance to a second Region, and switch the application to the replica during a disaster.
D.Use an Amazon Aurora global database with a secondary Region, and perform a managed planned failover or unplanned failover to the secondary Region.
AnswerD

Aurora global database replicates data to a secondary Region with a typical latency of under one second, meeting the 1-second RPO. Failover to the secondary Region can be completed in under a minute, meeting the 1-minute RTO. This is the only option that provides both low RPO and low RTO across Regions for MySQL-compatible workloads.

Why this answer

Amazon Aurora global database is designed for cross-Region disaster recovery with low latency replication and fast failover. Replication typically completes in under one second, satisfying the RPO. A managed failover promotes the secondary Region to primary in under a minute, satisfying the RTO.

This is the appropriate choice for a MySQL-compatible workload with stringent recovery requirements.

Exam trap

The trap here is assuming that cross-Region read replicas or snapshot copies can meet a 1-second RPO and 1-minute RTO, when their replication lag and restore times are much higher.

108
MCQmedium

A payments platform requires disaster recovery across Regions. Requirements: RPO of 15 minutes and RTO of about 1 hour. The business cannot afford full duplicate capacity in both Regions all the time, but the team wants automated readiness so failover is mostly operationally guided rather than a slow rebuild. Which DR strategy is the best fit?

A.Backup and restore only, relying on scheduled snapshots and manual restores during incidents.
B.Pilot light, keeping only minimal infrastructure in the secondary Region and starting full services after failover.
C.Warm standby, keeping core infrastructure and a partially provisioned environment ready in the secondary Region with frequent data replication.
D.Active/active, routing production traffic to both Regions continuously and accepting dual-region complexity.
AnswerC

Warm standby keeps core services and a scaled-down but functional environment running in the secondary Region with continuous replication, achieving the 15-minute RPO and roughly one-hour RTO. It avoids the cost of full duplicate capacity while remaining far faster than a rebuild.

Why this answer

Warm standby is the best fit because it maintains a partially provisioned environment in the secondary Region with core infrastructure (e.g., a smaller EC2 instance fleet, a replicated database) and uses frequent data replication (e.g., Amazon RDS cross-Region replication or DynamoDB global tables) to achieve an RPO of 15 minutes. The RTO of about 1 hour is achievable by scaling up the standby environment and redirecting traffic, which is faster than a full rebuild but avoids the cost of full duplicate capacity. This balances the business constraint of not affording active/active with the need for automated readiness and guided failover.

Exam trap

The trap here is that candidates often confuse pilot light with warm standby, assuming minimal infrastructure is sufficient for a 1-hour RTO, but pilot light requires provisioning compute resources after failover, which adds significant time, whereas warm standby already has compute running and only needs scaling.

How to eliminate wrong answers

Option A is wrong because backup and restore only relies on scheduled snapshots (e.g., EBS snapshots or RDS automated backups) and manual restores, which typically cannot achieve an RPO of 15 minutes (snapshots are often taken every few hours) and would result in an RTO far exceeding 1 hour due to manual intervention and data restoration time. Option B is wrong because pilot light keeps only minimal infrastructure (e.g., a small database replica and no application servers) in the secondary Region, and starting full services after failover requires provisioning compute resources, which would likely exceed the 1-hour RTO target. Option D is wrong because active/active requires full duplicate capacity in both Regions all the time, which contradicts the business constraint that they cannot afford this, and it introduces dual-region complexity that is unnecessary for the stated RPO/RTO goals.

109
MCQhard

A financial services firm runs a latency-sensitive trading application on Amazon EC2 instances distributed across three Availability Zones behind a Network Load Balancer. The application must continue serving traffic with no manual intervention if an entire Availability Zone becomes impaired, and each instance must receive a fair share of connections. Which combination of features meets these requirements?

A.Enable cross-zone load balancing on the Network Load Balancer and register targets in all three Availability Zones
B.Use an Application Load Balancer with sticky sessions enabled and register targets in all three Availability Zones
C.Attach an Elastic IP address to each instance and have clients connect directly using a published list of addresses
D.Create a Network Load Balancer with one target group per Availability Zone and rely on DNS failover between the groups
AnswerA

Cross-zone load balancing on a Network Load Balancer distributes traffic evenly across registered targets in every enabled Availability Zone, so a single impaired zone does not concentrate load on the surviving instances. With targets registered in all three zones, the load balancer's health checks automatically remove unhealthy targets and continue routing to the rest. This delivers both even connection distribution and unattended zone-failure tolerance.

Why this answer

Cross-zone load balancing on a Network Load Balancer spreads incoming connections evenly across registered targets in all enabled Availability Zones, which is exactly what the fair-share requirement demands. Registering targets in three zones lets health checks detect an impaired zone and remove its targets so the remaining instances absorb traffic without operator action. The other options either bypass health checking, bind clients to single targets, or rely on slow DNS-based failover.

Exam trap

The trap here is assuming that a Network Load Balancer distributes traffic evenly across zones by default; cross-zone load balancing must be enabled, and forgetting it can leave one zone carrying a disproportionate share.

110
MCQhard

A company runs a stateful web application on a fleet of Amazon EC2 instances in an Auto Scaling group. The application stores session state locally on each instance. During an Availability Zone failure, the Auto Scaling group replaces the unhealthy instances in a different AZ, but users lose their sessions and must log in again. The company wants to make the application resilient to AZ failures without requiring users to re-authenticate. Which solution should a solutions architect recommend?

A.Use Amazon S3 to store session data and configure the application to read and write session files directly to an S3 bucket.
B.Configure the Auto Scaling group to launch instances in only one Availability Zone and use a larger instance type to handle the load.
C.Store session state in an Amazon ElastiCache for Redis cluster with Multi-AZ enabled, and configure the application to use it.
D.Enable sticky sessions on the Application Load Balancer and increase the Auto Scaling group's desired capacity.
AnswerC

ElastiCache for Redis with Multi-AZ provides automatic failover to a replica in another AZ, and the session data is externalized from the EC2 instances. When an instance is replaced, it can retrieve the session from the Redis cluster, so users remain authenticated. This design decouples session state from the compute layer and survives an AZ failure.

Why this answer

Externalizing session state to a Multi-AZ ElastiCache for Redis cluster removes the dependency on local instance storage. When an instance is replaced after an AZ failure, the new instance can retrieve the session from Redis, so users stay logged in. This decouples state from compute and provides automatic failover, meeting the resilience requirement.

Exam trap

The trap here is thinking that sticky sessions or increased capacity can preserve session state when an Availability Zone fails.

111
MCQeasy

Based on the exhibit, the database must continue serving if the current Availability Zone fails. What should you change?

A.Create a read replica in another Availability Zone and promote it manually if needed.
B.Modify the DB instance to use a Multi-AZ deployment.
C.Increase the automated backup retention period to 30 days.
D.Resize the DB instance to a larger class.
AnswerB

A Multi-AZ RDS deployment provides synchronous standby replication in another Availability Zone and automatic failover if the primary AZ becomes unavailable. This directly matches the requirement to keep the database serving after an AZ failure. It is the simplest resilient design change when the application needs high availability rather than just backups.

Why this answer

Multi-AZ deployment automatically provisions and maintains a synchronous standby replica in a different Availability Zone. If the primary AZ fails, Amazon RDS automatically fails over to the standby, ensuring database availability without manual intervention. This meets the requirement of continuing service during an AZ failure.

Exam trap

The trap here is confusing read replicas (which are for read scaling and asynchronous replication) with Multi-AZ (which is for high availability and synchronous replication), leading candidates to choose Option A for failover scenarios.

Why the other options are wrong

A

Creating a read replica in another AZ does not provide automatic failover; manual promotion introduces downtime, failing the requirement for continuous serving if the current AZ fails.

C

Increasing the automated backup retention period does not provide high availability or automatic failover; it only extends the point-in-time recovery window, which does not help if the current Availability Zone fails.

D

Resizing the DB instance to a larger class improves performance but does not provide failover capability if the current Availability Zone fails. It does not address high availability across AZs.

112
MCQeasy

A inventory service exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most?

A.CloudFront caching with appropriate TTLs
B.AWS Backup Vault Lock
C.IAM Access Analyzer
D.S3 Select
AnswerA

CloudFront caching with appropriate TTLs is correct because it stores static content at edge locations, allowing the website to be served to users from cache even when the S3 origin is temporarily unavailable. As long as the TTL has not expired, users experience no interruption, and the cached content provides a buffer against origin outages. This makes the service more resilient, provided the TTL balances freshness and availability.

Why this answer

CloudFront caches responses from the S3 origin based on configured TTLs (Cache-Control or Expires headers). If the S3 origin experiences a short outage, CloudFront can still serve cached pages to users from its edge locations, maintaining availability. This is the most direct way to ensure users receive content during origin failures.

Exam trap

The trap here is that candidates may confuse CloudFront's caching with other AWS services like S3 Transfer Acceleration or S3 Cross-Region Replication, which do not provide cached responses during origin outages.

How to eliminate wrong answers

Option B (AWS Backup Vault Lock) is wrong because it is a data protection feature for backups, enforcing retention policies and preventing deletion, not related to serving cached web content during origin outages. Option C (IAM Access Analyzer) is wrong because it analyzes resource-based policies to identify unintended public or cross-account access, not for caching or origin failover. Option D (S3 Select) is wrong because it is a feature to retrieve subsets of object data using SQL queries, not for caching or serving static content during outages.

113
MCQmedium

Based on the exhibit, a faulty deployment corrupted production data at 10:30 UTC and the issue was discovered at 10:55 UTC. The team needs to recover the database to the last good state before the corruption. Which action should they take?

A.Restore the latest manual snapshot and accept data loss since the snapshot was taken overnight.
B.Use point-in-time restore to create a new database instance at 10:29 UTC, then switch the application to it.
C.Restart the database instance so the transaction log replays the failed migration cleanly.
D.Create a read replica and promote it, because replicas always contain the previous transaction state.
AnswerB

Point-in-time restore is the correct recovery method when automated backups are enabled and the team needs the database just before a known corruption event. Restoring to 10:29 UTC brings the data back to the last safe moment before the migration began. Creating a new instance first avoids modifying the damaged database until the restored copy is validated.

Why this answer

Amazon RDS for MySQL (and other engines) supports point-in-time recovery (PITR), which allows you to restore a database to any second within the backup retention period, up to the last five minutes. By restoring to 10:29 UTC (one minute before the corruption at 10:30 UTC), the team can recover the database to its last good state with minimal data loss. After restoring, the application can be pointed to the new instance, avoiding the corrupted data.

Exam trap

The trap here is that candidates may confuse point-in-time restore with snapshot restore, assuming snapshots are the only recovery option, or incorrectly believe that restarting or promoting a replica can undo a logical corruption that has already been written to disk.

How to eliminate wrong answers

Option A is wrong because restoring the latest manual snapshot would revert the database to the time the snapshot was taken (likely overnight), causing significant data loss of all transactions between that snapshot and 10:30 UTC, which is unacceptable when a more precise recovery is available. Option C is wrong because restarting the database instance does not replay transaction logs to undo a faulty deployment; it only replays committed transactions from the binary logs to ensure consistency, which would reapply the corruption. Option D is wrong because a read replica contains the same data as the primary at the time of replication lag, not a previous transaction state; promoting it would still include the corrupted data if the corruption occurred before the replica caught up.

114
MCQeasy

An order system receives events and uses a Lambda function to write each order into a database. During traffic spikes, the database sometimes throttles, and Lambda retries lead to occasional message loss in the event flow. The team wants buffering, automatic retries, and a way to isolate messages that repeatedly fail so they can be inspected later. What design change best meets this need?

A.Send events directly from EventBridge to Lambda without any queue to simplify the flow.
B.Use Amazon SQS as a buffer between the event source and Lambda, with an SQS dead-letter queue (DLQ).
C.Use SNS fan-out to multiple Lambda functions, but keep no retry logic and no DLQ.
D.Store events in an S3 bucket and trigger Lambda immediately after each upload, without using DLQs.
AnswerB

Using SQS as a buffer between the event source and Lambda is correct because SQS decouples event producers from the consumer, smoothing traffic spikes by buffering messages until Lambda can poll them. The visibility timeout provides built-in retry logic: if Lambda fails to process a message, it becomes visible again for another attempt, and after a configured Maximum Receives threshold, the message is automatically diverted to a dead-letter queue for later analysis. This pattern ensures that no order event is lost, supports backpressure, and lets you isolate recurring processing failures without disrupting the continuous flow of valid events.

Why this answer

Amazon SQS acts as a durable buffer between the event source and Lambda, absorbing traffic spikes and decoupling the producer from the consumer. The SQS dead-letter queue (DLQ) automatically captures messages that exceed the configured maximum retries, allowing the team to inspect and reprocess them later without loss. This design provides the required buffering, automatic retries via the Lambda event source mapping, and isolation of repeatedly failing messages.

Exam trap

The trap here is that candidates often assume a direct event-driven flow (like EventBridge to Lambda) is simpler and sufficient, but they overlook the need for buffering and a DLQ to handle throttling and isolate persistent failures, which SQS explicitly provides.

Why the other options are wrong

A

Sending events directly from EventBridge to Lambda without a queue provides no buffering during traffic spikes, so Lambda retries still cause throttling and message loss. It also lacks a dead-letter queue to isolate repeatedly failing messages.

C

SNS fan-out without retry logic or a DLQ does not provide buffering, automatic retries, or isolation of failed messages, which are explicitly required to handle throttling and prevent message loss.

D

S3 event notifications have no built-in retry logic or dead-letter queue; if Lambda fails, the event is lost after the retry limit, failing to isolate repeatedly failing messages for inspection.

115
MCQmedium

A company uses Amazon RDS for a PostgreSQL database powering a customer-facing application. The application’s availability depends on fast database failover with minimal manual intervention. The RDS instance currently runs as a single-AZ deployment in one DB subnet group. Which change most directly meets the goal?

A.Create a read replica in a different Availability Zone and configure the application to fail over manually.
B.Enable Multi-AZ for the RDS DB instance so AWS manages a standby in another Availability Zone with automatic failover.
C.Switch the database to use EBS snapshots more frequently and restore in case of failure.
D.Pin the DB to a specific instance type with higher CPU credits to prevent CPU-related disconnects.
AnswerB

RDS Multi-AZ creates a synchronous standby replica in a separate Availability Zone and automatically flips the DNS endpoint to the standby within seconds if the primary instance fails, a network partition occurs, or an AZ becomes unavailable. This is the only option that provides true high availability with no manual intervention, making it the correct answer because it directly addresses resilience to an entire Availability Zone failure.

Why this answer

Enabling Multi-AZ for the RDS DB instance creates a synchronous standby replica in a different Availability Zone. AWS automatically handles failover to the standby with no manual intervention required, which directly meets the goal of fast database failover with minimal manual intervention.

Exam trap

The trap here is that candidates often confuse read replicas (which are for read scaling and manual promotion) with Multi-AZ (which provides automatic failover and high availability), leading them to choose Option A incorrectly.

Why the other options are wrong

A

This option requires manual failover, which contradicts the requirement for minimal manual intervention and fast failover.

C

Frequent EBS snapshots provide point-in-time recovery but do not enable fast, automatic failover; restoring from a snapshot requires manual intervention and takes significant time, failing to meet the minimal manual intervention and fast failover requirements.

D

Higher CPU credits prevent CPU-related throttling but do not address database failover speed or minimize manual intervention, which is the core requirement for fast failover in a single-AZ deployment.

116
MCQmedium

A trading dashboard runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include?

A.A single EC2 instance with detailed monitoring
B.Subnets in at least two Availability Zones with health checks enabled
C.All instances in one larger subnet
D.A Network Load Balancer in one subnet
AnswerB

Placing subnets in at least two Availability Zones (AZs) and attaching health checks to the Auto Scaling group allows the group to detect and replace unhealthy instances while maintaining desired capacity across AZs. If one AZ fails, the remaining healthy instances in the other AZ continue serving traffic, and Auto Scaling launches new instances in the surviving AZs to compensate. Health checks (ELB or EC2 status checks) drive the replacement process, and distributing subnets across AZs makes the architecture resilient to both instance-level and AZ-level failures, which is the key requirement for a fault-tolerant trading dashboard.

Why this answer

Distributing EC2 instances across subnets in at least two Availability Zones ensures that if one AZ fails, the Auto Scaling group can maintain capacity using instances in the remaining AZ(s). Enabling health checks allows the group to detect and replace unhealthy instances, which is essential for fault tolerance. This configuration meets the requirement to tolerate the failure of one Availability Zone.

Exam trap

The trap here is that candidates often confuse high availability with fault tolerance, thinking a single large subnet or a single instance with monitoring is sufficient, when in fact distributing across multiple Availability Zones is the key to surviving an AZ failure.

How to eliminate wrong answers

Option A is wrong because a single EC2 instance, even with detailed monitoring, cannot tolerate the failure of an entire Availability Zone; if that AZ goes down, the instance becomes unavailable. Option C is wrong because placing all instances in one larger subnet confines them to a single Availability Zone, providing no redundancy if that AZ fails. Option D is wrong because a Network Load Balancer in one subnet does not solve the AZ failure requirement; the Auto Scaling group must span multiple AZs, and the load balancer itself should be cross-zone enabled to distribute traffic across AZs.

117
MCQmedium

A trading dashboard stores uploaded documents in S3. The business requires a copy in another AWS Region for disaster recovery. What should be configured?

A.An EBS snapshot schedule
B.S3 Cross-Region Replication with versioning enabled
C.S3 lifecycle transition to Glacier Flexible Retrieval
D.A CloudFront distribution
AnswerB

Cross-Region Replication copies objects asynchronously to a bucket in another AWS Region, satisfying the disaster-recovery constraint. Versioning must be enabled on both source and destination because replication requires it to track object versions and propagate deletions or overwrites correctly.

Why this answer

S3 Cross-Region Replication (CRR) with versioning enabled automatically replicates objects to a destination bucket in a different AWS Region, providing a durable, low-latency disaster recovery copy. Versioning must be enabled on both source and destination buckets to track object changes and ensure consistency during replication. This meets the requirement for a cross-region copy without manual intervention.

Exam trap

The trap here is that candidates may confuse lifecycle transitions (which change storage class within the same region) with cross-region replication (which copies data to a different region), or assume EBS snapshots apply to S3 storage.

How to eliminate wrong answers

Option A is wrong because EBS snapshots are used for backing up EC2 block storage volumes, not for S3 objects, and they are region-specific unless manually copied. Option C is wrong because S3 lifecycle transition to Glacier Flexible Retrieval moves objects to a cold storage tier for cost savings, not to a different AWS Region for disaster recovery. Option D is wrong because CloudFront is a content delivery network that caches data at edge locations for low-latency access, not a mechanism for replicating data to another region for DR.

118
MCQeasy

An internal service is hosted behind an Application Load Balancer (ALB) with targets spread across two Availability Zones. If the targets in one Availability Zone become unhealthy, the service must continue serving traffic from the healthy AZ. What change most directly improves resilience at the load-balancing layer?

A.Turn off health checks and rely only on instance CPU utilization to route traffic.
B.Configure ALB listener rules to route all traffic to a single target group in one Availability Zone.
C.Configure target group health checks so the ALB stops sending traffic to unhealthy targets and continues routing to healthy targets in the other Availability Zone.
D.Store requests in an SQS queue before routing them to the ALB.
AnswerC

With target group health checks enabled and configured correctly, the ALB evaluates each target's health and stops routing requests to targets marked unhealthy. As long as healthy targets exist in the other AZ, the ALB preserves reachability.

Why this answer

Configuring target group health checks allows the ALB to automatically detect unhealthy targets and stop sending traffic to them, while continuing to route requests to healthy targets in the other Availability Zone. This directly improves resilience at the load-balancing layer by ensuring traffic is only forwarded to healthy instances, maintaining service availability even when an entire AZ fails.

Exam trap

The trap here is that candidates may think SQS decoupling (Option D) improves resilience at the load-balancing layer, but SQS operates at the application layer and does not affect how the ALB routes traffic to unhealthy targets.

How to eliminate wrong answers

Option A is wrong because turning off health checks removes the ALB's ability to detect unhealthy targets, which would cause traffic to be sent to failed instances, breaking resilience. Option B is wrong because routing all traffic to a single target group in one AZ creates a single point of failure and defeats the purpose of multi-AZ redundancy. Option D is wrong because storing requests in an SQS queue before routing to the ALB adds unnecessary latency and complexity, and does not address the immediate need for the ALB to stop sending traffic to unhealthy targets.

119
Multi-Selecthard

A regional web application for a inventory service must fail over automatically to a secondary Region if the primary endpoint becomes unhealthy. Which two services or features are required? The design must avoid adding custom operational scripts.

Select 2 answers
A.Route 53 failover routing with health checks
B.S3 Transfer Acceleration
C.A deployed standby application stack in the secondary Region
D.AWS Organizations service control policies
AnswersA, C

Route 53 failover routing with health checks is the control-plane mechanism that implements active-passive regional failover. You configure a primary record with a health check that probes the primary endpoint's HTTP or TCP status, and a secondary record pointing to the standby region; when the health check fails, Route 53 stops returning the primary and starts resolving to the standby. The health check must be created separately and attached to the primary record set, and low TTLs or alias records are recommended so DNS clients quickly pick up the failed-over answer.

Why this answer

Route 53 failover routing with health checks (Option A) is required because it automatically evaluates the health of the primary endpoint and, upon detecting failure, updates DNS resolution to direct traffic to the secondary Region. This is the native AWS mechanism for DNS-based failover without custom scripts, relying on Route 53 health checkers to assess endpoint health via HTTP/HTTPS/TCP or calculated health checks.

Exam trap

The trap here is that candidates may think a single service like Route 53 alone can handle failover, but without a pre-deployed standby application stack in the secondary Region, there is no infrastructure to route traffic to, making both Route 53 failover routing and the standby stack required together.

120
MCQmedium

A media company runs a stateless transcoding fleet on Amazon EC2 instances spread across three Availability Zones behind a Network Load Balancer. The fleet must keep processing jobs even if an entire Availability Zone becomes unavailable, and the architect wants to minimize manual intervention. Which combination of actions should the architect take to meet these requirements?

A.Launch the instances as a Spot Fleet with a capacity-optimized allocation strategy and attach an Application Load Balancer in front of the fleet.
B.Create an Auto Scaling group that spans the three Availability Zones, enable ELB health checks, and configure the group to replace unhealthy instances automatically.
C.Configure the Network Load Balancer with cross-zone load balancing disabled and register all instances in a single target group.
D.Place all instances in a single Availability Zone and create an Amazon Route 53 latency routing policy pointing at the Network Load Balancer.
AnswerB

An Auto Scaling group spanning all three Availability Zones keeps capacity balanced across zones and, with ELB health checks enabled, automatically terminates and replaces instances that fail load balancer health checks. If one Availability Zone fails, the group launches replacement capacity in the remaining zones without operator involvement, which is exactly the resilience the scenario requires.

Why this answer

The requirement is automatic recovery from the loss of a whole Availability Zone. An Auto Scaling group distributed across three zones, combined with ELB health checks, continuously replaces unhealthy instances and rebalances capacity into surviving zones. That removes the need for manual intervention.

The other configurations either concentrate risk in one zone, rely on interruptible capacity, or alter load balancing behavior without providing replacement capacity.

Exam trap

The trap here is assuming that a load balancer by itself provides high availability, when in fact it only distributes traffic and cannot restore compute capacity after an Availability Zone fails.

121
MCQeasy

An event consumer sometimes processes the same SQS message more than once due to timeouts and retries. The consumer must ensure the payment is not charged twice. What design choice best addresses this requirement?

A.Assume messages are processed exactly once because SQS uses durable storage.
B.Make the payment operation idempotent by using an idempotency key and skipping side effects when the key indicates the payment already succeeded.
C.Increase the consumer visibility timeout to several days so messages are not redelivered.
D.Delete the message immediately even if processing fails validation.
AnswerB

Idempotency ensures that repeated processing attempts produce the same result. The consumer should use a stable idempotency key (for example, a business transaction ID) and record completion in durable storage. If the key already indicates the payment succeeded, the consumer skips charging again.

Why this answer

Making the payment operation idempotent using an idempotency key ensures that even if the same SQS message is processed multiple times due to timeouts and retries, the payment will only be charged once. The consumer checks the idempotency key before executing the payment; if the key indicates the payment already succeeded, the consumer skips the side effect. This pattern directly addresses the requirement of not charging twice without relying on SQS's at-least-once delivery guarantee.

Exam trap

The trap here is that candidates assume SQS provides exactly-once delivery or that increasing the visibility timeout is a reliable solution, but the exam tests understanding that SQS is at-least-once and that idempotency is the correct architectural pattern to handle duplicates.

How to eliminate wrong answers

Option A is wrong because SQS guarantees at-least-once delivery, not exactly-once processing; messages can be duplicated due to network issues or consumer timeouts, so assuming exactly-once processing is incorrect. Option C is wrong because increasing the visibility timeout to several days does not prevent redelivery; it only delays it, and if the consumer crashes or fails to delete the message, it will still be redelivered after the timeout expires. Option D is wrong because deleting a message immediately even if processing fails validation means the message is lost permanently, preventing any retry or dead-letter queue handling, which can lead to data loss or incomplete processing.

122
MCQmedium

A inventory service uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The architecture review board prefers a managed AWS-native control.

A.Lambda reserved concurrency set to zero
B.A Lambda dead-letter queue or failure destination
C.A larger deployment package
D.CloudFront error pages
AnswerB

Configuring a Lambda dead-letter queue (DLQ) or an asynchronous failure destination is the correct solution because Lambda will automatically send events that have exhausted built-in retries to a specified SQS queue, SNS topic, or other destination. This preserves the failed event payload and metadata, enabling you to inspect, replay, or alert on the failure. For asynchronous invocations, this is the standard pattern for building resilient, observable serverless workflows.

Why this answer

Lambda dead-letter queues (DLQs) or failure destinations are the correct AWS-native mechanism to retain failed events after retries are exhausted. When a Lambda function fails to process an event (e.g., due to an unreliable third-party API), the function can be configured to send the failed event payload to an SQS queue or SNS topic for later investigation. This ensures no data loss and aligns with the requirement for a managed, AWS-native solution.

Exam trap

The trap here is that candidates often confuse Lambda DLQs with SQS DLQs or assume that increasing retries (via reserved concurrency or package size) solves the retention problem, but the key is the explicit configuration to capture events after retries are exhausted.

How to eliminate wrong answers

Option A is wrong because setting Lambda reserved concurrency to zero would prevent the function from executing at all, not retain failed events. Option C is wrong because a larger deployment package has no impact on error handling or event retention; it only affects cold start times and deployment size. Option D is wrong because CloudFront error pages are for HTTP-level errors in front of web applications, not for Lambda function invocation failures or event retention.

123
MCQmedium

A media company runs a transcoding fleet on Amazon EC2 instances behind an Application Load Balancer in a single Availability Zone. The business requires the workload to survive the loss of that Availability Zone with no manual intervention and minimal downtime. The instances store intermediate files on instance store volumes and the fleet is managed by an Auto Scaling group. Which change should a solutions architect make to meet the requirement?

A.Increase the desired capacity of the Auto Scaling group so that more instances run in the existing Availability Zone.
B.Convert the instances to Spot Instances and configure a placement group to keep them close together.
C.Add a second Availability Zone to the Auto Scaling group and configure the Application Load Balancer with subnets in both Availability Zones.
D.Enable detailed CloudWatch monitoring and create an alarm that notifies operators when an instance becomes unhealthy.
AnswerC

Distributing the Auto Scaling group across two Availability Zones and attaching the load balancer to subnets in both zones lets the fleet continue serving traffic if one zone fails. The load balancer health checks remove unhealthy instances automatically, and the group replaces them, so the workload survives the loss without manual action.

Why this answer

Resilience to an Availability Zone failure requires spreading compute across multiple zones and giving the load balancer subnets in each of those zones. The Auto Scaling group then replaces failed instances in the surviving zone while the load balancer routes only to healthy targets, delivering continuous service without human action.

Exam trap

The trap here is assuming that adding more instances inside one Availability Zone provides high availability, when zone-level isolation is what actually protects the workload.

124
MCQmedium

A media company runs a video transcoding pipeline on Amazon EC2 instances in a single Availability Zone. The pipeline writes intermediate files to an Amazon EBS volume attached to each instance. The company needs the pipeline to survive the failure of any single Availability Zone and to recover automatically with minimal data loss. Which change should a solutions architect make?

A.Create an Auto Scaling group that spans multiple Availability Zones and configure each instance to store intermediate files on an instance store volume.
B.Create an Auto Scaling group that spans multiple Availability Zones and use Amazon EBS Multi-Attach volumes shared across all instances.
C.Create an Auto Scaling group in one Availability Zone and take scheduled EBS snapshots every hour to another Availability Zone.
D.Create an Auto Scaling group that spans multiple Availability Zones and store intermediate files in an Amazon S3 bucket.
AnswerD

Amazon S3 is a regional service with high durability, so intermediate files survive an Availability Zone failure. An Auto Scaling group spanning multiple Availability Zones replaces failed instances automatically, and new instances can read the intermediate files from the S3 bucket. This meets both the resilience and minimal data loss requirements.

Why this answer

Storing intermediate files in Amazon S3 decouples the pipeline from any single Availability Zone because S3 is a regional, highly durable service. An Auto Scaling group spanning multiple Availability Zones automatically replaces instances when a zone fails, and replacement instances retrieve the intermediate files from S3, minimizing data loss and manual recovery effort.

Exam trap

The trap here is assuming that an Auto Scaling group spanning multiple Availability Zones alone provides resilience, when the storage layer must also be decoupled from a single zone.

125
MCQmedium

A inventory service uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured?

A.Lambda reserved concurrency set to zero
B.A Lambda dead-letter queue or failure destination
C.A larger deployment package
D.CloudFront error pages
AnswerB

Configuring a dead-letter queue (an SQS queue or SNS topic) or an asynchronous failure destination causes Lambda to route the event to that target after the configured retry attempts are exhausted. The original event payload is preserved in the DLQ or delivered to a destination such as SQS, SNS, EventBridge, or another Lambda function, allowing operators to inspect and later reprocess the failed inventory events. This is the correct mechanism for capturing failed asynchronous invocations for analysis.

Why this answer

Lambda dead-letter queues (DLQs) or failure destinations are the correct mechanism to retain failed events after all retries are exhausted. When a Lambda function fails to process an event (e.g., from an asynchronous invocation), the service automatically retries twice. If those retries fail, the event can be sent to an SQS queue or SNS topic (DLQ) or to a specified destination (failure destination) for later investigation.

This ensures no data loss and provides a durable storage for post-mortem analysis.

Exam trap

The trap here is that candidates may confuse DLQs with retry mechanisms or think that increasing function resources (like memory or package size) will prevent failures, when in fact DLQs are the only way to durably capture events after retries are exhausted.

How to eliminate wrong answers

Option A is wrong because setting reserved concurrency to zero would prevent the Lambda function from executing at all, not retain failed events. Option C is wrong because a larger deployment package does not affect error handling or event retention; it only increases cold start latency and storage overhead. Option D is wrong because CloudFront error pages are for HTTP-level errors from a web distribution, not for capturing asynchronous Lambda invocation failures.

126
MCQmedium

Based on the exhibit, the application team wants the database to keep the same connection endpoint during failover and to reconnect automatically after the primary instance becomes unavailable. Which change best meets the requirement?

A.Keep the IP address and increase the JDBC connection timeout so the application waits longer during failover.
B.Replace the IP address with the RDS DNS endpoint and add client retry logic that re-resolves DNS after connection loss.
C.Create an additional read replica and point the application to it so failover is faster.
D.Place a Network Load Balancer in front of the database and use the load balancer target IP to avoid DNS changes.
AnswerB

RDS Multi-AZ failover preserves the database endpoint name, not the underlying IP address. When the standby is promoted, AWS updates the DNS record to point to the new primary. Using the RDS endpoint allows the application to follow that change, and retry logic helps the client recover from the short disconnect that occurs during failover.

Why this answer

Using the RDS DNS endpoint ensures that the application connects to the current primary instance, even after a failover. When the primary becomes unavailable, RDS promotes a standby (or read replica) to a new primary and updates the DNS record to point to the new instance's IP. By adding client retry logic that re-resolves DNS after a connection loss, the application automatically picks up the new IP and reconnects without manual intervention, meeting both requirements of a stable endpoint and automatic reconnection.

Exam trap

The trap here is that candidates assume a static IP or a load balancer can provide a stable endpoint, but AWS RDS does not support static IPs for Multi-AZ failover, and NLB cannot front RDS instances—the only reliable way is to use the RDS DNS endpoint with retry logic that re-resolves DNS after a connection loss.

How to eliminate wrong answers

Option A is wrong because keeping the IP address is unreliable—after a failover, the new primary instance will have a different IP address, so the application would connect to a stale IP and fail. Increasing the JDBC connection timeout only delays the failure; it does not resolve the underlying IP mismatch. Option C is wrong because creating an additional read replica does not change the connection endpoint for the primary; the application still connects to the original primary endpoint, which becomes unavailable during failover.

Read replicas are for read scaling, not for providing a failover endpoint. Option D is wrong because placing a Network Load Balancer in front of an RDS database is not a supported architecture—RDS does not integrate with NLB for database traffic, and the load balancer target IP would still change after failover, requiring DNS re-resolution anyway, making the solution unnecessarily complex and non-compliant with AWS best practices.

127
MCQmedium

A content publishing system uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The architecture review board prefers a managed AWS-native control.

A.Lambda reserved concurrency set to zero
B.A larger deployment package
C.CloudFront error pages
D.A Lambda dead-letter queue or failure destination
AnswerD

A dead-letter queue (SQS/SNS) or a failure destination is the correct mechanism because Lambda, after exhausting its default two retries, can route the original event payload to a configured target. A DLQ preserves the raw event and allows a separate process to consume, inspect, and reprocess it, while a failure destination offers richer metadata, such as the request ID and response context, and can send to SQS, SNS, Lambda, or EventBridge. This gives a durable record of failed async invocations and decouples error handling from the main processing function.

Why this answer

Lambda dead-letter queues (DLQs) or failure destinations are the managed AWS-native way to capture events that have exhausted all retry attempts from an asynchronous invocation. When the Lambda function fails after the configured number of retries (default 3), the event is automatically sent to an SQS queue or SNS topic (DLQ) or to a specified destination (e.g., SQS, SNS, EventBridge) for later investigation and reprocessing.

Exam trap

The trap here is that candidates may confuse Lambda's synchronous invocation retry behavior (which is controlled by the caller) with asynchronous invocation retries (which are managed by Lambda itself and require a DLQ or failure destination for post-retry capture).

How to eliminate wrong answers

Option A is wrong because setting reserved concurrency to zero would prevent the Lambda function from executing at all, not handle failed events after retries. Option B is wrong because a larger deployment package does not affect retry or failure handling; it only increases the function's code size and cold start latency. Option C is wrong because CloudFront error pages are for HTTP-level errors from a web distribution, not for capturing failed asynchronous Lambda invocations from a third-party API call.

128
MCQhard

Based on the exhibit, the application tier is not replacing unhealthy instances even though the Auto Scaling group spans two Availability Zones. What change most directly improves automatic recovery when the application process fails?

A.Increase the ASG desired capacity so that extra instances absorb the failed ones.
B.Set the Auto Scaling group health check type to ELB so target group health determines replacement.
C.Replace the Application Load Balancer with a Network Load Balancer to improve failover speed.
D.Increase the HealthCheckGracePeriod to the maximum value so the instances have more time to stabilize.
AnswerB

This makes Auto Scaling replace instances that fail the load balancer health check even when EC2 status checks still pass. The exhibit shows the application health endpoint returns 500 while EC2 checks remain passing, so EC2-only health checks miss the failure. ELB-based health checks align replacement with real application availability.

Why this answer

Setting the Auto Scaling group health check type to ELB allows the ASG to use the target group's health checks, which monitor application-level health (e.g., HTTP 200 responses). When the application process fails, the ELB marks the instance as unhealthy, and the ASG immediately terminates and replaces it. This directly addresses the issue of unhealthy instances not being replaced, as the default EC2 health check only verifies instance status (e.g., running vs. stopped), not application responsiveness.

Exam trap

The trap here is that candidates assume the default EC2 health check is sufficient for application-level failures, but it only checks instance state (running/stopped), not the application process, so the ASG never triggers replacement for application crashes.

Why the other options are wrong

A

Increasing desired capacity adds more instances but does not fix the health check configuration; the ASG still uses EC2 status checks, which may not detect application-level failures, so unhealthy instances are not replaced.

C

The question is about replacing unhealthy instances based on application process failure, not about failover speed. An NLB does not provide application-level health checks, so it would not detect application process failures.

D

Increasing HealthCheckGracePeriod only delays the start of health checks, but does not fix the root cause: the ASG is using EC2 status checks (default) instead of ELB health checks, so it never detects application-level failures.

129
Multi-Selecthard

A payments API requires point-in-time recovery and accidental-delete protection for a DynamoDB table. Which two settings should the architect enable? The architecture review board prefers a managed AWS-native control.

Select 2 answers
A.Deletion protection or tightly controlled delete permissions
B.Point-in-time recovery
C.Global secondary indexes
D.DAX
AnswersA, B

Deletion protection blocks table deletion outright, satisfying the accidental-delete requirement with a managed AWS-native control needing no custom tooling. Pair it with point-in-time recovery for continuous backups, enabling restore to any second within the 35-day window and meeting the low RPO.

Why this answer

Point-in-time recovery (PITR) enables continuous backups of the DynamoDB table, allowing restoration to any point within the last 35 days, which satisfies the requirement for point-in-time recovery. Deletion protection prevents accidental deletion of the table by blocking drop-table operations, meeting the accidental-delete protection requirement. Both are managed AWS-native controls that require no custom scripting or external tooling.

Exam trap

The trap here is that candidates often confuse operational features like DAX (caching) or GSIs (indexing) with data protection mechanisms, but neither provides backup/restore or deletion safeguards required for resilience and data durability.

130
MCQmedium

A trading dashboard uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered?

A.A single-AZ Aurora cluster
B.Aurora Global Database
C.Manual snapshots copied monthly
D.An ElastiCache Redis replica
AnswerB

Aurora Global Database replicates storage across Regions with typical lag under one second, giving low RPO and rapid promotion of a secondary Region to primary during disaster recovery. Single-Region Aurora replicas cannot deliver cross-Region failover at that recovery speed.

Why this answer

Aurora Global Database is designed for cross-Region disaster recovery with a typical RPO of 1 second and RTO of less than 1 minute, using storage-based replication that does not impact database performance. This meets the low RPO requirement for a trading dashboard, where data loss must be minimized.

Exam trap

The trap here is that candidates might choose manual snapshots (Option C) thinking they are sufficient for DR, but they overlook the critical requirement of low RPO, which snapshots copied monthly cannot satisfy.

How to eliminate wrong answers

Option A is wrong because a single-AZ Aurora cluster provides no cross-Region replication and offers no disaster recovery across AWS Regions, resulting in potentially high RPO if the primary Region fails. Option C is wrong because manual snapshots copied monthly have an RPO of up to one month, which is far too high for a trading dashboard requiring low RPO. Option D is wrong because ElastiCache Redis is an in-memory cache, not a persistent database, and cannot serve as a cross-Region disaster recovery solution for Aurora MySQL data.

131
MCQeasy

A startup runs a stateless web application on a single Amazon EC2 instance in one Availability Zone. The application has become popular, and the startup wants to ensure that the application can survive the failure of an Availability Zone and can handle increased traffic. Which architecture change should the startup make FIRST?

A.Resize the existing EC2 instance to a larger instance type to handle increased traffic and improve availability.
B.Create an Amazon Machine Image (AMI) of the instance and store it in Amazon S3 for quick recovery after a failure.
C.Enable detailed monitoring on the EC2 instance and create Amazon CloudWatch alarms to notify the team of failures.
D.Move the application to an Auto Scaling group that spans multiple Availability Zones, and place an Application Load Balancer in front of it.
AnswerD

An Auto Scaling group spanning multiple Availability Zones distributes instances across AZs, so the application survives an AZ failure. An Application Load Balancer distributes traffic and performs health checks, routing only to healthy instances. This provides both high availability and horizontal scalability, directly addressing the requirements with minimal changes.

Why this answer

Moving to an Auto Scaling group across multiple Availability Zones with an Application Load Balancer provides both high availability and elasticity. The load balancer health checks ensure traffic goes only to healthy instances, and the Auto Scaling group replaces failed instances automatically. This architecture eliminates single points of failure and allows the application to scale with demand, which is the fundamental first step for resilience.

Exam trap

The trap here is focusing on vertical scaling or monitoring as a way to achieve high availability, when the real solution requires distributing the workload across multiple Availability Zones.

132
MCQmedium

An events service publishes critical notifications using Amazon SNS. Three independent downstream systems (A, B, and C) subscribe to the topic. Downstream system B sometimes fails to process certain messages (for example, it times out or returns an error while handling the message), and you want: 1) failures in B to be isolated so A and C keep processing unaffected, and 2) messages that B cannot successfully process after retries to be sent to a DLQ for B. Which design best meets these requirements?

A.Subscribe each downstream directly with HTTPS endpoints and configure a single SNS dead-letter queue (DLQ) for the topic.
B.For each downstream system, create its own SQS queue, subscribe each SQS queue to the SNS topic, and configure a redrive policy with a DLQ for each SQS queue.
C.Use one shared SQS queue for all three downstream systems and configure a single DLQ only when all three downstream systems fail.
D.Use EventBridge rules to invoke A, B, and C synchronously with retries enabled, and send failures to a common DLQ.
AnswerB

SNS delivers the message independently to each subscribed SQS queue. If downstream B fails to process a message, B can avoid deleting it from its own queue; after visibility timeout and retry attempts, SQS redrives messages to B’s DLQ. A and C are isolated because they have separate queues and DLQs, so B’s failures do not prevent deliveries to A and C.

Why this answer

It creates a dedicated SQS queue for each downstream system, which isolates failures: if system B fails, its SQS queue will accumulate messages while systems A and C continue processing from their own queues. Each SQS queue can have a redrive policy that moves messages to a per-queue DLQ after the configured maximum retries are exhausted, satisfying the requirement for a B-specific DLQ without affecting the other subscribers.

Exam trap

The trap here is that candidates assume a single DLQ at the SNS topic level is sufficient, but SNS DLQs only apply to the SNS delivery failure (e.g., HTTP endpoint unreachable), not to downstream processing failures after the message is delivered to SQS.

How to eliminate wrong answers

Option A is wrong because a single SNS DLQ applies to the entire topic, not per-subscriber; if B fails, messages would be sent to the common DLQ for all subscribers, and A and C would still receive the message from SNS, but the DLQ is not isolated to B. Option C is wrong because a shared SQS queue for all three systems means a failure in B could block or delay messages for A and C, and a single DLQ would trigger only when all three fail, not when B alone fails. Option D is wrong because EventBridge synchronous invocation with a common DLQ would cause failures in B to potentially block or delay A and C (since synchronous calls are sequential), and the DLQ is shared, not isolated to B.

133
MCQmedium

A stateless web API runs on EC2 instances behind an Application Load Balancer (ALB). The Auto Scaling group (ASG) currently uses subnets from only one Availability Zone, even though the ALB spans two Availability Zones. During maintenance of that single AZ, the ALB remains up but clients see timeouts because there are no healthy targets. Which change most directly improves resilience against an AZ failure?

A.Keep the ASG in one subnet/AZ, but enable ALB stickiness to reduce session interruption.
B.Update the ASG to launch instances across subnets in at least two Availability Zones and ensure ALB health checks target an application-ready path.
C.Add a NAT gateway in the public subnets so instances can reach the internet during maintenance events.
D.Create a second ALB in the same Availability Zone and route traffic using DNS failover.
AnswerB

Spreading the ASG across subnets in two Availability Zones removes the single-AZ failure domain, so the ALB can route to healthy targets in the surviving zone during maintenance. Pointing health checks at an application-ready path ensures instances register only when genuinely serving traffic, preventing the timeouts seen previously.

Why this answer

The most direct fix for AZ failure resilience is to distribute the ASG across multiple Availability Zones. With the ALB already spanning two AZs, if the ASG only launches instances in one AZ, a failure of that AZ leaves the ALB with zero healthy targets, causing timeouts. By configuring the ASG to launch instances in at least two AZs and setting ALB health checks to an application-ready path, the ALB can route traffic to healthy instances in the surviving AZ, maintaining availability.

Exam trap

The trap here is that candidates may think adding a second ALB or enabling stickiness solves the problem, when the real issue is that the ASG is not distributing instances across multiple Availability Zones, leaving the ALB with no healthy targets during an AZ outage.

Why the other options are wrong

A

Enabling ALB stickiness does not address the root cause: the ASG has no healthy targets in the only AZ, so the ALB cannot route traffic to any instance, causing timeouts regardless of stickiness.

C

The issue is that EC2 instances are only in one AZ, so adding a NAT gateway does not provide healthy targets in the other AZ; NAT gateways enable outbound internet access but do not affect ALB target availability.

D

Creating a second ALB in the same AZ does not address the lack of healthy targets in the other AZ; the ALB is already spanning two AZs, but the ASG only has instances in one AZ. DNS failover would still route to the same AZ if both ALBs are in the same AZ, failing to provide resilience against an AZ failure.

134
MCQmedium

An orders service publishes payment instructions to an Amazon SQS Standard queue. A downstream consumer sometimes times out or crashes after it has partially completed processing, causing the same instruction to be processed more than once. You must keep the design resilient without attempting to guarantee exactly-once processing. Which approach best handles duplicates safely?

A.Set the SQS visibility timeout extremely long so the message cannot be retried even after processing failures.
B.Make the consumer idempotent by deriving a deterministic idempotency key from the payment instruction (for example, the instruction ID), persisting the result of successful processing, and skipping re-processing when that key is already marked successful.
C.Switch to an SQS FIFO queue but remove error handling in the consumer so duplicates never occur.
D.Send all failed messages to a DLQ and rely on it to deduplicate messages that were already successfully processed.
AnswerB

SQS Standard provides at-least-once delivery, so duplicates are expected. Idempotency ensures that re-processing the same instruction does not create incorrect side effects. Persisting a deterministic key/result allows the consumer to safely short-circuit duplicates after retries/timeouts.

Why this answer

Making the consumer idempotent ensures that even if the same payment instruction is processed multiple times due to timeouts or crashes, the system remains consistent. By deriving a deterministic idempotency key (e.g., the instruction ID) and persisting the result of successful processing, the consumer can skip re-processing when the key is already marked as successful. This approach aligns with the requirement to keep the design resilient without guaranteeing exactly-once processing, as it safely handles duplicates at the application level.

Exam trap

The trap here is that candidates often assume SQS FIFO queues or DLQs inherently solve duplicate processing, but the exam tests understanding that Standard queues require application-level idempotency for safe duplicate handling, and FIFO queues do not eliminate the need for idempotent consumers in crash scenarios.

How to eliminate wrong answers

Option A is wrong because setting the SQS visibility timeout extremely long does not prevent duplicates; it only delays retries, and if the consumer crashes after partially processing, the message will eventually become visible again and be reprocessed, leading to duplicates. Option C is wrong because switching to an SQS FIFO queue provides exactly-once processing within a five-minute deduplication window, but removing error handling in the consumer does not prevent duplicates from timeouts or crashes; FIFO queues still allow retries, and without error handling, the system becomes fragile. Option D is wrong because a Dead-Letter Queue (DLQ) is used to capture messages that fail after multiple retries, not to deduplicate messages; relying on a DLQ for deduplication is a misconception, as DLQs do not track successful processing and cannot prevent duplicates from being processed again.

135
MCQmedium

A claims workflow uses an RDS MySQL database and must remain available during an Availability Zone failure with minimal application changes. What should the architect enable? The design must avoid adding custom operational scripts.

A.S3 Cross-Region Replication
B.Multi-AZ deployment for the RDS DB instance
C.EBS snapshots every hour
D.Read replicas only
AnswerB

A Multi-AZ deployment is the correct RDS feature for availability because Amazon RDS automatically provisions and maintains a synchronous standby replica in a different Availability Zone within the same Region. The primary DB instance writes data synchronously to the standby before committing, and if the primary fails or its AZ becomes unavailable, RDS automatically performs failover to the standby by updating the DNS endpoint, typically within 60–120 seconds without requiring manual intervention. This eliminates the need for application-level failover logic and provides a clear improvement in availability for an RDS MySQL instance.

Why this answer

Multi-AZ deployment for RDS MySQL provides automatic failover to a standby replica in a different Availability Zone. This ensures high availability during an AZ failure with minimal application changes, as the DNS endpoint remains the same and failover is handled by AWS without custom scripts.

Exam trap

The trap here is that candidates often confuse read replicas with Multi-AZ deployments, assuming read replicas provide automatic failover, but they require manual promotion and do not maintain the same endpoint.

How to eliminate wrong answers

Option A is wrong because S3 Cross-Region Replication is for object storage replication across regions, not for database availability within a region, and it does not address RDS MySQL failover. Option C is wrong because EBS snapshots every hour provide point-in-time backups but do not enable automatic failover or maintain availability during an AZ failure; recovery would require manual intervention and data loss. Option D is wrong because read replicas are for read scaling and do not provide automatic failover for the primary instance; promoting a read replica requires manual steps or custom scripts, violating the 'no custom operational scripts' constraint.

136
MCQhard

Based on the exhibit, the team must restore an Amazon RDS for PostgreSQL database to the exact state just before a bad delete happened. What is the best recovery approach?

A.Restore the latest automated snapshot and accept data loss from the last backup window.
B.Perform a point-in-time restore to 2026-04-27 15:10 UTC into a new DB instance, then cut over after validation.
C.Promote a read replica because it will contain the deleted rows and can replace the primary immediately.
D.Enable Multi-AZ on the current database and wait for automatic failover to reverse the delete.
AnswerB

Point-in-time restore uses the automated backups and transaction logs to rebuild the database to an exact time before the bad change. The exhibit confirms the requested restore time is within the restorable window, and the business wants to validate the restored copy before switching traffic. Restoring to a new instance first is the safest way to recover without risking the current production database.

Why this answer

Point-in-time recovery (PITR) allows you to restore an Amazon RDS for PostgreSQL database to any second within the backup retention period, using automated backups and transaction logs. By restoring to 2026-04-27 15:10 UTC, just before the bad delete occurred, you can recover the exact state without data loss, then cut over after validation.

Exam trap

The trap here is that candidates often confuse read replicas or Multi-AZ as solutions for logical data corruption, when in fact they only protect against infrastructure failures, not user errors like a bad delete.

Why the other options are wrong

A

Restoring the latest automated snapshot would not recover to the exact state just before the bad delete at 2026-04-27 15:10 UTC; it would only restore to the last snapshot time, which could be hours earlier, causing data loss beyond the deleted rows.

C

A read replica is asynchronous and may not contain the exact state before the delete; it cannot be promoted to reverse a specific point-in-time deletion without data loss or inconsistency.

137
MCQhard

A patient portal must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable?

A.Use an in-memory queue on one EC2 instance
B.Use UDP messages sent directly to workers
C.Use Amazon SQS standard queue and design consumers to be idempotent
D.Use CloudFront signed URLs
AnswerC

Amazon SQS standard queues provide durable storage across multiple Availability Zones and at-least-once delivery, meaning every message is guaranteed to be delivered eventually, but duplicates can occasionally occur due to the distributed architecture. Consumers must be idempotent so that processing the same event multiple times has the same effect as processing it once, preventing duplicate medical records or actions. This design satisfies the requirement while maintaining the high throughput needed by a patient portal.

Why this answer

Amazon SQS standard queues provide at-least-once delivery, meaning each message is delivered at least once but may occasionally be delivered more than once. This aligns with the requirement that every event must be processed at least once, and since duplicate processing is acceptable when consumers are idempotent, the standard queue is the most suitable choice. SQS handles the decoupling and durability of messages without requiring custom infrastructure.

Exam trap

The trap here is that candidates may confuse 'at-least-once' with 'exactly-once' and incorrectly choose a FIFO queue or another option, but the question explicitly accepts duplicates if the consumer handles idempotency, making the standard queue the correct choice.

How to eliminate wrong answers

Option A is wrong because an in-memory queue on a single EC2 instance is not durable, cannot survive instance failures, and does not provide at-least-once delivery guarantees across distributed consumers. Option B is wrong because UDP is a connectionless, unreliable protocol that does not guarantee message delivery, ordering, or duplicate detection, making it unsuitable for at-least-once processing. Option D is wrong because CloudFront signed URLs are used for secure content delivery and access control, not for event messaging or queue-based processing.

138
MCQeasy

A worker consumes messages from an Amazon SQS queue. Some messages consistently fail validation and are retried until the worker can no longer process them. What is the most appropriate AWS mechanism to handle these poison messages while keeping the queue usable?

A.Enable SQS long polling and increase the maximum message size for the queue.
B.Send failing messages to an SQS dead-letter queue (DLQ) using a redrive policy based on receive count.
C.Change the queue to a FIFO queue and handle duplicates in the worker code without DLQs.
D.Delete the queue and recreate it hourly to clear out any problematic messages.
AnswerB

A DLQ with a redrive policy isolates poison messages. After a message is received and fails processing more than the configured maxReceiveCount, SQS moves it to the DLQ, preventing it from continually blocking retries in the source queue.

Why this answer

An SQS dead-letter queue (DLQ) with a redrive policy based on receive count allows messages that repeatedly fail processing (poison pills) to be moved out of the main queue after a specified number of retries. This keeps the main queue operational for valid messages and isolates problematic messages for later analysis or manual intervention.

Exam trap

The trap here is that candidates may think increasing retries or message size (Option A) solves the problem, but the exam specifically tests the concept of isolating poison messages via a DLQ with a receive-count-based redrive policy to maintain queue availability.

How to eliminate wrong answers

Option A is wrong because enabling long polling and increasing maximum message size does not address the core issue of messages that consistently fail validation; long polling reduces empty responses and larger message size allows bigger payloads, but neither prevents poison messages from blocking processing. Option C is wrong because changing to a FIFO queue does not inherently handle poison messages; FIFO queues preserve order and deduplicate based on message deduplication ID, but they still require a DLQ or explicit error handling to remove failing messages, and the worker code alone cannot prevent retries from exhausting resources. Option D is wrong because deleting and recreating the queue hourly is a disruptive, non-scalable approach that loses all messages (including valid ones) and does not provide a mechanism to isolate or analyze poison messages; it also violates the requirement to keep the queue usable.

139
MCQmedium

A trading dashboard runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include? The team wants the control to be enforceable during normal operations.

A.A single EC2 instance with detailed monitoring
B.Subnets in at least two Availability Zones with health checks enabled
C.All instances in one larger subnet
D.A Network Load Balancer in one subnet
AnswerB

Deploying subnets (and therefore EC2 instances) in at least two Availability Zones removes any single AZ as a point of failure. An Auto Scaling group distributed across those subnets works with a load balancer's health checks to detect unhealthy instances, terminate them, and launch replacement capacity in a healthy AZ. If one AZ fails entirely, the ASG can continue to meet desired capacity in the surviving AZ and the load balancer routes traffic away from unhealthy targets, keeping the dashboard available.

Why this answer

Distributing EC2 instances across at least two Availability Zones (AZs) ensures that if one AZ fails, the Auto Scaling group can maintain capacity in the remaining AZ(s). Enabling health checks allows the group to detect instance failures and automatically replace them, providing fault tolerance. This configuration meets the requirement to tolerate a single AZ failure while remaining enforceable during normal operations.

Exam trap

The trap here is that candidates often confuse high availability (spanning multiple AZs) with fault tolerance at the instance level, mistakenly thinking a single instance with monitoring or a single subnet can survive an AZ failure.

How to eliminate wrong answers

Option A is wrong because a single EC2 instance, even with detailed monitoring, cannot tolerate the failure of an entire Availability Zone; if that AZ goes down, the instance becomes unavailable. Option C is wrong because placing all instances in one larger subnet within a single AZ creates a single point of failure; an AZ failure would take down all instances. Option D is wrong because a Network Load Balancer in one subnet does not provide AZ-level fault tolerance; it still relies on that single AZ, and the Auto Scaling group must span multiple AZs for resilience.

140
MCQmedium

Based on the exhibit, the company wants DNS traffic to fail over automatically from the primary Region to a secondary Region when the primary endpoint is unhealthy. Which Route 53 change is best?

A.Keep simple routing and lower the TTL to 10 seconds.
B.Use weighted routing with equal weights for both ALBs.
C.Use geolocation routing so users in each continent reach a closer ALB.
D.Create Route 53 failover records with health checks for the primary and secondary ALBs.
AnswerD

Failover routing is the Route 53 policy intended for this use case. Route 53 returns the primary record while its health check passes, and automatically serves the secondary record when the primary health check fails. That provides DNS-based Regional failover without manual intervention.

Why this answer

Route 53 failover routing with health checks is the only option that automatically directs DNS traffic away from an unhealthy primary endpoint to a healthy secondary endpoint. When the health check for the primary ALB fails, Route 53 returns the secondary ALB's IP address in DNS responses, providing automatic failover across regions. Simple, weighted, and geolocation routing do not natively support automatic failover based on endpoint health.

Exam trap

The trap here is that candidates often confuse weighted routing with failover, assuming equal weights will somehow cause automatic failover, but weighted routing does not consider health status and requires manual intervention to shift traffic.

Why the other options are wrong

A

Simple routing does not support health checks or automatic failover; lowering TTL only speeds up DNS propagation but does not enable failover to a secondary endpoint when the primary is unhealthy.

B

Weighted routing distributes traffic based on weights, not health; it does not automatically failover to a healthy endpoint when the primary is unhealthy.

C

Geolocation routing directs traffic based on the user's geographic location, not health. It does not automatically failover to a secondary region when the primary endpoint is unhealthy; it would continue sending traffic from that region to the unhealthy endpoint.

141
MCQhard

A financial analytics platform stores results in an Amazon S3 bucket. Compliance requires that objects be recoverable for 30 days after deletion and that no user, including administrators, be able to permanently erase them during that window. Objects must also remain readable throughout the retention period. Which approach should the architect implement?

A.Enable S3 Object Lock in governance mode with a 30-day retention period on the bucket.
B.Enable S3 Versioning and add a bucket policy that denies s3:DeleteObject to all principals.
C.Configure an S3 Lifecycle rule to transition objects to S3 Glacier Instant Retrieval after one day and expire them after 30 days.
D.Enable S3 Object Lock in compliance mode with a 30-day retention period on the bucket.
AnswerD

Compliance mode enforces a write-once-read-many model in which no principal, including the root user, can overwrite or delete a locked object version before the retention date expires. Objects stay readable throughout the period, satisfying the availability clause. A 30-day retention directly matches the stated recovery window, and the protection is enforced by S3 itself rather than by policy that could be changed.

Why this answer

Meeting the requirement means using a control enforced by the storage service that no identity can override. S3 Object Lock in compliance mode provides exactly that: locked object versions cannot be deleted or overwritten by any principal until the retention period lapses, and they remain readable. Governance mode, versioning with policies, and lifecycle expiration all leave a deletion path open for privileged users.

Exam trap

The trap here is treating governance mode as equivalent to compliance mode, when governance mode still permits privileged users to bypass retention.

142
MCQhard

A warehouse integration service must use shared file storage across Linux EC2 instances in multiple Availability Zones. The storage must remain available during an AZ failure. Which service should be used?

A.Amazon EFS with mount targets in multiple Availability Zones
B.S3 mounted as a POSIX file system without a file gateway
C.Instance store volumes
D.An EBS volume attached to all instances
AnswerA

Amazon EFS is a fully managed, regional file storage service that provides a standard POSIX file system interface (NFSv4.1) for EC2 instances. By configuring mount targets in multiple Availability Zones, instances across the VPC can concurrently read and write to the same file system with shared file locking and metadata guarantees. This architecture delivers both high availability and elastic scalability, making it the appropriate shared storage layer for a warehouse integration service that must be accessed by multiple compute resources simultaneously.

Why this answer

Amazon EFS provides a fully managed, scalable, and elastic NFS file system that can be mounted concurrently on multiple Linux EC2 instances across different Availability Zones. By configuring mount targets in each AZ, the file system remains accessible even if one AZ fails, because the other mount targets continue to serve traffic. This meets the requirement for shared, highly available file storage across AZs.

Exam trap

The trap here is that candidates may confuse EBS multi-attach (which has strict limitations and is not suitable for shared file systems across AZs) with a true distributed file system like EFS, or assume that S3 with a FUSE driver can replace a POSIX-compliant shared file system.

How to eliminate wrong answers

Option B is wrong because mounting S3 as a POSIX file system without a file gateway (e.g., using s3fs-fuse) does not provide true POSIX semantics (e.g., file locking, atomic operations) and introduces performance and consistency issues; it is not a native shared file system for Linux EC2 instances. Option C is wrong because instance store volumes are ephemeral and tied to a single EC2 instance; they are lost if the instance stops or fails, and cannot be shared across instances or survive an AZ failure. Option D is wrong because a single EBS volume can only be attached to one EC2 instance at a time (multi-attach EBS is limited to specific io1/io2 volumes and is not designed for shared file system workloads across multiple instances in different AZs).

143
MCQmedium

A web application runs on an Auto Scaling group (ASG) behind an Application Load Balancer (ALB). The ASG is currently attached to subnets in only two Availability Zones (AZs). During a planned maintenance window, one AZ becomes unavailable for about 25 minutes. Monitoring shows that targets in the remaining AZ go healthy, and the ALB/target group health checks report normal. However, users still experience intermittent connection failures and slower responses during the AZ outage. What change will most directly improve resilience against an AZ loss while keeping the same ALB-based design?

A.Set the ASG min capacity to 0 so instances can be recreated faster when an AZ recovers.
B.Extend the ASG to use subnets in three AZs so there is placement redundancy during an AZ outage, while continuing to keep traffic behind the ALB.
C.Increase the ALB idle timeout to 120 seconds to reduce connection drops.
D.Disable health checks on the target group so instances are not deregistered during the maintenance window.
AnswerB

An AZ outage reduces the number of AZs where the ASG can place instances. With only two AZs, losing one significantly limits capacity and can cause temporary shortages and uneven load distribution, even if existing targets are marked healthy. Expanding the ASG to subnets in three (or more) AZs provides additional placement options so the ASG can maintain the desired number of instances across the remaining AZ(s). The ALB will continue routing only to healthy targets, and the system is more likely to sustain stable response times during the outage.

Why this answer

B is correct because deploying the ASG across three Availability Zones (AZs) ensures that when one AZ becomes unavailable, the remaining two AZs can handle the full traffic load without overloading the instances. This placement redundancy directly addresses the intermittent connection failures and slower responses, as the ALB can distribute traffic only to healthy targets in the remaining AZs, maintaining capacity and performance. The current two-AZ setup lacks sufficient buffer capacity, causing the single remaining AZ to become overwhelmed during the outage.

Exam trap

The trap here is that candidates may focus on connection-level settings (idle timeout) or health check behavior, missing the fundamental architectural need for multi-AZ redundancy to maintain capacity during an AZ outage.

How to eliminate wrong answers

Option A is wrong because setting the ASG min capacity to 0 does not help during an AZ outage; it would actually allow all instances to be terminated, making the application unavailable, and it does not address the lack of capacity in the remaining AZ. Option C is wrong because increasing the ALB idle timeout to 120 seconds only keeps idle connections open longer, which does not prevent connection failures or slow responses caused by insufficient capacity in the remaining AZ; it may even mask underlying issues. Option D is wrong because disabling health checks on the target group would prevent the ALB from deregistering unhealthy instances, causing traffic to be routed to failed instances in the unavailable AZ, leading to more connection failures and no improvement in resilience.

144
MCQmedium

A warehouse integration service receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers? The architecture review board prefers a managed AWS-native control.

A.AWS WAF
B.Amazon Route 53 weighted routing
C.Amazon SQS queue
D.Amazon CloudFront
AnswerC

Amazon SQS is a fully managed message queue that decouples the order ingestion service from the backend workers. Producers send order messages to the queue, and consumers poll messages at their own sustainable rate, which naturally absorbs bursts—the queue itself acts as a temporary, durable buffer. The visibility timeout ensures in-flight messages are not processed by multiple workers, and a dead-letter queue captures poison messages after repeated failures. Because SQS retains messages for up to 14 days and integrates with Auto Scaling via the ApproximateNumberOfMessages metric, it is the correct service for handling unpredictable spikes in order volume.

Why this answer

Amazon SQS is the correct choice because it acts as a fully managed message queue that decouples the web tier from the fulfilment workers, buffering incoming order bursts. It provides at-least-once delivery and allows workers to poll messages at their own pace, ensuring no requests are lost even during spikes. SQS also supports retries via a dead-letter queue (DLQ) for messages that fail processing, meeting the requirement for resilient, managed AWS-native control.

Exam trap

The trap here is that candidates may confuse AWS WAF or CloudFront as tools for handling traffic spikes, but neither provides the decoupling, buffering, and retry capabilities of a queue; they are designed for security and content delivery, respectively, not for asynchronous processing.

How to eliminate wrong answers

Option A is wrong because AWS WAF is a web application firewall that filters HTTP/S traffic based on rules (e.g., SQL injection, XSS) and does not provide message buffering, queuing, or retry logic for downstream services. Option B is wrong because Amazon Route 53 weighted routing distributes DNS traffic across multiple endpoints based on weights, but it does not absorb spikes or provide retry mechanisms; it only controls which endpoint receives a request, and a failed request is lost unless the client retries. Option D is wrong because Amazon CloudFront is a content delivery network (CDN) that caches static and dynamic content at edge locations to reduce latency, but it cannot buffer or retry requests for a downstream fulfilment service; it is designed for accelerating content delivery, not for decoupling or absorbing processing spikes.

145
MCQhard

A patient portal must process every event at least once, but duplicate processing is acceptable if the consumer handles idempotency. Which eventing approach is most suitable? The design must avoid adding custom operational scripts.

A.Use an in-memory queue on one EC2 instance
B.Use UDP messages sent directly to workers
C.Use Amazon SQS standard queue and design consumers to be idempotent
D.Use CloudFront signed URLs
AnswerC

Using an Amazon SQS standard queue meets the at-least-once requirement: SQS stores messages durably across multiple Availability Zones and redelivers any message that is not deleted before the visibility timeout expires, so every event is processed at least once. Because standard queues can occasionally deliver duplicate copies, your consumer design must be idempotent—e.g., by tracking a unique event ID or a deterministic processing key—so that repeated handling of the same event does not corrupt patient data. This combination gives you high throughput, operational simplicity, and no single point of failure, which is why it is the correct choice.

Why this answer

Amazon SQS standard queues guarantee at-least-once delivery, which satisfies the requirement that every event is processed at least once. The design avoids custom operational scripts by leveraging a fully managed service, and the acceptance of duplicate processing is handled by making consumers idempotent. This combination provides a scalable, resilient, and cost-effective event-driven architecture without the need for custom infrastructure management.

Exam trap

The trap here is that candidates may confuse 'at-least-once' delivery with 'exactly-once' delivery and incorrectly choose a solution like a FIFO queue or a custom retry mechanism, but the question explicitly allows duplicate processing, making the standard queue the correct choice.

How to eliminate wrong answers

Option A is wrong because an in-memory queue on a single EC2 instance creates a single point of failure, lacks durability, and requires custom operational scripts for management and recovery, violating the 'avoid adding custom operational scripts' constraint. Option B is wrong because UDP is a connectionless, unreliable protocol that does not guarantee message delivery, so it cannot ensure at-least-once processing; it also requires custom application-level handling for reliability. Option D is wrong because CloudFront signed URLs are used for securing content delivery and controlling access to files, not for event processing or message queuing; they do not provide any event delivery guarantee or queue semantics.

146
Multi-Selecthard

A payments API requires point-in-time recovery and accidental-delete protection for a DynamoDB table. Which two settings should the architect enable? The design must avoid adding custom operational scripts.

Select 2 answers
A.Deletion protection or tightly controlled delete permissions
B.Point-in-time recovery
C.Global secondary indexes
D.DAX
AnswersA, B

Deletion protection is a DynamoDB table attribute that blocks DeleteTable operations until explicitly disabled, while tightly scoped IAM policies that deny dynamodb:DeleteTable reduce the risk of a single misissued CLI command or console action taking down a payments table. Without this guard, an accidental deletion is immediate and permanent, taking all data, indexes, and backup history with it. It is therefore a required complement to backup features.

Why this answer

Deletion protection (option A) prevents accidental deletion of the DynamoDB table itself, which is critical for the accidental-delete protection requirement. Point-in-time recovery (option B) enables restoring the table to any point within the last 35 days, satisfying the point-in-time recovery requirement. Both features are native DynamoDB capabilities that require no custom scripts.

Exam trap

The trap here is that candidates may confuse deletion protection (which protects the table resource) with item-level delete prevention, or think that GSIs or DAX provide data durability or recovery features when they do not.

147
MCQmedium

A company hosts an internal API behind an Application Load Balancer (ALB) in two AWS Regions. They want Amazon Route 53 to automatically fail over to the secondary Region when the primary Region’s ALB is unhealthy. Health checks for the primary ALB are already configured, but the DNS record currently uses a latency-based routing policy. Which Route 53 configuration most directly provides automatic failover based on health status?

A.Keep latency-based routing, and set the weights so the secondary Region rarely receives traffic unless manual changes are made.
B.Use a Route 53 failover routing policy: configure two alias records for the ALBs where the primary record is marked PRIMARY, the secondary is marked SECONDARY, and each record has an associated health check.
C.Use an alias A record that returns both ALBs simultaneously so clients automatically load balance across Regions during outages.
D.Use geolocation routing to route users to the primary Region and rely on ALB health checks to shift requests between Regions.
AnswerB

Route 53 failover routing is designed specifically for active-passive failover: you create two alias records pointing to the two ALBs, one marked PRIMARY and the other SECONDARY, each with an associated health check. Under normal conditions, Route 53 returns the PRIMARY record, but if that record's health check fails, Route 53 automatically stops returning it and returns the SECONDARY record instead. This gives the company an automated, DNS-level failover that does not depend on client latency, manual weight changes, or client-side load balancing.

Why this answer

Route 53 failover routing policy is specifically designed to automatically route traffic away from an unhealthy resource to a healthy one. By creating two alias records (one PRIMARY with an associated health check for the primary ALB, and one SECONDARY for the secondary ALB), Route 53 will automatically fail over to the secondary record when the primary health check fails. This directly meets the requirement for automatic failover based on health status, unlike latency-based routing which only optimizes for response time.

Exam trap

The trap here is that candidates often confuse latency-based routing with failover routing, assuming latency-based routing inherently provides health-based failover, but it only optimizes for latency and does not automatically reroute based on health status.

How to eliminate wrong answers

Option A is wrong because latency-based routing does not support automatic failover based on health status; weights only control traffic distribution and manual changes would be required to shift traffic, which contradicts the 'automatic failover' requirement. Option C is wrong because an alias A record cannot return multiple ALBs simultaneously; Route 53 alias records point to a single AWS resource, and returning multiple IPs would require a non-alias record with multiple values, which still does not provide health-based failover. Option D is wrong because geolocation routing routes based on user location, not health; ALB health checks alone cannot shift requests between Regions because the DNS record itself does not change based on health status without a failover routing policy.

148
MCQmedium

An orders service publishes payment instructions to an Amazon SQS Standard queue. A downstream consumer sometimes times out and retries the work, causing the consumer to process the same instruction more than once. Operationally, the team must ensure that duplicate processing does not create duplicate charges. The queue type cannot be changed. What is the most resilient application-side approach?

A.Rely on SQS Standard to provide exactly-once delivery for each message, since the consumer uses retries.
B.Implement idempotent processing using a persistent deduplication key (for example, paymentInstructionId) so repeated messages are ignored or safely merged.
C.Increase the queue’s visibility timeout to 24 hours so messages never reappear even if the consumer times out.
D.Delete and recreate the queue with a different name whenever duplicates are detected in production.
AnswerB

Because SQS Standard is at-least-once, the consumer must assume duplicates are possible. Persisting a record keyed by paymentInstructionId (or using a database unique constraint) lets the consumer detect that a given instruction was already processed successfully and safely skip the charge or merge results deterministically.

Why this answer

Implementing idempotent processing with a persistent deduplication key (e.g., paymentInstructionId) ensures that even if SQS Standard delivers the same message multiple times due to consumer timeouts and retries, the downstream logic will detect and ignore or safely merge duplicate charges. This is the most resilient application-side approach as it does not rely on queue configuration changes and works within the constraints of SQS Standard's at-least-once delivery model.

Exam trap

The trap here is that candidates often assume SQS Standard can provide exactly-once delivery if retries are handled properly, but the exam tests the understanding that SQS Standard inherently allows duplicates and that idempotency is the only reliable application-side solution.

How to eliminate wrong answers

Option A is wrong because SQS Standard queues provide at-least-once delivery, not exactly-once delivery; retries and timeouts can cause duplicate messages, and relying on exactly-once is a misconception. Option C is wrong because increasing the visibility timeout to 24 hours does not prevent duplicates if the consumer times out and retries before the timeout expires, and it can delay processing unnecessarily, making it impractical and not resilient. Option D is wrong because deleting and recreating the queue with a different name is a disruptive, manual, and non-scalable approach that does not address the root cause of duplicate processing and would cause data loss and operational chaos.

149
MCQmedium

A ticket booking system runs on EC2 instances behind an Application Load Balancer. The design must tolerate the failure of one Availability Zone. What should the Auto Scaling group configuration include? The architecture review board prefers a managed AWS-native control.

A.Subnets in at least two Availability Zones with health checks enabled
B.All instances in one larger subnet
C.A Network Load Balancer in one subnet
D.A single EC2 instance with detailed monitoring
AnswerA

Spreading the Auto Scaling group across subnets in at least two Availability Zones with health checks enabled ensures instances in a surviving AZ absorb traffic when one fails. This satisfies the requirement to tolerate a single AZ failure using AWS-native controls.

Why this answer

An Auto Scaling group configured with subnets in at least two Availability Zones ensures that if one AZ fails, the remaining AZ(s) can continue to serve traffic. Health checks on the EC2 instances allow the Auto Scaling group to detect and replace unhealthy instances, maintaining the desired capacity across the surviving AZs. This aligns with the requirement for a managed AWS-native control to tolerate an AZ failure.

Exam trap

The trap here is that candidates might think a single large subnet or a Network Load Balancer provides AZ resilience, but subnets are AZ-scoped and an NLB is a separate load-balancing component, not an Auto Scaling group configuration setting.

How to eliminate wrong answers

Option B is wrong because placing all instances in one larger subnet, even if it spans multiple AZs (which is not possible as subnets are AZ-specific), does not provide AZ failure tolerance; a single AZ failure would take down all instances. Option C is wrong because a Network Load Balancer (NLB) is not a component of an Auto Scaling group configuration; the question asks what the Auto Scaling group should include, and an NLB is a separate resource, not a configuration setting within the group. Option D is wrong because a single EC2 instance, even with detailed monitoring, cannot tolerate the failure of one Availability Zone; if that instance resides in the failed AZ, the application becomes unavailable, and detailed monitoring does not provide redundancy.

150
Multi-Selecteasy

A service processes messages from an Amazon SQS queue. Sometimes the worker finishes the business logic but does not delete the message before the visibility timeout expires, so the message is delivered again. Which two changes improve resilience and reduce the impact of duplicate processing? Select two.

Select 2 answers
A.Make the message handler idempotent.
B.Set the SQS visibility timeout long enough for normal processing to complete.
C.Switch from SQS to Amazon SNS for reliable buffering.
D.Shorten the queue retention period so messages expire quickly.
E.Disable retries in the consumer application.
AnswersA, B

SQS provides at-least-once delivery, meaning the same message can be delivered to a consumer more than once, especially during network timeouts or consumer crashes. An idempotent handler design ensures that processing a duplicate message does not cause duplicate side effects, such as creating duplicate database records or processing the same payment twice. This is typically achieved by storing a unique message identifier or business key and checking for prior processing before executing the business logic. Idempotency is the fundamental corrective control for SQS's inherent lack of exactly-once semantics.

Why this answer

Making the message handler idempotent ensures that even if a message is processed multiple times (due to visibility timeout expiry), the business outcome remains the same. Idempotency is a key design pattern for resilient architectures when using at-least-once delivery systems like SQS. Option B is correct because setting the visibility timeout long enough for normal processing prevents premature redelivery, reducing the chance of duplicate processing in the first place.

Exam trap

The trap here is that candidates often think disabling retries or switching to SNS will solve the duplicate processing issue, but they fail to recognize that SQS's at-least-once delivery model inherently requires idempotent consumers and proper visibility timeout configuration.

← PreviousPage 2 of 4 · 257 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Design Resilient questions.