Courseiva

CCNA Design Resilient Questions

32 of 257 questions · Page 4/4 · Design Resilient topic · Answers revealed

226
MCQeasy

A production Amazon RDS database has automated backups enabled with sufficient retention. At 10:30 UTC, a release corrupts specific rows. The issue is detected at 10:45 UTC. The team wants to restore the database state to before the corruption with minimal complexity. What should they do?

A.Perform a point-in-time restore (PITR) to a timestamp just before 10:30 UTC and create a restored DB instance/cluster.
B.Change the VPC route tables so the database restarts in a clean state.
C.Relaunch the same DB instance in the same Availability Zone and rely on caching to revert the changes.
D.Enable a DLQ on the database to store invalid SQL statements until the system is fixed.
AnswerA

PITR uses automated backups to restore the database to a specific point in time. Selecting a timestamp just before the corruption (for example, slightly before 10:30 UTC) restores the affected data state as it existed before the bad release.

Why this answer

Amazon RDS Point-in-Time Restore (PITR) allows you to restore the database to any second within the backup retention period, using automated backups and transaction logs. By restoring to a timestamp just before 10:30 UTC, you can recover the database to a state before the corruption occurred, creating a new DB instance/cluster with minimal complexity and no data loss from the uncorrupted period.

Exam trap

The trap here is that candidates may confuse database recovery methods with network or application-level fixes, or incorrectly assume that restarting or relaunching an instance will clear data changes, when in fact only a restore from backup or PITR can revert committed transactions.

How to eliminate wrong answers

Option B is wrong because changing VPC route tables affects network traffic routing, not database state or data integrity; it cannot revert corrupt rows or restart the database in a clean state. Option C is wrong because relaunching the same DB instance in the same Availability Zone does not revert data changes; it simply creates a new instance with the same underlying storage, which still contains the corrupt rows. Option D is wrong because a Dead Letter Queue (DLQ) is a concept for message queues (like Amazon SQS) to handle failed message processing, not a feature of Amazon RDS; it cannot store or revert SQL statements.

227
MCQmedium

A Multi-AZ Amazon RDS database experiences incorrect writes at 10:15 UTC due to a buggy release. The team detects the problem at 10:25 UTC. They want to restore the data to a known-good point around 10:15 UTC, and validate the recovered data, without taking the current production instance offline during the recovery process. What is the most appropriate AWS action?

A.Immediately reboot the RDS instance and rely on the reboot to roll back the bad writes.
B.Perform a point-in-time restore (PITR) to a new DB instance using a restore time around 10:15 UTC, then test the restored instance before cutting over.
C.Create a new Read Replica from the current primary and use it as the recovered database after applying reverse migrations.
D.Temporarily disable Multi-AZ to speed up storage rollback, then re-enable Multi-AZ.
AnswerB

PITR restores to a specific timestamp using backups and transaction logs. Importantly, it creates a recovered copy (typically a new DB instance), which allows validation and cutover decisions without stopping or directly impacting the existing production instance.

Why this answer

Amazon RDS point-in-time recovery (PITR) allows you to restore a DB instance to any second within the backup retention period, creating a new, independent DB instance. This lets you validate the recovered data without affecting the current production instance, which remains online and serving traffic. The team can then cut over to the restored instance after confirming it is clean.

Exam trap

The trap here is that candidates may assume a reboot or Read Replica can undo bad writes, but neither provides a rollback mechanism; only PITR or a manual restore from a snapshot can recover to a specific point in time without affecting the live instance.

How to eliminate wrong answers

Option A is wrong because rebooting an RDS instance does not roll back writes; it only restarts the database engine and applies any pending maintenance or parameter changes, leaving the bad data intact. Option C is wrong because a Read Replica is an asynchronous copy of the primary that replicates all writes, including the buggy ones, so it cannot serve as a point-in-time recovery target without manual, error-prone reverse migrations. Option D is wrong because disabling Multi-AZ does not provide a storage rollback mechanism; it only removes the standby replica, and the primary's storage still contains the incorrect writes.

228
MCQhard

A financial analytics platform ingests events into an Amazon Kinesis Data Stream with four shards. During month-end peaks, producers receive ProvisionedThroughputExceededException errors and consumers fall behind. The architects want to increase capacity without changing producer code and must preserve the order of records that share the same partition key. What should they do?

A.Enable server-side encryption on the stream and increase the retention period to 365 days.
B.Switch the consumers to enhanced fan-out and raise the number of registered consumers.
C.Replace the Kinesis Data Stream with an Amazon SQS FIFO queue and have consumers poll it.
D.Increase the number of open shards using UpdateShardCount to a higher count.
AnswerD

UpdateShardCount increases the shard count by splitting existing shards, which raises the stream's write and read capacity. Records with the same partition key continue to map to a single shard, preserving order for that key. Producers need no code change because they keep writing with the same partition key to the same stream name.

Why this answer

Producer throttling on a Kinesis data stream is resolved by adding shards, since each shard provides a fixed write and read capacity. UpdateShardCount performs this online, and because the partition key still hashes to one shard, per-key ordering is maintained and no producer changes are required.

Exam trap

The trap here is confusing consumer-side read capacity with producer-side write capacity, so enhanced fan-out looks like a fix for throttling that actually originates on the write path.

229
MCQmedium

A production team accidentally deletes critical rows in an Amazon RDS for PostgreSQL database. The deletion occurred about 6 hours ago. The team wants to recover to a specific point in time with minimal disruption. Assuming automated backups are enabled, which approach provides the best resilience outcome?

A.Restore the current DB instance in place by overwriting it with only the latest automated backup.
B.Use point-in-time recovery (PITR) to restore a new DB instance to a timestamp shortly before the deletion, then switch application traffic to the restored instance.
C.Create a manual snapshot and restore from it only if the snapshot date exactly matches today.
D.Perform a database-level rollback using transaction logs from the application server without using RDS restore features.
AnswerB

With automated backups enabled, PITR allows restoring to a precise timestamp within the retention window. Creating a new DB instance (rather than overwriting production) enables verification of data correctness and then a controlled cutover, minimizing disruption while meeting the “specific point in time” requirement.

Why this answer

Point-in-time recovery (PITR) allows you to restore a new DB instance to any second within the automated backup retention period, which includes transaction logs. By restoring to a timestamp just before the deletion, you recover the lost rows without affecting the current production instance, then switch traffic to the new instance for minimal disruption.

Exam trap

The trap here is that candidates may think restoring in place (Option A) is faster or simpler, but they overlook that PITR provides granular recovery without overwriting the production instance, which is the key to minimal disruption.

Why the other options are wrong

A

Restoring the current DB instance in place by overwriting it with the latest automated backup would revert all data to the backup time, losing all changes made in the last 6 hours, including the critical rows that were deleted. It does not allow recovery to a specific point in time before the deletion.

C

Creating a manual snapshot today and restoring from it would not recover data from 6 hours ago; it would only restore to the snapshot creation time, which is after the deletion.

D

RDS does not expose transaction logs for direct database-level rollback; point-in-time recovery is the only supported method to restore to a specific time using automated backups and transaction logs stored by AWS.

230
MCQeasy

A company wants a disaster recovery setup for a web application. They want to keep costs low but still recover within a couple of hours after a regional disruption. They are willing to run only minimal infrastructure in the secondary location and scale it up during the outage. Which DR approach best matches this requirement?

A.Active-active, where both Regions run full production at all times.
B.Pilot light, where the secondary Region keeps minimal core components ready and scales up during failover.
C.Cold standby, where no infrastructure is running in the secondary Region until an outage occurs.
D.Backups-only, where recovery relies solely on manually restoring snapshots during an outage.
AnswerB

Pilot light keeps a minimal but always-on core in the secondary Region — for example, an RDS cross-Region read replica or replicated DynamoDB tables — while application servers and other scale-out components stay shut down. On failover, you use pre-built AMIs or CloudFormation templates to quickly scale up the remaining infrastructure, change Route 53 routing, and start serving traffic. This gives a low RTO (often under an hour) and lower steady-state cost than active-active, making it the best fit for the stated couple-hour RTO.

Why this answer

The Pilot light approach is correct because it keeps minimal core components (e.g., a small database, a scaled-down application server) running in the secondary Region, allowing rapid failover by scaling up those resources during an outage. This meets the requirement of low cost during normal operations while achieving recovery within a couple of hours, as the core infrastructure is already provisioned and can be scaled horizontally (e.g., using Auto Scaling groups and pre-configured AMIs) without needing to rebuild from scratch.

Exam trap

The trap here is confusing Pilot light with Cold standby, as both involve minimal infrastructure, but Pilot light has core components already running (e.g., a small database instance) while Cold standby has nothing provisioned, leading to significantly longer recovery times.

How to eliminate wrong answers

Option A is wrong because Active-active runs full production in both Regions at all times, which incurs high costs and does not match the requirement to keep costs low. Option C is wrong because Cold standby has no infrastructure running in the secondary Region until an outage occurs, which would typically require more than a couple of hours to provision and configure resources (e.g., launching EC2 instances, restoring databases) and thus fails the recovery time objective. Option D is wrong because Backups-only relies on manually restoring snapshots (e.g., EBS snapshots, RDS snapshots) during an outage, which is slow and error-prone, often exceeding the couple-of-hours recovery window due to manual intervention and data transfer times.

231
Multi-Selecthard

A financial services company is designing a new payment processing platform. The platform must continue to accept and process transactions even if an entire AWS Region becomes unavailable, and it must not lose any accepted transaction. The architects have decided to run active-active deployments in two Regions and use Amazon Route 53 for traffic management. Which two additional design elements are required to meet the durability and availability goals? (Choose two.)

Select 2 answers
A.Store all transaction records in a single Amazon S3 bucket in the primary Region and enable S3 Versioning for durability.
B.Use an Amazon Route 53 latency-based routing policy with health checks on both Regional endpoints so traffic shifts away from an unhealthy Region.
C.Replicate transaction data across both Regions using a multi-Region, multi-active database such as Amazon Aurora Global Database or DynamoDB global tables, and design writes to be idempotent.
D.Configure an Amazon Route 53 failover routing policy with a primary record in one Region and a secondary record in the other, and take hourly Amazon EBS snapshots of the application servers.
E.Deploy the application tier into a single Region and use AWS Global Accelerator to route European users through the nearest edge location for lower latency.
AnswersB, C

Latency-based routing with health checks directs users to the lowest-latency healthy Region and automatically removes an endpoint that fails its health check. This is what keeps the platform reachable when one Region is impaired, satisfying the availability goal in an active-active topology. Without health-check-driven failover, clients could continue being sent to a failed Region.

Why this answer

An active-active, multi-Region platform needs two things beyond compute in each Region: a routing layer that detects a failed Region and steers traffic to the healthy one, and a data layer that keeps a writable, replicated copy of transactions in both Regions. Health-checked latency routing handles the first, while a multi-Region database with idempotent writes handles the second and protects accepted transactions from loss.

Exam trap

The trap here is assuming that running application servers in two Regions is sufficient, when the data layer and the health-checked routing are what actually deliver Regional failover.

232
MCQeasy

A inventory service exposes a static website from S3 and CloudFront. Users should still receive cached pages if the S3 origin has a short outage. Which feature helps most? The design must avoid adding custom operational scripts.

A.CloudFront caching with appropriate TTLs
B.AWS Backup Vault Lock
C.IAM Access Analyzer
D.S3 Select
AnswerA

CloudFront caches objects at edge locations, so even when the S3 origin becomes temporarily unavailable, requests for cached content can be served from the edge as long as the TTL has not expired. If configured with error caching or a sufficiently long TTL, CloudFront can continue serving stale content during an origin failure, acting as a resilience buffer rather than merely a latency optimization. Selecting appropriate TTLs is therefore critical to making the static website tolerant to brief S3 outages.

Why this answer

CloudFront caches responses at edge locations based on configured TTLs (Cache-Control or Expires headers). If the S3 origin becomes temporarily unavailable, CloudFront can still serve stale or cached content to users, maintaining availability without any custom scripts or failover logic. This directly addresses the requirement to serve cached pages during short S3 outages.

Exam trap

The trap here is that candidates might think AWS Backup Vault Lock (Option B) provides some form of data availability or failover, but it is purely a compliance and retention tool with no impact on serving cached web content during origin outages.

How to eliminate wrong answers

Option B is wrong because AWS Backup Vault Lock is a data protection feature for backup vaults, enforcing retention policies (WORM) to prevent deletion; it does not provide caching or origin failover for web content. Option C is wrong because IAM Access Analyzer helps identify unintended resource access policies, not caching or availability during origin outages. Option D is wrong because S3 Select is a query-in-place feature to retrieve subsets of object data using SQL expressions; it has no role in caching or serving cached pages during origin failures.

233
MCQhard

A financial services company runs a critical application on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer. The application must be able to survive the failure of an entire AWS Region. The company wants a cost-effective solution that minimizes operational overhead. Which approach should the architect recommend?

A.Deploy the application in one Region and take regular Amazon EBS snapshots copied to another Region.
B.Deploy the application in two Regions with an Auto Scaling group in each, and use Amazon Route 53 failover routing with health checks.
C.Deploy the application in one Region and use AWS Global Accelerator to route traffic to the nearest edge location.
D.Deploy the application in two Regions and use an Amazon S3 cross-Region replication bucket to store application logs.
AnswerB

A multi-Region active-passive or active-active deployment with Route 53 failover routing and health checks automatically shifts traffic to the healthy Region if the primary fails. This provides regional resilience with minimal operational overhead compared to custom DNS or manual failover.

Why this answer

To survive a regional failure, the application must be deployed in at least two Regions with independent compute capacity. Route 53 failover routing with health checks automatically directs traffic to the healthy Region, providing resilience with minimal manual intervention and operational overhead.

Exam trap

The trap here is assuming that Global Accelerator or cross-Region log replication provides regional failover, when only a full deployment in a second Region with DNS failover can keep the application available.

234
MCQmedium

Based on the exhibit, which Route 53 configuration should be used so traffic automatically returns to the secondary Region only when the primary Region becomes unhealthy?

A.Use latency-based routing with both ALB records enabled.
B.Use failover routing with a primary alias record, a secondary alias record, and a Route 53 health check on the primary target.
C.Use geolocation routing so users are always sent to the closest Region.
D.Use a CNAME record that points to both ALBs so DNS can round-robin between Regions.
AnswerB

Failover routing is designed for this pattern: Route 53 returns the primary alias while the primary endpoint is healthy, and switches to the secondary alias when the primary health check fails. Alias records integrate cleanly with ALB targets, and the health check provides the signal that drives the failover decision.

Why this answer

Failover routing in Amazon Route 53 is designed for active-passive configurations. By creating a primary alias record pointing to the ALB in the primary Region and a secondary alias record pointing to the ALB in the secondary Region, and attaching a Route 53 health check to the primary target, traffic automatically fails over to the secondary Region only when the health check detects the primary as unhealthy. This meets the requirement of returning traffic to the secondary Region only upon primary failure.

Exam trap

The trap here is that candidates often confuse failover routing with latency-based or geolocation routing, assuming that 'closest' or 'fastest' automatically implies health awareness, but Route 53 health checks must be explicitly associated with failover records to trigger automatic traffic redirection.

How to eliminate wrong answers

Option A is wrong because latency-based routing directs users based on lowest latency, not health status, so it would not automatically fail over only when the primary is unhealthy; traffic could still be sent to an unhealthy primary if latency is low. Option C is wrong because geolocation routing sends users based on their geographic location, not the health of the endpoint, so it cannot automatically redirect traffic to the secondary Region when the primary becomes unhealthy. Option D is wrong because a CNAME record cannot point to multiple ALBs for round-robin; CNAME records can only point to a single DNS name, and DNS round-robin does not consider health checks, so traffic would still be sent to an unhealthy primary.

235
MCQmedium

A web application runs on an Auto Scaling group (ASG) behind an Application Load Balancer (ALB). After a new release, instances begin failing ALB health checks with errors like 502 while the application is still starting up. CloudWatch shows that the ASG replaces the instances before they finish initializing, so traffic never reaches healthy targets. Which change most directly prevents premature replacement during startup so traffic can resume as soon as the instances are actually healthy?

A.Reduce the ALB health check timeout to 1 second so failures are detected faster.
B.Increase the Auto Scaling group health check grace period to cover application startup and initialization time.
C.Enable connection draining on the ALB target group but set deregistration delay to 0 seconds.
D.Switch the ALB target group health checks from HTTP to TCP so the application does not need to return HTTP 200.
AnswerB

The ASG health check grace period tells Auto Scaling to ignore failing health checks for a period after instance launch. This prevents newly launched instances from being replaced before the application has finished booting and can pass ALB health checks.

Why this answer

B is correct because the Auto Scaling group health check grace period allows instances a specified amount of time to initialize before the ASG starts checking their health status. By increasing this grace period to cover the application startup time, the ASG will not prematurely replace instances that are still initializing, allowing them to pass the ALB health checks and begin receiving traffic once they are actually healthy.

Exam trap

The trap here is that candidates often confuse the ALB health check timeout or interval with the ASG health check grace period, thinking that adjusting ALB settings will fix the premature replacement issue, when in fact the ASG grace period is the direct control for delaying health check evaluation during startup.

How to eliminate wrong answers

Option A is wrong because reducing the ALB health check timeout to 1 second would cause health checks to fail even faster, exacerbating the problem of premature instance replacement. Option C is wrong because connection draining controls how existing connections are closed during deregistration, not how quickly instances are replaced during startup; setting deregistration delay to 0 seconds would abruptly terminate active connections, causing user disruption. Option D is wrong because switching to TCP health checks would bypass the application layer, allowing the ALB to consider an instance healthy even if the application is not fully initialized, which could lead to serving 502 errors to users.

236
Multi-Selectmedium

A company is designing a multi-Region disaster recovery (DR) strategy for a stateless web application running on Amazon EC2 instances behind an Application Load Balancer (ALB). The application uses an Amazon RDS for MySQL database as its data store. The architecture must provide rapid failover with the lowest possible Recovery Point Objective (RPO) and Recovery Time Objective (RTO). Which of the following design choices will help achieve these objectives? (Choose four.)

Select 4 answers
.Configure an active-passive failover strategy by deploying the application stack in two AWS Regions and using Amazon Route 53 health checks with a failover routing policy.
.Set up Amazon RDS Multi-AZ deployment to enable automatic failover to a standby replica in a different Availability Zone within the primary Region.
.Use Amazon RDS cross-Region read replicas with automatic failover to promote a read replica to a primary instance in the secondary Region.
.Deploy the application and ALB in an active-active configuration across two AWS Regions using Amazon Route 53 latency-based routing.
.Store static assets and application state in Amazon S3 with cross-Region replication enabled, and serve them via Amazon CloudFront.
.Use an Amazon RDS for MySQL single-AZ deployment in the primary Region and take daily snapshots copied to the secondary Region.

Why this answer

An active-passive failover strategy with Route 53 failover routing policy is correct because it provides rapid failover by directing traffic to the secondary Region only when health checks fail in the primary, minimizing RTO. Cross-Region read replicas with automatic failover are correct because they allow promoting a read replica to a primary in the secondary Region with low RPO (typically seconds) and automated failover, reducing RTO. Active-active configuration with latency-based routing is correct because it distributes traffic across both Regions, enabling immediate failover without DNS propagation delays, achieving very low RTO.

Storing static assets and application state in S3 with cross-Region replication and CloudFront is correct because it ensures data durability and low-latency access, supporting rapid recovery with minimal RPO.

Exam trap

The trap here is that candidates often confuse Multi-AZ (single-Region high availability) with cross-Region DR, or they assume daily snapshots provide adequate RPO for a DR strategy requiring the lowest possible RPO and RTO.

237
MCQeasy

A startup runs a stateless image-resizing API on a fleet of EC2 instances behind an Application Load Balancer. The instances store uploaded source images on their own instance store volumes before processing. During a routine scale-in event, an instance was terminated and several in-flight uploads were lost. The architect must make the design resilient to instance loss without changing the API code. What should the architect do?

A.Store uploaded images in Amazon S3 and have instances read from and write to the bucket instead of local disk
B.Increase the Auto Scaling group's minimum capacity so that instances are rarely terminated
C.Enable detailed CloudWatch monitoring and create an alarm that notifies operators before scale-in occurs
D.Configure an EC2 Auto Scaling lifecycle hook to delay instance termination until uploads complete
AnswerA

Moving the uploaded source images to Amazon S3 decouples the data from any single EC2 instance, so terminating an instance no longer destroys in-flight uploads. S3 provides durable, highly available object storage that all instances can access concurrently. Because the API already treats instances as stateless workers, replacing local storage with S3 is the standard resilience pattern and requires no change to the fleet's scaling behavior.

Why this answer

The root cause is that uploaded images live on instance store volumes, which are ephemeral and destroyed when an instance terminates. Relocating the source images to Amazon S3 removes the dependency on any single instance and gives the fleet shared, durable storage, so scale-in no longer causes loss. The remaining options delay or observe termination or reduce its frequency, but none preserve the data itself.

Exam trap

The trap here is treating instance store as durable because it is physically attached to the instance; instance store data is lost on stop, terminate, or host failure, so it cannot back resilient state.

238
MCQmedium

A inventory service uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The design must avoid adding custom operational scripts.

A.Lambda reserved concurrency set to zero
B.A Lambda dead-letter queue or failure destination
C.A larger deployment package
D.CloudFront error pages
AnswerB

A dead-letter queue or failure destination captures invocation records after Lambda exhausts its asynchronous retries, retaining failed events for later investigation. This is native Lambda configuration, so no custom operational scripts are needed, satisfying the requirement to retain failures without extra tooling.

Why this answer

A Lambda dead-letter queue (DLQ) or failure destination allows you to capture events that have exhausted all retry attempts from an asynchronous invocation. When the Lambda function fails after the maximum retries (default 3), the event is sent to the configured SQS queue or SNS topic for later investigation, without requiring custom scripts or manual polling.

Exam trap

The trap here is that candidates may confuse Lambda's DLQ/failure destination with other error-handling mechanisms like SQS redrive policies or CloudFront custom error pages, which serve different purposes and operate at different layers of the architecture.

How to eliminate wrong answers

Option A is wrong because setting reserved concurrency to zero would completely disable the Lambda function, preventing any invocations and thus failing to process or retain any events. Option C is wrong because a larger deployment package does not affect error handling or event retention; it only increases cold start latency and deployment size. Option D is wrong because CloudFront error pages are for HTTP-level errors from a web distribution, not for capturing asynchronous Lambda invocation failures or dead-letter events.

239
MCQmedium

A company runs a customer portal on an Amazon Aurora PostgreSQL cluster. The application currently connects directly to the writer instance endpoint and keeps long-lived connections open. During a maintenance failover, writes fail until clients are restarted. The team wants the application to reconnect to the correct Aurora endpoint automatically and reduce user-visible write interruptions. Which change is most likely to achieve this?

A.Use the Aurora cluster endpoint for write traffic, use the reader endpoint for read-only traffic, and implement connection retry or reconnect logic on failover.
B.Keep using the original writer instance endpoint so the database host name never changes during failover.
C.Convert the Aurora cluster to Single-AZ so there is only one database node to connect to.
D.Place Route 53 in front of the database and manually update DNS records whenever failover occurs.
AnswerA

The cluster endpoint always resolves to the current writer, so reconnecting after failover targets the promoted instance rather than a stale writer address; retry logic lets long-lived connections re-establish without restarting clients, reducing write interruption.

Why this answer

The Aurora cluster endpoint automatically points to the current writer instance, so using it for write traffic ensures that after a failover, new writes are directed to the new writer without needing to change the connection string. Implementing connection retry or reconnect logic in the application is essential because the existing long-lived connections will be broken during failover; the application must detect the failure and re-establish connections to the cluster endpoint to resume writes seamlessly.

Exam trap

The trap here is that candidates assume the writer instance endpoint remains constant during failover (Option B), but in Aurora, the writer instance endpoint changes because it is tied to the specific DB instance, not the cluster.

Why the other options are wrong

B

The writer instance endpoint points to a specific Aurora node, which changes during failover. Keeping it does not automatically redirect traffic to the new writer, so writes still fail until clients are restarted.

C

Converting to Single-AZ removes the standby replica, eliminating high availability. During a failover, there is no standby to promote, causing longer downtime and potential data loss, which contradicts the goal of reducing write interruptions.

D

Manually updating Route 53 DNS records during failover is not automated and would still cause write interruptions until the manual update is completed, failing to meet the requirement of automatic reconnection and reduced downtime.

240
MCQmedium

A developer accidentally deletes important rows in an RDS database. The mistake is discovered 45 minutes later. The database has automated backups enabled with a retention period of 7 days. What is the best way to restore the database to a point just before the deletion?

A.Restore the latest manual snapshot and then run SQL scripts to revert the deletion.
B.Use point-in-time restore (PITR) to restore the database to a specific timestamp before the deletion, based on automated backups.
C.Promote an existing read replica to be the primary and then copy the missing rows from logs.
D.Recreate the instance using the most recent CloudWatch metric alarm snapshot of storage metrics.
AnswerB

With automated backups enabled, RDS supports PITR within the retention window. PITR lets you restore to any second within that window, so you can select a timestamp just before the destructive deletion occurred. This avoids restoring a potentially stale snapshot and eliminates the need for risky manual compensating scripts.

Why this answer

Point-in-time restore (PITR) allows you to restore an RDS DB instance to any second within the automated backup retention period (here, 7 days). Since the deletion occurred 45 minutes ago, you can specify a timestamp just before the deletion, and RDS will replay the transaction logs to bring the database to that exact state. This is the most precise and efficient recovery method for accidental data modifications.

Exam trap

The trap here is that candidates may assume manual snapshots or read replicas can be used for granular point-in-time recovery, but only automated backups with transaction logs enable restoring to a specific second within the retention period.

How to eliminate wrong answers

Option A is wrong because manual snapshots capture the entire instance at a point in time, but they do not provide the granularity to restore to a specific moment just before the deletion; you would lose all changes made after the snapshot, and running SQL scripts to revert deletions is error-prone and not a built-in RDS feature. Option C is wrong because promoting a read replica makes it a new primary, but it does not revert data; it simply becomes a writable copy of the current state, which still contains the deletion. Option D is wrong because CloudWatch metric alarms monitor performance metrics, not database row-level data; they cannot be used to restore or recover deleted rows.

241
MCQmedium

A payments API uses an RDS MySQL database and must remain available during an Availability Zone failure with minimal application changes. What should the architect enable?

A.S3 Cross-Region Replication
B.Multi-AZ deployment for the RDS DB instance
C.Read replicas only
D.EBS snapshots every hour
AnswerB

A Multi-AZ deployment creates a synchronized standby replica in a different Availability Zone of the same Region. Amazon RDS automatically fails over to the standby when a problem is detected on the primary, providing high availability for the DB instance. This is the correct solution for automatic, synchronous failover with minimal data loss.

Why this answer

Multi-AZ deployment for RDS MySQL automatically provisions and maintains a synchronous standby replica in a different Availability Zone. In the event of an AZ failure, Amazon RDS automatically fails over to the standby, providing high availability with minimal application changes (the application only needs to reconnect to the same endpoint). This meets the requirement for availability during an AZ outage without requiring code modifications.

Exam trap

The trap here is that candidates often confuse read replicas (which are for read scaling and manual promotion) with Multi-AZ (which provides automatic failover for high availability), leading them to select read replicas as a cheaper or simpler alternative.

How to eliminate wrong answers

Option A is wrong because S3 Cross-Region Replication is for object storage in S3, not for RDS MySQL databases, and it does not provide automatic failover for a relational database. Option C is wrong because read replicas are designed for read scaling, not for automatic failover during an AZ failure; they require manual promotion and application changes to redirect writes. Option D is wrong because EBS snapshots every hour provide point-in-time backup and recovery, not high availability; restoring from a snapshot would involve significant downtime and manual intervention, not minimal application changes.

242
MCQmedium

You host a public API using Amazon API Gateway in two AWS Regions: us-east-1 (primary) and us-west-2 (secondary). You want Route 53 to send client traffic to the secondary region only when the primary API is unhealthy. Which Route 53 setup best meets this requirement?

A.Use latency-based routing with one routing policy per region, and use CloudWatch alarms to update traffic weights between regions.
B.Use Route 53 failover routing with two ALIAS records (same DNS name) pointing to the API Gateway regional endpoints: one record is configured as PRIMARY with an associated health check, and the other is configured as SECONDARY.
C.Use weighted routing across both regions and rely on Route 53 health checks to automatically set the secondary to 100% weight when the primary fails.
D.Use geolocation routing to map some client geographies to the secondary region and the rest to the primary region.
AnswerB

Route 53 failover routing is purpose-built for active-passive architecture across two endpoints. The PRIMARY ALIAS record is tied to a health check that monitors the API Gateway regional endpoint in the primary region; when that health check fails, Route 53 automatically returns the SECONDARY record's endpoint, redirecting traffic to the secondary region. This is the correct, fully managed way to achieve automatic regional failover for public APIs.

Why this answer

Route 53 failover routing is designed for active-passive setups where traffic is sent to a primary resource unless it is unhealthy, in which case traffic is routed to a secondary resource. By creating two ALIAS records with the same DNS name, one marked PRIMARY with an associated health check and the other marked SECONDARY, Route 53 will automatically fail over to the secondary region when the health check for the primary API Gateway endpoint fails. This directly meets the requirement of sending traffic to the secondary region only when the primary API is unhealthy.

Exam trap

The trap here is that candidates often confuse weighted routing with failover routing, mistakenly believing that Route 53 health checks can automatically adjust weights to achieve active-passive failover, when in fact weighted routing does not support dynamic weight adjustment based on health.

How to eliminate wrong answers

Option A is wrong because latency-based routing directs traffic based on lowest latency, not health, and using CloudWatch alarms to manually update weights is not an automatic failover mechanism; it also requires custom automation and does not natively support health-check-driven failover. Option C is wrong because weighted routing distributes traffic based on assigned weights and does not automatically set the secondary to 100% weight when the primary fails; Route 53 health checks can mark a record as unhealthy but do not dynamically adjust weights—they would cause the primary record to be excluded from responses, but the secondary would only receive traffic if its weight is non-zero, and the behavior is not a clean active-passive failover. Option D is wrong because geolocation routing directs traffic based on the geographic location of the client, not the health of the endpoint, and it cannot automatically fail over traffic from one region to another when the primary becomes unhealthy.

243
MCQmedium

An ECS service runs on EC2 instances and is fronted by an ALB. The ALB spans two Availability Zones, and the ECS service desired count is 2 tasks. The underlying EC2 capacity uses an Auto Scaling group (ASG) with min size set to 1, and the ASG also spans only one subnet in practice. What is the most effective change to meet the requirement that the service continues during a single-AZ instance loss?

A.Set the ECS deployment configuration to maximum percent 100 so tasks replace instances faster during rollouts.
B.Increase ASG min size to at least 2 and ensure the ASG uses subnets in at least two Availability Zones.
C.Enable ALB connection draining longer than expected so existing connections survive longer during an AZ event.
D.Reduce task memory reservations to pack both tasks onto a single EC2 instance.
AnswerB

Increasing the Auto Scaling group minimum to at least 2 and using subnets in at least two Availability Zones guarantees baseline EC2 capacity in multiple AZs. If one AZ becomes unavailable, the ALB can route traffic to healthy targets in the remaining AZ, and ECS has compute available to reschedule tasks. This directly provides the cross-AZ redundancy needed for the service to continue operating.

Why this answer

The current architecture has a single point of failure because the Auto Scaling group (ASG) spans only one subnet (one Availability Zone). If that AZ fails, all EC2 instances are lost, and the ECS service cannot run any tasks. Increasing the ASG min size to at least 2 and configuring it to use subnets in at least two AZs ensures that EC2 instances are distributed across AZs, allowing the ECS service to maintain at least one task in the surviving AZ during a single-AZ failure.

Exam trap

The trap here is that candidates often focus on ECS-specific settings (like deployment configuration or task placement) rather than recognizing that the root cause is the ASG's single-AZ limitation, which is a fundamental infrastructure resilience issue.

Why the other options are wrong

A

Setting maximum percent to 100 does not address the lack of multi-AZ redundancy; it only affects deployment speed, not availability during an AZ failure.

C

Connection draining helps preserve existing connections during a rolling update or instance deregistration, but it does not prevent service disruption when an entire Availability Zone fails. The ALB would still lose all healthy targets in that AZ, and new connections cannot be established to instances in the failed AZ.

D

Reducing task memory reservations does not address the single-AZ failure risk because both tasks could still be placed in the same Availability Zone, and the ASG only spans one subnet, so losing that AZ would still cause total service outage.

244
Multi-Selectmedium

A solutions architect is designing a highly available relational database tier for a customer-facing order system that must survive the loss of an entire Availability Zone with minimal administrative effort and no application connection-string changes during failover. (Choose two.)

Select 2 answers
A.Create manual read replicas in each Availability Zone and repoint the application to a replica after failover
B.Enable a Multi-AZ DB cluster deployment for the RDS database so a writer and two readable standbys span three Availability Zones
C.Store the database on an EC2 instance with an EBS volume and take nightly snapshots to a second Availability Zone
D.Deploy the database as an Amazon RDS Multi-AZ DB instance so a standby is maintained in a second Availability Zone
E.Rely on the RDS automated backup retention window to restore the instance into a different Availability Zone during an outage
AnswersB, D

A Multi-AZ DB cluster runs one writer and two readable standbys across three Availability Zones and provides a single writer endpoint plus a reader endpoint. Failover is automatic and typically faster than a Multi-AZ DB instance, and the endpoint remains stable, so the application is unaffected. This meets the zone-failure and no-connection-string-change requirements.

Why this answer

Both Multi-AZ deployment types keep a synchronous standby copy of the data in another Availability Zone and expose a stable endpoint that survives automatic failover, which removes the need to edit connection strings. A Multi-AZ DB cluster goes further by adding two readable standbys across three zones and offering faster failover. The remaining choices rely on asynchronous replicas, manual promotion, or snapshot restores that all require operator action.

Exam trap

The trap here is assuming read replicas provide automatic failover; RDS read replicas replicate asynchronously and promoting one requires a manual step and can lose recent writes.

245
MCQmedium

A ticket booking system uses Aurora MySQL. The company wants fast cross-Region disaster recovery with low RPO. Which architecture should be considered? The design must avoid adding custom operational scripts.

A.Aurora Global Database
B.A single-AZ Aurora cluster
C.An ElastiCache Redis replica
D.Manual snapshots copied monthly
AnswerA

Aurora Global Database is the correct choice because it uses a primary Region plus up to five secondary Regions with dedicated storage-level replication typically under one second. This gives a low recovery point objective (RPO) and fast promoted read replica failover for disaster recovery, unlike snapshot-based or single-Region options. It also allows local reads in secondary Regions for low-latency access, directly meeting the ticket system's requirement for fast access and rapid DR.

Why this answer

Aurora Global Database is designed for cross-Region disaster recovery with a typical RPO of 1 second and RTO of 1 minute, using storage-based replication that does not require custom scripts. It replicates data from a primary Region to up to five secondary Regions with minimal impact on database performance, meeting the low RPO requirement without operational overhead.

Exam trap

The trap here is that candidates may confuse cross-Region read replicas (which require manual promotion and have higher RPO) with Aurora Global Database, which provides automated failover and lower RPO without custom scripts.

How to eliminate wrong answers

Option B is wrong because a single-AZ Aurora cluster lacks any cross-Region replication or failover capability, providing no disaster recovery across Regions. Option C is wrong because ElastiCache Redis is an in-memory cache, not a persistent database, and cannot serve as a primary data store for ticket bookings or provide cross-Region DR with low RPO. Option D is wrong because manual snapshots copied monthly result in an RPO of up to one month, which is far too high for fast disaster recovery, and the process requires custom scripting to automate cross-Region copy.

246
MCQhard

A healthcare data platform stores patient documents in an Amazon S3 bucket in us-east-1. Regulations require that the data remain readable with low latency even if the entire us-east-1 Region becomes unavailable, and that writes continue in a secondary Region. The team wants object-level replication with minimal operational overhead and must preserve version history. Which solution BEST meets these requirements?

A.Use S3 Transfer Acceleration on the existing bucket and rely on edge locations to serve reads during the outage.
B.Enable S3 Cross-Region Replication (CRR) with versioning enabled on both buckets, and configure the application to read from the replica bucket during a regional failure.
C.Enable S3 Same-Region Replication to a second bucket in us-east-1 and serve reads from that bucket during a regional failure.
D.Configure an S3 Lifecycle policy to transition objects to S3 Glacier Deep Archive in a second Region.
AnswerB

S3 CRR replicates objects asynchronously to a bucket in another Region and requires versioning on both source and destination, preserving version history. If us-east-1 becomes unavailable, the application can read from the replica bucket, and writes can be redirected there, satisfying the low-latency read and continued-write requirements.

Why this answer

Cross-Region Replication copies objects to a bucket in another Region and requires versioning on both buckets, which preserves version history. The replica bucket remains accessible if the primary Region fails, and writes can be directed there, meeting both the availability and data-durability goals with minimal operational effort.

Exam trap

The trap here is confusing S3 Transfer Acceleration or archival lifecycle rules with actual cross-Region data replication, which are unrelated features.

247
MCQmedium

Based on the exhibit, an administrator accidentally deleted data from Amazon RDS for PostgreSQL about 90 minutes ago. Which recovery approach best restores the database to the exact required point in time?

A.Restore the latest automated snapshot back onto the existing DB instance.
B.Restore the database to the specified point in time into a new DB instance.
C.Create a read replica and promote it after the deletion is noticed.
D.Enable Multi-AZ so the database can automatically undo application mistakes.
AnswerB

Point-in-time restore uses automated backups plus transaction logs to recreate the database at a specific moment. For accidental deletion, this is the correct RDS recovery method because it can recover the database to just before the bad change while preserving all legitimate data up to that point.

Why this answer

Amazon RDS for PostgreSQL supports Point-in-Time Recovery (PITR), which allows you to restore a DB instance to any second within the backup retention period, up to the last five minutes. Since the deletion occurred approximately 90 minutes ago, you can restore to that exact point in time by specifying the timestamp, and RDS will create a new DB instance from automated backups and transaction logs. This is the only option that recovers the exact state before the accidental deletion.

Exam trap

The trap here is that candidates confuse automated snapshots with point-in-time recovery, assuming a snapshot restore can target a specific time, when in fact snapshots are point-in-time captures and cannot replay transaction logs to reach an arbitrary second.

Why the other options are wrong

A

Restoring the latest automated snapshot would not recover data to a specific point in time 90 minutes ago; it only restores to the snapshot creation time, which is likely much earlier.

C

Creating a read replica does not restore data; it only creates a live copy of the current database, which already has the deletion. Promoting it would not recover the deleted data from 90 minutes ago.

D

Multi-AZ provides high availability through synchronous replication to a standby instance, but it does not enable point-in-time recovery or undo application mistakes like accidental data deletion.

248
MCQmedium

A warehouse integration service receives bursts of orders that sometimes overwhelm a downstream fulfilment service. The architecture must absorb spikes and retry processing without losing requests. Which service should be placed between the web tier and fulfilment workers?

A.AWS WAF
B.Amazon Route 53 weighted routing
C.Amazon SQS queue
D.Amazon CloudFront
AnswerC

Amazon SQS is a fully managed message queue that decouples the warehouse integration service from order-producing applications. Producers enqueue each order, and consumers poll for messages at their own pace, allowing the queue to absorb sudden bursts and hold messages until processing capacity is available. Visibility timeout temporarily hides in-flight messages to prevent duplicate processing, while a dead-letter queue captures messages that repeatedly fail, enabling safe retries and preserving the order data during transient outages.

Why this answer

Amazon SQS is the correct choice because it acts as a durable, fully managed message queue that decouples the web tier from the fulfilment workers. It can absorb bursts of orders by storing messages durably, and workers can poll the queue at their own pace, with built-in retry logic via visibility timeouts and dead-letter queues to ensure no requests are lost.

Exam trap

The trap here is that candidates may confuse load-balancing or caching services (like Route 53 or CloudFront) with message queuing, failing to recognize that only a durable queue like SQS provides the necessary buffering, decoupling, and retry semantics for asynchronous order processing.

How to eliminate wrong answers

Option A is wrong because AWS WAF is a web application firewall that filters HTTP/S traffic based on rules, not a queuing or buffering mechanism; it cannot absorb spikes or retry processing. Option B is wrong because Amazon Route 53 weighted routing distributes DNS traffic across multiple endpoints based on weights, but it does not provide durable storage or retry capabilities for individual requests. Option D is wrong because Amazon CloudFront is a content delivery network (CDN) that caches static and dynamic content at edge locations; it can reduce load on origins but cannot queue or retry individual order messages.

249
MCQhard

A financial analytics platform runs a stateless API on Amazon EC2 instances in an Auto Scaling group behind a Network Load Balancer. The API reads from an Amazon Aurora MySQL cluster that has a single writer instance and one reader instance in a different Availability Zone. During a recent Availability Zone event, the writer instance failed and the API saw several minutes of failed writes. The team wants writes to resume automatically with minimal downtime and no application code changes. What should the solutions architect do?

A.Ensure the Aurora cluster has at least one Aurora Replica in a different Availability Zone and rely on the cluster's automatic failover, which updates the writer endpoint to point at the promoted replica.
B.Add a second Aurora reader instance in a third Availability Zone and update the application connection string to use the reader endpoint for all queries.
C.Create an Aurora Replica in a second Availability Zone and promote it manually to writer using the AWS Management Console when a failure occurs.
D.Convert the cluster to an Aurora multi-master cluster so every instance can accept writes, and keep the existing writer endpoint in the connection string.
AnswerA

Aurora automatically promotes an Aurora Replica to writer when the current writer fails, and the cluster writer endpoint is repointed to the new writer, so the application keeps using the same endpoint and resumes writes without code changes. Having the replica in a separate Availability Zone ensures the failover target survives the zone event.

Why this answer

Aurora provides automatic failover by promoting an existing Aurora Replica to the writer role and updating the cluster writer endpoint, so applications that connect through that endpoint reconnect without modification. The prerequisite is that a replica exists in a different Availability Zone from the writer, which makes it a viable failover target when the writer's zone is lost.

Exam trap

The trap here is reaching for manual promotion or multi-master when Aurora already performs automatic failover to a replica, provided the replica exists in a different Availability Zone.

250
MCQeasy

An internal API is hosted in two AWS Regions behind Route 53. Under normal conditions, clients should use the primary region. If the primary endpoint becomes unhealthy, traffic must automatically switch to the secondary region. Which Route 53 setup best meets this requirement?

A.Use latency-based routing with one record per region and no health checks.
B.Use failover routing policy: create two alias records for the same name (primary and failover) and associate health checks with the primary record.
C.Use weighted routing and manually change the weights during incidents.
D.Create a single alias record only for the primary region and rely on client-side DNS retries.
AnswerB

Failover routing with two alias records for the same DNS name (one marked primary, one marked secondary) gives you deterministic active-passive failover. You attach a health check to the primary alias; when that health check fails, Route 53 automatically returns the secondary record's endpoint in the next DNS response, without manual intervention. Alias records allow you to point directly to regional load balancers or other AWS resources, and the secondary record ensures that all traffic moves to the healthy region once the primary is considered unhealthy.

Why this answer

Route 53 failover routing policy is designed for active-passive failover scenarios. By creating two alias records (primary and secondary) for the same DNS name and associating a health check with the primary record, Route 53 automatically directs traffic to the secondary region if the primary health check fails. This meets the requirement of automatic failover without manual intervention.

Exam trap

The trap here is that candidates often confuse failover routing with latency-based routing, assuming latency routing inherently handles failover, but latency routing does not automatically switch traffic when an endpoint becomes unhealthy unless health checks are explicitly configured.

How to eliminate wrong answers

Option A is wrong because latency-based routing distributes traffic based on lowest latency, not active-passive failover, and without health checks it cannot detect endpoint failures. Option C is wrong because weighted routing requires manual weight changes during incidents, which violates the requirement for automatic failover. Option D is wrong because a single alias record with no secondary endpoint provides no failover capability; client-side DNS retries do not redirect to a different region.

251
MCQeasy

An orders service currently sends HTTP requests directly to two downstream services (inventory and shipping). During peak load, inventory slows down, causing the orders service to slow as well. The team wants the orders service to remain responsive even when a downstream service is temporarily slow or restarted. Which design change best achieves this resiliency goal?

A.Keep HTTP calls but add longer client timeouts so orders requests wait for slow downstream responses.
B.Introduce Amazon SQS as a buffer between orders and downstream services, with consumers processing from the queue.
C.Replace the downstream services with AWS Lambda functions that are invoked synchronously by the orders service.
D.Call the downstream services in parallel threads to reduce waiting time during peak load.
AnswerB

SQS decouples the producer (orders service) from the consumers (inventory/shipping processors). The orders service can quickly enqueue work and return to the caller, even if a downstream service is slow or restarted. Messages remain in the queue until consumers can process them, preventing cascading latency/backpressure from propagating to the orders API.

Why this answer

Introducing Amazon SQS as a buffer decouples the orders service from the downstream inventory and shipping services. The orders service can immediately enqueue messages and respond to the client, while downstream consumers process messages at their own pace. This prevents backpressure from a slow or restarting downstream service from blocking the orders service, achieving the desired resiliency.

Exam trap

The trap here is that candidates may think parallelizing calls (Option D) or increasing timeouts (Option A) solves the problem, but they fail to recognize that true resiliency requires decoupling via asynchronous messaging, not just concurrency or tolerance of delays.

How to eliminate wrong answers

Option A is wrong because adding longer client timeouts does not prevent the orders service from being blocked; it only increases the wait time before a timeout occurs, still causing the orders service to slow down during peak load. Option C is wrong because replacing downstream services with synchronously invoked Lambda functions does not decouple the services; the orders service would still block waiting for the Lambda invocation to complete, and Lambda has a 15-minute timeout limit, which does not solve the slowdown issue. Option D is wrong because calling downstream services in parallel threads reduces latency only if both services are responsive; if one service is slow or restarting, the orders service still waits for that slow response, and thread pool exhaustion can occur under peak load, leading to resource contention and slowdown.

252
MCQmedium

A content publishing system uses Lambda functions that call an unreliable third-party API. Failed events must be retained for later investigation after retries are exhausted. What should be configured? The design must avoid adding custom operational scripts.

A.Lambda reserved concurrency set to zero
B.A larger deployment package
C.CloudFront error pages
D.A Lambda dead-letter queue or failure destination
AnswerD

Configuring a Lambda dead-letter queue (SQS or SNS) or an on-failure destination for asynchronous invocations ensures that events that exhaust Lambda's retry policy are captured and stored. Lambda sends the failed event payload to the chosen DLQ or destination, allowing downstream consumers to analyze, quarantine, or reprocess it. This is the standard, built-in mechanism for persisting asynchronous invocation failures.

Why this answer

A Lambda dead-letter queue (DLQ) or failure destination is the correct solution because it captures events that have exhausted all retry attempts from an asynchronous Lambda invocation. This allows failed events to be retained in an Amazon SQS queue or SNS topic for later investigation, without requiring custom operational scripts. The DLQ or failure destination integrates directly with Lambda's built-in retry behavior, ensuring that only events that fail after the configured number of retries are sent to the destination.

Exam trap

The trap here is that candidates may confuse a DLQ with other error-handling mechanisms like CloudFront error pages or reserved concurrency, but only a DLQ or failure destination directly captures failed asynchronous Lambda events without custom code.

How to eliminate wrong answers

Option A is wrong because setting reserved concurrency to zero would prevent the Lambda function from executing at all, which stops all invocations and does not retain failed events. Option B is wrong because a larger deployment package does not affect error handling or event retention; it only increases the function's storage size and cold start time. Option C is wrong because CloudFront error pages are used for HTTP error responses from a web distribution, not for capturing failed Lambda invocations from asynchronous event sources.

253
Multi-Selectmedium

A healthcare company runs a batch ingestion pipeline on Amazon EC2 instances that read messages from an Amazon SQS queue and write results to Amazon DynamoDB. The pipeline must be resilient so that a single instance failure does not stop processing and no messages are lost. Which two architectural changes should a solutions architect make to meet these requirements? (Choose two.)

Select 2 answers
A.Increase the number of EC2 instances manually to five and place them all in the same Availability Zone for low latency.
B.Convert the SQS queue to a FIFO queue with content-based deduplication enabled to guarantee no message loss.
C.Move the ingestion workers into an Auto Scaling group across multiple Availability Zones with a launch template, so failed instances are replaced automatically.
D.Enable DynamoDB Accelerator (DAX) on the target table to cache writes so failed instances do not lose data.
E.Configure the SQS queue with a visibility timeout longer than the maximum processing time and have workers delete a message only after successful processing.
AnswersC, E

An Auto Scaling group spanning multiple Availability Zones maintains desired capacity and replaces unhealthy instances automatically, removing the single-instance failure point. Combined with a queue-based workload, replacement workers resume consuming messages, so processing continues without manual intervention and the pipeline remains resilient to instance loss.

Why this answer

Resilience here has two parts: redundant compute that self-heals, and queue semantics that return unprocessed messages to the queue. A multi-AZ Auto Scaling group replaces failed workers, while a visibility timeout longer than processing time plus delete-after-success ensures a message is reprocessed rather than lost if an instance dies mid-flight.

Exam trap

The trap here is treating caching or FIFO ordering as durability mechanisms, when message safety actually depends on visibility timeout and delete-after-success semantics.

254
MCQeasy

A media startup stores user-uploaded video files in an Amazon S3 bucket in the us-east-1 Region. The compliance team requires that the data remain recoverable if an entire AWS Region becomes unavailable, and that recovery can be performed by pointing applications at a different endpoint. Cost should be minimized while still meeting the requirement. Which solution should a solutions architect recommend?

A.Configure an S3 Lifecycle rule to transition objects to S3 Glacier Deep Archive after 30 days for regional durability.
B.Enable S3 Cross-Region Replication to a bucket in a second Region and configure the application to use the replica bucket endpoint during a regional failure.
C.Enable S3 Transfer Acceleration on the bucket so uploads and downloads are faster during a regional event.
D.Enable S3 Versioning on the existing bucket and rely on version history to restore objects if the Region fails.
AnswerB

Cross-Region Replication copies objects asynchronously to a bucket in another Region, giving a durable copy that survives a full regional outage. Because replication is object-level and pay-as-you-go, cost stays lower than continuously active multi-Region compute, and the application can be repointed to the replica bucket's endpoint during failover.

Why this answer

Protecting S3 data against a full regional outage requires a copy in a different Region, which Cross-Region Replication provides asynchronously and cost-effectively. Versioning, lifecycle transitions, and Transfer Acceleration all operate within the source Region and therefore cannot deliver recovery when that Region is lost.

Exam trap

The trap here is confusing durability features that live inside one Region, such as versioning or Glacier transitions, with actual cross-Region resilience.

255
MCQhard

Based on the exhibit, the current disaster recovery design misses the RTO target even though the database replica is current. Which deployment model best meets the requirements with the least always-on cost?

A.Pilot light, because only the database needs to be running in the secondary Region.
B.Warm standby, because a scaled-down application stack stays running in the secondary Region and can take over faster.
C.Active-active, because both Regions should always serve traffic to guarantee the RTO.
D.Backup and restore, because restoring from backups is the least expensive DR model available.
AnswerB

Warm standby is the best fit when you need faster recovery than pilot light but do not want the cost of full active-active capacity. The exhibit shows that starting the application stack from zero consumes most of the recovery time. Keeping a reduced but functional stack running in the secondary Region removes that startup delay and should bring the total recovery time within the 15-minute RTO while still keeping always-on cost below full production duplication.

Why this answer

Warm standby is the correct choice because it keeps a scaled-down application stack running in the secondary Region, which can be scaled up quickly to handle production traffic. This design meets the RTO target by reducing failover time compared to a pilot light, while avoiding the higher always-on cost of an active-active deployment.

Exam trap

The trap here is that candidates confuse pilot light with warm standby, assuming that only the database needs to be running to meet the RTO, but they overlook the time required to provision the application stack on failover.

Why the other options are wrong

A

The pilot light model only keeps the database running in the secondary Region, not the application stack. This means the application must be provisioned and scaled up after a disaster, which takes too long to meet the RTO, even if the database is current.

C

Active-active requires both regions to serve live traffic continuously, which incurs higher always-on costs than warm standby. The question asks for the 'least always-on cost' while meeting RTO, and active-active is more expensive because it runs full production capacity in both regions.

D

Backup and restore typically has a high RTO because it involves restoring data from backups, which is slower than having a running replica. The question states the database replica is current, so a warm standby with a scaled-down application stack can meet the RTO faster.

256
MCQeasy

A retail platform needs disaster recovery across AWS Regions. The business requirement is: RTO up to 6 hours, RPO up to 1 hour, and they want the ability to start serving quickly during a Region outage but do not want to run full production capacity continuously. Which DR strategy best fits these requirements?

A.Backup and restore only, with no continuously running infrastructure in the secondary Region.
B.Pilot light, keeping only the minimum resources needed to bootstrap the environment.
C.Warm standby, keeping a reduced but ready-to-scale environment in the secondary Region.
D.Multi-site active-active, serving production traffic from both Regions at all times.
AnswerC

Warm standby in the secondary Region runs a scaled-down but fully functional copy of your production stack, typically with key databases and services already deployed and synchronized. You can provision extra compute capacity on demand via Auto Scaling or pre-provisioned cluster resizing to reach full production load within the 6-hour RTO. This balances cost and recovery speed by keeping idle but ready infrastructure that can be quickly scaled up, making it the most appropriate choice for the stated requirement.

Why this answer

Warm standby is the correct strategy because it maintains a scaled-down but fully functional copy of the production environment in the secondary Region, which can be scaled up within the 6-hour RTO. The RPO of 1 hour is met by continuous replication (e.g., Amazon RDS cross-Region read replicas or DynamoDB global tables), and the reduced footprint avoids the cost of full production capacity while still enabling rapid failover.

Exam trap

The trap here is that candidates confuse pilot light with warm standby, assuming that any minimal running infrastructure qualifies as pilot light, but warm standby specifically requires a scaled-down but fully functional environment that can serve traffic immediately after scaling, whereas pilot light requires significant provisioning before it can serve traffic.

Why the other options are wrong

A

Backup and restore typically has RPOs of hours or days and RTOs of 24+ hours, failing to meet the 1-hour RPO and 6-hour RTO requirements.

B

The pilot light strategy typically has RTO of 10-15 minutes and RPO of a few minutes, which is faster than the required 6-hour RTO and 1-hour RPO, but it does not meet the requirement to 'start serving quickly' during a Region outage because it requires provisioning and scaling resources after failover, leading to longer recovery time than warm standby.

D

Multi-site active-active requires running full production capacity in both Regions continuously, which contradicts the requirement to not run full production capacity continuously.

257
MCQmedium

A media company stores original video assets in an Amazon S3 bucket in the us-east-1 Region. Editors in Europe report slow downloads, and the legal team requires that a copy of every asset exist in eu-west-1 within 15 minutes of upload, with the ability to fail over reads to the European copy during a Regional impairment. Which S3 feature should the architects enable?

A.S3 Multi-Region Access Points configured against the existing us-east-1 bucket, with no replication rule created.
B.S3 Object Lambda access points in eu-west-1 that transform objects on retrieval from the us-east-1 bucket.
C.S3 Cross-Region Replication with a replication rule that targets the eu-west-1 bucket and uses S3 Replication Time Control.
D.S3 Same-Region Replication to a second bucket in us-east-1, with an S3 Transfer Acceleration endpoint used by the European editors.
AnswerC

Cross-Region Replication copies new objects to the destination bucket automatically, and S3 Replication Time Control provides a service level agreement that most objects replicate within 15 minutes. This meets both the latency goal for European readers and the requirement for a failover copy in eu-west-1. It is the native, managed mechanism for the stated RPO.

Why this answer

The requirement is a durable, timely copy of each object in a second Region, which is exactly what Cross-Region Replication provides. Adding S3 Replication Time Control turns the timing expectation into a defined service level, aligning the solution with the 15-minute RPO the legal team specified. Other S3 features change routing or transform data but do not create the required replica.

Exam trap

The trap here is confusing features that change how clients reach S3 with features that actually copy object data to another Region.

← PreviousPage 4 of 4 · 257 questions total

Ready to test yourself?

Try a timed practice session using only Design Resilient questions.