Courseiva

CCNA Design Resilient Architectures Questions

37 questions · Design Resilient Architectures topic · All types, answers revealed

1
Multi-Selecthard

A company is designing a high-throughput, fault-tolerant payment processing system running on Amazon EC2 instances across multiple Availability Zones. The architecture must withstand the sudden failure of any single AZ without dropping active database transactions or corrupting state. Which TWO design practices should the solutions architect implement?

Select 2 answers
A.Deploy an Amazon Aurora database cluster configured across multiple Availability Zones with an automatic failover replica.
B.Store application state files on a single Amazon EBS General Purpose SSD attached to the primary EC2 instance.
C.Configure an Amazon Route 53 latency-based routing policy to direct user traffic directly to EC2 instances in each zone.
D.Distribute Amazon EC2 instances across an Auto Scaling group configured with subnets spanning at least three Availability Zones.
E.Pin all payment processing worker threads to a single static EC2 instance to simplify debugging and transaction logging.
AnswersA, D

Amazon Aurora automatically replicates storage across three Availability Zones and promotes a read replica to primary within seconds if the primary instance fails. This prevents database downtime and protects against data corruption during infrastructure failures.

Why this answer

Deploying the EC2 instances across an Auto Scaling group spanning multiple Availability Zones ensures compute capacity automatically rebalances during failures. Utilizing Amazon Aurora Multi-AZ deployments provides high availability and automatic failover for relational databases. Together, these practices protect both the compute and data tiers from localized outages.

Exam trap

Candidates select Multi-AZ RDS instances instead of Aurora clusters, missing the specific requirement for high-throughput payment processing systems running across multiple Availability Zones with automatic failover.

2
MCQmedium

A high-traffic application uses an Application Load Balancer (ALB) and an Auto Scaling group (ASG). During peak hours, some EC2 instances fail, but the ASG does not replace them immediately because the instances are still in a 'running' state despite the application being unresponsive. How should the architect fix this?

A.Change the ASG health check type from EC2 to ELB.
B.Increase the ASG health check grace period to 600 seconds.
C.Create a CloudWatch alarm for high CPU usage to trigger replacement.
D.Configure a Lambda function to manually terminate unresponsive instances.
AnswerA

The ELB health check type ensures that the Auto Scaling group considers an instance unhealthy if the Application Load Balancer determines the application is not responding correctly. This provides a more accurate view of application health than basic EC2 status checks, which only monitor the underlying hardware and network connectivity of the virtual machine.

Why this answer

By default, an Auto Scaling group uses EC2 health checks, which only detect instance-level failures (e.g., stopped or terminated instances) and not application-level unresponsiveness. Changing the health check type to ELB makes the ASG rely on the load balancer's health checks, which actively probe the application on each instance; if an instance fails the ELB health check, the ASG marks it unhealthy and replaces it. This directly resolves the issue of running-but-unresponsive instances.

Exam trap

SAA-C03 often tests the difference between EC2 and ELB health check types, and candidates mistakenly believe that increasing the grace period or adding CloudWatch alarms will make the ASG replace unresponsive instances.

How to eliminate wrong answers

Option B is wrong because increasing the health check grace period only delays when health checks begin after launch; it does not change the health check type, so unresponsive instances would still be considered healthy by EC2 checks. Option C is wrong because a CloudWatch alarm on CPU usage does not automatically trigger instance replacement unless paired with a scaling policy, and high CPU may not correlate with unresponsiveness; it is an indirect and unreliable fix. Option D is wrong because a Lambda function manually terminating instances is an operational workaround, not a native ASG health management solution, and it adds unnecessary complexity and latency.

3
MCQmedium

A web application runs on Amazon EC2 instances managed by an Auto Scaling group across three Availability Zones behind an Application Load Balancer. The application performs heavy computations that consume high CPU resources. Users report intermittent 504 Gateway Timeouts during peak hours. The solutions architect notices that newly launched instances take several minutes to bootstrap before they can serve traffic. How should the architect resolve this issue?

A.Configure an Amazon SQS queue in front of the EC2 instances to offload computational tasks asynchronously.
B.Implement pre-baked Amazon Machine Images (AMIs) with application code pre-installed and configure an Auto Scaling health check grace period.
C.Switch the load balancer from an Application Load Balancer to a Network Load Balancer for faster TCP packet routing.
D.Decrease the Auto Scaling group health check grace period to zero so failing instances are replaced immediately.
AnswerB

Pre-baked AMIs remove the multi-minute bootstrap, so new instances serve traffic almost immediately and scale out before peak load causes 504s. The health check grace period prevents the Auto Scaling group terminating instances that are still initialising, avoiding premature replacement.

Why this answer

Pre-baked AMIs eliminate the lengthy bootstrap process by having the application code and dependencies already installed, so new instances can serve traffic almost immediately. Configuring an appropriate health check grace period ensures the ASG does not terminate instances prematurely during any remaining startup time, while allowing ELB health checks to replace truly unhealthy instances. This combination reduces the window where the ASG is scaling out but new instances are not yet ready, mitigating 504 errors.

Exam trap

SAA-C03 often tests whether candidates understand that 504 errors during scaling are caused by slow instance readiness, and candidates incorrectly choose load balancer changes or SQS decoupling instead of reducing bootstrap time.

How to eliminate wrong answers

Option A is wrong because adding an SQS queue decouples and offloads tasks but does not address the root cause of slow instance bootstrapping or the 504 timeouts caused by insufficient capacity during peak; it changes the architecture rather than fixing the scaling latency. Option C is wrong because switching to a Network Load Balancer (NLB) improves TCP performance but does not reduce instance bootstrap time; 504 errors are HTTP-level timeouts from the ALB, and NLB operates at layer 4 without HTTP awareness, so it would not resolve application-level timeouts. Option D is wrong because decreasing the health check grace period to zero would cause the ASG to terminate instances before they finish bootstrapping, leading to a continuous cycle of instance replacement and worsening the issue.

4
Multi-Selecthard

A financial application stores data in Amazon Aurora. The architecture must survive a regional outage. Which TWO steps should the architect take to meet this requirement?

Select 2 answers
A.Enable Aurora Multi-AZ deployment in the primary region.
B.Create an Aurora Global Database with a secondary region.
C.Schedule daily Amazon RDS snapshots and copy them to an S3 bucket in the same region.
D.Promote the secondary region to primary if the primary region becomes unavailable.
E.Enable Performance Insights on all Aurora instances.
AnswersB, D

Aurora Global Database is specifically designed for regional disaster recovery. It uses dedicated infrastructure to replicate data to secondary regions with minimal latency, typically under one second. This configuration allows for rapid promotion of a secondary region to read-write status in the event of a primary regional failure.

Why this answer

Option B is correct because an Aurora Global Database replicates data to a secondary region with typical cross-region replication latency under one second, which is the AWS-recommended design for surviving a full regional outage. Option D is correct because, during a regional failure, the secondary region's Aurora cluster must be promoted (via the managed planned or unplanned failover) to become a standalone read/write primary, restoring write capability in the surviving region. Option A is not sufficient because Aurora Multi-AZ (Aurora Replicas across AZs) only protects against an Availability Zone failure within a single region, not a regional outage.

Option C is wrong because RDS snapshots copied to an S3 bucket in the same region are also lost if that region fails, and daily snapshots give a poor RPO. Option E is irrelevant because Performance Insights is a monitoring/performance-tuning feature and provides no disaster-recovery capability.

Exam trap

The trap is confusing Aurora Multi-AZ (intra-region HA) with Aurora Global Database (cross-region DR); candidates often pick Multi-AZ thinking it covers regional outages, but it only protects against AZ failures.

5
MCQhard

A company hosts a monolithic application on EC2. They want to move to a microservices architecture to improve fault isolation. Which approach is best?

A.Increase the size of the existing EC2 instances.
B.Use a load balancer to split traffic between two monolithic instances.
C.Decompose the monolith into independent services using containers.
D.Move the entire application code to a single Lambda function.
AnswerC

Decomposing the monolith into independent services using containers provides the best fault isolation. If one service fails, the other services can continue to operate, limiting the impact of the failure. Using containers also allows for faster deployment, better resource utilization, and easier management, all of which contribute to a more resilient overall architecture.

Why this answer

Decomposing the monolith into independent services running in containers gives each service its own lifecycle, scaling, and failure domain, which is exactly what improves fault isolation. Containers package each service with its dependencies and can be orchestrated by ECS or EKS for resilience.

Exam trap

SAA-C03 often tests the confusion between high availability (multiple copies of the same monolith) and fault isolation (independent services), tempting candidates to pick load balancing or vertical scaling.

How to eliminate wrong answers

Option A is wrong because vertically scaling EC2 instances keeps the monolith intact and does nothing for fault isolation — a failure in one component still takes down the whole application. Option B is wrong because running two copies of the same monolith behind a load balancer improves availability but not fault isolation, since both instances share the same monolithic failure modes. Option D is wrong because moving the entire application into a single Lambda function preserves the monolith and adds Lambda limits (15-minute timeout, memory caps) without any service decomposition.

6
MCQmedium

An architect is designing a system that must be resilient against AWS regional failure. Which AWS service is best suited to manage global DNS traffic and provide health-based routing?

A.Elastic Load Balancing (ELB).
B.Amazon Route 53.
C.AWS Direct Connect.
D.Amazon CloudFront.
AnswerB

Route 53 is the primary tool for global traffic management. Its health checks continuously monitor the status of regional endpoints. If an endpoint fails, Route 53 can update DNS records to route traffic to healthy infrastructure in other regions, providing the necessary resilience against regional-wide service interruptions.

Why this answer

Route 53 is a highly available and scalable cloud Domain Name System (DNS) web service. It provides health checks and routing policies, such as failover, latency-based, and geolocation routing. By using health checks, Route 53 can detect the failure of an endpoint in one region and automatically route traffic to a healthy endpoint in another, fulfilling the requirement for regional resilience.

Exam trap

Candidates sometimes choose Global Accelerator or CloudFront. While these provide global traffic management, Route 53 is the primary DNS-based service specifically designed for health-based routing and regional failover.

7
MCQhard

A company needs a Disaster Recovery (DR) strategy with an RTO of 30 minutes and an RPO of 15 minutes. They want to minimize costs while having a scaled-down version of their core environment always running in a second region. Which DR strategy should they use?

A.Backup and Restore.
B.Pilot Light.
C.Warm Standby.
D.Multi-site Active-Active.
AnswerC

Warm Standby keeps a minimized version of the full environment running in the second region. This ensures that the application is always ready to handle traffic, and only needs to be scaled up to meet production loads during a failover. This strategy easily fits the 30-minute RTO and 15-minute RPO requirements while remaining cost-conscious.

Why this answer

A Warm Standby DR strategy involves running a functional but scaled-down version of the application in a separate region. This allows for a very low Recovery Time Objective (RTO) because the core infrastructure is already live. Data is replicated frequently to meet the Recovery Point Objective (RPO) of 15 minutes, providing a balance between cost and speed.

Exam trap

Candidates often confuse Warm Standby with Pilot Light; they choose Pilot Light because it sounds cheaper, missing that Pilot Light's RTO is too high for a 30-minute requirement.

8
MCQmedium

An application uses Amazon DynamoDB. To ensure the application remains resilient to regional outages, what is the best approach?

A.Enable DynamoDB backups in the same region.
B.Use DynamoDB Global Tables.
C.Use AWS Database Migration Service to sync data.
D.Manually replicate data to another region every hour.
AnswerB

DynamoDB Global Tables enable multi-region, multi-active replication, allowing the database to survive a full regional failure. It provides automatic synchronization of data across multiple regions, enabling low-latency access for globally distributed users and ensuring that the application can remain operational even if one of the AWS regions experiences a complete outage.

Why this answer

DynamoDB Global Tables provide multi-region, active-active replication with automatic conflict resolution, ensuring the application can read and write to the table in any region and remain resilient to regional outages. This is the only option that offers automatic, low-latency replication across regions without manual intervention.

Exam trap

The trap is assuming that backups or manual replication provide regional resilience. Candidates often overlook that backups are region-scoped and manual replication lacks automatic failover and conflict resolution.

How to eliminate wrong answers

Option A is wrong because backups in the same region do not protect against regional outages; they are stored redundantly within the region but are unavailable if the region fails. Option C is wrong because AWS DMS is designed for database migration, not for continuous bidirectional synchronization with conflict resolution. Option D is wrong because manual hourly replication is not automatic, has high RPO, and does not provide a seamless failover mechanism.

9
MCQmedium

Refer to the exhibit. An application needs to store highly durable data in S3. The architect wants to ensure data is protected against accidental deletion while maintaining cost-effective storage. What should be done?

A.Enable S3 Versioning on the bucket and use Lifecycle policies to manage older versions.
B.Enable S3 Cross-Region Replication for all objects.
C.Modify the IAM policy to include s3:DeleteObject in the Deny statement.
D.Use S3 Intelligent-Tiering for all objects to reduce storage costs.
AnswerA

S3 Versioning keeps multiple variants of an object in the same bucket, providing protection against accidental deletion. Lifecycle policies automatically handle the transition of older versions to S3 Glacier or perform expirations, ensuring that storage costs remain optimized while maintaining high data durability and recovery capabilities for the bucket.

Why this answer

S3 Versioning preserves every prior version of an object, so an accidental overwrite or delete does not destroy the underlying data — the previous version remains retrievable. Pairing Versioning with Lifecycle policies lets you transition or expire noncurrent versions after a set number of days, which controls the cost of retaining all those versions. Together they deliver durability against accidental deletion while keeping storage spend predictable.

Exam trap

SAA-C03 often tests the misconception that Cross-Region Replication protects against accidental deletion, when in fact replication propagates deletions — Versioning is the correct control for recoverability.

How to eliminate wrong answers

Option B is wrong because Cross-Region Replication copies objects to another region for disaster recovery and latency, but it does not protect against accidental deletion — a delete marker or overwrite propagates to the replica. Option C is wrong because denying s3:DeleteObject via IAM blocks legitimate deletion operations too, which is a blunt control that does not preserve already-deleted data and can break application workflows. Option D is wrong because S3 Intelligent-Tiering only optimizes storage class costs based on access patterns; it provides no versioning or deletion protection whatsoever.

10
MCQmedium

A global e-commerce enterprise runs its checkout workflow using AWS Lambda functions integrated with Amazon API Gateway. During flash sales, traffic spikes cause downstream payment APIs to timeout, leading to lost transactions and frustrated customers. The solutions architect needs to redesign the architecture to decouple the frontend from the payment processor and ensure no transaction requests are lost. What should the architect do?

A.Configure API Gateway to cache responses using an Amazon ElastiCache Redis cluster deployed in multi-AZ mode.
B.Integrate an Amazon SQS FIFO queue between API Gateway and the downstream payment processing Lambda functions.
C.Increase the timeout and memory allocation of the API Gateway integration and the AWS Lambda functions.
D.Use AWS Step Functions with standard workflows to execute the payment processing steps sequentially.
AnswerB

An SQS FIFO queue buffers checkout requests when payment Lambdas are throttled, decoupling API Gateway from the processor so requests persist rather than being lost to timeouts. FIFO ordering and exactly-once processing preserve transaction sequence, directly meeting the no-lost-transactions requirement during flash-sale spikes.

Why this answer

Integrating an Amazon SQS FIFO queue between API Gateway and the payment Lambda decouples the frontend from the downstream payment processor, buffering requests during traffic spikes so no transaction is lost. FIFO queues preserve order and provide exactly-once processing, which is critical for payment transactions. The Lambda functions can then poll the queue at a controlled rate, preventing downstream timeouts.

Exam trap

SAA-C03 often tests the misconception that increasing Lambda timeout/memory or adding caching solves decoupling problems, when the real requirement is durable buffering with SQS.

How to eliminate wrong answers

Option A is wrong because ElastiCache caching only stores responses and does not decouple or buffer write transactions — it cannot prevent lost payment requests. Option C is wrong because increasing timeouts and memory does not solve the fundamental problem of downstream overload; it merely delays failures and can worsen throttling. Option D is wrong because Step Functions Standard workflows orchestrate steps but do not provide durable buffering or decoupling from the payment processor under burst load.

11
Multi-Selectmedium

A company is designing a mission-critical database architecture using Amazon RDS. Which TWO steps should the architect take to ensure high availability and data durability? (Select TWO.)

Select 2 answers
A.Enable Multi-AZ deployment for the RDS instance.
B.Configure a Read Replica in the same Availability Zone.
C.Enable automated backups with a defined retention period.
D.Set the storage type to General Purpose SSD (gp2) only.
E.Disable the deletion protection feature for all instances.
AnswersA, C

Multi-AZ deployment ensures that a synchronous standby instance is maintained in a separate Availability Zone. In the event of a primary instance failure, RDS automatically promotes the standby to primary, significantly reducing downtime and ensuring the database remains available for critical application traffic without manual intervention or data loss.

Why this answer

Option A is correct because an RDS Multi-AZ deployment maintains a synchronous standby replica in a different Availability Zone, and RDS automatically fails over to that standby if the primary instance or its AZ fails, which directly provides high availability for a mission-critical database. Option C is correct because enabling automated backups with a defined retention period gives point-in-time recovery (PITR) to any second within the retention window (up to 35 days), which ensures data durability and recoverability from corruption or accidental deletion. Option B is not appropriate because a Read Replica in the same AZ is asynchronous, does not provide automatic failover, and would not survive an AZ-level outage.

Option D is incorrect because gp2 is just one storage option and choosing it does not by itself deliver high availability or durability. Option E is incorrect because disabling deletion protection increases the risk of accidental data loss rather than protecting the database.

Exam trap

SAA-C03 often tests the confusion between Read Replicas (read scaling, asynchronous) and Multi-AZ (high availability, synchronous failover), causing candidates to select a Read Replica when HA is required.

12
MCQmedium

Refer to the exhibit. The current IAM policy allows an EC2 instance to read backups. The architect wants to ensure the data is recoverable if the source bucket is compromised by an external actor with write access. What should be done?

A.Enable S3 Object Lock on the bucket.
B.Apply a Bucket Policy to deny s3:PutObject for all users.
C.Create a read-only IAM user for the EC2 instance.
D.Enable S3 Server-Side Encryption (SSE-S3).
AnswerA

Object Lock provides WORM (Write Once, Read Many) protection. By preventing objects from being overwritten or deleted for a set period, it ensures that even an account compromise cannot destroy the backup data. This is the most effective way to secure backups against malicious actors and accidental data loss.

Why this answer

S3 Object Lock enforces a Write Once Read Many (WORM) model, preventing objects from being deleted or overwritten for a specified retention period — even by users with full administrative permissions. This directly addresses the scenario where an external actor with write access could otherwise overwrite or delete backups, ensuring the data remains recoverable. Object Lock requires versioning to be enabled on the bucket.

Exam trap

SAA-C03 often tests whether candidates know that Object Lock provides WORM immutability even against administrators, and candidates mistakenly pick encryption or IAM policies as deletion protection.

How to eliminate wrong answers

Option B is wrong because a bucket policy denying s3:PutObject for all users would also block legitimate backup writes and does not protect existing objects from deletion; it is a blunt instrument that breaks the backup process. Option C is wrong because creating a read-only IAM user for the EC2 instance only restricts that instance's permissions; it does not prevent a compromised external actor with separate write credentials from deleting or overwriting objects. Option D is wrong because SSE-S3 provides encryption at rest but does not prevent deletion or modification of objects — encryption and immutability are orthogonal controls.

13
Multi-Selecthard

An application uses a monolithic architecture on EC2 instances. The company wants to improve resilience against regional failures. Which TWO actions should the architect take to achieve this? (Select TWO.)

Select 2 answers
A.Deploy the application to multiple Availability Zones within a single region.
B.Configure Amazon Route 53 with failover routing policies to a secondary region.
C.Implement Amazon Aurora Global Database for cross-region data replication.
D.Use an AWS Direct Connect connection to link the two regions.
E.Increase the minimum number of EC2 instances in the Auto Scaling group.
AnswersB, C

Route 53 failover routing policies monitor the health of primary region endpoints. If health checks fail, traffic is automatically routed to the secondary region. This is a foundational step in building a resilient multi-region architecture that provides automated business continuity during a regional disaster.

Why this answer

Option B is correct because Route 53 failover routing policies with health checks automatically redirect traffic to a secondary region when the primary region becomes unhealthy, directly providing regional failover capability. Option C is correct because Aurora Global Database replicates data across regions with typical latency under one second, enabling a secondary region to serve reads and support fast promotion during a regional failure, which is essential for cross-region resilience of a stateful monolith. Option A is not correct because deploying across multiple Availability Zones within a single region only protects against AZ-level failures, not regional failures.

Option D is not correct because Direct Connect is a dedicated network link between on-premises and AWS (or between regions via Direct Connect Gateway) and does not by itself provide application failover or resilience. Option E is not correct because increasing the minimum EC2 instance count in an Auto Scaling group only adds capacity within a single region and does nothing for regional failure resilience.

Exam trap

SAA-C03 often tests the misconception that multi-AZ deployment equals regional resilience, when true regional failover requires cross-region replication and DNS failover.

14
MCQmedium

An application is deployed across two AWS Regions to ensure resilience. The architect needs to direct users to the healthy region with the lowest latency. If the primary region fails, traffic should fail over automatically. Which Route 53 routing policy should be implemented?

A.Simple routing policy.
B.Latency routing policy with health checks.
C.Failover routing policy.
D.Weighted routing policy.
AnswerB

Latency routing optimizes the user experience by serving requests from the AWS Region that provides the lowest network latency. When combined with health checks, it provides both performance optimization and high availability. Route 53 monitors the health of the endpoints and automatically skips any region that fails health checks during the routing decision.

Why this answer

Latency-based routing with health checks directs each user to the AWS Region that provides the lowest latency while the health check removes an unhealthy region's records from DNS responses, achieving automatic failover. This satisfies both requirements: lowest-latency selection under normal conditions and automatic redirection when the primary region fails.

Exam trap

SAA-C03 often tests the misconception that Failover routing automatically picks the lowest-latency region — it does not; it is strictly active-passive, so candidates who overlook the 'lowest latency' requirement choose C incorrectly.

How to eliminate wrong answers

Option A is wrong because Simple routing returns a single static record with no health checking or latency awareness, so it cannot fail over or optimize latency. Option C is wrong because Failover routing is active-passive and always sends traffic to the primary unless it is unhealthy; it does not select the lowest-latency region for each user. Option D is wrong because Weighted routing distributes traffic by fixed proportions (e.g., 70/30) and does not consider latency or provide automatic health-based failover by itself.

15
MCQmedium

A company's web application requires a highly available database layer to survive an Availability Zone failure without manual intervention. The database must remain accessible during maintenance windows and ensure zero data loss. Which solution meets these requirements with the least operational overhead?

A.Configure RDS Read Replicas and manually promote them during an outage.
B.Deploy a Single-AZ RDS instance and use AWS Backup for hourly snapshots.
C.Implement an Amazon RDS Multi-AZ deployment.
D.Host a MySQL database on an Amazon EC2 instance with an EBS volume.
AnswerC

Amazon RDS Multi-AZ uses synchronous replication to a standby instance in a different AZ, providing automatic failover and data redundancy. If the primary instance fails, AWS automatically updates the DNS record to point to the standby. This meets the high availability and zero-intervention requirements while handling maintenance windows with minimal impact on application uptime.

Why this answer

Amazon RDS Multi-AZ maintains a synchronous standby replica in a different Availability Zone and automatically fails over during an AZ failure or maintenance event, with zero data loss (RPO=0) and no manual intervention. It meets high availability, maintenance resilience, and zero data loss with minimal operational overhead.

Exam trap

SAA-C03 often tests the distinction between Multi-AZ (synchronous, automatic failover, HA) and Read Replicas (asynchronous, manual promotion, read scaling); the trap is choosing Read Replicas for a zero-data-loss HA requirement.

How to eliminate wrong answers

Option A is wrong because Read Replicas use asynchronous replication and require manual promotion, so they can lose recent transactions and do not provide automatic failover. Option B is wrong because a Single-AZ instance with hourly snapshots has no automatic failover and can lose up to an hour of data, violating zero data loss. Option D is wrong because self-managed MySQL on EC2 requires the customer to architect, patch, and manage replication and failover, adding significant operational overhead and not guaranteeing zero data loss without careful synchronous setup.

16
MCQhard

A company runs a mission-critical web application on EC2. The database is on RDS. How should the architect ensure the fastest possible recovery time objective (RTO) if the primary database instance fails?

A.Configure an RDS Read Replica in a different region and promote it if the primary fails.
B.Enable RDS Multi-AZ deployment.
C.Take hourly automated snapshots and restore from the latest snapshot upon failure.
D.Back up the database to S3 and use a standby EC2 instance to run a database engine.
AnswerB

RDS Multi-AZ is specifically designed for high availability. It maintains a synchronous secondary standby in another AZ. Upon detecting a failure, RDS automatically initiates a failover process. This is the fastest, most reliable way to maintain database availability and minimize downtime for mission-critical applications without manual intervention.

Why this answer

RDS Multi-AZ deployment provides synchronous replication to a standby instance in a different Availability Zone (AZ) within the same region. If the primary instance fails, RDS automatically performs a failover to the standby, updating the DNS record to point to the new primary. This process typically happens in under 60 seconds, which is the most effective native way to achieve a low RTO for RDS.

Exam trap

Candidates often confuse Multi-AZ (high availability within a region) with Cross-Region Read Replicas (disaster recovery across regions).

17
MCQmedium

An application consists of a web tier and an application tier. The web tier must be accessible from the internet, but the application tier must remain private. How should the architect configure the network for maximum resilience?

A.Deploy all tiers in a single public subnet in one AZ.
B.Deploy web and app tiers across multiple subnets in multiple AZs, with app tiers in private subnets.
C.Put everything in a public subnet and use Security Groups for isolation.
D.Use a single private subnet for both tiers and a NAT Gateway in the same subnet.
AnswerB

This architecture provides both high availability and security. By spreading instances across multiple AZs, the application can survive a data center failure. Placing the application tier in private subnets ensures that sensitive back-end components are not directly reachable from the internet, significantly hardening the application against external attacks.

Why this answer

For maximum resilience, the web and application tiers should be deployed across multiple subnets in multiple Availability Zones, with the web tier in public subnets (accessible from the internet) and the application tier in private subnets (not directly accessible). This design ensures high availability and fault tolerance, and follows the principle of least privilege by isolating the private tier.

Exam trap

SAA-C03 often tests the misconception that Security Groups alone can provide sufficient isolation for a private tier in a public subnet, but network placement (private subnets) is required for true isolation.

How to eliminate wrong answers

Option A is wrong because deploying all tiers in a single public subnet in one AZ creates a single point of failure and exposes the application tier to the internet, violating security best practices. Option C is wrong because placing everything in a public subnet, even with Security Groups, still exposes the application tier to potential inbound traffic if Security Groups are misconfigured, and it does not provide network-level isolation. Option D is wrong because using a single private subnet for both tiers and a NAT Gateway in the same subnet is not possible; NAT Gateways must be in a public subnet, and a single private subnet lacks multi-AZ resilience.

18
MCQmedium

A company is using AWS Direct Connect to connect their on-premises data center to AWS. What is the most resilient way to connect?

A.Use a single 10 Gbps Direct Connect connection.
B.Use two Direct Connect connections in different locations.
C.Use a single VPN connection over the internet.
D.Configure the connection to use only the AWS public IP space.
AnswerB

Redundant Direct Connect connections in different locations provide protection against local failures, such as fiber cuts or data center outages. This ensures that even if one physical path is entirely compromised, connectivity can be maintained through the alternate path, meeting the requirements for a highly available and resilient hybrid cloud network infrastructure.

Why this answer

To achieve high resiliency and maximum redundancy with AWS Direct Connect, a company should establish at least two Direct Connect connections in different AWS locations (or use AWS Direct Connect resiliency recommendations which include multiple connections across multiple locations). Option B provides geographical redundancy, eliminating single points of failure.

Exam trap

Candidates often assume a single high-bandwidth connection (like 10 Gbps) is sufficient for resiliency, but bandwidth does not equal redundancy.

19
MCQhard

A mission-critical application requires a recovery time objective (RTO) of near-zero. Which disaster recovery strategy should the architect implement?

A.Backup and Restore.
B.Pilot Light.
C.Warm Standby.
D.Multi-Site Active-Active.
AnswerD

The Multi-Site Active-Active approach runs the application in multiple regions simultaneously. Because both regions are already handling production traffic, a failure in one region does not require scaling or recovery time; traffic is simply routed to the remaining active region, resulting in the lowest possible RTO for the application.

Why this answer

The explanation is accurate, but the question requires a defined exam trap and common trap note to help students avoid pitfalls.

Exam trap

Candidates often confuse Warm Standby with Multi-Site Active-Active. While Warm Standby allows for fast recovery, it is not 'near-zero' because it requires scaling up resources or promoting a database, whereas Active-Active is already running at full capacity.

20
Multi-Selecthard

A media streaming company stores high-value video assets in an Amazon S3 bucket. The company requires a resilient storage architecture that protects against accidental deletion, malicious overwrites, and zonal or regional outages. Which combinations of features should a Solutions Architect implement to achieve these requirements? (Choose TWO.)

Select 2 answers
A.Enable S3 Versioning on the bucket and configure MFA Delete to prevent accidental or malicious deletion of object versions.
B.Configure Amazon S3 Intelligent-Tiering to automatically move infrequently accessed video assets to lower-cost storage tiers.
C.Implement S3 Cross-Region Replication to asynchronously copy all video assets to an S3 bucket in a different AWS Region.
D.Attach a resource-based bucket policy that explicitly denies all delete object requests from any IAM user except the root account.
E.Enable S3 Object Lock in governance mode with a fixed retention period of 30 days for all uploaded video files.
AnswersA, C

S3 Versioning preserves every version of every object, ensuring deleted or overwritten files can be easily restored. Multi-Factor Authentication Delete adds an extra layer of authorization security, requiring physical or virtual MFA tokens for permanent deletion operations.

Why this answer

Enabling S3 Versioning ensures that previous versions are retained when objects are deleted or overwritten, protecting against accidental data loss. Furthermore, replicating objects to another AWS Region using S3 Cross-Region Replication guarantees durability and availability against catastrophic regional disruptions, satisfying robust disaster recovery best practices for critical media assets.

Exam trap

Candidates often confuse S3 versioning with replication, or pick features like cross-origin resource sharing (CORS) which are unrelated to regional disaster recovery and data protection.

21
MCQmedium

A web application stores static files in Amazon S3. Due to regulatory requirements, the company must ensure that the files are replicated to a secondary region. What is the most efficient way to achieve this?

A.Use AWS DataSync to copy files periodically.
B.Enable S3 Cross-Region Replication on the source bucket.
C.Create an EventBridge rule to trigger a Lambda function for every upload.
D.Use AWS Storage Gateway to replicate local data to S3.
AnswerB

S3 Cross-Region Replication is the native, AWS-recommended way to handle automated, asynchronous replication of objects between buckets in different regions. It ensures data durability and compliance with disaster recovery policies by keeping a secondary copy of the data in a separate geographic region with minimal latency and no manual intervention.

Why this answer

S3 Cross-Region Replication (CRR) is the most efficient and native way to automatically replicate objects from a source bucket to a destination bucket in a different region. It replicates new objects as they are uploaded and can also replicate existing objects if configured, providing a managed, low-latency replication solution that meets regulatory requirements.

Exam trap

The trap is selecting custom or hybrid solutions like Lambda triggers or DataSync when a native, automated S3 feature (CRR) directly addresses the requirement with less operational overhead.

How to eliminate wrong answers

Option A is wrong because AWS DataSync is designed for migrating or copying data between on-premises storage and AWS, or between AWS storage services, but it is not a continuous replication mechanism for S3-to-S3 and requires scheduling and management overhead. Option C is wrong because triggering a Lambda function on every upload to copy objects is a custom, serverless approach that adds complexity, potential failures, and does not handle large objects or metadata as efficiently as native CRR. Option D is wrong because AWS Storage Gateway is for hybrid cloud storage integration, not for replicating S3 data between regions.

22
MCQhard

Refer to the exhibit. An application running on EC2 is receiving signature errors when accessing S3. The application uses an IAM role. What should the architect investigate first?

A.The IAM role permissions attached to the EC2 instance.
B.The local time on the EC2 instance.
C.The S3 bucket policy for explicit Deny statements.
D.The VPC endpoint for S3 for correct routing.
AnswerB

AWS authentication requires that the timestamp included in a request matches the current time within a small margin of error (usually five minutes). If the EC2 instance's clock is drifting or significantly offset, the signature will be rejected. Synchronizing the clock using NTP is the standard fix for signature mismatch issues.

Why this answer

AWS Signature Version 4, used for authenticating requests to S3, includes a timestamp and requires the client's clock to be within 15 minutes of AWS's time. If the EC2 instance's local time is skewed, the signature will be considered expired, resulting in SignatureDoesNotMatch errors. Since the application uses an IAM role, the credentials are automatically rotated and valid, so the most likely cause is time drift on the instance.

Exam trap

SAA-C03 often tests the misconception that signature errors are always due to permission issues, when in fact time skew is a frequent cause.

How to eliminate wrong answers

Option A is wrong because if the IAM role permissions were insufficient, the error would be AccessDenied, not a signature error. Option C is wrong because an explicit Deny in the bucket policy would also produce an AccessDenied error, not a signature mismatch. Option D is wrong because a misconfigured VPC endpoint would typically cause connectivity issues or timeouts, not signature errors.

23
Multi-Selecteasy

A security team is concerned about accidental or malicious deletion of critical objects in an Amazon S3 bucket. Which TWO features should be enabled to prevent permanent data loss and require additional authentication for deletions?

Select 2 answers
A.S3 Versioning.
B.MFA Delete.
C.S3 Object Lock in compliance mode.
D.S3 Inventory.
E.S3 Transfer Acceleration.
AnswersA, B

Enabling S3 Versioning ensures that whenever an object is deleted, S3 inserts a delete marker instead of permanently removing the data. This allows administrators to easily restore previous versions of the object. It protects against accidental overwrites and provides a history of changes, which is essential for data durability and recovery.

Why this answer

S3 Versioning (A) is correct because it keeps multiple variants of an object in the same bucket, so when an object is overwritten or deleted, a delete marker is placed and the prior version remains recoverable, preventing permanent data loss from accidental or malicious deletion. MFA Delete (B) is correct because it adds an additional authentication factor requirement for permanently deleting object versions or changing the versioning state of the bucket, directly satisfying the requirement for extra authentication on deletions. Together, versioning preserves the data and MFA Delete protects the destructive operations.

S3 Object Lock in compliance mode (C) is not among the marked answers and, while it prevents deletion for a retention period, it does not itself require additional authentication for deletions. S3 Inventory (D) only provides scheduled reports of objects and metadata, and S3 Transfer Acceleration (E) only speeds up uploads/downloads via edge locations; neither prevents permanent deletion or adds authentication.

Exam trap

SAA-C03 often tests the confusion between Object Lock (immutability/retention) and MFA Delete (authentication for deletion) — candidates who pick Object Lock miss that the question explicitly asks for additional authentication.

24
MCQmedium

A company has a legacy application that requires a persistent file system shared across multiple EC2 instances. The file system must be highly available and support standard file system protocols. Which solution is most appropriate?

A.Mount an Amazon EBS volume to multiple EC2 instances simultaneously.
B.Use Amazon EFS with mount targets in multiple Availability Zones.
C.Create an S3 bucket and use it as a mounted file system via S3FS.
D.Set up a self-managed NFS server on an EC2 instance.
AnswerB

EFS provides a managed, scalable NFS file system that supports simultaneous access from multiple instances. By configuring mount targets in multiple AZs, the solution ensures high availability, as EFS handles data replication across those zones, making it resilient to single AZ failure while meeting shared file system application requirements.

Why this answer

Amazon EFS provides a fully managed, scalable, and highly available NFS file system that can be mounted simultaneously by multiple EC2 instances across multiple AZs. It is designed to be resilient, automatically replicating data across AZs within a region. This makes it the ideal choice for legacy applications that require shared storage without the management overhead of self-managed storage servers or distributed file systems.

Exam trap

Candidates mistakenly select Amazon EBS because it is the default block storage, ignoring that EBS cannot be shared simultaneously across multiple instances in different Availability Zones.

25
MCQmedium

A company hosts a web application on Amazon EC2 instances behind an Application Load Balancer. The application stores session state in the local memory of the EC2 instances. Users report losing their sessions when the Auto Scaling group scales in or replaces unhealthy instances. Which solution ensures session persistence while maintaining high availability?

A.Configure the Application Load Balancer to use stickiness with a duration-based cookie.
B.Store the session state in an Amazon S3 bucket with versioning enabled.
C.Move the session state to an Amazon ElastiCache for Redis cluster.
D.Replicate the session data across all EC2 instances using a broadcast script.
AnswerC

ElastiCache for Redis offers low-latency, in-memory performance ideal for session management. By centralizing the session data, the application becomes stateless, allowing the Auto Scaling group to terminate or replace instances without impacting the user experience, as the state is preserved independently of the compute resources.

Why this answer

Storing session state in ElastiCache for Redis externalizes it from the EC2 instances, so any instance in the Auto Scaling group can serve any user's request regardless of scale-in or instance replacement. Redis provides low-latency, highly available shared state with replication and Multi-AZ failover, ensuring session persistence and high availability simultaneously.

Exam trap

SAA-C03 often tests the misconception that ALB stickiness solves session loss — stickiness only helps while the instance lives, so it fails during scale-in or instance replacement, which is exactly the scenario described.

How to eliminate wrong answers

Option A is wrong because ALB stickiness only pins a user to one instance; when that instance is terminated during scale-in or fails, the session is still lost, so it does not solve the underlying problem. Option B is wrong because S3 is object storage with high latency and eventual consistency characteristics for overwrites; it is not designed for high-frequency session reads/writes and would degrade application performance. Option D is wrong because broadcasting session data across instances is fragile, does not scale, and creates race conditions and consistency issues as the group changes size.

26
MCQhard

An enterprise financial application stores critical transaction records in Amazon Aurora MySQL. The company requires a cross-Region disaster recovery strategy with a Recovery Point Objective of less than 1 minute and a Recovery Time Objective of less than 5 minutes. Which architecture achieves this?

A.Take daily automated snapshots of the primary Aurora database and copy the snapshot to a secondary AWS Region.
B.Implement Amazon Aurora Global Databases with a primary writer Region and a cross-Region reader Region.
C.Configure standard MySQL asynchronous replication from an Amazon EC2 instance running MySQL in one Region to another.
D.Deploy a multi-AZ Aurora cluster and use AWS Database Migration Service for continuous replication to another Region.
AnswerB

Aurora Global Database replicates committed writes to a secondary Region with typical latency under one second, meeting the sub-one-minute RPO. Promoting the secondary reader to writer completes in under a minute, satisfying the five-minute RTO for cross-Region disaster recovery.

Why this answer

Amazon Aurora Global Database is designed for cross-Region disaster recovery with typical replication latency under 1 second, easily meeting the RPO of less than 1 minute. It supports fast failover to a secondary Region, typically completing in under 1 minute, which satisfies the RTO of less than 5 minutes. This is the purpose-built solution for low-RPO, low-RTO cross-Region DR on Aurora.

Exam trap

SAA-C03 often tests RPO/RTO numbers — candidates pick Multi-AZ or DMS thinking they provide cross-Region DR, but only Aurora Global Database delivers sub-second RPO and sub-minute RTO across Regions.

How to eliminate wrong answers

Option A is wrong because daily snapshots yield an RPO of up to 24 hours and an RTO measured in hours (snapshot restore time), far exceeding the <1 minute RPO and <5 minute RTO requirements. Option C is wrong because self-managed MySQL on EC2 with asynchronous replication lacks Aurora's managed failover, has higher operational overhead, and cannot guarantee sub-minute RPO/RTO reliably. Option D is wrong because Multi-AZ Aurora only provides HA within a single Region — it does not provide cross-Region DR, and DMS is not designed for sub-minute RPO continuous replication of an Aurora cluster.

27
MCQhard

An application processes large batches of data from an S3 bucket. The process can take hours. If an instance fails, the batch is lost. Which design pattern ensures the most resilient batch processing?

A.Use an Auto Scaling group with a single instance to process the S3 files sequentially.
B.Use an SQS queue to store tasks and an Auto Scaling group to process them.
C.Directly trigger Lambda functions from S3 and increase the timeout.
D.Use a single large EC2 instance to process the data in memory.
AnswerB

This architecture decouples the workload from the compute. If an instance fails, the message remains in the queue to be retried. The Auto Scaling group ensures that sufficient compute power is always available to clear the queue, providing a highly resilient and scalable solution for long-running batch processing.

Why this answer

Using an SQS queue to decouple the producer and consumer is the ideal pattern. The S3 event notification sends a message to the queue, and instances poll the queue for tasks. If an instance fails, the message becomes visible again in the queue after the visibility timeout, allowing another instance to pick up the task and resume processing.

Exam trap

Candidates mistakenly choose SNS or direct S3 notifications to EC2, missing that without SQS decoupling, failed instances will permanently lose the batch processing tasks.

28
MCQhard

Refer to the exhibit. A solutions architect reviewed an AWS CloudFormation template used to deploy an Auto Scaling group for a production web application. During an AWS Availability Zone outage in us-east-1, users experienced partial application downtime even though the Auto Scaling group reported instances running. Why is this architecture failing resiliency best practices?

A.The Launch Template version is hardcoded to version 1 instead of using $Latest or $Default.
B.The Auto Scaling group defines Availability Zones directly instead of referencing subnets across multiple zones.
C.The MinSize and DesiredCapacity are set too low to handle sudden spikes in traffic during an outage.
D.The Auto Scaling group is missing a health check type configuration set to ELB instead of EC2.
AnswerB

Hardcoding Availability Zones rather than referencing subnets across multiple zones prevents the Auto Scaling group from launching replacement instances elsewhere during a zone outage. Referencing multi-AZ subnets satisfies the resiliency constraint, maintaining capacity when one zone fails.

Why this answer

For an Auto Scaling group to survive an Availability Zone outage, it must launch instances across multiple subnets in different AZs. When the ASG specifies Availability Zones directly (or uses a single subnet), all instances land in one AZ, so an AZ failure takes down the entire fleet even though the ASG still reports 'running' instances in the failed zone. Referencing subnets across multiple AZs lets the ASG rebalance and launch replacements in healthy zones.

Exam trap

SAA-C03 often tests the misconception that 'instances running' means the architecture is resilient — candidates overlook that all instances may be in one AZ, so an AZ outage causes downtime despite the ASG reporting healthy.

How to eliminate wrong answers

Option A is wrong because a hardcoded Launch Template version affects configuration drift and update rollout, not AZ resiliency — the instances would still be distributed across AZs if subnets are correct. Option C is wrong because MinSize/DesiredCapacity affect capacity and scaling headroom, not whether instances are spread across AZs; low capacity alone does not cause an AZ outage to take down the app. Option D is wrong because health check type (EC2 vs ELB) affects instance replacement when health checks fail, but it does not control AZ distribution — even with ELB health checks, a single-AZ ASG fails during an AZ outage.

29
MCQmedium

A company hosts a web application on EC2 instances behind an Application Load Balancer (ALB). The application requires high availability across two Availability Zones (AZs). Which architectural design ensures the most resilient traffic distribution?

A.Deploy all instances in one AZ and use a Route 53 failover policy to a secondary region.
B.Use an Auto Scaling group with a single AZ and enable Cross-Zone Load Balancing on the ALB.
C.Place EC2 instances in an Auto Scaling group distributed across two AZs behind the ALB.
D.Use a Network Load Balancer (NLB) with static IP addresses for each instance in one AZ.
AnswerC

Distributing Auto Scaling groups across multiple AZs ensures that if one AZ experiences an outage, instances in the other AZ continue serving traffic. The ALB automatically routes requests to healthy targets across the available zones, maintaining service continuity and ensuring the application remains resilient against localized hardware or power failures.

Why this answer

High availability across two AZs requires resources to be distributed across both AZs so that the failure of one AZ does not take down the application. An Auto Scaling group spanning two AZs behind an ALB ensures that healthy targets exist in both AZs, and the ALB automatically routes traffic only to healthy targets in each enabled AZ.

Exam trap

SAA-C03 often tests the misconception that enabling Cross-Zone Load Balancing alone provides high availability, when in fact the underlying compute resources must be distributed across multiple AZs.

How to eliminate wrong answers

Option A is wrong because deploying all instances in one AZ creates a single point of failure; a Route 53 failover to a secondary region is a disaster recovery strategy, not high availability within a region, and it introduces significant RTO/RPO. Option B is wrong because a single-AZ Auto Scaling group still concentrates all instances in one AZ; Cross-Zone Load Balancing only distributes traffic evenly across healthy targets in enabled AZs and cannot compensate for the loss of the entire AZ. Option D is wrong because an NLB with static IPs in one AZ does not provide multi-AZ resilience, and static IPs per instance are not how NLB targets are addressed.

30
MCQhard

A media streaming company stores millions of video assets in an Amazon S3 bucket. Compliance regulations mandate that objects must be protected against accidental deletion, malicious tampering, and region-wide disasters. The solution must prevent permanent deletion even by users with root credentials during a retention period. Which combination of features meets these requirements?

A.Enable S3 Versioning, apply an S3 Lifecycle rule to transition objects to S3 Glacier, and configure AWS Backup vaults.
B.Enable S3 Versioning, configure S3 Object Lock in compliance mode, and set up S3 Cross-Region Replication to another AWS Region.
C.Apply strict AWS Identity and Access Management (IAM) bucket policies denying s3:DeleteObject, and enable server-side encryption.
D.Configure Amazon S3 Intelligent-Tiering, enable Multi-Factor Authentication (MFA) Delete, and attach bucket policies.
AnswerB

S3 Object Lock in compliance mode enforces a WORM retention period that no user, including the root account, can shorten or bypass, directly satisfying the mandate against permanent deletion. Versioning preserves prior object versions against tampering, while Cross-Region Replication provides the required resilience against region-wide disasters.

Why this answer

Enabling S3 Versioning preserves every version of every object, protecting against accidental overwrites and deletions. Configuring S3 Object Lock in compliance mode ensures that objects cannot be deleted or modified by any user, including the root account, until the retention period expires. Finally, S3 Cross-Region Replication protects against catastrophic regional disasters by maintaining synchronized copies.

Exam trap

Candidates forget to include S3 Object Lock in compliance mode, mistakenly believing that standard versioning or bucket policies alone can protect data from being deleted by root users.

31
Multi-Selecthard

An organization needs to improve the resilience of its EC2-based application. Which THREE actions should the architect perform? (Select THREE.)

Select 3 answers
A.Distribute EC2 instances across multiple Availability Zones.
B.Configure an Auto Scaling group to replace unhealthy instances.
C.Use a single large EC2 instance instead of multiple smaller ones.
D.Implement health checks for the load balancer to monitor instances.
E.Place all instances in a single Auto Scaling group with a fixed capacity.
AnswersA, B, D

Spreading instances across multiple Availability Zones ensures that the application remains operational even if an entire data center experiences an outage. This is a primary requirement for high availability, as it mitigates the risk of localized infrastructure failures affecting the entire application availability and preventing end-user service disruption during regional events.

Why this answer

Option A is correct because distributing EC2 instances across multiple Availability Zones ensures the application survives the failure of an entire AZ, since each AZ has independent power, cooling, and networking. Option B is correct because an Auto Scaling group continuously performs health checks and automatically replaces unhealthy instances, restoring capacity without manual intervention. Option D is correct because load balancer health checks detect unhealthy targets and stop routing traffic to them, so users are only sent to functioning instances.

Option C is wrong because a single large instance is a single point of failure and does not improve resilience. Option E is wrong because a fixed-capacity Auto Scaling group cannot scale out to absorb load or compensate for lost capacity, so it does not meaningfully improve resilience.

Exam trap

SAA-C03 often tests the misconception that simply increasing instance size or count in one AZ provides high availability, when true resilience requires distribution across multiple AZs plus automated health checks and replacement.

32
MCQhard

A global gaming company wants to reduce latency for its players who are distributed worldwide. The application uses UDP-based traffic and requires a static entry point to simplify firewall management. Which service should the architect recommend to optimize the network path?

A.Amazon CloudFront.
B.AWS Global Accelerator.
C.Amazon Route 53 with Geolocation routing.
D.AWS Direct Connect.
AnswerB

AWS Global Accelerator is the best choice for this scenario as it provides two static Anycast IP addresses and supports UDP traffic. It routes player traffic over the high-speed AWS global private network instead of the public internet, which reduces jitter and latency, providing a more consistent and resilient gaming experience for global users.

Why this answer

AWS Global Accelerator uses the AWS global network to route traffic to the optimal regional endpoint based on health and proximity. Unlike CloudFront, which is primarily for HTTP/S content, Global Accelerator supports non-HTTP protocols like UDP. It provides static IP addresses that act as a fixed entry point, improving performance by keeping traffic on the AWS backbone.

Exam trap

Candidates often choose Amazon CloudFront because it is the most well-known global routing service, ignoring that CloudFront primarily handles HTTP/S traffic whereas Global Accelerator natively supports UDP-based protocols.

33
Multi-Selectmedium

A database administrator needs to ensure that an Amazon RDS for MySQL instance can withstand an Availability Zone failure and provide low-latency reads for a reporting application. Which TWO steps should be taken?

Select 2 answers
A.Enable Multi-AZ deployment for the RDS instance.
B.Create a Read Replica in a different Availability Zone.
C.Use EBS Snapshots to back up the database every 5 minutes.
D.Enable Enhanced Monitoring with a 1-second granularity.
E.Implement a Network Load Balancer in front of the RDS instance.
AnswersA, B

Multi-AZ deployment is the standard feature for high availability in Amazon RDS. It creates a standby instance in a different Availability Zone and uses synchronous replication. In the event of a failure, RDS automatically fails over to the standby, ensuring the database remains available without manual intervention or data loss during the transition.

Why this answer

Enabling Multi-AZ provides the high availability needed to survive an AZ failure by maintaining a synchronous standby. Creating a Read Replica in a different AZ offloads read traffic from the primary instance, improving performance for the reporting application while also providing an additional layer of data redundancy across the region.

Exam trap

Candidates often forget that a Multi-AZ standby is not accessible for reads. To offload read traffic, a separate Read Replica is mandatory, as the standby instance is strictly for failover.

34
MCQmedium

A solutions architect is designing a mission-critical web application running on Amazon EC2 instances behind an Application Load Balancer. The application must remain highly available even during an Availability Zone outage. Which deployment strategy meets this requirement with optimal resilience?

A.Deploy all EC2 instances in a single Availability Zone and configure an AWS Auto Scaling group with a maximum capacity of one.
B.Distribute EC2 instances across multiple Availability Zones within a single AWS Region using an Auto Scaling group.
C.Provision EC2 instances in multiple AWS Regions and use Route 53 with a latency routing policy.
D.Use a single EC2 instance of a larger size to handle peak traffic loads without scaling.
AnswerB

Spreading instances across multiple Availability Zones guarantees high availability by ensuring that an outage in one zone does not disrupt the entire application workload. The Auto Scaling group automatically maintains the desired capacity across remaining zones.

Why this answer

Distributing EC2 instances across multiple Availability Zones within a single Region using an Auto Scaling group provides high availability against an AZ outage while maintaining low-latency intra-Region communication. If one AZ fails, the ASG can launch replacement instances in surviving AZs, and the ALB routes traffic only to healthy targets. This is the standard AWS-recommended HA pattern.

Exam trap

SAA-C03 often tests the difference between Multi-AZ (HA within a Region) and Multi-Region (DR/global latency) — candidates over-select multi-Region solutions when the requirement only mentions an Availability Zone outage.

How to eliminate wrong answers

Option A is wrong because deploying all instances in a single AZ creates a single point of failure — an AZ outage takes down the entire application, and a max capacity of one prevents any scaling or redundancy. Option C is wrong because multi-Region deployment with Route 53 latency routing addresses geographic latency and Region-level disasters, but it is over-engineered and more expensive for an AZ-outage requirement; it also introduces cross-Region data consistency challenges. Option D is wrong because a single large EC2 instance has no redundancy at all — if the instance or its AZ fails, the application is completely unavailable, and vertical scaling does not provide high availability.

35
MCQmedium

A financial services company is hosting a critical web application on Amazon EC2 instances behind an Application Load Balancer. The application must remain available even if an entire AWS Region experiences a major outage. The database tier uses Amazon Aurora Global Databases. Which solution provides the most resilient multi-Region architecture with automated failover?

A.Configure an Application Load Balancer in a single Region with EC2 instances distributed across multiple Availability Zones, and back the application with a standard Aurora MySQL database instance.
B.Use Amazon Route 53 with weighted routing policies to distribute traffic between two AWS Regions, and configure standard cross-region database replication using Amazon RDS snapshots taken every hour.
C.Deploy the application stack across two AWS Regions, use Amazon Route 53 with active-passive failover and automated health checks, and configure Amazon Aurora Global Databases with cross-region replication.
D.Set up an AWS Global Accelerator standard accelerator with endpoint groups in two AWS Regions, and use Amazon DynamoDB global tables with local secondary indexes for data storage.
AnswerC

Route 53 active-passive routing with health checks automatically detects regional failures and redirects user traffic to the secondary Region. Aurora Global Databases provide low-latency cross-region replication and fast failover capabilities to ensure uninterrupted application availability.

Why this answer

Deploying Route 53 with active-passive failover and automated health checks combined with a secondary Region ensures automatic traffic rerouting during a regional disaster. Aurora Global Databases minimize replication lag and permit fast promotion of the secondary region database, maintaining data consistency and business continuity for critical financial applications requiring minimal recovery time objectives.

Exam trap

Candidates often choose solutions that provide multi-region high availability without considering automated cross-region database failover or active-passive failover mechanisms necessary for disaster recovery in financial applications.

36
MCQmedium

A company hosts a web application on EC2 instances behind an Application Load Balancer. The application stores session state in local memory, causing users to be logged out whenever the load balancer routes requests to different instances. Which solution provides the most resilient, scalable architecture to resolve this?

A.Configure the Application Load Balancer to use source IP stickiness.
B.Use Amazon EBS multi-attach to share session files across instances.
C.Store session state in Amazon ElastiCache for Redis.
D.Replicate session files using a cron job across all web instances.
AnswerC

ElastiCache for Redis provides a high-performance, distributed, and managed key-value store that persists session data outside the web server memory. This ensures that session state remains available even when instances are replaced, enabling seamless horizontal scaling and maintaining high availability across multiple Availability Zones.

Why this answer

Storing session state in Amazon ElastiCache for Redis externalizes session data so any EC2 instance behind the ALB can serve any request, eliminating the logout problem while providing high availability and scalability. Redis is purpose-built for low-latency session stores and supports replication and failover. This is the AWS-recommended pattern for stateless web tiers.

Exam trap

SAA-C03 often tests whether candidates recognize that stickiness and file replication are anti-patterns for scalable session management; the trap is choosing ALB stickiness as a quick fix instead of externalizing state to a managed in-memory store.

How to eliminate wrong answers

Option A is wrong because source IP stickiness ties a user to one instance based on IP, which breaks for users behind NAT/proxies and does not survive instance failure — it is a workaround, not a resilient architecture. Option B is wrong because EBS Multi-Attach is limited to specific instance types in the same AZ and is not designed for shared session files across a scalable, multi-AZ web tier; it also introduces filesystem consistency challenges. Option D is wrong because replicating session files via cron is brittle, non-real-time, and operationally heavy; it does not provide the low-latency, consistent session store needed and fails during instance failures.

37
MCQmedium

A solutions architect is designing a mission-critical web application that runs on Amazon EC2 instances behind an Application Load Balancer. The application must remain available even if an entire AWS Region experiences a catastrophic failure. The RPO must be near zero, and the RTO must be less than 15 minutes. Which disaster recovery strategy should the architect implement?

A.Implement a backup and restore strategy using automated daily Amazon EBS snapshots copied to a secondary AWS Region.
B.Deploy a pilot light strategy by maintaining a minimal footprint of core database and configuration services in a secondary Region.
C.Configure a multi-site active-active deployment spanning two AWS Regions with continuous data replication and Route 53 health checks.
D.Set up a warm standby architecture running a scaled-down version of the entire application stack in a secondary Region.
AnswerC

Running active workloads across two regions with continuous data replication ensures that traffic can be rerouted instantly upon failure. Route 53 health checks monitor endpoint availability, delivering an RTO of under 15 minutes and minimal RPO.

Why this answer

A multi-site active-active deployment spanning two AWS Regions with continuous data replication and Route 53 health checks provides near-zero RPO and sub-15-minute RTO by allowing both regions to serve traffic simultaneously. If one region fails, Route 53 automatically redirects traffic to the healthy region, ensuring continuous availability.

Exam trap

SAA-C03 often tests the trade-offs between DR strategies, and candidates may confuse warm standby or pilot light with active-active, not realizing that only active-active can meet near-zero RPO and <15 min RTO.

How to eliminate wrong answers

Option A is wrong because backup and restore typically has an RTO of hours and RPO of hours, not near-zero. Option B is wrong because pilot light maintains only core services and requires significant time to scale up, often exceeding 15 minutes RTO. Option D is wrong because warm standby runs a scaled-down stack that must be scaled up, which can take longer than 15 minutes and may not achieve near-zero RPO without continuous replication.

Ready to test yourself?

Try a timed practice session using only Design Resilient Architectures questions.