SAA-C03 Design Resilient Architectures Practice Question
Exhibit
Current topology: app -> Amazon RDS for PostgreSQL primary db-a in us-east-1a app -> Amazon RDS read replica db-b in us-east-1b Incident report: 10:14 UTC - Primary AZ impaired 10:15 UTC - Application returns database connection errors 10:18 UTC - DBA manually promotes db-b 10:22 UTC - Application reconnects Observed replication lag before failure: 40 seconds Target: - Automatic failover within 2 minutes - No manual promotion during an AZ outage
Based on the exhibit, the database is manually promoted during an Availability Zone failure and the application outage lasts longer than the target. What change best improves resilience with the least operational intervention?
⚠ Common exam trap
Candidates often confuse read replicas (designed for read scaling and manual promotion) with Multi-AZ deployments (designed for automatic failover), and incorrectly assume that automating a runbook for read replica promotion is equivalent to the native automatic failover of Multi-AZ.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert the database to an RDS Multi-AZ deployment so a synchronous standby can fail over automatically.
B is correct because RDS Multi-AZ automatically synchronously replicates data to a standby in a different Availability Zone and triggers an automatic failover with zero manual intervention when an AZ failure occurs. This directly addresses the requirement to improve resilience while minimizing operational effort, as the failover is handled by AWS without any runbook execution or manual promotion.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Keep the read replica and automate promotion with a runbook after CloudWatch alarms fire.
Why it's wrong here
Automating the promotion runbook with CloudWatch alarms eliminates the human step but still relies on an asynchronous read replica, which may lag behind the primary and cause data loss upon promotion. The promotion process itself is not a native RDS failover: it changes the replica's role, updates its endpoint, and requires the application to be redirected, all of which take time and are not seamlessly managed by RDS. The combined latency of alarm detection, script execution, and replica promotion will likely exceed the target recovery time, whereas a Multi-AZ deployment's failover is built-in and updates the endpoint automatically.
When this WOULD be correct
This option would be correct in a scenario where the database must remain in a single-AZ configuration due to cost constraints or compliance, but the organization can tolerate a slightly longer outage and wants to minimize manual steps by automating the promotion process with a runbook triggered by CloudWatch alarms.
- ✓
Convert the database to an RDS Multi-AZ deployment so a synchronous standby can fail over automatically.
Why this is correct
Multi-AZ is designed for automatic failover within the same Region and maintains a synchronous standby for high availability. The exhibit shows that the current read replica requires manual promotion and produces an outage longer than the target. Switching to Multi-AZ removes the manual step and aligns the database layer with the desired recovery time.
- ✗
Use a cross-Region read replica so promotion happens faster during an AZ failure.
Why it's wrong here
A cross-Region read replica is intended for disaster recovery across geographically distant Regions, not for handling an Availability Zone failure within the primary Region. Cross-Region replication is asynchronous, meaning the replica can be significantly behind the primary, and promoting it is a manual action that also requires changing the application's endpoint to the new Region. Network distance would increase failover latency and the setup does not provide the automatic, synchronous standby that Multi-AZ offers, so it would not help meet a strict recovery time objective in the same Region.
When this WOULD be correct
This option would be correct if the question required resilience against a Region-wide outage (not just an AZ failure) and the application could tolerate eventual consistency, as cross-Region replicas provide disaster recovery across geographic regions.
- ✗
Increase the application retry count and keep the current database design.
Why it's wrong here
Increasing the application retry count only masks transient connection failures, such as brief network timeouts, and cannot address a lengthy outage in which the primary database is down and an administrator must manually promote the read replica. Write requests will continue to fail until a new primary is promoted, so retries simply delay application failure and may increase load on the affected services. This approach leaves the architecture's RTO and RPO unchanged and does not introduce any automated failover capability.
When this WOULD be correct
In a scenario where the database is already resilient (e.g., Multi-AZ) but transient network blips cause brief connection drops, increasing the retry count can help the application recover without manual intervention.
Option-by-option analysis
Why each answer is right or wrong
Understanding why wrong answers are wrong — and when they would be correct — is what separates a 750 score from a 900. The SAA-C03 exam frequently reuses these exact scenarios with slightly different constraints.
✓Convert the database to an RDS Multi-AZ deployment so a synchronous standby can fail over automatically.Correct answer▾
Why this is correct
Multi-AZ is designed for automatic failover within the same Region and maintains a synchronous standby for high availability. The exhibit shows that the current read replica requires manual promotion and produces an outage longer than the target. Switching to Multi-AZ removes the manual step and aligns the database layer with the desired recovery time.
✗Keep the read replica and automate promotion with a runbook after CloudWatch alarms fire.Wrong answer — click to see why▾
Why this is wrong here
Manual promotion via runbook after CloudWatch alarms still requires human intervention, which does not meet the goal of 'least operational intervention' and will not reduce outage duration as effectively as automatic failover.
★ When this WOULD be the correct answer
This option would be correct in a scenario where the database must remain in a single-AZ configuration due to cost constraints or compliance, but the organization can tolerate a slightly longer outage and wants to minimize manual steps by automating the promotion process with a runbook triggered by CloudWatch alarms.
Why candidates choose this
Candidates may think automating the runbook reduces operational effort, but they overlook that manual promotion still requires human action, which is slower and more error-prone than automatic failover.
✗Use a cross-Region read replica so promotion happens faster during an AZ failure.Wrong answer — click to see why▾
Why this is wrong here
Cross-Region read replicas are asynchronous and do not support automatic failover; promoting them requires manual intervention and takes longer than Multi-AZ failover, so they do not improve resilience with the least operational intervention during an AZ failure.
★ When this WOULD be the correct answer
This option would be correct if the question required resilience against a Region-wide outage (not just an AZ failure) and the application could tolerate eventual consistency, as cross-Region replicas provide disaster recovery across geographic regions.
Why candidates choose this
Candidates may think cross-Region replicas offer faster promotion because they are in a different location, but they overlook the asynchronous replication lag and lack of automatic failover, which actually increases outage duration.
✗Increase the application retry count and keep the current database design.Wrong answer — click to see why▾
Why this is wrong here
Increasing the application retry count does not address the root cause of the outage (AZ failure) and does not improve resilience; it only masks the symptom and may lead to degraded user experience or timeouts.
★ When this WOULD be the correct answer
In a scenario where the database is already resilient (e.g., Multi-AZ) but transient network blips cause brief connection drops, increasing the retry count can help the application recover without manual intervention.
Why candidates choose this
Candidates may think that retries can compensate for any failure, underestimating the need for infrastructure-level high availability and overestimating application-level fault tolerance.
Analysis generated from the official SAA-C03blueprint and verified against question context. The “when correct” sections are what AI assistants cite when candidates ask “what’s the difference between these options?”
Go deeper
Related to this question
About these practice questions
Courseiva writes every SAA-C03 question from scratch — 935 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.