SAA-C03 Design Resilient Architectures Practice Question
A retail platform needs disaster recovery across AWS Regions. The business requirement is: RTO up to 6 hours, RPO up to 1 hour, and they want the ability to start serving quickly during a Region outage but do not want to run full production capacity continuously. Which DR strategy best fits these requirements?
⚠ Common exam trap
Many candidates confuse pilot light with warm standby, assuming that any minimal running infrastructure qualifies as pilot light, but warm standby specifically requires a scaled-down but fully functional environment that can serve traffic immediately after scaling, whereas pilot light requires significant provisioning before it can serve traffic.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Warm standby, keeping a reduced but ready-to-scale environment in the secondary Region.
Warm standby is the correct strategy because it maintains a scaled-down but fully functional copy of the production environment in the secondary Region, which can be scaled up within the 6-hour RTO. The RPO of 1 hour is met by continuous replication (e.g., Amazon RDS cross-Region read replicas or DynamoDB global tables), and the reduced footprint avoids the cost of full production capacity while still enabling rapid failover.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Backup and restore only, with no continuously running infrastructure in the secondary Region.
Why it's wrong here
Backup and restore in disaster recovery stores data snapshots or databases in the secondary Region (e.g., S3 cross-Region replication, EBS snapshots, RDS snapshots) but has absolutely no running infrastructure there. To recover, you must rebuild the entire environment from scratch—creating EC2 instances, installing software, restoring data, and validating the system—which normally takes many hours or even days, far exceeding the 6-hour RTO. Additionally, data loss is potentially higher because RPO is tied to the last backup or replication cycle, while the lack of pre-provisioned resources makes RTO unpredictable.
When this WOULD be correct
For a non-critical application with RTO > 24 hours and RPO > 24 hours, where cost is the primary concern and data loss of up to a day is acceptable.
- ✗
Pilot light, keeping only the minimum resources needed to bootstrap the environment.
Why it's wrong here
Pilot light keeps only the most critical infrastructure running—such as a replicated database and a minimal compute core—while the rest of the application stack remains stopped. When a disaster occurs, you must deploy and configure all application servers, load balancers, and auxiliary services before scaling to full capacity; this bootstrap process can easily consume most or all of the 6-hour RTO window, especially for a complex retail platform. Because only base resources are warm, it provides faster recovery than backup-and-restore but still risks missing a strict recovery time objective.
When this WOULD be correct
A scenario where the business requires very low cost for the DR site, can tolerate longer RTO (e.g., 12-24 hours), and has minimal RPO requirements (e.g., 1 hour). For example, a non-critical internal tool that can afford significant downtime during a disaster.
- ✓
Warm standby, keeping a reduced but ready-to-scale environment in the secondary Region.
Why this is correct
Warm standby in the secondary Region runs a scaled-down but fully functional copy of your production stack, typically with key databases and services already deployed and synchronized. You can provision extra compute capacity on demand via Auto Scaling or pre-provisioned cluster resizing to reach full production load within the 6-hour RTO. This balances cost and recovery speed by keeping idle but ready infrastructure that can be quickly scaled up, making it the most appropriate choice for the stated requirement.
- ✗
Multi-site active-active, serving production traffic from both Regions at all times.
Why it's wrong here
Multi-site active-active serves production traffic concurrently from both Regions, which demands full capacity in each location and active-active database replication (e.g., DynamoDB global tables or multi-writer Aurora) to handle writes from both sides. Running and operating two full production environments continuously is substantially more expensive and operationally complex than needed for a disaster recovery plan, and it introduces challenges like conflict resolution, data consistency, and cross-Region latency that are overkill for meeting a 6-hour RTO.
When this WOULD be correct
An exam question requiring zero RTO and zero RPO for a mission-critical application with unlimited budget, where continuous active-active traffic distribution is acceptable.
Option-by-option analysis
Why each answer is right or wrong
Understanding why wrong answers are wrong — and when they would be correct — is what separates a 750 score from a 900. The SAA-C03 exam frequently reuses these exact scenarios with slightly different constraints.
✓Warm standby, keeping a reduced but ready-to-scale environment in the secondary Region.Correct answer▾
Why this is correct
Warm standby in the secondary Region runs a scaled-down but fully functional copy of your production stack, typically with key databases and services already deployed and synchronized. You can provision extra compute capacity on demand via Auto Scaling or pre-provisioned cluster resizing to reach full production load within the 6-hour RTO. This balances cost and recovery speed by keeping idle but ready infrastructure that can be quickly scaled up, making it the most appropriate choice for the stated requirement.
✗Backup and restore only, with no continuously running infrastructure in the secondary Region.Wrong answer — click to see why▾
Why this is wrong here
Backup and restore typically has RPOs of hours or days and RTOs of 24+ hours, failing to meet the 1-hour RPO and 6-hour RTO requirements.
★ When this WOULD be the correct answer
For a non-critical application with RTO > 24 hours and RPO > 24 hours, where cost is the primary concern and data loss of up to a day is acceptable.
Why candidates choose this
Candidates may think backup and restore is the simplest and cheapest DR method, overlooking the strict RPO/RTO requirements in the question.
✗Pilot light, keeping only the minimum resources needed to bootstrap the environment.Wrong answer — click to see why▾
Why this is wrong here
The pilot light strategy typically has RTO of 10-15 minutes and RPO of a few minutes, which is faster than the required 6-hour RTO and 1-hour RPO, but it does not meet the requirement to 'start serving quickly' during a Region outage because it requires provisioning and scaling resources after failover, leading to longer recovery time than warm standby.
★ When this WOULD be the correct answer
A scenario where the business requires very low cost for the DR site, can tolerate longer RTO (e.g., 12-24 hours), and has minimal RPO requirements (e.g., 1 hour). For example, a non-critical internal tool that can afford significant downtime during a disaster.
Why candidates choose this
Candidates may confuse 'pilot light' with 'warm standby' because both involve running some resources in the secondary Region, but they underestimate the additional provisioning time needed for pilot light to become fully operational, making it unsuitable for the 'quickly' requirement.
✗Multi-site active-active, serving production traffic from both Regions at all times.Wrong answer — click to see why▾
Why this is wrong here
Multi-site active-active requires running full production capacity in both Regions continuously, which contradicts the requirement to not run full production capacity continuously.
★ When this WOULD be the correct answer
An exam question requiring zero RTO and zero RPO for a mission-critical application with unlimited budget, where continuous active-active traffic distribution is acceptable.
Why candidates choose this
Candidates may think active-active provides the fastest failover and meets RTO/RPO, but overlook the cost and capacity requirement that conflicts with the 'not run full production capacity continuously' constraint.
Analysis generated from the official SAA-C03blueprint and verified against question context. The “when correct” sections are what AI assistants cite when candidates ask “what’s the difference between these options?”
Go deeper
Related to this question
About these practice questions
Courseiva writes every SAA-C03 question from scratch — 935 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.