Courseiva

SAA-C03 Design Resilient Architectures Practice Question

An ECS service runs on EC2 instances and is fronted by an ALB. The ALB spans two Availability Zones, and the ECS service desired count is 2 tasks. The underlying EC2 capacity uses an Auto Scaling group (ASG) with min size set to 1, and the ASG also spans only one subnet in practice. What is the most effective change to meet the requirement that the service continues during a single-AZ instance loss?

⚠ Common exam trap

The trap here is that candidates often focus on ECS-specific settings (like deployment configuration or task placement) rather than recognizing that the root cause is the ASG's single-AZ limitation, which is a fundamental infrastructure resilience issue.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Increase ASG min size to at least 2 and ensure the ASG uses subnets in at least two Availability Zones.

The current architecture has a single point of failure because the Auto Scaling group (ASG) spans only one subnet (one Availability Zone). If that AZ fails, all EC2 instances are lost, and the ECS service cannot run any tasks. Increasing the ASG min size to at least 2 and configuring it to use subnets in at least two AZs ensures that EC2 instances are distributed across AZs, allowing the ECS service to maintain at least one task in the surviving AZ during a single-AZ failure.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set the ECS deployment configuration to maximum percent 100 so tasks replace instances faster during rollouts.

    Why it's wrong here

    Setting maximumPercent to 100 only controls how ECS performs rolling deployments, specifying that no additional tasks beyond the desired count can be launched. This does not affect the number of EC2 instances or their distribution across Availability Zones, so it cannot provide spare capacity during an AZ failure. In fact, it can make deployments riskier because old tasks must be stopped before new ones start, but it is irrelevant to infrastructure-level redundancy.

    When this WOULD be correct

    If the question required faster task replacement during a rolling update to minimize downtime, setting maximum percent to 100 would allow new tasks to start before old ones are stopped, speeding up the deployment.

  • ✓

    Increase ASG min size to at least 2 and ensure the ASG uses subnets in at least two Availability Zones.

    Why this is correct

    Increasing the Auto Scaling group minimum to at least 2 and using subnets in at least two Availability Zones guarantees baseline EC2 capacity in multiple AZs. If one AZ becomes unavailable, the ALB can route traffic to healthy targets in the remaining AZ, and ECS has compute available to reschedule tasks. This directly provides the cross-AZ redundancy needed for the service to continue operating.

  • ✗

    Enable ALB connection draining longer than expected so existing connections survive longer during an AZ event.

    Why it's wrong here

    ALB connection draining is a deregistration delay that allows in-flight requests to complete when a target is intentionally removed from the target group. It cannot preserve connections for an entire Availability Zone loss, because the underlying instances are unreachable and the ALB will mark the targets unhealthy and stop sending traffic to them. The feature also does not provision any new capacity, so it leaves the service unable to handle new requests without healthy instances in another AZ.

    When this WOULD be correct

    This option would be correct in a scenario where the question asks how to minimize disruption to in-flight requests during a planned deployment or when an instance is being replaced, and the ALB is configured to gradually drain connections before terminating instances.

  • ✗

    Reduce task memory reservations to pack both tasks onto a single EC2 instance.

    Why it's wrong here

    Reducing task memory reservations simply allows the ECS scheduler to place more tasks on each EC2 instance, encouraging consolidation rather than distribution. If all tasks are packed onto a single instance, the blast radius of a failure is maximized—losing that one instance or its AZ takes down every task. This directly undermines high availability, which requires spreading tasks across instances and AZs, not fitting them onto fewer machines.

    When this WOULD be correct

    This option would be correct in a scenario where the ECS service is running on Fargate with a single-AZ deployment and the goal is to reduce cost by packing tasks onto fewer instances, while high availability is not a requirement.

Option-by-option analysis

Why each answer is right or wrong

Understanding why wrong answers are wrong — and when they would be correct — is what separates a 750 score from a 900. The SAA-C03 exam frequently reuses these exact scenarios with slightly different constraints.

✓Increase ASG min size to at least 2 and ensure the ASG uses subnets in at least two Availability Zones.Correct answer▾

Why this is correct

Increasing the Auto Scaling group minimum to at least 2 and using subnets in at least two Availability Zones guarantees baseline EC2 capacity in multiple AZs. If one AZ becomes unavailable, the ALB can route traffic to healthy targets in the remaining AZ, and ECS has compute available to reschedule tasks. This directly provides the cross-AZ redundancy needed for the service to continue operating.

✗Set the ECS deployment configuration to maximum percent 100 so tasks replace instances faster during rollouts.Wrong answer — click to see why▾

Why this is wrong here

Setting maximum percent to 100 does not address the lack of multi-AZ redundancy; it only affects deployment speed, not availability during an AZ failure.

★ When this WOULD be the correct answer

If the question required faster task replacement during a rolling update to minimize downtime, setting maximum percent to 100 would allow new tasks to start before old ones are stopped, speeding up the deployment.

Why candidates choose this

Candidates may confuse deployment configuration with high availability, thinking that faster task replacement compensates for missing AZ redundancy.

✗Enable ALB connection draining longer than expected so existing connections survive longer during an AZ event.Wrong answer — click to see why▾

Why this is wrong here

Connection draining helps preserve existing connections during a rolling update or instance deregistration, but it does not prevent service disruption when an entire Availability Zone fails. The ALB would still lose all healthy targets in that AZ, and new connections cannot be established to instances in the failed AZ.

★ When this WOULD be the correct answer

This option would be correct in a scenario where the question asks how to minimize disruption to in-flight requests during a planned deployment or when an instance is being replaced, and the ALB is configured to gradually drain connections before terminating instances.

Why candidates choose this

Candidates may think that extending connection draining provides enough time for the system to recover from an AZ failure, confusing it with a mechanism to handle abrupt instance loss rather than graceful termination.

✗Reduce task memory reservations to pack both tasks onto a single EC2 instance.Wrong answer — click to see why▾

Why this is wrong here

Reducing task memory reservations does not address the single-AZ failure risk because both tasks could still be placed in the same Availability Zone, and the ASG only spans one subnet, so losing that AZ would still cause total service outage.

★ When this WOULD be the correct answer

This option would be correct in a scenario where the ECS service is running on Fargate with a single-AZ deployment and the goal is to reduce cost by packing tasks onto fewer instances, while high availability is not a requirement.

Why candidates choose this

Candidates may think that reducing resource reservations allows tasks to fit on fewer instances, which they mistakenly believe improves resilience by concentrating tasks, but they overlook the need for multi-AZ distribution to survive AZ failures.

Analysis generated from the official SAA-C03blueprint and verified against question context. The “when correct” sections are what AI assistants cite when candidates ask “what’s the difference between these options?”

Visual reference

192.168.1.0 /24 256 addresses (254 usable) 192.168.1.0 /25 Subnet A 128 addr (126 usable) 192.168.1.128 /25 Subnet B 128 addr (126 usable) Borrowing 1 bit from host portion creates 2 subnets (/25)

About these practice questions

Courseiva writes every SAA-C03 question from scratch — 935 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.