Courseiva

SAA-C03 Design High-Performing Architectures Practice Question

Exhibit

CloudWatch metrics for the Auto Scaling group (5-minute period):
- CPUUtilization: 28% average
- NetworkIn: 190 MB/min average, no saturation
- GroupDesiredCapacity: 4
- ALBRequestCountPerTarget: 4,800 during peaks
- TargetResponseTime p95: 2.7 seconds during peaks

ALB access log sample:
2026-04-28T09:02:11Z app/prod-alb 203.0.113.10:443 10.0.1.21:8080 0.000 2.698 0.000 200 200 1843 1920 "GET https://app.example.com/search?q=aws HTTP/1.1"

Based on the exhibit, a web application runs on an Amazon EC2 Auto Scaling group behind an Application Load Balancer. During traffic surges, the average CPU utilization stays below 35%, but request latency increases sharply and the ALB access logs show far more requests per target than expected. Which change is the best way to improve scaling behavior?

⚠ Common exam trap

Test-takers frequently assume CPU utilization is the universal scaling metric, but the question explicitly states CPU stays low while latency spikes, indicating the bottleneck is request throughput, not compute, making RequestCountPerTarget the correct metric to scale on.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure target tracking scaling on ALB RequestCountPerTarget for the Auto Scaling group.

The issue is that request latency increases sharply and the ALB logs show far more requests per target than expected, indicating that the Auto Scaling group is not scaling based on the actual load per instance. By configuring target tracking scaling on ALB RequestCountPerTarget, the Auto Scaling group will launch new instances when the average number of requests per target exceeds a defined threshold, directly addressing the root cause of high request volume per instance. This approach ensures scaling is driven by the actual workload distribution rather than CPU utilization, which remains low due to the application being I/O-bound or network-bound.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Lower the CPU target tracking threshold so the Auto Scaling group launches more instances sooner.

    Why it's wrong here

    Lowering the CPU target tracking threshold is ineffective because the architecture is not CPU-bound; the exhibit shows healthy but low CPU usage while request latency climbs, indicating the bottleneck is per-instance request concurrency or another saturation point. Auto Scaling with a CPU metric would not detect this demand increase, and a lower threshold could trigger arbitrary scale-outs on minimal CPU blips, leading to instability and cost without resolving the queueing that drives latency.

  • ✗

    Replace the Application Load Balancer with a Network Load Balancer to reduce request latency.

    Why it's wrong here

    Replacing the Application Load Balancer with a Network Load Balancer does not address the root cause because an NLB operates at Layer 4 and simply forwards TCP/UDP traffic without inspecting HTTP request paths or headers. The exhibit's issue is application-layer capacity pressure on the EC2 targets themselves—the ALB is not the bottleneck; the instances are failing to keep up with request volume, so changing the balancer type neither increases target capacity nor improves how Auto Scaling measures real user demand.

  • ✓

    Configure target tracking scaling on ALB RequestCountPerTarget for the Auto Scaling group.

    Why this is correct

    RequestCountPerTarget directly reflects how many requests each instance is serving, which matches the symptom in the exhibit. It scales the fleet based on actual per-target demand instead of CPU, so the group can add capacity before queueing and latency grow.

  • ✗

    Increase the ALB idle timeout so requests can wait longer before timing out.

    Why it's wrong here

    Increasing the ALB idle timeout merely extends the period a connection can remain open without data transfer, which masks slow application responses and can actually worsen the problem by encouraging clients to hold connections longer, occupying server threads and memory. The exhibit shows latency growth from insufficient instance capacity, not from connections timing out prematurely; idle timeout adjustments do nothing to scale the fleet or reduce the application's processing time, so they only delay the inevitable failure under load.

About these practice questions

One of 935 original SAA-C03 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.