Courseiva

How to Reduce Latency During Peak Traffic Without Overprovisioning Using Target Tracking Scaling

Exhibit

ALB and ASG snapshot (15-minute peak):
- RequestCountPerTarget: 1,920
- TargetResponseTime p95: 2.9 seconds
- HTTPCode_Target_5XX_Count: 0
EC2 application metrics from CloudWatch agent:
- CPUUtilization: 33%
- MemoryUtilization: 46%
- NetworkIn/Out: steady
Application logs:
[WARN] worker queue depth reached 5,000
[INFO] rejecting requests after thread pool saturation
Current Auto Scaling policy:
- Target tracking on CPUUtilization = 55%

Based on the exhibit, which change best reduces latency during peak traffic without overprovisioning the fleet?

⚠ Common exam trap

Many candidates confuse 'reducing latency' with 'improving network throughput' (Option D) or 'static capacity increases' (Option A), missing that dynamic scaling based on per-target request count directly addresses the latency caused by overloaded instances during peak traffic.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Change the Auto Scaling policy to target tracking on ALB RequestCountPerTarget.

Using a target tracking scaling policy on ALB RequestCountPerTarget dynamically adjusts the fleet size based on the actual load per instance, ensuring that capacity scales with demand during peak traffic without manual intervention or overprovisioning. This approach directly addresses latency caused by high request rates per instance by maintaining a target request count, which reduces response time without adding unnecessary instances during off-peak periods.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Replace the instances with a larger instance family so each server has more headroom.

    Why it's wrong here

    Enlarging the instance family increases available CPU and memory per server, but the exhibit indicates the bottleneck is queue buildup from request concurrency, not CPU saturation. The Auto Scaling group still relies on CPU-based alarms, so it will not add capacity until CPU rises, leaving the thread pool exhausted and latency high. If requests are I/O-bound or waiting on backend dependencies, larger instances may not reduce per-request response time at all.

  • ✓

    Change the Auto Scaling policy to target tracking on ALB RequestCountPerTarget.

    Why this is correct

    RequestCountPerTarget matches the actual demand reaching each instance and scales capacity before the thread pool saturates. Because CPU is still low, CPU-based scaling would react too late or not at all. Target tracking on request count helps keep queue depth and latency down while avoiding unnecessary overprovisioning during quieter periods.

  • ✗

    Use scheduled scaling to add instances only during the business hours peak window.

    Why it's wrong here

    Scheduled scaling adds or removes instances at fixed times, assuming traffic follows a predictable business-hours pattern. It cannot react to unexpected surges, early peaks, or changes in traffic mix, so a flash crowd before the scheduled window will still overwhelm the existing thread pools. Target tracking on ALB RequestCountPerTarget adjusts capacity continuously in response to actual demand per instance, making latency consistent without predicting the clock.

  • ✗

    Replace the ALB with a Network Load Balancer to reduce request latency.

    Why it's wrong here

    An NLB works at Layer 4 and forwards TCP connections directly to targets, which can reduce load balancer forwarding overhead and latency, but the evidence points to thread exhaustion inside the application instances, not to ALB processing time. The ALB's RequestCountPerTarget metric is not emitted by an NLB, so switching would remove the scaling signal that matches demand to capacity. Without a load balancer-level request count, the Auto Scaling policy would have less precise feedback, and the underlying thread pool saturation would remain unchanged.

About these practice questions

This SAA-C03 question is part of Courseiva's 935-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.