Courseiva

SAA-C03 Design High-Performing Architectures Practice Question

Your web application runs on EC2 instances behind an Application Load Balancer (ALB). During traffic spikes, p95 response time increases, but average CPU utilization remains below 40%. The current Auto Scaling policy scales based on average CPU%. What should you change to improve performance during spikes?

⚠ Common exam trap

Watch out — candidates often assume high latency always means high CPU, but AWS tests the understanding that p95 latency can spike due to request queueing even when CPU is idle, making request-based scaling the correct choice over CPU-based scaling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Scale on a request-driven metric such as ALB RequestCount per target (or target-group request rate)

The p95 response time is increasing during traffic spikes while CPU utilization remains low, indicating that the bottleneck is not compute capacity but rather request handling or connection overhead. By scaling on ALB RequestCountPerTarget, you directly target the metric causing latency—each target's request load—rather than an indirect metric like CPU. This ensures that new instances are launched precisely when individual targets are overwhelmed by requests, reducing queueing delays and improving response times.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Keep scaling on CPU% to avoid over-scaling

    Why it's wrong here

    CPU-based scaling can lag or fail when the bottleneck is not CPU saturation (for example, thread/connection limits, queueing, downstream dependency slowness, or ALB target response time). Your symptom already shows CPU is not the limiting factor.

  • ✓

    Scale on a request-driven metric such as ALB RequestCount per target (or target-group request rate)

    Why this is correct

    A request-driven metric correlates directly with incoming workload pressure. Scaling on request rate helps ensure enough capacity is added before request queues build up, which can reduce p95 response time even when CPU remains low.

  • ✗

    Disable scaling and manually increase capacity during business hours

    Why it's wrong here

    Disabling auto scaling and manually adjusting capacity during business hours trades elasticity for static provisioning, requiring accurate forecasts of peak traffic. In practice, web traffic can spike unpredictably from marketing campaigns, partner referrals, or viral content, so manual changes often leave the fleet either overprovisioned (cost waste) or underprovisioned (high p95 latency) outside the predicted window. It also introduces operational overhead and human delay, whereas Auto Scaling can react in minutes to changing demand.

  • ✗

    Scale only when network packet drops fall below a threshold

    Why it's wrong here

    Scaling only when network packet drops fall below a threshold focuses on a low-level infrastructure signal that does not reflect application-layer saturation. The p95 latency is driven by queued requests, blocked worker threads, or dependent service slowness, all of which can occur while packet drops remain near zero. Moreover, packet drops are often transient and caused by network bursts or security-group limits, so they can produce noisy or inverted scaling decisions that either over-provision or leave the app short-handed exactly when latency is rising.

About these practice questions

One of 935 original SAA-C03 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.