SAA-C03 Design High-Performing Architectures Practice Question
A company runs a stateless web API on Amazon EC2 behind an Application Load Balancer. The team notices that during business hours, the ALB starts queueing requests and the average request latency rises. They want to scale out quickly and reliably based on demand, not CPU alone. Which Auto Scaling approach best matches this requirement?
⚠ Common exam trap
Candidates often assume CPU utilization is the best metric for all scaling scenarios, but for a stateless web API behind an ALB, request count per target is a more direct and reliable indicator of demand and latency issues.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use target tracking scaling based on ALB request count per target.
Target tracking scaling based on ALB request count per target directly aligns with the requirement to scale out based on demand (request queuing and latency) rather than CPU alone. This policy automatically adjusts the Auto Scaling group size to maintain a target value for the average number of requests per instance, which is a more reliable indicator of load for a stateless web API than CPU utilization.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a fixed-size Auto Scaling group and increase capacity manually once per hour.
Why it's wrong here
A fixed-size Auto Scaling group with manual hourly adjustments requires an operator to predict future load and act proactively, but a stateless web API can experience sudden traffic spikes at any time. Between hourly checks, a surge of requests will saturate the existing instances, causing degraded response times and rejected connections until the next manual change. This approach lacks any automated feedback loop and cannot adapt to real-time demand.
When this WOULD be correct
For a stateless application with predictable, gradual traffic changes where cost control is critical and automated scaling is not desired, a fixed-size group with manual adjustments might be acceptable.
- ✓
Use target tracking scaling based on ALB request count per target.
Why this is correct
Target tracking scaling with ALB request count per target directly measures actual demand by dividing incoming requests by the number of healthy EC2 instances. The policy automatically adjusts capacity to keep the average near a target value, and because the ALB metric updates in near real time, it can add instances within minutes when a traffic spike begins. This ties scaling directly to the user-facing load that causes queuing and latency, making it the most responsive and precise option.
- ✗
Scale based only on EC2 instance memory utilization, regardless of load.
Why it's wrong here
For a stateless web API, memory utilization often remains flat because no session state is stored locally, so it does not reveal CPU saturation, request queue depth, or user-facing latency. Scaling solely on memory would ignore the actual bottleneck—compute capacity for processing requests—and could leave the fleet under-sized during CPU-intensive spikes. Additionally, if instances have ample memory but are CPU-bound, the scaling policy will never add capacity when it is most needed.
When this WOULD be correct
This option would be correct for an application that is memory-bound, such as an in-memory cache or a data processing job where high memory usage indicates the need for more instances, and CPU or request metrics are not the primary drivers.
- ✗
Use step scaling with a single threshold on average network-in bytes.
Why it's wrong here
Step scaling with a single threshold on average network-in bytes is a poor proxy for application load because API request bodies are often small, while responses and processing demands are not captured by network-in. A single threshold only triggers one level of adjustment, which can be too coarse to counter a rapidly growing queue, and network-in metrics tend to lag behind CPU or request-rate spikes. This can lead to either over-provisioning for bulk uploads or under-provisioning for many small requests.
When this WOULD be correct
A company runs a data ingestion service that receives large file uploads, and scaling should trigger when network throughput exceeds a critical threshold to avoid packet loss. Step scaling with a single threshold on average network-in bytes would be appropriate to add capacity quickly when network input spikes.
Option-by-option analysis
Why each answer is right or wrong
Understanding why wrong answers are wrong — and when they would be correct — is what separates a 750 score from a 900. The SAA-C03 exam frequently reuses these exact scenarios with slightly different constraints.
✓Use target tracking scaling based on ALB request count per target.Correct answer▾
Why this is correct
Target tracking scaling with ALB request count per target directly measures actual demand by dividing incoming requests by the number of healthy EC2 instances. The policy automatically adjusts capacity to keep the average near a target value, and because the ALB metric updates in near real time, it can add instances within minutes when a traffic spike begins. This ties scaling directly to the user-facing load that causes queuing and latency, making it the most responsive and precise option.
✗Use a fixed-size Auto Scaling group and increase capacity manually once per hour.Wrong answer — click to see why▾
Why this is wrong here
Manual scaling once per hour cannot respond quickly to sudden demand spikes during business hours, leading to request queuing and increased latency.
★ When this WOULD be the correct answer
For a stateless application with predictable, gradual traffic changes where cost control is critical and automated scaling is not desired, a fixed-size group with manual adjustments might be acceptable.
Why candidates choose this
Candidates may think manual scaling is simpler and more controllable, underestimating the need for rapid, automated scaling in response to real-time demand.
✗Scale based only on EC2 instance memory utilization, regardless of load.Wrong answer — click to see why▾
Why this is wrong here
Scaling based on EC2 instance memory utilization is not appropriate for a stateless web API where the bottleneck is request queuing and latency, not memory. Memory utilization may not correlate with demand, leading to under- or over-scaling.
★ When this WOULD be the correct answer
This option would be correct for an application that is memory-bound, such as an in-memory cache or a data processing job where high memory usage indicates the need for more instances, and CPU or request metrics are not the primary drivers.
Why candidates choose this
Candidates might think memory utilization is a good proxy for load, especially if they have experience with memory-intensive applications, but for a stateless web API, request count is a more direct indicator of demand.
✗Use step scaling with a single threshold on average network-in bytes.Wrong answer — click to see why▾
Why this is wrong here
Network-in bytes is not a reliable indicator of request queuing or latency for a stateless web API; it can spike due to large payloads without corresponding load, and step scaling with a single threshold lacks the smooth, demand-responsive behavior needed to prevent queueing.
★ When this WOULD be the correct answer
A company runs a data ingestion service that receives large file uploads, and scaling should trigger when network throughput exceeds a critical threshold to avoid packet loss. Step scaling with a single threshold on average network-in bytes would be appropriate to add capacity quickly when network input spikes.
Why candidates choose this
Candidates may think network-in bytes reflects incoming request load, but it ignores request count and latency, and step scaling seems simpler to configure than target tracking, leading them to overlook the need for a metric that directly measures demand.
Analysis generated from the official SAA-C03blueprint and verified against question context. The “when correct” sections are what AI assistants cite when candidates ask “what’s the difference between these options?”
Go deeper
Related to this question
About these practice questions
This SAA-C03 question is part of Courseiva's 935-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.