SOA-C02 Cost and Performance Optimization Practice Question
A company uses an Application Load Balancer (ALB) to distribute traffic to EC2 instances. The SysOps team wants to reduce costs by ensuring that idle capacity is minimized. Which configuration should they implement?
⚠ Common exam trap
SOA-C02 often tests the difference between reactive scaling (step/CPU-based) and demand-based scaling (target tracking on request count) — candidates pick CPU-based step scaling out of habit, missing that request count per target is the most direct and cost-efficient metric for ALB workloads.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure an Auto Scaling group with a target tracking scaling policy based on ALB request count per target.
A target tracking scaling policy based on ALB request count per target automatically adjusts the number of EC2 instances to keep the request count per instance at a specified target value. This directly ties scaling to actual traffic demand, minimizing idle capacity while maintaining performance. It is the most cost-efficient and responsive configuration for an ALB-fronted workload.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Configure an Auto Scaling group with a step scaling policy based on CPU utilization.
Why it's wrong here
Step scaling policies require you to manually define multiple CloudWatch alarms and adjustment steps, and they only react after a threshold is breached, which introduces lag and often leads to over-provisioning because the policy holds the current capacity until the metric drops below a separate scale-in threshold. For ALB-based workloads, CPU utilization is also a poor proxy for actual user request load because many web applications are I/O-bound or network-bound, so request spikes may not immediately show up in CPU usage. This makes step scaling on CPU slower and less cost-effective than directly tracking the request count per target.
- ✓
Configure an Auto Scaling group with a target tracking scaling policy based on ALB request count per target.
Why this is correct
A target tracking scaling policy with the ALBRequestCountPerTarget metric is purpose-built for web workloads behind an ALB because it continuously adjusts capacity to maintain a specified target value (e.g., 1000 requests per instance), automatically creating and managing the required CloudWatch alarms and scaling activities. This policy responds directly to the metric that reflects actual user traffic per instance, so it scales out quickly when request volume increases and scales in when traffic decreases, significantly reducing idle instance capacity. AWS recommends this approach over CPU or memory based scaling for HTTP(S) applications because it aligns scaling with the real bottleneck—request throughput per instance.
- ✗
Increase the number of EC2 instances in the Auto Scaling group to handle peak load.
Why it's wrong here
Manually increasing the number of EC2 instances in the Auto Scaling group to handle peak load is a form of static scaling that permanently raises the minimum, maximum, or desired capacity to a high value. While this ensures the application can survive a known peak, it leaves that extra capacity running 24/7, wasting compute cost when traffic is low, and it provides no automated response if the actual peak exceeds your estimate. This approach also requires constant manual monitoring and reconfiguration as workload patterns change, unlike dynamic scaling policies that adjust capacity automatically.
- ✗
Set the Auto Scaling group desired capacity to the maximum expected load.
Why it's wrong here
Setting the Auto Scaling group's desired capacity to the maximum expected load forces the ASG to launch and maintain that many instances at all times, completely ignoring load variations and disabling the benefit of any attached scaling policies because the desired capacity always overrides them. This guarantees maximum availability but at the highest possible cost, because every instance runs even during zero-traffic periods, and it offers no protection if traffic exceeds your assumed maximum—the ASG cannot scale beyond the configured max. It is the opposite of elasticity and creates a single fixed compute footprint that does not adapt to either seasonality or unpredictable spikes.
Go deeper
Related to this question
About these practice questions
Courseiva writes every SOA-C02 question from scratch — 1,169 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This SOA-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SOA-C02 exam.