SOA-C02 Cost and Performance Optimization Practice Question
A company is running a stateful web application on EC2 instances in an Auto Scaling group. The application requires low latency and high throughput. Currently, the application is experiencing performance degradation during peak hours. Which scaling strategy should the SysOps administrator implement to improve performance and optimize cost?
⚠ Common exam trap
A common mix-up: candidates choose scheduled scaling (Option B) because they see 'peak hours' and assume a fixed schedule, but they miss that predictive scaling uses ML to handle variable peak patterns more efficiently than rigid schedules.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Predictive scaling policy
Predictive scaling is the correct choice because it uses machine learning to analyze historical traffic patterns and proactively adjust capacity before demand spikes, which is ideal for a stateful web application experiencing predictable peak-hour performance degradation. This approach ensures low latency and high throughput by pre-warming instances, while optimizing cost by avoiding over-provisioning during off-peak periods.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Step scaling policy based on memory utilization
Why it's wrong here
A step scaling policy using memory utilization is flawed on two counts. First, the EC2 instance's memory utilization is not published as a standard CloudWatch metric, so you must install and configure the CloudWatch agent to emit custom metrics, adding operational overhead and a delay before data is available. Second, step scaling is inherently reactive: it responds only after the alarm threshold is breached, and its stepped adjustments combined with cooldown timers can cause oscillations. Even with well-tuned steps, it cannot pre-provision capacity for a sudden incoming spike in user traffic, which is the core issue for a stateful web application.
- ✗
Scheduled scaling with fixed times
Why it's wrong here
Scheduled scaling with fixed times works only if the application's traffic pattern is absolutely predictable, such as a known 9-to-5 business workload. It does not take into account real-time utilization or unexpected events like a viral marketing campaign or a flash crowd, so if the traffic pattern shifts, the schedule will either add unneeded instances or fail to add them in time. For a stateful web application, this rigid approach can cause unnecessary cost without guaranteeing performance, and it cannot adapt to the organic growth or changing seasonality that typical web apps experience.
- ✗
Simple scaling policy based on CPU utilization
Why it's wrong here
A simple scaling policy based on CPU utilization is reactive and uses a single adjustment step with a mandatory cooldown period. During a rapid spike, it will take several minutes for the alarm to fire, then the cooldown prevents any further scaling action until the period expires, so capacity lags far behind demand. In a stateful web application, this delay can lead to dropped sessions or degraded user experience. Furthermore, CPU utilization alone often does not capture the application's actual bottleneck—memory pressure, database connection pool exhaustion, or request queue depth—so the scaling trigger may be misaligned with the real workload.
- ✓
Predictive scaling policy
Why this is correct
A predictive scaling policy uses machine learning to analyze historical traffic patterns and forecast future demand, allowing Auto Scaling to launch instances ahead of the actual spike. This proactive approach is ideal for a stateful web application because it eliminates the cold start lag and reduces the risk of insufficient capacity during sudden bursts. It also improves cost efficiency by avoiding over-provisioning, and when combined with a dynamic scaling policy, it can handle both anticipated trends and unexpected deviations. The key is that predictive scaling learns from recurring patterns (daily, weekly, or monthly) and smooths out the capacity curve before the load actually arrives.
Go deeper
Related to this question
About these practice questions
Courseiva writes every SOA-C02 question from scratch — 1,169 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This SOA-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SOA-C02 exam.