SOA-C02 Monitoring, Logging, and Remediation Practice Question
A company runs a critical web application on Amazon EC2 instances in an Auto Scaling group across three Availability Zones. The application uses an Application Load Balancer (ALB) for traffic distribution. The SysOps administrator has configured a CloudWatch alarm to monitor the ALB's `TargetResponseTime` metric, with a threshold of 5 seconds. The alarm triggers when the average response time exceeds 5 seconds for 2 consecutive periods. Recently, the alarm has been triggering frequently during peak hours, but the application team reports that the response time is acceptable and the application is performing normally. The administrator investigates and finds that a small number of requests are taking a very long time (over 30 seconds), skewing the average. The administrator needs to reduce the number of false alarms while still being alerted if the overall application performance degrades. Which course of action should the administrator take?
⚠ Common exam trap
Watch out — candidates often think increasing the threshold or evaluation periods is the solution, but they fail to recognize that the average metric is inherently sensitive to outliers, and the correct fix is to change the statistic to a percentile like p95 or p99.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Change the statistic to p95 and keep the threshold at 5 seconds
Using the p95 (95th percentile) statistic instead of the average filters out the impact of the small number of outlier requests that take over 30 seconds. The p95 metric shows the response time below which 95% of requests fall, providing a more accurate representation of typical application performance. This reduces false alarms from skewed averages while still alerting if the majority of users experience degraded response times exceeding 5 seconds.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Change the statistic to p95 and keep the threshold at 5 seconds
Why this is correct
Changing the statistic to p95 means the alarm triggers when the 95th percentile latency exceeds 5 seconds, effectively ignoring the slowest 5% of requests. This reflects the experience of the majority of users, because rare outliers like a single bad request won't cause false alarms. Keeping the threshold at 5 seconds ensures the alarm still detects genuine degradation affecting the typical request. p95 is a robust metric for latency monitoring because it filters out transient spikes while staying responsive to systemic issues.
- ✗
Increase the threshold to 30 seconds
Why it's wrong here
Increasing the threshold to 30 seconds would only fire the alarm for extremely severe latency, such as a stuck request or a failing backend. Moderate degradation, where page loads take 8-10 seconds instead of the expected 5, would go completely unnoticed. This defeats the purpose of early warning and allows user-perceived performance to deteriorate badly before an alert is raised. The threshold must align with the application's performance goal, not with the worst-case outlier, and 30 seconds is far too lenient to protect the user experience.
- ✗
Decrease the period to 60 seconds and lower the threshold to 3 seconds
Why it's wrong here
Decreasing the period to 60 seconds and lowering the threshold to 3 seconds makes the alarm extremely sensitive to any short burst of slowness. With a one-minute period, a single slow request or a brief network hiccup can push the average above 3 seconds, causing a breach even though normal traffic is fine. This configuration produces constant false alarms and alert fatigue, and the 3-second target may not match the application’s actual baseline. A tighter period and lower threshold do not fix outlier sensitivity; they amplify it, making the alarm less trustworthy.
- ✗
Increase the evaluation periods to 5 consecutive periods
Why it's wrong here
Requiring 5 consecutive evaluation periods means the alarm must observe a breach for 5 straight minutes before triggering, which introduces a significant delay in notification. Moreover, if sporadic outliers occur once every minute, each period could still breach the average threshold, so the alarm would fire despite the occasional slow requests not affecting the majority. This approach neither filters out outliers nor speeds up detection; it only delays the alarm. The intent of evaluation periods is to prevent transient flapping, not to address the statistic selection problem, so the underlying outlier-driven behavior remains unchanged.
Visual reference
Go deeper
Related to this question
About these practice questions
This SOA-C02 question is part of Courseiva's 1,169-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This SOA-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SOA-C02 exam.