A cloud engineer is troubleshooting a performance issue where a web server cluster experiences high latency during peak hours. The cluster uses an auto-scaling group behind a load balancer. Which THREE steps should the engineer take to identify the root cause?
High latency during peak hours may stem from CPU or memory saturation on the web servers themselves. Monitoring both metrics reveals whether instances are resource-constrained, which would prevent the auto-scaling group from serving requests promptly behind the load balancer.
Why this answer
Option A is correct because monitoring CPU and memory utilization on the web servers reveals whether the instances are resource-saturated during peak hours, which is a common cause of high latency in an auto-scaling cluster. Option B is correct because analyzing web server access logs for slow requests helps pinpoint which endpoints, queries, or upstream dependencies are contributing to the latency, providing concrete evidence of the bottleneck. Option C is correct because checking the load balancer's backend instance health status verifies whether unhealthy or failing instances are being served or whether health checks are flapping, which directly impacts response times.
Option D is incorrect because reducing the number of instances in the auto-scaling group would decrease capacity and likely worsen latency rather than help diagnose the root cause. Option E is incorrect because reviewing security group rules for the load balancer addresses connectivity and access control, not performance degradation during peak hours.
Exam trap
The trap here is that candidates may think reducing instances (Option D) is a valid troubleshooting step, but it is a remediation action that can mask the root cause and potentially crash the application under load.