Troubleshooting ALB 5xx Errors
A company runs a critical web application on EC2 instances behind an Application Load Balancer (ALB) with Auto Scaling. Users report intermittent 503 errors. CloudWatch metrics show that the ALB's 'RequestCount' is normal, but 'HTTPCode_ELB_5XX_Count' spikes. The 'TargetResponseTime' metric shows occasional high latency. Which troubleshooting step should the DevOps engineer take FIRST?
Quick Answer
The correct first step is to enable and analyze the ALB access logs stored in S3, filtering for 503 errors and correlating them with target response times. This approach directly addresses the intermittent 503 errors by allowing you to pinpoint whether the errors align with spikes in latency from specific targets, revealing issues like resource exhaustion or application bottlenecks that CloudWatch metrics alone cannot isolate. On the AWS Certified DevOps Engineer Professional DOP-C02 exam, this scenario tests your ability to choose the most granular diagnostic tool before making configuration changes—a common trap is jumping to scaling actions or misreading CloudTrail logs, which record API calls, not HTTP traffic. Remember, access logs are your forensic microscope for ALB errors; CloudWatch shows the fever, but access logs find the infection. A useful memory tip: "Logs before knobs"—always inspect the data before turning any dials.
⚠ Common exam trap
DOP-C02 often tests whether candidates jump to scaling or configuration changes instead of first using observability data (ALB access logs, CloudWatch metrics) to isolate whether errors originate from the ALB or the targets.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable and analyze the ALB access logs stored in S3, filtering for 503 errors and correlating with target response times.
ALB access logs capture detailed per-request information including target IP, response time, and the specific error (e.g., 503 due to target connection errors or target timeouts). Analyzing these logs filtered for 503s and correlated with TargetResponseTime reveals whether the errors originate from unhealthy targets, connection limits, or slow responses, making it the correct first diagnostic step.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable and analyze the ALB access logs stored in S3, filtering for 503 errors and correlating with target response times.
Why this is correct
ALB access logs record per-request target status, target IP and timing, so filtering 503s reveals whether targets returned errors or the load balancer itself failed. This directly correlates the ELB 5XX spike with backend latency, satisfying the stem's need to identify the failing tier first.
- ✗
Increase the desired capacity of the Auto Scaling group to handle more requests.
Why it's wrong here
Scaling out adds targets, yet the 5XX count originates at the load balancer itself, indicating targets failing health checks or timing out rather than insufficient capacity. It tempts because rising latency and errors often signal overload, making capacity increases the reflexive first response in autoscaling environments.
- ✗
Disable connection draining on the target group to prevent slow-draining instances from causing errors.
Why it's wrong here
Disabling connection draining would terminate in-flight requests abruptly, worsening 503s rather than resolving them; draining exists to let targets finish requests during deregistration. It tempts because deregistering instances during scale-in is a known source of intermittent errors, so the setting appears relevant to the symptom.
- ✗
Review AWS CloudTrail logs for any recent configuration changes to the ALB.
Why it's wrong here
CloudTrail records control-plane API calls, not the target health or timeout behaviour producing the ELB-generated 5XX responses. It tempts because recent configuration changes are a common cause of sudden load balancer failures, and auditing them is a sound practice once health-check and target metrics have been examined.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DOP-C02 question from scratch — 1,298 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on DOP-C02
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A company's application running on EC2 instances behind an Application Load Balancer (ALB) is returning intermittent 504 errors. The instances are in an Auto Scaling group with a health check grace period of 300 seconds. What should the DevOps engineer check first to troubleshoot the issue?
medium- A.Review the Auto Scaling group scaling policies.
- B.Verify the target group health checks are passing.
- C.Check security group rules for the ALB.
- ✓ D.Check ALB access logs for target response times.
Why D: A 504 error indicates the load balancer did not receive a response from the target within the idle timeout period. Checking ALB access logs for target response times is the first step to determine if the backend is slow or unresponsive. Option A is wrong because scaling policies affect the number of instances, not response times. Option B is wrong because health checks verify instance availability, but intermittent slow responses may not cause health check failures. Option C is wrong because security group rules would cause different errors (e.g., connection timeouts) rather than 504s.
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.