A company is running a production web application on Auto Scaling EC2 instances behind an ALB. They have enabled detailed CloudWatch metrics on the EC2 instances and enabled CloudTrail. Recently, users reported intermittent 503 errors. The operations team reviews CloudWatch dashboards but sees no spike in CPU or memory. What is the MOST likely cause of the 503 errors?
The ALB routes requests only to targets that have successfully passed their configured health checks. If health check failures occur—due to an incorrect health check path, a timeout threshold being too low, or an application dependency failing—the corresponding instances are marked unhealthy and removed from the rotation. When the number of healthy targets drops below the minimum needed (typically zero healthy targets in a target group), the ALB returns HTTP 503 Service Unavailable. This can happen without CPU or memory utilization rising, because health checks validate application-level readiness, not just resource utilization.
Why this answer
ALB returns HTTP 503 when no healthy targets are available in the target group to serve the request. If health checks are failing intermittently — due to application-level issues, misconfigured health check paths, or slow responses — targets are marked unhealthy and removed from rotation, leaving insufficient capacity and producing 503s. Because CPU and memory show no spike, the cause is not resource saturation but target availability.
Exam trap
DOP-C02 often tests the distinction between ALB 503 (no healthy targets) and 502 (bad gateway from a target) — candidates chase resource metrics or security group misconfigurations when the real signal is target health check failures.
How to eliminate wrong answers
Option A is wrong because CloudTrail logs API activity for auditing, not application request handling; insufficient trail configuration cannot cause 503 errors. Option C is wrong because detailed monitoring (1-minute metrics) affects metric granularity, not target health — disabling it would not cause 503s, and the scenario says detailed metrics are enabled. Option D is wrong because a misconfigured ALB security group would typically cause connection timeouts or refused connections (and would affect all traffic consistently), not intermittent 503s from the ALB itself.