Diagnosing 503 Errors from ALB
A company runs a multi-tier web application on EC2 instances behind an Application Load Balancer. The application experiences intermittent 503 errors during peak traffic. The Auto Scaling group is configured with a step scaling policy based on CPU utilization. CloudWatch metrics show that CPU utilization never exceeds 70%, but the ALB target group reports that some targets are unhealthy. What is the MOST likely cause?
⚠ Common exam trap
DOP-C02 often tests the assumption that 503 errors always mean capacity problems, leading candidates to blame Auto Scaling policies when the real cause is failing health checks removing targets from rotation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The application health check endpoint is returning HTTP 5xx or timing out.
ALB target groups mark targets unhealthy when the configured health check fails — typically because the application endpoint returns 5xx or times out. Unhealthy targets are removed from rotation, and if too few healthy targets remain to serve peak traffic, the ALB returns 503 Service Unavailable. Since CPU never exceeds 70%, the ASG is not scaling out, confirming the bottleneck is health-check failures rather than capacity.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
The application health check endpoint is returning HTTP 5xx or timing out.
Why this is correct
The ALB health check is performing HTTP requests against the configured health check path on each target. When the application returns any 5xx status code (or the request times out because the app hangs under load), the ALB marks that target as unhealthy. With an intermittent application bug (e.g., a memory leak or a connection pool exhaustion), the health check will fail sporadically, causing the ALB to periodically stop routing traffic to that instance. If all targets become unhealthy at the same time, the ALB returns HTTP 503 Service Unavailable to clients, which exactly matches the intermittent nature of the reported errors.
- ✗
The ALB is misconfigured with an incorrect security group blocking traffic to the targets.
Why it's wrong here
An incorrectly configured ALB security group that blocks traffic to the targets would prevent health check requests from ever reaching the applications. This condition is persistent, not intermittent: the security group rules do not change on their own, so every health check would fail consistently, and the ALB would return 503 continuously. The question explicitly says the errors are intermittent, which rules out a static security group misconfiguration. Additionally, a security group it too broad or too narrow affects all traffic uniformly; it does not cause occasional 5xx responses only during certain time windows.
- ✗
The ALB connection draining settings are too short, causing in-flight requests to fail.
Why it's wrong here
Connection draining (also called deregistration delay) is an ALB feature that controls how long the load balancer keeps existing connections open to instances that are being deregistered or that have become unhealthy. It does not influence how health checks are performed or whether the health check endpoint returns a successful status code. Misconfiguring the draining delay can drop in-flight requests during scaling events, but it will not cause health checks to intermittently return 5xx or time out. The symptom described in the question is specifically about health check failures, not about connection termination during instance replacement.
- ✗
The step scaling policy is too aggressive and is terminating instances prematurely.
Why it's wrong here
A step scaling policy (via Auto Scaling or a custom scaling solution) adjusts the desired capacity based on CloudWatch alarms, but it does not directly interact with the ALB health check mechanism. If the policy were too aggressive, it might reduce the number of instances too quickly, potentially leading to insufficient capacity and 503s from the ALB when no healthy targets remain. However, that would be a capacity shortage symptom, not a health check failing on its own — the health check would still report the remaining instances as healthy. The question refers to instances being marked unhealthy by health checks, not simply being terminated, so an aggressive scaling policy is not the correct explanation.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DOP-C02 question from scratch — 1,298 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.