DOP-C02 Incident and Event Response Practice Question
An application running on Amazon ECS (Fargate) experiences intermittent HTTP 503 errors. The application uses an Application Load Balancer. The ECS service has a desired count of 2. CPU and memory utilization are below 50%. What is the most likely cause?
⚠ Common exam trap
The trap here is that candidates often attribute 503 errors to resource exhaustion (CPU/memory) or scaling issues, but the question explicitly states low utilization, forcing you to focus on health check configuration as the root cause of intermittent availability.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The target group health check threshold is set too low.
Intermittent HTTP 503 errors from an Application Load Balancer (ALB) typically indicate that the target group health checks are failing, causing the ALB to stop routing traffic to the affected tasks. With a desired count of 2 and low CPU/memory utilization, the most likely cause is a health check threshold set too low (e.g., a low unhealthy threshold count), which makes the ALB prematurely mark tasks as unhealthy during transient issues, leading to no healthy targets and 503 responses. This aligns with the symptom of intermittent errors, as tasks may briefly fail a health check but recover quickly, yet the low threshold causes them to be deregistered.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The ECS service Auto Scaling is too aggressive.
Why it's wrong here
ECS Service Auto Scaling uses CloudWatch alarms on CPU/memory utilization; the question states metrics are below thresholds, so no scaling activity would occur, neither scaling out nor scaling in. Even aggressive step-scaling policies cannot invoke a scaling action without breaching the alarm threshold, so the 503s cannot be attributed to Auto Scaling. Furthermore, Auto Scaling actions are not immediate and operate on a cooldown, making them an unlikely source of intermittent, short-lived service disruptions.
- ✓
The target group health check threshold is set too low.
Why this is correct
The target group's unhealthy threshold is the number of consecutive failed health checks required before ALB deregisters a task and stops sending it traffic. If the threshold is set too low (e.g., 2), a single transient failure—such as a slow response during Fargate instance startup or a momentary network blip—will cause the task to be marked unhealthy and removed from rotation. While it is out of service, if other tasks are also briefly unhealthy or in a draining state, the ALB has no healthy targets and returns HTTP 503 errors until the next successful health check (possibly just seconds later), producing intermittent 503s.
- ✗
The ALB listener rule is misconfigured.
Why it's wrong here
An ALB listener rule controls routing based on host/path conditions, but its configuration is static and evaluated consistently for every request. If a listener rule were misconfigured, requests would either always match the wrong target group (yielding persistent routing errors, 404s, or 503s) or never match (falling to the default action), not intermittently. The symptom described—intermittent 503 errors while CPU/memory are low—points to health check flapping rather than a static routing misconfiguration, since a misconfigured rule wouldn't alternate between healthy and unhealthy target states.
- ✗
The task definition has an incorrect memory hard limit.
Why it's wrong here
The task definition's memory hard limit (memory) is a cgroup limit; if the container exceeds it, the kernel kills the container, causing the task to stop and be replaced by ECS, which would produce outages. However, the metrics show memory utilization is below 50%, meaning the container is nowhere near the configured hard limit, so the task is not being OOM-killed. An incorrect hard limit that is simply low relative to some assumed value has no effect unless actual memory use approaches it; therefore, it cannot explain the intermittent 503s when utilization is well under the limit.
Go deeper
Related to this question
About these practice questions
One of 1,298 original DOP-C02 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.