DOP-C02 Monitoring and Logging Practice Question
A DevOps engineer is troubleshooting an application that runs on Amazon EC2 instances behind an Application Load Balancer. Users report intermittent 503 errors. CloudWatch metrics for the ALB show an increase in 'HTTPCode_ELB_5XX_Count' but the backend 'HealthyHostCount' remains stable. Which action should the engineer take to identify the root cause?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable and review the ALB access logs stored in Amazon S3 to analyze the HTTP response codes and request patterns.
ALB access logs capture detailed information about each request, including HTTP status codes, request processing times, and target response times. By analyzing these logs in Amazon S3, the engineer can identify which specific requests are resulting in 503 errors, examine if there are patterns such as high latency or target unavailability, and determine the root cause. Option A is incorrect because increasing instance size does not address the intermittent 503 errors; if the instances are healthy (HealthyHostCount stable), the issue is likely not capacity but something else like configuration or request handling. Option B is incorrect because detailed CloudWatch metrics on instances, while useful for performance, would not directly reveal why the ALB is returning 503 errors; the backend might be healthy but returning errors or timing out. Option D is incorrect because increasing the idle timeout would only help if requests are being dropped due to idle connections, but 503 errors typically indicate that the targets are not responding or are returning errors, not idle timeouts.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the size of the EC2 instances to handle more requests.
Why it's wrong here
Increasing the size of the EC2 instances addresses compute capacity, but the symptom of intermittent 503s while all targets remain healthy points to a load balancer-level fault, not resource exhaustion. Vertical scaling would only mask the issue if the application is erroring on a request path; it introduces downtime and cost without giving diagnostic data. The ALB is the component generating the error, so you need to inspect its logs before making capacity changes.
- ✗
Enable detailed CloudWatch metrics on the EC2 instances to monitor CPU and memory.
Why it's wrong here
Detailed CloudWatch metrics—including memory, which requires the CloudWatch agent—monitor host resource utilization, but 503 responses are issued by the ALB itself when it cannot route a request to a target, irrespective of CPU or memory headroom. Since the target group shows a healthy host count, host-level metrics are unlikely to correlate with the failures. You need ALB access logs or ALB CloudWatch metrics (like HealthyHostCount and RequestCountPerTarget) to identify which requests are failing and why.
- ✓
Enable and review the ALB access logs stored in Amazon S3 to analyze the HTTP response codes and request patterns.
Why this is correct
Enabling ALB access logs and storing them in Amazon S3 records every request with fields such as elb_status_code, target_status_code, request_processing_time, target_processing_time, and response_processing_time, as well as the exact request path and user agent. By querying these logs, you can identify the specific HTTP response codes (including the 503s), observe patterns by path or client, and determine whether the timeouts occur in the load balancer or at the target. This is the definitive diagnostic step to isolate the cause of intermittent 503s.
- ✗
Increase the idle timeout setting on the ALB.
Why it's wrong here
The ALB idle timeout controls how long the load balancer keeps an idle connection open before closing it; when this expires during a request, the ALB returns a 504 Gateway Timeout, not a 503. A 503 Service Unavailable indicates the ALB has no healthy targets, is over its connection limit, or encountered a configuration error like a deregistered target. Extending the idle timeout would neither remediate the root cause nor change the error code, and it may actually mask slow-application behavior, so it is not the correct fix.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,298 original DOP-C02 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.