DOP-C02 Monitoring and Logging Practice Question
A company is running a batch processing job on Amazon EMR that writes results to an Amazon S3 bucket. The job runs daily and takes about 2 hours. The DevOps team wants to be alerted if the job fails or takes longer than 3 hours. Which solution is the MOST cost-effective and operationally efficient?
⚠ Common exam trap
Many candidates assume CloudWatch Logs metric filters (Option D) are the simplest way to detect failures, but they miss the timeout requirement and require log-based failure patterns, whereas event-driven state monitoring (Option B) inherently captures both failure and duration scenarios without custom logging.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Amazon CloudWatch Events to trigger an AWS Lambda function when the EMR cluster changes to 'TERMINATED' state, then check the job duration and send an alert if it exceeded 3 hours.
It uses CloudWatch Events to detect the EMR cluster's 'TERMINATED' state, which triggers a Lambda function that can check the job duration against the 3-hour threshold and send an alert via SNS if needed. This approach is cost-effective (no polling, event-driven) and operationally efficient, as it decouples monitoring from the job itself and handles both failure and timeout scenarios without modifying the EMR job code.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Configure Amazon Simple Notification Service (SNS) directly from the EMR job to send notifications on completion.
Why it's wrong here
Configuring SNS directly from the EMR job means embedding AWS SDK calls in the application code itself, which couples business logic to notification logic and requires custom error handling. EMR does not provide a native, managed hook to publish to SNS on cluster completion, so if the job fails before reaching the publish statement, no notification is sent. Additionally, this approach still does not inherently calculate or compare job duration against a threshold—it only sends raw completion signals.
- ✓
Use Amazon CloudWatch Events to trigger an AWS Lambda function when the EMR cluster changes to 'TERMINATED' state, then check the job duration and send an alert if it exceeded 3 hours.
Why this is correct
This is correct because Amazon EventBridge (formerly CloudWatch Events) can natively capture EMR cluster state-change events, such as the transition to TERMINATED, and invoke a Lambda function without any polling or custom code. The Lambda function can call emr:describeCluster to obtain the cluster's start and end times, compute the total runtime, and post an SNS message only when the duration exceeds three hours. This event-driven architecture is cost-effective, serverless, and decoupled, making it the recommended pattern for alerting on abnormal job completion times.
- ✗
Create a CloudWatch alarm on the EMR cluster's EC2 instance CPUUtilization metric to detect abnormal runtime.
Why it's wrong here
A CloudWatch alarm on CPUUtilization is fundamentally mismatched because CPU metrics reflect resource consumption, not elapsed job time, and a job may run for over three hours with low average CPU (e.g., while waiting for external dependencies or during shuffle bottlenecks). Conversely, a short-lived but compute-intensive job could spike CPU and trigger a false positive. Furthermore, an EMR cluster has multiple EC2 instances, so you would have to decide which instance's metric to alarm on, and none of these metrics indicate whether the job has actually completed, let alone exceeded a specific duration threshold.
- ✗
Use Amazon CloudWatch Logs to monitor the job's log stream and create a metric filter for 'FAILED' messages.
Why it's wrong here
This approach requires EMR logs to be streamed to CloudWatch Logs, which is not enabled by default—you must manually install and configure the CloudWatch agent on each core and task instance, adding operational overhead. Even with logs in place, a metric filter matching 'FAILED' messages only captures explicit failure keywords; it cannot detect a successful job that ran for over three hours, nor does it provide the job's start and end times needed to compute duration. Therefore, it fails both to measure duration and to alert on jobs that exceed the time limit but complete without a FAILED log line.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DOP-C02 question is part of Courseiva's 1,298-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.