A company is running a batch processing job on Amazon EMR that writes results to an Amazon S3 bucket. The job runs daily and takes about 2 hours. The DevOps team wants to be alerted if the job fails or takes longer than 3 hours. Which solution is the MOST cost-effective and operationally efficient?
Trap 1: Configure Amazon Simple Notification Service (SNS) directly from…
Configuring SNS directly from the EMR job means embedding AWS SDK calls in the application code itself, which couples business logic to notification logic and requires custom error handling. EMR does not provide a native, managed hook to publish to SNS on cluster completion, so if the job fails before reaching the publish statement, no notification is sent. Additionally, this approach still does not inherently calculate or compare job duration against a threshold—it only sends raw completion signals.
Trap 2: Create a CloudWatch alarm on the EMR cluster's EC2 instance…
A CloudWatch alarm on CPUUtilization is fundamentally mismatched because CPU metrics reflect resource consumption, not elapsed job time, and a job may run for over three hours with low average CPU (e.g., while waiting for external dependencies or during shuffle bottlenecks). Conversely, a short-lived but compute-intensive job could spike CPU and trigger a false positive. Furthermore, an EMR cluster has multiple EC2 instances, so you would have to decide which instance's metric to alarm on, and none of these metrics indicate whether the job has actually completed, let alone exceeded a specific duration threshold.
Trap 3: Use Amazon CloudWatch Logs to monitor the job's log stream and…
This approach requires EMR logs to be streamed to CloudWatch Logs, which is not enabled by default—you must manually install and configure the CloudWatch agent on each core and task instance, adding operational overhead. Even with logs in place, a metric filter matching 'FAILED' messages only captures explicit failure keywords; it cannot detect a successful job that ran for over three hours, nor does it provide the job's start and end times needed to compute duration. Therefore, it fails both to measure duration and to alert on jobs that exceed the time limit but complete without a FAILED log line.
- A
Configure Amazon Simple Notification Service (SNS) directly from the EMR job to send notifications on completion.
Why it fails: Configuring SNS directly from the EMR job means embedding AWS SDK calls in the application code itself, which couples business logic to notification logic and requires custom error handling. EMR does not provide a native, managed hook to publish to SNS on cluster completion, so if the job fails before reaching the publish statement, no notification is sent. Additionally, this approach still does not inherently calculate or compare job duration against a threshold—it only sends raw completion signals.
- B
Use Amazon CloudWatch Events to trigger an AWS Lambda function when the EMR cluster changes to 'TERMINATED' state, then check the job duration and send an alert if it exceeded 3 hours.
This is correct because Amazon EventBridge (formerly CloudWatch Events) can natively capture EMR cluster state-change events, such as the transition to TERMINATED, and invoke a Lambda function without any polling or custom code. The Lambda function can call emr:describeCluster to obtain the cluster's start and end times, compute the total runtime, and post an SNS message only when the duration exceeds three hours. This event-driven architecture is cost-effective, serverless, and decoupled, making it the recommended pattern for alerting on abnormal job completion times.
- C
Create a CloudWatch alarm on the EMR cluster's EC2 instance CPUUtilization metric to detect abnormal runtime.
Why it fails: A CloudWatch alarm on CPUUtilization is fundamentally mismatched because CPU metrics reflect resource consumption, not elapsed job time, and a job may run for over three hours with low average CPU (e.g., while waiting for external dependencies or during shuffle bottlenecks). Conversely, a short-lived but compute-intensive job could spike CPU and trigger a false positive. Furthermore, an EMR cluster has multiple EC2 instances, so you would have to decide which instance's metric to alarm on, and none of these metrics indicate whether the job has actually completed, let alone exceeded a specific duration threshold.
- D
Use Amazon CloudWatch Logs to monitor the job's log stream and create a metric filter for 'FAILED' messages.
Why it fails: This approach requires EMR logs to be streamed to CloudWatch Logs, which is not enabled by default—you must manually install and configure the CloudWatch agent on each core and task instance, adding operational overhead. Even with logs in place, a metric filter matching 'FAILED' messages only captures explicit failure keywords; it cannot detect a successful job that ran for over three hours, nor does it provide the job's start and end times needed to compute duration. Therefore, it fails both to measure duration and to alert on jobs that exceed the time limit but complete without a FAILED log line.