Courseiva

How to Count Errors in CloudWatch Logs and Alert with Metric Filters

A company uses Amazon CloudWatch Logs to store application logs. The SysOps administrator needs to detect when the number of log entries containing the string 'ERROR' exceeds 100 in any 5-minute window. When this threshold is breached, an email should be sent to the operations team. Which combination of AWS services should be used with the least operational overhead?

⚠ Common exam trap

The trap here is that candidates may overcomplicate the solution by choosing a Lambda-based or custom agent approach, not realizing that CloudWatch metric filters and alarms provide a fully managed, serverless way to monitor log patterns with minimal operational overhead.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Create a metric filter on the log group for 'ERROR', then create a CloudWatch alarm on that metric with an SNS action to send email.

It uses CloudWatch metric filters to extract the count of 'ERROR' log entries as a custom metric, then a CloudWatch alarm on that metric triggers an SNS topic to send email notifications. This approach requires no custom code or additional infrastructure, minimizing operational overhead while meeting the requirement of detecting >100 errors in any 5-minute period.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    CloudWatch Logs Insights scheduled query with SNS action.

    Why it's wrong here

    CloudWatch Logs Insights is an interactive query engine for analyzing log data retrospectively, not a real-time monitoring service. While you can schedule a query through an EventBridge rule to run periodically, the schedule has no built-in integration with CloudWatch Alarms—you would need a Lambda intermediary to evaluate results and publish to SNS. Moreover, scheduled queries often run on intervals of minutes and incur per-GB scan costs, so they lack the low-latency, threshold-based alerting that metric filters and alarms provide natively.

  • ✓

    Create a metric filter on the log group for 'ERROR', then create a CloudWatch alarm on that metric with an SNS action to send email.

    Why this is correct

    This is the native, fully managed approach: a metric filter is attached to the log group and processes new log events in real-time as CloudWatch Logs ingests them, extracting each occurrence of the pattern 'ERROR' and incrementing a custom CloudWatch metric. A CloudWatch alarm continuously evaluates that metric against the specified threshold (e.g., more than 100 errors in a 5-minute period); when the alarm transitions to the ALARM state, it invokes an SNS topic that sends the email notification. This pattern uses only built-in CloudWatch features and requires no custom code, external agents, or additional queueing services.

  • ✗

    Use a Lambda function that reads the log stream and sends an email via Amazon Simple Email Service (SES) when errors exceed 100.

    Why it's wrong here

    This option requires you to build and operate a custom log‑processing pipeline: you must create a Lambda function, subscribe it to the log group as a CloudWatch Logs subscription, parse incoming log events, and maintain state to track the count of 'ERROR' messages over time—then trigger SES only after exceeding your threshold. That logic duplicates the metric filter and alarm functionality that CloudWatch already offers, adds significant development and maintenance overhead, and risks missing alerts if the Lambda function fails or throttles. You also lose the built-in alarm evaluation and SNS integration, making this far more fragile and complex than the native solution.

  • ✗

    Install an agent on the application server that sends logs to Amazon SQS, then poll the queue with a Lambda function to trigger an email.

    Why it's wrong here

    This architecture introduces an unnecessary chain of components: you would have to install an agent on the application server to forward logs to SQS, then run a Lambda function that polls the queue, parses log events, and calls SES to send an email when errors exceed a threshold. Such a design adds multiple failure points (agent, SQS, Lambda, SES) and duplicated effort—CloudWatch Logs already has an agentless ingestion path and native filtering/alarming capabilities. It also lacks the integrated, automatic retries and state machine of CloudWatch Alarms, so you would need to manually handle poller timeouts, queue retention, and idempotency for duplicate messages.

About these practice questions

One of 1,169 original SOA-C02 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

2 more ways this is tested on SOA-C02

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. An application writes error logs to Amazon CloudWatch Logs. The SysOps administrator needs to monitor for the occurrence of the string 'ERROR' in the logs and trigger an Amazon SNS notification if more than 10 errors occur within a 5-minute window. The administrator also wants to visualize the error count over time. Which approach should be used to meet these requirements with the least operational overhead?

medium
  • ✓ A.Create a CloudWatch Logs metric filter to count 'ERROR' entries, then create a CloudWatch alarm on that metric with a period of 5 minutes and a threshold of 10.
  • B.Use CloudWatch Logs Insights to run a query every 5 minutes and send notifications via a scheduled AWS Lambda function.
  • C.Create an AWS Lambda function that processes log events in real-time and publishes to Amazon SNS when the error count exceeds 10 in 5 minutes.
  • D.Use Amazon EventBridge to match log events with the pattern 'ERROR' and send them to an SNS topic.

Why A: CloudWatch Logs metric filters can extract a count of 'ERROR' occurrences from incoming log events and emit a custom metric. A CloudWatch alarm on that metric with a period of 5 minutes and a threshold of 10 directly triggers an SNS notification when the error count exceeds 10 within the window, and the metric itself can be graphed in CloudWatch dashboards for visualization—all with minimal configuration and no custom code.

Variation 2. A company uses Amazon CloudWatch Logs to store application logs. The SysOps administrator needs to count the occurrences of the string 'ERROR' in the logs and trigger an Amazon SNS notification when more than 10 errors occur within a 5-minute window. Which steps should the administrator take?

easy
  • ✓ A.Create a metric filter on the log group and then create a CloudWatch alarm on the resulting metric
  • B.Create a CloudWatch alarm directly on the log group
  • C.Create an AWS Lambda function to parse the logs and send a notification to Amazon SNS
  • D.Create an Amazon EventBridge rule to filter log events and send to SNS

Why A: A metric filter on a CloudWatch Logs log group extracts a numeric metric (e.g., count of 'ERROR' occurrences) and publishes it to a CloudWatch custom metric. A CloudWatch alarm can then be configured on that metric to evaluate a threshold (e.g., >10) over a specified period (e.g., 5 minutes) and trigger an SNS notification when breached. This is the native, serverless, and cost-effective approach for counting log patterns and alerting.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This SOA-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SOA-C02 exam.