Courseiva

CCNA Monitoring, Logging, and Remediation Questions

25 of 250 questions · Page 4/4 · Monitoring, Logging, and Remediation · Answers revealed

226
MCQhard

An application writes logs to a file on an EC2 instance. The SysOps team needs to send these logs to Amazon CloudWatch Logs in real time. The logs must be encrypted at rest in CloudWatch Logs using a customer-managed KMS key. Which steps are required?

A.Use AWS CloudTrail to deliver logs to CloudWatch Logs with KMS encryption.
B.Store logs in S3 with KMS encryption and use S3 event notifications to trigger Lambda to put logs in CloudWatch Logs.
C.Install the CloudWatch Logs agent and enable encryption on the EC2 instance volume using KMS.
D.Install the CloudWatch Logs agent and associate a KMS key with the log group using the 'associate-kms-key' API.
AnswerD

This enables encryption at rest with a customer-managed key.

Why this answer

The CloudWatch Logs agent can send log data from an EC2 instance to CloudWatch Logs in real time, and the 'associate-kms-key' API (or the equivalent AWS CLI command 'put-log-group-encryption') allows you to associate a customer-managed KMS key with a log group, encrypting the logs at rest. This meets both the real-time delivery and customer-managed KMS encryption requirements without additional services or workarounds.

Exam trap

The trap here is that candidates often confuse encrypting the log file on the EC2 instance volume (Option C) with encrypting the logs at rest in CloudWatch Logs, or they overcomplicate the solution by introducing unnecessary services like S3 and Lambda (Option B) instead of using the native KMS integration with CloudWatch Logs.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail delivers API activity logs, not application log files from an EC2 instance, and it cannot be used to send arbitrary application logs to CloudWatch Logs in real time. Option B is wrong because storing logs in S3 and using S3 event notifications to trigger a Lambda function introduces latency and complexity, and does not provide real-time streaming to CloudWatch Logs; it also requires additional services and is not the standard method for real-time log ingestion. Option C is wrong because enabling encryption on the EC2 instance volume using KMS encrypts the log file at rest on the instance, but does not encrypt the logs at rest in CloudWatch Logs; the CloudWatch Logs agent sends data over the network, and the log group itself must be encrypted with a KMS key to meet the requirement.

227
MCQmedium

A company uses Amazon RDS for MySQL with Multi-AZ deployment. The SysOps administrator notices that the DB instance's CPU utilization spikes to 100% every few minutes. CloudWatch alarms have been set to trigger when CPU exceeds 90% for 5 minutes, but no alarm state changes are observed. The administrator checks the CloudWatch metrics and sees that the CPU utilization metric shows periodic spikes but they last only 2-3 minutes each. What is the most likely cause and what should the administrator do to receive notifications?

A.Set the alarm to evaluate over 1 minute instead of 5 minutes.
B.The Multi-AZ failover is causing the spikes; disable Multi-AZ.
C.The DB instance is not publishing CPU metrics at a high enough resolution.
D.The CPU metric is not accurate; use the CPU credit metric instead.
AnswerA

A CloudWatch alarm with a 5-minute period averages all CPUUtilization samples in that window, so a brief spike in utilization can be smoothed out and never crosses the threshold. Changing the period to 1 minute, which matches the standard metric publishing interval for RDS, ensures the short spike is evaluated as a discrete data point. That is the correct fix for detecting short-duration CPU spikes.

Why this answer

The CPU utilization spikes last only 2-3 minutes, which is shorter than the alarm's evaluation period of 5 consecutive minutes. Since the alarm requires the metric to exceed the 90% threshold for 5 minutes before triggering, these brief spikes never satisfy the alarm's duration condition. Setting the alarm to evaluate over 1 minute will match the spike duration and allow the alarm to trigger on these short-lived bursts.

Exam trap

The trap here is that candidates assume the alarm should trigger because the metric exceeds the threshold, but they overlook the requirement that the breach must persist for the entire evaluation period, not just momentarily.

How to eliminate wrong answers

Option B is wrong because Multi-AZ failover does not cause periodic CPU spikes; failover is a rare event triggered by planned maintenance or failure, not a recurring every-few-minutes pattern, and disabling Multi-AZ would reduce availability without addressing the underlying CPU issue. Option C is wrong because Amazon RDS for MySQL publishes CPU utilization metrics at 1-minute resolution by default (with detailed monitoring enabled), and the administrator can already see the spikes in CloudWatch, so resolution is not the problem. Option D is wrong because CPU credit metrics apply only to burstable performance instances (e.g., T2/T3), not to standard RDS instances, and the CPU metric is accurate; the issue is the alarm evaluation period, not the metric's validity.

228
MCQeasy

A SysOps administrator needs to monitor the application logs of a web server and receive an email notification when the number of 'ERROR' log entries exceeds 100 in a 5-minute window. The logs are already being sent to Amazon CloudWatch Logs. Which combination of AWS services should be used to meet this requirement with the least operational overhead?

A.CloudWatch Logs metric filter, CloudWatch alarm, and Amazon SNS
B.Amazon Kinesis Data Firehose and AWS Lambda
C.AWS CloudTrail and Amazon EventBridge
D.AWS Config managed rule and Amazon SNS
AnswerA

A CloudWatch Logs metric filter continuously scans incoming log events for a pattern you define (e.g., ERROR or Exception) and incrementally publishes a custom metric to CloudWatch. A CloudWatch alarm then evaluates that metric against a threshold, and when the threshold is breached, the alarm state changes to ALARM and triggers an Amazon SNS topic to send an email notification. This is a native, fully managed, near-real-time monitoring solution with no custom code to maintain, making it the correct architecture.

Why this answer

CloudWatch Logs metric filters can parse log events for the string 'ERROR' and count them in real time. A CloudWatch alarm can then trigger when the metric exceeds 100 in a 5-minute period, and Amazon SNS sends the email notification. This combination requires no custom code or additional infrastructure, minimizing operational overhead.

Exam trap

The trap here is that candidates may confuse CloudTrail (which logs API calls) with CloudWatch Logs (which stores application logs), leading them to choose Option C, but CloudTrail cannot inspect application log content.

How to eliminate wrong answers

Option B is wrong because Amazon Kinesis Data Firehose is designed for streaming large volumes of data to destinations like S3 or Redshift, not for real-time metric extraction and alerting; adding AWS Lambda would introduce custom code and increase complexity. Option C is wrong because AWS CloudTrail records API activity, not application log entries, and Amazon EventBridge is for event-driven workflows, not for counting log patterns. Option D is wrong because AWS Config managed rules evaluate resource compliance against desired configurations, not log content; they cannot parse log entries for 'ERROR' strings.

229
MCQhard

Refer to the exhibit. A SysOps administrator runs the CloudWatch Logs Insights query shown. What does this query do?

A.Groups ERROR and FATAL entries by log stream name.
B.Counts the number of ERROR and FATAL log entries per 5-minute interval and displays them in descending order by time.
C.Displays the full log messages of all ERROR and FATAL entries.
D.Deletes all log entries containing ERROR or FATAL older than 5 minutes.
AnswerB

This is the correct interpretation because the query uses the aggregation function 'stats count() by bin(5m)' which counts all matching ERROR/FATAL events within each 5-minute window, then applies 'sort @timestamp desc' to order the resulting time buckets from the most recent to the oldest. The combined pipeline first filters entries using a pattern match, then aggregates counts over fixed time intervals, and finally sorts the aggregated rows by the timestamp field in descending order. Therefore the output is a time series of error/fatal counts per five-minute bucket, not raw messages or stream-level groupings.

Why this answer

The CloudWatch Logs Insights query uses `stats count(*) by bin(5m)` to aggregate log events into 5-minute time buckets, then filters with `filter @message like /ERROR|FATAL/` to include only those severity levels. The `sort @timestamp desc` orders the resulting time buckets in descending chronological order, producing a count of ERROR and FATAL entries per 5-minute interval. This matches option B exactly.

Exam trap

The trap here is that candidates see `ERROR|FATAL` and `sort @timestamp desc` and assume the query returns raw log messages in reverse chronological order, overlooking that `stats count(*)` aggregates the data into counts per time bucket.

How to eliminate wrong answers

Option A is wrong because the query does not include `by @logStream` or any grouping on log stream name; it groups only by the 5-minute time bin. Option C is wrong because the query uses `stats count(*)` which returns counts, not the full log messages; to display full messages you would use `fields @message` without aggregation. Option D is wrong because CloudWatch Logs Insights is a read-only query engine that cannot delete log entries; deletion requires a separate API call or retention policy.

230
MCQmedium

A SysOps administrator needs to monitor the CPU utilization of an Amazon EC2 instance fleet and send an alert when the average CPU utilization exceeds 80% for 10 consecutive minutes. The administrator also wants to automatically stop the instance if the CPU utilization remains above 90% for 30 minutes to prevent runaway costs. Which combination of AWS services should be used?

A.Amazon CloudWatch alarm + AWS Lambda + AWS Systems Manager Automation
B.Amazon CloudWatch alarm + Amazon Simple Notification Service (SNS) + AWS Lambda
C.Amazon CloudWatch Logs + Amazon EventBridge + AWS Step Functions
D.AWS CloudTrail + Amazon EventBridge + AWS CodePipeline
AnswerB

A CloudWatch alarm monitors the CPU metric and publishes to an SNS topic when the threshold is breached. The SNS topic triggers a Lambda function that calls the EC2 StopInstances API to stop the instance. This is a clean, low-overhead solution.

Why this answer

It uses Amazon CloudWatch alarms to monitor CPU utilization metrics and trigger an SNS topic, which then invokes an AWS Lambda function. The Lambda function can execute the logic to stop the EC2 instance when the alarm state indicates CPU utilization above 90% for 30 minutes, providing automated cost control without manual intervention.

Exam trap

The trap here is that candidates may assume Systems Manager Automation (Option A) is required for instance stop actions, but Lambda is simpler and directly triggered by SNS, while Automation is better suited for complex multi-step workflows like patching or AMI creation.

How to eliminate wrong answers

Option A is wrong because AWS Systems Manager Automation is designed for predefined runbook-style remediation (e.g., patching, configuration changes) and is not directly triggered by CloudWatch alarms to stop an instance based on a metric threshold; it requires additional orchestration and does not natively support the stop action from an alarm. Option C is wrong because Amazon CloudWatch Logs is for log data, not metric monitoring, and Amazon EventBridge with Step Functions is overkill for a simple stop action; CloudWatch Logs cannot directly trigger alarms on CPU utilization metrics. Option D is wrong because AWS CloudTrail records API activity, not CPU metrics, and Amazon EventBridge with CodePipeline is for CI/CD pipelines, not for monitoring or stopping instances based on utilization thresholds.

231
MCQhard

A SysOps administrator manages multiple AWS accounts and wants to create a single Amazon CloudWatch dashboard that displays real-time metrics from all accounts in one view. The administrator needs to avoid managing separate dashboards for each account. Which solution should the administrator implement?

A.Use CloudWatch cross-account observability by setting up a monitoring account and sharing metrics from source accounts.
B.Export CloudWatch metrics to Amazon QuickSight and create a dashboard there.
C.Use AWS Config aggregator to collect metrics and display in CloudWatch.
D.Create a Lambda function that periodically pulls metrics from each account and publishes to a central account's CloudWatch.
AnswerA

CloudWatch cross-account observability uses a designated monitoring account in AWS Organizations to search, visualize, and create dashboards from source-account metrics without duplicating or exporting metric data. Enabling resource discovery from source accounts exposes their CloudWatch telemetry read-only, so dashboards show live data and remain useful for real-time operations rather than static exports.

Why this answer

CloudWatch cross-account observability allows you to designate a monitoring account that can view metrics, logs, and traces from multiple source accounts. This feature uses AWS Organizations or CloudWatch cross-account links to share observability data in real time, enabling a single dashboard that aggregates metrics from all accounts without needing separate dashboards.

Exam trap

The trap here is that candidates may confuse AWS Config aggregator (which aggregates configuration data) with CloudWatch cross-account observability (which aggregates monitoring metrics), leading them to choose a service that does not handle real-time metric visualization.

How to eliminate wrong answers

Option B is wrong because Amazon QuickSight is a business analytics service for interactive dashboards, not a real-time CloudWatch metrics viewer; it requires exporting metrics via API calls and cannot provide the low-latency, native CloudWatch dashboard experience. Option C is wrong because AWS Config aggregator collects configuration and compliance data, not real-time CloudWatch metrics; it is designed for resource inventory and rule evaluation, not for monitoring metric streams. Option D is wrong because creating a Lambda function to periodically pull metrics introduces latency, complexity, and potential data staleness; CloudWatch cross-account observability provides native, real-time streaming without custom code or polling overhead.

232
MCQeasy

A company has enabled AWS CloudTrail in all regions and is logging to an S3 bucket. The security team needs to be alerted within minutes if any IAM user creates a new access key. What is the MOST efficient way to achieve this?

A.Enable S3 event notifications on the CloudTrail bucket to trigger a Lambda function that parses logs and sends an alert.
B.Use AWS Config rules to detect changes to IAM access keys and trigger an SNS notification.
C.Configure CloudTrail to send logs to CloudWatch Logs. Create a metric filter for the IAM event 'CreateAccessKey' and set a CloudWatch alarm that sends an SNS notification.
D.Run a script on an EC2 instance that polls CloudTrail API for new events every minute and sends alerts.
AnswerC

CloudTrail can be configured to deliver all management events to a CloudWatch Logs log group in near-real-time, allowing you to analyze them with metric filters. By creating a metric filter that matches the event name "CreateAccessKey" and setting a CloudWatch alarm on the resulting metric, you get an SNS notification as soon as the event occurs, with minimal latency and no need for custom code or compute resources. This is the serverless best practice for security event monitoring on AWS.

Why this answer

CloudTrail can be configured to deliver events to CloudWatch Logs, where a metric filter can be created to match the 'CreateAccessKey' event. A CloudWatch alarm based on that metric filter can then trigger an SNS notification within minutes, providing the most efficient and native AWS solution for real-time alerting without custom code or polling.

Exam trap

The trap here is that candidates often choose S3 event notifications (Option A) because they think it's the simplest, but they overlook the built-in CloudWatch Logs integration which provides faster, more reliable, and fully managed alerting without custom code.

How to eliminate wrong answers

Option A is wrong because S3 event notifications on the CloudTrail bucket are not real-time; they can have delays and require a Lambda function to parse logs, which is less efficient than using CloudWatch metric filters. Option B is wrong because AWS Config rules are designed for compliance and configuration tracking, not for real-time event-driven alerting; they evaluate resources periodically or on configuration changes, not within minutes of an API call. Option D is wrong because running a script on an EC2 instance that polls the CloudTrail API every minute introduces latency, operational overhead, and a single point of failure, making it less efficient than the serverless, event-driven approach in option C.

233
MCQmedium

A company uses AWS CloudFormation to deploy its infrastructure. The SysOps administrator needs to be notified when a stack creation fails. Which solution meets this requirement with the LEAST effort?

A.Create a CloudWatch alarm that triggers when the CloudFormation stack status is 'CREATE_FAILED'.
B.Use AWS CloudTrail to monitor CreateStack API calls and trigger an SNS notification.
C.Configure an SNS topic in the CloudFormation stack's 'NotificationARNs' parameter.
D.Write a custom script that polls the CloudFormation API every minute and sends an SNS notification on failure.
AnswerC

The NotificationARNs parameter is a native CloudFormation feature that accepts one or more Amazon SNS topic ARNs, causing CloudFormation to publish all stack lifecycle events—including CREATE_FAILED, ROLLBACK_COMPLETE, and resource-level failures—to the topic. This is the built-in, real-time mechanism designed exactly for this scenario, requiring no custom code, polling, or metric configuration. Simply specify the SNS topic ARN when creating or updating the stack, and ensure the topic policy trusts CloudFormation's publishing role.

Why this answer

CloudFormation natively supports specifying an SNS topic in the 'NotificationARNs' parameter, which automatically sends notifications on stack events such as creation failure. This requires no additional infrastructure, scripting, or monitoring setup, making it the least-effort solution.

Exam trap

The trap here is that candidates often overthink and choose CloudWatch alarms or CloudTrail, not realizing that CloudFormation's built-in SNS notification parameter provides a zero-configuration, event-driven solution for stack failure alerts.

How to eliminate wrong answers

Option A is wrong because CloudWatch cannot directly alarm on CloudFormation stack status; CloudWatch alarms are designed for metrics (e.g., EC2 CPU utilization) and not for CloudFormation stack state changes. Option B is wrong because CloudTrail logs API calls but does not trigger SNS notifications directly; you would need additional services like EventBridge to route the event to SNS, adding complexity. Option D is wrong because writing a custom script to poll the CloudFormation API every minute introduces unnecessary overhead, latency, and maintenance effort, contradicting the 'least effort' requirement.

234
MCQeasy

A SysOps administrator is troubleshooting a Lambda function that does not write logs to CloudWatch Logs. The IAM role attached to the function includes the policy shown. What is the most likely reason the logs are not being created?

A.The log group name in the Resource ARN does not match the actual log group created by the Lambda function.
B.The IAM role is not assigned to the Lambda function's execution role.
C.The Lambda function is in a VPC without a VPC endpoint for CloudWatch Logs.
D.The policy does not include the logs:PutLogEvents permission.
AnswerA

Lambda automatically creates a log group named /aws/lambda/<function-name> in the same Region as the function. If the Resource ARN in the IAM policy points to a different log group name (e.g., a typo or a custom group like /aws/lambda/my-function-v2), CloudWatch Logs rejects the write attempt even though the role and actions are correct. The resulting CloudWatch Logs error typically indicates that the specified log group does not exist or the resource ARN does not match, which matches this root cause.

Why this answer

The IAM policy shown in the question includes a Resource ARN that specifies a specific log group name (e.g., `/aws/lambda/MyFunction`). If the Lambda function is configured to write to a different log group (e.g., `/aws/lambda/MyOtherFunction` or a custom log group), the `logs:CreateLogGroup` and `logs:CreateLogStream` permissions will fail because the ARN does not match. This mismatch prevents the function from creating the log group or stream, so no logs are written to CloudWatch Logs.

Exam trap

The trap here is that candidates often overlook the Resource ARN mismatch and instead focus on missing permissions or VPC connectivity, but the core issue is that the IAM policy's log group ARN does not match the actual log group name Lambda tries to use.

How to eliminate wrong answers

Option B is wrong because the IAM role is already attached to the Lambda function's execution role (the policy is part of that role), so the assignment is not the issue. Option C is wrong because a Lambda function in a VPC can still write logs to CloudWatch Logs via the public internet or a NAT gateway; a VPC endpoint for CloudWatch Logs is not required unless the VPC has no internet access and no NAT gateway. Option D is wrong because the policy shown includes `logs:PutLogEvents` (the question states the policy includes it), so the missing permission is not the cause.

235
MCQmedium

A company's application running on EC2 instances is experiencing intermittent errors. The SysOps team needs to collect and analyze application logs from all instances centrally. The logs must be stored durably and searchable with minimal latency. Which solution meets these requirements?

A.Enable AWS CloudTrail and store logs in an S3 bucket.
B.Use Amazon Kinesis Data Firehose to send logs directly from each instance to Amazon Redshift.
C.Install the CloudWatch Logs agent on each EC2 instance and stream logs to Amazon CloudWatch Logs.
D.Store logs locally on each instance and periodically copy them to Amazon S3.
AnswerC

Installing the CloudWatch Logs agent on each EC2 instance enables near-real-time streaming of application and system log files to Amazon CloudWatch Logs, providing a centralized, scalable, and searchable log store. With CloudWatch Logs you can create metric filters to trigger CloudWatch alarms, use Logs Insights to run queries across log groups, and perform live tailing to watch logs as they arrive. The agent also handles log rotation, multi-line log records, and timestamp parsing automatically, which simplifies log management and removes the need for manual log harvesting or extra infrastructure.

Why this answer

The CloudWatch Logs agent (or unified CloudWatch agent) installed on each EC2 instance can stream application logs in near real-time to Amazon CloudWatch Logs, which provides durable storage, automatic encryption at rest, and a searchable interface via the console, CLI, or API with minimal latency. This centralized logging solution meets the requirements for collecting logs from all instances, storing them durably, and enabling immediate querying without additional infrastructure.

Exam trap

The trap here is that candidates may confuse CloudTrail (API logging) with application logging, or assume that S3 periodic uploads are sufficient for 'minimal latency' searchability, when in fact CloudWatch Logs is the native AWS service designed for real-time log ingestion and querying from EC2 instances.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail records API activity for governance and auditing, not application-level logs generated by processes running on EC2 instances; it cannot capture stdout, stderr, or custom application log files. Option B is wrong because Amazon Kinesis Data Firehose is a streaming ingestion service that delivers data to destinations like S3 or Redshift, but sending logs directly from each instance to Firehose without an agent or SDK is not a standard pattern, and Amazon Redshift is a data warehouse optimized for analytical queries, not a low-latency log search engine; this adds unnecessary complexity and cost. Option D is wrong because storing logs locally on each instance risks data loss on instance termination or failure, and periodically copying logs to S3 introduces latency that prevents real-time searchability, failing the 'minimal latency' requirement.

236
MCQeasy

A SysOps administrator needs to track changes to security groups in the AWS account. Which AWS service should be used to record configuration changes and provide a history of security group modifications?

A.AWS Trusted Advisor
B.Amazon CloudWatch
C.AWS Config
D.AWS CloudTrail
AnswerC

AWS Config is the correct service because it continuously records configuration items for supported resources, including security groups, and maintains a configuration history. When a security group rule is added or removed, AWS Config generates a configuration item and allows you to review the previous and new state using the configuration timeline. It also enables compliance rules to detect and evaluate changes, making it the definitive service for tracking and auditing security group changes.

Why this answer

AWS Config is the correct service because it provides a detailed inventory of AWS resources, records configuration changes, and maintains a historical timeline of those changes. For security groups, AWS Config can track modifications such as rule additions, deletions, or updates, and it can trigger evaluations against desired configurations. This makes it the ideal service for auditing and compliance use cases involving security group changes.

Exam trap

The trap here is that candidates often confuse AWS CloudTrail (which logs API calls) with AWS Config (which records resource configuration state and history), leading them to choose CloudTrail for change tracking when Config is the service designed for configuration history and compliance auditing.

How to eliminate wrong answers

Option A is wrong because AWS Trusted Advisor is an advisory service that inspects your AWS environment and makes recommendations based on AWS best practices, but it does not record or maintain a history of configuration changes to resources like security groups. Option B is wrong because Amazon CloudWatch is a monitoring service for metrics, logs, and alarms; it can detect and alert on changes via CloudWatch Events, but it does not natively store a historical record of configuration changes or provide a timeline of modifications. Option D is wrong because AWS CloudTrail records API calls and events, including those that modify security groups, but it focuses on who made the call and when, not on the state or configuration history of the resource itself; CloudTrail does not provide a point-in-time configuration snapshot or a change timeline for the resource's configuration.

237
Multi-Selecthard

A SysOps administrator is troubleshooting a Lambda function that is not processing messages from an SQS queue. The function is subscribed to the queue via an event source mapping. The function has a reserved concurrency of 0. Which TWO actions will resolve the issue?

Select 2 answers
A.Add SQS permissions to the Lambda execution role.
B.Configure a dead-letter queue for the Lambda function.
C.Set the reserved concurrency to a value greater than 0.
D.Enable the event source mapping if it is disabled.
E.Increase the batch size in the event source mapping.
AnswersC, D

Reserved concurrency of 0 is an intentional hard throttle that blocks all function invocations, causing Lambda to return a throttling error for every request. By raising reserved concurrency to any positive value, you grant the function a dedicated pool of execution capacity, allowing the event source mapping to successfully invoke it. This is the most direct fix when the function cannot execute at all, because it removes the service-level blocking condition that prevents both synchronous and event source mapping invocations.

Why this answer

Reserved concurrency of 0 means the Lambda function has no available execution capacity, so it cannot process any invocations, including those from SQS. Setting reserved concurrency to a value greater than 0 (e.g., 1 or more) allocates the necessary execution slots for the function to run. This directly resolves the issue because the event source mapping will successfully invoke the function only when concurrency is available.

Exam trap

The trap here is that candidates often overlook reserved concurrency of 0 as a valid configuration that completely blocks invocations, and instead focus on permissions or queue settings, not realizing that a concurrency limit of 0 is a deliberate disablement mechanism.

238
MCQmedium

A company uses Amazon CloudWatch Logs to store application logs. The security team needs to be alerted when any log group contains a specific error pattern. The solution must minimize latency and operational overhead. What should a SysOps administrator do?

A.Stream the logs to Amazon Kinesis Data Firehose, which then triggers a Lambda function to check for errors.
B.Create a CloudWatch metric filter on the log group and set an alarm that triggers an SNS notification.
C.Create a Lambda function subscribed to the CloudWatch Logs log group, which checks for the error pattern and publishes to an SNS topic.
D.Use CloudWatch Logs Insights to run a query every minute and send results via SNS.
AnswerC

When you configure a subscription filter on a CloudWatch Logs group, CloudWatch Logs asynchronously invokes a Lambda function as log events are ingested, delivering a gzip-compressed batch of data that you can decode and search for error signatures. The Lambda function can be written in Python, Node.js, or another supported runtime, applying arbitrary regexes, aggregating events, and then directly publishing to an SNS topic for immediate fan-out to email, SMS, or Chatbot. This pattern provides real-time, event-driven alerting with minimal latency and no extra infrastructure, which is why it is the recommended approach for this requirement.

Why this answer

Subscribing a Lambda function directly to a CloudWatch Logs log group allows real-time, low-latency processing of log events as they arrive. The Lambda function can parse each log event for the specific error pattern and publish to an SNS topic to alert the security team, minimizing operational overhead by avoiding additional streaming or polling services.

Exam trap

The trap here is that candidates may choose Option B (metric filter and alarm) because it seems simpler, but they overlook that metric filters only count occurrences over time and cannot trigger immediate, per-event alerts, which is required for minimizing latency in security alerting.

How to eliminate wrong answers

Option A is wrong because streaming logs to Kinesis Data Firehose adds unnecessary latency and operational complexity; Firehose is designed for batch delivery to destinations like S3 or Redshift, not for real-time alerting with minimal latency. Option B is wrong because a CloudWatch metric filter counts occurrences of a pattern but cannot trigger an alarm on a per-log-event basis; alarms are evaluated periodically (e.g., every minute) and require a threshold, introducing latency and potential missed alerts for sporadic errors. Option D is wrong because CloudWatch Logs Insights queries are on-demand or scheduled at intervals (minimum 1 minute), not real-time, and require manual or scheduled execution, increasing latency and operational overhead compared to event-driven processing.

239
MCQhard

A SysOps administrator is troubleshooting a slow-running application on an EC2 instance. CloudWatch metrics show high CPU utilization but low disk I/O. The instance type is t3.medium. Which action would most likely improve performance?

A.Change the instance type to c5.large, which provides dedicated CPU performance.
B.Increase the instance memory by changing to r5.large.
C.Enable EBS-optimized on the instance and use provisioned IOPS SSD volumes.
D.Increase the size of the EBS volume to improve disk throughput.
AnswerA

c5.large instances are compute-optimized with dedicated vCPUs, so they do not rely on CPU credits or burst capacity. When the current burstable instance exhausts its CPU credit balance, it is throttled to the baseline utilization level, causing the application to run slowly. By moving to c5.large, you get consistent, full CPU performance for sustained periods, directly addressing the CPU-bound bottleneck.

Why this answer

The t3.medium is a burstable instance that relies on CPU credits. High CPU utilization with low disk I/O indicates the application is CPU-bound and the instance has likely exhausted its CPU credits, causing performance throttling. Changing to a c5.large provides dedicated, consistent CPU performance without credit-based limitations, directly addressing the bottleneck.

Exam trap

The trap here is that candidates may focus on disk or memory improvements because the application is 'slow,' but the CloudWatch metrics clearly point to a CPU bottleneck, and the t3 family's credit-based performance model is a common exam pitfall.

How to eliminate wrong answers

Option B is wrong because increasing memory (r5.large) does not resolve CPU starvation; the metrics show high CPU utilization, not memory pressure. Option C is wrong because EBS optimization and provisioned IOPS improve disk throughput, but disk I/O is already low, indicating the bottleneck is not storage-related. Option D is wrong because increasing EBS volume size does not inherently improve disk throughput; throughput depends on volume type and IOPS, not size alone, and disk I/O is not the issue.

240
MCQeasy

A SysOps administrator is troubleshooting an issue where an EC2 instance's CPU utilization is consistently above 90%, but no CloudWatch alarm is triggered. The alarm is configured to monitor the 'CPUUtilization' metric with a threshold of 80% for 2 consecutive periods of 5 minutes. What is the most likely cause?

A.The alarm is in the 'OK' state and not 'INSUFFICIENT_DATA'.
B.The CPUUtilization metric is not enabled by default for EC2 instances.
C.The CPU utilization spikes above 80% for less than 10 minutes at a time.
D.The alarm period is set to 5 minutes, but the metric is reported every 1 minute.
AnswerC

This is correct because CloudWatch alarms evaluate a metric against the threshold over a specified number of consecutive periods. If the alarm is configured with a period of 5 minutes and evaluation periods of 2, the CPU utilization must exceed 80% for the entire 10-minute span covered by two consecutive data points. When the CPU spikes above 80% for less than 10 minutes, it never passes two full evaluation periods, so the alarm state remains OK. Thus, short-lived spikes, even if severe, will not trigger the alarm.

Why this answer

The CloudWatch alarm requires 2 consecutive periods of 5 minutes (i.e., 10 minutes total) where the CPU utilization exceeds 80%. If the CPU utilization spikes above 80% for less than 10 minutes at a time, the alarm will not trigger because it never meets the consecutive evaluation period requirement. The alarm evaluates each 5-minute period independently, and only when both consecutive periods breach the threshold does the alarm state change to ALARM.

Exam trap

The trap here is that candidates often assume any breach of the threshold triggers the alarm immediately, but they overlook the 'consecutive periods' requirement, which means the alarm only fires after the condition persists for the full evaluation window (e.g., 10 minutes for 2 periods of 5 minutes).

How to eliminate wrong answers

Option A is wrong because the alarm being in the 'OK' state is the result of the condition not being met, not the cause of the alarm not triggering; the question asks for the cause of no alarm being triggered, and the alarm state is a symptom, not a root cause. Option B is wrong because the CPUUtilization metric is enabled by default for all EC2 instances and is available in CloudWatch without any additional configuration; it is a standard metric that is automatically sent every 5 minutes (or 1 minute with detailed monitoring). Option D is wrong because the metric being reported every 1 minute (with detailed monitoring) does not prevent the alarm from triggering; the alarm period of 5 minutes means CloudWatch aggregates the 1-minute data points into 5-minute averages, and the alarm evaluates those averages against the threshold, so the reporting interval does not cause the alarm to fail.

241
MCQhard

An application running on EC2 instances occasionally throws 'Connection refused' errors when connecting to an RDS database. The SysOps administrator needs to determine if the issue is due to database connection limits or network security groups. Which metrics and logs should the administrator examine?

A.Check CloudWatch RDS CPUUtilization and CloudTrail logs for RDS API calls.
B.Review RDS error logs in CloudWatch Logs and check the EC2 instance's system log.
C.Look at the EC2 instance's CloudWatch NetworkIn and NetworkOut metrics and RDS FreeableMemory metric.
D.Examine the RDS CloudWatch metric DatabaseConnections and analyze VPC Flow Logs for the EC2 instance's network interface.
AnswerD

DatabaseConnections shows whether the RDS instance has hit its connection limit, while VPC Flow Logs reveal whether security groups or network ACLs are rejecting traffic. Together they distinguish a connection-limit problem from a network-blocking problem.

Why this answer

'Connection refused' errors typically stem from either the database exhausting its maximum connections or network-level security groups blocking traffic. The RDS CloudWatch metric `DatabaseConnections` directly shows the current number of active connections against the instance's `max_connections` limit, while VPC Flow Logs capture whether packets are being accepted or rejected by security groups or network ACLs, pinpointing network blockages.

Exam trap

The trap here is that candidates confuse aggregate network metrics (like NetworkIn/NetworkOut) or CPU metrics with the specific indicators needed to differentiate between connection limits and security group denials, leading them to choose options that measure volume rather than connection state or packet acceptance.

How to eliminate wrong answers

Option A is wrong because `CPUUtilization` does not indicate connection limits or security group blocks, and CloudTrail logs record API calls (e.g., creating DB instances) not real-time connection or network failures. Option B is wrong because RDS error logs in CloudWatch Logs may show authentication or query errors but not connection limit exhaustion or network-level rejections, and the EC2 instance's system log (console output) does not capture network flow data. Option C is wrong because `NetworkIn`/`NetworkOut` show aggregate traffic volume, not whether connections are accepted or rejected, and `FreeableMemory` indicates memory pressure but not connection count or security group rules.

242
MCQeasy

A SysOps administrator needs to monitor the health of an Amazon RDS for MySQL DB instance. The administrator wants to receive an alert when the database connection count exceeds a threshold of 500 for more than 5 minutes. Which AWS service should be used to create this alert?

A.Amazon CloudWatch
B.Amazon Simple Notification Service (SNS)
C.AWS CloudTrail
D.AWS Config
AnswerA

Amazon CloudWatch is the native AWS monitoring service that retrieves and stores metrics from RDS, including the DatabaseConnections metric. You can create a CloudWatch alarm that evaluates this metric every minute, and if it exceeds 500 concurrently for 5 consecutive minutes, the alarm transitions to ALARM state and can trigger SNS notifications or Auto Scaling actions. This makes CloudWatch the correct service for detecting and responding to database connection pressure.

Why this answer

Amazon CloudWatch is the correct service because it can monitor RDS metrics such as DatabaseConnections and trigger an alarm when the value exceeds a threshold of 500 for a specified duration (e.g., 5 consecutive evaluation periods). CloudWatch alarms evaluate metric data against a defined threshold and can then publish to an SNS topic to send notifications.

Exam trap

The trap here is that candidates confuse the service that evaluates metrics (CloudWatch) with the service that delivers notifications (SNS), leading them to select SNS because they think of alerts as notifications, but CloudWatch is the service that creates and evaluates the alarm based on the metric threshold.

How to eliminate wrong answers

Option B (Amazon SNS) is wrong because SNS is a notification service, not a monitoring or alert evaluation service; it cannot itself evaluate metric thresholds or create alarms. Option C (AWS CloudTrail) is wrong because CloudTrail records API calls for auditing and governance, not real-time performance metrics like database connection counts. Option D (AWS Config) is wrong because Config evaluates resource configurations and compliance rules, not operational metrics such as connection counts.

243
MCQeasy

A company stores critical data in an S3 bucket and wants to be notified immediately when any object is deleted from the bucket. Which combination of services should the SysOps administrator use?

A.Configure an S3 event notification for 's3:ObjectRemoved:*' events to send to an SNS topic.
B.Enable S3 server access logging and send logs to CloudWatch Logs, then create a metric filter and alarm.
C.Use S3 event notifications to invoke a Lambda function that checks the object and sends an email.
D.Use AWS CloudTrail to log DeleteObject calls and create a CloudWatch Events rule to send an SNS notification.
AnswerA

S3 event notifications are the direct, native mechanism for reacting to object lifecycle events. Configuring a notification for the `s3:ObjectRemoved:*` prefix covers both individual `DeleteObject` and multi-object `DeleteObjects` API calls, and SNS can deliver an email immediately upon event publication. With an SNS topic subscribed to by email, the notification is pushed in near real-time, typically within seconds, with no polling or compute layer required. This makes it the simplest and most appropriate option for instant alerting.

Why this answer

S3 event notifications can be configured to trigger on 's3:ObjectRemoved:*' events, which cover both permanent and versioned object deletions. These notifications can be sent directly to an SNS topic, enabling immediate notification without additional compute or logging overhead. This is the simplest and most direct approach for real-time alerts on object deletions.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing CloudTrail or Lambda, not realizing that S3 event notifications can directly trigger SNS for immediate alerts without additional services or delays.

How to eliminate wrong answers

Option B is wrong because S3 server access logs are delivered on a best-effort basis, often with delays of several hours, making them unsuitable for immediate notification. Option C is wrong because invoking a Lambda function to check the object and send an email adds unnecessary complexity and latency; the event notification can directly send to SNS without custom code. Option D is wrong because CloudTrail logs are typically delivered within 5-15 minutes, not in real time, and using CloudTrail for this purpose introduces additional cost and complexity compared to native S3 event notifications.

244
MCQhard

A SysOps administrator is troubleshooting an issue where an EC2 instance's CloudWatch agent is not sending memory metrics. The agent is installed and configured to collect memory metrics. The IAM role attached to the instance has the following policy: { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": "cloudwatch:PutMetricData", "Resource": "*" }, { "Effect": "Allow", "Action": "cloudwatch:ListMetrics", "Resource": "*" } ] } What is the most likely reason the memory metrics are not appearing?

A.The CloudWatch agent requires the SSM Agent to be installed.
B.The CloudWatch agent must be configured to send metrics to CloudWatch Logs first.
C.The IAM role is missing the ssm:GetParameter permission.
D.Detailed monitoring must be enabled on the EC2 instance.
AnswerC

This is the root cause because the CloudWatch agent is configured to retrieve its metrics configuration from the SSM Parameter Store. To read that configuration, the EC2 instance's IAM role must include the ssm:GetParameter permission; without it, the agent cannot fetch the JSON that defines which memory metrics to collect. Even though the role includes cloudwatch:PutMetricData, the agent has no configuration to act on, so no memory metrics are published.

Why this answer

The IAM policy includes cloudwatch:PutMetricData and cloudwatch:ListMetrics, but if the CloudWatch agent retrieves its memory metrics configuration from Systems Manager Parameter Store (a common setup), the instance requires the ssm:GetParameter permission. Without this permission, the agent cannot fetch the configuration, and memory metrics will not be sent. Therefore, option C correctly identifies the missing IAM permission.

Exam trap

Candidates often focus solely on the cloudwatch:PutMetricData permission and overlook other required permissions like ssm:GetParameter. The policy already includes PutMetricData, so option C is a distractor if not read carefully; however, the real issue is the missing ssm:GetParameter permission needed for configuration retrieval from Parameter Store.

How to eliminate wrong answers

Option A is wrong because the CloudWatch agent does not require the SSM Agent to be installed; the CloudWatch agent can run independently and communicate directly with CloudWatch via HTTPS. Option B is wrong because the CloudWatch agent sends metrics directly to CloudWatch Metrics, not to CloudWatch Logs first; memory metrics are custom metrics, not log data. Option D is wrong because detailed monitoring on the EC2 instance only enables 1-minute frequency for hypervisor-level metrics (CPU, network, disk), not memory metrics; memory metrics are collected by the CloudWatch agent and require the agent to be running and properly configured.

245
MCQmedium

Multiple microservices each write structured JSON logs to separate CloudWatch log groups. The operations team needs to find all ERROR-level log entries across all log groups for the past 24 hours and count errors by service name. Which approach achieves this with the least operational overhead?

A.Run a CloudWatch Logs Insights query selecting all relevant log groups, filter where level = 'ERROR', and use stats count(*) by service
B.Export each log group to S3 and run an Athena query joining all exported files
C.Subscribe all log groups to a Kinesis Data Firehose stream and query the aggregated data in OpenSearch
D.Use the AWS CLI to download and grep log events from each log group separately, then sum the results
AnswerA

Logs Insights accepts a comma-separated list of log group names (or a log group name prefix pattern) in the query scope. The filter and stats commands work across all selected groups in a single query execution. No additional pipeline or aggregation layer is needed.

Why this answer

CloudWatch Logs Insights natively supports querying multiple log groups in a single query. By specifying all relevant log groups in the query scope, filtering for `level = 'ERROR'` using the `filter` command, and using `stats count(*) by service`, the operations team can directly aggregate error counts per service without any data movement, additional infrastructure, or manual scripting. This approach has the least operational overhead because it leverages existing CloudWatch capabilities with no setup or maintenance.

Exam trap

The trap here is that candidates may overcomplicate the solution by assuming cross-log-group analysis requires data aggregation pipelines (like Kinesis or S3/Athena), when CloudWatch Logs Insights natively supports querying multiple log groups with a single query, making it the simplest and most cost-effective option.

How to eliminate wrong answers

Option B is wrong because exporting logs to S3 and querying with Athena introduces significant operational overhead: you must set up S3 buckets, configure export schedules (which can take hours for large volumes), and manage Athena table definitions and partitions, all of which are unnecessary when CloudWatch Logs Insights can query the same data directly. Option C is wrong because subscribing all log groups to a Kinesis Data Firehose stream and querying in OpenSearch requires provisioning and managing a Firehose delivery stream, an OpenSearch cluster, and index management, which adds complexity and cost far beyond the simple Insights query. Option D is wrong because using the AWS CLI to download and grep log events from each log group separately is manual, error-prone, and does not scale; it requires scripting to handle pagination, rate limits, and log group enumeration, and it lacks the built-in aggregation and filtering capabilities of CloudWatch Logs Insights.

246
MCQeasy

A SysOps administrator wants to be alerted when an EC2 instance is terminated unexpectedly. Which CloudWatch event should be used to trigger a notification?

A.A CloudWatch alarm on the CPUUtilization metric dropping to zero.
B.A CloudTrail trail that logs TerminateInstances API calls.
C.A CloudWatch alarm on the StatusCheckFailed metric.
D.An Amazon EventBridge rule that matches EC2 Instance State-change Notification events.
AnswerD

An EventBridge rule can match the EC2 Instance State-change Notification event, which is emitted synchronously as an instance transitions among states such as pending, running, stopping, stopped, and terminated. By configuring an event pattern for detail.state = "terminated" and targeting an SNS topic or Lambda function, you receive a near-real-time notification that is the native, purpose-built mechanism for alerting on lifecycle changes.

Why this answer

Amazon EventBridge can capture EC2 Instance State-change Notification events, which are emitted whenever an EC2 instance transitions between states (e.g., running, stopped, terminated). By creating a rule that matches the 'terminated' state, the administrator can trigger an SNS notification or Lambda function to alert on unexpected termination, providing a real-time, event-driven response.

Exam trap

The trap here is that candidates confuse CloudWatch alarms on metrics (like CPUUtilization or StatusCheckFailed) with event-driven notifications, overlooking that EventBridge rules directly capture state-change events for immediate, precise alerting.

How to eliminate wrong answers

Option A is wrong because a CloudWatch alarm on CPUUtilization dropping to zero is not a reliable indicator of termination; an instance could be idle or stopped without being terminated, and the alarm would not fire immediately upon termination. Option B is wrong because a CloudTrail trail logging TerminateInstances API calls records the API action but does not directly trigger a notification; it requires additional integration (e.g., CloudWatch Logs metric filter or EventBridge rule) to generate alerts, and it only captures API-initiated terminations, not those from Auto Scaling or AWS Health events. Option C is wrong because a CloudWatch alarm on StatusCheckFailed monitors system or instance status checks (e.g., OS-level issues), not termination events; an instance can fail status checks without being terminated, and termination does not necessarily cause a status check failure.

247
MCQhard

A SysOps administrator is troubleshooting an issue where an Amazon RDS DB instance's storage space is running out. The administrator has enabled CloudWatch alarms for FreeStorageSpace, but the alarm did not trigger before the storage was exhausted. What is the most likely reason?

A.The FreeStorageSpace metric is not available for the selected DB instance class.
B.The alarm was configured to use a static threshold but the metric is not emitted during storage operations.
C.The alarm's evaluation period was too long and the storage filled up faster than the alarm could trigger.
D.The alarm was monitoring the wrong metric, such as 'Storage' instead of 'FreeStorageSpace'.
AnswerC

An evaluation period (the number of consecutive 60-second periods the metric must breach the threshold before the alarm enters ALARM state) lengthens the required detection time. If the database's storage fills up completely within, say, 15 minutes, an alarm configured with 10 evaluation periods at 5-minute periods would need 50 minutes of data and thus never fire before the instance goes read-only. Shortening the evaluation period and using a threshold well above zero (e.g., 10% free space) avoids this pitfall.

Why this answer

CloudWatch alarms evaluate metrics based on a specified evaluation period (e.g., 5 minutes). If the storage fills up faster than the alarm's evaluation period, the alarm may not have enough data points to trigger before the storage is exhausted. This is a common issue when the rate of storage consumption exceeds the alarm's evaluation frequency.

Exam trap

The trap here is that candidates assume CloudWatch alarms trigger instantly when a metric crosses a threshold, but in reality, alarms require multiple data points over the evaluation period to change state, which can delay detection if storage fills rapidly.

How to eliminate wrong answers

Option A is wrong because FreeStorageSpace is a standard metric available for all RDS DB instance classes, including those with General Purpose (gp2/gp3) or Provisioned IOPS (io1/io2) storage. Option B is wrong because the FreeStorageSpace metric is emitted continuously during storage operations, regardless of whether the alarm uses a static threshold or anomaly detection. Option D is wrong because 'Storage' is not a valid CloudWatch metric name for RDS; the correct metric is FreeStorageSpace, and monitoring the wrong metric would not cause the alarm to fail to trigger—it would simply not reflect storage exhaustion.

248
MCQeasy

A company wants to centrally collect and analyze logs from all AWS accounts in an organization. The logs include CloudTrail, VPC Flow Logs, and AWS Config logs. Which solution is the most scalable and cost-effective?

A.Stream all logs to a central CloudWatch Logs account using cross-account subscriptions.
B.Use CloudWatch Logs Insights to query logs from each account individually.
C.Use Amazon Kinesis Data Firehose to deliver logs to an Amazon Elasticsearch Service cluster.
D.Configure each account to deliver logs to a centralized S3 bucket and use Amazon Athena to query them.
AnswerD

This option centralizes logs by having each account deliver log files to a single S3 bucket—either directly or via S3 Cross-Account Replication—and then uses Amazon Athena to run SQL queries across that centralized data lake. S3 offers cheap, durable, and scalable storage for large volumes of logs, while Athena is serverless and charges only for the data actually scanned per query. By partitioning the S3 paths by account, region, and date and using AWS Glue Data Catalog (or a crawler), analysts can run cross-account queries from one place without managing any query infrastructure, making it the most cost-effective and operationally simple solution.

Why this answer

It uses a centralized S3 bucket to aggregate logs from all accounts, which is highly scalable and cost-effective due to S3's low storage costs and lifecycle policies. Amazon Athena then allows serverless, pay-per-query analysis of the logs without needing to provision or manage any infrastructure, making it ideal for ad-hoc and cross-account log analysis.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing managed services like CloudWatch Logs or Elasticsearch, overlooking the simplicity, scalability, and cost-effectiveness of S3 + Athena for centralized log analysis across multiple accounts.

How to eliminate wrong answers

Option A is wrong because streaming all logs to a central CloudWatch Logs account via cross-account subscriptions incurs high ingestion and storage costs, and CloudWatch Logs is not designed for long-term, cost-effective storage of large volumes of logs from multiple accounts. Option B is wrong because CloudWatch Logs Insights can only query logs within a single account and cannot aggregate or query logs across multiple accounts, failing the centralization requirement. Option C is wrong because using Amazon Kinesis Data Firehose to deliver logs to an Amazon Elasticsearch Service cluster introduces significant operational overhead for managing the Elasticsearch cluster, and the cost scales with the volume of data indexed and stored, making it less cost-effective than S3 + Athena for infrequent or ad-hoc queries.

249
MCQeasy

A company uses Amazon RDS for MySQL and wants to monitor the number of database connections in real time. Which CloudWatch metric should the SysOps administrator use?

A.DatabaseConnections
B.CPUUtilization
C.ActiveConnections
D.ConnectionCount
AnswerA

DatabaseConnections is the correct standard CloudWatch metric because Amazon RDS publishes it in the AWS/RDS namespace as a gauge that reports the number of client network connections to the DB instance. It includes all active sessions from applications, readers, and management tools, and it is the authoritative counter to compare against the instance's max_connections, which varies by instance class and DB engine. This metric directly answers the question of how many concurrent connections are present.

Why this answer

Amazon RDS for MySQL exposes a CloudWatch metric named `DatabaseConnections` that reports the number of client connections to the DB instance. This metric is derived from the MySQL `Threads_connected` status variable and is updated every minute, providing real-time visibility into connection counts for monitoring and alarming.

Exam trap

The trap here is that candidates confuse the MySQL status variable `Threads_connected` (which maps to `DatabaseConnections`) with `Threads_running` (which counts only actively executing queries) or assume a generic name like `ActiveConnections` or `ConnectionCount` exists as a CloudWatch metric.

How to eliminate wrong answers

Option B is wrong because `CPUUtilization` measures the percentage of CPU used by the DB instance, not the number of database connections. Option C is wrong because `ActiveConnections` is not a standard CloudWatch metric for RDS; it may be confused with a MySQL status variable (`Threads_running`) but is not exposed as a CloudWatch metric. Option D is wrong because `ConnectionCount` is not a valid CloudWatch metric name for RDS; the correct metric is `DatabaseConnections`.

250
MCQhard

A SysOps administrator is investigating why a CloudWatch alarm did not trigger an SNS notification. The alarm state changed to ALARM, but the notification was not sent. The SNS topic has a subscription to an email endpoint. What is the most likely cause?

A.The alarm's evaluation period is too short.
B.The email subscription to the SNS topic has not been confirmed.
C.The alarm's actions are not configured to send to the SNS topic.
D.The SNS topic is encrypted with a KMS key that the alarm does not have permissions to use.
AnswerB

For SNS email subscriptions, the subscriber must confirm the subscription by clicking the link in the initial confirmation message that Amazon SNS sends to the email address. Until that confirmation is completed, the subscription remains in the 'Pending Confirmation' state and Amazon SNS will not deliver messages published to the topic, even if the CloudWatch alarm successfully executes its action. This is a common and often overlooked cause of missing email notifications, and it directly explains the symptom described in the question.

Why this answer

The most likely cause is that the email subscription to the SNS topic has not been confirmed. When an SNS topic has an email subscription, AWS sends a confirmation email to the endpoint, and the subscriber must click the confirmation link before notifications can be delivered. Until the subscription is confirmed, the SNS topic will not send any messages to that endpoint, even if the CloudWatch alarm enters the ALARM state and successfully publishes to the topic.

Exam trap

The trap here is that candidates assume configuring the alarm to publish to an SNS topic is sufficient, overlooking the mandatory subscription confirmation step for email endpoints, which is a distinct and separate requirement.

How to eliminate wrong answers

Option A is wrong because the evaluation period affects how many consecutive data points must breach the threshold before the alarm state changes, but the alarm did change to ALARM, so the evaluation period is not the issue. Option C is wrong because if the alarm's actions were not configured to send to the SNS topic, the alarm would not have published to the topic at all, but the question states the alarm state changed to ALARM, implying the action was configured; the problem is on the subscription side. Option D is wrong because if the SNS topic were encrypted with a KMS key that the alarm did not have permissions to use, the alarm would fail to publish to the topic, and the alarm state would not change to ALARM; the alarm successfully published, so KMS permissions are not the issue.

← PreviousPage 4 of 4 · 250 questions total

Ready to test yourself?

Try a timed practice session using only Monitoring, Logging, and Remediation questions.