Courseiva

CCNA Monitoring, Logging, and Remediation Questions

75 of 250 questions · Page 3/4 · Monitoring, Logging, and Remediation · Answers revealed

151
MCQeasy

Refer to the exhibit. The command returns no datapoints for CPUUtilization for the specified instance. What is the most likely reason?

A.The instance was stopped or did not emit metrics during the specified time range.
B.The metric name is incorrect.
C.The instance does not have detailed monitoring enabled.
D.The period of 300 seconds is too short.
AnswerA

CloudWatch retains CPUUtilization only while the instance runs; a stopped instance emits no datapoints, so the queried period returns an empty result set. The metric namespace and dimensions remain valid, making the instance's stopped state the reason no datapoints appear for the specified range.

Why this answer

The most likely reason for no datapoints is that the instance was stopped or did not emit metrics during the specified time range. CloudWatch only retains and returns metric data when the instance is running and the CloudWatch agent or EC2 hypervisor is actively publishing CPUUtilization. If the instance was in a stopped state, no metrics are generated, resulting in an empty response from the GetMetricStatistics API call.

Exam trap

The trap here is that candidates often assume missing datapoints are due to a configuration issue (like detailed monitoring not enabled or wrong period), when in fact the instance simply wasn't running during the queried time window.

How to eliminate wrong answers

Option B is wrong because CPUUtilization is a standard EC2 metric name; if the metric name were incorrect, the API would return an error message (e.g., 'InvalidParameterValue') rather than an empty dataset. Option C is wrong because basic monitoring (5-minute granularity) still emits CPUUtilization datapoints; detailed monitoring only affects the frequency (1-minute granularity), not the existence of data. Option D is wrong because a period of 300 seconds is a valid and common value for basic monitoring (matching the default 5-minute interval) and does not cause missing datapoints.

152
MCQmedium

A SysOps administrator is troubleshooting an application that runs on AWS Lambda. The application occasionally fails with timeout errors. The administrator needs to identify the exact lines of code that are causing the delays. Which AWS service or feature should be used to gather this information?

A.Enable detailed CloudWatch Logs and search for 'timeout' strings.
B.Use AWS X-Ray to trace the Lambda function and view segment details.
C.Set a CloudWatch Metric Filter for 'Duration' and create an alarm.
D.Enable AWS CloudTrail data events for the Lambda function.
AnswerB

AWS X-Ray is the correct choice because it provides end-to-end distributed tracing for Lambda invocations, capturing a segment for the entire execution and subsegments for each downstream operation, such as DynamoDB queries, HTTP calls, or custom code blocks. When the X-Ray SDK is installed and the function is instrumented with wrappers or decorators, you can view segment details to see the exact duration of each subsegment, pinpointing which line or API call is slow. Even without custom instrumentation, X-Ray shows the overall execution time and any downstream service traces, but adding subsegments yields the precise line-level breakdown needed for this troubleshooting.

Why this answer

AWS X-Ray provides end-to-end tracing for Lambda functions, capturing segment details and subsegments that pinpoint the exact lines of code causing delays. By analyzing the trace timeline and annotations, the administrator can identify which specific function calls or operations exceed the timeout threshold, unlike CloudWatch Logs which only show aggregate duration or error strings without code-level granularity.

Exam trap

The trap here is that candidates often confuse CloudWatch Logs or Metrics (which show aggregate data) with the code-level tracing capability of X-Ray, assuming that searching for 'timeout' strings or monitoring 'Duration' metrics will reveal the exact lines of code causing the delay.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs can show 'timeout' strings but cannot trace the exact lines of code causing delays; they only log output from the function, not internal execution flow. Option C is wrong because a CloudWatch Metric Filter for 'Duration' and an alarm only monitors the overall execution time, not the specific code segments or lines responsible for the timeout. Option D is wrong because AWS CloudTrail data events record API calls to the Lambda service (e.g., Invoke, UpdateFunctionCode) but do not capture the internal execution trace or code-level timing within the function.

153
MCQmedium

A company needs to continuously scan Amazon EC2 instances for software vulnerabilities and unintended network exposure. Which AWS service should be used?

A.AWS Config
B.Amazon Inspector
C.AWS Trusted Advisor
D.Amazon GuardDuty
AnswerB

Amazon Inspector is purpose-built for vulnerability management, using an agent installed on EC2 instances to continuously collect telemetry about software packages, network configurations, and running processes. It correlates this data against a database of known Common Vulnerabilities and Exposures (CVEs) and performs network reachability checks to identify exposure of ports and protocols. This provides ongoing, deep scanning of software vulnerabilities and makes Inspector the correct choice.

Why this answer

Amazon Inspector is the correct service because it is specifically designed to automatically scan Amazon EC2 instances for software vulnerabilities (CVEs) and unintended network exposure (network reachability). It uses a combination of a managed agent (for OS-level assessment) and network configuration analysis to produce a detailed findings report, directly meeting the requirement for continuous scanning.

Exam trap

The trap here is that candidates often confuse Amazon Inspector with Amazon GuardDuty, mistakenly thinking GuardDuty performs vulnerability scanning when it actually focuses on threat detection from network and account activity, not on scanning EC2 instances for CVEs or network exposure.

How to eliminate wrong answers

Option A is wrong because AWS Config is a service for evaluating and auditing the configuration of AWS resources against desired policies (e.g., compliance rules), not for scanning for software vulnerabilities or network exposure. Option C is wrong because AWS Trusted Advisor provides high-level best-practice recommendations across cost, performance, security, and fault tolerance, but it does not perform deep vulnerability scanning of EC2 instances. Option D is wrong because Amazon GuardDuty is a threat detection service that analyzes VPC Flow Logs, DNS logs, and CloudTrail events for malicious activity, not for scanning EC2 instances for software vulnerabilities or unintended network exposure.

154
MCQeasy

A company has an AWS Lambda function that processes files uploaded to an S3 bucket. The function fails intermittently with a timeout error. What should the SysOps administrator do to monitor and resolve this issue?

A.Place the S3 bucket in a different AWS region to reduce latency.
B.Enable provisioned concurrency for the Lambda function.
C.Increase the Lambda function timeout and review the function logs in CloudWatch Logs.
D.Configure EC2 Auto Scaling to launch more instances for the Lambda function.
AnswerC

A Lambda timeout error indicates that the function exceeded its configured execution time limit, so increasing the timeout gives the process more time to complete its work. Reviewing CloudWatch Logs is essential to determine why the function is slow, such as inefficient logic, a large download, or a blocking call to another service. After increasing the timeout, verify both the function's memory allocation and the maximum allowed timeout (15 minutes) to ensure the new setting is valid, and monitor logs to confirm the root cause is resolved.

Why this answer

Increasing the Lambda function timeout directly addresses the intermittent timeout error by allowing the function more time to complete execution before Lambda terminates it. Reviewing CloudWatch Logs is essential to identify the root cause, such as slow downstream dependencies or inefficient code, enabling targeted optimization. This approach aligns with standard troubleshooting for Lambda timeout issues.

Exam trap

The trap here is that candidates may confuse Lambda's scaling behavior with EC2 Auto Scaling, or assume that provisioned concurrency fixes execution duration issues, when in fact it only addresses cold start latency.

How to eliminate wrong answers

Option A is wrong because moving the S3 bucket to a different region does not resolve Lambda timeout errors; it may increase latency due to cross-region data transfer and does not affect the function's execution duration. Option B is wrong because provisioned concurrency initializes execution environments to reduce cold starts, but it does not extend the maximum execution time allowed by the Lambda timeout setting. Option D is wrong because EC2 Auto Scaling manages EC2 instances, not Lambda functions; Lambda scales automatically based on incoming requests, and Auto Scaling has no role in Lambda execution.

155
Multi-Selecthard

A SysOps administrator is troubleshooting an issue where an EC2 instance has failed a status check. The instance is still running but is unresponsive. Which THREE actions should the administrator take to diagnose and resolve the issue? (Choose THREE.)

Select 3 answers
A.Reboot the instance.
B.Check the system status checks in the EC2 console.
C.Review the instance system log (console output).
D.Restore the instance from the latest AMI.
E.Stop and start the instance (recovery action).
AnswersB, C, E

Checking the system status checks in the EC2 console is the correct first triage step because these checks directly report on the health of the underlying physical host, network connectivity, and power delivery. A failed system status check (e.g., 'System reachability' or 'Instance connectivity') indicates a hardware-level issue that is independent of the guest OS, guiding you toward a stop/start recovery rather than a reboot. CloudWatch metrics for status checks should also be reviewed to confirm the duration and pattern of the failure.

Why this answer

System status checks (Option B) monitor the underlying physical host for issues like network or power loss, while instance status checks (like system log in Option C) detect OS-level problems. Reviewing the system log helps identify kernel panics or boot failures. Stopping and starting the instance (Option E) forces a migration to a new physical host, which can resolve host-level impairments without losing the instance's configuration or data.

Exam trap

The trap here is that candidates confuse 'reboot' with 'stop/start' — rebooting does not change the underlying host, while stopping and starting does, which is the key recovery action for host-level failures.

156
MCQmedium

A company has deployed a web application on EC2 instances with an Auto Scaling group. The SysOps administrator needs to automatically replace any instance that is in a 'failed' status as reported by the EC2 status checks. Which action should the administrator take?

A.Create an AWS Config rule to detect failed status checks and trigger a remediation action.
B.Create an AWS Lambda function that stops and starts the failed instance.
C.Configure the Auto Scaling group to use EC2 status checks for health checks.
D.Create a CloudWatch alarm on the StatusCheckFailed metric and trigger an SNS notification.
AnswerC

Configuring the Auto Scaling group to use EC2 status checks for health checks is the correct native mechanism because Auto Scaling continuously monitors the StatusCheckFailed metrics (system status and instance status) for each instance in the group. When a status check fails, the ASG marks the instance as unhealthy, terminates it, and launches a replacement instance to maintain the desired capacity—all without custom code or manual intervention. This approach leverages Amazon's built-in health check integration and is the intended method for automatically replacing instances that have failed EC2 status checks, ensuring the web application remains available.

Why this answer

An Auto Scaling group can be configured to use EC2 status checks (both system and instance) as the health check type. When the status check reports a failed status, the Auto Scaling group automatically terminates the unhealthy instance and launches a new one to replace it, ensuring self-healing without manual intervention.

Exam trap

The trap here is that candidates often confuse monitoring (CloudWatch alarms or SNS notifications) with automated remediation, forgetting that Auto Scaling groups have a built-in health check replacement feature that directly addresses the requirement to automatically replace failed instances.

How to eliminate wrong answers

Option A is wrong because AWS Config rules are designed for compliance and resource configuration auditing, not for real-time health monitoring or automatic replacement of failed instances; they lack the native ability to trigger instance replacement in an Auto Scaling group. Option B is wrong because stopping and starting a failed instance does not replace it; the instance retains its private IP and may still be unhealthy, and this approach bypasses the Auto Scaling group's lifecycle management, potentially causing state inconsistencies. Option D is wrong because a CloudWatch alarm on StatusCheckFailed with an SNS notification only sends an alert; it does not automatically replace the instance, requiring manual or additional automation to trigger the replacement.

157
MCQhard

A company uses Amazon CloudWatch Logs to collect logs from multiple EC2 instances. The SysOps administrator needs to create a metric filter that counts the number of ERROR-level log entries per hour and triggers an alarm when the count exceeds 100 in any 5-minute period. Which metric filter pattern should be used?

A.Use the pattern "ERROR" and set the metric value to 100.
B.Use the pattern "ERROR" and set the metric value to 1.
C.Use the pattern "ERROR *" to match any log entry starting with ERROR.
D.Use the pattern "[ERROR, 5]" to match 5 consecutive ERROR entries.
AnswerB

This is correct because CloudWatch Logs metric filters increment the target metric by the specified metric value each time a log event matches the given pattern. The pattern 'ERROR' is a simple, case-sensitive term that CloudWatch matches against the entire log event, not just the beginning, so it will catch every log line that contains the substring ERROR. Setting the metric value to 1 ensures that each matching event adds exactly 1, giving you an accurate real-time count of ERROR-level log entries.

Why this answer

A CloudWatch Logs metric filter counts each log event that matches the pattern. Setting the metric value to 1 ensures that each matching log entry increments the metric by 1, allowing the alarm to evaluate the sum over a 5-minute period against the threshold of 100. The pattern "ERROR" matches any log entry containing the string "ERROR" anywhere in the message.

Exam trap

The trap here is that candidates often think the metric value should match the alarm threshold (e.g., 100) or that wildcards or special syntax are needed, when in fact the metric value should be 1 and the threshold is set in the alarm definition.

How to eliminate wrong answers

Option A is wrong because setting the metric value to 100 would cause each matching log entry to increment the metric by 100, making the alarm trigger after a single ERROR entry (100/100 = 1), not after 100 entries. Option C is wrong because the pattern "ERROR *" uses a wildcard that is not valid in CloudWatch Logs metric filter syntax; CloudWatch Logs uses space-delimited token matching with brackets for positional patterns, not glob-style wildcards. Option D is wrong because the pattern "[ERROR, 5]" is not valid metric filter syntax; CloudWatch Logs does not support counting consecutive entries or specifying a count within the pattern itself.

158
MCQmedium

A company wants to receive alerts when an Auto Scaling group launches or terminates instances. They already have a CloudTrail trail enabled. What is the simplest way to achieve this?

A.Create a CloudWatch Events rule that matches Auto Scaling event patterns and sends notifications to an SNS topic.
B.Write a script on each EC2 instance to call the CloudWatch Logs API on launch/termination.
C.Configure the Auto Scaling group to publish lifecycle hooks and use Lambda to send notifications.
D.Enable CloudWatch detailed monitoring on the Auto Scaling group and create alarms.
AnswerA

A CloudWatch Events rule (EventBridge) can filter Auto Scaling group state-change events, such as EC2 Instance Launch Successful and EC2 Instance Terminate Successful, which are published automatically by the ASG service. The rule targets an SNS topic, delivering near real-time notifications to subscribed endpoints like email or SMS. This is the simplest, fully managed approach because it requires no instances, scripts, or custom code, and leverages an existing, reliable event stream.

Why this answer

CloudWatch Events (now part of Amazon EventBridge) can automatically capture Auto Scaling group state changes (launch and terminate) via CloudTrail API calls. By creating a rule that matches the specific event pattern for Auto Scaling events (e.g., 'EC2 Instance Launch Successful' and 'EC2 Instance Terminate Successful'), you can directly route those events to an SNS topic, which then sends notifications (e.g., email or SMS). This requires no custom code, lifecycle hooks, or additional monitoring configuration, making it the simplest solution.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing lifecycle hooks (Option C) or custom scripts (Option B), not realizing that CloudWatch Events can directly consume CloudTrail API events for Auto Scaling without additional infrastructure.

How to eliminate wrong answers

Option B is wrong because writing a script on each EC2 instance to call the CloudWatch Logs API on launch/termination is overly complex, requires custom code on every instance, and does not inherently trigger alerts; it only logs data, not send notifications. Option C is wrong because lifecycle hooks are designed for custom actions during scaling events (e.g., running a script before termination), but they add unnecessary complexity and cost (Lambda execution) when the goal is simply to receive alerts; CloudWatch Events can do this without hooks. Option D is wrong because CloudWatch detailed monitoring on an Auto Scaling group provides metrics like CPU utilization, not instance launch/termination events; alarms based on these metrics cannot detect scaling actions directly, and they would require threshold-based logic that is indirect and unreliable for this use case.

159
MCQeasy

A company wants to receive a notification when the root account is used to perform any action in the AWS account. Which service should be used to monitor this?

A.AWS Config
B.Amazon CloudWatch Logs
C.AWS CloudTrail with Amazon CloudWatch Events
D.AWS Trusted Advisor
AnswerC

AWS CloudTrail records every API call made on the account, including root user sign-ins and root-access-key usage, and these are delivered as event history. Amazon CloudWatch Events (now Amazon EventBridge) can monitor CloudTrail events in real time and use a rule to match a pattern such as a ConsoleLogin event with a root user identity. When matched, the rule can invoke an SNS topic to send the desired notification, making this the correct combination for real-time root activity alerts.

Why this answer

AWS CloudTrail logs all API activity in the account, including root user actions. By creating a CloudTrail trail and sending those logs to Amazon CloudWatch Events (now part of Amazon EventBridge), you can define a rule that matches the `RootAccess` event or any API call where `userIdentity.type` is `Root` and trigger a notification via SNS or Lambda. This combination provides real-time monitoring and alerting for root account usage.

Exam trap

The trap here is that candidates confuse AWS Config (which audits resource configurations) with CloudTrail (which audits API activity), or they think CloudWatch Logs alone can trigger alerts without CloudWatch Events, missing the requirement for event-driven notification.

How to eliminate wrong answers

Option A is wrong because AWS Config is designed for resource inventory, configuration history, and compliance rules (e.g., checking if S3 buckets are public), not for monitoring real-time API calls or user actions. Option B is wrong because Amazon CloudWatch Logs can store and query log data, but it cannot directly generate notifications from CloudTrail events without a CloudWatch Events rule or metric filter; the question requires a service that monitors root actions, and CloudWatch Logs alone lacks the event-driven alerting capability. Option D is wrong because AWS Trusted Advisor provides best-practice checks for cost optimization, performance, security, and fault tolerance, but it does not monitor or alert on specific user actions like root account usage.

160
MCQmedium

Refer to the exhibit. An EC2 instance is running the CloudWatch Logs agent and uses the IAM policy shown. The agent is configured to send logs to the log group 'MyAppLogGroup'. However, logs are not appearing. What is the MOST likely cause?

A.The log group name in the policy does not match the agent configuration.
B.The policy does not allow the 'logs:PutLogEvents' action.
C.The policy is missing permission to create the log group if it does not exist.
D.The Resource ARN incorrectly specifies a wildcard after the log group name.
AnswerC

CloudWatch agent requires 'logs:CreateLogGroup' to create the specified log group when it does not already exist. The policy in question grants only 'logs:CreateLogStream' and 'logs:PutLogEvents', but not 'logs:CreateLogGroup', so the agent cannot provision the MyAppLogGroup on startup. As a result, the agent fails to deliver any log data, making the missing CreateLogGroup permission the correct explanation for the issue.

Why this answer

The CloudWatch Logs agent cannot automatically create a log group; it requires explicit permission via the `logs:CreateLogGroup` action in the IAM policy. Without this permission, if the log group 'MyAppLogGroup' does not already exist, the agent will fail to send logs, even though the `logs:PutLogEvents` action is allowed. The policy shown only grants `logs:PutLogEvents` and `logs:DescribeLogStreams`, missing the necessary `logs:CreateLogGroup` and `logs:CreateLogStream` actions for initial setup.

Exam trap

The trap here is that candidates assume `logs:PutLogEvents` alone is sufficient for sending logs, overlooking the fact that the agent must first create the log group and log stream if they do not exist, which requires additional permissions.

How to eliminate wrong answers

Option A is wrong because the log group name in the policy ('MyAppLogGroup') matches the agent configuration, so there is no mismatch. Option B is wrong because the policy explicitly includes the `logs:PutLogEvents` action, so that permission is present. Option D is wrong because the Resource ARN `arn:aws:logs:us-east-1:123456789012:log-group:MyAppLogGroup:*` correctly uses a wildcard after the log group name to match all log streams within that group, which is standard practice.

161
MCQmedium

A SysOps administrator needs to automatically restart an Amazon RDS DB instance when the 'DatabaseConnections' metric exceeds a threshold of 200 for 5 consecutive minutes. The administrator wants a solution that uses minimal custom code and leverages AWS managed services. Which combination of services should be used?

A.Amazon CloudWatch alarm with an Auto Scaling policy.
B.Amazon CloudWatch alarm with an Amazon Simple Notification Service (SNS) topic that triggers an AWS Lambda function to restart the instance.
C.Amazon CloudWatch alarm with an AWS Systems Manager Automation action.
D.Amazon RDS event subscription that triggers an AWS Lambda function.
AnswerC

CloudWatch alarms support a native 'Systems Manager Automation' action that directly invokes an SSM automation runbook, such as the pre-built AWS-RestartRDSInstance runbook, when the alarm enters an alarm state. This runbook encapsulates the RDS reboot API call and includes appropriate wait/verify steps, all without requiring you to write or maintain any custom code. The integration is purpose-built for metric-driven remediation and is the minimal-effort, fully managed way to automatically restart an RDS DB instance.

Why this answer

AWS Systems Manager Automation provides a built-in 'AWSSupport-StartRDSInstance' or 'AWSSystemsManager-RestartRDSInstance' runbook that can be triggered directly by a CloudWatch alarm action, requiring no custom code. This leverages a managed service to restart the RDS instance automatically when the 'DatabaseConnections' metric exceeds 200 for 5 consecutive minutes, meeting the minimal custom code requirement.

Exam trap

The trap here is that candidates often assume a Lambda function is always required for custom remediation actions, but AWS Systems Manager Automation provides a no-code alternative for many common operations like restarting RDS instances, which directly meets the 'minimal custom code' constraint.

How to eliminate wrong answers

Option A is wrong because Auto Scaling policies are designed to scale EC2 instances or other Auto Scaling group resources, not to restart RDS instances; they cannot directly trigger a database restart. Option B is wrong because while it uses a Lambda function to restart the instance, it introduces custom code (the Lambda function logic) which violates the 'minimal custom code' requirement. Option D is wrong because RDS event subscriptions are for notification of events like instance creation or failure, not for triggering automated remediation based on CloudWatch metric thresholds; they lack the direct integration with CloudWatch alarms needed for this metric-based condition.

162
MCQeasy

A company wants to receive an email notification when an EC2 instance's status check fails. What AWS service should be used?

A.AWS CloudTrail
B.Amazon CloudWatch Alarm
C.AWS Config
D.AWS Trusted Advisor
AnswerB

Amazon CloudWatch Alarm continuously monitors the StatusCheckFailed, StatusCheckFailed_System, and StatusCheckFailed_Instance metrics that EC2 automatically publishes at one-minute frequency. When one of these alarm metrics crosses a configured threshold, the alarm transitions to ALARM state and publishes a message to an Amazon SNS topic, which can be configured to send an email notification. This directly satisfies the requirement to receive an email when an EC2 instance's health changes.

Why this answer

Amazon CloudWatch Alarms can monitor EC2 instance status checks (both system and instance checks) and trigger an action, such as sending an email via Amazon SNS, when a status check fails. This is the correct service because it directly integrates with EC2 metrics and supports alarm-based notifications for status check failures.

Exam trap

The trap here is that candidates often confuse AWS Config or CloudTrail with monitoring services, but only CloudWatch Alarms can directly monitor EC2 status check metrics and trigger notifications.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail records API activity and management events, not instance-level health metrics or status checks; it cannot trigger email notifications for status check failures. Option C is wrong because AWS Config evaluates resource configurations against rules and tracks configuration changes, but it does not monitor real-time operational metrics like EC2 status checks or send direct email alerts. Option D is wrong because AWS Trusted Advisor provides best-practice recommendations and checks for cost optimization, security, and fault tolerance, but it does not monitor EC2 status checks or send notifications for status check failures.

163
Matchingmedium

Match each AWS service to its primary function.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Monitoring and observability

API activity logging

Resource compliance and configuration tracking

Best practice recommendations

Operational management and automation

Why these pairings

AWS CloudTrail records API calls for auditing; AWS Config evaluates resource configurations; AWS Trusted Advisor gives recommendations. Common confusions involve swapping CloudTrail and Config functions.

164
MCQhard

A company has an application running on Amazon EC2 instances behind an Application Load Balancer (ALB). The application logs show intermittent 503 errors. The ALB access logs show that the errors occur when the target response time exceeds 30 seconds. Which configuration change should the SysOps administrator make to reduce the number of 503 errors without affecting the application's behavior?

A.Increase the deregistration delay on the target group
B.Increase the idle timeout setting on the ALB
C.Enable cross-zone load balancing on the ALB
D.Decrease the health check interval on the target group
AnswerB

The ALB idle timeout is the maximum time the load balancer waits for a response from the target after forwarding a request. If the application takes longer than this timeout to send a response, the ALB closes the connection and the client receives an error. Increasing the idle timeout gives long-running requests more time to complete, so this is the correct fix.

Why this answer

The 503 errors occur when the target response time exceeds 30 seconds, which matches the default idle timeout of the Application Load Balancer. By increasing the idle timeout setting on the ALB to a value higher than the application's maximum expected response time (e.g., 60 or 120 seconds), the ALB will wait longer before closing the connection, preventing premature 503 errors while the backend is still processing the request.

Exam trap

The trap here is that candidates often confuse the ALB idle timeout with the target group deregistration delay or health check settings, mistakenly thinking that adjusting health checks or deregistration will fix timeout-related 503 errors, when the root cause is the ALB's connection timeout parameter.

How to eliminate wrong answers

Option A is wrong because increasing the deregistration delay on the target group controls how long the ALB waits before sending new connections to a deregistering target, which does not address the 30-second response timeout causing 503 errors. Option C is wrong because enabling cross-zone load balancing distributes traffic evenly across all targets in all Availability Zones, which improves load distribution but does not affect the ALB's idle timeout or prevent 503 errors from slow responses. Option D is wrong because decreasing the health check interval on the target group makes health checks more frequent, which can detect unhealthy targets faster but does not change the ALB's connection timeout behavior that is causing the 503 errors.

165
Multi-Selecteasy

Which TWO AWS services can be used to monitor the performance of an Amazon RDS database and set alarms based on metrics?

Select 2 answers
A.AWS Lambda
B.Amazon RDS Performance Insights
C.Amazon S3
D.Amazon CloudWatch
E.AWS CloudTrail
AnswersB, D

Amazon RDS Performance Insights is a database performance tuning and monitoring feature that gives you a visual dashboard of the database load, wait states, and top SQL statements. It helps you quickly identify bottlenecks such as CPU, memory, or lock contention and can publish metrics to Amazon CloudWatch for threshold-based alarms. This makes it a purpose-built service for monitoring the performance of an RDS instance, so it is a correct answer.

Why this answer

Amazon CloudWatch (Option D) is the primary monitoring service for AWS resources, including Amazon RDS. It collects metrics like CPU utilization, database connections, and read/write latency, and allows you to set CloudWatch Alarms that trigger actions (e.g., SNS notifications) when thresholds are breached. Amazon RDS Performance Insights (Option B) provides deeper database performance analysis by visualizing database load and identifying bottlenecks, and it can also publish metrics to CloudWatch for alarm purposes.

Exam trap

The trap here is that candidates often confuse AWS CloudTrail (audit logging) with CloudWatch (monitoring), or think Lambda can be used for monitoring when it is actually a compute trigger, not a monitoring service.

166
Multi-Selectmedium

A company uses Amazon CloudWatch to monitor its AWS infrastructure. The operations team wants to receive notifications when any EC2 instance's status check fails. Which TWO steps should be taken to achieve this?

Select 2 answers
A.Create a CloudWatch alarm on the CPUUtilization metric with a threshold of 90%.
B.Configure the alarm to send a notification to an Amazon SNS topic.
C.Enable AWS CloudTrail to capture instance state changes.
D.Install the CloudWatch agent on the instance and publish custom memory metrics.
E.Create a CloudWatch alarm on the StatusCheckFailed (or StatusCheckFailed_Instance) metric with a threshold of greater than 0.
AnswersB, E

An Amazon SNS topic is the correct notification mechanism when paired with a CloudWatch alarm. When you configure an alarm, you can set an action to publish a message to an SNS topic, which then delivers alerts through email, SMS, or other subscribed endpoints. This is the essential step that ensures administrators are actually notified when an instance fails its status checks. Without an SNS action, the alarm would only remain in the ALARM state silently.

Why this answer

Amazon CloudWatch alarms can be configured to send notifications to an Amazon SNS topic when the alarm state changes. This allows the operations team to receive immediate notifications (e.g., via email, SMS, or HTTP) when an EC2 instance's status check fails. Option E is also correct because the StatusCheckFailed metric (or StatusCheckFailed_Instance) directly reflects the result of the EC2 status check; setting an alarm with a threshold of greater than 0 triggers when any status check fails.

Exam trap

The trap here is that candidates may confuse CloudWatch metrics like CPUUtilization or custom metrics with the built-in StatusCheckFailed metric, or think that CloudTrail can be used for real-time monitoring and alerting, when in fact it is designed for auditing API activity, not for instance health checks.

167
MCQmedium

Refer to the exhibit. A SysOps administrator runs this CloudWatch Logs Insights query against an application log group. The query returns no results, even though the administrator knows that errors occurred in the last hour. What is the most likely cause?

A.The 'stats' command requires a 'by' clause with a field name, but 'bin(5m)' is invalid.
B.The log group contains too many log events, causing the query to time out.
C.The @message field is not a valid field in CloudWatch Logs Insights.
D.The log group's retention policy is set to 1 day and the data is older than the retention period.
AnswerB

CloudWatch Logs Insights queries have a maximum execution time of 60 seconds, and when a log group contains a very high volume of log events within the queried time range, the query engine may exceed that limit. A timeout causes the query to return no results, even though the data exists. This matches the symptom described in the question, making it the correct explanation.

Why this answer

In CloudWatch Logs Insights, `stats count() by bin(5m)` is valid syntax; `bin()` does not require a preceding field. Therefore option A is not the cause. The most likely cause among the options is that the log group contains too many log events, causing the query to time out before results are returned.

Retention policy does not affect events from the last hour, and `@message` is a valid field.

Exam trap

Candidates often assume a query returning no results is due to a syntax error. However, CloudWatch Logs Insights queries can time out on very large log groups, and a timeout may produce no results; `bin(5m)` is valid syntax.

How to eliminate wrong answers

Option A is wrong because the 'stats' command in CloudWatch Logs Insights does not require a 'by' clause; 'bin(5m)' is a valid function that groups timestamps into 5-minute intervals, so the syntax is correct. Option B is wrong because CloudWatch Logs Insights queries have a 10,000-event limit per query, but they do not time out due to too many log events; instead, they return partial results or a message indicating the limit was reached. Option C is wrong because @message is a reserved field in CloudWatch Logs Insights that contains the raw log event text, and it is always available for querying.

168
MCQmedium

An application logs user authentication attempts to Amazon CloudWatch Logs. The SysOps administrator needs to create a custom metric that counts the number of failed authentication attempts every 5 minutes and trigger an alarm when the count exceeds 5. Which combination of actions should the administrator take?

A.Use the PutMetricData API in the application to publish the number of failed attempts as a custom metric, then create an alarm.
B.Create a metric filter on the log group for the string 'FAILED_AUTH', set the metric value to 1, then create an alarm on the resulting metric.
C.Use AWS CloudTrail to track authentication events and create a metric filter on the CloudTrail log group.
D.Use Amazon Athena to query the logs every 5 minutes and publish results to a CloudWatch metric.
AnswerB

A CloudWatch Logs metric filter applies a pattern match to incoming log events, such as the string 'FAILED_AUTH', and increments a specified metric value (e.g., 1) for every matching event. The resulting metric is automatically published to CloudWatch, and an alarm can be configured to trigger when the count exceeds a threshold over a period, such as the sum in 5 minutes. Because this operates directly on the existing log stream, it requires no application changes and provides near-real-time monitoring.

Why this answer

CloudWatch Logs metric filters allow you to extract a numeric value from log events and publish it as a custom metric. By creating a filter that matches the string 'FAILED_AUTH' and setting the metric value to 1, each failed attempt increments the metric. You can then set the metric's period to 5 minutes and create an alarm that triggers when the sum exceeds 5, meeting the requirement without modifying the application code.

Exam trap

The trap here is that candidates often confuse CloudTrail (which logs AWS API calls) with application-level logging, leading them to choose Option C, or they assume the application must be modified to publish metrics (Option A), missing the serverless metric filter approach.

How to eliminate wrong answers

Option A is wrong because it requires modifying the application to call the PutMetricData API, which adds complexity and couples the application to CloudWatch, whereas the requirement can be met without application changes using a metric filter. Option C is wrong because CloudTrail logs management events, not application-level authentication logs; it tracks API calls to AWS services, not user authentication attempts within an application. Option D is wrong because Amazon Athena is an interactive query service for analyzing data in S3, not a real-time or scheduled metric publisher; it cannot automatically publish results to CloudWatch every 5 minutes without custom orchestration, and it introduces unnecessary latency and cost.

169
MCQmedium

A SysOps administrator needs to analyze application logs stored in Amazon CloudWatch Logs to find specific error patterns across multiple log groups. The administrator wants to run queries to filter and parse the logs. Which feature should the administrator use?

A.CloudWatch Logs subscriptions
B.CloudWatch Logs Insights
C.CloudWatch Metric Filters
D.CloudWatch Contributor Insights
AnswerB

CloudWatch Logs Insights is the correct choice because it is a dedicated, fully managed query engine for log data stored in CloudWatch Logs. It uses a SQL-like query language (fields, filter, stats, sort, etc.) to run interactive, ad-hoc searches across one or more log groups, allowing you to discover error patterns, aggregate results, and visualize findings in the console. Unlike the other options, Logs Insights is specifically designed to answer arbitrary questions on historical log data without requiring any external pipeline.

Why this answer

CloudWatch Logs Insights is the correct feature because it enables you to interactively search and analyze log data stored in CloudWatch Logs using a purpose-built query language. It allows you to run queries across multiple log groups, filter, parse, and aggregate logs to identify specific error patterns, making it ideal for ad-hoc log analysis and troubleshooting.

Exam trap

The trap here is that candidates often confuse CloudWatch Metric Filters (which can filter logs for metric extraction) with the interactive querying capability of CloudWatch Logs Insights, but Metric Filters cannot parse or analyze log content across multiple log groups in a query-like manner.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs subscriptions are used to stream log data in real-time to other services like Lambda, Kinesis, or Elasticsearch for processing or storage, not for running interactive queries to filter and parse logs. Option C is wrong because CloudWatch Metric Filters extract metric data from logs to create CloudWatch metrics and trigger alarms, but they do not support running queries to parse and analyze log content for error patterns across multiple log groups. Option D is wrong because CloudWatch Contributor Insights analyzes time-series data to identify top contributors and understand traffic patterns, but it is not designed for querying and parsing raw log messages to find specific error patterns.

170
MCQeasy

A company needs to retain API call logs for 7 years for compliance. Which AWS service should be used to store these logs?

A.AWS CloudTrail
B.AWS Config
C.Amazon CloudWatch Logs
D.Amazon VPC Flow Logs
AnswerA

CloudTrail records API activity and delivers event logs to Amazon S3, where lifecycle policies retain objects for the mandated seven years. This satisfies the compliance retention constraint, since CloudTrail itself is the capture service feeding durable storage.

Why this answer

AWS CloudTrail is the correct service because it records API activity across your AWS infrastructure and can be configured to store logs in an S3 bucket with lifecycle policies that retain data for 7 years. CloudTrail is specifically designed for auditing and compliance, capturing management and data plane API calls, and supports long-term retention via S3 object locking or lifecycle rules.

Exam trap

The trap here is that candidates confuse CloudTrail (API auditing) with CloudWatch Logs (operational logs) or Config (configuration history), but only CloudTrail captures the specific API call logs required for compliance retention.

How to eliminate wrong answers

Option B (AWS Config) is wrong because it tracks resource configuration changes and compliance over time, not API call logs; it stores configuration history, not API activity. Option C (Amazon CloudWatch Logs) is wrong because it is intended for real-time monitoring and operational logging from applications and services, with a default retention of indefinite but not optimized for 7-year compliance archiving; it lacks native long-term retention controls like S3 lifecycle policies. Option D (Amazon VPC Flow Logs) is wrong because it captures IP traffic metadata (source/destination IPs, ports, protocols) for network analysis, not API call logs; it is not designed for auditing API-level actions.

171
Multi-Selectmedium

A company wants to receive real-time notifications when specific API calls are made in their AWS account. Which TWO services can be used together to achieve this? (Choose TWO.)

Select 2 answers
A.AWS CloudTrail
B.AWS Config
C.AWS Lambda
D.Amazon CloudWatch Logs
E.Amazon CloudWatch Events
AnswersA, E

AWS CloudTrail records API calls and delivers log files to an S3 bucket or CloudWatch Logs, but it does not send proactive notifications in real time. It is the authoritative source for who made which API calls, yet it requires a separate consumer—such as EventBridge, CloudWatch Logs, or a custom script—to parse those logs and trigger alerts. Thus CloudTrail provides the data for auditing but lacks the built-in event-matching and routing capabilities needed for immediate notification.

Why this answer

AWS CloudTrail captures API calls and can send events directly to Amazon CloudWatch Events (now Amazon EventBridge). CloudWatch Events rules can match specific API call patterns and trigger real-time notifications via targets like SNS or Lambda. No CloudWatch Logs involvement is required for this integration, so the solution uses exactly two services: CloudTrail and CloudWatch Events.

Exam trap

The trap here is that candidates often pick AWS Lambda alone, thinking it can directly monitor API calls, but Lambda requires a triggering event source such as CloudWatch Events or S3, not raw CloudTrail logs.

172
MCQmedium

A SysOps administrator is troubleshooting an issue where an Amazon EC2 instance running Amazon Linux 2 is not sending logs to CloudWatch Logs. The CloudWatch agent is installed and configured. Which step should the administrator take FIRST to diagnose the issue?

A.Reinstall the CloudWatch agent from the AWS Systems Manager.
B.Update the IAM role attached to the instance to include CloudWatch Logs permissions.
C.Verify the EC2 instance status in the AWS Management Console.
D.Check the CloudWatch agent status and review its log files on the instance.
AnswerD

Checking the CloudWatch agent status and reviewing its log files is the correct first step because it provides direct, actionable diagnostic information about the agent's runtime and integration with CloudWatch. Use commands like 'systemctl status amazon-cloudwatch-agent' or '/opt/aws/amazon-cloudwatch-agent/bin/amazon-cloudwatch-agent-ctl -a status' to confirm the agent is running, then inspect '/opt/aws/amazon-cloudwatch-agent/logs/amazon-cloudwatch-agent.log' and 'agent.log' entries for errors such as 'Error occurred during config', 'dial tcp: i/o timeout', or 'ResourceNotFoundException'. These logs reveal whether the problem is a malformed configuration file, missing IAM permissions, a network/firewall block, or an issue with the log file path itself, enabling a targeted resolution instead of trial-and-error.

Why this answer

The CloudWatch agent is already installed and configured, so the first step is to check its operational status and review its own log files (typically in /var/log/aws/amazon-cloudwatch-agent/). This directly reveals whether the agent is running, encountering configuration errors, or failing to connect to the CloudWatch Logs service endpoint, without making unnecessary changes.

Exam trap

The trap here is that candidates assume the IAM role is the most common cause of log delivery failure and jump to updating it, but the question specifies the agent is already installed and configured, so the first logical step is to check the agent's own logs to confirm the exact error before making any changes.

How to eliminate wrong answers

Option A is wrong because reinstalling the agent from Systems Manager is a disruptive step that should only be taken after verifying the agent is not functioning due to corruption or misconfiguration, not as a first diagnostic step. Option B is wrong because updating the IAM role is a valid fix if permissions are missing, but the question states the agent is installed and configured; checking the agent's logs first will confirm whether the issue is permission-related (e.g., an AccessDeniedException) before modifying the role. Option C is wrong because verifying the EC2 instance status in the console only confirms the instance is running and reachable, but does not provide any insight into the CloudWatch agent's internal state or log delivery failures.

173
MCQmedium

A company runs a web application on EC2 instances behind an Application Load Balancer. The SysOps administrator creates a CloudWatch alarm on the ALB's HTTPCode_ELB_5XX_Count metric to trigger an SNS notification when there are many 5xx errors. However, the alarm remains in INSUFFICIENT_DATA state. What is a likely cause?

A.The HTTPCode_ELB_5XX_Count metric is not available for ALBs.
B.The ALB and the CloudWatch alarm are in different AWS Regions.
C.The IAM role for CloudWatch does not have permission to read the ALB metrics.
D.The alarm's period is set to 1 minute, but the metric is published every 5 minutes.
AnswerB

CloudWatch alarms can only evaluate metrics that exist in the exact same AWS Region as the alarm. Standard ALB metrics are published into the Region where the Application Load Balancer is deployed; an alarm created in a different Region cannot see or reference those metric data points. Cross-region metric aggregation is not a native feature for ALB CloudWatch metrics, so placing the alarm in a different Region from the ALB would indeed prevent the alarm from ever receiving data and leave it in INSUFFICIENT_DATA.

Why this answer

CloudWatch alarms can only evaluate metrics from the same AWS Region in which the alarm is created. If the ALB and the CloudWatch alarm are in different Regions, the alarm will never receive metric data points, causing it to remain in INSUFFICIENT_DATA state. This is a common cross-Region limitation for CloudWatch metrics.

Exam trap

The trap here is that candidates often overlook the regional scope of CloudWatch metrics and alarms, assuming that metrics are globally accessible, when in fact they are strictly regional.

How to eliminate wrong answers

Option A is wrong because HTTPCode_ELB_5XX_Count is a standard metric emitted by Application Load Balancers and is fully available in CloudWatch. Option C is wrong because CloudWatch does not require a separate IAM role to read ALB metrics; metrics are published automatically by AWS services, and the alarm uses the service-linked role or the caller's permissions, which are not the cause of INSUFFICIENT_DATA. Option D is wrong because even if the alarm's period is shorter than the metric's publication interval, CloudWatch will still receive data points (just less frequently) and the alarm would eventually show OK or ALARM, not remain in INSUFFICIENT_DATA indefinitely.

174
MCQhard

A SysOps administrator is investigating a security incident where an EC2 instance was used to launch an attack. The administrator needs to determine the source IP addresses that were used to access the instance prior to the attack. Which AWS service and feature should be used to capture this information?

A.Amazon GuardDuty
B.AWS CloudTrail with data events enabled for EC2
C.VPC Flow Logs
D.AWS Config with recording of security groups
AnswerC

VPC Flow Logs capture metadata about IP traffic flowing to and from network interfaces in a VPC, including source and destination IP addresses, ports, protocol, and whether the traffic was accepted or rejected. These logs can be published to Amazon CloudWatch Logs or Amazon S3, and they are the standard AWS service for retracing actual network activity during a security investigation. Unlike API logs or configuration records, VPC Flow Logs directly answer the need to identify which source IPs communicated with your resources.

Why this answer

VPC Flow Logs capture metadata about IP traffic flowing to and from network interfaces in a VPC, including the source IP address, destination IP address, port, protocol, and timestamps. This allows the administrator to identify the source IP addresses that accessed the EC2 instance prior to the attack, as the logs record all accepted and rejected traffic at the network interface level.

Exam trap

The trap here is that candidates confuse AWS CloudTrail (which logs API calls) with network traffic logging, mistakenly thinking CloudTrail captures IP-level traffic data, when in fact only VPC Flow Logs record the actual source IP addresses of network connections to an EC2 instance.

How to eliminate wrong answers

Option A is wrong because Amazon GuardDuty is a threat detection service that analyzes logs (like VPC Flow Logs, DNS logs, and CloudTrail events) to identify malicious activity, but it does not natively store or provide raw source IP address logs for historical analysis of traffic to an instance. Option B is wrong because AWS CloudTrail with data events for EC2 records API calls (e.g., RunInstances, DescribeInstances) and does not capture network-level traffic metadata such as source IP addresses of packets reaching the instance. Option D is wrong because AWS Config records configuration changes to resources like security groups, but it does not log network traffic flows or the source IP addresses of connections to an instance.

175
MCQeasy

A SysOps administrator wants to receive a notification when an EC2 instance's status check fails. Which AWS service should the administrator use to set up an alarm based on the status check metric?

A.Amazon CloudWatch Alarms
B.Amazon EventBridge
C.AWS Config
D.AWS Systems Manager
AnswerA

Amazon CloudWatch Alarms is the correct service because EC2 publishes the `StatusCheckFailed` metric (a composite of `StatusCheckFailed_System` and `StatusCheckFailed_Instance`) to CloudWatch every minute. You can create a metric alarm that watches this value and transitions to ALARM when it exceeds a threshold (e.g., > 0 for consecutive periods), then publishes to an SNS topic to trigger an email or SMS notification. This is purpose-built for metric-based threshold monitoring, unlike event-driven services.

Why this answer

Amazon CloudWatch Alarms is the correct service because EC2 instance status checks are published as CloudWatch metrics (e.g., StatusCheckFailed, StatusCheckFailed_Instance, StatusCheckFailed_System). A CloudWatch Alarm can be configured to monitor these metrics and trigger an action, such as sending a notification via Amazon SNS, when the alarm state changes to ALARM. This directly meets the requirement to receive a notification when a status check fails.

Exam trap

The trap here is that candidates may confuse EventBridge's ability to react to EC2 instance state changes (e.g., running, stopped) with the need to monitor a continuous metric like status check failures, which requires CloudWatch Alarms for threshold-based evaluation.

How to eliminate wrong answers

Option B (Amazon EventBridge) is wrong because EventBridge is a serverless event bus used to route events from AWS services or custom sources to targets like Lambda or SQS; it does not natively evaluate metric thresholds or trigger alarms based on status check metrics. Option C (AWS Config) is wrong because Config is used for resource inventory, configuration history, and compliance auditing, not for real-time monitoring of operational metrics like status checks. Option D (AWS Systems Manager) is wrong because Systems Manager provides operational management capabilities (e.g., patching, automation, Run Command) but does not offer metric-based alarm functionality for EC2 status checks.

176
MCQmedium

An application running on EC2 instances in an Auto Scaling group is experiencing intermittent errors. The errors correlate with periods of high memory usage. The SysOps administrator wants to set up a CloudWatch alarm to scale out when memory usage exceeds 80%. What should the administrator do to enable monitoring of memory usage?

A.Enable detailed monitoring on the Auto Scaling group.
B.Install the CloudWatch agent on the EC2 instances to publish memory metrics.
C.Use the default EC2 memory metric provided by CloudWatch.
D.Create a Lambda function to pull memory data from the EC2 instances.
AnswerB

The CloudWatch agent is the recommended mechanism for publishing EC2 guest-level metrics such as memory utilization, swap usage, and disk space to CloudWatch. Running on each instance, it collects counters from the OS and sends them under the CWAgent namespace, making them available for CloudWatch alarms and dashboards. To use it, install the agent on every instance in the Auto Scaling group, typically via user data or a bootstrap script, so that newly launched instances also begin reporting memory metrics.

Why this answer

The CloudWatch agent is required to publish custom memory metrics from EC2 instances because memory utilization is not a standard metric provided by default. By installing and configuring the agent, the administrator can collect memory usage data and create a CloudWatch alarm to trigger Auto Scaling actions when usage exceeds 80%.

Exam trap

The trap here is that candidates often assume memory metrics are automatically available in CloudWatch like CPU utilization, but AWS intentionally excludes guest-level metrics (memory, disk, swap) by default, requiring the CloudWatch agent to be installed and configured.

How to eliminate wrong answers

Option A is wrong because enabling detailed monitoring on the Auto Scaling group only increases the frequency of standard EC2 metrics (e.g., CPU, network) to 1-minute intervals, but does not add memory metrics. Option C is wrong because CloudWatch does not provide a default EC2 memory metric; memory is a guest-level metric that must be explicitly reported by an agent. Option D is wrong because while a Lambda function could theoretically pull memory data, it is an unnecessarily complex and indirect approach compared to the straightforward, supported method of using the CloudWatch agent, and it would require custom code and permissions management.

177
MCQhard

The monitoring team needs to collect per-process CPU and memory utilization for a specific Java process (named 'app.jar') running on EC2 Linux instances. Standard EC2 metrics show aggregate CPU but not per-process details. Which CloudWatch agent configuration section enables this?

A.Add a procstat section under metrics_collected in the CloudWatch agent config, specifying process_name = 'app.jar' to collect per-process CPU and memory
B.Enable enhanced monitoring on the EC2 instance and select 'per-process metrics' from the console
C.Configure a CloudWatch Logs metric filter on the Java GC log output to derive CPU and memory figures
D.Use the aws ec2 describe-instance-status API on a schedule to pull process metrics from the instance's system status checks
AnswerA

The procstat plugin uses the Linux /proc filesystem to sample per-process resource usage. With process_name set to 'app.jar', the agent matches the running JVM process and publishes metrics like procstat_cpu_usage and procstat_memory_rss to CloudWatch every collection interval. These metrics carry instance ID and process name dimensions.

Why this answer

The CloudWatch agent's `procstat` plugin is specifically designed to collect per-process metrics such as CPU and memory utilization. By adding a `procstat` section under `metrics_collected` in the agent configuration file and specifying the process name (e.g., `process_name = 'app.jar'`), the agent will gather the required per-process metrics and send them to CloudWatch. This is the only native method within the CloudWatch ecosystem to achieve per-process monitoring on EC2 Linux instances.

Exam trap

The trap here is that candidates may confuse 'enhanced monitoring' (a hypervisor-level feature) with OS-level per-process monitoring, or incorrectly assume that CloudWatch Logs metric filters can derive CPU/memory metrics from application logs, when in fact only the CloudWatch agent's `procstat` plugin can collect actual OS-level per-process resource utilization.

How to eliminate wrong answers

Option B is wrong because 'enhanced monitoring' is a feature of EC2 that provides additional hypervisor-level metrics (e.g., CPU credit usage, network throughput) but does not expose per-process metrics; it cannot see inside the guest OS. Option C is wrong because CloudWatch Logs metric filters can parse log patterns and create numerical metrics from log data, but Java GC logs do not contain CPU or memory utilization figures for the process; they only contain garbage collection timing and heap usage, not OS-level resource consumption. Option D is wrong because the `aws ec2 describe-instance-status` API returns instance health and status checks (e.g., system reachability, instance status) and has no capability to retrieve per-process metrics from within the instance.

178
MCQhard

A company has a VPC with public and private subnets. An EC2 instance in a private subnet needs to send logs to CloudWatch Logs. Which steps are necessary to allow this without traversing the internet? (Select TWO.)

A.Place the instance in a public subnet with a public IP.
B.Attach an IAM role to the instance with permissions to call PutLogEvents.
C.Create a VPC endpoint for CloudWatch Logs (com.amazonaws.region.logs).
D.Ensure the instance is in the default VPC.
E.Attach a NAT gateway to the private subnet's route table.
AnswerB, C

Attaching an IAM role with permissions to call PutLogEvents is the mandatory identity-layer step: the CloudWatch agent or SDK retrieves temporary credentials from the instance metadata service, and those credentials must include actions such as logs:CreateLogStream and logs:PutLogEvents. Without this role, every API request to CloudWatch Logs is rejected with an access-denied error, even if the network path is perfectly reachable. This role is required whether you reach CloudWatch Logs via a VPC endpoint, a NAT gateway, or the internet.

Why this answer

The EC2 instance must have an IAM role attached that includes permissions for `logs:PutLogEvents` to authenticate and authorize log delivery to CloudWatch Logs. Without this role, the instance cannot sign API requests, even if network connectivity exists. Option C is correct because a VPC endpoint for CloudWatch Logs (com.amazonaws.region.logs) provides private connectivity via AWS PrivateLink, allowing the instance to send logs without traversing the internet or a NAT gateway.

Exam trap

The trap here is that candidates often assume a NAT gateway is required for private subnet outbound traffic, but for AWS services like CloudWatch Logs, a VPC endpoint provides a more secure and direct path without internet traversal.

How to eliminate wrong answers

Option A is wrong because placing the instance in a public subnet with a public IP would expose it to the internet, which violates the requirement to avoid traversing the internet and introduces unnecessary security risks. Option D is wrong because being in the default VPC does not inherently provide private connectivity to CloudWatch Logs; the default VPC still requires either a NAT gateway or a VPC endpoint for outbound traffic to AWS services. Option E is wrong because a NAT gateway provides internet access for private subnets, but it forces traffic to traverse the internet, which contradicts the requirement to avoid internet traversal; a VPC endpoint is the correct solution for private connectivity.

179
MCQmedium

A SysOps administrator is investigating a security incident and needs to determine who deleted an S3 bucket. Which AWS service should be used to find this information?

A.AWS CloudTrail
B.AWS Trusted Advisor
C.S3 server access logs
D.CloudWatch Logs for the S3 service
AnswerA

AWS CloudTrail is the correct choice because it records management events, including the DeleteBucket API call, as part of its default audit trail. Each event captures the identity of the caller, the source IP address, the time of the action, and request parameters, which is exactly the evidence needed to investigate an unauthorized deletion. Since DeleteBucket is a control-plane operation, it appears in CloudTrail even if data-event logging is not configured.

Why this answer

AWS CloudTrail is the correct service because it records API activity across AWS services, including S3 bucket deletion events. When a user or role deletes an S3 bucket, CloudTrail logs the `DeleteBucket` API call with details such as the IAM user or role, source IP address, and timestamp, enabling the administrator to identify the responsible entity.

Exam trap

The trap here is that candidates often confuse S3 server access logs (which log object-level requests) with CloudTrail (which logs management-plane API calls), leading them to incorrectly choose S3 server access logs for bucket deletion events.

How to eliminate wrong answers

Option B (AWS Trusted Advisor) is wrong because it provides best-practice recommendations for cost, performance, security, and fault tolerance, but does not record or log API calls like S3 bucket deletions. Option C (S3 server access logs) is wrong because these logs record object-level access requests (e.g., GET, PUT, DELETE on objects) and HTTP status codes, but they do not capture management-plane API calls such as `DeleteBucket`; they are also not enabled by default and require a target bucket. Option D (CloudWatch Logs for the S3 service) is wrong because CloudWatch Logs can ingest log data from various sources, but S3 does not natively emit management-plane API logs to CloudWatch Logs; CloudTrail logs are the source for such events, and CloudWatch Logs can be used as a destination for CloudTrail, but the service itself does not directly capture `DeleteBucket` events.

180
MCQmedium

A company runs a web application on Amazon EC2 instances. The SysOps administrator needs to monitor two metrics: high CPU utilization (greater than 90%) and high memory utilization (greater than 85%). An alarm should trigger when both conditions are true simultaneously for a period of 5 minutes. Which CloudWatch feature should the administrator use to create this alarm?

A.Metric math
B.Composite alarm
C.Anomaly detection
D.Logs Insights
AnswerB

A composite alarm evaluates a rule that combines the states of multiple underlying CloudWatch alarms using AND/OR logic, so it can trigger only when every specified alarm is in the desired state. For example, you can set it to alarm only when both the CPU utilization and memory usage alarms are in ALARM, which is exactly the kind of multi-condition trigger the scenario requires. With composite alarms, you can also incorporate up to 10 alarms or other composite alarms, and you can define a period within which the conditions must remain met. This provides a managed, declarative way to create condition-based escalation without writing custom metric math.

Why this answer

A composite alarm in Amazon CloudWatch allows you to create an alarm that evaluates multiple conditions using logical operators (AND, OR, NOT). In this scenario, the administrator needs the alarm to trigger only when both CPU utilization > 90% AND memory utilization > 85% are true simultaneously for 5 minutes. Composite alarms evaluate the state of underlying metric alarms (e.g., two separate simple alarms for CPU and memory) and combine them with an AND condition, making it the correct feature for this requirement.

Exam trap

The trap here is that candidates often confuse Metric Math with composite alarms, thinking that arithmetic operations can simulate logical AND, but Metric Math cannot evaluate alarm states or combine them with logical operators—it only produces a new numeric metric series.

How to eliminate wrong answers

Option A is wrong because Metric Math is used to perform arithmetic operations on multiple metrics (e.g., sum, average, rate) to create a single time series, but it cannot evaluate logical conditions like AND across separate alarms; it only produces a new metric, not an alarm with combined states. Option C is wrong because Anomaly Detection uses machine learning to detect deviations from expected behavior based on historical patterns, but it does not support combining two separate conditions with an AND operator; it creates a single band for a metric. Option D is wrong because Logs Insights is a query engine for analyzing CloudWatch Logs data, not for creating alarms based on real-time metric thresholds; it cannot trigger alarms directly and does not support composite logic.

181
MCQhard

A company has a CloudWatch alarm that monitors the CPU utilization of an EC2 instance. The alarm is set to trigger when CPU utilization exceeds 80% for 5 consecutive minutes. The alarm state is 'INSUFFICIENT_DATA'. What does this mean?

A.The CPU utilization is below 80% for 5 minutes.
B.The CPU utilization has exceeded 80% for 5 minutes.
C.The alarm does not have enough data to determine the state.
D.The alarm is missing data points for the past 5 minutes.
AnswerC

The INSUFFICIENT_DATA state is specifically defined as the alarm being unable to evaluate the metric against the threshold because the number of data points received during the alarm's evaluation period is below the required minimum number. This occurs, for example, when no metrics are published, when a client app stops emitting metric values, or when a math expression returns no results. It is a distinct alarm state independent of OK and ALARM.

Why this answer

The INSUFFICIENT_DATA state in CloudWatch indicates that the alarm has not received enough metric data points to evaluate whether the threshold (CPU utilization > 80% for 5 consecutive minutes) has been breached. This typically occurs when the EC2 instance is newly launched, the CloudWatch agent is not reporting, or there are gaps in metric collection due to network issues or instance stops. It does not imply any conclusion about the CPU utilization level itself.

Exam trap

The trap here is that candidates confuse INSUFFICIENT_DATA with missing data points for the exact evaluation period, but the state actually means the alarm cannot determine whether the threshold is breached due to a lack of sufficient data across the configured evaluation periods, not just the last 5 minutes.

How to eliminate wrong answers

Option A is wrong because INSUFFICIENT_DATA does not mean the CPU utilization is below 80%; that would correspond to the ALARM state being false (OK state). Option B is wrong because exceeding 80% for 5 minutes would trigger the ALARM state, not INSUFFICIENT_DATA. Option D is wrong because while missing data points can cause INSUFFICIENT_DATA, the alarm state specifically indicates insufficient data to determine the state, not merely that data points are missing for the past 5 minutes—the alarm evaluates based on the configured evaluation periods and may have partial data.

182
MCQmedium

A company uses CloudWatch Logs to store application logs from EC2 instances. The SysOps team needs to search for specific error patterns across all log groups. What is the most efficient way to perform this search?

A.Define a CloudWatch metric filter to count errors and view the metric.
B.Use CloudWatch Logs Insights to run a query across the log groups.
C.Create a subscription filter to stream logs to Amazon ES and use Kibana.
D.Export the logs to Amazon S3 and use S3 Select to search.
AnswerB

CloudWatch Logs Insights is purpose-built for interactive analysis of log data stored in CloudWatch Logs, and it can query multiple log groups at once using a que CWL query language with parse, filter, stats, and sort commands. It is the fastest native option for a one-time search because it runs directly over the already-retained log events, returning the matching event details and summaries without requiring any additional infrastructure. For ad hoc troubleshooting, this avoids the setup time and operational overhead of forwarding or exporting logs to other services.

Why this answer

CloudWatch Logs Insights is purpose-built for ad-hoc querying and analysis of log data across multiple log groups. It allows you to run SQL-like queries (using a query language) to search for specific patterns, filter results, and aggregate data without needing to set up additional infrastructure. This is the most efficient method for the SysOps team's requirement because it provides immediate, interactive search capabilities directly within the AWS Management Console or via the AWS CLI.

Exam trap

The trap here is that candidates often confuse metric filters (which only aggregate counts) with the ability to search actual log content, leading them to choose Option A, or they over-engineer the solution by selecting Option C or D, not realizing that CloudWatch Logs Insights provides a native, serverless, and cost-effective query capability for exactly this scenario.

How to eliminate wrong answers

Option A is wrong because a metric filter only counts occurrences of a pattern and stores the count as a CloudWatch metric; it does not allow you to view the actual log events or search for specific error messages across log groups. Option C is wrong because creating a subscription filter to stream logs to Amazon Elasticsearch Service (now Amazon OpenSearch Service) and using Kibana adds significant setup overhead, latency, and cost; it is overkill for a simple search task and not the most efficient approach for an ad-hoc search. Option D is wrong because exporting logs to S3 and using S3 Select is inefficient for this use case: S3 Select is designed for querying structured data in files (like CSV or JSON) and requires a full export of logs, which is time-consuming and not suitable for real-time or frequent searches across multiple log groups.

183
Multi-Selecteasy

A SysOps administrator needs to track changes to security group rules in a VPC. Which AWS services can be used to monitor and log these changes? (Choose TWO.)

Select 2 answers
A.VPC Flow Logs.
B.AWS Trusted Advisor.
C.AWS CloudTrail.
D.AWS Config.
E.Amazon CloudWatch Logs.
AnswersC, D

AWS CloudTrail records all API calls made within the account as discrete events, including AuthorizeSecurityGroupIngress, RevokeSecurityGroupEgress, and DeleteSecurityGroup. Each event contains the requester, source IP, time, and request/response details, providing exactly the evidence needed to track who modified a security group and when. CloudTrail is the authoritative service for control-plane auditing and is the direct answer to this requirement.

Why this answer

AWS CloudTrail is correct because it records API calls made to the Amazon EC2 service, including AuthorizeSecurityGroupIngress, RevokeSecurityGroupEgress, and CreateSecurityGroup. These events capture the identity, source IP, and timestamp of every change to security group rules, providing an audit trail for compliance and troubleshooting.

Exam trap

The trap here is that candidates confuse VPC Flow Logs (which monitor traffic flows) with CloudTrail (which monitors API calls), or assume CloudWatch Logs alone can capture changes without a source like CloudTrail or Config.

184
MCQhard

An application running on Amazon ECS (Fargate) is experiencing intermittent failures. The logs show 'CannotPullContainerError: error pulling image configuration: download failed after attempts=6'. The SysOps team has verified that the image exists in Amazon ECR and the task role has permissions to pull from ECR. What is the most likely cause?

A.The container image is corrupted.
B.The ECR repository policy is not allowing the task role.
C.The ECS task definition has an incorrect memory allocation.
D.The ECS tasks are in a private subnet without a NAT gateway or VPC endpoints for ECR and S3.
AnswerD

Fargate tasks running in a private subnet do not have public IP addresses, so they cannot reach the internet without a NAT gateway. To pull images from ECR, the task also needs access to the Amazon S3 endpoints that store image layers, which requires either a NAT gateway or VPC endpoints for both ECR and S3. Without these routes, the image pull attempts time out or fail with a network connectivity error, exactly matching the scenario. This is the correct cause.

Why this answer

The error 'CannotPullContainerError: error pulling image configuration: download failed after attempts=6' indicates that the ECS task is unable to download the image layers from Amazon ECR. Since the image exists and the task role has permissions, the most likely cause is a network connectivity issue. When ECS tasks run in a private subnet without a NAT gateway or VPC endpoints for ECR and S3, they cannot reach the public ECR API endpoints or the S3 buckets that store image layers, causing the pull to fail after multiple retries.

Exam trap

The trap here is that candidates often assume the error is due to missing IAM permissions or a corrupted image, overlooking the fact that ECS tasks in private subnets require explicit network paths (NAT gateway or VPC endpoints) to reach ECR and S3, even when permissions are correctly configured.

How to eliminate wrong answers

Option A is wrong because a corrupted image would typically produce a different error, such as a manifest or layer integrity failure, not a download failure after multiple attempts. Option B is wrong because the task role already has permissions to pull from ECR, and the repository policy is a separate mechanism that would cause an authorization error (e.g., 'AccessDeniedException') rather than a download failure. Option C is wrong because incorrect memory allocation would cause the task to fail at launch with a resource-related error (e.g., 'CannotStartContainerError: ResourceInitializationError') or OOM kill, not an image pull failure.

185
MCQhard

An organization uses AWS CloudTrail to log API calls across multiple accounts in AWS Organizations. The logs are delivered to a central S3 bucket. The security team wants to receive near-real-time notifications whenever an IAM user creates a new access key. Which solution is the MOST operationally efficient?

A.Create an Amazon EventBridge rule that matches the CreateAccessKey event from CloudTrail and publishes to an Amazon SNS topic.
B.Enable S3 Event Notifications on the CloudTrail bucket to trigger a Lambda function that scans new objects for access key creation.
C.Configure CloudTrail to send logs to Amazon CloudWatch Logs and set up a metric filter that triggers an alarm.
D.Use Amazon CloudWatch Logs Insights to run a query every minute and send results via SNS.
AnswerA

EventBridge is fully integrated with CloudTrail and receives management events as soon as they occur, typically within seconds. A rule with an event pattern such as `source: aws.iam` and `eventName: CreateAccessKey` will match the IAM API call and trigger an SNS topic as a push target, delivering a notification in near-real-time with no polling, log scanning, or waiting for CloudTrail's 5-minute S3 log delivery.

Why this answer

Amazon EventBridge can directly consume CloudTrail events (including `CreateAccessKey`) in near-real-time without polling or custom code. By creating a rule that matches this specific event and targets an SNS topic, the security team gets immediate notifications with minimal operational overhead. This approach is serverless, event-driven, and requires no intermediate services or custom functions.

Exam trap

The trap here is that candidates often default to CloudWatch Logs metric filters or S3 Event Notifications because they are familiar, but they overlook that EventBridge provides the most direct, low-latency, and operationally efficient path for CloudTrail event-driven notifications.

How to eliminate wrong answers

Option B is wrong because S3 Event Notifications are object-creation events, not real-time; they introduce latency (typically minutes) and require a Lambda function to parse each log file, which is less efficient and adds complexity. Option C is wrong because sending CloudTrail logs to CloudWatch Logs and using metric filters adds latency (logs are delivered in batches, metric filters poll every minute) and requires additional configuration for alarms, making it less near-real-time than EventBridge. Option D is wrong because running a CloudWatch Logs Insights query every minute is polling-based, incurs query costs, and is not event-driven; it also introduces up to 60 seconds of delay and is operationally inefficient compared to a push-based EventBridge rule.

186
MCQhard

A company runs a critical web application on a fleet of EC2 instances behind an Application Load Balancer (ALB). The instances are in an Auto Scaling group. The operations team uses CloudWatch alarms to monitor the application's health. Recently, they noticed that the application's error rate has increased sporadically, but the CPU utilization and memory usage remain normal. The team suspects that the issue is related to a specific HTTP endpoint returning 5xx errors. They want to set up monitoring that will alert them when the error rate exceeds 5% of total requests over a 5-minute period. The application logs are already sent to CloudWatch Logs. Which combination of steps should the SysOps administrator take to meet this requirement?

A.Create a metric filter in CloudWatch Logs to extract error codes and total requests from the application logs. Create two custom metrics: one for error count and one for total requests. Then create a CloudWatch alarm using a math expression that calculates error rate (error count / total requests) and triggers when >0.05 for 5 minutes.
B.Enable AWS X-Ray on the application to trace requests and identify error patterns. Create a CloudWatch alarm on the X-Ray error rate metric.
C.Install the CloudWatch agent on the EC2 instances to collect application-level metrics. Configure the agent to emit a custom metric for error rate. Then create an alarm on that metric.
D.Enable detailed monitoring on the ALB and create a CloudWatch alarm on the HTTPCode_ELB_5XX metric with a threshold of 5% of the request count. Use the ALB's RequestCount metric to compute the percentage.
AnswerA

This is the correct approach because the application logs are already flowing into CloudWatch Logs, and a metric filter can parse them in real time to extract both the number of error codes (e.g., status codes or application-specific errors) and the total request count. By creating two custom metrics—ErrorCount and TotalRequests—you can then define a CloudWatch alarm using a metrics math expression such as e1/e2, with the alarm triggering when the ratio exceeds 0.05 for a 5-minute period. This leverages the existing log data without requiring additional instrumentation or external services, and it accurately reflects application-level error rates as observed in the logs.

Why this answer

The requirement is to alert when the 5xx error rate exceeds 5% of total requests over a 5-minute window, and the logs are already in CloudWatch Logs. A metric filter extracts the error count and total request count into custom metrics, and a CloudWatch alarm with a math expression (errors/requests) evaluates the ratio against the 0.05 threshold. This is the only option that produces a true percentage-based alarm from the existing log data.

Exam trap

SOA-C02 often tests whether candidates know that ALB's HTTPCode_ELB_5XX only counts load-balancer-generated errors, and that CloudWatch alarms need metric math (not a simple threshold) to evaluate a ratio like error percentage.

How to eliminate wrong answers

Option B is wrong because X-Ray traces requests for latency and dependency analysis but does not natively emit an 'error rate' CloudWatch metric that can be alarmed on as a percentage of total requests. Option C is wrong because the CloudWatch agent collects OS-level and application metrics but does not parse application logs to derive an error-rate metric; it would require custom application instrumentation that isn't described. Option D is wrong because HTTPCode_ELB_5XX counts only errors generated by the load balancer itself (e.g., 502/503/504 from unhealthy targets), not application-generated 5xx responses, and CloudWatch alarms cannot directly compute a percentage of another metric without a math expression.

187
MCQeasy

A SysOps administrator wants to receive alerts when the root user performs an action in the AWS account. Which service should be used?

A.AWS Identity and Access Management (IAM)
B.Amazon CloudWatch Metrics
C.AWS Config
D.AWS CloudTrail and Amazon CloudWatch Logs
AnswerD

AWS CloudTrail captures a full history of API activity and management events, including the root user's sign-in attempt, which is recorded as a 'ConsoleLogin' event with a 'userIdentity.type' of 'Root'. Sending that CloudTrail trail to Amazon CloudWatch Logs allows you to create a CloudWatch Logs metric filter that matches the JSON pattern for root-level sign-in events, and then attach a CloudWatch alarm to that metric with a threshold of one or more events. When the metric filter sees a matching root sign-in, the alarm transitions to ALARM and sends a notification to the SNS topic you configured. Together, CloudTrail supplies the detailed event data and CloudWatch Logs/alarms provide the real-time detection and alerting, making this the correct and recommended AWS solution.

Why this answer

AWS CloudTrail captures all API calls made by the root user as events. By sending these events to Amazon CloudWatch Logs, you can create a metric filter that matches root user activity and trigger an alarm via CloudWatch Alarms. This combination enables real-time notification when the root user performs any action.

Exam trap

The trap here is that candidates often choose AWS Config because it monitors resource changes, but they fail to realize that root user actions are API calls, not configuration changes, and thus require CloudTrail and CloudWatch Logs for detection.

How to eliminate wrong answers

Option A is wrong because AWS IAM manages users, roles, and permissions but does not provide event monitoring or alerting capabilities for root user actions. Option B is wrong because Amazon CloudWatch Metrics alone cannot capture or alert on specific API actions; it requires CloudTrail logs and metric filters to detect root user activity. Option C is wrong because AWS Config evaluates resource configurations and compliance rules, not API call activity; it cannot detect when the root user performs an action.

188
MCQeasy

A SysOps administrator needs to monitor the CPU utilization of an Amazon EC2 instance and receive an alert if it exceeds 80% for 10 consecutive minutes. The instance is in a VPC with no Internet access. What is the MOST efficient way to meet these requirements?

A.Use AWS Systems Manager to run a script on the instance that checks CPU and sends an SNS notification.
B.Create a CloudWatch alarm on the CPUUtilization metric with a period of 5 minutes and an evaluation period of 2.
C.Enable detailed monitoring on the EC2 instance and create a CloudWatch alarm on the CPUUtilization metric.
D.Install the CloudWatch agent on the EC2 instance to collect CPU metrics and create a CloudWatch alarm.
AnswerB

This is the correct approach because EC2 instances automatically emit CPUUtilization metrics every 5 minutes under basic monitoring, so a CloudWatch alarm with a period of 5 minutes will have the data it needs without any extra configuration. An evaluation period of 2 means the alarm only enters ALARM state after two consecutive data points breach the threshold, which reduces false positives from transient spikes. This is the simplest, most cost-effective way to monitor CPU utilization on an EC2 instance.

Why this answer

A CloudWatch alarm with a period of 5 minutes and an evaluation period of 2 means the alarm evaluates two consecutive 5-minute data points, totaling 10 minutes. Since the EC2 instance is in a VPC with no Internet access, CloudWatch metrics are still reported via the CloudWatch service endpoint within the VPC (or via VPC endpoints), so no additional agent or script is needed. This is the most efficient approach as it uses native CloudWatch functionality without requiring any custom scripts or additional software.

Exam trap

The trap here is that candidates often assume they need detailed monitoring or the CloudWatch agent to meet a specific time window, but the default 5-minute period with multiple evaluation periods can achieve the same result more efficiently and at lower cost.

How to eliminate wrong answers

Option A is wrong because using AWS Systems Manager to run a script on the instance that checks CPU and sends an SNS notification introduces unnecessary complexity and overhead; it requires the instance to have outbound internet access or a VPC endpoint for Systems Manager, and it is not the most efficient native solution. Option C is wrong because enabling detailed monitoring (1-minute metrics) is not required for this scenario; a 5-minute period with 2 evaluation periods already meets the 10-minute requirement, and detailed monitoring would incur additional cost without benefit. Option D is wrong because installing the CloudWatch agent is unnecessary; the EC2 instance already publishes the CPUUtilization metric by default (basic monitoring) without any agent, and the agent is only needed for custom or OS-level metrics, not for standard CPU utilization.

189
MCQhard

Refer to the exhibit. A SysOps administrator created this IAM policy for an application that sends custom metrics to CloudWatch and writes logs to CloudWatch Logs. The application reports that it cannot publish logs. What is the most likely reason?

A.The resource ARN for the logs actions is incorrect; it should include 'log-group:' before the wildcard.
B.The policy requires a condition key to restrict access to specific log groups.
C.The policy does not allow the cloudwatch:PutMetricData action for the specific metric.
D.The application must assume an IAM role to write logs.
AnswerA

The ARN for CloudWatch Logs actions must follow the format arn:aws:logs:region:account-id:log-group:log-group-name:*. Omitting the 'log-group:' prefix produces an invalid ARN that cannot match any log group, so the policy would not grant the necessary permissions. Without this prefix, IAM cannot resolve the resource to a specific log group, causing the logs actions to fail at runtime.

Why this answer

The IAM policy uses `arn:aws:logs:us-east-1:123456789012:*` for the `Resource` element of the `logs:PutLogEvents` and `logs:CreateLogGroup` actions. For CloudWatch Logs, the resource ARN must include the `log-group:` prefix before the log group name or wildcard, such as `arn:aws:logs:us-east-1:123456789012:log-group:*`. Without this prefix, the ARN does not match any valid CloudWatch Logs resource, causing the application to fail when attempting to publish logs.

Exam trap

The trap here is that candidates often assume a wildcard resource ARN like `arn:aws:logs:region:account:*` is sufficient for CloudWatch Logs actions, but they overlook the required `log-group:` prefix in the ARN structure, leading them to incorrectly suspect missing conditions or role assumption issues.

How to eliminate wrong answers

Option B is wrong because the policy does not require a condition key to restrict access to specific log groups; the issue is the malformed resource ARN, not the absence of conditions. Option C is wrong because the policy includes `cloudwatch:PutMetricData` with a wildcard resource (`*`), which is correct for CloudWatch custom metrics, and the application's failure is specifically about publishing logs, not metrics. Option D is wrong because the application can use IAM user credentials or an IAM role attached to an EC2 instance profile; the policy itself does not require assuming a role, and the error is due to the incorrect resource ARN, not the authentication method.

190
MCQmedium

Refer to the exhibit. The alarm has been in INSUFFICIENT_DATA state for several hours. What is the most likely cause?

A.The alarm evaluation period is too long.
B.The EC2 instance is stopped or terminated.
C.The instance has no CloudWatch agent installed.
D.The instance is running but the CPU utilization is below the threshold.
AnswerB

If the instance is stopped, no metrics are emitted.

Why this answer

The INSUFFICIENT_DATA state for several hours indicates that CloudWatch has not received any metric data points for the specified period. If the EC2 instance is stopped or terminated, the CloudWatch agent stops sending metrics, and the default CPU utilization metric (which is published by AWS, not the agent) also ceases because the instance is no longer running. This causes the alarm to remain in INSUFFICIENT_DATA indefinitely until the instance is started again or the metric resumes.

Exam trap

The trap here is that candidates often confuse INSUFFICIENT_DATA with ALARM or OK states, mistakenly thinking low CPU utilization or missing CloudWatch agent would cause this state, when in fact INSUFFICIENT_DATA strictly means no metric data has been received at all for the evaluation period.

How to eliminate wrong answers

Option A is wrong because the alarm evaluation period being too long would only delay transitions between states, but it would not cause a permanent INSUFFICIENT_DATA state; data would still be collected and eventually evaluated. Option C is wrong because the CPU utilization metric is a default EC2 metric published by AWS automatically without requiring the CloudWatch agent; the agent is only needed for custom or OS-level metrics. Option D is wrong because if the instance is running and CPU utilization is below the threshold, the alarm would be in ALARM or OK state (depending on the comparison operator), not INSUFFICIENT_DATA; INSUFFICIENT_DATA specifically means no data points are available, not that data exists but is below a threshold.

191
MCQhard

A company uses AWS CloudTrail to log API calls in a multi-account environment. The security team wants to be alerted when an IAM user in the production account modifies a security group to allow inbound SSH from 0.0.0.0/0. Which combination of actions should be taken to meet this requirement?

A.Use AWS Config managed rule 'restricted-ssh' to detect the security group change and trigger an SNS notification.
B.Enable AWS Security Hub and configure a custom insight to detect the security group modification.
C.Create an AWS Lambda function that is triggered by CloudTrail events and publishes to SNS.
D.Stream CloudTrail logs to CloudWatch Logs, create a metric filter for the specific API call, and set a CloudWatch Alarm that sends a notification to an SNS topic.
AnswerD

Streaming CloudTrail logs to CloudWatch Logs is the foundation for real-time monitoring of API activity. You can then create a CloudWatch Logs metric filter that matches the specific API call, such as AuthorizeSecurityGroupIngress or RevokeSecurityGroupIngress, and increments a custom metric. Finally, set a CloudWatch alarm on that metric with a threshold of one, which triggers an SNS notification to the designated topic. This is the standard, event-driven method that provides immediate alerting with minimal overhead and is the recommended pattern for API call monitoring.

Why this answer

CloudTrail logs can be streamed to CloudWatch Logs, where a metric filter can be created to match the specific API call (e.g., AuthorizeSecurityGroupIngress with a CIDR of 0.0.0.0/0 and port 22). A CloudWatch Alarm based on that metric can then trigger an SNS notification, providing a real-time alert for the exact security group modification.

Exam trap

The trap here is that candidates may confuse AWS Config rules (which are reactive and evaluate configuration state) with CloudWatch metric filters (which provide real-time alerting on API calls), leading them to choose Option A despite its inability to trigger immediate notifications on the specific event.

How to eliminate wrong answers

Option A is wrong because the AWS Config managed rule 'restricted-ssh' only checks whether security groups allow unrestricted SSH access at the time of evaluation; it does not provide real-time alerting on the API call itself and cannot trigger an SNS notification directly without additional configuration. Option B is wrong because Security Hub custom insights are used for querying and visualizing findings, not for real-time alerting; they do not directly send notifications to SNS. Option C is wrong because CloudTrail events cannot directly trigger a Lambda function; CloudTrail delivers events to an S3 bucket or CloudWatch Logs, and Lambda can be triggered from those sources, but the option states 'triggered by CloudTrail events' which is technically incorrect without an intermediary.

192
MCQeasy

A SysOps administrator needs to send a notification when an EC2 instance's CPU utilization exceeds 90% for 5 consecutive minutes. Which AWS service should be used to create the alarm?

A.AWS Config
B.Amazon CloudWatch Alarms
C.AWS Trusted Advisor
D.Amazon EventBridge
AnswerB

Amazon CloudWatch Alarms are the native mechanism for monitoring a single metric or a metric math expression over a specified time period. When the metric breaches a defined threshold for consecutive periods, the alarm transitions to the ALARM state and can trigger an action, such as publishing to an SNS topic to send a notification. This directly fulfills the requirement to notify when a metric condition is met.

Why this answer

Amazon CloudWatch Alarms are the correct service because they allow you to monitor a specific metric (e.g., EC2 CPUUtilization) and trigger an action when the metric crosses a defined threshold for a specified number of consecutive evaluation periods. In this case, you can create a CloudWatch alarm with a statistic of 'Average', a threshold of 90%, and set the 'Datapoints to Alarm' and 'Evaluation Periods' to 5 (with a period of 1 minute) to achieve the '5 consecutive minutes' requirement. The alarm can then send a notification via Amazon SNS.

Exam trap

The trap here is that candidates often confuse Amazon EventBridge with CloudWatch Alarms, thinking EventBridge can directly evaluate metric thresholds, but EventBridge requires a CloudWatch Alarm to generate the event, and it cannot perform the metric evaluation itself.

How to eliminate wrong answers

Option A is wrong because AWS Config is a service for evaluating and auditing the configuration of AWS resources against desired policies (e.g., compliance rules), not for monitoring real-time performance metrics like CPU utilization. Option C is wrong because AWS Trusted Advisor provides best-practice recommendations for cost optimization, security, fault tolerance, and performance, but it does not create metric-based alarms or send notifications for threshold breaches. Option D is wrong because Amazon EventBridge is a serverless event bus for routing events between AWS services and custom applications, but it cannot directly evaluate a metric over a time window; it relies on CloudWatch Alarms or other sources to generate the events that trigger its rules.

193
MCQmedium

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The SysOps administrator needs to monitor the application's HTTP 5xx error rate and set an alarm when the error rate exceeds 5% over a 5-minute period. The alarm must trigger an Amazon SNS notification. Which metric should be used for the alarm?

A.HTTPCode_ELB_5XX_Count
B.HTTPCode_Target_5XX_Count
C.RequestCount
D.TargetResponseTime
AnswerB

The HTTPCode_Target_5XX_Count metric reports the number of HTTP 5xx responses returned directly by the registered EC2 instances, capturing errors such as 500 Internal Server Error from the application. Because these are the actual responses sent to clients from your web application, this metric is the correct measurement for an alarm that detects application-level failures. You can then create a CloudWatch alarm on this metric to trigger when the count exceeds a threshold, possibly combined with RequestCount to derive an error rate.

Why this answer

The alarm must monitor the error rate from the application targets (EC2 instances) behind the ALB. HTTPCode_Target_5XX_Count tracks HTTP 5xx responses generated by the targets themselves, which directly reflects application-level errors. To calculate the error rate, you would divide this metric by RequestCount, but the metric itself is the correct source for target-side 5xx errors.

Exam trap

The trap here is that candidates confuse HTTPCode_ELB_5XX_Count with HTTPCode_Target_5XX_Count, assuming all 5xx errors originate from the load balancer, when in fact the ALB separates its own errors from target-generated errors to provide precise fault isolation.

How to eliminate wrong answers

Option A is wrong because HTTPCode_ELB_5XX_Count tracks 5xx errors generated by the ALB itself (e.g., due to load balancer failures or configuration issues), not the application targets, so it would not reflect the application's error rate. Option C is wrong because RequestCount is a count of all requests processed by the ALB, not a measure of error rate; it is used as a denominator in rate calculations but cannot trigger an alarm on error percentage alone. Option D is wrong because TargetResponseTime measures the time taken for targets to respond, not error codes, and is unrelated to HTTP 5xx error rate monitoring.

194
MCQmedium

A company uses AWS Lambda functions that process data from an Amazon SQS queue. The Lambda function is failing intermittently due to timeouts. The SysOps administrator needs to be notified immediately when the function times out. What is the most efficient way to achieve this?

A.Modify the Lambda function to catch the timeout exception and log it to CloudWatch Logs
B.Create a CloudWatch alarm on the Lambda Errors metric that sends an SNS notification
C.Set up a CloudWatch Logs subscription filter to send error logs to an SNS topic
D.Configure an Amazon SNS topic as a Lambda destination for the function
AnswerB

A CloudWatch alarm on the AWS/Lambda Errors metric directly monitors every failed invocation, including timeouts, and transitions to ALARM when the error count exceeds the configured threshold. The alarm then publishes to an SNS topic, which can deliver email, SMS, or trigger a webhook or incident-management tool. This is the native, low-latency path because the metric is emitted automatically by the Lambda service without any code changes or log processing.

Why this answer

A CloudWatch alarm on the Lambda Errors metric directly monitors function invocations that result in errors, including timeouts. When the alarm state changes to ALARM, it can immediately trigger an SNS notification, providing the fastest and most efficient notification mechanism without requiring code changes or additional infrastructure.

Exam trap

The trap here is that candidates often assume they can catch a timeout exception inside the function code (Option A) or that Lambda destinations (Option D) are a catch-all for all errors, when in fact they only apply to specific invocation types and do not cover all timeout scenarios.

How to eliminate wrong answers

Option A is wrong because catching a timeout exception inside the Lambda function is not possible — Lambda enforces a hard timeout at the configured limit, and the runtime cannot catch it; the function simply terminates with an error. Option C is wrong because a CloudWatch Logs subscription filter requires the function to first write a log entry, which may not occur reliably on timeout, and it adds latency and complexity compared to a direct metric alarm. Option D is wrong because Lambda destinations are triggered only on successful invocation or explicit failure states like 'OnFailure' for asynchronous invocations, but they do not capture all timeout errors reliably and require additional configuration; also, SNS as a destination is not supported for synchronous invocations like those from SQS.

195
MCQeasy

A SysOps administrator needs to track changes made to an Amazon S3 bucket policy and receive notifications when changes occur. Which AWS service should be used?

A.AWS Trusted Advisor
B.Amazon CloudWatch Events
C.AWS Config
D.AWS CloudTrail
AnswerC

AWS Config is the correct choice because it continuously records and evaluates configuration changes to supported AWS resources, including S3 bucket policies. You can set up AWS Config rules to detect specific changes and configure Amazon SNS to send notifications when a resource deviates from a desired configuration or when a change occurs. It maintains a configuration history and timeline, enabling both auditing and alerting for modifications, which directly matches the requirement.

Why this answer

AWS Config is designed to record configuration changes to AWS resources, including S3 bucket policies, and can evaluate them against desired configurations. It can send notifications via Amazon SNS when changes occur, and it provides a history of configuration changes. This directly meets the requirement to track changes and receive notifications.

Exam trap

The trap is confusing CloudTrail (API activity logging) with AWS Config (configuration change tracking); candidates often pick CloudTrail because it 'tracks changes', but it does not provide configuration state or compliance notifications.

How to eliminate wrong answers

Option A is wrong because AWS Trusted Advisor provides best-practice checks and recommendations, not change tracking or notifications for specific resource policies. Option B is wrong because Amazon CloudWatch Events (now Amazon EventBridge) can react to API calls via CloudTrail, but it does not track configuration changes or provide a configuration history; it's an event bus. Option D is wrong because AWS CloudTrail records API activity (who made the change) but does not track resource configuration state or send notifications on configuration changes by itself; it logs events but requires additional services for alerting.

196
MCQmedium

A company is running a web application on EC2 instances behind an Application Load Balancer. The application experiences intermittent latency spikes. The SysOps administrator needs to identify the root cause. Which set of CloudWatch metrics should be analyzed first?

A.ALB TargetResponseTime and EC2 CPUUtilization
B.EC2 CPUUtilization and NetworkIn
C.EC2 StatusCheckFailed and ALB UnhealthyHostCount
D.ALB RequestCount and HealthyHostCount
AnswerA

TargetResponseTime isolates ALB-to-target latency, showing whether spikes originate at the application tier, while CPUUtilization reveals host saturation driving slow responses. Together they distinguish backend compute pressure from network or client-side delay, the fastest first step before drilling into ELB 5xx counts or request queues.

Why this answer

Intermittent latency spikes in a web application behind an Application Load Balancer (ALB) are most directly investigated by correlating ALB TargetResponseTime (which measures the time taken for the target to respond to the ALB) with EC2 CPUUtilization (which indicates whether the instance is under compute pressure). A spike in TargetResponseTime alongside high CPUUtilization suggests the EC2 instance is struggling to process requests, pointing to a compute bottleneck as the root cause.

Exam trap

The trap here is that candidates often confuse latency metrics with availability metrics, choosing options like C or D that indicate failures or traffic volume, rather than the performance-specific metrics needed to diagnose intermittent slowness.

How to eliminate wrong answers

Option B is wrong because while EC2 CPUUtilization is relevant, NetworkIn alone does not directly indicate latency; high network input could be normal traffic and does not measure response time or processing delays. Option C is wrong because EC2 StatusCheckFailed and ALB UnhealthyHostCount indicate instance or health check failures, not intermittent latency spikes; these metrics would show binary health states, not gradual performance degradation. Option D is wrong because ALB RequestCount and HealthyHostCount measure traffic volume and target health, not response latency; high request count alone does not explain why responses are slow.

197
MCQmedium

A SysOps administrator needs to monitor the disk usage on Amazon EC2 instances running Linux. The administrator wants to collect disk utilization metrics every 5 minutes and set up an alarm when disk usage exceeds 80%. Which solution meets these requirements?

A.Use the EC2 detailed monitoring feature to collect disk metrics.
B.Install the Amazon CloudWatch Agent on the instances and configure it to collect disk metrics.
C.Use AWS Systems Manager Patch Manager to check disk space.
D.Configure an Amazon CloudWatch metric filter on the system log.
AnswerB

The Amazon CloudWatch Agent runs inside the guest OS and gathers custom metrics from the operating system, including disk utilization metrics like disk_used_percent, disk_free, and inode usage. You define the metrics to collect in the agent configuration file, and the agent sends them to CloudWatch, where you can create alarms on thresholds. This is the standard method for monitoring filesystem space on EC2 instances.

Why this answer

The Amazon CloudWatch Agent is required to collect custom metrics like disk utilization from EC2 instances. It can be configured to gather disk space metrics every 5 minutes and publish them to CloudWatch, where an alarm can be set to trigger when usage exceeds 80%. EC2 detailed monitoring only collects hypervisor-level metrics (CPU, network, disk I/O), not guest OS-level disk usage.

Exam trap

The trap here is that candidates confuse EC2 detailed monitoring (which collects hypervisor-level metrics) with the ability to collect guest OS metrics like disk usage, leading them to choose Option A incorrectly.

How to eliminate wrong answers

Option A is wrong because EC2 detailed monitoring provides hypervisor-level metrics such as CPU, network, and disk I/O, but it does not collect guest OS-level disk utilization (e.g., filesystem usage percentage). Option C is wrong because AWS Systems Manager Patch Manager is used for patching operating systems and applications, not for monitoring disk space or setting CloudWatch alarms. Option D is wrong because CloudWatch metric filters parse log data from log groups to create metrics, but they cannot extract disk usage metrics from system logs unless the logs contain structured disk usage data, and they do not replace the need for a CloudWatch agent to collect guest OS metrics.

198
MCQeasy

A SysOps administrator wants to be notified when an Auto Scaling group launches a new instance. Which AWS service can be used to capture the Auto Scaling lifecycle events and send a notification?

A.AWS Config
B.Amazon CloudWatch Logs
C.Amazon Simple Notification Service (SNS)
D.AWS CloudTrail
AnswerC

Amazon Simple Notification Service (SNS) is a fully managed pub/sub messaging service that Auto Scaling can directly integrate with for lifecycle events. You create an SNS topic, subscribe endpoints like email or Lambda, and then configure the Auto Scaling group to publish notifications such as 'autoscaling:EC2_INSTANCE_LAUNCH' and 'autoscaling:EC2_INSTANCE_TERMINATE'. This provides immediate, push-based alerts to administrators, making it the correct choice.

Why this answer

Amazon SNS is the correct choice because it can receive lifecycle notifications from Auto Scaling groups via Amazon EventBridge (formerly CloudWatch Events) and then deliver those notifications to subscribers (e.g., email, SMS, HTTP endpoints). Auto Scaling groups emit lifecycle events (e.g., `EC2 Instance-launch Lifecycle Action`) that can be captured by EventBridge rules, which then invoke an SNS topic to send the notification.

Exam trap

The trap here is that candidates often confuse AWS CloudTrail (which records API calls) with EventBridge (which captures service events), leading them to choose CloudTrail for event-driven notifications, but CloudTrail does not handle lifecycle events or push notifications directly.

How to eliminate wrong answers

Option A is wrong because AWS Config is a service for evaluating resource configurations against desired policies and tracking configuration changes, not for capturing real-time lifecycle events or sending notifications. Option B is wrong because Amazon CloudWatch Logs is used to store, monitor, and access log files from AWS resources; it does not natively send notifications for Auto Scaling lifecycle events without additional integration (e.g., metric filters to SNS). Option D is wrong because AWS CloudTrail records API calls for auditing and compliance, but it does not capture Auto Scaling lifecycle events (which are not API calls) and cannot directly send notifications.

199
MCQeasy

A company uses Amazon CloudWatch Logs to store application logs. The SysOps administrator needs to detect when the number of log entries containing the string 'ERROR' exceeds 100 in any 5-minute window. When this threshold is breached, an email should be sent to the operations team. Which combination of AWS services should be used with the least operational overhead?

A.CloudWatch Logs Insights scheduled query with SNS action.
B.Create a metric filter on the log group for 'ERROR', then create a CloudWatch alarm on that metric with an SNS action to send email.
C.Use a Lambda function that reads the log stream and sends an email via Amazon Simple Email Service (SES) when errors exceed 100.
D.Install an agent on the application server that sends logs to Amazon SQS, then poll the queue with a Lambda function to trigger an email.
AnswerB

This is the native, fully managed approach: a metric filter is attached to the log group and processes new log events in real-time as CloudWatch Logs ingests them, extracting each occurrence of the pattern 'ERROR' and incrementing a custom CloudWatch metric. A CloudWatch alarm continuously evaluates that metric against the specified threshold (e.g., more than 100 errors in a 5-minute period); when the alarm transitions to the ALARM state, it invokes an SNS topic that sends the email notification. This pattern uses only built-in CloudWatch features and requires no custom code, external agents, or additional queueing services.

Why this answer

It uses CloudWatch metric filters to extract the count of 'ERROR' log entries as a custom metric, then a CloudWatch alarm on that metric triggers an SNS topic to send email notifications. This approach requires no custom code or additional infrastructure, minimizing operational overhead while meeting the requirement of detecting >100 errors in any 5-minute period.

Exam trap

The trap here is that candidates may overcomplicate the solution by choosing a Lambda-based or custom agent approach, not realizing that CloudWatch metric filters and alarms provide a fully managed, serverless way to monitor log patterns with minimal operational overhead.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs Insights scheduled queries do not natively support triggering actions like SNS; they are designed for ad-hoc analysis and can only output results to S3 or other destinations, not directly invoke SNS. Option C is wrong because using a Lambda function to read log streams and send email via SES introduces unnecessary complexity, custom code, and potential latency, increasing operational overhead compared to the native metric filter and alarm approach. Option D is wrong because installing an agent to send logs to SQS and polling with Lambda adds significant operational overhead, requires managing custom infrastructure, and is not a native CloudWatch solution for real-time log monitoring.

200
MCQeasy

A SysOps administrator notices that an Amazon RDS instance's CPU utilization is consistently above 90% during peak hours. The application is read-heavy and can tolerate eventual consistency. Which action would MOST effectively reduce CPU load?

A.Increase the instance size from db.r5.large to db.r5.xlarge.
B.Enable Multi-AZ deployment for the RDS instance.
C.Create one or more read replicas and direct read traffic to them.
D.Change the storage type to Provisioned IOPS.
AnswerC

Creating one or more read replicas and redirecting SELECT queries to them offloads a significant portion of the read workload from the primary DB instance, directly lowering its CPU utilization. The primary continues to handle writes and any critical read-your-own-writes traffic, while the replicas asynchronously apply changes and can serve large numbers of reads. This horizontal scaling approach is the standard and most cost-effective solution for read-heavy or OLTP workloads experiencing high CPU. Read replicas can even be placed across Availability Zones or regions to also improve read latency.

Why this answer

The application is read-heavy and can tolerate eventual consistency, making read replicas the ideal solution. By offloading read traffic to one or more read replicas, the primary RDS instance's CPU load is reduced because it no longer has to process all read queries. This directly addresses the high CPU utilization during peak hours without requiring a larger instance or other changes.

Exam trap

The trap here is that candidates often confuse Multi-AZ standby replicas with read replicas, assuming Multi-AZ can also offload read traffic, but AWS explicitly prohibits using the Multi-AZ standby for reads—it only supports failover.

How to eliminate wrong answers

Option A is wrong because increasing the instance size (e.g., from db.r5.large to db.r5.xlarge) adds more CPU and memory capacity, but it does not offload read traffic; it only scales the existing single instance, which may still be overwhelmed by the same read-heavy workload. Option B is wrong because enabling Multi-AZ deployment provides high availability and automatic failover via a synchronous standby replica, but that standby replica cannot serve read traffic (it is not a read replica), so it does not reduce CPU load on the primary instance. Option D is wrong because changing the storage type to Provisioned IOPS improves I/O performance and reduces latency, but it does not directly reduce CPU utilization; CPU load is driven by query processing, not storage throughput.

201
Multi-Selecteasy

A SysOps administrator needs to monitor the CPU and memory utilization of an EC2 instance running a legacy application that cannot be modified. Which TWO methods can be used to collect this information? (Choose TWO.)

Select 2 answers
A.Enable detailed monitoring on the instance to get memory metrics.
B.Install the CloudWatch agent on the instance and configure it to collect memory metrics.
C.Use a custom script to push memory data to CloudWatch via the PutMetricData API.
D.Use the EC2 hypervisor metrics available from CloudWatch.
E.Use AWS Systems Manager Inventory to collect memory utilization.
AnswersB, C

The CloudWatch agent runs inside the EC2 instance's guest OS and reads memory utilization directly from OS sources, such as /proc/meminfo on Linux or performance counters on Windows. After installation, you configure the agent with a JSON file or the console wizard, and it publishes memory metrics under the reserved CWAgent namespace—for example, mem_used_percent or mem_available. This is the standard, fully supported method for collecting memory metrics because the agent has privileged access to the OS internals, something the hypervisor cannot provide.

Why this answer

The CloudWatch agent can be installed on an EC2 instance to collect custom metrics, including memory utilization, which is not available by default from the hypervisor. The agent sends these metrics to CloudWatch using the PutMetricData API, enabling monitoring of in-guest resources like memory and disk.

Exam trap

The trap here is that candidates often assume detailed monitoring or hypervisor metrics include memory utilization, not realizing that memory is an in-guest metric requiring an agent or custom script to collect.

202
MCQmedium

Refer to the exhibit. A SysOps administrator creates the CloudWatch Alarm shown. However, the alarm never enters ALARM state even though the CPU utilization of the EC2 instance is consistently above 90%. What is the most likely reason?

A.The alarm is missing the InstanceId dimension.
B.The threshold is set too high; it should be 80%.
C.The evaluation periods are too few.
D.The statistic should be Maximum instead of Average.
AnswerA

A CloudWatch EC2 metric such as CPUUtilization is published with an InstanceId dimension, and the metric name alone is ambiguous because multiple instances in the account can have similar CPU data. Without the InstanceId dimension, the alarm cannot select the specific instance's metric stream; in practice, the alarm will either fail validation or remain in INSUFFICIENT_DATA because no datapoints match the metric signature. Adding the InstanceId dimension scopes the alarm to the instance shown in the exhibit, which lets CloudWatch evaluate that instance's CPU utilization against the threshold.

Why this answer

The alarm never enters ALARM state because it is missing the required `InstanceId` dimension. CloudWatch metrics for EC2, such as `CPUUtilization`, are published with the `InstanceId` dimension to uniquely identify the data stream. Without specifying this dimension in the alarm configuration, CloudWatch cannot match the alarm to the metric data emitted by the EC2 instance, so the alarm remains in INSUFFICIENT_DATA or OK state regardless of actual CPU usage.

Exam trap

The trap here is that candidates focus on tuning threshold or evaluation periods, overlooking the fundamental requirement that CloudWatch alarms must include the correct dimensions to match the metric data stream.

How to eliminate wrong answers

Option B is wrong because the threshold being set to 90% is not the issue; the alarm never evaluates the metric at all due to the missing dimension, so adjusting the threshold would not fix the problem. Option C is wrong because the evaluation periods being too few would cause the alarm to take longer to trigger or require more consecutive datapoints, but it would still eventually enter ALARM state if the metric were being evaluated; here the alarm never evaluates due to the missing dimension. Option D is wrong because using Average vs Maximum might affect sensitivity, but the alarm is not receiving any metric data to evaluate, so changing the statistic would not resolve the missing dimension issue.

203
Multi-Selectmedium

A company has a CloudWatch dashboard that displays metrics for several EC2 instances. The SysOps administrator wants to share the dashboard with external stakeholders who do not have AWS accounts. Which actions should the administrator take? (Select TWO.)

Select 2 answers
A.Use the CloudWatch dashboard sharing feature to make the dashboard public.
B.Create a web application that reads CloudWatch metrics and host it on EC2.
C.Create IAM users for each stakeholder and assign appropriate permissions.
D.Generate a shareable URL for the dashboard and send it to the stakeholders.
E.Export the dashboard to Amazon QuickSight and share it via email.
AnswersA, D

The CloudWatch dashboard sharing feature provides a built-in mechanism to expose a dashboard to anyone via a public link without requiring AWS credentials or sign-in. This is the most direct solution for external stakeholders because they only need the URL to view the live metrics, and you can revoke access by disabling sharing at any time. It avoids the need to create AWS identities or build custom applications.

Why this answer

Option A is correct because CloudWatch dashboards support a built-in sharing feature that lets you share a dashboard publicly with anyone, including people who do not have AWS accounts, by configuring the dashboard's share settings. Option D is correct because after enabling sharing, CloudWatch generates a shareable URL that can be sent to external stakeholders so they can view the dashboard without signing in to AWS. Options B, C, and E are not appropriate: building a custom EC2 web application is unnecessary and overly complex, creating IAM users requires stakeholders to have AWS accounts and credentials, and exporting to QuickSight is not a supported CloudWatch dashboard sharing mechanism for anonymous external viewers.

Exam trap

The trap here is that candidates often confuse the CloudWatch dashboard sharing feature with IAM-based access, assuming external users must have AWS credentials, when in fact the public URL mechanism is designed specifically for sharing with non-AWS users.

204
MCQeasy

A SysOps administrator needs to centralize logs from multiple AWS accounts into a single S3 bucket for analysis. Which solution is the MOST operationally efficient?

A.Use S3 replication to copy logs from each account's bucket to a central bucket.
B.Use CloudWatch Logs subscription filter to stream logs to a central account.
C.Use Amazon Kinesis Data Firehose to deliver logs from each account to a central S3 bucket.
D.Configure CloudTrail in each account to deliver logs to the same S3 bucket in the central account.
AnswerD

Configure each account's CloudTrail trail to deliver logs to the same central S3 bucket by creating a shared destination bucket with a bucket policy that trusts the CloudTrail service principal (cloudtrail.amazonaws.com) and grants write permissions to each source account. Set a distinct log-file prefix per account or enable partition-based prefixes to keep logs organized, and ensure the bucket versioning is enabled for log integrity. This is the standard, AWS-supported method for centralizing CloudTrail logs across multiple accounts, requiring no replication, streaming, or additional services.

Why this answer

AWS CloudTrail can be configured in each account to deliver logs directly to the same S3 bucket in a central account by specifying the central bucket's ARN and setting appropriate bucket policies. This approach is operationally efficient as it eliminates the need for intermediate services or replication, reducing complexity and cost while ensuring logs are centralized without manual intervention.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing Kinesis or replication, missing that CloudTrail natively supports cross-account S3 delivery, which is the simplest and most cost-effective method for centralizing logs.

How to eliminate wrong answers

Option A is wrong because S3 replication requires logs to first be delivered to separate buckets in each account, then replicated to a central bucket, adding latency, storage costs, and management overhead; it is not the most operationally efficient. Option B is wrong because CloudWatch Logs subscription filters are designed to stream logs to a central account for real-time processing, but they require additional configuration and do not directly deliver to S3 without an intermediary like Kinesis or Lambda, making it less efficient for simple centralized storage. Option C is wrong because Amazon Kinesis Data Firehose introduces an unnecessary streaming layer and additional cost for a batch-logging use case, whereas CloudTrail can directly write to S3 without extra services.

205
Multi-Selectmedium

A company is using Amazon CloudWatch Logs to centralize logs from multiple EC2 instances. The SysOps administrator needs to ensure that log data is encrypted at rest and in transit. Which TWO actions should the administrator take? (Choose TWO.)

Select 2 answers
A.Encrypt the CloudWatch Logs log group using an AWS KMS customer managed key.
B.Enable server-side encryption with S3-managed keys (SSE-S3) on the CloudWatch Logs log group.
C.Apply an S3 bucket policy that requires encryption for objects uploaded to the bucket.
D.Configure the CloudWatch Logs agent to use TLS/SSL for log delivery.
E.Install an AWS Certificate Manager (ACM) certificate on the EC2 instances.
AnswersA, D

CloudWatch Logs supports encryption at rest by associating an AWS KMS customer managed key with the log group. When this key is attached, every incoming log event is encrypted before storage and decrypted only by authorized readers; you must grant the CloudWatch Logs service permission to use the key via key policies. Using a customer managed key allows you to control rotation, access, and lifecycle independently of the default AWS-managed aws/logs key.

Why this answer

CloudWatch Logs supports encryption at rest using AWS KMS customer managed keys (CMKs). By associating a KMS key with a log group, all log data stored in that log group is encrypted at rest, meeting the encryption-at-rest requirement. Option D is correct because the CloudWatch Logs agent can be configured to use TLS/SSL (port 443) for log delivery, ensuring encryption in transit between the EC2 instances and CloudWatch Logs.

Exam trap

The trap here is that candidates may confuse CloudWatch Logs encryption with S3 server-side encryption options (SSE-S3 or SSE-KMS) or assume that ACM certificates are needed for agent-to-service encryption, when in fact the CloudWatch Logs agent uses AWS SDK-managed TLS automatically.

206
MCQmedium

A SysOps administrator needs to monitor Amazon S3 for object-level operations such as PUT and DELETE events in a specific bucket. The administrator wants these events to be sent to an Amazon SQS queue for downstream processing by an application. Which solution should be used to achieve this with the least operational overhead?

A.Use Amazon CloudWatch Events to match S3 API calls from CloudTrail and route to an SQS queue.
B.Configure an S3 event notification on the bucket to send events to an SQS queue.
C.Deploy an application that periodically polls S3 for changes using ListObjects.
D.Use AWS CloudTrail to deliver logs to CloudWatch Logs, then create a metric filter and trigger a Lambda function to send to SQS.
AnswerB

S3 bucket event notifications natively publish a rich event envelope (bucket, object key, size, etag) to a designated SQS queue as soon as the object operation completes, making them the most direct, low-latency mechanism for this requirement. The integration is fully managed: you enable the notification in the bucket configuration, attach an SQS resource policy that allows S3 to send messages, and S3 handles retries and batching automatically. This eliminates the need for custom intermediaries or polling, and supports prefix/suffix filters so the queue only receives relevant object events.

Why this answer

Amazon S3 can directly publish event notifications for object-level operations (e.g., PUT, DELETE) to an SQS queue without any intermediate services. This native integration requires no custom code or additional infrastructure, minimizing operational overhead while meeting the requirement to send events to SQS for downstream processing.

Exam trap

The trap here is that candidates may overcomplicate the solution by choosing CloudTrail or Lambda-based approaches, forgetting that S3 has a built-in, direct event notification feature for SQS that requires no additional services.

How to eliminate wrong answers

Option A is wrong because Amazon CloudWatch Events can match S3 API calls from CloudTrail, but this approach requires enabling CloudTrail and incurs additional cost and complexity; it also introduces latency and is not the simplest native method for S3 event delivery to SQS. Option C is wrong because periodically polling S3 with ListObjects is inefficient, does not capture real-time events, and introduces significant operational overhead for an application that must detect changes. Option D is wrong because it chains CloudTrail logs to CloudWatch Logs, then uses a metric filter and Lambda to send to SQS, adding multiple layers of complexity, cost, and potential failure points compared to the direct S3 event notification.

207
MCQeasy

A company runs containerized applications on Amazon ECS using the Fargate launch type. The SysOps administrator needs to monitor CPU and memory utilization at the task level. Which AWS service provides pre-built dashboards and metrics for this purpose?

A.Amazon CloudWatch Container Insights
B.Amazon CloudWatch custom metrics
C.Amazon CloudWatch Logs
D.Amazon CloudWatch Synthetics
AnswerA

Amazon CloudWatch Container Insights automatically collects and aggregates CPU, memory, disk, network, and task-level metrics from ECS clusters, services, and tasks, then displays them on prebuilt, out-of-the-box dashboards without requiring custom instrumentation. For ECS on EC2, you run the CloudWatch agent as a daemon service; for Fargate, you simply enable the cluster setting. This is exactly the native, low-effort monitoring solution that matches the scenario.

Why this answer

Amazon CloudWatch Container Insights provides pre-built dashboards and metrics specifically for monitoring containerized applications, including CPU and memory utilization at the task level for Amazon ECS with the Fargate launch type. It automatically collects, aggregates, and summarizes metrics and logs from containerized applications, offering out-of-the-box visualizations without requiring custom setup.

Exam trap

The trap here is that candidates may confuse CloudWatch Logs (which stores logs) with Container Insights (which provides pre-built dashboards and metrics), or assume custom metrics are required when a managed solution already exists.

How to eliminate wrong answers

Option B is wrong because Amazon CloudWatch custom metrics require manual creation and publishing of metrics via the PutMetricData API, which is not a pre-built dashboard solution and adds operational overhead. Option C is wrong because Amazon CloudWatch Logs is designed for log storage, search, and analysis, not for providing pre-built dashboards or metrics for CPU and memory utilization. Option D is wrong because Amazon CloudWatch Synthetics is used for monitoring application endpoints and APIs through canary tests, not for collecting or visualizing CPU and memory metrics from ECS tasks.

208
MCQhard

An organization has a production AWS environment with multiple VPCs and hundreds of EC2 instances. The security team wants to be alerted when any security group is modified. Which approach should a SysOps administrator use to meet this requirement with minimal overhead?

A.Enable CloudTrail and create a CloudWatch alarm for each security group modification event.
B.Use AWS Config rules to detect security group changes and trigger an SNS notification.
C.Enable VPC Flow Logs and analyze them with Amazon Athena for security group changes.
D.Deploy Amazon GuardDuty to monitor for security group modifications.
AnswerB

AWS Config records every change to security group resources as a configuration item and can evaluate those changes using managed or custom rules. When a rule marks a security group as noncompliant, AWS Config can send an alert through an SNS topic that you configure, giving you immediate notification of the modification. This approach both detects the change and enforces your security policies, unlike simply logging API activity or monitoring traffic.

Why this answer

AWS Config rules can continuously evaluate security group configurations against desired settings and trigger an SNS notification when a change is detected. This approach provides automated, event-driven monitoring with minimal operational overhead, as it does not require custom scripts or manual log analysis.

Exam trap

The trap here is confusing CloudTrail event monitoring with AWS Config's configuration change detection, leading candidates to choose CloudTrail-based alarms despite the higher overhead and lack of direct compliance evaluation.

How to eliminate wrong answers

Option A is wrong because CloudTrail logs API calls for security group modifications, but creating a CloudWatch alarm for each event would require custom metric filters and alarms, adding complexity and overhead; it also does not provide direct configuration compliance evaluation. Option C is wrong because VPC Flow Logs capture network traffic metadata (IP addresses, ports, protocols), not security group configuration changes, so they cannot detect modifications to security group rules. Option D is wrong because Amazon GuardDuty is a threat detection service that analyzes network and account activity for malicious behavior, not for tracking configuration changes like security group modifications.

209
Multi-Selectmedium

Which TWO actions should a SysOps admin take to troubleshoot an Amazon RDS instance that is experiencing high CPU utilization? (Choose 2.)

Select 2 answers
A.Enable Enhanced Monitoring to view OS-level metrics
B.Enable Multi-AZ to distribute the load
C.Delete the slow query log to free up CPU
D.Enable Performance Insights to identify high-load queries
E.Increase the instance size to reduce CPU utilization
AnswersA, D

Enhanced Monitoring on Amazon RDS delivers a stream of OS-level telemetry, including per-process CPU utilization, memory usage, disk I/O, and network traffic, via the CloudWatch console. This granular view lets you distinguish whether the high CPU is being consumed by the database engine itself or by auxiliary processes such as backups, log writers, or the OS, narrowing the root-cause search. It is a purely diagnostic action—it does not alter the workload but provides the raw data needed to formulate a targeted fix.

Why this answer

Enhanced Monitoring provides OS-level metrics (e.g., CPU, memory, disk I/O) for the RDS instance, which helps identify resource contention at the operating system level. This granularity is essential for diagnosing high CPU utilization that may be caused by OS processes (e.g., backup, patching) rather than database queries alone.

Exam trap

The trap here is that candidates often confuse Multi-AZ with read replicas, assuming it distributes load, or they think deleting logs frees CPU, when in fact logs are written asynchronously and have negligible CPU impact.

210
Multi-Selecthard

A company uses AWS CloudTrail to log API activity. The security team needs to be notified immediately when an IAM user creates a new access key. Which combination of steps should a SysOps administrator take? (Choose TWO.)

Select 2 answers
A.Create a CloudWatch Logs metric filter for the 'CreateAccessKey' event.
B.Enable AWS Config rules to detect changes to IAM users.
C.Configure CloudTrail to send notifications directly to Amazon SNS.
D.Create a CloudWatch alarm based on the metric filter and publish to an SNS topic.
E.Create an Amazon EventBridge rule to match the 'CreateAccessKey' API call.
AnswersA, D

A CloudTrail trail that is integrated with CloudWatch Logs streams each API event as a JSON log record. A metric filter using a pattern such as `{ $.eventName = "CreateAccessKey" }` scans every incoming log event and increments a CloudWatch metric whenever a new access key is created. This is the core detection step because it transforms raw CloudTrail API activity into a trackable, numerical signal that downstream alarms can act on.

Why this answer

CloudWatch Logs metric filters allow you to extract specific patterns from CloudTrail log data, such as the 'CreateAccessKey' event, and convert them into a metric. This enables you to monitor for this specific API call and trigger an alarm when it occurs, meeting the requirement for immediate notification.

Exam trap

The trap here is that candidates often think CloudTrail can directly send notifications to SNS (Option C) or that AWS Config rules are suitable for real-time event-driven alerts (Option B), but neither is correct for immediate notification of a specific API call.

211
MCQmedium

An application running on an EC2 instance writes logs to a local file. The operations team needs to monitor these logs in near real-time for troubleshooting. Which solution provides the most efficient way to stream these logs to CloudWatch Logs?

A.Use the AWS CLI to periodically upload the log file using the put-log-events command.
B.Install the CloudWatch Logs agent on the instance and configure it to tail the log file.
C.Install the Amazon Kinesis Agent on the instance and configure it to send logs to CloudWatch Logs.
D.Configure the application to write logs to an S3 bucket and use S3 Event Notifications to trigger a Lambda function that puts logs to CloudWatch.
AnswerB

The CloudWatch Logs agent (or the newer unified CloudWatch agent) runs as a daemon on the EC2 instance and uses the `tail` mechanism to monitor the specified log file, each time a new log line is written it is picked up and pushed to CloudWatch Logs in near real time. The agent manages checkpointing, so if the process restarts it can resume from the last-read position without re-sending old lines or losing new ones, and it also handles batching and `put-log-events` calls to the CloudWatch Logs API automatically. This is the purpose-built solution for streaming application logs to CloudWatch and is the correct choice for a near real-time requirement.

Why this answer

The CloudWatch Logs agent (or the newer unified CloudWatch agent) is designed specifically to tail log files from EC2 instances and stream them to CloudWatch Logs in near real-time. This provides the most efficient solution because it continuously monitors the file for new entries and sends them with minimal latency, without requiring periodic uploads or complex event-driven pipelines.

Exam trap

The trap here is that candidates may confuse the Kinesis Agent with the CloudWatch Logs agent, assuming both can send directly to CloudWatch Logs, but the Kinesis Agent only supports Kinesis destinations natively.

How to eliminate wrong answers

Option A is wrong because using the AWS CLI to periodically upload logs via put-log-events introduces significant latency (since it must run on a schedule) and is inefficient for near real-time monitoring; it also requires manual scripting to track the last uploaded position. Option C is wrong because the Amazon Kinesis Agent is designed to send data to Amazon Kinesis Data Streams or Firehose, not directly to CloudWatch Logs; sending logs to CloudWatch would require an additional intermediary (e.g., a Lambda function), adding complexity and cost. Option D is wrong because writing logs to S3 and using S3 Event Notifications with Lambda adds unnecessary latency (S3 is object storage, not a streaming target) and complexity; this approach is better suited for batch or archival processing, not near real-time monitoring.

212
MCQhard

A SysOps administrator needs to monitor a custom application metric 'OrdersPerMinute' published to Amazon CloudWatch. The metric should trigger an alarm when the count exceeds 100 for more than 2 consecutive data points, but only during business hours (9 AM to 5 PM weekdays). The alarm must evaluate the metric as a rate per minute. How should the administrator configure the alarm?

A.Create a CloudWatch alarm with a period of 1 minute, evaluation periods of 2, datapoints to alarm of 2, and use a math expression to filter time range.
B.Create a CloudWatch alarm with a period of 1 minute, evaluation periods of 2, datapoints to alarm of 2, and disable the alarm outside business hours using a Lambda function triggered by CloudWatch Events.
C.Create a CloudWatch alarm with a period of 1 minute, evaluation periods of 2, datapoints to alarm of 2, and use a metric math expression 'IF(IN_BUSINESS_HOURS(), OrdersPerMinute, 0)' but CloudWatch does not have IN_BUSINESS_HOURS function.
D.Create a CloudWatch alarm with a period of 1 minute, evaluation periods of 1, datapoints to alarm of 2 (impossible).
AnswerB

This solution uses a scheduled Lambda function (via CloudWatch Events) to enable/disable the alarm. The alarm itself is configured with the correct evaluation criteria (2 out of 2 datapoints above 100). This meets the requirement while using automated remediation.

Why this answer

To trigger when 'OrdersPerMinute' exceeds 100 for more than 2 consecutive data points, the alarm must require 3 consecutive breaching data points. This is achieved by setting evaluation periods to 3 and datapoints to alarm to 3. The Lambda/EventBridge approach to disable the alarm outside business hours is correct, but option B's evaluation period settings are incorrect.

As written, no option fully satisfies the requirement.

Exam trap

The trap here is that candidates assume CloudWatch has a built-in time-based filtering function (like IN_BUSINESS_HOURS) or that math expressions can evaluate time, when in reality AWS requires external scheduling via Lambda or EventBridge to manage alarm activation windows.

How to eliminate wrong answers

Option A is wrong because CloudWatch math expressions do not include a function like 'IN_BUSINESS_HOURS()' to filter by time range; math expressions operate on metric values, not time-based conditions. Option C is wrong because it incorrectly claims CloudWatch has an 'IN_BUSINESS_HOURS' function, which does not exist; this would cause the alarm to fail or evaluate incorrectly. Option D is wrong because it sets evaluation periods to 1 and datapoints to alarm to 2, which is impossible—the number of datapoints to alarm cannot exceed the number of evaluation periods.

213
MCQmedium

An organization is using AWS CloudFormation to deploy infrastructure. The SysOps administrator needs to receive notifications when stack creation fails. What is the simplest way to achieve this?

A.Create a CloudWatch alarm on the 'StackCreationFailure' metric.
B.Provide an SNS topic ARN in the --notification-arns parameter when creating the stack.
C.Create an EventBridge rule that triggers on CloudFormation events.
D.Enable CloudTrail and create a metric filter for 'CreateStack' failures.
AnswerB

The correct approach is to specify an SNS topic ARN in the --notification-arns parameter when creating the stack (or the equivalent NotificationARNs property in the CloudFormation template). CloudFormation publishes all stack-level events, including CREATE_FAILED, to that SNS topic, and subscribers such as email, SMS, or Lambda functions receive immediate notifications. This native integration is purpose-built for CloudFormation lifecycle monitoring and requires no additional resources or custom logic.

Why this answer

The `--notification-arns` parameter in the AWS CLI `create-stack` command directly associates an SNS topic with the stack, causing CloudFormation to publish notifications for all stack events, including failures. This is the simplest method as it requires no additional services or configuration beyond specifying the SNS topic ARN at stack creation.

Exam trap

The trap here is that candidates often over-engineer the solution by choosing EventBridge or CloudTrail, missing the fact that CloudFormation has a built-in, one-step SNS notification feature that is the simplest and most direct way to receive stack failure alerts.

How to eliminate wrong answers

Option A is wrong because CloudFormation does not emit a 'StackCreationFailure' metric to CloudWatch; CloudWatch metrics for CloudFormation are limited to 'Drift' and 'ResourceCount', not stack status events. Option C is wrong because while an EventBridge rule can capture CloudFormation events, it requires additional setup to filter for stack creation failures and is not the simplest approach compared to directly using SNS notifications. Option D is wrong because enabling CloudTrail and creating a metric filter for 'CreateStack' failures is overly complex and indirect; CloudTrail logs API calls but does not natively trigger notifications without additional CloudWatch Alarms or EventBridge rules.

214
MCQeasy

A company wants to be able to query application logs in near real-time using a SQL-like syntax. Which AWS service should be used?

A.CloudWatch Logs Insights
B.CloudWatch Metrics Insights
C.CloudWatch Events
D.CloudWatch Logs subscription filters
AnswerA

CloudWatch Logs Insights is a purpose-built query engine for log data stored in CloudWatch Logs. It uses a SQL-like query language to parse, filter, aggregate, and visualize log events in near real-time. Queries can be run across one or more log groups and saved for reuse, making it the correct tool for directly querying application logs.

Why this answer

CloudWatch Logs Insights is the correct service because it enables interactive querying of log data stored in CloudWatch Logs using a purpose-built SQL-like query language. It is designed for ad-hoc analysis of logs in near real-time, allowing users to filter, aggregate, and visualize log events without needing to export data to another analytics platform.

Exam trap

The trap here is that candidates confuse CloudWatch Logs Insights with CloudWatch Metrics Insights, assuming both can query logs, but Metrics Insights only works with numeric metric data and cannot parse or search log message content.

How to eliminate wrong answers

Option B is wrong because CloudWatch Metrics Insights is used to query and analyze metric data (numerical time-series data), not log data, and does not support SQL-like syntax for log content. Option C is wrong because CloudWatch Events (now part of Amazon EventBridge) is a service for routing events to targets based on rules, not for querying log data with SQL-like syntax. Option D is wrong because CloudWatch Logs subscription filters are used to stream log data in real-time to other destinations (e.g., Lambda, Kinesis, Elasticsearch) for processing, but they do not provide a query interface or SQL-like syntax for interactive analysis.

215
MCQmedium

Refer to the exhibit. A Lambda function is unable to write logs to CloudWatch Logs. The IAM role attached to the Lambda function includes the policy shown. What is the issue?

A.The log group name is incorrect.
B.The policy effect is Deny.
C.The 'logs:PutLogEvents' action is not allowed.
D.The resource ARN does not include a log-stream component.
AnswerD

For PutLogEvents, IAM must authorize against the log-stream ARN, which follows the pattern arn:aws:logs:region:account-id:log-group:log-group-name:log-stream:log-stream-name. The policy's resource ARN stops at 'log-group:my-log-group' and lacks the ':log-stream:...' component, so the permission does not match the API call's resource. To allow writes, the resource must be 'arn:aws:logs:region:account-id:log-group:my-log-group:log-stream:*' or a specific stream name. This is the root cause of the Lambda function's inability to write logs.

Why this answer

The Lambda function's IAM policy grants permissions for `logs:CreateLogGroup`, `logs:CreateLogStream`, and `logs:PutLogEvents` on the resource ARN `arn:aws:logs:us-east-1:123456789012:log-group:/aws/lambda/MyFunction:*`. However, this ARN only specifies the log group and a wildcard for log streams, which is insufficient for the `logs:PutLogEvents` action. The `logs:PutLogEvents` action requires a resource ARN that includes a specific log-stream component (e.g., `arn:aws:logs:us-east-1:123456789012:log-group:/aws/lambda/MyFunction:log-stream:*`), because the API call targets a particular log stream within the log group.

Without this, the Lambda function cannot write logs, resulting in a permissions error.

Exam trap

The trap here is that candidates assume a wildcard on the log group ARN (e.g., `log-group:*`) covers all actions, but AWS requires the log-stream component for write operations like `PutLogEvents`, causing a subtle permissions failure.

How to eliminate wrong answers

Option A is wrong because the log group name `/aws/lambda/MyFunction` is the default naming convention for Lambda functions and is correct; the issue is not with the name but with the resource ARN structure. Option B is wrong because the policy effect is explicitly `Allow`, not `Deny`, and there is no `Deny` statement present to override permissions. Option C is wrong because the `logs:PutLogEvents` action is listed in the policy's `Action` array, so it is allowed; the problem is that the resource ARN does not match the required format for that action.

216
MCQhard

A company runs a critical web application on EC2 instances behind an Application Load Balancer (ALB) in an Auto Scaling group. The application experiences intermittent latency spikes. The SysOps administrator has enabled detailed CloudWatch metrics on the ALB and the EC2 instances. The administrator notices that during the latency spikes, the ALB's TargetResponseTime metric increases, but the EC2 instance's CPU utilization and memory usage remain normal. The administrator also observes that the number of concurrent connections to the ALB spikes during these periods. Which action should the administrator take to identify the root cause?

A.Analyze the ALB's ActiveConnectionCount and RequestCountPerTarget metrics to see if the load balancer is reaching its connection limit.
B.Check the RDS database's DatabaseConnections metric to see if the database is overwhelmed.
C.Enable VPC Flow Logs and analyze the traffic patterns for dropped packets.
D.Enable detailed monitoring on the EC2 instances to capture CPU credits for burstable instances.
AnswerA

The ALB's ActiveConnectionCount metric tracks the total number of concurrent TCP connections (including idle keep-alive connections) flowing through the load balancer, while RequestCountPerTarget measures the number of requests distributed to each backend instance. If the ALB approaches its per-node connection ceiling or the targets reach their connection backlog, new requests can queue in the ALB's network stack, producing latency even when CPU utilization is low. This is the correct first diagnostic step because the symptom is at the application entry point, and these metrics directly reveal whether the load balancer or targets are saturated with connections rather than compute capacity.

Why this answer

The ALB's ActiveConnectionCount and RequestCountPerTarget metrics directly indicate whether the load balancer is approaching its connection limit (default 50,000 for ALBs). During latency spikes, if concurrent connections spike but instance CPU/memory are normal, the bottleneck is likely at the load balancer level, not the instances. Analyzing these metrics helps determine if the ALB is queuing or dropping requests due to connection limits, causing increased TargetResponseTime.

Exam trap

The trap here is that candidates assume latency spikes always indicate backend instance issues (CPU/memory) and overlook the ALB's connection limits, which can cause increased TargetResponseTime even when instances are underutilized.

How to eliminate wrong answers

Option B is wrong because the question states CPU and memory on EC2 instances remain normal, and there is no mention of database-related latency or errors; checking RDS DatabaseConnections would only be relevant if the application was database-bound, which is not indicated. Option C is wrong because VPC Flow Logs capture network traffic metadata (source/destination IPs, ports, protocol, packets) but do not measure ALB connection limits or application-layer latency; dropped packets would indicate network issues, not ALB connection saturation. Option D is wrong because detailed monitoring on EC2 instances provides 1-minute metrics (vs. 5-minute default) but does not expose CPU credit exhaustion for burstable instances; the question already has normal CPU/memory, so this would not identify the root cause of ALB-level connection spikes.

217
MCQeasy

A SysOps administrator needs to monitor the CPU utilization of an Amazon RDS DB instance and receive an alarm when CPU utilization exceeds 80% for 5 consecutive minutes. Which AWS service should be used to create this alarm?

A.AWS CloudTrail
B.Amazon CloudWatch
C.AWS Config
D.AWS Trusted Advisor
AnswerB

Amazon CloudWatch is the native monitoring service for RDS, automatically collecting the CPUUtilization metric at one-minute or five-minute granularity depending on the database instance class. You can create an alarm that evaluates the metric over consecutive periods and then publishes to an SNS topic or invokes an action when the threshold is breached. This makes CloudWatch the appropriate tool for monitoring CPU utilization and alerting.

Why this answer

Amazon CloudWatch is the native AWS monitoring service that can track RDS DB instance metrics, such as CPU utilization, and trigger alarms based on thresholds and time periods. In this scenario, you would create a CloudWatch alarm on the `CPUUtilization` metric for the specific DB instance, with a threshold of 80% and a period of 5 consecutive minutes (e.g., 5 datapoints of 1-minute periods).

Exam trap

The trap here is that candidates often confuse CloudTrail (for auditing API calls) with CloudWatch (for monitoring metrics and logs), leading them to select CloudTrail when the question explicitly asks about creating an alarm on a performance metric.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail records API activity and governance events, not real-time performance metrics like CPU utilization; it cannot create metric alarms. Option C is wrong because AWS Config evaluates resource configurations against rules and compliance standards, not operational metrics; it cannot monitor CPU utilization or trigger alarms based on threshold breaches. Option D is wrong because AWS Trusted Advisor provides best-practice recommendations and cost optimization checks, but it does not support custom metric alarms or real-time monitoring of CPU utilization.

218
MCQeasy

An EC2 instance runs a Java application. The operations team wants to monitor heap memory utilization in CloudWatch and set alarms when it exceeds 85 percent. EC2 does not natively publish memory metrics to CloudWatch. What is the simplest way to get this metric into CloudWatch?

A.Install the CloudWatch agent on the instance and configure it to collect mem_used_percent; publish JVM heap metrics from the application using PutMetricData
B.Enable detailed monitoring on the EC2 instance to increase metric resolution to 1-minute intervals
C.Configure a CloudWatch Logs metric filter on the application log stream to count lines containing 'OutOfMemoryError'
D.Use AWS Systems Manager Inventory to collect memory data and sync it to CloudWatch
AnswerA

The CloudWatch agent handles OS-level memory automatically once configured. For JVM heap, the application publishes a custom namespace metric via PutMetricData. Both appear in CloudWatch within minutes and can be graphed and alarmed like any native metric.

Why this answer

The CloudWatch agent can collect custom metrics like memory utilization from the EC2 instance, and the Java application can directly publish JVM heap metrics to CloudWatch using the PutMetricData API. This combination provides the simplest and most direct way to monitor heap memory utilization and set alarms at the 85% threshold, as EC2 does not natively expose memory metrics.

Exam trap

The trap here is that candidates often assume detailed monitoring or Systems Manager Inventory can provide memory metrics, but neither feature collects or publishes memory utilization data to CloudWatch.

How to eliminate wrong answers

Option B is wrong because enabling detailed monitoring increases the resolution of standard EC2 metrics (like CPU, network) to 1-minute intervals, but it does not add memory or JVM heap metrics, which are not published by EC2 at all. Option C is wrong because a CloudWatch Logs metric filter on 'OutOfMemoryError' only detects when the application has already crashed, not proactive heap utilization levels, and it cannot measure the percentage of heap memory used. Option D is wrong because AWS Systems Manager Inventory collects software inventory and configuration data, not real-time memory utilization metrics, and it does not sync data to CloudWatch as a metric for alarm purposes.

219
MCQmedium

A company uses AWS CloudTrail to log all management events. The SysOps administrator needs to be notified when an IAM user creates a new access key. Which configuration is the MOST efficient?

A.Create a CloudWatch Logs metric filter on the CloudTrail log group for CreateAccessKey events and set an alarm.
B.Use AWS Trusted Advisor to check for excessive access keys.
C.Create a CloudWatch Events rule that matches the CreateAccessKey API call and triggers an SNS notification.
D.Use AWS Config to monitor the 'iam-user' resource type for changes to access keys.
AnswerC

CloudWatch Events (now Amazon EventBridge) directly intercepts the CloudTrail API event as soon as it occurs. By defining a rule with an event pattern that filters the 'CreateAccessKey' API call and linking it to an SNS topic, the architecture provides a real-time, serverless notification pipeline. This is the simplest and most efficient way to alert on security-sensitive IAM actions because it consumes the event stream directly without requiring log parsing or polling.

Why this answer

CloudWatch Events (now Amazon EventBridge) can directly match the CreateAccessKey API call from CloudTrail in real time and trigger an SNS notification. This is the most efficient solution as it requires no additional log analysis or polling, and it reacts immediately when the API call occurs.

Exam trap

The trap here is that candidates often default to CloudWatch Logs metric filters (Option A) because they are familiar with log-based monitoring, but they overlook the more efficient and real-time event-driven approach using CloudWatch Events/EventBridge for API call notifications.

How to eliminate wrong answers

Option A is wrong because it requires creating a metric filter on CloudTrail logs in CloudWatch Logs, which introduces latency from log ingestion and metric evaluation, and is less efficient than a direct event-driven approach. Option B is wrong because AWS Trusted Advisor checks for excessive access keys based on best practices (e.g., keys older than 90 days), not for the creation event itself, so it cannot provide real-time notification when a new key is created. Option D is wrong because AWS Config monitors configuration changes and can detect access key modifications, but it is designed for compliance and resource tracking, not for real-time event notification, and it would require additional setup to trigger a notification.

220
Multi-Selectmedium

A SysOps administrator is setting up centralized logging for multiple AWS accounts using CloudWatch Logs. Which TWO actions should the administrator take to ensure that logs from all accounts are aggregated in a single account?

Select 2 answers
A.In the central account, create an IAM role that trusts the source accounts and allows PutLogEvents.
B.In the central account, create a CloudWatch Logs destination and attach a resource policy that grants the source accounts permission to write logs.
C.In each source account, configure a subscription filter on the log groups to send log events to the central account's CloudWatch Logs destination.
D.In the central account, create a log group with the same name as the source accounts' log groups.
E.In each source account, create a Kinesis Data Firehose delivery stream that sends logs to the central account's S3 bucket.
AnswersB, C

This is the correct approach because a CloudWatch Logs destination is a logical target that can receive log events from other accounts. The central account must attach a resource-based policy to the destination, explicitly listing the source account IDs and allowing them to call PutLogEvents. This policy grants the necessary write access without requiring cross-account IAM roles or complex key management, enabling centralized aggregation of logs from multiple source accounts.

Why this answer

A CloudWatch Logs destination in the central account, combined with a resource policy that grants the source accounts permission to write logs, is the standard mechanism for cross-account log aggregation. The destination acts as a target for subscription filters, and the resource policy explicitly allows the source accounts to call the PutLogEvents API against that destination. This setup ensures that log events from source accounts are delivered to the central account without requiring IAM roles or additional infrastructure.

Exam trap

The trap here is that candidates often confuse IAM cross-account roles with CloudWatch Logs destinations, assuming that a role with PutLogEvents permissions is sufficient, when in fact CloudWatch Logs requires a destination resource policy for cross-account delivery.

221
MCQhard

A SysOps administrator manages a fleet of EC2 instances that run a batch processing job. The job runs every hour and takes about 45 minutes to complete. The administrator wants to be notified if any job takes longer than 1 hour. Currently, the administrator uses CloudWatch Logs to capture job start and end times from application logs. The job writes a log message at start with 'JOB_START' and at end with 'JOB_END'. The administrator wants to create a metric filter that counts jobs that exceed 1 hour. However, the administrator is unsure how to achieve this with CloudWatch Logs. What should the administrator do?

A.Use CloudWatch Logs Insights to run a query every hour and check the duration.
B.Use CloudWatch Events to capture the log events and trigger a Lambda function to compute duration.
C.Create a metric filter that extracts the timestamp of JOB_START and JOB_END and computes the duration in a custom metric.
D.Create a Lambda function that is triggered by S3 to process the logs and publish a custom metric.
AnswerB

CloudWatch Events (EventBridge) can deliver CloudWatch Log events to a Lambda function in near real-time via a subscription filter, enabling event-driven processing. The Lambda function can parse the JOB_START and JOB_END entries, correlate them by job ID, calculate the duration, and publish a custom metric or trigger an alarm. This serverless architecture avoids polling and reacts immediately to each logged job, making it the recommended pattern.

Why this answer

CloudWatch Events (now part of Amazon EventBridge) can capture log events in real-time and trigger a Lambda function. The Lambda function can then compute job duration by correlating JOB_START and JOB_END events (e.g., using a DynamoDB table to store start times) and publish a custom metric or trigger an alarm if duration exceeds 1 hour. This approach handles the per-job correlation that metric filters cannot achieve.

Exam trap

Candidates often think metric filters can compute duration by extracting timestamps from JOB_START and JOB_END, but metric filters operate on individual log events and cannot correlate two events for the same job. The correct solution uses CloudWatch Events with Lambda for stateful computation.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs Insights is a query-based analysis tool for ad-hoc or scheduled queries, but it cannot directly trigger alarms or continuously monitor for durations exceeding 1 hour without custom scripting and additional services. Option B is wrong because CloudWatch Events (now Amazon EventBridge) can capture log events and trigger a Lambda function, but this approach adds unnecessary complexity and cost compared to a native metric filter, and it requires custom code to compute duration and publish metrics. Option D is wrong because S3 is not involved in the described workflow; the logs are in CloudWatch Logs, not S3, and using S3 triggers would require exporting logs to S3 first, adding latency and complexity.

222
MCQhard

A company runs a multi-tier application that uses an Amazon RDS for PostgreSQL database. The SysOps administrator needs to monitor the database for performance anomalies, such as sudden spikes in connections or query latencies. The administrator wants to receive alerts when metrics deviate from their expected baseline. The solution must automatically adjust to changes in normal behavior over time, such as seasonal patterns. Which AWS service or feature should the administrator use?

A.Configure Amazon CloudWatch Anomaly Detection on the relevant RDS metrics (e.g., DatabaseConnections, ReadLatency, WriteLatency) and set an alarm to notify when the metric breaches the anomaly band.
B.Use Amazon RDS Performance Insights to analyze database load and set CloudWatch alarms on the DBLoad metric with static thresholds.
C.Enable Amazon CloudWatch Metrics Explorer to create a dashboard that visualizes the metrics and manually review for anomalies.
D.Use AWS X-Ray to trace database queries and set alarms on trace segment durations.
AnswerA

CloudWatch Anomaly Detection uses machine learning to automatically model the expected patterns of RDS metrics such as DatabaseConnections, ReadLatency, and WriteLatency, including daily and weekly seasonal trends. It builds a dynamic baseline band around the metric's normal behavior and can trigger an alarm when data points breach that band, with no need to manually define static thresholds. The alarm action can notify via SNS, providing the automated, adaptive monitoring required to detect unusual RDS behavior without operator intervention.

Why this answer

Amazon CloudWatch Anomaly Detection uses machine learning to continuously analyze metric patterns and establish a dynamic baseline that adapts to seasonal trends and gradual changes in normal behavior. By applying anomaly detection to RDS metrics like DatabaseConnections, ReadLatency, and WriteLatency, the administrator can set an alarm that triggers when a metric deviates outside the calculated anomaly band, automatically adjusting to evolving traffic patterns without manual threshold updates.

Exam trap

The trap here is that candidates often confuse Performance Insights (a diagnostic tool for analyzing database load) with a monitoring and alerting solution, overlooking that it does not provide adaptive baselines or automatic anomaly detection.

How to eliminate wrong answers

Option B is wrong because RDS Performance Insights provides database load analysis and the DBLoad metric, but it relies on static thresholds for CloudWatch alarms, which cannot automatically adapt to changing baselines or seasonal patterns. Option C is wrong because CloudWatch Metrics Explorer is a visualization and query tool for exploring metrics, not a monitoring or alerting feature; it requires manual review and does not provide automated anomaly detection or adaptive baselines. Option D is wrong because AWS X-Ray is designed for tracing and analyzing application requests end-to-end, not for monitoring database-level metrics like connection counts or query latencies, and it cannot set alarms on RDS performance metrics.

223
MCQeasy

Refer to the exhibit. A SysOps administrator runs the command shown to investigate a CloudWatch alarm named 'HighCPU'. What does the output indicate?

A.The alarm entered the ALARM state and then returned to OK.
B.The alarm was deleted and recreated.
C.The alarm is currently in INSUFFICIENT_DATA state.
D.The alarm never entered the ALARM state.
AnswerA

CloudWatch alarm history records each state transition with a timestamp. In the exhibit, the alarm changed from OK to ALARM at 10:25 and then changed back to OK afterward, which is exactly what the history shows. Therefore, the correct interpretation is that the alarm entered the ALARM state and subsequently returned to OK.

Why this answer

The output shows two state transition datapoints: one at timestamp 2021-03-15T10:30:00Z with 'oldState' OK and 'newState' ALARM, and another at 2021-03-15T10:35:00Z with 'oldState' ALARM and 'newState' OK. This sequence confirms the alarm entered the ALARM state and then returned to OK, which is exactly what the describe-alarm-history command reveals when an alarm has experienced a full ALARM-to-OK cycle.

Exam trap

The trap here is that candidates may misinterpret the two datapoints as separate unrelated events rather than recognizing them as a complete ALARM-to-OK cycle, leading them to incorrectly choose that the alarm never entered ALARM state or that it was recreated.

How to eliminate wrong answers

Option B is wrong because deleting and recreating an alarm would produce a new alarm name or ARN, and the history would show a creation event, not a transition from OK to ALARM and back to OK. Option C is wrong because INSUFFICIENT_DATA state would appear as a transition from OK to INSUFFICIENT_DATA or ALARM to INSUFFICIENT_DATA, but the output only shows transitions between OK and ALARM. Option D is wrong because the first datapoint explicitly shows a transition from OK to ALARM, proving the alarm did enter the ALARM state.

224
MCQeasy

A SysOps administrator wants to receive a notification when an EC2 instance's status check fails. Which AWS service should be used to achieve this?

A.Amazon CloudWatch Alarms
B.AWS Config
C.AWS CloudTrail
D.AWS Trusted Advisor
AnswerA

Amazon CloudWatch Alarms is the correct service because it directly consumes the EC2 StatusCheckFailed metric, which is emitted every minute by the instance hypervisor. You can configure an alarm on this metric with a threshold (e.g., >=1 for one or more consecutive evaluation periods) to transition to ALARM state, and then invoke an SNS topic to send notifications via email, SMS, or Lambda. CloudWatch also supports separate alarms for StatusCheckFailed_System (host-level issues) and StatusCheckFailed_Instance (guest-OS level issues), giving you granular, near-real-time health monitoring.

Why this answer

Amazon CloudWatch Alarms can monitor EC2 instance status checks (both system and instance checks) and trigger an action, such as sending a notification via Amazon SNS, when a status check fails. This is the native AWS service designed for real-time monitoring and alerting on metric thresholds, making it the correct choice for this use case.

Exam trap

The trap here is that candidates often confuse AWS Config (which evaluates configuration compliance) with CloudWatch Alarms (which monitor metric thresholds), leading them to select AWS Config for real-time health alerts instead of the correct monitoring service.

How to eliminate wrong answers

Option B (AWS Config) is wrong because it is used for evaluating and recording resource configurations against desired policies, not for monitoring real-time status check failures. Option C (AWS CloudTrail) is wrong because it captures API activity and management events, not instance-level health metrics like status checks. Option D (AWS Trusted Advisor) is wrong because it provides best-practice recommendations and cost optimization checks, not real-time monitoring or alerting on EC2 status checks.

225
Multi-Selectmedium

A SysOps administrator needs to monitor the disk space utilization on a fleet of EC2 instances running Windows Server. Which TWO steps should the administrator take to collect and visualize this data? (Choose TWO.)

Select 2 answers
A.Enable detailed monitoring on the EC2 instances.
B.Install the CloudWatch agent on each EC2 instance to collect disk space metrics.
C.Use AWS CloudTrail to log disk space changes.
D.Enable default EC2 monitoring to collect disk space metrics automatically.
E.Create a CloudWatch dashboard to visualize the disk space metrics.
AnswersB, E

The CloudWatch agent is the correct solution because it runs inside the guest OS and directly reads file system utilization from the instance. You configure the agent with a JSON file to collect custom metrics such as disk_space_used, disk_space_free, and disk_space_utilization for each mount point, and it publishes these to CloudWatch under the custom namespace. The agent can be installed via Systems Manager or user data, and it also supports memory and log collection, making it the comprehensive approach for OS-level monitoring including disk space.

Why this answer

The CloudWatch agent is required to collect custom metrics like disk space utilization from EC2 instances running Windows Server. Default EC2 monitoring only collects hypervisor-level metrics (CPU, network, disk I/O), not guest OS metrics such as disk space. Installing the CloudWatch agent and configuring it to collect disk space metrics is the correct step.

Creating a CloudWatch dashboard then allows visualization of those collected metrics.

Exam trap

The trap here is that candidates assume default or detailed EC2 monitoring includes guest OS metrics like disk space, when in fact those metrics require the CloudWatch agent to be installed and configured on the instance.

← PreviousPage 3 of 4 · 250 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Monitoring, Logging, and Remediation questions.