Courseiva

CCNA Monitoring and Logging Questions

75 of 197 questions · Page 1/3 · Monitoring and Logging · Answers revealed

1
MCQeasy

A company wants to visualize the performance of their application running on EC2. They need to create a dashboard that shows CPU utilization, memory usage, and disk I/O. Which AWS service should they use?

A.Amazon CloudWatch Dashboards.
B.AWS CloudTrail.
C.AWS Systems Manager.
D.Amazon QuickSight.
AnswerA

Amazon CloudWatch Dashboards is the correct choice because it natively integrates with CloudWatch, which collects and stores EC2 standard metrics (CPU, network, disk) and, when the CloudWatch agent is installed on the instance, also collects custom operating-system-level metrics such as memory utilization and disk space. These dashboards can be customized to display multiple metrics in a single graph, aggregated across instances, and automatically refreshed for near-real-time monitoring. Additionally, CloudWatch Dashboards support alarm overlays and anomaly detection, making it purpose-built for visualizing application performance as a monitoring and observability tool.

Why this answer

Amazon CloudWatch Dashboards is the correct choice because it is the native AWS service for visualizing operational metrics such as CPU utilization, memory usage, and disk I/O. CloudWatch collects these metrics from EC2 instances (with the CloudWatch agent required for memory and disk metrics at the OS level) and allows you to create customizable, shareable dashboards. It supports cross-account, cross-region views and integrates with alarms and logs, making it the standard tool for performance monitoring.

The other services do not provide metric visualization for EC2 performance.

Exam trap

DOP-C02 often tests the distinction between monitoring (CloudWatch) and auditing (CloudTrail), or between operational dashboards and business intelligence (QuickSight), so candidates must remember that memory and disk metrics require the CloudWatch agent, not just default EC2 monitoring.

How to eliminate wrong answers

Option B is wrong because AWS CloudTrail records API activity and audit logs, not performance metrics like CPU or memory; it is used for governance, compliance, and operational auditing. Option C is wrong because AWS Systems Manager is a management suite for patching, automation, and inventory, not a dashboarding or metric visualization service, though it can collect some inventory data. Option D is wrong because Amazon QuickSight is a business intelligence service for analyzing data from various sources, not a real-time operational monitoring tool for EC2 metrics; it lacks native integration with CloudWatch metrics for live dashboards.

2
MCQhard

An e-commerce application runs on Amazon ECS with Fargate. The operations team notices that the application's latency increases during peak hours. The engineer needs to correlate high CPU usage with increased request latency to identify the root cause. Which approach should be used?

A.Use CloudWatch Logs Insights to query container logs
B.Enable Container Insights and ServiceLens to correlate metrics and traces
C.Configure CloudWatch Synthetics canaries to measure latency
D.Set up a Prometheus server on an EC2 instance to scrape container metrics
AnswerB

For Fargate, Container Insights must be enabled to collect infrastructure telemetry such as CPU and memory utilization at the cluster, service, task, and container levels; AWS then ships these metrics to CloudWatch automatically. ServiceLens builds on that by combining CloudWatch metrics, logs, and X-Ray traces into a single service map and console, allowing you to select a trace and see the corresponding CPU/memory metric trends for that underlying container. This native integration is exactly what is required to prove whether a CPU saturation event is causing an observed latency increase in a traced request.

Why this answer

Amazon CloudWatch Container Insights collects, aggregates, and summarizes metrics and logs from containerized applications on ECS/Fargate, including CPU utilization. ServiceLens integrates CloudWatch with AWS X-Ray traces, allowing you to correlate metrics with request traces and latency. Together they provide the end-to-end visibility needed to link high CPU usage with increased request latency.

Exam trap

DOP-C02 often tests whether candidates confuse monitoring tools that collect metrics (Container Insights) with those that trace requests (X-Ray/ServiceLens) or measure external latency (Synthetics), missing the need for correlation.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs Insights queries log data but does not natively correlate metrics with distributed traces or provide the integrated latency-to-CPU correlation. Option C is wrong because Synthetics canaries measure endpoint latency from outside but do not correlate that latency with internal CPU metrics or traces. Option D is wrong because a self-managed Prometheus server adds operational overhead and does not natively integrate with X-Ray traces or CloudWatch ServiceLens for correlation.

3
MCQhard

Refer to the exhibit. A security team reviews this CloudTrail log entry. Which finding is most concerning?

A.The event occurred in us-east-1.
B.The instance was terminated by an assumed role.
C.The source IP is from a public IP.
D.The user did not authenticate with MFA.
AnswerD

The absence of MFA in a CloudTrail session is a direct security-control failure because AWS relies on multi-factor authentication as a second factor to verify the caller's identity beyond long-term credentials. For sensitive actions like terminating an EC2 instance, missing MFA increases the risk that the request came from compromised keys or a stolen session. This makes the MFA status the key anomaly that warrants investigation, as it violates the security best practice of enforcing MFA for destructive operations.

Why this answer

The session was created without MFA (mfaAuthenticated: false). This is a security concern because the role allows console access and the user did not use MFA, increasing risk of unauthorized access. The termination is the action, but the lack of MFA is a security gap.

4
MCQmedium

A company is running a production application on Amazon ECS with AWS Fargate. The application has unpredictable traffic patterns and occasionally experiences increased latency. The DevOps team needs to configure scaling based on a custom metric that tracks the number of active user sessions in real time. Which solution will allow the team to scale the ECS service based on this custom metric?

A.Use AWS Auto Scaling to scale the ECS service based on the custom metric.
B.Create a CloudWatch dashboard to visualize the metric and manually adjust the service count.
C.Publish the custom metric to Amazon CloudWatch, then create a target tracking scaling policy in Application Auto Scaling for the ECS service.
D.Use an AWS Lambda function to directly update the desired count of the ECS service based on the metric.
AnswerC

The recommended approach is to first publish the custom application metric to Amazon CloudWatch using the PutMetricData API, ensuring it has the appropriate ECS service dimensions. Then, in Application Auto Scaling, create a target tracking scaling policy that references that custom metric; the policy automatically calculates the required desired count to maintain the target value, and adjusts the ECS service accordingly. This provides native, automated scaling with built-in cooldowns and no additional compute, making it the most operationally efficient solution.

Why this answer

Application Auto Scaling is the service that scales ECS services, and it supports target tracking policies based on custom CloudWatch metrics. Publishing the custom session-count metric to CloudWatch and then creating a target tracking scaling policy for the ECS service is the correct, fully managed approach.

Exam trap

DOP-C02 often tests the confusion between 'AWS Auto Scaling' and 'Application Auto Scaling' — candidates who pick the generic AWS Auto Scaling option miss that ECS services are scaled specifically through Application Auto Scaling with target tracking policies.

How to eliminate wrong answers

Option A is wrong because 'AWS Auto Scaling' is a separate service for scaling groups of related resources (EC2, DynamoDB, etc.) and is not the mechanism used to scale ECS service desired count — that is Application Auto Scaling. Option B is wrong because a CloudWatch dashboard is a visualization tool and manual adjustment does not meet the requirement for automatic scaling based on a custom metric. Option D is wrong because a Lambda function directly updating the desired count is a custom, unmanaged workaround that bypasses Application Auto Scaling's built-in cooldowns, alarms, and policy evaluation — it is not the recommended solution and adds operational overhead.

5
MCQmedium

A company is using Amazon CloudWatch Synthetics to monitor the availability of a web application. The canary runs every 5 minutes from multiple locations. Recently, the canary has been failing intermittently with HTTP 503 errors, but the application team reports that the application is healthy. Which step should the DevOps engineer take to identify the cause of the false positives?

A.Increase the canary timeout setting to allow more time for the application to respond.
B.Add more canary locations to increase coverage.
C.Review the canary's CloudWatch Logs to check for network errors or timeouts.
D.Increase the canary run frequency to every 1 minute.
AnswerC

The CloudWatch Synthetics canary's logs contain detailed request/response records, error stack traces, and screenshots captured during the run, making them essential for pinpointing whether a failure is caused by network errors, timeouts, or the application itself. By reviewing these logs, you can check for DNS resolution failures, TCP handshake timeouts, TLS certificate errors, or 4xx/5xx responses that indicate the problem is not with the client-side script. This is the first diagnostic step because it separates infrastructure-level issues from application-level defects.

Why this answer

Reviewing the canary's CloudWatch Logs is the correct first step because Synthetics canaries write detailed logs and screenshots for each run, including network errors, timeouts, and HTTP response details. Since the application team reports the app is healthy, the 503 errors are likely caused by the canary's network path, DNS resolution, or a transient issue that the logs will reveal. This is the diagnostic step that identifies the root cause of the false positives.

Exam trap

DOP-C02 often tests the instinct to 'fix' a monitoring problem by changing the monitor's configuration (timeout, frequency, locations) rather than investigating the logs to find the actual cause of false positives.

How to eliminate wrong answers

Option A is wrong because increasing the timeout does not diagnose the cause of 503 errors; if the app is healthy and responding quickly, a longer timeout will not fix a false positive and may mask the real issue. Option B is wrong because adding more canary locations increases coverage but does not explain why existing runs fail intermittently — it may even add more false positives. Option D is wrong because increasing run frequency to every 1 minute increases noise and cost without diagnosing the root cause, and could exacerbate throttling or rate-limit issues that cause 503s.

6
MCQmedium

A company runs an internal REST API on Amazon API Gateway backed by AWS Lambda. The operations team wants a single CloudWatch alarm that fires when the API's overall error rate exceeds 5% during any 5-minute period, without creating one alarm per resource. The API is deployed as a REST API (not HTTP API). Which approach should a DevOps engineer take?

A.Enable AWS X-Ray tracing on the stage and create a CloudWatch alarm on the X-Ray ErrorRate metric for the API.
B.Create a CloudWatch metric math alarm using the expression m1/m2 where m1 is the sum of the 4XXError and 5XXError metrics and m2 is the Count metric for the stage, all with period 300 and statistic Sum.
C.Create one CloudWatch alarm per method with a 5% threshold and combine them into a composite alarm using an OR rule.
D.Create a CloudWatch Logs metric filter on the API Gateway access logs and alarm on the resulting custom metric with a 5% threshold.
AnswerB

Metric math lets a single alarm combine several API Gateway metrics. Summing 4XXError and 5XXError and dividing by Count at a 300-second period yields the stage-level error rate, so one alarm covers the whole API rather than one alarm per method or resource. This directly satisfies the 5% threshold requirement.

Why this answer

API Gateway publishes 4XXError, 5XXError, and Count per stage. A metric math alarm can sum the error metrics and divide by Count at a five-minute period, producing a single stage-level error-rate alarm. Per-method alarms, X-Ray metrics, and log metric filters either fragment the view or do not directly yield a percentage.

Exam trap

The trap here is assuming API Gateway exposes a ready-made percentage error metric, when in fact you must compute the ratio yourself with metric math.

7
MCQmedium

A company uses AWS Lambda functions behind an Amazon API Gateway REST API. The DevOps team wants to monitor the end-to-end latency of API requests, including the time spent in API Gateway and Lambda. Which approach provides the most granular breakdown?

A.Enable Lambda Insights to get per-request latency breakdown.
B.Enable AWS X-Ray tracing on API Gateway and Lambda.
C.Enable VPC Flow Logs to capture network round-trip times.
D.Use CloudWatch metrics for API Gateway and Lambda, then add them together.
AnswerB

X-Ray propagates a trace ID through API Gateway and the downstream Lambda invocation, producing segment-level timings for each component. This granular breakdown isolates API Gateway overhead from Lambda init and execution duration, satisfying the requirement to measure end-to-end latency across both services.

Why this answer

AWS X-Ray provides end-to-end tracing with detailed segments for API Gateway and Lambda, allowing per-request breakdown of latency across each component. Option A (Lambda Insights) offers OS-level metrics like CPU and memory, not request-level latency breakdown. Option C (VPC Flow Logs) captures network traffic metadata but not application layer latency.

Option D (combining CloudWatch metrics) only gives aggregate statistics and cannot break down latency per request or within a single request's path.

8
MCQeasy

A DevOps team is deploying a new web application on AWS Elastic Beanstalk. They want to monitor the application's health and receive notifications when the environment's health status changes to 'Degraded' or 'Severe'. What is the simplest way to achieve this?

A.Use the Elastic Beanstalk management console to manually check the health status twice a day.
B.Create a CloudWatch alarm on the 'EnvironmentHealth' metric published by the Elastic Beanstalk environment.
C.Write a custom script that polls the Elastic Beanstalk DescribeEnvironmentHealth API and sends an email using Amazon SES.
D.Configure an AWS CloudTrail trail to monitor Elastic Beanstalk API calls and create a CloudWatch alarm on the trail.
AnswerB

Elastic Beanstalk with enhanced health reporting publishes the 'EnvironmentHealth' metric to CloudWatch, and creating an alarm on this metric allows you to react automatically when the environment transitions to states such as severe or degraded. The alarm can publish to an SNS topic to email or page the team, enabling real-time incident response without any custom code. This approach is the native, supported integration between Elastic Beanstalk health and CloudWatch observability.

Why this answer

Elastic Beanstalk automatically publishes environment health metrics to Amazon CloudWatch, including the 'EnvironmentHealth' metric which reports values like 0 (Ok), 1 (Info), 5 (Unknown), 10 (NoData), 15 (Warning), 20 (Degraded), and 25 (Severe). Creating a CloudWatch alarm on this metric with a threshold of >= 20 (or specifically for Degraded/Severe) and configuring an SNS notification is the simplest, most direct, and fully managed solution. This requires no custom code, no polling infrastructure, and leverages native AWS integration.

Exam trap

DOP-C02 often tests the difference between CloudWatch (metrics and alarms) and CloudTrail (API activity logging), so candidates may incorrectly choose CloudTrail for monitoring health status changes.

How to eliminate wrong answers

Option A is wrong because manual console checks are not automated, are error-prone, and cannot provide real-time notifications, violating the requirement for automatic alerts on health changes. Option C is wrong because writing a custom polling script with SES introduces unnecessary complexity, requires managing credentials, scheduling, and error handling, and is not the simplest approach when CloudWatch alarms already exist. Option D is wrong because CloudTrail records API calls for auditing, not environment health metrics; it does not capture the 'EnvironmentHealth' metric and cannot be used to alarm on health status changes.

9
Multi-Selecteasy

A DevOps engineer is setting up monitoring for an Amazon DynamoDB table that experiences high read traffic. They want to monitor the read capacity consumption and be alerted when the consumed read capacity exceeds 80% of the provisioned capacity for 5 consecutive minutes. Which TWO steps should they take? (Select TWO.)

Select 2 answers
A.Enable AWS CloudTrail to log DynamoDB read requests.
B.Set up an AWS Lambda function to monitor the DynamoDB ReadThrottleEvents metric.
C.Use CloudWatch to monitor the ConsumedReadCapacityUnits and ProvisionedReadCapacityUnits metrics.
D.Configure DynamoDB to stream all read events to CloudWatch Logs.
E.Create a CloudWatch alarm with a metric math expression that calculates (ConsumedReadCapacityUnits / ProvisionedReadCapacityUnits) and set the threshold to 0.8.
AnswersC, E

DynamoDB natively publishes ConsumedReadCapacityUnits and ProvisionedReadCapacityUnits to CloudWatch at a one-minute granularity. Monitoring these metrics directly gives you a continuous view of how much read capacity is being used relative to what the table has provisioned. This is the foundational data source for any read-capacity alarm or scaling policy, and it lets you detect capacity pressure before throttling begins.

Why this answer

CloudWatch directly exposes the ConsumedReadCapacityUnits and ProvisionedReadCapacityUnits metrics for DynamoDB, which are the exact metrics needed to calculate read capacity utilization. Monitoring these metrics allows the engineer to track how much of the provisioned capacity is being consumed over time, which is the foundation for setting up the desired alert.

Exam trap

The trap here is that candidates often confuse throttling metrics (like ReadThrottleEvents) with capacity utilization metrics, leading them to select Option B, which only detects throttling after it happens rather than providing a proactive alert based on capacity consumption.

10
MCQeasy

A DevOps engineer sets up a CloudWatch dashboard to monitor an application's performance. The application runs on EC2 instances in an Auto Scaling group. The engineer wants to display the average CPU utilization across all instances in the group. Which CloudWatch metric and statistic should be used?

A.CPUUtilization metric with the Sum statistic, filtered by Auto Scaling group.
B.CPUUtilization metric with the Average statistic, filtered by Auto Scaling group.
C.StatusCheckFailed metric with the Average statistic, filtered by Auto Scaling group.
D.NetworkOut metric with the Average statistic, filtered by Auto Scaling group.
AnswerB

The Average statistic applied to CPUUtilization with the AutoScalingGroupName dimension gives the mean CPU percentage across all instances in the group, which is exactly what the dashboard needs to assess overall load. CloudWatch computes this by taking the CPUUtilization data points from each instance and calculating the arithmetic mean over the selected period. This is the standard and CloudFormation-idiomatic approach for monitoring an Auto Scaling group's aggregate CPU usage, and it directly answers the requirement without mixing units or extrapolating to a nonsensical total.

Why this answer

To display the average CPU utilization across all instances in an Auto Scaling group, you should use the CPUUtilization metric with the Average statistic, filtered by the Auto Scaling group dimension. This aggregates the CPU utilization of all instances, providing a representative average.

Exam trap

DOP-C02 often tests the appropriate statistic for a given metric, and candidates may incorrectly choose Sum for utilization metrics, not realizing it produces a meaningless total.

How to eliminate wrong answers

Option A is wrong because the Sum statistic would add up CPU utilization percentages across instances, resulting in a meaningless total (e.g., 500% for 5 instances at 100%). Option C is wrong because StatusCheckFailed measures instance status checks, not CPU utilization, and is not the desired metric. Option D is wrong because NetworkOut measures network traffic, not CPU utilization.

11
MCQmedium

A company runs a serverless application using AWS Lambda, Amazon API Gateway, and Amazon DynamoDB. The application is used by thousands of users. Recently, the operations team noticed an increase in 5xx errors from API Gateway. The team has enabled CloudWatch Logs for the Lambda functions and API Gateway. They see the errors are sporadic and not correlated with high traffic. The Lambda function's error count in CloudWatch is also increasing. The team wants to identify the specific requests that are failing and understand the error details. Which solution should the team implement?

A.Use CloudWatch Logs Insights to query the Lambda logs for ERROR messages and correlate with API Gateway logs
B.Enable VPC Flow Logs for the Lambda function's VPC to capture network traffic
C.Enable AWS X-Ray active tracing on the Lambda functions and API Gateway to capture detailed request traces and error details
D.Enable AWS CloudTrail to log API Gateway API calls and analyze the logs
AnswerC

AWS X-Ray active tracing on both API Gateway and the Lambda functions provides end-to-end request tracing that captures every segment and subsegment — including Lambda invocation details, function execution time, HTTP status codes, exceptions, and downstream AWS service calls — allowing you to pinpoint the exact component that caused the error. When active tracing is enabled, API Gateway sends trace headers to Lambda, and Lambda automatically records the function's execution, with error details visible in the X-Ray console under the service map and trace list. This lets you filter traces by HTTP 5xx status or Lambda ERROR, drill into the affected trace to see the precise failing segment, and inspect the exception stack trace or error message, directly answering where and why the API request failed. X-Ray is purpose-built for distributed serverless applications, making it the most efficient and comprehensive option for this diagnostic task.

Why this answer

AWS X-Ray active tracing on Lambda and API Gateway captures end-to-end request traces, including downstream calls to DynamoDB, latency breakdowns, and exception details for each failing invocation. Because the errors are sporadic and not traffic-correlated, X-Ray's per-request trace view is the fastest way to pinpoint which specific requests fail and why. CloudWatch Logs alone would require manual correlation across services and lacks the trace context X-Ray provides.

Exam trap

DOP-C02 often tests the confusion between control-plane logging (CloudTrail), network logging (VPC Flow Logs), and request-level tracing (X-Ray), tempting candidates to pick CloudTrail or Flow Logs when the question asks for per-request error detail.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs Insights can query Lambda logs, but it cannot natively correlate a specific API Gateway request with its downstream Lambda and DynamoDB calls; the team would have to stitch together request IDs manually and still lack trace-level timing and error propagation detail. Option B is wrong because VPC Flow Logs capture IP-level network metadata (source/dest, ports, accept/reject), not application errors or Lambda exceptions, so they cannot explain 5xx responses. Option D is wrong because CloudTrail logs control-plane API calls (e.g., who changed a configuration), not data-plane request execution or Lambda runtime errors, so it will not reveal failing user requests.

12
MCQeasy

A company uses Amazon CloudWatch Logs to store application logs from EC2 instances. The security team requires that logs be retained for 5 years for compliance. Which action should be taken to meet this requirement cost-effectively?

A.Export the logs to Amazon S3 and use S3 Glacier Deep Archive for long-term storage.
B.Set a log retention policy of 5 years on the CloudWatch Logs log groups.
C.Disable log retention and let CloudWatch Logs keep the logs indefinitely.
D.Use AWS CloudTrail to store the logs for 5 years.
AnswerA

Exporting application logs from CloudWatch Logs to Amazon S3 and applying a lifecycle policy that transitions objects to S3 Glacier Deep Archive is the most cost-efficient way to meet a 5-year retention requirement. CloudWatch Logs storage costs roughly $0.03 per GB-month, while Glacier Deep Archive costs about $0.00099 per GB-month — a savings of nearly 97%. This makes sense for rarely accessed logs where the 12-hour retrieval time of Glacier Deep Archive is acceptable for compliance audits.

Why this answer

Exporting logs to Amazon S3 and using lifecycle policies to transition them to S3 Glacier Deep Archive is the most cost-effective way to retain logs for 5 years. S3 provides low-cost storage, and Glacier Deep Archive offers the lowest storage cost for long-term archival, making it ideal for compliance requirements. In contrast, retaining logs in CloudWatch Logs beyond a short period is more expensive due to higher storage costs.

Option B is incorrect because setting a 5-year retention policy in CloudWatch Logs incurs significant costs compared to exporting to S3. Option C is incorrect because disabling retention causes logs to expire after 90 days, not indefinite. Option D is incorrect because AWS CloudTrail captures API activity, not application logs.

13
MCQmedium

A company runs a serverless application using AWS Lambda, Amazon API Gateway, and Amazon DynamoDB. The application processes financial transactions. The DevOps team needs to monitor for duplicate transactions that could occur due to retries. The team wants to set up an alert when the number of duplicate transaction attempts exceeds 10 in a 5-minute window. The application logs each transaction attempt with a unique transaction ID to CloudWatch Logs. What is the most efficient way to achieve this?

A.Create a CloudWatch Logs metric filter that counts log events containing 'DuplicateTransaction' and set an alarm on the metric with a threshold of 10.
B.Use DynamoDB Streams to trigger a Lambda function that counts duplicates and publishes metrics.
C.Stream the CloudWatch Logs to Amazon Kinesis Data Analytics and use SQL queries to detect duplicates.
D.Modify the Lambda function to publish a custom metric to CloudWatch for each duplicate transaction, then set an alarm.
AnswerA

CloudWatch Logs metric filters evaluate log events in real time against a filter pattern and can emit a custom metric that increments for every matching event, here 'DuplicateTransaction'. This avoids changing the Lambda function or introducing additional services, and it leverages the transaction logs already being written to CloudWatch. A CloudWatch alarm can then monitor that metric over a 5-minute period and trigger when the count exceeds 10, providing an accurate and low-cost detection mechanism.

Why this answer

CloudWatch Logs metric filters can scan log events for specific terms like 'DuplicateTransaction' and convert matches into a custom metric. You can then create a CloudWatch alarm on that metric with a threshold of 10 over a 5-minute period. This requires no code changes, no additional services, and is the most efficient and native solution for monitoring log-based patterns.

Exam trap

DOP-C02 often tests the misconception that you need to modify application code or use additional services to monitor log patterns, when native CloudWatch Logs metric filters and alarms are sufficient and more efficient.

How to eliminate wrong answers

Option B is wrong because DynamoDB Streams capture item-level changes, not duplicate transaction attempts; duplicates are logged in CloudWatch Logs, not necessarily reflected as separate DynamoDB writes, and this adds unnecessary Lambda and Streams complexity. Option C is wrong because Kinesis Data Analytics is for real-time stream processing with SQL, which is overkill for simple counting and requires streaming logs to Kinesis, adding cost and latency. Option D is wrong because modifying the Lambda function to publish custom metrics requires code changes and deployment, and it still needs an alarm; it's less efficient than using existing logs with a metric filter.

14
MCQmedium

A DevOps engineer needs a centralized view of operational metrics and logs for applications running in three AWS Regions. The security team requires that the data be queryable together, that access be governed by existing IAM roles, and that no data be copied into a separate third-party account. Which approach should the engineer use?

A.Create an Amazon CloudWatch cross-account, cross-Region observability configuration in a central monitoring account and share the source accounts' metrics and log groups with it.
B.Use AWS CloudFormation StackSets to deploy identical CloudWatch dashboards and alarms into every Region and account, and have operators switch roles to view each one.
C.Enable AWS Config in all Regions and accounts and use the AWS Config advanced query feature to search metrics and logs.
D.Configure each application to write metrics and logs to an Amazon Kinesis Data Firehose delivery stream that lands data in a central Amazon S3 bucket, then query with Amazon Athena.
AnswerA

CloudWatch cross-account observability supports linking source accounts and Regions to a monitoring account, letting users query metrics, logs, and traces from one place while data stays in the source accounts. IAM roles control who can view the shared data, satisfying governance without copying data to a third party.

Why this answer

CloudWatch cross-account observability links source accounts and Regions to a monitoring account, enabling a single query experience for metrics, logs, and traces while data remains in place. IAM roles govern access, meeting the governance constraint. Data pipelines to S3, StackSet dashboards, and AWS Config do not provide the same in-place, unified query capability.

Exam trap

The trap here is assuming observability data must be copied to a central account, when CloudWatch can share it in place through a monitoring account link.

15
MCQmedium

A company uses Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer. The operations team notices that some instances are failing health checks but are not being terminated by Auto Scaling. What should be investigated to resolve this issue?

A.Confirm that the load balancer's health check target is pointing to the correct port and path on the instances.
B.Check the health check grace period setting in the Auto Scaling group. If it is too long, instances failing health checks may not be terminated quickly.
C.Ensure the instances are sending health check requests to the load balancer.
D.Verify that the security group for the instances allows inbound traffic from the load balancer on the health check port.
AnswerB

The health check grace period tells Auto Scaling how long to wait after an instance launches before it starts evaluating the instance's health status. If this value is set too high, an instance that is already failing health checks will continue running without being terminated because the elapsed time since launch has not yet exceeded the grace period. Extending the grace period can therefore leave unhealthy instances in service, which matches the described issue.

Why this answer

The health check grace period in an Auto Scaling group controls how long Auto Scaling waits before checking the health of a newly launched instance. If this grace period is set too long, instances that fail health checks before the grace period expires will not be terminated promptly, leading to the observed behavior where unhealthy instances remain in service even though they are failing health checks. The default grace period is 300 seconds, but if it is excessively long, it delays the termination of instances that fail health checks during that period.

Exam trap

The trap here is that candidates often confuse the health check grace period with the load balancer's health check settings, assuming that if health checks are failing, the issue must be with the health check configuration (Option A) rather than the Auto Scaling group's delay in acting on those failures.

How to eliminate wrong answers

Option A is wrong because the load balancer's health check target pointing to the correct port and path is essential for health checks to work, but the issue here is that instances are failing health checks and not being terminated, not that health checks are misconfigured. Option C is wrong because instances do not send health check requests to the load balancer; the load balancer sends health check requests to the instances, so this option describes a reverse and incorrect flow. Option D is wrong because while security group rules allowing inbound traffic from the load balancer on the health check port are necessary for health checks to succeed, the problem is that instances are failing health checks and not being terminated, not that health checks are failing due to blocked traffic.

16
MCQeasy

A company uses AWS CloudFormation to deploy infrastructure. The DevOps team wants to receive notifications when a stack fails to create or update. What is the MOST efficient way to achieve this?

A.Configure an SNS topic in the stack's notification options.
B.Create a custom resource in the stack that publishes to Amazon SNS.
C.Create a CloudWatch alarm on the StackStatus metric.
D.Use Amazon EventBridge to capture CloudFormation events and publish to SNS.
AnswerA

Configuring an SNS topic in the stack's notification options is the native CloudFormation mechanism for sending stack lifecycle events (e.g., CREATE_COMPLETE, UPDATE_FAILED, DELETE_IN_PROGRESS) to a pub/sub channel. You simply provide the topic ARN (up to five) when you create or update the stack, and CloudFormation publishes every stack event to it without requiring any Lambda code, custom resources, or additional infrastructure. This is the simplest and most direct way to get notified about stack changes.

Why this answer

AWS CloudFormation natively supports specifying Amazon SNS topic ARNs in the stack's notification options. When a stack operation (create, update, or delete) fails, CloudFormation automatically publishes a notification to the configured SNS topic without requiring any custom code, additional resources, or external event processing. This is the most efficient approach as it leverages built-in functionality with zero maintenance overhead.

Exam trap

The trap here is that candidates overthink the solution by choosing EventBridge or custom resources, missing the fact that CloudFormation has a built-in, one-step SNS notification feature that is both simpler and more reliable for failure alerts.

How to eliminate wrong answers

Option B is wrong because creating a custom resource in the stack to publish to SNS introduces unnecessary complexity, requires a Lambda function or other compute resource to handle the custom resource lifecycle, and does not reliably capture all stack failure scenarios (e.g., failures during stack creation before the custom resource is processed). Option C is wrong because CloudFormation does not emit a 'StackStatus' metric to CloudWatch; CloudWatch alarms cannot directly monitor CloudFormation stack status without custom metrics or log-based metrics, making this approach infeasible. Option D is wrong because while EventBridge can capture CloudFormation events (e.g., via AWS API calls or CloudTrail), this requires additional configuration, incurs EventBridge costs, and is less efficient than the native SNS notification option that requires no extra services or rules.

17
MCQmedium

A DevOps engineer executes the above CloudWatch Logs Insights query. What will the output contain?

A.The count of ERROR messages per 1-minute interval for the most recent 20 intervals
B.The count of ERROR messages per 5-minute interval for the most recent 20 intervals
C.The total count of ERROR messages in the log group
D.A list of the 20 most recent log entries that contain the word 'ERROR'
AnswerB

The query filters for events containing 'ERROR' and then performs stats count(*) by bin(5m), which aggregates the matching events into consecutive 5-minute time buckets. The results are ordered by bucket timestamp in descending order and limited to 20 rows, so the output is exactly the count of ERROR messages for each of the 20 most recent 5-minute intervals. The bin(5m) function and the LIMIT clause together determine the time span and number of rows returned.

Why this answer

The CloudWatch Logs Insights query uses 'stats count(*) by bin(5m)' which aggregates log events into 5-minute buckets, and 'limit 20' restricts the output to the 20 most recent buckets. Therefore the output is a count of ERROR messages per 5-minute interval for the most recent 20 intervals. The bin() function defines the time bucket size, and the limit applies to the number of result rows returned.

Exam trap

The trap is misreading the bin() interval or confusing 'limit 20' (which limits result rows/buckets) with limiting raw log entries — candidates often assume limit applies to individual log lines rather than aggregated output.

How to eliminate wrong answers

Option A is wrong because bin(5m) creates 5-minute buckets, not 1-minute buckets — a 1-minute interval would require bin(1m). Option C is wrong because the query groups results by time bucket rather than returning a single total; a total count would require removing the 'by bin()' clause. Option D is wrong because the query uses 'stats count(*)' which returns aggregated counts, not raw log entries — returning raw entries would require removing the stats command and using fields/display instead.

18
Multi-Selectmedium

A company is using Amazon CloudWatch Logs to collect logs from multiple applications. The DevOps team wants to create a metric filter to count the number of ERROR log entries and trigger an alarm when the count exceeds 10 in 5 minutes. Which TWO steps must the team take? (Choose TWO.)

Select 2 answers
A.Create a subscription filter to stream logs to Amazon Kinesis Data Firehose.
B.Create a metric filter on the log group that extracts ERROR count.
C.Create a CloudWatch alarm on the metric with the threshold of 10.
D.Set a log group retention policy to retain logs indefinitely.
E.Create a CloudWatch dashboard to visualize the ERROR count.
AnswersB, C

Creating a metric filter on the log group that extracts ERROR count is the first required action. A metric filter defines a pattern, such as the literal string 'ERROR', and evaluates each incoming log event against that pattern; every match increments a specified CloudWatch metric value. The filter can also output a custom value and unit, producing a metric that can be used for alarms and dashboards. This is the native CloudWatch Logs mechanism to translate unstructured log text into a numerical time-series metric.

Why this answer

Option B is correct because a CloudWatch Logs metric filter must be defined on the log group to parse log events and publish a custom metric that counts occurrences of the ERROR pattern; this is the mechanism that turns raw log data into a numeric CloudWatch metric. Option C is correct because once the metric exists, a CloudWatch alarm is created against that metric with a threshold of 10 and an evaluation period of 5 minutes so it can trigger when the count exceeds the limit. Option A is not needed because subscription filters stream logs to destinations like Kinesis Data Firehose for processing, not for creating metrics or alarms.

Option D is irrelevant because retention policy only controls how long log events are stored and does not affect metric filtering or alarming. Option E is unnecessary because a dashboard only visualizes metrics and does not create the metric or trigger the alarm.

Exam trap

DOP-C02 often tests whether candidates confuse the roles of metric filters, subscription filters, and dashboards — the trap is selecting subscription filters (for streaming) or dashboards (for visualization) when the requirement is specifically to count and alarm on log events.

19
MCQhard

A company runs a critical e-commerce application on AWS. The architecture includes an Application Load Balancer (ALB) in front of an Auto Scaling group of EC2 instances running a web server, and an Amazon RDS MySQL Multi-AZ database. The DevOps team has implemented CloudWatch dashboards to monitor key metrics. Recently, customers have reported that the website becomes unresponsive for a few minutes during peak traffic hours. The team reviews the CloudWatch metrics and observes that during the incidents, the ALB's 'TargetResponseTime' metric spikes, and the RDS 'ReadLatency' and 'WriteLatency' metrics also spike. However, the EC2 CPU utilization and memory usage remain normal. The ALB health check shows 'Healthy' for all targets. The team needs to identify the root cause. Which course of action should the team take?

A.Configure the ALB to add a second listener and distribute traffic across multiple target groups.
B.Review the ALB access logs to identify if there are any unusual request patterns causing the latency.
C.Enable Performance Insights on the RDS instance to analyze database performance and identify slow queries.
D.Increase the desired capacity of the Auto Scaling group to add more EC2 instances to handle the load.
AnswerC

Performance Insights for Amazon RDS is the direct diagnostic tool for database performance. It visualizes database load by wait states, SQL statements, and hosts, allowing you to quickly identify slow queries, lock waits, or I/O bottlenecks that are driving the ALB backend latency. This matches the scenario where application instances are healthy but the database is the constraint, making it the correct next step for root-cause analysis.

Why this answer

Option C is correct because the symptoms — spiking ALB TargetResponseTime alongside spiking RDS ReadLatency and WriteLatency, with normal EC2 CPU/memory and healthy targets — point to a database bottleneck, most likely slow or inefficient queries. Performance Insights is the AWS-native tool for identifying top SQL statements, wait events, and database load contributors, so it directly addresses the root cause. Adding instances or listeners would not help because the bottleneck is downstream at the database, not at the application tier.

Exam trap

The trap is chasing the loudest symptom — the ALB TargetResponseTime spike — and adding application-tier capacity, when the correlated RDS latency metrics reveal the database as the true bottleneck.

How to eliminate wrong answers

Option A is wrong because adding a second ALB listener and target group only changes traffic routing; it does nothing to reduce database latency, which is the actual bottleneck. Option B is wrong because ALB access logs show request-level patterns (URLs, status codes, latencies) but cannot reveal SQL-level causes inside RDS — they would confirm the symptom, not diagnose the root cause. Option D is wrong because EC2 CPU and memory are normal, so the application tier is not resource-constrained; adding instances would multiply concurrent database connections and could worsen the problem.

20
MCQhard

A company runs a critical application on Amazon RDS for PostgreSQL. The database experiences periodic slowdowns. The team wants to monitor the number of active connections and the query execution time. Which approach is most cost-effective?

A.Install the CloudWatch agent on the RDS instance to collect custom metrics.
B.Use the RDS console to view the 'DatabaseConnections' and 'QueryExecutionTime' metrics.
C.Enable Performance Insights and set up CloudWatch alarms on the 'DBLoad' metric.
D.Enable Enhanced Monitoring and publish metrics to CloudWatch, then create alarms on relevant metrics.
AnswerC

Performance Insights exposes the DBLoad metric, which represents the average number of active sessions and directly reflects query execution workload, and it can be streamed to CloudWatch for alarm creation. Per-query execution time is visible through the top SQL dashboard, allowing you to correlate load spikes with specific queries. The basic retention is included with RDS at no additional cost, making this a cost-effective and fully managed solution for monitoring query performance.

Why this answer

Performance Insights provides detailed query execution time metrics, enabling monitoring of query performance. It also includes the 'DBLoad' metric which reflects database load from active connections and queries. Combined with CloudWatch alarms, this approach is cost-effective as Performance Insights is included with RDS at no additional cost for up to 7 days of retention (longer retention has a fee).

Options A and B are invalid because the CloudWatch agent cannot be installed on RDS instances, and 'QueryExecutionTime' is not a standard CloudWatch metric for RDS. Option D is less suitable for the specific requirement of query execution time; Enhanced Monitoring provides OS-level metrics like processes and memory but does not track query-level execution time.

Exam trap

Candidates often assume Enhanced Monitoring is the most cost-effective because it is free, but it lacks query-level performance data. The key is that Performance Insights tracks execution time natively and is included at no extra cost for the first 7 days, making it the correct choice for this requirement.

How to eliminate wrong answers

Option A is wrong because the CloudWatch agent cannot be installed on an RDS instance; RDS is a managed service and does not allow direct OS access or agent installation. Option B is wrong because 'QueryExecutionTime' is not a standard metric available in the RDS console; the console provides 'DatabaseConnections' but not query execution time. Option C is wrong because Performance Insights focuses on database load (DBLoad) and query performance analysis, but it does not directly expose the number of active connections as a metric for CloudWatch alarms; additionally, enabling Performance Insights incurs extra costs beyond the basic RDS pricing.

21
MCQmedium

Refer to the exhibit. An administrator wants to be notified when any EC2 instance is launched in the account. Which combination of services would provide the most efficient and cost-effective solution?

A.Enable detailed billing reports and create a cost anomaly detection monitor.
B.Use AWS Config rules to detect non-compliant instances and trigger an SNS notification.
C.Create an Amazon EventBridge rule that matches RunInstances events from CloudTrail and sends to an SNS topic.
D.Set up a CloudWatch alarm on the RunInstances metric in the EC2 namespace.
AnswerC

An EventBridge rule can match event patterns against CloudTrail records, for example source 'aws.ec2', detail-type 'AWS API Call via CloudTrail', and eventName 'RunInstances'. When a matching API call occurs, EventBridge invokes an SNS topic in near real time, and the event payload contains instance IDs, user identity, request parameters, and response elements. This is the native AWS event-driven pattern for reacting to EC2 API activity.

Why this answer

The most efficient and cost-effective solution is to create an Amazon EventBridge rule that matches RunInstances events from CloudTrail and sends notifications to an SNS topic. EventBridge provides near real-time event processing with minimal overhead, and CloudTrail captures EC2 instance launch events. This avoids the need for continuous polling or additional services.

Exam trap

DOP-C02 often tests the confusion between CloudWatch metrics and CloudTrail events, and candidates may incorrectly assume CloudWatch has a RunInstances metric.

How to eliminate wrong answers

Option A is wrong because detailed billing reports and cost anomaly detection are for cost monitoring, not for real-time instance launch notifications. Option B is wrong because AWS Config rules evaluate compliance periodically, not in real-time, and may incur additional costs. Option D is wrong because CloudWatch does not have a RunInstances metric in the EC2 namespace; such events are not metrics but API calls logged by CloudTrail.

22
MCQmedium

A company is using AWS Lambda functions for data processing. The operations team needs to monitor the number of invocations, duration, and error counts for each function. They also want to set alarms when the error rate exceeds 5% in a 5-minute period. Which combination of AWS services should the team use to achieve this with minimal effort?

A.Use AWS CloudTrail to log Lambda invocations and configure CloudWatch alarms on the log events.
B.Enable Lambda Insights to collect detailed metrics and use CloudWatch dashboards to monitor error rates.
C.Stream Lambda logs to CloudWatch Logs and use CloudWatch Logs Insights to query error rates, then create alarms.
D.Use CloudWatch metrics published by Lambda and create a CloudWatch alarm on the ErrorCount metric with a math expression to calculate error rate.
AnswerD

Lambda automatically emits a set of standard metrics to CloudWatch, including 'Invocations', 'Errors', and 'ErrorCount' (via enhanced metrics if enabled), so you can directly build an error-rate expression without additional setup. The correct approach is to use a CloudWatch math expression, such as 'm1/m2*100' where m1 is ErrorCount and m2 is Invocations, and then create an alarm on that expression to alert when the error rate exceeds a threshold. This leverages the native, low-latency monitoring pipeline and avoids the overhead of log-based or third-party tooling, making it the simplest and most operationally sound solution.

Why this answer

Lambda automatically publishes metrics to CloudWatch, including Invocations, Duration, and Errors. To monitor error rate, you can create a CloudWatch alarm using a math expression that calculates the error rate from the ErrorCount and Invocations metrics. This approach requires minimal effort because it leverages built-in metrics and CloudWatch's native alarm capabilities.

Exam trap

DOP-C02 often tests the distinction between metrics and logs; candidates may choose log-based solutions, but the exam favors using built-in CloudWatch metrics and math expressions for minimal effort and real-time monitoring.

How to eliminate wrong answers

Option A is wrong because CloudTrail logs API calls, not Lambda invocations, and does not provide metrics for duration or error counts. Option B is wrong because Lambda Insights provides enhanced metrics but requires additional setup and is not necessary for basic error rate monitoring. Option C is wrong because streaming logs to CloudWatch Logs and using Logs Insights to query error rates is more complex and does not directly create alarms; it would require custom metric filters and alarms.

23
Multi-Selectmedium

A company is using Amazon CloudWatch to monitor its production environment. The operations team receives alerts for the same underlying issue from multiple alarms, causing alert fatigue. The team wants to reduce noise and consolidate alerts into actionable notifications. Which TWO steps should the team take? (Choose two.)

Select 2 answers
A.Configure the CloudWatch alarms to publish to an SNS topic, and use SNS subscription filter policies to route only critical notifications.
B.Use CloudWatch Evidently to run experiments and filter out false alarms.
C.Use CloudWatch composite alarms to combine multiple alarms into a single alarm that triggers only when certain conditions are met.
D.Use CloudWatch Logs Insights to query logs and create alarms based on the query results.
E.Use AWS Config rules to automatically suppress alarms that are not compliant.
AnswersA, C

Publishing CloudWatch alarms to an SNS topic and applying subscription filter policies is a valid approach because SNS supports content-based filtering on message attributes. When an alarm transitions to ALARM state, it publishes a message with attributes like `state` and `severity`; each subscriber can define filter policies that match only their desired subset (e.g., `state = ALARM` with high severity). This effectively routes critical notifications to the right team while suppressing non-critical messages at the subscription level, without altering the alarm logic itself.

Why this answer

You can configure CloudWatch alarms to publish to an SNS topic and use SNS subscription filter policies to route only critical notifications, thereby reducing noise. Option C is correct because CloudWatch composite alarms allow you to combine multiple alarms into a single alarm that triggers only when specific conditions (e.g., AND/OR logic) are met, consolidating alerts for the same underlying issue. Option B is incorrect because CloudWatch Evidently is used for running experiments and feature flags, not for alert consolidation.

Option D is incorrect because CloudWatch Logs Insights is a tool for querying log data, not for combining alarms. Option E is incorrect because AWS Config rules are designed to evaluate resource compliance, not to suppress alarms.

24
MCQhard

A DevOps engineer is troubleshooting an AWS Lambda function that processes messages from an Amazon SQS queue. The function is invoked successfully, but it frequently times out after 15 seconds. The function's CloudWatch Logs show that the timeout occurs while the function is making an HTTP request to an external API. The function's reserved concurrency is set to 5, and the SQS queue has a visibility timeout of 30 seconds. Which change would MOST effectively reduce the number of timeouts?

A.Increase the Lambda function's timeout to 30 seconds.
B.Increase the SQS queue's visibility timeout to 60 seconds.
C.Decrease the SQS batch size to 1.
D.Increase the Lambda function's reserved concurrency to 10.
AnswerA

The Lambda service enforces a configurable timeout that caps how long a single invocation can run, with a default of 3 seconds. If the function calls a downstream HTTP endpoint that is slower than the remaining execution time, the runtime terminates the invocation and returns a timeout error to the caller. Raising the timeout to 30 seconds grants the HTTP request enough time to return a response, directly resolving the symptom described in the troubleshooting scenario.

Why this answer

The function times out at 15 seconds while waiting on an external HTTP call, so the most direct fix is to raise the Lambda timeout to 30 seconds, giving the external API enough time to respond. The SQS visibility timeout (30s) already exceeds the current Lambda timeout, so it is not the bottleneck. Increasing the timeout is the change that most directly reduces the number of timeouts.

Exam trap

The trap is that candidates focus on SQS visibility timeout or concurrency, but the symptom (timeout during an HTTP call) points squarely at the Lambda execution timeout, which is the only setting that directly controls how long the function may run.

How to eliminate wrong answers

Option B is wrong because the visibility timeout (30s) is already greater than the Lambda timeout (15s), so extending it to 60s does not address the root cause — the function itself is being killed by Lambda, not by SQS redelivery. Option C is wrong because decreasing batch size to 1 reduces per-invocation work but does not change the fact that a single external HTTP call is exceeding 15 seconds. Option D is wrong because reserved concurrency controls how many concurrent invocations run, not how long each invocation may run — it would not prevent individual timeouts.

25
MCQeasy

A company wants to receive real-time notifications when their Auto Scaling group launches or terminates EC2 instances. Which AWS service should they use?

A.Amazon CloudWatch alarm on the GroupTotalInstances metric.
B.AWS Config rules to detect changes in Auto Scaling groups.
C.AWS CloudTrail to monitor Auto Scaling API calls.
D.Amazon SNS notifications from the Auto Scaling group.
AnswerD

Auto Scaling groups have a built-in notification feature that publishes messages to an SNS topic on lifecycle events like ec2-instance-launch and ec2-instance-terminate. This provides immediate, event-driven delivery of the instance ID, the Auto Scaling group name, and the event type to all subscribers (email, SMS, Lambda, HTTP endpoints). It is the simplest native way to receive real-time notifications for the exact scaling actions the company cares about.

Why this answer

Auto Scaling groups natively support lifecycle hooks and notification configurations that publish events (EC2 Instance-launch, EC2 Instance-terminate, EC2 Instance-launch-lifecycle-action, etc.) directly to an Amazon SNS topic. This is the built-in, real-time mechanism designed specifically for launch/terminate notifications, requiring no polling or metric evaluation.

Exam trap

The trap here is confusing CloudWatch alarms (threshold-based, aggregate metrics) with native ASG lifecycle notifications (event-based, per-instance) — candidates often pick CloudWatch because it 'feels' like the monitoring answer.

How to eliminate wrong answers

Option A is wrong because a CloudWatch alarm on GroupTotalInstances only fires when the aggregate count crosses a threshold — it cannot notify on individual instance launch/terminate events and is not real-time per-instance. Option B is wrong because AWS Config rules evaluate resource configuration compliance over time (typically minutes), not real-time lifecycle events, and would require custom rules to detect scaling activity. Option C is wrong because CloudTrail logs Auto Scaling API calls for auditing, but it is not a real-time push notification service — you would need to build EventBridge/CloudWatch Logs filters on top of it.

26
MCQmedium

A company uses AWS CloudFormation to deploy infrastructure. The DevOps team wants to receive notifications when a stack creation fails due to a resource limit exceeded error. Which approach should be used?

A.Create an Amazon EventBridge rule that matches CloudFormation resource limit exceeded events and sends to SQS.
B.Configure an SNS topic as a notification option in the CloudFormation stack, and subscribe an email endpoint.
C.Use AWS Config to detect when a stack is in a failed state.
D.Enable CloudTrail and create a CloudWatch alarm on the CreateStack API call.
AnswerB

CloudFormation natively supports an SNS topic as a stack notification option: when you specify an SNS topic ARN in the stack's NotificationARNs property, CloudFormation publishes every stack event—such as CREATE_FAILED and ROLLBACK_COMPLETE—to that topic. Subscribing an email endpoint to the topic delivers these events in near real time, giving operations teams immediate insight into resource limit failures or any other stack error. This is the only answer that directly leverages CloudFormation's built-in notification channel without needing extra services to infer failure.

Why this answer

CloudFormation natively supports sending stack events (including creation failures) to an SNS topic. By configuring an SNS topic as a notification option in the stack creation request, the DevOps team can subscribe an email endpoint to receive real-time notifications when a resource limit exceeded error occurs, without needing additional services or custom logic.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing EventBridge or CloudTrail-based monitoring, missing the fact that CloudFormation has a built-in, straightforward SNS notification feature specifically designed for real-time stack event alerts.

How to eliminate wrong answers

Option A is wrong because Amazon EventBridge does not natively emit a specific 'resource limit exceeded' event from CloudFormation; CloudFormation events in EventBridge are generic stack-level events (e.g., CREATE_FAILED) and would require custom filtering and parsing to detect the specific error message, making it less direct than using SNS. Option C is wrong because AWS Config is designed for resource compliance and configuration tracking, not for real-time monitoring of CloudFormation stack creation failures; it cannot trigger notifications for transient stack events like resource limit exceeded errors. Option D is wrong because enabling CloudTrail and creating a CloudWatch alarm on the CreateStack API call would only detect that a CreateStack call was made, not whether the stack creation failed due to a resource limit exceeded error; the alarm would fire on every CreateStack call, not on failures, and would require additional log filtering and metric filters to isolate the specific error.

27
Multi-Selectmedium

A company is using Amazon CloudWatch to monitor a production environment. The DevOps team wants to receive notifications when the CPU utilization of an EC2 instance exceeds 90% for 5 consecutive minutes. Which TWO steps should the team take to achieve this? (Choose TWO.)

Select 2 answers
A.Enable detailed monitoring on the EC2 instance to get 1-minute metrics.
B.Configure an Amazon SNS topic and subscribe the team's email address to it, then set the alarm to send notifications to the SNS topic.
C.Create a CloudWatch alarm on the CPUUtilization metric with a threshold of 90% and an evaluation period of 5 consecutive minutes.
D.Create a CloudWatch Logs metric filter to count CPU utilization errors.
E.Create a CloudWatch dashboard to visualize CPU utilization.
AnswersB, C

An Amazon SNS topic acts as the delivery channel for CloudWatch alarm actions. By creating a topic, subscribing the team's email address, and confirming the subscription, the alarm can publish messages to that topic whenever it transitions to the ALARM state. This is the standard method to send email notifications and is a required component for the team to be alerted. Without this, the alarm would only change state and not proactively reach the team.

Why this answer

Option B is correct because CloudWatch alarms cannot send email directly; they must publish to an Amazon SNS topic, and the team's email address must be subscribed to that topic so the notification is delivered. Option C is correct because the alarm must be defined on the EC2 CPUUtilization metric with a threshold of 90% and an evaluation period of 5 consecutive 1-minute datapoints (5 minutes) to match the requirement. Option A is not required because basic monitoring already provides CPUUtilization at 5-minute intervals, which is sufficient for a 5-minute evaluation period; detailed monitoring only changes the granularity to 1 minute.

Option D is incorrect because CloudWatch Logs metric filters operate on log events, not on the CPUUtilization metric. Option E is incorrect because a dashboard only visualizes metrics and does not generate notifications.

Exam trap

DOP-C02 often tests whether candidates overcomplicate the solution by enabling detailed monitoring unnecessarily — standard 5-minute metrics are sufficient for a 5-minute evaluation period.

28
MCQeasy

A company is using AWS CloudTrail to track API calls. They want to be notified immediately when an IAM user creates a new access key. Which combination of AWS services should be used?

A.Amazon CloudWatch Logs with a metric filter and alarm.
B.AWS Config with an AWS Lambda function.
C.Amazon CloudWatch Events (Amazon EventBridge) with an AWS Lambda function that sends an email via Amazon SES.
D.Amazon CloudWatch Events (Amazon EventBridge) with an Amazon SNS topic.
AnswerD

Amazon EventBridge is the native event router for CloudTrail API activity: CloudTrail automatically delivers every call to an event bus, and a rule with a JSON pattern can match the specific API (e.g., an unauthorized or sensitive call). The rule immediately invokes an Amazon SNS topic, which then fans out notifications via email, SMS, or other subscribers. This event-driven flow provides sub-second, real-time alerts with no polling, no custom code, and direct integration, making it the correct architecture.

Why this answer

To be notified immediately when an IAM user creates a new access key, the most efficient approach is to use Amazon CloudWatch Events (Amazon EventBridge) with an Amazon SNS topic. CloudTrail records the 'CreateAccessKey' API call as an event. An EventBridge rule can be configured to match this specific event pattern and send the event to an SNS topic, which can then send notifications via email, SMS, etc.

This provides real-time notification without additional services. Option A (CloudWatch Logs with metric filter and alarm) requires sending CloudTrail logs to CloudWatch Logs, which adds latency and complexity; it is not as direct as EventBridge. Option B (AWS Config with Lambda) is not designed for real-time event notification.

Option C (EventBridge with Lambda and SES) adds unnecessary Lambda processing since SNS can directly send email when subscribed to the topic.

29
MCQhard

A company runs a critical application on Amazon EKS. The DevOps team uses Prometheus for monitoring and Grafana for visualization. The team has set up a Prometheus server on an EC2 instance to scrape metrics from the EKS cluster. However, they are experiencing high memory usage on the Prometheus server, and some metrics are being dropped because of the retention period. The team wants to implement a scalable and managed monitoring solution that can store metrics for longer durations without the operational overhead of managing the Prometheus server. The team also wants to retain the ability to use PromQL queries and Grafana dashboards. What should the team do?

A.Use Amazon Managed Grafana to visualize metrics directly from the EKS cluster without a Prometheus server.
B.Migrate to Amazon Managed Service for Prometheus to ingest and store metrics, and use Amazon Managed Grafana for visualization.
C.Increase the EC2 instance size for the Prometheus server and extend the retention period.
D.Set up Amazon CloudWatch Container Insights to collect metrics from the EKS cluster and store them in CloudWatch Logs.
AnswerB

Amazon Managed Service for Prometheus is a scalable, fully managed service compatible with Prometheus query language (PromQL) and remote write. It eliminates the operational overhead of running Prometheus at scale, provides durable storage, and integrates with Amazon Managed Grafana for dashboards. This is the recommended AWS-native approach for long-term metrics retention on EKS.

Why this answer

Amazon Managed Service for Prometheus is a scalable, fully managed service that ingests and stores Prometheus metrics, supports PromQL queries, and integrates seamlessly with Amazon Managed Grafana. This eliminates the operational overhead of managing a Prometheus server while enabling longer retention and scaling. Option A is incorrect because Amazon Managed Grafana is a visualization tool only; it does not store metrics.

Option C is incorrect because increasing the EC2 instance size does not address the operational overhead or scalability issues; it only postpones the problem. Option D is incorrect because Amazon CloudWatch Container Insights does not support PromQL natively, and migrating to CloudWatch would require rewriting queries and dashboards, losing compatibility with existing Prometheus and Grafana setups.

30
Multi-Selectmedium

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer. The operations team wants to analyze application access logs and error rates. They need to identify the top IP addresses making requests, as well as the distribution of HTTP status codes over time. Which THREE steps should the team take to achieve this? (Select THREE.)

Select 3 answers
A.Enable access logs on the Application Load Balancer and store them in an Amazon S3 bucket.
B.Use Amazon CloudWatch Logs Insights to run queries on the access logs.
C.Enable AWS CloudTrail to log all API calls.
D.Enable VPC Flow Logs to capture IP traffic data.
E.Use Amazon CloudWatch Contributor Insights to analyze the top IP addresses.
AnswersA, B, E

Enabling access logs on the Application Load Balancer is the foundational step; it delivers a raw, per-request log file containing the client IP, request URI, User-Agent, and HTTP status code for every request handled by the ALB. These logs are written in gzip-compressed files to an S3 bucket you specify, and S3 provides durable, queryable storage. Without this first step, no HTTP-level transaction data exists in S3, making any subsequent log analysis or querying impossible. This is why enabling ALB access logs is the correct primary action.

Why this answer

Enabling access logs on the Application Load Balancer and storing them in an S3 bucket captures detailed HTTP request data, including client IPs, request paths, and HTTP status codes. This raw log data is essential for analyzing top IP addresses and status code distributions over time.

Exam trap

The trap here is confusing AWS CloudTrail (management plane logging) with application-level access logging, leading candidates to select CloudTrail instead of ALB access logs for HTTP request analysis.

31
MCQmedium

A company runs a containerized order-processing service on Amazon EKS. The DevOps team wants to detect when the service's CPU utilization exceeds 80% for 5 consecutive minutes and then automatically notify an on-call engineer. They also need to see both application logs and CPU metrics side-by-side in a single dashboard for troubleshooting. Which combination of AWS services should the team use to meet these requirements with the LEAST operational overhead?

A.AWS X-Ray with CloudWatch Logs Insights and CloudWatch alarms
B.Amazon CloudWatch Container Insights with CloudWatch alarms and CloudWatch dashboards
C.AWS CloudTrail with Amazon EventBridge rules and Amazon SNS topics
D.Amazon Managed Service for Prometheus with AWS Distro for OpenTelemetry and Grafana dashboards
AnswerB

Container Insights automatically collects CPU, memory, network, and log data from EKS nodes and pods without installing custom agents beyond the CloudWatch agent DaemonSet. CloudWatch alarms evaluate the CPUUtilization metric over 5-minute periods, and CloudWatch dashboards can display both metrics and log insights queries together, satisfying the monitoring and alerting needs with minimal operational effort.

Why this answer

Container Insights provides native EKS monitoring with automatic metric and log collection, and CloudWatch alarms and dashboards integrate directly to alert on CPU thresholds and visualize logs and metrics together. The other options either lack CPU metric collection or require substantial additional tooling and management, making them less suitable for a low-overhead solution.

Exam trap

The trap here is assuming that X-Ray or CloudTrail can provide CPU utilization metrics, when they are designed for tracing and auditing respectively.

32
MCQeasy

A company uses Amazon RDS for PostgreSQL and wants to monitor database performance metrics such as CPU utilization, memory, and disk I/O. Which AWS service should be used to set up custom dashboards and alarms for these metrics?

A.AWS X-Ray
B.Amazon VPC Flow Logs
C.AWS CloudTrail
D.Amazon CloudWatch
AnswerD

Amazon CloudWatch natively ingests Amazon RDS performance metrics—including CPU utilization, freeable memory, database connections, read/write IOPS, and replica lag—through the AWS/RDS metric namespace. These metrics are published automatically and can be visualized in CloudWatch dashboards or used to trigger alarms for proactive alerting and auto scaling actions. CloudWatch also supports publishing custom metrics and streaming RDS logs, making it the comprehensive monitoring service for RDS PostgreSQL.

Why this answer

Amazon CloudWatch (Option D) is the correct service for monitoring Amazon RDS performance metrics such as CPU utilization, memory, and disk I/O. CloudWatch provides built-in metrics for RDS, allows creation of custom dashboards to visualize these metrics, and supports setting alarms for proactive notifications. Option A (AWS X-Ray) is used for tracing and analyzing requests through applications, not for infrastructure metrics.

Option B (Amazon VPC Flow Logs) captures IP traffic information for network troubleshooting, not database performance. Option C (AWS CloudTrail) logs API calls for auditing, not performance monitoring. Therefore, CloudWatch is the appropriate choice for this use case.

33
MCQmedium

A company runs a web application behind an Application Load Balancer (ALB) in a production AWS account. The DevOps team needs to analyze HTTP request patterns and identify the top IP addresses generating errors. They want to store the data cost-effectively for querying with SQL. Which solution meets these requirements?

A.Use CloudWatch Metrics to monitor error rates and top IPs via custom metrics.
B.Enable CloudWatch Logs for the ALB and use CloudWatch Logs Insights to query the logs.
C.Stream the ALB logs to Amazon Kinesis Data Analytics and use SQL applications.
D.Enable ALB access logs and store them in Amazon S3, then use Amazon Athena to query the logs with SQL.
AnswerD

Enabling ALB access logs to be delivered to Amazon S3 creates immutable, row-based log files that are ideal for large-scale retrospective analysis. Amazon Athena lets you run standard SQL directly on that S3 data using a serverless engine that charges only for the bytes scanned, and you can further optimize costs and performance by partitioning S3 objects by date or using compression. This combination is the industry-standard cost-effective approach when the goal is to perform flexible SQL queries over historical ALB logs without pre-provisioning infrastructure or paying for continuous ingestion.

Why this answer

ALB access logs provide detailed HTTP request data (including source IP, request URI, response code, etc.) and are stored in Amazon S3, which is cost-effective for long-term storage. Amazon Athena allows querying these logs directly with standard SQL without needing to load data into a database, meeting the requirement for SQL-based analysis of top IP addresses generating errors.

Exam trap

The trap here is that candidates often confuse CloudWatch Logs (which for ALB only contain error logs, not full request details) with ALB access logs (which are stored in S3 and contain all request data), leading them to choose Option B instead of D.

How to eliminate wrong answers

Option A is wrong because CloudWatch Metrics cannot capture individual HTTP request details like source IP addresses; custom metrics are aggregated and cannot be used to identify top IPs generating errors. Option B is wrong because CloudWatch Logs for ALB capture only error-level logs (e.g., 5xx responses) and do not include request-level details such as source IP; CloudWatch Logs Insights cannot query for top IP addresses from these logs. Option C is wrong because Kinesis Data Analytics is designed for real-time stream processing with SQL, but the requirement is to store data cost-effectively for querying, not real-time analysis; streaming logs to Kinesis incurs ongoing costs and is overkill for batch querying of historical patterns.

34
MCQeasy

A company wants to monitor the number of messages in an Amazon SQS queue and send an alert if the queue depth exceeds 1000 for more than 5 minutes. Which AWS service should be used to create the alarm?

A.Amazon EventBridge
B.Amazon CloudWatch Alarms
C.AWS X-Ray
D.Amazon CloudWatch Logs
AnswerB

Amazon CloudWatch Alarms are the correct choice because they continuously monitor a specified CloudWatch metric—such as ApproximateNumberOfMessagesVisible for an SQS queue—against a defined threshold over a configured period. When the metric crosses the threshold, the alarm changes state (OK, ALARM, or INSUFFICIENT_DATA) and can trigger an action like an SNS notification, EC2 Auto Scaling, or an arbitrary EC2 stop/terminate action. This provides the exact metric-threshold monitoring required.

Why this answer

Amazon CloudWatch Alarms is the correct service because it can monitor SQS queue metrics (such as ApproximateNumberOfMessagesVisible) and trigger an alarm when the metric exceeds a threshold (e.g., 1000) for a specified evaluation period (e.g., 5 minutes). CloudWatch Alarms directly integrate with SQS via the AWS/SQS namespace and support actions like sending notifications through Amazon SNS.

Exam trap

The trap here is that candidates may confuse EventBridge's ability to react to SQS metric changes (via CloudWatch metric streams) with the actual alarm evaluation logic, but EventBridge cannot perform threshold-based monitoring over a time window—only CloudWatch Alarms can.

How to eliminate wrong answers

Option A is wrong because Amazon EventBridge is a serverless event bus used for routing events between services (e.g., reacting to state changes), not for monitoring metric thresholds over time or creating alarms based on sustained conditions. Option C is wrong because AWS X-Ray is a distributed tracing service for analyzing and debugging application requests, not for monitoring queue depth or setting metric alarms. Option D is wrong because Amazon CloudWatch Logs is used for storing, monitoring, and querying log data, not for creating alarms on numeric metrics like SQS queue depth.

35
MCQeasy

A company wants to monitor CPU utilization of its EC2 instances and receive an alert when utilization exceeds 80% for 5 consecutive minutes. Which AWS service should be used to create this alarm?

A.AWS CloudTrail
B.VPC Flow Logs
C.Amazon CloudWatch Alarms
D.AWS Config
AnswerC

Amazon CloudWatch Alarms are the correct service because they directly evaluate CloudWatch metrics, and CPUUtilization is a built-in metric emitted by EC2 instances. You can configure an alarm to transition to an ALARM state when the average CPU utilization exceeds a threshold, and then trigger actions like Auto Scaling, Amazon SNS notifications, or EC2 actions (reboot, stop, terminate). Alarm evaluation periods and statistic types (e.g., average, maximum) allow precise control over when an alarm fires, making it the natural choice for CPU monitoring.

Why this answer

Amazon CloudWatch Alarms is the correct service to monitor CPU utilization of EC2 instances and trigger an alert when the metric exceeds a threshold for a specified period. CloudWatch collects metrics from EC2 instances, and alarms can be configured to watch a metric over a time period and perform actions (e.g., send an SNS notification) when the threshold is breached. This directly matches the requirement.

Exam trap

DOP-C02 often tests the difference between monitoring, logging, and auditing services, and candidates may incorrectly choose CloudTrail or Config for performance monitoring.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail logs API activity, not performance metrics like CPU utilization. Option B is wrong because VPC Flow Logs capture IP traffic information, not instance performance metrics. Option D is wrong because AWS Config evaluates resource configurations against desired policies, not real-time performance metrics.

36
MCQmedium

An application running on AWS Lambda is experiencing cold starts. The team wants to monitor the cold start duration. What should they do?

A.Monitor the 'InitDuration' metric in CloudWatch for the Lambda function.
B.Use CloudWatch Logs Insights to query log groups for 'REPORT' lines and calculate duration.
C.Publish a custom metric from the Lambda code that measures initialization time.
D.Enable AWS X-Ray and trace the Lambda invocation to see cold start duration.
AnswerA

Monitoring the 'InitDuration' metric is the direct, built-in approach because CloudWatch automatically receives this metric for every Lambda invocation that experiences a cold start. This metric reports the time the runtime spends initializing the execution environment (downloading the code, starting the runtime, and running initialization code) before the handler is invoked. It requires no custom instrumentation or querying, making it the simplest and most authoritative source for cold start duration.

Why this answer

AWS Lambda automatically publishes the 'InitDuration' metric in CloudWatch for cold starts, which measures the time spent initializing the runtime and code. Option B is incorrect because while CloudWatch Logs Insights can query for 'REPORT' lines, it is more complex and unnecessary since the metric is already available. Option C is incorrect because publishing a custom metric from the Lambda code is redundant; the InitDuration metric is automatically provided.

Option D is incorrect because AWS X-Ray can trace cold starts, but the dedicated metric is simpler and directly available.

37
MCQeasy

Refer to the exhibit. The IAM policy above is attached to a Lambda function's execution role. The Lambda function is supposed to publish custom metrics to CloudWatch using PutMetricData. However, the metrics are not appearing. What is the most likely reason?

A.The policy does not include the 'cloudwatch:PutMetricData' action.
B.The policy includes unnecessary actions that conflict with each other.
C.The function needs to specify a metric name and value when calling PutMetricData.
D.The policy uses a wildcard resource, which is not allowed for the PutMetricData action.
AnswerC

Granting cloudwatch:PutMetricData only authorizes the API call; it does not construct or complete it. The Lambda function must send a MetricDatum object that includes both a MetricName and a numeric Value (or StatisticValues/Values), along with other optional fields like Dimensions and Timestamp. If these required fields are missing or empty, the API call cannot create a metric data point, regardless of how permissive the attached IAM policy is.

Why this answer

The most likely reason the metrics are not appearing is that the Lambda function is not providing the required parameters—specifically a metric name and value—when calling PutMetricData. The IAM policy correctly grants the cloudwatch:PutMetricData action, but the API call itself must include at least a MetricName and a Value (or StatisticValues) in the MetricDatum array; otherwise, CloudWatch silently drops the request without publishing any metric. This is a common coding error where the execution role permissions are sufficient but the function logic is incomplete.

Exam trap

The trap here is that candidates assume the issue is always a missing IAM permission (Option A) when the real problem is often a missing required parameter in the API call, especially since PutMetricData returns success even with invalid data.

How to eliminate wrong answers

Option A is wrong because the policy explicitly includes 'cloudwatch:PutMetricData' as an action, so the Lambda function has the necessary permission to publish metrics. Option B is wrong because the presence of multiple actions (like PutMetricData, GetMetricData, ListMetrics) does not cause conflicts; IAM policies allow multiple actions and they do not interfere with each other unless there are explicit deny statements. Option D is wrong because the PutMetricData action supports a wildcard resource ('*') in the policy; CloudWatch metrics are global and do not require a specific resource ARN, so using '*' is both allowed and standard practice.

38
MCQhard

A company runs a microservices application on Amazon EKS. The DevOps team wants to collect and visualize metrics such as pod CPU and memory usage, and set up alerts. Which combination of AWS services should be used?

A.Prometheus and Grafana on EC2
B.AWS X-Ray and Amazon CloudWatch ServiceLens
C.AWS CloudTrail and Amazon CloudWatch Logs
D.Amazon CloudWatch Container Insights and CloudWatch Alarms
AnswerD

Amazon CloudWatch Container Insights automatically discovers and collects pod-level metrics from Amazon EKS, including CPU, memory, network, and disk utilization, and presents them through pre-built dashboards. By defining CloudWatch Alarms on these metrics, operators receive proactive notifications when thresholds are breached, making this a fully managed, integrated solution for monitoring and alerting on container resource usage.

Why this answer

Amazon CloudWatch Container Insights and CloudWatch Alarms. Container Insights collects, aggregates, and summarizes metrics and logs from containerized applications and microservices running on Amazon EKS. It provides pre-built dashboards for pod CPU, memory, network, and disk metrics.

CloudWatch Alarms can be set on these metrics to trigger notifications or automated actions. Option A (Prometheus and Grafana on EC2) is possible but not an AWS service combination; it requires manual setup and maintenance. Option B (AWS X-Ray and Amazon CloudWatch ServiceLens) is for distributed tracing and service maps, not for collecting CPU/memory metrics.

Option C (AWS CloudTrail and Amazon CloudWatch Logs) is for auditing API calls and storing log data, not for metrics visualization and alerting.

39
MCQeasy

A company wants to receive notifications when an EC2 instance's CPU utilization exceeds 90% for 10 consecutive minutes. Which AWS service should be used?

A.Amazon CloudWatch alarm
B.AWS Config rule
C.AWS CloudTrail event
D.Amazon Inspector
AnswerA

A CloudWatch alarm is the correct mechanism because it continuously evaluates an EC2 instance's time-series metrics, such as CPUUtilization or StatusCheckFailed, against a defined threshold. When the metric crosses the threshold, the alarm state transitions to ALARM and automatically publishes a message to an SNS topic, which can deliver email, SMS, or invoke a Lambda function. Alarms can also be configured to perform EC2 actions like stop or reboot, making them the native AWS service for threshold-based metric notification.

Why this answer

Amazon CloudWatch alarms monitor specified metrics (like CPU utilization) and trigger actions (e.g., SNS notification) when a threshold is breached for a given period. Option A is correct. Option B is incorrect because AWS Config rules evaluate configuration compliance, not metric thresholds.

Option C is incorrect because AWS CloudTrail records API activity, not metric monitoring. Option D is incorrect because Amazon Inspector assesses security vulnerabilities, not performance metrics.

40
MCQmedium

A company is running a critical web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The DevOps team wants to monitor HTTP 5xx errors and receive alerts when the error rate exceeds 5% over a 5-minute period. Which combination of services and configurations should be used to meet these requirements?

A.Enable CloudWatch Logs for the ALB and use CloudWatch Logs Insights to query 5xx logs, then create a metric filter and alarm.
B.Configure AWS Config rules to check ALB 5xx error counts and trigger alarms.
C.Use CloudWatch ALB metrics (HTTPCode_ELB_5XX_Count) and create a CloudWatch Alarm on the Sum statistic with a threshold based on total request count.
D.Use AWS X-Ray to trace requests and create a CloudWatch alarm based on X-Ray error rate.
AnswerC

Correct: The Application Load Balancer natively emits the HTTPCode_ELB_5XX_Count metric to CloudWatch, representing the number of 5xx responses returned by the load balancer itself. Create a CloudWatch Alarm on this metric using the Sum statistic over a period (e.g., 5 minutes) and set a threshold, optionally using a math expression to divide by RequestCount to track the error ratio. This is the simplest and most direct method because it uses existing metrics with no additional setup, latency, or cost.

Why this answer

ALB automatically publishes the `HTTPCode_ELB_5XX_Count` metric to CloudWatch, and you can create a CloudWatch alarm using the `Sum` statistic over a 5-minute period. To detect when the error rate exceeds 5%, you need to combine this metric with the `RequestCount` metric in a math expression (e.g., `m1/m2*100 > 5`) or use a composite alarm, as the alarm threshold must be based on the ratio of 5xx errors to total requests, not just the raw count.

Exam trap

The trap here is that candidates often assume they need to parse logs (Option A) or use a separate tracing service (Option D) for error rate monitoring, when in fact the ALB's built-in CloudWatch metrics and metric math provide a simpler, real-time, and cost-effective solution without additional log ingestion or query overhead.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs Insights is a query tool for analyzing log data, not a real-time alerting mechanism; while you can create a metric filter from ALB logs to count 5xx errors, this approach introduces latency and additional cost, and it is not the simplest or most direct method when ALB metrics are already available. Option B is wrong because AWS Config rules are designed for compliance and resource configuration auditing (e.g., checking if ALB is configured with a specific security policy), not for monitoring real-time error rates or triggering alarms on metric thresholds. Option D is wrong because AWS X-Ray traces individual requests to identify latency and errors, but it does not aggregate HTTP 5xx error rates over a time window or natively publish a metric that can be used directly in a CloudWatch alarm for this specific requirement.

41
MCQhard

Refer to the exhibit. The CloudWatch alarm is set on CPUUtilization. The instance's CPU at 10:00 is 75%, at 10:05 is 82%, and at 10:10 is 85%. Will the alarm trigger?

A.Yes, because the CPU exceeded 80% at 10:05.
B.Yes, but only after 10:10 when two consecutive periods are above the threshold.
C.No, because the alarm uses Average statistic and the average over the entire time is below 80%.
D.No, because the first data point at 10:00 is below the threshold.
AnswerB

Correct. When the alarm evaluates at 10:10, it considers the two most recent consecutive periods—10:05 and 10:10—both of which have CPU utilization above 80%. Since the alarm is configured with two evaluation periods and the threshold of 80%, both consecutive periods satisfy the condition, causing the alarm to enter ALARM state. At 10:05, only one period had breached, so the alarm was still OK.

Why this answer

The alarm evaluates 2 consecutive periods (each 5 minutes). At 10:05, the average for that period is 82% (above 80), but the previous period (10:00) average is 75% (below 80). So only one period is breached.

At 10:10, the average for that period is 85% (above 80), and the previous period (10:05) average is 82% (above 80). So two consecutive periods are breached, triggering the alarm.

42
MCQhard

A company uses AWS CloudTrail to monitor API activity. The DevOps team needs to ensure that any deletion of an S3 bucket is detected in real time and triggers an automated response. Which combination of AWS services should be used to meet these requirements?

A.Use CloudWatch Logs to monitor the logs, and create a metric filter to trigger an alarm when the DeleteBucket event appears.
B.Configure S3 event notifications to send to an SQS queue, and poll the queue with a Lambda function.
C.Configure CloudTrail to deliver logs to an S3 bucket, and use S3 event notifications to invoke a Lambda function.
D.Send CloudTrail logs to CloudWatch Logs, create a CloudWatch Events rule matching the DeleteBucket event, and target a Lambda function.
AnswerD

Configuring CloudTrail to stream events to CloudWatch Logs provides near-real-time delivery of management events, including DeleteBucket. A CloudWatch Events (EventBridge) rule can use an event pattern to match the specific DeleteBucket API call and directly target a Lambda function as its target. This creates a low-latency, event-driven pipeline that automatically invokes the Lambda function without polling or human intervention, enabling immediate remediation or alerting.

Why this answer

CloudTrail logs can be sent to CloudWatch Logs, and a CloudWatch Events rule (now Amazon EventBridge) can be created to match the DeleteBucket event and trigger a Lambda function for automated response in real time. Option A is incorrect because CloudWatch Logs monitors log data, but a metric filter on DeleteBucket events in CloudTrail logs can trigger an alarm, but that is not as direct as EventBridge. More importantly, CloudWatch Logs metric filters have latency and are not the best for real-time response.

Option B is incorrect because S3 event notifications are for operations on objects, not for bucket-level operations like deletion. Option C is incorrect because while CloudTrail can deliver logs to S3, S3 event notifications are not triggered by CloudTrail log file delivery; they are for object-level events. Using EventBridge directly with CloudTrail is the correct real-time approach.

43
MCQmedium

Refer to the exhibit. A DevOps engineer checks the CloudWatch alarm configuration and state. The alarm is in ALARM state for CPUUtilization averaging 90% over 5 minutes, but no notification was received. What is the most likely reason?

A.The SNS topic does not have any confirmed subscriptions.
B.The EC2 instance is stopped.
C.The alarm period is set to 300 seconds, which is too long.
D.The alarm has insufficient data to evaluate.
AnswerA

CloudWatch alarm actions publish notifications to the configured SNS topic, but the publish operation only succeeds if the topic has at least one active, confirmed subscriber endpoint, such as an email address that has clicked the confirmation link. When subscriptions are still in PendingConfirmation state or no subscription exists, the message is silently dropped while the alarm continues to show ALARM. Therefore, even though the alarm is firing correctly, no notification reaches the engineer because the delivery path lacks a confirmed subscriber. This is the root cause.

Why this answer

A CloudWatch alarm in ALARM state triggers its configured actions, typically publishing to an SNS topic. If no notification is received, the most likely cause is that the SNS topic has no confirmed subscriptions—meaning no endpoint (email, SMS, Lambda, etc.) has completed the subscription confirmation handshake. Without a confirmed subscription, SNS accepts the publish but delivers to no one, so the alarm fires silently.

This is a common operational oversight.

Exam trap

DOP-C02 often tests the assumption that an alarm in ALARM state automatically sends notifications, ignoring that SNS subscriptions must be confirmed and that alarm actions can be disabled or misconfigured.

How to eliminate wrong answers

Option B is wrong because if the EC2 instance were stopped, the CPUUtilization metric would either stop being published or drop to zero, not average 90%; the alarm is already in ALARM state, so the instance is running and consuming CPU. Option C is wrong because a 300-second period is a standard and valid evaluation window; it does not prevent notifications—it only affects how quickly the alarm evaluates. Option D is wrong because the alarm is explicitly in ALARM state, meaning it has sufficient data to evaluate and has breached the threshold; insufficient data would result in INSUFFICIENT_DATA state, not ALARM.

44
MCQmedium

A DevOps engineer is troubleshooting a production AWS Lambda function that occasionally times out. The function has a timeout of 30 seconds and uses a synchronous invocation. The engineer wants to capture invocation logs to identify the cause. Which approach will provide the MOST detailed diagnostic information?

A.Enable AWS CloudTrail data events for Lambda.
B.Create a CloudWatch dashboard with function duration metrics.
C.Add more logging statements to the function code and check CloudWatch Logs.
D.Enable AWS X-Ray tracing on the Lambda function.
AnswerD

AWS X-Ray tracing on the Lambda function provides an end-to-end view of the invocation, showing each subsegment's duration, including the time spent in external HTTP calls, AWS SDK operations, and service integrations. It automatically captures the trace for every invocation without requiring code changes (unless you need custom subsegments), and the trace timeline reveals exactly which downstream call or code block is consuming the most time. This makes X-Ray the correct tool to diagnose performance bottlenecks in a production Lambda, as it gives the per-invocation, subsegment-level detail that the other options lack.

Why this answer

AWS X-Ray provides end-to-end tracing for Lambda functions, capturing detailed timing information for each invocation, including subsegments for downstream calls, function initialization, and execution phases. This allows the engineer to pinpoint exactly where time is being spent, which is essential for diagnosing intermittent timeouts in synchronous invocations.

Exam trap

The trap here is that candidates often confuse CloudWatch Logs (which show custom log output) with X-Ray tracing (which provides automatic, detailed timing of every subcomponent), leading them to choose option C instead of the more diagnostic X-Ray approach.

How to eliminate wrong answers

Option A is wrong because CloudTrail data events for Lambda only record API calls (e.g., Invoke, UpdateFunctionConfiguration) and do not capture function execution logs or timing details needed to diagnose timeouts. Option B is wrong because a CloudWatch dashboard with duration metrics shows aggregated statistics (e.g., average, p99) over time, but cannot reveal per-invocation breakdowns or pinpoint the specific phase causing a timeout. Option C is wrong because adding more logging statements to the function code and checking CloudWatch Logs provides only custom log output without automatic tracing of downstream calls or sub-millisecond timing, making it insufficient for identifying intermittent timeout causes.

45
MCQmedium

A DevOps engineer supports a Python application on AWS Lambda that writes structured JSON logs to CloudWatch Logs. When a request fails, the engineer must be able to search all invocations across a log group for a specific request ID and see only lines containing that ID, returning results in seconds. Which approach meets this requirement with the least effort?

A.Create a CloudWatch Logs metric filter that matches the request ID pattern and graph the resulting metric on a dashboard.
B.Subscribe the log group to a Lambda function that scans events and writes matches to an Amazon DynamoDB table.
C.Use CloudWatch Logs Insights with a query that filters on the request ID field parsed from the JSON log events.
D.Export the log group to Amazon S3 and query it with Amazon Athena using a Glue table.
AnswerC

CloudWatch Logs Insights parses JSON log events automatically, so a query can filter on the request ID field and return matching lines in seconds across the entire log group. This requires no additional infrastructure, making it the lowest-effort solution for the stated search requirement.

Why this answer

CloudWatch Logs Insights is purpose-built for interrogating log groups with a query language that understands JSON, so filtering on a request ID field returns the exact log lines quickly. It works against the existing log group without new infrastructure, which satisfies both the speed and minimal-effort constraints.

Exam trap

The trap here is confusing metric filters, which only emit numeric metrics, with query tools that can actually return matching log lines.

46
MCQeasy

A DevOps engineer needs to monitor the number of 4xx and 5xx HTTP errors returned by an Application Load Balancer (ALB). They want to set up a dashboard that shows the error count over the last 24 hours. Which CloudWatch metrics should they use?

A.Use the 'HTTPCode_Target_4XX_Count' and 'HTTPCode_Target_5XX_Count' metrics.
B.Use the 'RequestCount' metric with a statistic of 'ErrorCount'.
C.Use the 'HTTPCode_ELB_4XX_Count' and 'HTTPCode_ELB_5XX_Count' metrics.
D.Use the 'TargetResponseTime' metric and count the number of responses above 4 seconds.
AnswerA

These CloudWatch metrics are emitted specifically for responses returned by the registred targets, such as EC2 instances or ECS tasks, after the load balancer forwards the request. By applying a Sum statistic over your monitoring window, you get the exact number of 4xx and 5xx status codes produced by your backend application, which is precisely what you need to detect target-side errors. Unlike ELB-level metrics, these don't include errors generated by the load balancer itself.

Why this answer

The correct metrics are 'HTTPCode_Target_4XX_Count' and 'HTTPCode_Target_5XX_Count', which track HTTP errors returned by the targets behind the Application Load Balancer. Option B is incorrect because 'RequestCount' does not have an 'ErrorCount' statistic. Option C is incorrect because 'HTTPCode_ELB_4XX_Count' and 'HTTPCode_ELB_5XX_Count' are load balancer-level metrics that track errors generated by the ALB itself (e.g., due to misconfiguration), not the errors returned by targets.

Option D is incorrect because 'TargetResponseTime' measures latency, not error counts.

47
Multi-Selectmedium

A company is using Amazon CloudWatch Logs to store application logs. The security team requires that logs are encrypted at rest using a customer-managed KMS key. Which TWO steps must be taken to achieve this?

Select 2 answers
A.Recreate the log group after associating the key.
B.Add a statement to the KMS key policy that allows CloudWatch Logs to use the key.
C.Create a KMS grant to allow CloudWatch Logs to use the key.
D.Specify the KMS key ARN when creating each log stream.
E.Use the put-log-group-encryption API to associate the KMS key with the log group.
AnswersB, E

To use a customer-managed KMS key, the key policy must explicitly grant the CloudWatch Logs service principal (logs.<region>.amazonaws.com) the required cryptographic permissions: kms:Encrypt, kms:Decrypt, kms:ReEncrypt, kms:GenerateDataKey, and kms:DescribeKey. Without an Allow statement, CloudWatch Logs cannot decrypt the key for the log group, and the association or ingest operations will fail. The key policy is the sole authorization mechanism for this integration.

Why this answer

To encrypt CloudWatch Logs at rest with a customer-managed KMS key, you must add a statement to the KMS key policy granting CloudWatch Logs permission to use the key (option B). Then, use the put-log-group-encryption API to associate the KMS key with the log group (option E). Option C is incorrect because CloudWatch Logs uses key policies, not grants.

Option D is incorrect because you specify the key ARN at the log group level, not per log stream. Option A is incorrect because you do not need to recreate the log group; you can associate the key with an existing log group.

48
MCQhard

A DevOps team is troubleshooting a performance issue where an Amazon RDS for PostgreSQL instance's CPU utilization spikes every hour. The team suspects a specific query from an application. Which combination of tools can identify the problematic query?

A.CloudWatch Logs Insights and CloudWatch metrics.
B.Amazon RDS Performance Insights and Enhanced Monitoring.
C.VPC Flow Logs and Lambda.
D.CloudTrail and CloudWatch alarms.
AnswerB

RDS Performance Insights computes a Database Load metric in 'average active sessions' and breaks it down by wait state, host, and individual SQL statement, so you can immediately rank queries by their contribution to the bottleneck. Enhanced Monitoring supplements this by gathering hypervisor and operating-system metrics (CPU, memory, file descriptor pressure, and network throughput) at one-second to one-minute intervals through a separate agent on the RDS instance. Together they reveal not only which query is responsible, but whether the root cause is CPU saturation, memory contention, or disk I/O, enabling a targeted fix such as adding an index or scaling the instance class.

Why this answer

Amazon RDS Performance Insights is purpose-built to visualize database load (DBLoad) broken down by SQL statement, wait event, and user, so it can pinpoint the exact query causing hourly CPU spikes. Enhanced Monitoring provides OS-level metrics (per-process CPU, memory, disk I/O) at up to 1-second granularity, letting the team correlate the spike with the specific PostgreSQL backend process. Together they identify both the offending SQL and the resource it consumes.

Exam trap

DOP-C02 often tests the distinction between control-plane visibility (CloudTrail, CloudWatch metrics) and database-internal observability (Performance Insights, Enhanced Monitoring) — candidates pick CloudWatch because it 'monitors everything' but it cannot attribute CPU to a SQL statement.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs Insights only queries log text (e.g., PostgreSQL logs) and CloudWatch metrics only show aggregate instance-level CPU — neither attributes load to a specific SQL statement. Option C is wrong because VPC Flow Logs capture IP-level network metadata (source/dest/port/bytes) and Lambda is compute, neither of which inspects SQL inside the database. Option D is wrong because CloudTrail records AWS API calls (control plane), not database query execution, and CloudWatch alarms only notify on thresholds rather than identify the query.

49
MCQeasy

A DevOps team wants to monitor the disk space utilization on their EC2 instances. What is the simplest way to achieve this?

A.Use AWS Systems Manager Inventory to collect disk space data.
B.Install the CloudWatch agent on the EC2 instances and configure the disk metric.
C.Enable EC2 detailed monitoring in CloudWatch.
D.Use EC2 basic monitoring in CloudWatch.
AnswerB

The CloudWatch agent runs inside the guest OS and reads disk space metrics such as disk_used_percent, disk_free, and disk_used from the filesystem, then publishes them as custom metrics under the System/Linux namespace (or appropriate custom namespace) to CloudWatch. You configure the agent via the amazon-cloudwatch-agent-config.json file or AWS Systems Manager Parameter Store, and once installed it can deliver both default host metrics and in-guest disk metrics. This is the only option that actually exposes the guest-OS-level disk space data to CloudWatch, enabling you to create alarms and dashboards on values like available disk capacity.

Why this answer

The CloudWatch agent can collect disk space metrics from EC2 instances, which is not available with default CloudWatch metrics. Option A is incorrect because AWS Systems Manager Inventory is used for software inventory and configuration, not real-time disk space monitoring. Option C is incorrect because EC2 detailed monitoring only provides more frequent metrics for standard EC2 metrics (CPU, network, etc.), not disk metrics.

Option D is incorrect because EC2 basic monitoring also does not include disk metrics.

50
MCQmedium

A company is using a centralized logging solution with Amazon OpenSearch Service. The DevOps team notices that logs from some EC2 instances are missing. The CloudWatch agent is installed and configured on all instances. What should the team do to troubleshoot the issue?

A.Check the CloudWatch agent status using the CloudWatch agent status command.
B.Configure a Lambda function to poll the CloudWatch agent for logs.
C.Verify that the EC2 instances have an SQS queue configured for log delivery.
D.Check the CloudWatch agent log file located at /var/log/amazon/amazon-cloudwatch-agent/amazon-cloudwatch-agent.log.
AnswerD

The CloudWatch agent logs its own operational activity—including configuration errors, permission issues, network timeouts, and partial failures—to `/var/log/amazon/amazon-cloudwatch-agent/amazon-cloudwatch-agent.log`. When log events are not appearing in CloudWatch Logs, this file is the authoritative source for finding the root cause. It records detailed error messages and stack traces that are not available in the agent status output or anywhere else, making it the first place to look when troubleshooting a missing log delivery.

Why this answer

The CloudWatch agent writes detailed operational logs to /var/log/amazon/amazon-cloudwatch-agent/amazon-cloudwatch-agent.log. This file contains errors, warnings, and debug messages that can reveal why logs from specific EC2 instances are not being delivered to Amazon OpenSearch Service. Checking this log is the first and most direct troubleshooting step because it captures agent-level issues such as configuration errors, network connectivity failures, or permission problems.

Exam trap

The trap here is that candidates may assume a 'status' command exists for the CloudWatch agent (Option A) because many other AWS services have such commands, but the agent uses a control script instead, and the real diagnostic starting point is the agent's own log file.

How to eliminate wrong answers

Option A is wrong because the CloudWatch agent does not have a 'status' command; the correct command to check the agent's operational state is 'sudo /opt/aws/amazon-cloudwatch-agent/bin/amazon-cloudwatch-agent-ctl -m ec2 -a status'. Option B is wrong because polling the CloudWatch agent with a Lambda function is unnecessary and inefficient; the agent already pushes logs to CloudWatch Logs, and the issue is about missing logs, not about needing to pull them. Option C is wrong because SQS queues are not used for log delivery from the CloudWatch agent; logs are sent directly to CloudWatch Logs via the HTTPS API, and SQS is unrelated to this data path.

51
Multi-Selecthard

A platform team runs Amazon EKS clusters across three AWS accounts and wants centralized observability. They need (1) container-level CPU and memory metrics with custom dimensions for namespace and workload, and (2) application logs from all pods shipped to a single destination, queryable with a structured query language. Which TWO actions should the team take? (Choose two.)

Select 2 answers
A.Enable AWS CloudTrail data events on the EKS cluster to capture pod-level container metrics from the Kubernetes API server.
B.Deploy Prometheus with the CloudWatch agent's Prometheus scraping feature and rely on CloudWatch Logs for log aggregation without Fluent Bit.
C.Configure Fluent Bit on each node to send pod logs to a CloudWatch Logs log group, and use CloudWatch Logs Insights for structured queries.
D.Deploy the CloudWatch agent with the Amazon CloudWatch Observability EKS add-on to collect Container Insights metrics with the pod and namespace dimensions.
E.Use AWS Config conformance packs to aggregate pod logs from all three accounts into a single queryable destination.
AnswersC, D

Fluent Bit shipping pod logs to a centralized CloudWatch Logs log group provides a single destination across accounts when combined with cross-account log sharing. CloudWatch Logs Insights supports a structured query language with fields, filter, stats, and parse commands, meeting the structured-query requirement for application logs.

Why this answer

Container Insights via the CloudWatch Observability EKS add-on supplies the CPU and memory metrics with namespace and workload dimensions, while Fluent Bit log shipping to CloudWatch Logs provides the centralized, queryable log destination. Together they satisfy both the metric-dimension and structured-log-query requirements across the three accounts using native AWS observability services.

Exam trap

The trap here is assuming Prometheus scraping alone satisfies the log requirement, when the scenario explicitly needs pod logs aggregated and queryable.

52
MCQmedium

A DevOps engineer is designing a monitoring solution for an application that runs on Amazon EC2 instances in an Auto Scaling group. The engineer needs to collect memory utilization metrics and visualize them in a dashboard. What should the engineer do?

A.Create a custom CloudWatch metric namespace and publish memory data using the AWS CLI.
B.Install the Amazon CloudWatch agent on the EC2 instances to collect memory metrics and publish them to CloudWatch.
C.Enable detailed monitoring on the EC2 instances to collect memory metrics.
D.Use AWS CloudTrail to capture memory utilization events from the EC2 instances.
AnswerB

The Amazon CloudWatch agent runs inside the guest OS and collects memory metrics such as mem_used_percent, mem_available, and mem_total, publishing them as custom metrics under the CWAgent namespace. It uses an IAM role for secure PutMetricData calls, can be configured via JSON, and can be centrally managed with AWS Systems Manager, making it the standard solution for OS-level metric collection.

Why this answer

The Amazon CloudWatch agent is required to collect memory utilization metrics from EC2 instances because memory usage is not a default metric provided by CloudWatch. By installing the CloudWatch agent on the instances, it can gather memory metrics (and other custom metrics) and publish them to CloudWatch, where they can be visualized in dashboards. This is the standard AWS-recommended approach.

Exam trap

DOP-C02 often tests the difference between default and custom metrics; the trap is assuming that detailed monitoring includes memory, or that CloudTrail can capture resource utilization, when in fact only the CloudWatch agent can provide memory metrics.

How to eliminate wrong answers

Option A is wrong because while you can create a custom metric namespace and publish memory data using the AWS CLI, this requires you to manually collect and push the data, which is not scalable or automated; the CloudWatch agent does this automatically. Option C is wrong because enabling detailed monitoring only increases the frequency of default metrics (from 5 minutes to 1 minute) but does not include memory metrics; memory is not a default metric. Option D is wrong because AWS CloudTrail records API activity, not memory utilization events; it is not a monitoring tool for resource utilization.

53
MCQeasy

A company uses AWS X-Ray to trace requests through its microservices application. The DevOps engineer notices that some traces are incomplete. What is a possible reason?

A.The X-Ray daemon is not running on the application servers.
B.X-Ray cannot trace requests that cross multiple AWS services.
C.The X-Ray SDK sampling rate is configured too low, causing many requests to be skipped.
D.X-Ray requires the CloudWatch agent to be installed on all EC2 instances.
AnswerC

The sampling rate directly controls the proportion of incoming requests that produce trace data. If the X-Ray SDK is configured with a low sampling rate (e.g., a small fixed percentage or a small reservoir per second), the majority of requests are intentionally skipped, resulting in sparse and often incomplete trace coverage. This matches the symptom of 'some requests traced, others not,' and is a common root cause for missing segments. Adjusting the sampling rule to a higher rate, or using a rate-based strategy tailored to traffic volume, would capture more complete traces.

Why this answer

The X-Ray SDK uses a sampling rate to decide which requests to record. If the sampling rate is set too low, a large percentage of requests are skipped, leading to incomplete traces. The DevOps engineer would observe missing segments for requests that were not sampled, even though the daemon and SDK are functioning correctly.

Exam trap

The trap here is that candidates often assume incomplete traces are due to infrastructure issues (daemon not running) rather than a configuration parameter (sampling rate), which is a subtle but common cause in distributed tracing.

How to eliminate wrong answers

Option A is wrong because if the X-Ray daemon were not running, the engineer would likely see no traces at all or errors in the SDK logs, not just incomplete traces. Option B is wrong because X-Ray is specifically designed to trace requests across multiple AWS services (e.g., API Gateway, Lambda, DynamoDB) using trace headers and service maps. Option D is wrong because X-Ray does not require the CloudWatch agent; it uses its own daemon and SDK to send trace data directly to the X-Ray API.

54
MCQmedium

A company runs a production web application on Amazon EC2 instances in an Auto Scaling group behind an Application Load Balancer (ALB). The application is deployed across three Availability Zones. The DevOps team recently noticed that the application's error rate is spiking periodically, but they cannot correlate the spikes with any known deployments or changes. The team has enabled detailed CloudWatch metrics for the ALB and EC2, and they are using CloudWatch Logs for application logs. They also have AWS X-Ray enabled for tracing. The team observes that during error spikes, the ALB's 5XX count increases, but the EC2 instance-level CPU and memory metrics remain normal. The application logs show 'Connection timed out' errors. The team suspects the issue is related to network connectivity but is not sure. Which course of action should the DevOps team take to identify the root cause of the periodic error spikes?

A.Enable VPC Flow Logs for the subnets and analyze the logs to identify dropped connections during the error spikes.
B.Increase the EC2 instance size to handle higher traffic and reduce timeouts.
C.Configure a step scaling policy for the Auto Scaling group based on ALB 5XX count.
D.Enable ALB access logs and analyze the 5xx response patterns.
AnswerA

VPC Flow Logs capture interface-level metadata for all IP traffic, including source and destination IPs, ports, protocol, and whether the action was accepted or rejected. During error spikes, analyzing Flow Logs via CloudWatch Logs Insights or Athena can reveal if connections to the instances are being blocked by security group rules or network ACLs, or if packets are being dropped before reaching the target. This directly identifies the root cause of ALB 5xx errors caused by network connectivity failures, rather than application-level issues.

Why this answer

VPC Flow Logs capture metadata about IP traffic going to and from network interfaces in a VPC, including whether the traffic was accepted or rejected. Since the application logs show 'Connection timed out' errors and instance-level metrics are normal, the issue likely lies in the network path (e.g., security groups, NACLs, or subnet routing) rather than the application or compute layer. Analyzing VPC Flow Logs during the error spikes will reveal if connections are being dropped or rejected, pinpointing the root cause of the timeouts.

Exam trap

The trap here is that candidates often jump to scaling or access logs (options C or D) because they focus on the 5XX error symptom, but the question specifically points to network-level timeouts, making VPC Flow Logs the only diagnostic tool that can reveal dropped or rejected packets at the network layer.

How to eliminate wrong answers

Option B is wrong because increasing EC2 instance size addresses compute resource constraints (CPU/memory), but the metrics show those are normal, so the timeouts are not due to resource exhaustion. Option C is wrong because configuring a step scaling policy based on ALB 5XX count would only react to the symptom (error rate) by adding instances, but it does not diagnose the underlying network connectivity issue causing the timeouts. Option D is wrong because ALB access logs record HTTP request/response details (e.g., status codes, timestamps) but do not capture network-level drops or rejections; they would show 5xx errors but not explain why connections are timing out at the network layer.

55
MCQmedium

A DevOps engineer is setting up monitoring for an Amazon S3 bucket that stores sensitive data. The engineer needs to be notified whenever an object in the bucket is accessed by a user or application, including read and write operations. Which AWS service should the engineer use to capture these events and trigger notifications?

A.Configure S3 event notifications to send events to an SNS topic for object-level operations.
B.Enable AWS CloudTrail data events for the S3 bucket and configure CloudWatch alarms on the log group.
C.Use AWS Config to record S3 resource changes and trigger an SNS notification.
D.Use Amazon CloudWatch metrics for the S3 bucket and set an alarm on the NumberOfObjects metric.
AnswerA

S3 event notifications are the most direct service for this use case: they emit near-real-time events for object-level operations such as s3:ObjectCreated:*, s3:ObjectRemoved:*, and s3:ObjectRestore:*, and can deliver them to an SNS topic without any polling or forensic analysis. By configuring an SNS topic as the destination, the DevOps engineer can receive immediate notifications for each action on objects, with optional prefix/suffix filtering to reduce noise. This provides operational awareness of who/what performed an operation, which CloudTrail and Config cannot do in real time.

Why this answer

Amazon S3 can be configured to send event notifications to SNS, SQS, or Lambda for object-level operations (e.g., PutObject, GetObject). This provides real-time notifications for read and write access. Option B is incorrect because CloudTrail data events capture object-level API calls but do not provide real-time notifications; they require additional setup with CloudWatch Logs and alarms.

Option C is incorrect because AWS Config records configuration changes, not individual object access events. Option D is incorrect because CloudWatch metrics track bucket-level statistics, not per-object access, and cannot trigger notifications for individual object access.

56
MCQeasy

A company is running a batch processing job on Amazon EMR that writes results to an Amazon S3 bucket. The job runs daily and takes about 2 hours. The DevOps team wants to be alerted if the job fails or takes longer than 3 hours. Which solution is the MOST cost-effective and operationally efficient?

A.Configure Amazon Simple Notification Service (SNS) directly from the EMR job to send notifications on completion.
B.Use Amazon CloudWatch Events to trigger an AWS Lambda function when the EMR cluster changes to 'TERMINATED' state, then check the job duration and send an alert if it exceeded 3 hours.
C.Create a CloudWatch alarm on the EMR cluster's EC2 instance CPUUtilization metric to detect abnormal runtime.
D.Use Amazon CloudWatch Logs to monitor the job's log stream and create a metric filter for 'FAILED' messages.
AnswerB

This is correct because Amazon EventBridge (formerly CloudWatch Events) can natively capture EMR cluster state-change events, such as the transition to TERMINATED, and invoke a Lambda function without any polling or custom code. The Lambda function can call emr:describeCluster to obtain the cluster's start and end times, compute the total runtime, and post an SNS message only when the duration exceeds three hours. This event-driven architecture is cost-effective, serverless, and decoupled, making it the recommended pattern for alerting on abnormal job completion times.

Why this answer

It uses CloudWatch Events to detect the EMR cluster's 'TERMINATED' state, which triggers a Lambda function that can check the job duration against the 3-hour threshold and send an alert via SNS if needed. This approach is cost-effective (no polling, event-driven) and operationally efficient, as it decouples monitoring from the job itself and handles both failure and timeout scenarios without modifying the EMR job code.

Exam trap

The trap here is that candidates often assume CloudWatch Logs metric filters (Option D) are the simplest way to detect failures, but they miss the timeout requirement and require log-based failure patterns, whereas event-driven state monitoring (Option B) inherently captures both failure and duration scenarios without custom logging.

How to eliminate wrong answers

Option A is wrong because configuring SNS directly from the EMR job requires modifying the job code to publish notifications, which is not operationally efficient and does not natively handle the 'takes longer than 3 hours' timeout condition—it only sends a completion notification, not an alert for excessive duration. Option C is wrong because CPUUtilization metrics are not a reliable indicator of job runtime or failure; a job can fail or run long without abnormal CPU usage, and this approach would require complex threshold tuning and still miss job-specific failures. Option D is wrong because using CloudWatch Logs metric filters for 'FAILED' messages only detects explicit failure log entries, not the timeout condition (job running >3 hours), and it requires the job to write specific log messages, which may not be present for all failure modes.

57
Multi-Selecteasy

A company is using Amazon CloudWatch Logs to collect application logs. They need to search and analyze the logs in near real-time. Which TWO AWS services can be used to achieve this?

Select 2 answers
A.Amazon CloudWatch Logs Insights
B.Amazon CloudWatch Synthetics
C.Amazon Kinesis Data Analytics
D.Amazon Athena
E.Amazon OpenSearch Service
AnswersA, E

Amazon CloudWatch Logs Insights is a native query engine that runs SQL-like queries directly against log data already ingested into CloudWatch Logs, without requiring any data movement or additional infrastructure. It automatically parses structured JSON log events, discovers fields, and lets you use commands such as filter, stats, sort, and time-based aggregation for interactive troubleshooting and pattern analysis. As the logs are already collected in CloudWatch, this is the most immediate and cost-effective way to analyze application log data.

Why this answer

Amazon CloudWatch Logs Insights (option A) is correct because it is a purpose-built, interactive query engine within CloudWatch Logs that lets you search, filter, and analyze log data in near real-time using a specialized query syntax, with results returned in seconds. Amazon OpenSearch Service (option E) is also correct because it can ingest CloudWatch Logs (typically via a subscription filter to Lambda or Kinesis) and provides near real-time search, dashboards, and analytics on that log data. Option B, Amazon CloudWatch Synthetics, is wrong because it creates canaries to monitor endpoints and availability, not to search or analyze log content.

Option C, Amazon Kinesis Data Analytics, is wrong because it is for running SQL/Flink stream processing over streaming data, not for ad-hoc log search and analysis. Option D, Amazon Athena, is wrong because it queries data in S3 with SQL on a batch/interactive basis and does not natively search CloudWatch Logs in near real-time.

Exam trap

DOP-C02 often tests whether candidates confuse log analysis (Logs Insights, OpenSearch) with log-generating or monitoring services (Synthetics) or assume Athena can query CloudWatch Logs directly without an S3 export.

58
MCQeasy

An organization wants to ensure that all API calls made in their AWS account are logged for security analysis. Which AWS service should be enabled to meet this requirement?

A.AWS CloudTrail
B.AWS Config
C.Amazon CloudWatch Logs
D.VPC Flow Logs
AnswerA

AWS CloudTrail records API activity across the account, capturing the identity, time, source IP and request details for each call. Enabling it satisfies the stem's requirement that all API calls be logged for security analysis, providing the audit trail other services do not.

Why this answer

AWS CloudTrail is the service that logs all API calls made in an AWS account, providing a record of actions taken by users, roles, or AWS services. It captures API activity across the AWS Management Console, SDKs, CLI, and other services. Enabling CloudTrail meets the requirement to log all API calls for security analysis.

Exam trap

DOP-C02 often tests the difference between CloudTrail, Config, and CloudWatch, and candidates may confuse API logging with configuration tracking or performance monitoring.

How to eliminate wrong answers

Option B is wrong because AWS Config records resource configurations and changes, not API calls. Option C is wrong because Amazon CloudWatch Logs is a log storage and analysis service, but it does not automatically log API calls; you would need to ingest CloudTrail logs into CloudWatch Logs for analysis. Option D is wrong because VPC Flow Logs capture IP traffic metadata, not API calls.

59
MCQmedium

A company uses Amazon CloudWatch to monitor its EC2 instances. A DevOps engineer notices that some metrics (e.g., memory utilization) are not available in the CloudWatch console. The engineer wants to collect these metrics. What should the engineer do?

A.Enable VPC Flow Logs for the instance's subnet.
B.Install and configure the CloudWatch Agent on the EC2 instances.
C.Send the metrics to CloudWatch Logs and create a metric filter.
D.Enable detailed monitoring on the EC2 instances.
AnswerB

The Amazon CloudWatch Agent is a software component installed inside the EC2 instance that collects operational metrics such as memory utilization, disk space and inode usage, and per-disk I/O statistics from the guest operating system. It can also collect custom application metrics and logs, and it sends them to CloudWatch via the PutMetricData API. This is the standard solution for monitoring metrics that are not exposed by the hypervisor, like memory usage, and it is the correct approach here.

Why this answer

Memory utilization and disk space metrics are not available by default in CloudWatch because they are OS-level metrics, not hypervisor-level. The CloudWatch Agent must be installed on EC2 instances to collect these custom metrics and publish them to CloudWatch.

Exam trap

DOP-C02 often tests the misconception that enabling detailed monitoring adds memory metrics — detailed monitoring only increases frequency, not metric scope; the CloudWatch Agent is required for OS-level metrics.

How to eliminate wrong answers

Option A is wrong because VPC Flow Logs capture IP traffic metadata for network analysis, not OS-level metrics like memory or disk. Option C is wrong because sending metrics to CloudWatch Logs and using metric filters is a workaround for log-derived metrics, not the standard approach for OS metrics — the CloudWatch Agent publishes metrics directly. Option D is wrong because detailed monitoring only increases the frequency of existing EC2 metrics (from 5 minutes to 1 minute); it does not add memory or disk metrics.

60
Multi-Selecthard

A company is running a production application on Amazon ECS with Fargate. The DevOps team needs to monitor the application's performance and set up alerts for high memory usage. Which THREE steps should the team take to achieve this?

Select 3 answers
A.Configure CloudWatch to automatically collect memory metrics from Fargate tasks
B.Create a custom CloudWatch metric and publish memory usage data from the application
C.Set up a CloudWatch alarm on the custom memory metric with appropriate threshold and actions
D.Enable the ECS task metadata endpoint and configure the application to publish memory metrics to CloudWatch
E.Enable Amazon CloudWatch Container Insights for the ECS cluster
AnswersB, C, D

The application itself can read its memory usage from the container's cgroup files (e.g., /sys/fs/cgroup/memory/memory.usage_in_bytes) or from runtime APIs, and then call CloudWatch PutMetricData to publish a custom metric. This gives you full control over the metric's namespace, dimensions (e.g., service, task ARN), and granularity. Unlike built-in or container-level metrics, custom metrics are the only way to capture memory from within the Fargate task process.

Why this answer

Options B, C, and D are correct. To monitor memory usage in ECS Fargate, you must enable the ECS task metadata endpoint (option D), which allows the container to access task metadata and publish custom memory metrics to CloudWatch. Then create a custom CloudWatch metric (option B) and set up a CloudWatch alarm on that metric (option C).

Option A is incorrect because CloudWatch does not automatically collect memory metrics from Fargate tasks; custom metrics are required. Option E is incorrect because CloudWatch Container Insights provides visibility into container-level metrics but does not collect memory metrics from Fargate tasks without additional configuration; custom metrics are still needed for memory monitoring.

61
Multi-Selectmedium

A DevOps team is setting up centralized logging for a multi-account AWS environment. They want to aggregate logs from all accounts into a single S3 bucket. Which services should be used to achieve this? (Choose TWO.)

Select 2 answers
A.AWS CloudTrail
B.AWS Config
C.Amazon CloudWatch Logs
D.Amazon Kinesis Data Firehose
E.Amazon S3 replication
AnswersA, C

AWS CloudTrail is the correct primary service for multi-account log aggregation because it supports organization trails, which automatically capture API activity from every account in an AWS Organization and deliver the event history to a single centralized S3 bucket. This eliminates the need to configure individual trails per account, making it the native and most efficient solution for aggregating AWS API logs across accounts.

Why this answer

AWS CloudTrail is correct because it can be configured to deliver log files from multiple accounts to a single S3 bucket by setting up a trail in the management account and using the 'Enable for all accounts in my organization' option, which automatically applies the trail to all member accounts in AWS Organizations. This centralizes CloudTrail logs without requiring per-account configuration.

Exam trap

The trap here is that candidates often confuse Amazon Kinesis Data Firehose as a log aggregation service, but it is a delivery stream that requires a log source to send data to it, not a service that natively collects logs from multiple accounts.

62
Multi-Selecteasy

A DevOps engineer is investigating a performance issue in a serverless application using AWS Lambda. The engineer wants to view the duration of each invocation and identify cold starts. Which TWO AWS services should be used? (Choose TWO.)

Select 2 answers
A.Amazon CloudWatch Metrics for Lambda (duration, invocations)
B.AWS X-Ray to trace invocations and detect cold starts
C.AWS CloudTrail to record Lambda API calls
D.Amazon CloudWatch Logs for Lambda execution logs
E.AWS Config to track Lambda configuration changes
AnswersB, D

AWS X-Ray can trace Lambda requests and annotate cold starts when tracing is enabled, but it requires the function to use the X-Ray SDK and sampling to be active. X-Ray focuses on distributed request flows across services, not on measuring the execution performance of the Lambda function itself in aggregate. For an initial performance investigation, CloudWatch Logs and Metrics give the definitive invocation and duration data without extra instrumentation.

Why this answer

Option B (AWS X-Ray) is correct because X-Ray traces individual Lambda invocations end-to-end, showing per-invocation duration via subsegments and explicitly flagging cold starts with an 'Initialization' subsegment, which is exactly what the engineer needs to identify cold starts. Option D (Amazon CloudWatch Logs) is correct because Lambda automatically streams each invocation's execution logs, including the REPORT line with Duration, Billed Duration, Memory Size, and Init Duration (the latter indicating a cold start), so per-invocation duration and cold-start evidence can be extracted from these logs. Option A (CloudWatch Metrics) is not correct here because Lambda metrics such as Duration and Invocations are aggregated at one-minute granularity and do not expose per-invocation duration or cold-start identification.

Option C (AWS CloudTrail) is not correct because it only records control-plane API calls (e.g., CreateFunction, Invoke via management events) and does not provide invocation duration or cold-start data. Option E (AWS Config) is not correct because it tracks Lambda configuration/resource changes for compliance, not runtime performance or invocation timing.

Exam trap

Candidates often choose Amazon CloudWatch Metrics because it includes a 'duration' metric, but this metric is aggregated across invocations and does not reveal individual invocation durations or definitively identify cold starts. The correct approach is to use X-Ray for tracing and CloudWatch Logs for per-invocation execution details.

63
MCQhard

A company runs a containerized application on Amazon ECS Fargate. The DevOps team wants to collect custom application metrics (e.g., request count, error rate) and send them to Amazon CloudWatch. The team wants to minimize changes to the application code. Which solution should be used?

A.Have the application call the CloudWatch PutMetricData API directly.
B.Run the CloudWatch agent as a sidecar container in the ECS task definition, configured to collect StatsD metrics from the application container.
C.Use the ECS agent's built-in metric collection feature.
D.Modify the application to send logs using the embedded metric format.
AnswerB

The CloudWatch agent can run as a sidecar container in the same ECS task, and with the default awsvpc network mode all containers share a localhost network namespace. Configure the agent with a StatsD stanza so it listens on port 8125, and the application simply sends UDP messages in the StatsD format—no SDK, no logging changes, and no application code modifications. The agent buffers, batches, and forwards these metric points to CloudWatch as a single API consumer, preserving the existing application behavior.

Why this answer

The CloudWatch agent can run as a sidecar container in the same ECS task definition and listen for StatsD metrics (over UDP port 8125) from the application container. This approach requires zero changes to the application code—the application simply emits StatsD-formatted metrics, and the CloudWatch agent forwards them to CloudWatch via the PutMetricData API. It minimizes operational overhead while enabling custom metric collection from containerized workloads on Fargate.

Exam trap

The trap here is that candidates often assume the ECS agent or CloudWatch agent must be installed on the host, but in Fargate, the sidecar pattern is the only way to run the CloudWatch agent without modifying the application code.

How to eliminate wrong answers

Option A is wrong because it requires modifying the application code to call the CloudWatch PutMetricData API directly, which contradicts the requirement to minimize code changes. Option C is wrong because the ECS agent's built-in metric collection feature only gathers infrastructure-level metrics (CPU, memory, network) for the task, not custom application metrics like request count or error rate. Option D is wrong because it requires modifying the application to emit logs in the embedded metric format and then configuring a log subscription filter to extract metrics, which still involves code changes and adds complexity compared to the sidecar approach.

64
Multi-Selecthard

A DevOps engineer needs to set up a monitoring solution for an application running on Amazon EKS. The application emits custom metrics that need to be stored in Amazon CloudWatch and visualized on a dashboard. Which THREE steps should the engineer take? (Choose THREE.)

Select 3 answers
A.Configure the CloudWatch agent to emit custom metrics to CloudWatch.
B.Use CloudWatch Logs Insights to analyze the custom metrics.
C.Create a CloudWatch dashboard to visualize the collected metrics.
D.Install the CloudWatch agent on the EKS cluster using a DaemonSet.
E.Use Amazon Managed Service for Prometheus to scrape the metrics.
AnswersA, C, D

The CloudWatch agent is the correct mechanism for collecting custom metrics from an EKS cluster and emitting them to CloudWatch. By configuring the agent with a metrics collection interval and a custom namespace, you can send application and cluster-level metrics (e.g., pod CPU, memory, or custom application counters) via the PutMetricData API. This is the foundational step that makes those metrics available for dashboards, alarms, and further analysis within CloudWatch, so it is a required part of a CloudWatch-centric monitoring solution.

Why this answer

The CloudWatch agent can be configured to emit custom application metrics to Amazon CloudWatch, which is the required destination for storing the metrics. The agent uses the CloudWatch PutMetricData API to send these metrics, enabling centralized monitoring and alerting within CloudWatch.

Exam trap

The trap here is that candidates may confuse CloudWatch Logs Insights (for logs) with CloudWatch Metrics (for numeric data), or assume Amazon Managed Service for Prometheus is a direct replacement for CloudWatch metrics, when the question specifically requires storing custom metrics in CloudWatch.

65
MCQeasy

A company is running a critical application on Amazon RDS for PostgreSQL. The DevOps team needs to set up monitoring to detect when database connections exceed 80% of the maximum connections for more than 5 minutes. Which CloudWatch metric should be used to create an alarm?

A.DatabaseConnections
B.FreeableMemory
C.CPUUtilization
D.DiskQueueDepth
AnswerA

DatabaseConnections reports the count of client sessions currently connected to the RDS PostgreSQL instance, so an alarm threshold set at 80% of max_connections with a five-minute evaluation period detects sustained connection saturation, exactly the condition the DevOps team must monitor.

Why this answer

Amazon RDS publishes the DatabaseConnections metric to CloudWatch, representing the number of client connections currently open to the DB instance. To alarm at 80% of max connections, you compare DatabaseConnections against the instance's max_connections parameter. This is the only listed metric that directly measures connection count.

Exam trap

The trap here is assuming CPUUtilization or FreeableMemory reflects connection saturation; candidates must recognize DatabaseConnections as the direct metric for connection limits.

How to eliminate wrong answers

Option B is wrong because FreeableMemory measures available RAM, which can correlate with connection pressure but does not count connections. Option C is wrong because CPUUtilization measures processor load and may stay low even when connections are near the limit. Option D is wrong because DiskQueueDepth measures pending disk I/O operations, unrelated to connection saturation.

66
MCQmedium

Your company runs a multi-tier web application on AWS. The application consists of an Application Load Balancer (ALB) that distributes traffic to a fleet of Amazon EC2 instances running a web server. The web servers write access logs to a shared Amazon EFS filesystem. The operations team needs to monitor the web server logs in real-time to detect and alert on 5xx error spikes. Currently, the team manually SSHes into instances to tail logs, which is inefficient and doesn't provide real-time alerting. The team wants a centralized, near-real-time logging solution with minimal operational overhead. They have asked you to design a solution that ingests logs from the EFS filesystem into a centralized log analytics platform. Which solution would you recommend?

A.Enable AWS CloudTrail data events for the EC2 instances to capture log file modifications.
B.Configure an Amazon EventBridge scheduled rule to invoke an AWS Lambda function that reads new log lines from EFS and publishes them to Amazon CloudWatch Logs.
C.Stream the log files to Amazon Kinesis Data Streams using a custom producer, then use a Lambda function to analyze and alert on 5xx errors.
D.Install and configure the Amazon CloudWatch Logs agent on each EC2 instance to tail the log files from the EFS mount and send them to CloudWatch Logs. Create a metric filter and alarm for 5xx errors.
AnswerD

The Amazon CloudWatch Logs agent (now part of the unified CloudWatch agent) can be installed on each EC2 instance to monitor the EFS-mounted log file and push new lines to CloudWatch Logs in near-real-time. After the log group receives the entries, a metric filter can extract the '5xx' HTTP status code pattern to create a custom metric, and a CloudWatch alarm on that metric will page the team when the error rate breaches a threshold. This is the purpose-built, low-overhead solution that supports tailing, rotation, and automatic delivery.

Why this answer

Installing the CloudWatch Logs agent on each EC2 instance allows it to tail the log files from the shared EFS mount point and stream them to CloudWatch Logs in near real-time. This provides centralized log ingestion with minimal operational overhead, and you can create a metric filter and alarm to detect and alert on 5xx error spikes without manual SSH access.

Exam trap

The trap here is that candidates may overcomplicate the solution by choosing Kinesis or Lambda-based approaches (Options B and C) when a simple agent-based solution (Option D) is sufficient, or they may confuse CloudTrail data events (Option A) with log file monitoring, not realizing CloudTrail captures API activity, not file content changes.

How to eliminate wrong answers

Option A is wrong because CloudTrail data events for EC2 instances capture API calls (e.g., RunInstances, TerminateInstances), not log file modifications on EFS; they cannot ingest or analyze web server log content. Option B is wrong because an EventBridge scheduled rule with a Lambda function that reads new log lines from EFS would introduce latency (scheduled intervals) and complexity in tracking file offsets, making it unsuitable for near-real-time monitoring. Option C is wrong because streaming logs to Kinesis Data Streams requires a custom producer to be deployed and managed, adding significant operational overhead compared to the agent-based approach, and it does not directly integrate with CloudWatch Logs for metric filtering and alerting without additional Lambda processing.

67
Multi-Selecthard

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The application logs show that some requests are timing out. The team needs to identify the source of the issue. Which TWO steps should they take?

Select 2 answers
A.Enable ALB access logs and analyze them.
B.Enable VPC Flow Logs to capture network traffic.
C.Enable AWS WAF logs to inspect HTTP requests.
D.Review CloudWatch metrics for the ALB, such as 'RequestCount' and 'TargetResponseTime'.
E.Enable AWS CloudTrail to log all API calls.
AnswersA, D

ALB access logs are the authoritative source for request-level diagnostics because they capture every HTTP request processed by the load balancer, including the target processing time, request processing time, HTTP status, and client/user-agent data. Analyzing these logs lets you identify exactly which requests experienced slow responses or timeouts, isolate problematic targets by IP or URL pattern, and correlate with backend health. Unlike aggregated metrics, access logs provide per-request granularity that is essential for root-causing intermittent timeout issues.

Why this answer

Option A is correct because ALB access logs capture detailed per-request information such as request processing time, target response time, and the specific error codes (e.g., 504 Gateway Timeout) returned by the load balancer, which directly helps pinpoint whether timeouts originate at the ALB or the backend targets. Option D is correct because CloudWatch metrics for the ALB, particularly 'TargetResponseTime' and 'RequestCount', reveal latency trends and traffic patterns that indicate whether targets are slow or overloaded, helping isolate the source of the timeouts. Option B is not appropriate because VPC Flow Logs only capture IP-level metadata (source/destination, ports, accept/reject) and cannot show HTTP-layer timing or application-level errors.

Option C is not appropriate because AWS WAF logs only record requests that match or are blocked by WAF rules, and the scenario does not indicate WAF is in use or that requests are being blocked. Option E is not appropriate because CloudTrail records AWS API calls for auditing, not application request behavior or latency.

Exam trap

DOP-C02 often tests the distinction between logging services: candidates may confuse VPC Flow Logs (network-level) with ALB access logs (application-level) or assume CloudTrail captures performance data, leading to wrong selections.

68
MCQhard

Refer to the exhibit. An alarm is configured as shown. The CPU utilization averages 85% for 10 minutes, then spikes to 95% for the next 5 minutes, and returns to 80%. How many times will the SNS topic receive a notification?

A.0
B.1
C.2
D.3
AnswerA

The CloudWatch alarm is configured with `EvaluationPeriods: 3` and `DatapointsToAlarm: 2`. During the monitoring window, the CPU utilization never produces two breaching datapoints within any three consecutive periods, so the alarm condition remains unsatisfied. Consequently, the alarm never leaves the OK state and thus never enters ALARM.

Why this answer

Per the exhibit: Period=300s (5 min), EvaluationPeriods=2, Threshold=90.0, ComparisonOperator=GreaterThanThreshold, no DatapointsToAlarm override (so 2 consecutive breaching 5-minute periods are required to enter ALARM). CPU averages 85% for two 5-minute periods (not breaching, since 85 is not > 90), then 95% for one 5-minute period (breaching), then returns to 80% (not breaching). The single breaching period is immediately preceded and followed by non-breaching periods, so the alarm never accumulates 2 consecutive breaches and never transitions into ALARM -- it stays in OK the entire time.

Since neither AlarmActions nor OKActions fire without an actual state transition, and none occurs, the SNS topic receives zero notifications.

Exam trap

The trap here is that candidates assume the alarm triggers immediately when the metric exceeds the threshold, but CloudWatch requires a specified number of consecutive evaluation periods (datapoints) to breach before changing state, and the threshold comparison is strict (greater than, not greater than or equal).

How to eliminate wrong answers

Option A (0) is wrong because the alarm does trigger after 3 consecutive periods above the threshold, sending one notification. Option C (2) is wrong because only one notification is sent when the alarm enters ALARM state; no notification is sent when it returns to OK because the scenario ends before the 3-period requirement for OK is met. Option D (3) is wrong because there is no repeated flapping or multiple state changes; the alarm transitions only once from OK to ALARM.

69
Multi-Selecteasy

A company uses AWS Lambda for data processing. The operations team wants to be alerted when a function fails. Which TWO methods can they use?

Select 2 answers
A.Configure S3 event notifications to trigger on Lambda errors.
B.Enable AWS CloudTrail to log Lambda invocations.
C.Configure a dead-letter queue (DLQ) for the Lambda function and monitor the queue.
D.Create a CloudWatch alarm on the 'Errors' metric for the Lambda function.
E.Use AWS Config to detect Lambda function failures.
AnswersC, D

For asynchronous invocations, Lambda can be configured with a dead-letter queue (an SQS queue or SNS topic) to receive event payloads that could not be processed after a function fails or exhausts retry attempts. The DLQ preserves the exact original event data, allowing you to inspect, replay, or archive failed records in a durable buffer. Monitoring the DLQ (for example, with a CloudWatch alarm on SQS ApproximateNumberOfMessagesVisible) directly alerts you to the presence of undelivered events. This is why the correct answer combines a DLQ with active monitoring of that queue.

Why this answer

Option D is correct because Lambda automatically publishes the 'Errors' metric to Amazon CloudWatch, and a CloudWatch alarm can be created on that metric to trigger an Amazon SNS notification when the error threshold is breached, directly alerting the operations team. Option C is correct because configuring a dead-letter queue (an Amazon SQS queue or SNS topic) for the Lambda function captures failed asynchronous invocations, and monitoring that queue (e.g., via CloudWatch metrics like ApproximateNumberOfMessagesVisible) provides an alerting mechanism for failures. Option A is incorrect because S3 event notifications only trigger Lambda invocations on object events; they do not detect or report Lambda execution errors.

Option B is incorrect because CloudTrail records API activity and management events, not Lambda function invocation failures, and it is not an alerting service. Option E is incorrect because AWS Config evaluates resource configuration compliance, not runtime function failures.

Exam trap

DOP-C02 often tests the confusion between monitoring services (CloudWatch) and logging/auditing services (CloudTrail, AWS Config), so candidates must recognize that only CloudWatch alarms and DLQ monitoring provide direct alerting on Lambda failures.

70
MCQeasy

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer. The DevOps team wants to receive an alert when the number of HTTP 5xx errors from the load balancer exceeds 100 in a 5-minute period. The team wants to use the most direct and least complex method. What should they do?

A.Install the CloudWatch agent on the EC2 instances to collect web server logs, and create a custom metric for 5xx errors with an alarm.
B.Configure AWS CloudTrail to log load balancer API calls and create an EventBridge rule to trigger an SNS notification when 5xx errors occur.
C.Create a CloudWatch alarm on the HTTPCode_ELB_5XX_Count metric for the load balancer with a threshold of 100 over 5 minutes, and configure an Amazon SNS topic for notifications.
D.Enable access logs on the load balancer, then use CloudWatch Logs metric filters to count 5xx errors and create an alarm on the resulting metric.
AnswerC

The Application Load Balancer publishes the HTTPCode_ELB_5XX_Count metric to CloudWatch. Creating an alarm directly on this metric with a 5-minute period and a threshold of 100, then attaching an SNS topic for email or SMS notifications, is the simplest and most direct way to meet the requirement. No additional agents or complex configurations are needed.

Why this answer

Application Load Balancers automatically publish the HTTPCode_ELB_5XX_Count metric to CloudWatch, so creating an alarm on that metric with an SNS notification is the most direct and least complex solution. The other options involve unnecessary complexity or use services that do not monitor HTTP errors.

Exam trap

The trap here is overcomplicating the solution by using access logs or CloudTrail, when a built-in CloudWatch metric already provides the needed data.

71
MCQeasy

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The application uses a custom health check endpoint '/health'. The DevOps team notices that the ALB is marking some instances as unhealthy even though the application is running fine. The team checks the security groups and network ACLs and confirms they allow traffic. What should the team check next?

A.Ensure the health check path is case-insensitive.
B.Increase the health check interval and timeout values.
C.Confirm that the health check path is correctly configured to '/health' on the target group.
D.Verify that the health check port matches the application port.
AnswerC

The target group health check path must exactly match the application endpoint that is designed to return an HTTP 200 OK when healthy. If the path is incorrectly set to something like '/' instead of '/health', the application may return a different status code, such as a redirect or a default page, which can cause intermittent health check failures as the application's behavior changes under load. Confirming the path resolves the root cause because the health check will then probe a purpose-built endpoint that consistently returns 200.

Why this answer

If security groups and NACLs are confirmed correct, the next most likely cause is a misconfigured health check path on the target group — for example, the path is set to '/' or '/healthz' instead of '/health', so the ALB receives a 404 and marks the instance unhealthy. Verifying the exact path string in the target group's health check settings is the correct next diagnostic step. The other options either do not apply or are premature tuning steps.

Exam trap

DOP-C02 often tests whether candidates jump to tuning health check timing parameters when the actual root cause is a simple configuration mismatch like the wrong path or success code, so candidates pick interval/timeout adjustments instead of verifying the path.

How to eliminate wrong answers

Option A is wrong because HTTP paths are case-sensitive by convention and ALB health checks do not treat paths as case-insensitive; this is not a real configuration knob. Option B is wrong because increasing interval and timeout values only delays failure detection — it does not fix a wrong path or a failing endpoint, and the team has no evidence of slow responses. Option D is wrong because the question states the application is running fine and security groups allow traffic; while port mismatch is a possible cause, the scenario points to path configuration, and the correct answer addresses the path specifically.

72
Multi-Selecteasy

A company wants to ensure that all changes to its Amazon S3 bucket policies are logged for auditing purposes. Which TWO AWS services should be enabled to capture these changes?

Select 2 answers
A.Amazon CloudWatch
B.AWS Config
C.Amazon GuardDuty
D.VPC Flow Logs
E.AWS CloudTrail
AnswersB, E

AWS Config is the correct service for this requirement because it continuously records and evaluates the configuration of AWS resources, including S3 bucket policies, and maintains a detailed configuration timeline. It can compare the recorded configuration against desired compliance rules (e.g., ensuring a bucket is not public). With AWS Config, you can see exactly when a bucket policy was last changed and what the previous configuration was, making it ideal for auditing and compliance.

Why this answer

Options B and E are correct because AWS Config records resource configuration changes, including S3 bucket policies, and AWS CloudTrail logs API calls such as PutBucketPolicy. Option A is incorrect because Amazon CloudWatch monitors operational metrics and logs, not auditing of policy changes. Option C is incorrect because Amazon GuardDuty provides threat detection, not audit logging.

Option D is incorrect because VPC Flow Logs capture network traffic, not configuration changes.

73
Multi-Selectmedium

A DevOps engineer is designing a monitoring solution for a multi-tier web application hosted on AWS. The application consists of an Application Load Balancer (ALB), EC2 instances, and an RDS database. The engineer needs to capture and analyze HTTP request logs from the ALB to understand client behavior and troubleshoot errors. Which THREE steps are necessary to achieve this?

Select 3 answers
A.Install the CloudWatch Agent on the ALB
B.Enable AWS CloudTrail for the ALB
C.Use Amazon Athena to query the access logs in S3
D.Enable access logs on the ALB
E.Create an Amazon S3 bucket to store the access logs
AnswersC, D, E

Amazon Athena is the correct service for interactively querying ALB access logs directly from S3 without loading data into a database. The logs are stored as gzipped text files in a partitionable layout (AWSLogs/account-id/elasticloadbalancing/region/yyyy/mm/dd), which can be registered as a table in Athena using JSON or CSV SerDe. With Athena's SQL, you can analyze request patterns, error rates, latency, and client behavior by writing queries against the log fields, and partitioning or partition projection keeps query costs low. This makes Athena the natural final step after enabling ALB access logs and storing them in S3.

Why this answer

Option D is correct because ALB access logs must be explicitly enabled on the load balancer, which captures detailed information about every HTTP/HTTPS request including client IP, request path, response codes, and latency. Option E is correct because ALB access logs are delivered to an Amazon S3 bucket, so a target S3 bucket (with the proper bucket policy allowing the ALB to write) must exist before enabling logging. Option C is correct because Amazon Athena can query the ALB access logs stored in S3 directly using SQL, enabling analysis of client behavior and troubleshooting of errors without loading data into a database.

Option A is incorrect because the CloudWatch Agent runs on EC2 instances or on-premises servers, not on ALBs, which are managed services that cannot host agents. Option B is incorrect because AWS CloudTrail records API activity and management events, not HTTP request logs, so it does not capture ALB access log data.

Exam trap

DOP-C02 often tests the confusion between CloudTrail (API audit logs) and ALB access logs (HTTP request logs), and the misconception that the CloudWatch Agent can be installed on managed services like ALB.

74
MCQhard

A company is migrating its on-premises applications to AWS and wants to maintain the same level of monitoring for its Linux-based EC2 instances. They currently use Nagios for monitoring. They want a managed AWS service that can monitor instance health, system metrics, and application logs. Which solution should they use?

A.Install the Amazon CloudWatch agent on each EC2 instance to collect system metrics and logs, and send them to CloudWatch.
B.Use AWS CloudTrail to monitor instance activity and capture log files.
C.Use AWS Systems Manager Inventory to collect system configuration and log files.
D.Use AWS Config to track instance configuration changes and trigger alerts.
AnswerA

The CloudWatch agent runs on each instance to collect system metrics and application logs, then ships them to CloudWatch. This delivers the managed monitoring service the company requires, replacing Nagios without self-managed infrastructure, and satisfies the instance health, metrics and logs monitoring requirement.

Why this answer

The CloudWatch agent is the managed AWS solution for collecting guest-level system metrics (memory, disk, swap) and application/system logs from EC2 instances and delivering them to CloudWatch. It replaces the need for a self-managed Nagios server by providing native metric and log ingestion, alarms, and dashboards. Installing it on each instance satisfies the requirement for instance health, system metrics, and application logs in a managed service.

Exam trap

The trap is confusing AWS monitoring services: CloudTrail (API audit), Config (configuration compliance), and Systems Manager Inventory (metadata) are often mistaken for a metrics/logs monitoring solution, but only the CloudWatch agent collects guest OS metrics and application logs.

How to eliminate wrong answers

Option B is wrong because CloudTrail records API activity and management events, not guest OS metrics or application log files, so it cannot replace Nagios-style monitoring. Option C is wrong because Systems Manager Inventory collects configuration metadata (installed applications, OS details) rather than real-time system metrics and application logs. Option D is wrong because AWS Config tracks resource configuration changes and compliance, not runtime health metrics or log content.

75
MCQhard

A company runs a containerized application on Amazon ECS with Fargate launch type. The application consists of three microservices: frontend, backend, and database. The ECS cluster is in a VPC with public and private subnets. The frontend service is publicly accessible via an Application Load Balancer (ALB) in public subnets. The backend service communicates with the database service, which runs as a stateful service with persistent storage using Amazon EFS. The DevOps team is using CloudWatch Container Insights and has enabled Prometheus metrics for the ECS cluster. Recently, the team observed that the frontend service's response time has increased significantly, and some requests are timing out. The team checked the ALB metrics and saw an increase in 5xx errors. They also noticed that the backend service's CPU utilization is high, and the database service's disk I/O is high. The team suspects a bottleneck in the backend service. Which course of action should the team take FIRST to identify the root cause?

A.Disable the health check for the backend service in the ALB target group.
B.Migrate the database service to Amazon RDS for better performance.
C.Check the backend service's application logs in CloudWatch Logs to identify errors or slow database queries.
D.Increase the desired count of the backend service to reduce load per task.
AnswerC

Checking the backend service's application logs in CloudWatch Logs is the correct initial action because it provides direct visibility into application errors, database query execution times, and slow transactional paths. These logs, combined with ECS task metrics and ALB access logs, help isolate whether the high latency is due to application code, database contention, or an upstream dependency. Logs are the least intrusive and most informative diagnostic step, enabling an evidence-based decision before changing infrastructure.

Why this answer

The first step is to analyze the backend service's application logs to identify any errors or slow operations. High CPU and disk I/O may be caused by inefficient queries or code issues. Option A is incorrect because disabling health checks would hide the problem and could route traffic to unhealthy tasks.

Option B is incorrect because migrating to RDS does not address the immediate issue and is a significant change without root cause analysis. Option D is incorrect because increasing the desired count without understanding the root cause may temporarily alleviate load but does not fix underlying performance issues and can increase costs.

Page 1 of 3 · 197 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Monitoring and Logging questions.