Courseiva

CCNA Monitoring and Logging Questions

47 of 197 questions · Page 3/3 · Monitoring and Logging · Answers revealed

151
MCQhard

A company runs a microservices application on Amazon ECS with Fargate. The operations team notices that some services are experiencing intermittent high latency, but CPU and memory metrics appear normal. They need to identify the root cause. Which approach should they use?

A.Enable detailed CloudWatch Logs and use CloudWatch Logs Insights to query logs for slow requests.
B.Use Amazon Managed Service for Prometheus to collect custom metrics and set up dashboards.
C.Set up CloudWatch Synthetics canaries to monitor the endpoints and measure response times.
D.Instrument the application with the AWS X-Ray SDK and use the X-Ray console to analyze traces.
AnswerD

Instrumenting with the AWS X-Ray SDK is the correct approach because X-Ray traces individual requests as they traverse services, generating a trace ID that propagates across HTTP headers and AWS SDK client calls. The X-Ray console provides a service map and trace timelines, segment and subsegment views that break down latency for each downstream call, letting you pinpoint the exact service or resource causing the slowdown. You can also annotate traces with request metadata and set sampling rules to balance overhead and observability.

Why this answer

Intermittent latency with normal CPU and memory metrics points to a distributed tracing problem — the bottleneck is likely in a downstream call, a network hop, or a specific service in the request chain. AWS X-Ray traces requests end-to-end across ECS tasks, Lambda functions, and downstream services, showing exactly where time is spent. This makes it the right tool to pinpoint the root cause of intermittent latency in a microservices architecture.

Exam trap

DOP-C02 often tests the distinction between metrics, logs, and traces — candidates pick CloudWatch Logs or Prometheus because they sound comprehensive, but only X-Ray provides the per-request, cross-service causality needed to diagnose intermittent latency.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs Insights can query logs for slow requests, but it requires the application to already log timing data and does not automatically correlate latency across service boundaries — it is reactive and manual rather than a tracing solution. Option B is wrong because Amazon Managed Service for Prometheus collects metrics, and metrics alone cannot explain why a specific request was slow — they show aggregates, not per-request causality. Option C is wrong because CloudWatch Synthetics canaries measure endpoint response times from the outside but cannot tell you which internal service or call caused the latency, so they detect the symptom rather than diagnose the cause.

152
MCQhard

A company uses AWS CloudFormation to deploy infrastructure. The security team wants to be notified whenever a stack is created, updated, or deleted. They also want to track who made the change. Which combination of services should be used to achieve this?

A.AWS Config rules and Amazon SNS
B.AWS CloudTrail and Amazon CloudWatch Events (now Events) with SNS
C.Amazon S3 event notifications and AWS Lambda
D.AWS Lambda and Amazon DynamoDB
AnswerB

AWS CloudTrail records all CloudFormation management-plane API calls as event payloads, including the calling identity, request parameters, and timestamp. Amazon CloudWatch Events (now EventBridge) can consume those CloudTrail events using a rule that matches on source: aws.cloudformation and specific event names like UpdateStack or DeleteStack. That rule can then route the matched event to an SNS topic, producing immediate, precise notifications of who performed the stack operation and what operation occurred. This is the native, event-driven pattern for CloudFormation activity monitoring.

Why this answer

CloudTrail captures CloudFormation API calls (CreateStack, UpdateStack, DeleteStack) and CloudWatch Events can trigger SNS notifications based on those API calls. Option A is wrong because Config rules evaluate resource compliance, not API events. Option C is wrong because S3 event notifications are for S3 objects.

Option D is wrong because Lambda alone cannot capture who made the change without CloudTrail integration.

153
Drag & Dropmedium

Drag and drop the steps to configure an AWS Auto Scaling group with a launch template and scaling policies.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

First create the launch template, then create the Auto Scaling group, then configure network, then set capacities, then add scaling policy.

154
MCQhard

A company is running a critical application on Amazon ECS with Fargate. The application generates custom metrics that are published to CloudWatch using the PutMetricData API. Recently, the metrics have been delayed by up to 5 minutes. The DevOps team needs to reduce the latency. What should the team do?

A.Install the CloudWatch agent on the Fargate tasks to collect metrics.
B.Set the StorageResolution parameter to 1 when calling PutMetricData.
C.Publish the metrics as structured logs to CloudWatch Logs and use metric filters.
D.Increase the frequency of PutMetricData calls to every 5 seconds.
AnswerB

Calling PutMetricData with StorageResolution=1 creates a high-resolution custom metric with 1-second granularity, making the data available for CloudWatch alarms in as little as 10 seconds instead of the 60-second standard resolution. This lower storage resolution is the key to detecting critical issues faster because CloudWatch can evaluate alarms at a 10- or 30-second period. Without this parameter, your metrics default to 60-second resolution and alarm latency remains as high as a minute.

Why this answer

Setting the StorageResolution parameter to 1 when calling PutMetricData enables high-resolution metrics with a 1-second granularity. This reduces the latency of metric ingestion and retrieval because CloudWatch processes high-resolution metrics more quickly than standard 60-second resolution metrics, addressing the 5-minute delay.

Exam trap

The trap here is that candidates may think increasing API call frequency or using log-based metrics will reduce latency, but the actual cause is the default 60-second storage resolution, which delays metric availability regardless of how often data is sent.

How to eliminate wrong answers

Option A is wrong because the CloudWatch agent cannot be installed on Fargate tasks; Fargate is a serverless compute engine that does not allow direct installation of agents, and metrics are already being published via PutMetricData, so the agent is unnecessary. Option C is wrong because publishing metrics as structured logs and using metric filters adds additional processing overhead and latency from log ingestion and filter evaluation, which would not reduce the delay and may increase it. Option D is wrong because increasing the frequency of PutMetricData calls to every 5 seconds does not change the resolution or ingestion latency; it may cause throttling from CloudWatch API limits and does not address the underlying delay caused by standard-resolution metric processing.

155
Multi-Selectmedium

A DevOps engineer needs to set up a monitoring solution that can detect and alert on unusual patterns in application metrics. Which TWO AWS services can be used together to achieve this? (Choose TWO.)

Select 2 answers
A.Amazon GuardDuty
B.Amazon CloudWatch Alarms
C.Amazon CloudWatch Anomaly Detection
D.AWS CloudTrail
E.AWS Config
AnswersB, C

Amazon CloudWatch Alarms are the action engine that watches a single CloudWatch metric, a math expression, or an anomaly detection band over a specified time period, then transitions to an ALARM state when the observed value breaches a defined threshold. You can configure the alarm to publish to an SNS topic, trigger Auto Scaling, or execute an EC2 action such as a reboot when the anomaly condition persists. In this solution, the alarm consumes the band produced by CloudWatch Anomaly Detection and calls the monitoring hook when unusual metric behavior is detected.

Why this answer

Amazon CloudWatch Anomaly Detection [CORRECT] is correct because it applies machine-learning algorithms to a metric's historical baseline and creates an expected-value band, so it can flag unusual patterns in application metrics without requiring manually tuned static thresholds. Amazon CloudWatch Alarms [CORRECT] is correct because it evaluates a metric or an anomaly-detection band against a defined condition and triggers actions such as Amazon SNS notifications, making it the alerting mechanism that works together with anomaly detection. Used together, Anomaly Detection identifies the deviation and CloudWatch Alarms fires the alert, which directly satisfies the requirement to detect and alert on unusual metric patterns.

Amazon GuardDuty is a threat-detection service that analyzes VPC Flow Logs, DNS logs, and CloudTrail events for malicious activity, not application metric patterns. AWS CloudTrail records API activity for auditing and governance, and AWS Config evaluates resource configuration compliance, so neither detects anomalies in application metrics or sends metric-based alerts.

Exam trap

DOP-C02 often tests the confusion between security monitoring services (GuardDuty, CloudTrail) and performance monitoring services (CloudWatch), where candidates incorrectly select security tools for anomaly detection in application metrics.

156
Multi-Selecteasy

A DevOps engineer wants to monitor the health of an Auto Scaling group and receive notifications when instances are launched or terminated. Which TWO AWS services can be used together to achieve this?

Select 2 answers
A.AWS CloudTrail.
B.AWS Config.
C.Amazon EventBridge.
D.Amazon SNS.
E.AWS Lambda.
AnswersC, D

Amazon EventBridge is a serverless event bus that can natively consume Auto Scaling group state changes such as EC2 instance launch, terminate, and lifecycle hook notifications. Using event patterns that filter on the aws.autoscaling source and detail types like EC2 Instance Launch Successful, EventBridge matches relevant events in near real time and routes them to targets such as SNS topics for email/SMS alerts. It is the correct service because it provides built-in event ingestion, filtering, and routing without requiring custom code or external monitoring agents, making it ideal for health and lifecycle monitoring.

Why this answer

Amazon EventBridge (option C) is correct because it can capture Auto Scaling group state-change events such as EC2 Instance Launch Successful and EC2 Instance Terminate Successful, and route them to a target for notification. Amazon SNS (option D) is correct because it serves as the notification target, delivering email, SMS, or HTTP messages to subscribers when EventBridge matches those Auto Scaling events. Together, EventBridge detects the launch/termination events and SNS publishes the alerts, which is exactly the monitoring-and-notification pattern requested.

AWS CloudTrail (A) only records API activity for auditing and does not natively push notifications, AWS Config (B) evaluates resource compliance and configuration history rather than real-time launch/terminate alerts, and AWS Lambda (E) is a compute target that could process events but is not itself a notification service.

Exam trap

The trap is selecting CloudTrail or Config for event notifications — candidates confuse auditing/compliance services with real-time event routing and notification, which is EventBridge + SNS.

157
MCQhard

A company runs a web application on EC2 instances behind an Application Load Balancer. They use Amazon CloudFront for content delivery. The DevOps team notices that some requests are returning HTTP 503 errors intermittently. After checking the CloudFront and ALB logs, they find that the errors originate from the ALB. What is the most likely cause?

A.The SSL certificate on the ALB is expired.
B.The security group for the ALB is blocking traffic from CloudFront.
C.CloudFront is configured to forward an HTTP method that the ALB does not support.
D.The ALB is experiencing a surge in traffic and is scaling up, but during the scaling activity, some requests are rejected.
AnswerD

An ALB returns 503 Service Unavailable when it cannot handle incoming requests due to scaling activity or when all targets are unhealthy. During a traffic surge, the ALB nodes scale up by provisioning additional capacity, and during this scaling activity the ALB may temporarily reject or fail to accept new requests, causing clients to receive 503 responses. This matches the scenario described, making it the correct explanation for the issue.

Why this answer

When an Application Load Balancer (ALB) experiences a sudden surge in traffic that exceeds its current capacity, it may temporarily reject requests with HTTP 503 errors while it scales up. During the scaling activity, the ALB's target group might not have enough healthy registered targets to handle the load, causing the ALB to return 503 responses until new instances are provisioned and pass health checks. This matches the intermittent nature of the errors described in the scenario.

Exam trap

The trap here is that candidates often confuse 503 errors with SSL certificate issues or security group misconfigurations, but the intermittent nature of the errors and the fact that they originate from the ALB (not CloudFront) points directly to capacity scaling limitations rather than configuration errors.

How to eliminate wrong answers

Option A is wrong because an expired SSL certificate on the ALB would cause SSL/TLS handshake failures (e.g., ERR_CERT_DATE_INVALID) and result in 502 Bad Gateway errors from CloudFront, not 503 errors from the ALB. Option B is wrong because if the security group for the ALB were blocking traffic from CloudFront, the ALB would not receive the requests at all, and CloudFront would return 502 errors (or connection timeouts) instead of the ALB returning 503 errors. Option C is wrong because CloudFront forwarding an unsupported HTTP method would cause the ALB to return a 405 Method Not Allowed error, not a 503 Service Unavailable error.

158
MCQmedium

A company has deployed a containerized application on Amazon ECS with Fargate. The application is fronted by an Application Load Balancer (ALB). The DevOps team is using CloudWatch Container Insights to monitor the ECS cluster. They notice that the 'MemoryUtilized' metric for the service is consistently above 80%, and the 'CPUUtilized' is around 50%. The ALB's 'TargetResponseTime' is increasing over time. The team wants to resolve the performance issue. Which action should the team take?

A.Increase the memory limit for the ECS task definition to allow the container to use more memory.
B.Increase the CPU limit for the ECS task definition to improve performance.
C.Increase the number of ALB targets by adding more availability zones.
D.Increase the desired count of the ECS service to distribute the load across more tasks.
AnswerA

The ECS task definition's memory limit is a hard limit enforced by Docker; when the container's memory utilization consistently exceeds 80%, it is likely approaching or hitting that ceiling, leading to OOM kills or severe performance degradation. Increasing the memory limit lets the container allocate more heap or working set, directly relieving the memory bottleneck. This is a vertical scaling action, and you must also ensure the EC2 instance has enough free memory to support the increased limit.

Why this answer

The metrics show MemoryUtilized consistently above 80% while CPUUtilized is only around 50%, and TargetResponseTime is rising — this is a classic memory-bound bottleneck. When a container approaches its memory limit, the kernel may reclaim page cache, trigger GC pressure, or begin swapping (if enabled), all of which degrade response time. Increasing the memory limit in the task definition gives the container headroom to operate without memory pressure, directly addressing the root cause.

Exam trap

DOP-C02 often tests whether candidates can diagnose the bottleneck from metrics — the trap is picking CPU scaling or horizontal scaling when the data clearly shows memory pressure as the root cause of rising response time.

How to eliminate wrong answers

Option B is wrong because CPUUtilized is only around 50%, so CPU is not the bottleneck — increasing the CPU limit wastes resources and does not address the memory pressure causing the latency. Option C is wrong because adding ALB targets in more availability zones increases redundancy and capacity at the load-balancer level, but the bottleneck is per-task memory, not insufficient targets or AZ coverage. Option D is wrong because scaling out the ECS service adds more tasks, which can help throughput, but each task still hits the same memory ceiling — without raising the memory limit, the underlying per-task memory pressure and rising response time persist.

159
MCQhard

A company runs a microservices architecture on Amazon ECS with Fargate. The operations team wants to collect custom application metrics (e.g., request latency per service) and visualize them in CloudWatch dashboards. The team also needs to set CloudWatch alarms based on these metrics. Which solution requires the LEAST amount of code changes and operational overhead?

A.Use the CloudWatch Embedded Metric Format to emit custom metrics as JSON log entries.
B.Deploy a StatsD daemon as a sidecar container and configure the application to send metrics to StatsD, then forward to CloudWatch.
C.Modify the application code to use the AWS SDK to call PutMetricData API directly.
D.Install the CloudWatch Agent on each Fargate task as a sidecar container to collect custom metrics.
AnswerA

The CloudWatch Embedded Metric Format encodes custom metric values inside a structured JSON log event; when the Fargate task's awslogs driver sends that log to CloudWatch Logs, CloudWatch automatically extracts the declared metrics into the specified namespace for graphing and alarms. This requires no separate daemon, sidecar, or SDK call—developers only add a serialization layer to application logging, making it the minimal-code path you asked for.

Why this answer

The CloudWatch Embedded Metric Format allows applications to emit metrics as structured JSON logs, which CloudWatch automatically extracts into metrics and logs. This requires minimal code changes (just log format). Option B is wrong because publishing to CloudWatch via PutMetricData requires the AWS SDK and more code changes.

Option C is wrong because CloudWatch Agent on Fargate is not supported (requires EC2). Option D is wrong because using a sidecar container for StatsD adds complexity and overhead.

160
Multi-Selectmedium

A company uses Amazon CloudWatch Logs to store application logs. The DevOps team wants to search across multiple log groups for a specific error pattern. Which TWO options can be used to achieve this? (Choose TWO.)

Select 2 answers
A.Use CloudWatch Logs Insights to run queries across multiple log groups.
B.Export the logs to Amazon S3 and use Amazon Athena to query the logs.
C.Install the CloudWatch Logs agent on an EC2 instance and tail the logs.
D.Create a Lambda function that reads logs from each log group and searches for the pattern.
E.Use Amazon Kinesis Data Analytics to process the log streams.
AnswersA, B

CloudWatch Logs Insights queries multiple log groups in one request using its query syntax, filtering and aggregating events directly within CloudWatch Logs. This satisfies the requirement to search several log groups for an error pattern without exporting data.

Why this answer

Option A is correct because CloudWatch Logs Insights natively supports querying across multiple log groups in a single query — you can select up to 50 log groups in the console or specify multiple log group ARNs/names in the StartQuery API, and use the query syntax (fields, filter, stats, parse) to search for a specific error pattern. Option B is correct because exporting CloudWatch Logs to Amazon S3 (via CreateExportTask or subscription filters) and then querying the exported data with Amazon Athena lets you run SQL across many log groups' data at once, which is a standard approach for cross-log-group searching and analysis. Option C is not correct because installing the CloudWatch Logs agent and tailing logs only reads logs on a single EC2 instance and does not provide cross-log-group search capability.

Option D is not correct because a custom Lambda function reading each log group would be a bespoke, inefficient workaround rather than a supported cross-log-group search feature, and it is not the intended solution. Option E is not correct because Kinesis Data Analytics processes streaming data for real-time analytics, not for searching historical log data across multiple CloudWatch log groups.

Exam trap

The trap here is that candidates may think Lambda or Kinesis are suitable for ad-hoc log searching, but they are designed for real-time processing or custom workflows, not for efficient cross-log-group querying like CloudWatch Logs Insights or Athena.

161
MCQhard

A company runs a containerized web application on Amazon ECS with AWS Fargate. The application is critical and requires high availability. The DevOps team has set up an Amazon CloudWatch alarm that triggers an auto scaling action when the average CPU utilization exceeds 75% for 5 minutes. However, during a recent traffic spike, the application became slow and some requests timed out, even though the CloudWatch alarm did not fire. The team checked the ECS service auto scaling configuration and found that the target tracking scaling policy based on average CPU utilization is set with a target value of 75%. The ECS service is configured with a minimum of 2 tasks and a maximum of 10 tasks. Upon investigation, they noticed that the CPU utilization metric for the service remained below 75% during the spike, but the memory utilization was high (over 90%). The application logs show that the tasks were running out of memory, causing garbage collection pauses and slow responses. Which course of action should the DevOps engineer take to prevent this issue in the future?

A.Add a second target tracking scaling policy based on average memory utilization with a target value of 75%.
B.Decrease the CPU target value to 50% to trigger scaling earlier.
C.Increase the minimum number of tasks from 2 to 5 to provide more capacity upfront.
D.Increase the task memory limit in the task definition to 8 GB.
AnswerA

Memory exhaustion caused the timeouts while CPU stayed below target, so CPU-based target tracking never scaled out. A second target tracking policy on memory utilisation triggers scaling on the actual bottleneck, preventing garbage collection pauses during future spikes.

Why this answer

The issue is that the application is running out of memory, causing performance degradation, but the scaling policy only monitors CPU. To prevent this, you should add a target tracking scaling policy based on memory utilization. This will trigger scaling when memory exceeds the target, providing additional tasks to handle the load and preventing memory exhaustion.

Exam trap

DOP-C02 often tests the misconception that CPU-based scaling is sufficient for all performance issues, when in fact memory bottlenecks require memory-based scaling policies.

How to eliminate wrong answers

Option B is wrong because decreasing the CPU target would scale earlier based on CPU, but CPU was not the bottleneck; memory was. Option C is wrong because increasing the minimum tasks provides more baseline capacity but does not dynamically scale based on memory, and may not be sufficient during spikes. Option D is wrong because increasing task memory might help individual tasks, but it does not address the scaling issue; if the load increases, tasks could still run out of memory, and it may be more costly.

162
Multi-Selectmedium

A company is using Amazon CloudWatch Logs to store application logs. The DevOps team wants to set up real-time monitoring for specific error patterns and trigger remediation actions. Which TWO services can process the log events in real time and invoke an AWS Lambda function for remediation? (Choose two.)

Select 2 answers
A.Stream log events to Amazon Kinesis Data Streams and configure a Lambda function to process the stream.
B.Create a CloudWatch Logs subscription filter that delivers log events to a Lambda function.
C.Create an Amazon EventBridge rule that matches on CloudWatch Logs log group events.
D.Configure the log group to send log events to an Amazon SQS queue, and have the Lambda function poll the queue.
E.Publish log events to an Amazon SNS topic and subscribe the Lambda function.
AnswersA, B

Amazon CloudWatch Logs subscription filters can deliver log events in real time to a Kinesis data stream, which then uses a Lambda event source mapping to process each record. The stream acts as a durably buffered, highly scalable ingestion layer that decouples producers from consumers and preserves event ordering per shard. This is a fully supported path for real-time log processing and allows multiple Lambda functions or other consumers to read the same stream independently.

Why this answer

Option A is correct because CloudWatch Logs subscription filters can stream log events in real time to Amazon Kinesis Data Streams, and a Lambda function can be configured with the Kinesis stream as an event source to process records and invoke remediation logic. Option B is correct because a CloudWatch Logs subscription filter can deliver matching log events directly to AWS Lambda in real time, which is the native pattern for real-time log processing and automated remediation. Option C is not correct because EventBridge rules match on CloudWatch Logs API events such as CreateLogGroup or PutRetentionPolicy, not on the contents of individual log events, so it cannot process error patterns in real time.

Option D is not correct because CloudWatch Logs does not natively send log events to Amazon SQS, and polling SQS is not a real-time push-based processing pattern. Option E is not correct because CloudWatch Logs does not publish log events directly to Amazon SNS, and even if it did, SNS-to-Lambda is not a supported real-time log event processing path for this scenario.

Exam trap

The trap is assuming CloudWatch Logs can natively target SQS or SNS (options D and E); in reality only subscription filters and Kinesis streaming are supported real-time paths to Lambda.

163
MCQhard

A DevOps engineer is configuring a centralized logging solution using Amazon CloudWatch Logs. They need to ensure that logs from multiple AWS accounts are aggregated into a single CloudWatch Logs account. Which approach meets this requirement?

A.Use Amazon Kinesis Data Firehose in each account to stream logs to a central Amazon S3 bucket, then use Amazon Athena to query.
B.Create a subscription filter in each account that delivers log events to a CloudWatch Logs destination in the central account.
C.Set up a cross-account destination using an Amazon Kinesis Data Streams stream in the central account and configure each account to send logs to that stream.
D.Configure each application to use the PutLogEvents API to send logs directly to the central account's log group.
AnswerB

Subscription filters apply a filter pattern and forward matching log events to a destination, which can be a CloudWatch Logs destination in another account. This satisfies the cross-account aggregation requirement by streaming each account's logs into the single central logging account.

Why this answer

CloudWatch Logs supports cross-account subscription filters that can deliver log events to a CloudWatch Logs destination in a central account. The destination is a logical resource that points to a Kinesis Data Stream or Lambda function in the central account, and the source account creates a subscription filter that sends matching log events to that destination. This allows centralized aggregation without requiring each account to manage separate streaming infrastructure.

Exam trap

The trap here is that candidates confuse the CloudWatch Logs destination (which is a cross-account subscription mechanism) with directly writing to a Kinesis stream or using PutLogEvents across accounts, both of which are not supported for cross-account log aggregation.

How to eliminate wrong answers

Option A is wrong because Amazon Kinesis Data Firehose cannot directly stream logs from CloudWatch Logs in multiple accounts to a central S3 bucket without additional cross-account permissions and intermediate services; it also introduces unnecessary complexity and latency for real-time log aggregation. Option C is wrong because while a cross-account Kinesis Data Streams destination can be used, the correct implementation requires creating a CloudWatch Logs destination in the central account that points to the Kinesis stream, not configuring each account to send logs directly to the stream via PutRecord. Option D is wrong because the PutLogEvents API requires the log group and log stream to exist in the same account as the API call; cross-account PutLogEvents is not supported, and applications cannot send logs directly to a central account's log group.

164
Matchingmedium

Match each AWS service health or performance concept to its meaning.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Maximum limits on resources per account

Shows events and changes affecting your AWS resources

Monitors a metric and performs actions based on thresholds

Provides recommendations for cost, performance, security, and fault tolerance

Recommends optimal AWS compute resources for workloads

Why these pairings

The correct matches are: Amazon CloudWatch monitors resources in real-time; AWS Trusted Advisor optimizes cost, security, and performance; AWS Health Dashboard provides personalized health alerts. Common confusions involve swapping monitoring (CloudWatch) with auditing (CloudTrail) and recommendations (Trusted Advisor) with monitoring.

165
MCQeasy

A DevOps engineer needs to set up an alert for when the CPU utilization of an EC2 instance exceeds 90% for 5 consecutive minutes. Which CloudWatch features should be used?

A.CloudWatch Logs with a metric filter on CPU utilization logs.
B.CloudTrail to monitor EC2 instance CPU usage.
C.CloudWatch alarm on the CPUUtilization metric with a period of 5 minutes and threshold of 90.
D.Amazon S3 server access logs to check CPU utilization.
AnswerC

CloudWatch alarms are the native way to react to metric changes: the CPUUtilization metric for EC2 is published every 5 minutes with basic monitoring, so setting a period of 5 minutes ensures the alarm evaluates the average CPU utilization over each interval. A threshold of 90 means the alarm state becomes ALARM when the CPU utilization statistic exceeds 90%, and you can attach an action such as an SNS notification or Auto Scaling policy. This leverages a built-in metric with a well-established alarm evaluation model, making it the correct and simplest approach.

Why this answer

CloudWatch alarms can be configured on the `CPUUtilization` metric (a standard EC2 metric emitted every 5 minutes by default) with a threshold of 90 and an evaluation period of 1 (since the period is set to 5 minutes, one evaluation period covers the 5 consecutive minutes). This directly meets the requirement without additional setup.

Exam trap

The trap here is that candidates may confuse CloudWatch Logs metric filters (used for custom log-based metrics) with native EC2 metrics, or mistakenly think CloudTrail or S3 logs can monitor system performance, when only CloudWatch alarms on the `CPUUtilization` metric directly satisfy the requirement.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs with a metric filter requires EC2 instances to send CPU utilization data to CloudWatch Logs via a custom agent or script, which is unnecessary since the `CPUUtilization` metric is already available natively. Option B is wrong because CloudTrail records API calls and management events, not system-level metrics like CPU usage; it cannot monitor CPU utilization. Option D is wrong because Amazon S3 server access logs track requests made to S3 buckets, not EC2 instance performance metrics.

166
MCQeasy

A DevOps engineer is tasked with ensuring that all Amazon S3 buckets in the account have server access logging enabled. The engineer needs to be automatically notified when a new bucket is created without logging enabled. Which AWS service should they use?

A.Use AWS CloudTrail to detect CreateBucket API calls and trigger a Lambda function to check logging.
B.Use AWS Trusted Advisor to check S3 bucket logging and send notifications via Amazon SNS.
C.Use Amazon S3 Event Notifications to trigger a Lambda function when a new bucket is created.
D.Use AWS Config with a managed rule to check if S3 bucket logging is enabled, and configure an SNS topic for notifications.
AnswerD

AWS Config continuously records configuration changes and evaluates them with managed rules, including s3-bucket-logging-enabled, which verifies server access logging is turned on for each bucket. When a bucket is created or its logging configuration changes, AWS Config re-evaluates in near real time and can publish compliance results to an SNS topic, triggering notifications. This provides a fully managed, automated, and near-real-time compliance check, making it the appropriate solution.

Why this answer

AWS Config provides continuous monitoring and evaluation of your AWS resource configurations. By using the managed rule 's3-bucket-server-access-logging-enabled', AWS Config can automatically check all S3 buckets (including newly created ones) for server access logging. When a bucket is non-compliant, AWS Config can trigger an SNS notification to alert the DevOps engineer, meeting the requirement for automatic notification without custom code.

Exam trap

The trap here is that candidates often confuse event-driven services like CloudTrail or S3 Event Notifications with configuration compliance services, mistakenly thinking they can directly detect and react to resource misconfigurations without the need for custom evaluation logic.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail records API calls but does not evaluate resource configurations; triggering a Lambda function from CloudTrail would require custom code to parse the event and check logging, which is not the most efficient or managed solution. Option B is wrong because AWS Trusted Advisor checks S3 bucket logging only for buckets in the 'S3 Bucket Logging' check, but it does not provide real-time notifications for new bucket creation; it runs periodic checks and requires manual setup or custom automation for alerts. Option C is wrong because Amazon S3 Event Notifications cannot be configured on a bucket that does not exist yet; you cannot set up event notifications for 'new bucket creation' events, as S3 Event Notifications are per-bucket and only support events like object creation or deletion within an existing bucket.

167
MCQhard

A company is using Amazon CloudWatch Logs to store application logs. The DevOps engineer needs to ensure that log data is encrypted at rest using a customer-managed KMS key. What step must be taken?

A.Use AWS CloudTrail to encrypt the log data before it is sent to CloudWatch Logs.
B.Create a KMS key and apply it to the IAM role used by the application.
C.Create a new KMS customer-managed key and associate it with the CloudWatch Logs log group.
D.Enable server-side encryption on the log group using the default CloudWatch Logs key.
AnswerC

This is the correct solution because CloudWatch Logs supports server-side encryption using a customer-managed KMS key that you can associate directly with the log group. When you create a log group or call the AssociateKmsKey API, you supply the key ARN, and CloudWatch Logs uses that key to encrypt all new incoming log events. This gives you full control over key rotation, access auditing, and lifecycle management, which is exactly what a requirement for customer-managed encryption requires. Note that the key must exist in the same AWS Region as the log group and its key policy must grant CloudWatch Logs permission to generate a data key for encryption.

Why this answer

CloudWatch Logs supports encryption at rest using a customer-managed KMS key, which must be explicitly associated with the log group. When you create or update a log group, you can specify a KMS key ID (via the AWS CLI, SDK, or console) to encrypt all log data stored in that group. This ensures that the log data is encrypted using a key you control, not the default AWS-managed key.

Exam trap

The trap here is that candidates often confuse associating a KMS key with an IAM role (which controls access) with associating it directly with the log group (which controls encryption at rest), leading them to select Option B instead of C.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail is an auditing service that records API calls, not an encryption mechanism; it cannot encrypt log data before it is sent to CloudWatch Logs. Option B is wrong because applying a KMS key to an IAM role does not encrypt log data at rest; the key must be associated directly with the CloudWatch Logs log group, not with an IAM role. Option D is wrong because enabling server-side encryption with the default CloudWatch Logs key uses an AWS-managed key, not a customer-managed KMS key, which does not meet the requirement for a customer-managed key.

168
MCQhard

A DevOps engineer is tasked with centralizing logs from multiple AWS accounts into a single Amazon OpenSearch Service domain. The engineer sets up Amazon Kinesis Data Firehose to deliver logs from each account to the OpenSearch domain. However, some accounts show failed deliveries in the Firehose console. Which configuration is MOST likely causing the failures?

A.The IAM role assumed by Firehose in each account does not have permissions to write to the cross-account OpenSearch domain
B.The source accounts do not have a CloudWatch Logs subscription filter to send logs to Firehose
C.The Kinesis Data Streams used as the Firehose source is not encrypted
D.The OpenSearch domain's access policy does not allow access from the S3 bucket used by Firehose
AnswerA

Firehose assumes an IAM role in each source account, and that role's permissions govern delivery. If it lacks an OpenSearch write policy for the domain in the central account, delivery fails only in those accounts, matching the intermittent failures. Cross-account access requires the role plus a domain access policy permitting it.

Why this answer

The most likely cause of failed deliveries is that the IAM role assumed by Kinesis Data Firehose in each source account lacks the necessary permissions to write to the cross-account Amazon OpenSearch Service domain. Firehose uses a service-linked or custom IAM role to perform actions such as `es:ESHttpPut` and `es:ESHttpPost` against the OpenSearch domain endpoint. Without explicit cross-account trust and resource-based policy allowing the Firehose role's ARN, the delivery will fail with an authorization error.

Exam trap

The trap here is that candidates often assume the failure is due to missing CloudWatch subscription filters or S3 bucket permissions, but the real issue is the missing cross-account IAM trust between the Firehose role and the OpenSearch domain's access policy.

How to eliminate wrong answers

Option B is wrong because CloudWatch Logs subscription filters are used to stream log data to Firehose, but the question states that logs are being delivered from multiple accounts; the failure is at the Firehose-to-OpenSearch stage, not at the ingestion stage. Option C is wrong because Kinesis Data Streams encryption (whether server-side or client-side) does not affect Firehose's ability to write to OpenSearch; Firehose can read encrypted streams as long as it has the proper KMS permissions. Option D is wrong because Firehose writes directly to the OpenSearch domain via HTTP/HTTPS, not through an S3 bucket; the OpenSearch domain's access policy must grant access to the Firehose IAM role or the source account's principal, not to an S3 bucket.

169
Multi-Selecthard

A company uses AWS Organizations to manage multiple accounts. The DevOps team needs to monitor for any IAM user creation across all accounts in the organization. Which THREE steps should be taken to implement this centralized monitoring?

Select 3 answers
A.Create a CloudWatch Logs metric filter on the organization's CloudTrail log group for 'CreateUser' events.
B.Enable CloudTrail in the management account with an organization trail that applies to all accounts.
C.Configure an S3 bucket to receive CloudTrail logs from all accounts and enable S3 event notifications for object creation.
D.Use AWS Config rules to detect IAM user creation across accounts.
E.Set a CloudWatch alarm on the metric to send notifications via SNS.
AnswersA, B, E

Metric filters in CloudWatch Logs can parse CloudTrail logs for specific event names, counting occurrences of CreateUser API calls. Since the organization trail delivers logs to a central log group in the management account, a single metric filter can monitor IAM user creation across all member accounts. This provides a real-time, event-driven signal to trigger a CloudWatch alarm, rather than relying on periodic scans or resource configuration evaluations.

Why this answer

Option B is correct because an organization trail created in the management account automatically applies to all member accounts, delivering CloudTrail events such as IAM CreateUser to a central S3 bucket and CloudWatch Logs, which is the foundation for centralized monitoring. Option A is correct because a CloudWatch Logs metric filter on the organization's CloudTrail log group can match the 'CreateUser' event name and turn each occurrence into a custom metric, enabling detection of IAM user creation across all accounts. Option E is correct because a CloudWatch alarm on that metric can trigger notifications through Amazon SNS, providing the DevOps team with real-time alerts when IAM users are created.

Option C is not needed because S3 event notifications on object creation only signal that a log file was delivered, not that a CreateUser event occurred, and CloudWatch Logs metric filters are the appropriate mechanism. Option D is not appropriate because AWS Config rules evaluate resource configuration compliance and are not the primary real-time event-detection mechanism for API calls like CreateUser; CloudTrail plus CloudWatch is the intended solution.

Exam trap

DOP-C02 often tests whether candidates confuse S3-based log archival with CloudWatch-based real-time alerting — the S3 event notification option sounds plausible but fires on log file delivery, not on the specific API event.

170
MCQeasy

A company runs a web application on an Auto Scaling group of EC2 instances. The operations team uses CloudWatch alarms to monitor the application. They have set up a CPUUtilization alarm that triggers when the average CPU exceeds 70% for 5 minutes. The alarm triggers a scaling policy to add instances. Recently, the team noticed that the alarm frequently triggers during the day, but the application performance is acceptable. They suspect the alarm is too sensitive and want to reduce the number of false alarms. The team wants to keep the alarm responsive to real CPU spikes but avoid triggering on short bursts. What should the team change in the alarm configuration?

A.Create a composite alarm that combines CPUUtilization with MemoryUtilization.
B.Reduce the metric period to 1 minute and keep evaluation periods at 1.
C.Increase the number of evaluation periods to 3, so the alarm triggers only if CPU is high for 3 consecutive periods.
D.Lower the threshold to 60% to catch more CPU spikes.
AnswerC

Increasing evaluation periods to 3 (with the same period of e.g., 5 minutes) means the alarm triggers only if CPU exceeds 70% for 3 consecutive periods (15 minutes). This filters out short bursts and reduces false alarms while still catching sustained high CPU.

Why this answer

To reduce false alarms while remaining responsive to real CPU spikes, increase the number of evaluation periods. With 3 consecutive periods, the alarm triggers only if CPU exceeds 70% for 15 minutes (assuming 5-minute periods), filtering out short bursts. This maintains responsiveness to sustained spikes while avoiding triggers on transient spikes.

Exam trap

The trap is confusing period with evaluation periods; candidates might think reducing the period makes the alarm less sensitive, but it actually makes it more granular and can increase false positives if evaluation periods are low.

How to eliminate wrong answers

Option A is wrong because a composite alarm combining CPU and memory would require both metrics to be high, which could miss CPU-only spikes and is not the intended fix for sensitivity. Option B is wrong because reducing the period to 1 minute and keeping evaluation periods at 1 would make the alarm more sensitive, triggering on even shorter bursts. Option D is wrong because lowering the threshold to 60% would make the alarm trigger more often, increasing false alarms.

171
Multi-Selectmedium

A company is deploying a new microservice on AWS Lambda. The DevOps team needs to monitor the function for errors and performance issues. Which TWO steps should the team take to set up effective monitoring?

Select 2 answers
A.Enable VPC Flow Logs to monitor network traffic to the function
B.Enable AWS Config rules to evaluate the function configuration
C.Enable active tracing with AWS X-Ray to trace requests through the function
D.Enable CloudWatch Logs for the Lambda function to capture application logs
E.Install the CloudWatch Agent on the Lambda execution environment
AnswersC, D

Enabling active tracing with AWS X-Ray gives you end-to-end visibility into requests as they pass through the Lambda function and any downstream AWS services or HTTP APIs. X-Ray automatically records a trace segment for each invocation, captures timing, errors, and subsegments for calls made with the AWS SDK, and supports sampling to control cost. It also propagates trace IDs across services, enabling you to follow a single user request through the entire distributed application.

Why this answer

Option C is correct because enabling active tracing with AWS X-Ray on a Lambda function instruments the function and its downstream calls, letting the team trace requests end-to-end, identify latency bottlenecks, and pinpoint errors across the microservice. Option D is correct because Lambda automatically streams function output to CloudWatch Logs, which captures application logs, stack traces, and custom metrics needed to diagnose errors and performance issues. VPC Flow Logs (A) only record IP-level network traffic metadata for VPC resources and do not provide function-level error or performance insight.

AWS Config rules (B) evaluate resource configuration compliance, not runtime errors or performance. The CloudWatch Agent (E) is installed on EC2 instances or on-premises servers and cannot be installed inside the managed Lambda execution environment.

Exam trap

DOP-C02 often tests the misconception that agents (like CloudWatch Agent) can be installed on Lambda — candidates must remember Lambda is a managed runtime where you rely on native CloudWatch Logs and X-Ray integration.

172
Multi-Selecthard

A DevOps team is troubleshooting a slow website that uses Amazon CloudFront with an Application Load Balancer as the origin. The team notices that cache hit ratio is low. Which THREE actions are most likely to improve the cache hit ratio?

Select 3 answers
A.Configure CloudFront to forward all cookies to the origin.
B.Enable CloudFront Origin Shield to reduce load on the origin and increase cache effectiveness.
C.Decrease the default TTL for objects.
D.Increase the minimum TTL for the CloudFront distribution.
E.Optimize the cache key to include only relevant headers.
AnswersB, D, E

Origin Shield is an optional intermediate cache layer that sits between CloudFront edge locations and the origin, aggregating requests from all edges. When multiple edges miss their local cache for the same object, Origin Shield consolidates those requests into a single origin fetch and caches the result globally, raising the effective hit ratio for the entire distribution and reducing origin traffic. It also adds resilience by absorbing request spikes and lowering latency for revalidation.

Why this answer

CloudFront Origin Shield acts as an additional caching layer that consolidates requests from multiple edge locations, reducing the load on the origin and increasing the likelihood of cache hits by serving cached content from the Origin Shield regional cache. This improves cache effectiveness, especially for origins with high latency or limited capacity.

Exam trap

The trap here is that candidates often confuse decreasing TTL with improving cache hit ratio, but in reality, shorter TTLs cause more frequent cache expirations and origin fetches, reducing cache effectiveness.

173
MCQmedium

A DevOps engineer notices that an Amazon RDS for MySQL instance's CPU is consistently high during business hours. The engineer wants to identify the specific queries causing the high CPU. Which combination of services should be used to capture and analyze the queries? (Choose the best answer.)

A.Enable RDS Performance Insights and analyze the top SQL queries
B.Enable RDS Enhanced Monitoring and view metrics in CloudWatch
C.Enable AWS X-Ray tracing on the application and database
D.Enable RDS audit logs and stream them to Amazon CloudWatch Logs
AnswerA

Performance Insights samples the database load and attributes it to individual SQL statements, exposing the top queries by database load during the high-CPU window. This directly identifies the offending statements on the MySQL instance without requiring slow query log parsing.

Why this answer

RDS Performance Insights provides a database performance tuning feature that visualizes database load and identifies the specific SQL queries causing high CPU. It captures query-level metrics such as wait events, SQL digest, and host/user information, allowing the DevOps engineer to pinpoint the exact queries responsible for the CPU spike during business hours.

Exam trap

The trap here is that candidates often confuse Enhanced Monitoring (OS-level metrics) with Performance Insights (query-level analysis), or assume audit logs or X-Ray can provide SQL-level performance data, when in fact they serve different purposes (compliance and tracing, respectively).

How to eliminate wrong answers

Option B is wrong because Enhanced Monitoring provides OS-level metrics (e.g., CPU, memory, disk I/O) but does not capture or identify individual SQL queries; it cannot show which specific queries are causing high CPU. Option C is wrong because AWS X-Ray traces application requests and can trace calls to the database, but it does not capture the actual SQL queries executed on the RDS instance; it is designed for distributed tracing, not query-level analysis. Option D is wrong because RDS audit logs record database activities (e.g., logins, schema changes) for compliance, not query performance metrics; streaming them to CloudWatch Logs does not provide the query-level CPU impact analysis needed to identify high-CPU queries.

174
MCQeasy

A company runs a production web application on Amazon EC2 instances that are part of an Auto Scaling group. The instances are behind an Application Load Balancer. The DevOps team has enabled detailed CloudWatch metrics and set up a CloudWatch dashboard to monitor the application. Recently, the team noticed that the CPU Utilization metric for the Auto Scaling group shows a spike every day at 2:00 PM, but the application performance remains normal. The team wants to investigate the cause of the CPU spike. What should the team do FIRST to identify the root cause?

A.Enable AWS CloudTrail to log all API calls to the instances.
B.Use CloudWatch Logs Insights to query the application logs on the instances to identify any scheduled tasks or jobs running at 2:00 PM.
C.Disable any scheduled tasks on the instances to see if the spike stops.
D.Increase the instance size to provide more CPU capacity to handle the spike.
AnswerB

CloudWatch Logs Insights enables you to run SQL-like queries across log groups that receive application and system logs from EC2 instances via the unified CloudWatch agent. By filtering for messages between 1:55 PM and 2:05 PM and searching for terms such as 'cron,' 'schedule,' or the job name, you can identify a recurring batch process that coincides with the spike. This is a non-invasive, first-step diagnostic that directly associates application behavior with the CPU metric.

Why this answer

The FIRST step in root-cause analysis is to gather evidence, not to change the environment. CloudWatch Logs Insights allows the team to query application and system logs on the EC2 instances to identify scheduled tasks (e.g., cron jobs, batch scripts, log rotation) that run at 2:00 PM and cause the CPU spike. This is non-disruptive and directly targets the symptom's timing, providing data to confirm or rule out hypotheses before making changes.

Exam trap

DOP-C02 often tests the instinct to 'fix' the problem immediately (disable tasks, resize instances) rather than first gathering diagnostic data, leading candidates to skip the investigative step.

How to eliminate wrong answers

Option A is wrong because CloudTrail logs AWS API calls, not in-instance process activity, so it would not reveal a cron job or application task causing CPU spikes. Option C is wrong because disabling scheduled tasks before identifying them is a premature change that could break business processes and does not confirm the root cause. Option D is wrong because increasing instance size masks the symptom rather than diagnosing the cause, and it incurs unnecessary cost.

175
MCQeasy

A company is using Amazon RDS for MySQL and wants to monitor database connections. They need to set up an alarm when the number of connections exceeds 80% of the maximum connections for more than 5 minutes. Which CloudWatch metric and statistic should be used?

A.DatabaseConnections metric with Maximum statistic
B.DatabaseConnections metric with Average statistic
C.DatabaseConnections metric with Sum and then divide by the number of data points
D.DatabaseConnections metric with Sum statistic
AnswerB

The Average statistic computes the mean DatabaseConnections over the 5-minute interval, which inherently dampens short-lived fluctuations and reveals the central tendency of connection concurrency. If the average exceeds the 80% threshold, it means the typical number of connections during the entire window was too high, matching the criterion of sustained usage for more than 5 minutes. This is the most appropriate aggregation for a threshold alarm aimed at detecting prolonged saturation of the connection pool.

Why this answer

The Average statistic of the DatabaseConnections metric over a 5-minute period provides a smoothed representation of connection usage, which is appropriate for detecting sustained breaches of the 80% threshold. Using Average reduces sensitivity to transient spikes, ensuring the alarm triggers only when the average number of connections remains above the threshold for the entire evaluation period, aligning with the requirement of 'more than 5 minutes'.

Exam trap

The trap here is that candidates often choose Maximum because they think it is the most conservative for detecting high usage, but they overlook that the requirement is for sustained breaches over 5 minutes, not instantaneous spikes, making Average the correct choice for avoiding false alarms.

How to eliminate wrong answers

Option A is wrong because the Maximum statistic captures the highest single data point within the period, which would trigger alarms on brief spikes even if the average stays below 80%, causing false positives. Option C is wrong because dividing the Sum by the number of data points is mathematically equivalent to the Average statistic, but this approach is unnecessarily complex and not a standard CloudWatch metric statistic; CloudWatch directly supports Average. Option D is wrong because the Sum statistic aggregates the total number of connections over the period, which is not meaningful for comparing against a percentage of maximum connections—Sum values scale with the number of data points and do not represent a per-moment connection count.

176
MCQmedium

A company uses AWS Lambda functions to process incoming events. The DevOps team notices that some functions are timing out after 30 seconds, but the configured timeout is 1 minute. They want to capture the actual invocation duration for all invocations to analyze performance. What is the most efficient way to achieve this?

A.Add custom metrics using the AWS SDK within the Lambda function code to record the duration.
B.Configure Amazon Kinesis Data Streams to receive Lambda invocation records and compute duration using a consumer application.
C.Enable detailed CloudWatch Logs for the Lambda functions and parse the 'REPORT' log entries to extract the 'Duration' value.
D.Use AWS CloudTrail to capture Lambda execution events and analyze the 'duration' field.
AnswerC

Enabling CloudWatch Logs is the correct approach because every Lambda invocation automatically emits a REPORT log entry containing a Duration field, for example 123.45 ms, along with billed duration and memory usage. This is generated by the Lambda managed runtime, so there is no need to instrument the application code. You can use CloudWatch Logs Insights with a filter such as filter @type = 'REPORT' to query and parse these entries across all invocations, making it an efficient and fully managed solution for extracting execution duration.

Why this answer

Lambda automatically writes a REPORT log entry to CloudWatch Logs at the end of each invocation, which includes the exact 'Duration' in milliseconds. Parsing these logs is the most efficient approach since it requires no code changes, no additional infrastructure, and leverages existing logging with no extra cost beyond standard CloudWatch Logs ingestion.

Exam trap

The trap here is that candidates may confuse CloudTrail's 'duration' field (which measures API call latency) with the actual function execution duration, leading them to incorrectly select option D.

How to eliminate wrong answers

Option A is wrong because adding custom metrics via the AWS SDK within the function code requires modifying every function, introduces latency from SDK calls, and incurs additional CloudWatch custom metrics costs, making it less efficient than using built-in logs. Option B is wrong because configuring Kinesis Data Streams to receive invocation records is overly complex and costly; Lambda does not natively send invocation records to Kinesis, and building a consumer application to compute duration from streamed data is far less efficient than parsing existing logs. Option D is wrong because CloudTrail captures API calls (e.g., Invoke actions) but does not record the actual function execution duration; the 'duration' field in CloudTrail events refers to the API call latency, not the function's runtime.

177
MCQmedium

A company uses AWS Lambda functions to process streaming data from Amazon Kinesis Data Streams. The Lambda function processes records in batches and writes the results to an Amazon DynamoDB table. Recently, the operations team noticed that the Lambda function is experiencing a high number of throttling errors (HTTP 400) when writing to DynamoDB. The DynamoDB table has on-demand capacity mode enabled. The CloudWatch metrics show that the DynamoDB consumed write capacity is well below the provisioned limits, but the Lambda function's error rate is increasing. The Lambda function's reserved concurrency is set to 100, and the function's timeout is 1 minute. The Kinesis stream has 10 shards. What is the MOST likely cause of the throttling errors?

A.The DynamoDB table is experiencing hot partitions due to uneven access patterns.
B.The Lambda function's timeout is too short, causing the function to retry and overload DynamoDB.
C.The Lambda function's reserved concurrency is too high, causing too many concurrent invocations.
D.The Kinesis stream's batch size is too large, causing the Lambda function to write too many records at once.
AnswerA

DynamoDB on-demand capacity protects against table-level throttling but still enforces a per-partition limit of about 1,000 write capacity units. When access patterns are uneven, such as a single partition key receiving a disproportionate share of writes from the Kinesis stream, that partition hits its limit and throttles requests even though the overall table has plenty of capacity. Adaptive capacity can redistribute some of this load, but it is not instantaneous and does not fully prevent throttling during sharp spikes.

Why this answer

The most likely cause is hot partitions in the DynamoDB table due to uneven access patterns. Even with on-demand capacity, DynamoDB partitions have throughput limits, and if the Lambda function writes to a few partitions heavily, throttling occurs. The consumed write capacity being below provisioned limits indicates that the issue is not overall capacity but partition-level contention.

Exam trap

The trap is assuming that on-demand capacity eliminates throttling, but hot partitions can still cause throttling even when overall capacity is sufficient.

How to eliminate wrong answers

Option B is wrong because a short timeout would cause Lambda retries, but the error is throttling from DynamoDB, not timeout errors. Option C is wrong because reserved concurrency of 100 is not too high; it limits concurrency, and the issue is DynamoDB throttling, not Lambda concurrency. Option D is wrong because batch size affects the number of records processed per invocation, but the throttling is due to DynamoDB partition limits, not batch size.

178
MCQeasy

A DevOps engineer is tasked with setting up monitoring for a serverless application that uses AWS Lambda, Amazon API Gateway, and Amazon DynamoDB. The engineer needs to create a centralized dashboard that displays the number of Lambda invocations, API Gateway request counts, and DynamoDB consumed read/write capacity units. The dashboard should be accessible to the operations team without requiring AWS Management Console login. The engineer also wants to set up email alerts when the DynamoDB consumed capacity exceeds 80% of the provisioned capacity. Which solution meets these requirements with the LEAST operational overhead?

A.Use Amazon QuickSight to connect to CloudWatch metrics and create a dashboard with email alerts.
B.Use CloudWatch Logs Insights to query the logs of each service and create a dashboard from the results.
C.Create a CloudWatch dashboard and share it using Amazon Cognito to grant access to the operations team.
D.Create a CloudWatch dashboard with the relevant metrics and set CloudWatch alarms on DynamoDB consumed capacity. Share the dashboard as a public read-only dashboard.
AnswerD

This is the correct solution because CloudWatch natively supports creating dashboards that display multiple operational metrics, and alarms on DynamoDB consumed capacity can be configured to trigger SNS notifications (e.g., email) when thresholds are exceeded. The dashboard can be shared as a public read-only dashboard using the CloudWatch console's 'Share' feature, which generates a URL that grants view-only access without requiring IAM credentials or Cognito. This directly addresses both the need for at-a-glance monitoring and threshold-based alerting on DynamoDB capacity, making it the most operationally sound and low-overhead option.

Why this answer

A CloudWatch dashboard can display metrics from Lambda, API Gateway, and DynamoDB in one place, and CloudWatch alarms on DynamoDB consumed capacity can trigger email notifications via Amazon SNS. Sharing the dashboard as a public read-only dashboard allows the operations team to view it without AWS Management Console login. This approach uses native AWS monitoring services with minimal setup and no custom code.

Exam trap

The trap is overcomplicating the solution with third-party or custom tools; candidates may choose QuickSight or Cognito when native CloudWatch dashboards and alarms already meet the requirements with least overhead.

How to eliminate wrong answers

Option A is wrong because QuickSight is a business intelligence service for data visualization, not a native CloudWatch metric dashboard, and it does not natively send CloudWatch alarm emails; it adds unnecessary complexity. Option B is wrong because CloudWatch Logs Insights queries log data, not metrics, and cannot create a centralized metric dashboard or capacity alarms. Option C is wrong because Cognito is for user authentication in applications, not for sharing CloudWatch dashboards; it would require building an authentication layer, increasing operational overhead.

179
Multi-Selecteasy

A company is using AWS CloudTrail to log API activity in their AWS account. They want to ensure that any modification to CloudTrail configuration itself is logged and that the logs are immutable. Which combination of actions should they take? (Choose TWO.)

Select 2 answers
A.Enable S3 Object Lock on the destination S3 bucket in governance mode.
B.Enable log file validation to guarantee integrity of log files.
C.Disable log file validation to reduce overhead.
D.Store CloudTrail logs in a CloudWatch Logs log group with a retention policy.
E.Enable CloudTrail Insights to detect configuration changes.
AnswersA, B

Enabling S3 Object Lock on the destination S3 bucket is the correct way to make CloudTrail logs immutable. In governance mode, you can set a retention period and object lock protects objects from being deleted or overwritten by any user—including the AWS account root user—unless they have the `s3:BypassGovernanceRetention` permission. This ensures that the audit log remains intact for the duration of the retention period, satisfying compliance mandates that require unauditable log preservation.

Why this answer

Enabling S3 Object Lock in governance mode on the destination S3 bucket prevents any user, including the root user, from overwriting or deleting CloudTrail log objects during the retention period, ensuring immutability. Option B is correct because enabling log file validation creates a digest file that uses SHA-256 hashing to verify that log files have not been modified, deleted, or tampered with after delivery, providing integrity assurance.

Exam trap

The trap here is that candidates often confuse CloudTrail Insights (which detects configuration changes) with the actual mechanisms for ensuring log immutability and integrity, leading them to select option E instead of the correct combination of S3 Object Lock and log file validation.

180
MCQmedium

A company uses Amazon CloudWatch Logs to store application logs from multiple EC2 instances. The security team requires that logs be encrypted at rest using a customer-managed KMS key. Which configuration step should the engineer perform to meet this requirement?

A.Configure the CloudWatch agent to encrypt logs before sending
B.Create a KMS key and attach a policy allowing CloudWatch Logs to use it
C.Enable default encryption on the S3 bucket that stores logs
D.Associate the KMS key with the CloudWatch log group using the console or CLI
AnswerD

Associating the KMS key with the CloudWatch log group is the required and sufficient step to encrypt log data at rest. When you use the console or the associate-kms-key AWS CLI command, CloudWatch Logs begins encrypting all log events in that group with the specified key and uses the same key to decrypt them when you view or export them. This association applies to the log group's data, including all log streams inside it, and is the only way to apply customer-managed KMS encryption to CloudWatch Logs at rest.

Why this answer

CloudWatch Logs supports encryption at rest with a customer-managed KMS key by associating that key with the log group. This is done via the console or CLI (e.g., aws logs associate-kms-key), and once set, all log data in that group is encrypted with the specified CMK. This directly satisfies the security team's requirement for customer-managed key encryption.

Exam trap

DOP-C02 often tests the misconception that creating a KMS key and granting permissions is enough — the actual required action is associating the key with the specific CloudWatch log group, which is the step candidates most often overlook.

How to eliminate wrong answers

Option A is wrong because the CloudWatch agent does not perform KMS encryption of log data before sending; encryption at rest is handled by the CloudWatch Logs service, not the agent. Option B is wrong because creating a KMS key and granting CloudWatch Logs permission is necessary but insufficient — the key must actually be associated with the log group to take effect. Option C is wrong because CloudWatch Logs does not store log data in S3; enabling S3 default encryption is irrelevant to log group encryption.

181
Multi-Selecthard

A company uses Amazon RDS for MySQL and wants to monitor slow queries to optimize performance. Which actions should the DevOps engineer take to capture and analyze slow query logs? (Choose THREE.)

Select 3 answers
A.Use AWS CloudTrail to capture SQL queries
B.Enable the slow query log parameter in the RDS DB parameter group
C.Enable RDS Performance Insights
D.Configure RDS to publish logs to Amazon CloudWatch Logs
E.Use CloudWatch Logs Insights to query and analyze the slow query logs
AnswersB, D, E

In the RDS DB parameter group for MySQL, set slow_query_log to 1 or ON and define long_query_time with the desired threshold (for example, 2 seconds) — these parameters control which queries are written to the slow query log. This log records the exact SQL text, query execution time, and timestamp for every query exceeding the threshold, giving you the raw data needed to calculate SLO metrics like the ratio of slow queries to total queries. Because slow_query_log is a dynamic parameter for MySQL in RDS, you can apply the change immediately without rebooting the database instance.

Why this answer

To capture and analyze slow query logs in Amazon RDS for MySQL, the DevOps engineer should enable the slow query log parameter in the DB parameter group (B), configure RDS to publish logs to Amazon CloudWatch Logs (D), and use CloudWatch Logs Insights to query and analyze the logs (E). Option A (CloudTrail) captures API activity, not SQL queries. Option C (Performance Insights) monitors database performance metrics but does not capture slow query logs.

182
MCQeasy

A DevOps engineer wants to receive an alert when the total number of error logs in an application exceeds 100 within a 5-minute period. The application writes logs to CloudWatch Logs. How can this be achieved?

A.Create a CloudWatch dashboard with a line chart for error count and manually monitor it.
B.Use CloudWatch Logs Insights to run a query every 5 minutes and trigger an alert based on the result.
C.Create a CloudWatch Logs subscription filter to send matching logs to a Lambda function, which counts errors and sends an alert.
D.Create a metric filter on the log group for 'ERROR', then create a CloudWatch alarm on the resulting metric with a threshold of 100.
AnswerD

A metric filter on the CloudWatch Logs group can count occurrences of the pattern 'ERROR', creating a custom metric. A CloudWatch alarm on that metric with a threshold of 100 and a 5-minute period will trigger when the error count exceeds 100.

Why this answer

A metric filter on the CloudWatch Logs group can count occurrences of the pattern 'ERROR', creating a custom metric. A CloudWatch alarm on that metric with a threshold of 100 and a 5-minute period will trigger when the error count exceeds 100. Option A is incorrect because dashboards are for visualization, not automated alerts.

Option B is incorrect because CloudWatch Logs Insights is for ad-hoc queries, not continuous real-time alerting. Option C is incorrect because subscription filters send logs to destinations like Lambda, but the Lambda would need custom logic to count and alert, which is not as direct as a metric filter. The metric filter with alarm is the simplest native solution.

183
MCQhard

A DevOps team uses AWS Lambda functions to process events from an SQS queue. The Lambda function occasionally fails due to transient errors, and the team wants to capture and analyze the full error details, including stack traces, for debugging. The errors are not always related to invocation failures (e.g., timeouts) but include exceptions thrown within the function code. Which approach will capture the MOST comprehensive error information?

A.Configure a DLQ on the SQS queue to capture failed messages and inspect them.
B.Enable CloudWatch Logs and rely on the automatic logging of invocation results.
C.Ensure the Lambda function code returns a meaningful error object (e.g., throws an exception) so that the error is logged in CloudWatch Logs with a stack trace.
D.Use AWS X-Ray to trace the function execution and analyze the traces.
AnswerC

By returning a meaningful error object (e.g., throwing an exception) within the Lambda handler, the error details and stack trace are automatically written to CloudWatch Logs. This gives the most comprehensive information for debugging application errors.

Why this answer

When a Lambda function throws an exception or returns an error object, AWS Lambda automatically logs the error details, including the stack trace, to CloudWatch Logs. This captures the full error information necessary for debugging transient errors. Option A is incorrect because a Dead Letter Queue (DLQ) on SQS captures the failed messages themselves, not the error details or stack traces of the function execution.

Option B is incorrect because CloudWatch Logs automatic invocation logging provides only basic information such as invocation time, duration, and status; it does not include the function's stack trace unless explicitly logged by the code. Option D is incorrect because AWS X-Ray provides tracing of requests and can show service maps and latency, but it does not necessarily capture the full stack trace of application-level exceptions; it focuses on request flow rather than detailed error logs.

184
Multi-Selectmedium

A company is using Amazon CloudWatch Logs to store application logs. The DevOps team needs to search and analyze logs from multiple EC2 instances in real time. Which TWO services can be used to achieve this? (Choose TWO.)

Select 2 answers
A.Amazon OpenSearch Service.
B.Amazon Athena.
C.Amazon QuickSight.
D.Amazon Kinesis Data Analytics.
E.CloudWatch Logs Insights.
AnswersA, E

Amazon OpenSearch Service ingests CloudWatch Logs via a subscription filter and Lambda, indexing them for real-time full-text search, aggregations, and Kibana visualization. This makes it purpose-built for interactive log analytics and operational dashboards, directly querying the live stream without S3 export latency. It also scales to handle massive log volumes with open-source Elasticsearch-compatible APIs.

Why this answer

CloudWatch Logs can stream logs to Amazon OpenSearch Service for real-time search and analytics. Option E is correct because CloudWatch Logs Insights allows real-time querying of log groups directly within CloudWatch. Option B is incorrect: Amazon Athena is designed for querying data in S3, not for real-time log search from EC2 instances.

Option C is incorrect: Amazon QuickSight is a business intelligence service for visualization, not real-time log search. Option D is incorrect: Amazon Kinesis Data Analytics is for analyzing streaming data, not directly searching CloudWatch Logs.

185
MCQhard

A company has a multi-account AWS environment using AWS Organizations. The security team needs to centrally monitor and analyze VPC Flow Logs from all accounts. The solution must be cost-effective and allow querying across accounts. Which approach should they take?

A.Use Amazon Elasticsearch Service (Amazon OpenSearch Service) with a cross-account ingestion pipeline.
B.Stream VPC Flow Logs from each account to Amazon Kinesis Data Analytics for real-time analysis.
C.Send VPC Flow Logs from each account to a centralized Amazon S3 bucket, then use Amazon Athena to query the logs.
D.Configure each account to send VPC Flow Logs to a central CloudWatch Logs group using cross-account subscription.
AnswerC

Sending VPC Flow Logs from each account to a centralized Amazon S3 bucket is correct because it creates a single, durable, cost-effective data lake that scales to petabytes. You configure each account's VPC Flow Logs to deliver to the same S3 bucket (with a bucket policy allowing cross-account delivery, ideally scoped to your AWS Organization ID). Then Amazon Athena can query these logs directly using standard SQL, with per-query pricing and no server to manage; using partition projection on account, region, and date drastically reduces scan costs and speeds up investigations.

Why this answer

It uses a centralized Amazon S3 bucket to aggregate VPC Flow Logs from all accounts, which is cost-effective (S3 storage costs are low) and enables cross-account querying via Amazon Athena using standard SQL. This approach avoids per-ingestion costs of services like CloudWatch Logs or Kinesis and provides a serverless, scalable query engine for analyzing logs across accounts.

Exam trap

The trap here is that candidates may overestimate the complexity of cross-account S3 access or underestimate the cost of CloudWatch Logs ingestion, leading them to choose Option D (central CloudWatch Logs group) which seems simpler but is actually more expensive and less query-friendly than S3+Athena.

How to eliminate wrong answers

Option A is wrong because Amazon OpenSearch Service (formerly Elasticsearch Service) incurs significant costs for ingestion and storage, and cross-account ingestion pipelines require complex setup with Lambda or Kinesis, making it less cost-effective than S3+Athena. Option B is wrong because Amazon Kinesis Data Analytics is designed for real-time stream processing, not for cost-effective historical querying across accounts; it would be overkill and expensive for periodic analysis of VPC Flow Logs. Option D is wrong because CloudWatch Logs cross-account subscriptions require each account to send logs to a central account's CloudWatch Logs group, which incurs per-ingestion costs and does not natively support SQL-based querying like Athena; querying across accounts would require additional tools or cross-account log group access, increasing complexity and cost.

186
MCQmedium

A company runs a serverless application using AWS Lambda and Amazon API Gateway. The application processes user uploads to an S3 bucket. The operations team uses CloudWatch Logs for monitoring, but they are finding it difficult to correlate logs across multiple Lambda functions that handle different parts of the workflow. The team wants to trace requests as they flow through the application and identify bottlenecks or errors. The team has already enabled CloudWatch Logs for all Lambda functions. What should the team do to achieve end-to-end request tracing?

A.Use CloudWatch Contributor Insights to analyze the log data and identify the top contributors to latency.
B.Use AWS CloudTrail to log all API calls and correlate them with CloudWatch Logs.
C.Create a CloudWatch ServiceLens service map to visualize the application components.
D.Enable AWS X-Ray on the Lambda functions and API Gateway to trace requests end-to-end.
AnswerD

AWS X-Ray provides distributed tracing by propagating a trace ID across instrumented services and recording segments and subsegments for each component. Enabling X-Ray on API Gateway and Lambda functions (e.g., via active tracing on the API Gateway stage and the Lambda execution role with X-Ray permissions) captures the full lifecycle of a request, including the API Gateway frontend, Lambda invocation, and any downstream SDK calls or HTTP requests. This allows you to follow a specific request through the entire architecture, view a service map, and drill into per-service latency, errors, and annotations, making it the correct solution for end-to-end tracing.

Why this answer

AWS X-Ray is the native distributed tracing service that integrates with Lambda and API Gateway to trace requests end-to-end across services. Enabling X-Ray on the Lambda functions and API Gateway propagates a trace ID through the workflow, letting the team visualize the full request path, latency per segment, and errors. CloudWatch ServiceLens can then surface the X-Ray traces alongside logs and metrics for correlation.

Exam trap

DOP-C02 often tests the difference between logging/auditing (CloudTrail, CloudWatch Logs) and distributed tracing (X-Ray), so candidates who focus on 'correlating logs' pick ServiceLens or Contributor Insights instead of enabling X-Ray.

How to eliminate wrong answers

Option A is wrong because CloudWatch Contributor Insights analyzes log data for top contributors but does not provide distributed tracing or end-to-end request correlation across services. Option B is wrong because CloudTrail records API calls for auditing and governance, not request-level tracing through application logic. Option C is wrong because ServiceLens is a visualization layer that depends on X-Ray traces; without X-Ray enabled, there is no trace data to map.

187
Multi-Selecteasy

A company is using AWS CloudFormation to deploy infrastructure. The DevOps team wants to receive notifications when a stack creation fails. Which services can be used together to send an email notification on stack failure? (Choose TWO.)

Select 2 answers
A.AWS Lambda
B.Amazon Simple Queue Service (SQS)
C.Amazon Simple Notification Service (SNS)
D.AWS CloudFormation
E.Amazon CloudWatch
AnswersC, D

Amazon SNS is a fully managed pub/sub messaging service that supports multiple subscription protocols, including email (as well as HTTP, Lambda, SQS, etc.). To receive CloudFormation stack event notifications, you create an SNS topic, subscribe an email address to it, and specify the topic ARN in the CloudFormation stack's `NotificationARNs` property. When the stack state changes, CloudFormation publishes the event to the topic, and SNS delivers an email to every confirmed subscriber.

Why this answer

Amazon SNS (Option C) is correct because it can send email notifications to subscribers when a CloudFormation stack creation fails. AWS CloudFormation (Option D) is correct because it can directly publish failure events to an SNS topic via the 'NotificationARNs' parameter in stack creation, enabling automated email alerts without additional services.

Exam trap

The trap here is that candidates might think CloudWatch (Option E) can send emails directly, but CloudWatch only publishes to SNS or other targets; it cannot natively deliver email notifications without SNS.

188
Multi-Selecthard

A company is using Amazon CloudWatch Synthetics canaries to monitor its web application endpoints. The canaries are failing intermittently with timeout errors. The DevOps team needs to troubleshoot the root cause. Which THREE actions should they take? (Select THREE.)

Select 3 answers
A.Use AWS CloudTrail to review Canary API calls.
B.Increase the canary timeout configuration to allow more time for the endpoint to respond.
C.Check the EC2 instance CPU utilization in the VPC where the canaries run.
D.Review VPC Flow Logs to see if requests are being dropped or denied.
E.Examine the canary logs in CloudWatch Logs for error messages.
AnswersB, D, E

If the timeout is too low, increasing it may resolve false positives.

Why this answer

Option B is correct because CloudWatch Synthetics canaries have a configurable timeout setting, and if the endpoint legitimately needs more time to respond, raising the timeout prevents intermittent timeout failures. Option D is correct because canaries can run inside a VPC, and VPC Flow Logs capture ACCEPT/REJECT records for traffic to and from the canary's ENIs, revealing whether security groups, NACLs, or routing are dropping the requests. Option E is correct because each canary writes execution logs, screenshots, and HAR files to CloudWatch Logs under /aws/lambda/cwsyn-* log groups, and these logs contain the exact error messages and timing data needed to diagnose the timeout.

Option A is not appropriate because CloudTrail records control-plane API calls (e.g., CreateCanary, UpdateCanary), not the canary's runtime HTTP request behavior. Option C is not appropriate because canaries run as Lambda functions managed by the Synthetics service, not on customer EC2 instances, so EC2 CPU utilization in the VPC is irrelevant to canary timeouts.

Exam trap

DOP-C02 often tests the scope of monitoring tools; the trap is confusing CloudTrail (API activity) with canary execution logs, or assuming canaries run on customer EC2 instances when they actually run on AWS-managed infrastructure.

189
MCQeasy

A company wants to detect and alert on unusual patterns in a CloudWatch Logs log group, such as a sudden spike in error messages, without writing complex queries. The team wants CloudWatch to learn normal patterns and notify when behavior deviates. Which feature should the DevOps engineer use?

A.CloudWatch Logs anomaly detection on a metric filter, which trains a model on the metric's historical behavior and alarms on deviations.
B.CloudWatch Logs Insights saved queries scheduled with an anomaly detection alarm.
C.CloudWatch Contributor Insights rules on the log group with a static alarm on the top contributor.
D.CloudWatch Logs metric filters with a static threshold alarm.
AnswerA

CloudWatch anomaly detection builds a model from up to two weeks of metric data and creates a band of expected values, alarming when the metric falls outside it. Applying it to a metric filter derived from the log group lets CloudWatch learn normal error patterns and notify on spikes without manual thresholds, matching the requirement.

Why this answer

CloudWatch anomaly detection learns a metric's normal pattern and alarms when values deviate from the expected band, removing the need to pick static thresholds. Applied to a metric filter on the log group, it can detect spikes in error messages. Static-threshold filters, Logs Insights queries, and Contributor Insights do not provide the same adaptive learning.

Exam trap

The trap here is treating a metric filter alarm as equivalent to anomaly detection, when the filter alarm still needs a manually chosen static threshold.

190
MCQmedium

A company wants to monitor network traffic to and from its VPC for security analysis. It needs to capture IP traffic information, including accepted and rejected connection attempts, and store the data in S3 for long-term analysis. Which AWS service should be used?

A.Amazon CloudWatch Logs
B.Amazon VPC Flow Logs
C.Amazon GuardDuty
D.AWS CloudTrail
AnswerB

Amazon VPC Flow Logs is the native service that captures IP traffic metadata for network interfaces in a VPC, including source and destination IPs, ports, protocol, and packet/byte counts for both accepted and rejected traffic. These logs can be published to Amazon S3 or CloudWatch Logs for long-term retention and analysis with services like Athena. As a result, VPC Flow Logs provides the flow-level visibility needed to monitor network traffic to and from a VPC.

Why this answer

Amazon VPC Flow Logs capture IP traffic metadata for ENIs, subnets, or VPCs, including ACCEPT and REJECT records, and can be published directly to Amazon S3 for long-term retention and Athena analysis. This matches the requirement to capture accepted and rejected connection attempts and store them in S3. Flow Logs are the standard DOP-C02 answer for VPC-level network traffic visibility.

Exam trap

The trap here is confusing GuardDuty (which analyzes traffic and produces findings) with VPC Flow Logs (which capture the raw traffic records), so candidates pick GuardDuty thinking it stores traffic in S3.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs is a log aggregation service, not a network traffic capture mechanism — it can receive Flow Logs as a destination but does not itself capture IP traffic. Option C is wrong because GuardDuty is a threat-detection service that consumes VPC Flow Logs, CloudTrail, and DNS logs internally; it produces findings, not raw traffic records, and does not let you store raw flow data in S3. Option D is wrong because CloudTrail records AWS API activity (control plane and optionally data plane), not IP-level network traffic between instances.

191
MCQmedium

A company is running a microservices application on Amazon ECS with AWS Fargate. The operations team needs to monitor application performance and troubleshoot slow API responses. They currently use Amazon CloudWatch Logs for container logs and have enabled Container Insights. However, they are unable to see detailed latency breakdowns per API endpoint. Which solution would provide the most granular visibility into API performance?

A.Enable detailed CloudWatch metrics for ECS and Fargate, including CPU and memory.
B.Enable CloudWatch Logs Insights to query API logs for slow requests.
C.Use AWS X-Ray to instrument the application and collect trace data.
D.Deploy the AWS Distro for OpenTelemetry collector on each task to send metrics to CloudWatch.
E.Set up VPC Flow Logs to analyze network latency between services.
AnswerC

X-Ray traces individual requests across microservices, producing segment and subsegment timings that break latency down per API endpoint and downstream call. Container Insights only aggregates task and service metrics, so it cannot satisfy the stem's requirement for granular per-endpoint latency visibility.

Why this answer

AWS X-Ray provides end-to-end tracing of requests as they travel through microservices, capturing detailed latency breakdowns per API endpoint, including downstream calls, database queries, and external HTTP requests. This gives the operations team the granular visibility needed to pinpoint exactly where slow responses occur, unlike aggregated metrics or log-based queries.

Exam trap

The trap here is that candidates confuse infrastructure-level metrics (CPU, memory, network) or log-based querying with the distributed tracing capability needed to break down latency per API endpoint, overlooking that only X-Ray provides end-to-end trace segments with sub-millisecond timing per service call.

How to eliminate wrong answers

Option A is wrong because enabling detailed CloudWatch metrics for ECS and Fargate (CPU, memory, network) provides infrastructure-level metrics, not per-endpoint latency breakdowns. Option B is wrong because CloudWatch Logs Insights can query logs for slow requests but cannot trace a single request across multiple services or show the latency contributed by each downstream call. Option D is wrong because the AWS Distro for OpenTelemetry collector sends metrics and traces to CloudWatch, but without X-Ray integration or trace sampling, it does not provide the per-endpoint latency breakdowns that X-Ray's service map and trace segments offer.

Option E is wrong because VPC Flow Logs capture network-level metadata (packet headers, timestamps) and can indicate network latency between ENIs, but they cannot reveal application-level latency per API endpoint or trace a request through microservices.

192
MCQmedium

A company is running a web application on Amazon EC2 instances behind an Application Load Balancer. The application is experiencing intermittent errors. The DevOps engineer needs to identify if the errors are caused by the application or the underlying infrastructure. Which solution provides the MOST detailed visibility into the application's behavior?

A.Enable VPC Flow Logs and analyze traffic patterns
B.Enable AWS CloudTrail and monitor for API errors
C.Instrument the application with AWS X-Ray SDK and analyze traces
D.Enable detailed CloudWatch metrics on the EC2 instances and ALB
AnswerC

Instrumenting the application with the X-Ray SDK adds tracing headers to each incoming HTTP request and tracks the request as it flows through the EC2-hosted application and any downstream calls to databases or other services. X-Ray segments and subsegments capture the exact service, operation, and any exceptions or faults that occur, enabling a trace map that pinpoints the failing component or code path. This gives per-request context and error stack traces needed to resolve application-level errors definitively.

Why this answer

AWS X-Ray traces individual requests as they travel through the application, showing latency, errors, and downstream calls at each segment. This provides the granular, request-level visibility needed to determine whether errors originate in application code or in dependencies like databases or external services. CloudWatch metrics and VPC Flow Logs operate at a coarser level and cannot pinpoint application logic failures.

Exam trap

DOP-C02 often tests whether candidates conflate infrastructure-level monitoring (CloudWatch, VPC Flow Logs, CloudTrail) with application-level tracing — the key discriminator is whether the tool can follow a single request through the application code.

How to eliminate wrong answers

Option A is wrong because VPC Flow Logs capture IP-level network metadata (source/destination, ports, accept/reject) and cannot reveal application-layer errors or code paths. Option B is wrong because CloudTrail records AWS API calls for auditing, not the runtime behavior of application requests. Option D is wrong because detailed CloudWatch metrics on EC2 and ALB show aggregate performance counters (CPU, latency, 5xx counts) but do not trace individual requests through the application stack.

193
Multi-Selecteasy

A DevOps engineer is troubleshooting a performance issue with an Amazon RDS for MySQL database. The engineer suspects that slow queries are causing high CPU utilization. Which TWO actions can the engineer take to identify the slow queries?

Select 2 answers
A.Create an RDS event subscription for 'low storage' events.
B.Monitor the 'CPUUtilization' metric in CloudWatch.
C.Enable the slow query log and publish it to CloudWatch Logs.
D.Enable Performance Insights to visualize database load and identify top SQL statements.
E.Enable Enhanced Monitoring to view process list and SQL queries.
AnswersC, D

Enabling the RDS slow query log captures every SQL statement whose execution exceeds the configured long_query_time threshold, recording details like query text, execution duration, rows examined, and timestamps. By publishing these log events to CloudWatch Logs, you gain a searchable, historical record that can be queried with CloudWatch Logs Insights to filter for the slowest statements, identify patterns, and correlate with other metrics. This directly exposes the exact SQL causing performance issues, making it a definitive diagnostic method for slow query analysis.

Why this answer

Enable the slow query log and publish it to CloudWatch Logs (Option C) allows you to capture and analyze slow SQL queries. Performance Insights (Option D) provides a dashboard to visualize database load and identify the top SQL statements causing performance issues. Option A is incorrect because event subscriptions for low storage notify about storage events, not slow queries.

Option B is incorrect because monitoring CPUUtilization only indicates high CPU usage but does not identify specific slow queries. Option E is incorrect because Enhanced Monitoring provides OS-level metrics like CPU and memory, not the actual SQL queries.

194
MCQeasy

A Lambda function is timing out. The log above shows a recent invocation. What is the most likely cause?

A.The function is running out of memory.
B.The function is being invoked too frequently.
C.The function is experiencing a cold start.
D.The function timeout is set too low.
AnswerD

The correct diagnosis is that the function's timeout setting is too low. The 3000 ms Duration exactly matches the default Lambda timeout of 3 seconds, and the 'Task timed out' error means Lambda terminated the handler at its configured limit. Raising the timeout, after reviewing the code for inefficiencies, would allow the function to complete successfully.

Why this answer

The log shows the function timed out at 3000 ms, which is the default Lambda timeout (3 seconds). The correct answer is D because the timeout value is set too low, causing the function to be terminated before it can complete. Option A is incorrect because memory usage is only 64 MB out of 128 MB, so insufficient memory is not the issue.

Option B is incorrect because there is only one invocation shown; frequent invocations would cause throttling, not a timeout. Option C is incorrect because the init duration is normal (e.g., 2.34 ms), indicating that a cold start is not the cause; cold starts increase latency but do not cause timeouts if the function runs within the timeout limit.

195
Multi-Selectmedium

A DevOps engineer is designing a centralized logging solution for a multi-account AWS environment. The solution must be cost-effective and provide real-time log analysis. Which THREE services should they consider?

Select 3 answers
A.Amazon OpenSearch Service (Elasticsearch)
B.Amazon Kinesis Data Firehose
C.Amazon S3
D.Amazon CloudWatch Logs
E.AWS CloudTrail
AnswersA, B, D

Amazon OpenSearch Service provides the interactive search and visualization layer for a centralized logging solution. Logs ingested from Firehose, CloudWatch Logs subscriptions, or Logstash are indexed and available for real-time queries, aggregations, and Kibana dashboards, making it the correct target for log analytics. Unlike object storage or API audit trails, OpenSearch supports full-text search, ad hoc filtering, and anomaly detection across massive log volumes.

Why this answer

Amazon CloudWatch Logs (D) is correct because it is the native AWS service for collecting, storing, and monitoring log data from EC2 instances, Lambda functions, and other AWS resources, and it supports real-time metric filters and subscription filters for analysis. Amazon Kinesis Data Firehose (B) is correct because it reliably streams log data in near real-time to destinations such as Amazon S3, Amazon OpenSearch Service, or Splunk, and it can transform and batch records cost-effectively. Amazon OpenSearch Service (A) is correct because it provides real-time search, analytics, and visualization of log data via Kibana, making it ideal for interactive log analysis in a centralized multi-account setup.

Amazon S3 (C) is not marked correct because, while it is a cost-effective storage destination for logs, it is not a real-time analysis service on its own. AWS CloudTrail (E) is not marked correct because it records API activity and account events for auditing, not application or system log aggregation and real-time analysis.

Exam trap

The trap is selecting S3 as an analysis service or CloudTrail as a general log source — candidates must distinguish storage vs. analysis and audit logs vs. application logs.

196
MCQhard

A company runs a fleet of EC2 instances behind an Auto Scaling group. The DevOps team wants to detect and respond to memory leaks in their application. They have configured CloudWatch agent to collect memory metrics. However, the metric shows unpredictable spikes. The team needs to correlate these spikes with application logs to identify the root cause. Which solution provides the BEST correlation?

A.Export the memory metric and application logs to Amazon S3 and use Amazon Athena to join them
B.Enable AWS X-Ray on the application to trace requests and identify memory-heavy requests
C.Use CloudWatch Logs Insights to query application logs for error patterns around the time of memory spikes
D.Use Amazon EventBridge to capture EC2 instance state changes and correlate with memory spikes
AnswerC

CloudWatch Logs Insights provides a queryable index of live application logs, so you can run interactive queries that filter for ERROR, FATAL, or exception patterns within the exact time window of a memory spike. Using the parse command on structured logs and grouping by timestamp or host, you can quickly identify a correlated error or log burst, making it the appropriate tool for real-time operational correlation between CloudWatch metrics and log data.

Why this answer

CloudWatch Logs Insights allows you to query application logs directly in CloudWatch Logs using a purpose-built query language. By filtering logs around the timestamps of memory spikes, you can correlate specific error patterns or log entries with the metric data, enabling root cause analysis without moving data or adding complexity.

Exam trap

The trap here is that candidates often confuse AWS X-Ray's request tracing with OS-level metric correlation, or assume that exporting to S3 and using Athena is a universal solution, when in fact CloudWatch Logs Insights provides the most direct and efficient correlation within the same monitoring ecosystem.

How to eliminate wrong answers

Option A is wrong because exporting metrics and logs to S3 and using Athena to join them introduces unnecessary latency, cost, and complexity; Athena is designed for ad-hoc analysis of structured data, not real-time correlation of streaming metrics and logs. Option B is wrong because AWS X-Ray traces requests and identifies latency or errors, but it does not capture memory metrics or correlate them with memory leaks; it focuses on distributed tracing, not OS-level resource usage. Option D is wrong because EventBridge captures EC2 instance state changes (e.g., start, stop, terminate), which are unrelated to memory spikes caused by application-level memory leaks; state changes do not provide the granular log correlation needed.

197
Multi-Selecteasy

A DevOps engineer needs to monitor the health of a web application running on EC2 instances behind an Application Load Balancer (ALB). Which TWO metrics from ALB should be monitored to detect application errors? (Choose TWO.)

Select 2 answers
A.RequestCount.
B.HTTPCode_ELB_5XX_Count.
C.HTTPCode_Target_5XX_Count.
D.HealthyHostCount.
E.TargetResponseTime.
AnswersC, E

HTTPCode_Target_5XX_Count is the definitive metric for monitoring application health because it directly counts HTTP 5xx responses returned by the registered targets (EC2 instances, containers, or Lambda). This value reflects actual application failures—unhandled exceptions, internal server errors, or gateway timeouts—making it the most precise indicator of unhealthy application logic. It is the recommended metric to alarm on for detecting when your web application is returning server-side errors to users.

Why this answer

(HTTPCode_Target_5XX_Count) is correct because it counts HTTP 5XX errors returned by the target (EC2 instances), indicating application errors. Option E (TargetResponseTime) is correct because elevated response times can indicate application performance issues or errors, and is a key metric for detecting application problems. Option A (RequestCount) is incorrect because it simply counts total requests, not errors.

Option B (HTTPCode_ELB_5XX_Count) is incorrect because it counts 5XX errors generated by the ALB itself (e.g., due to configuration issues), not application errors. Option D (HealthyHostCount) is incorrect because it indicates the number of healthy targets, not application-level errors.

← PreviousPage 3 of 3 · 197 questions total

Ready to test yourself?

Try a timed practice session using only Monitoring and Logging questions.