Courseiva

CCNA Monitoring and Logging Questions

75 of 197 questions · Page 2/3 · Monitoring and Logging · Answers revealed

76
Multi-Selectmedium

A DevOps engineer is setting up centralized logging for a multi-account environment using AWS Organizations. The engineer needs to aggregate logs from all accounts into a single Amazon S3 bucket. Which TWO steps are necessary?

Select 2 answers
A.Create IAM roles in each account to allow the central bucket to read logs.
B.Create a bucket policy on the central S3 bucket that grants permissions to the source accounts.
C.Enable CloudTrail organization trail in the management account to deliver logs to the central bucket.
D.Set up a cross-account subscription in CloudWatch Logs to forward logs to the central account.
E.Configure each account’s services (e.g., CloudTrail, VPC Flow Logs) to deliver logs to the central S3 bucket.
AnswersB, E

The central S3 bucket needs a resource-based policy that explicitly grants the source accounts' log-delivery services (e.g., CloudTrail, VPC Flow Logs) permission to write objects and read the bucket ACL. Without this bucket policy, cross-account writes from other accounts will be denied by default. The policy must reference the source account IDs or the organization ID and include conditions like aws:SourceAccount or aws:SourceArn to prevent confused deputy attacks. This is a mandatory step to enable centralized log collection.

Why this answer

A bucket policy on the central S3 bucket can grant cross-account permissions to source accounts to write logs. This allows services like CloudTrail and VPC Flow Logs from member accounts to deliver logs directly to the central bucket without requiring IAM roles in each account for reading logs.

Exam trap

The trap here is that candidates often confuse the need for IAM roles in each account (Option A) with the correct bucket policy approach, or they assume that enabling an organization trail (Option C) is mandatory when the question allows for individual account configuration.

77
MCQhard

A company uses Amazon S3 to store sensitive data. The security team wants to be notified when an S3 bucket policy is modified. Which approach is most efficient?

A.Create an Amazon EventBridge rule that matches the 'PutBucketPolicy' API call and sends a notification to an SNS topic.
B.Set up an AWS Config rule to detect changes to the bucket policy.
C.Configure S3 event notifications for 's3:PutBucketPolicy' on the bucket.
D.Enable S3 server access logs and use CloudWatch Logs Insights to run queries periodically.
AnswerA

EventBridge matches the PutBucketPolicy API call from CloudTrail and routes it to an SNS topic, giving near-real-time notification without polling. This is more efficient than scheduled log analysis or S3 event notifications, which do not cover bucket policy changes.

Why this answer

Amazon EventBridge can match AWS API calls such as PutBucketPolicy via CloudTrail management events and route them to an SNS topic for notification. This is the most efficient, near-real-time, event-driven approach for alerting on bucket policy changes.

Exam trap

DOP-C02 often tests the difference between event-driven detection (EventBridge) and periodic/compliance-based detection (Config, access logs) — candidates may pick S3 event notifications, which do not cover bucket policy API calls.

How to eliminate wrong answers

Option B is wrong because AWS Config rules detect configuration changes but are not designed for immediate event-driven notification and can be slower and less direct for this use case. Option C is wrong because S3 event notifications only fire for object-level events (like PutObject), not for bucket policy API calls such as PutBucketPolicy. Option D is wrong because S3 server access logs record data-plane requests and querying them periodically is not real-time or efficient for policy-change alerting.

78
MCQhard

A company is running a critical application on Amazon ECS with Fargate launch type. The application writes logs to Amazon CloudWatch Logs. The DevOps team needs to set up an alert when the application generates more than 100 error logs in any 5-minute window. Which configuration should be used?

A.Create a CloudWatch Logs Insights query that runs every 5 minutes and triggers an SNS notification
B.Create an Amazon EventBridge rule that matches CloudWatch Logs events for the word 'ERROR' and triggers an alarm
C.Create a CloudWatch Logs metric filter for 'ERROR' and a CloudWatch alarm on the resulting metric with a period of 5 minutes
D.Enable AWS CloudTrail logging for the ECS task and create a metric filter on CloudTrail logs
AnswerC

A CloudWatch Logs metric filter is applied to a log group in near-real time as log events are ingested, and it uses pattern syntax to count occurrences of the string 'ERROR' and emit a custom metric (e.g., ErrorCount) for that log group. The CloudWatch alarm is then configured on that custom metric with an evaluation period of 5 minutes, so if the number of ERROR log lines within a 5-minute interval exceeds the configured threshold, the alarm state changes to ALARM and can trigger an SNS notification. This is the standard, fully managed mechanism for log-pattern-based alerting because it does not require any custom code or additional infrastructure, and it integrates directly with CloudWatch alarm actions.

Why this answer

A CloudWatch Logs metric filter can be configured to count occurrences of the word 'ERROR' in log streams. This filter creates a custom metric that can be monitored by a CloudWatch alarm with a period of 5 minutes. When the metric exceeds the threshold of 100, the alarm triggers an action such as an SNS notification.

Option A is incorrect because CloudWatch Logs Insights is a query tool for interactive analysis, not for continuous real-time alerting. Option B is incorrect because EventBridge events are not generated from log content directly; you would need a metric filter or subscription filter to turn log data into events. Option D is incorrect because AWS CloudTrail logs API activities, not application error logs written to CloudWatch Logs.

79
MCQmedium

A DevOps engineer is setting up centralized logging for multiple AWS accounts. They need to collect VPC Flow Logs, CloudTrail logs, and application logs into a single Amazon S3 bucket. What is the most efficient approach?

A.Configure a Lambda function in each account to copy logs to a central S3 bucket.
B.Create an S3 bucket in each account and use S3 replication.
C.Use Amazon Kinesis Data Firehose to stream logs from all accounts to a central S3 bucket.
D.Use an S3 bucket in a centralized logging account with a bucket policy that grants write access from all other accounts.
AnswerD

Placing a single S3 bucket in a centralized logging account with a bucket policy that grants s3:PutObject to principals in all source accounts is the most efficient pattern because CloudTrail, VPC Flow Logs, and similar services can deliver logs cross-account natively without additional moving parts. The bucket policy can restrict writes using aws:SourceArn or aws:SourceAccount conditions to prevent the confused-deputy problem, and the central account owns objects for unified lifecycle and access control. This direct-write model achieves lower latency and zero maintenance compared to middleware or copying mechanisms.

Why this answer

It uses a centralized logging account with a single S3 bucket configured with a bucket policy that grants write access (s3:PutObject) to all other accounts. This approach avoids data duplication, eliminates the need for replication or intermediate compute resources, and is the most efficient and cost-effective method for aggregating logs from multiple accounts into a single destination.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing managed services like Kinesis or Lambda, when a simple S3 bucket policy with cross-account write access is the most efficient and AWS-recommended approach for centralized log aggregation.

How to eliminate wrong answers

Option A is wrong because using a Lambda function in each account to copy logs to a central S3 bucket introduces unnecessary complexity, potential single points of failure, and additional cost from Lambda invocations and data transfer, making it less efficient than a direct write approach. Option B is wrong because creating an S3 bucket in each account and using S3 replication results in data duplication, increased storage costs, and replication latency, and it requires managing multiple buckets and replication rules, which is less efficient than a single bucket with a cross-account policy. Option C is wrong because Amazon Kinesis Data Firehose is designed for streaming data ingestion and transformation, but it adds unnecessary complexity and cost for log aggregation when a simpler S3 bucket policy can achieve the same result; Firehose is better suited for real-time processing needs, not for batch log collection from multiple accounts.

80
Multi-Selecthard

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer. The DevOps team has enabled detailed CloudWatch metrics for the ALB and is using CloudWatch Logs for the EC2 instances. Recently, users report intermittent 503 errors. The team notices that the ALB's 'RequestCount' metric shows a sudden drop during error periods, while the 'ActiveConnectionCount' remains steady. Which TWO steps should the team take to diagnose the issue? (Choose two.)

Select 2 answers
A.Enable and analyze the ALB access logs to see the HTTP response codes and target processing time.
B.Check the Amazon Route 53 health checks for the ALB DNS name.
C.Review the EC2 instances' CloudWatch metrics for CPU utilization and network traffic.
D.Observe the ALB's 'UnhealthyHostCount' metric and check target group health checks.
E.Inspect AWS CloudTrail logs for the ALB to see if there are any configuration changes.
AnswersA, D

ALB access logs capture full request-level detail for every HTTP request the load balancer receives, including the exact HTTP response code returned to the client (e.g., 503) and the target processing time, which is the time the EC2 instance took to respond. By filtering for 503 responses and correlating them with target IPs and timestamps, you can confirm whether errors originate from unhealthy targets, overloaded instances, or misconfigured routing. This is the most direct diagnostic because access logs are designed to expose the request/response path, rather than merely indicating that a threshold was crossed.

Why this answer

ALB access logs contain detailed per-request data, including HTTP response codes (e.g., 503) and target processing time. Analyzing these logs will reveal whether the 503 errors are coming from the ALB itself (e.g., due to request queue overflow) or from the targets, and whether the sudden drop in RequestCount is due to clients aborting or the ALB throttling requests.

Exam trap

The trap here is that candidates often focus on EC2-level metrics (CPU, network) or CloudTrail config changes, missing that the ALB's own health check and access log data are the direct sources for diagnosing 503 errors tied to target unavailability or request queue limits.

81
MCQmedium

A DevOps engineer receives an alarm that an EC2 instance's StatusCheckFailed metric has been in ALARM state for 10 minutes. Which action should the engineer take first to investigate?

A.Review the instance's system log and application logs in CloudWatch Logs
B.Use AWS Config to check the instance's configuration compliance
C.Check AWS CloudTrail for any API calls that modified the instance
D.Restart the EC2 instance to clear the alarm
AnswerA

Instance status checks detect OS-level problems such as failed system boot, kernel panic, or filesystem corruption. The instance's system log (console output) shows the boot sequence and kernel messages, while application logs streamed to CloudWatch Logs reveal software-level errors that may have caused the failure. Reviewing both gives the evidence needed to identify the root cause before taking any corrective action.

Why this answer

When StatusCheckFailed is in ALARM, the first investigative step is to examine the instance's system log and application logs in CloudWatch Logs, because these reveal whether the failure is an OS-level crash, kernel panic, or application error causing the instance to fail its status checks. Reviewing logs is non-destructive and provides the diagnostic evidence needed before taking any remediation action.

Exam trap

The trap is treating a remediation action (restart the instance) as the first step — the exam tests whether you know to investigate and gather diagnostic evidence from logs before taking any corrective action that could destroy the evidence.

How to eliminate wrong answers

Option B is wrong because AWS Config tracks resource configuration changes and compliance, not runtime health or the cause of a status check failure — it would not explain why the instance is failing. Option C is wrong because CloudTrail records API activity (who changed what), which is useful if you suspect a recent configuration change, but it does not diagnose the instance's runtime failure and is not the first step for a health alarm. Option D is wrong because restarting the instance is a remediation action, not an investigation step, and it destroys the running state that could reveal the root cause — rebooting before gathering logs is a classic anti-pattern.

82
MCQmedium

A company uses AWS CloudTrail to log API activity across multiple accounts. The security team needs to ensure that all CloudTrail logs are delivered to a centralized S3 bucket in the audit account, and that any log file validation failures trigger an immediate notification. What should the engineer do to meet this requirement?

A.Enable CloudTrail log file validation and create a CloudWatch alarm on the DigestDeliveryFailed metric
B.Create a Lambda function that checks the integrity of logs and publishes to SNS
C.Configure CloudTrail to deliver logs to the S3 bucket and enable SNS notifications for all events
D.Send CloudTrail logs to CloudWatch Logs and create a metric filter for validation errors
AnswerA

CloudTrail's built-in log file validation creates hash-chained digest files that are digitally signed with a private key, and you can verify them with the public key published by AWS. The CloudTrail service emits the DigestDeliveryFailed metric to CloudWatch when it cannot deliver a digest file, which is often the first sign of tampering or a delivery problem. Creating a CloudWatch alarm on this metric gives you immediate, proactive notification without building custom code.

Why this answer

Enabling CloudTrail log file validation triggers the generation of digest files that contain hash values for verifying log file integrity. CloudTrail also emits the DigestDeliveryFailed metric to CloudWatch when a digest file delivery fails. Creating a CloudWatch alarm on this metric allows you to send immediate notifications via SNS when a validation failure occurs.

Option B (Lambda function) is unnecessary because CloudTrail already provides the necessary metrics for this alerting. Option C (SNS notifications for all events) would generate excessive notifications and does not directly address log file validation failures. Option D (CloudWatch Logs and metric filter) is not the standard approach; CloudTrail directly emits the DigestDeliveryFailed metric, which is simpler and more reliable.

83
MCQmedium

A DevOps team needs to monitor failed API calls in their AWS account. They want to receive notifications when specific IAM actions, such as DeleteBucket, fail. Which service should they use?

A.AWS CloudTrail and Amazon EventBridge.
B.AWS Config rules.
C.Amazon S3 server access logs.
D.CloudWatch Logs and metric filters.
AnswerA

AWS CloudTrail records all management API calls in the account, including failed attempts, with metadata such as the IAM principal, source IP, event name, and error codes. Amazon EventBridge can be configured with a rule whose event pattern matches CloudTrail's api_call events and a filter condition on errorCode, routing the matched events to an SNS topic for real-time alerting. This combination is purpose-built for monitoring failed API calls.

Why this answer

AWS CloudTrail captures API calls, and Amazon EventBridge (formerly CloudWatch Events) can be used to create rules that match specific failed API calls (e.g., DeleteBucket) and trigger notifications. Option B is incorrect because AWS Config rules monitor resource configuration compliance, not API call failures. Option C is incorrect because S3 server access logs log requests made to an S3 bucket, not IAM API calls.

Option D is incorrect because CloudWatch Logs and metric filters are used to monitor log data, but they are not the primary service for capturing API calls; CloudTrail is needed for that.

84
MCQmedium

A company is using Amazon CloudWatch Logs Insights to analyze application logs. The DevOps team needs to create a metric filter that counts occurrences of the word 'ERROR' in the log events. Which CloudWatch Logs Insights query should be used to test the metric filter?

A.fields @timestamp, @message | stats count() by bin(5m)
B.fields @timestamp, @message | filter @message like /ERROR/
C.fields @timestamp, @message | parse @message '[*] *' as @severity, @log
D.fields @timestamp, @message | sort @timestamp desc
AnswerB

fields @timestamp, @message | filter @message like /ERROR/ applies a regular expression filter to the raw message text, returning only the individual log events that contain the substring 'ERROR'. This directly mirrors the behavior of the CloudWatch Logs metric filter pattern "ERROR", allowing you to see the exact events that would increment the metric. It is the appropriate query for testing whether the pattern matches the intended production logs before creating or updating the metric filter.

Why this answer

To test a metric filter that counts occurrences of the word 'ERROR' in log events, you need a CloudWatch Logs Insights query that filters log events containing 'ERROR'. The query `fields @timestamp, @message | filter @message like /ERROR/` does exactly that: it selects the timestamp and message fields and filters for messages that match the regular expression /ERROR/. This allows you to verify that the filter pattern will correctly identify the relevant log events.

Exam trap

The trap is choosing a query that counts or sorts logs without filtering for the specific pattern. Candidates might think that any query that includes @message is sufficient, but the key is to filter for 'ERROR'. Also, some might forget that the filter must match the exact pattern used in the metric filter, which often includes a regex.

How to eliminate wrong answers

Option A is wrong because it uses `stats count() by bin(5m)`, which counts all log events in 5-minute bins without filtering for 'ERROR'; it does not test the metric filter's pattern. Option C is wrong because it parses the message into severity and log fields but does not filter for 'ERROR'; it would return all logs, not just those containing 'ERROR'. Option D is wrong because it sorts logs by timestamp descending but does not filter for 'ERROR'; it would return all logs in reverse chronological order.

85
MCQhard

A company runs a multi-region application on Amazon EC2 instances across us-east-1 and eu-west-1. The application uses an Amazon Aurora global database for writes in us-east-1 and reads in eu-west-1. The DevOps team wants to monitor the replication lag between the primary and secondary regions. They have set up a CloudWatch alarm on the AuroraReplicaLag metric in both regions. However, they notice that the alarm in eu-west-1 sometimes triggers false positives when the lag spikes briefly but then recovers. The team wants to reduce false alarms while still being alerted to sustained high lag that could impact read replicas. The team is already using a standard CloudWatch alarm with a period of 1 minute and evaluation periods of 1. What should the team change to reduce false positives?

A.Increase the alarm threshold to a higher value, such as 10 seconds.
B.Reduce the metric period to 30 seconds to get more granular data.
C.Increase the number of evaluation periods to 3, so the alarm triggers only if the lag is high for 3 consecutive minutes.
D.Create a composite alarm that triggers when both the AuroraReplicaLag and CPUUtilization metrics are high.
AnswerC

Increasing the number of evaluation periods to 3 with a 1-minute period means the alarm must observe the replication lag exceeding the threshold for three consecutive datapoints (three consecutive minutes) before entering ALARM state. This creates a temporal smoothing effect that filters out brief, self-correcting spikes in AuroraReplicaLag, which are often caused by momentary write bursts or replica catch-up delays. CloudWatch evaluates all three most recent datapoints when determining alarm state, so sustained high lag triggers the alarm, while isolated outliers do not, directly reducing false positives without changing the sensitivity to genuine long-duration lag.

Why this answer

The false positives occur because a single 1-minute evaluation period triggers on brief, transient lag spikes. Increasing the number of evaluation periods to 3 means the alarm fires only when the lag exceeds the threshold for 3 consecutive 1-minute periods, filtering out short spikes while still catching sustained high lag. This is the standard CloudWatch approach to reduce false positives without losing sensitivity to real issues.

Exam trap

DOP-C02 often tests the difference between changing a threshold (sensitivity) and changing evaluation periods/datapoints (transient filtering), and candidates frequently pick threshold changes or composite alarms when the real fix is sustained-breach evaluation.

How to eliminate wrong answers

Option A is wrong because raising the threshold to 10 seconds may reduce some false positives but also risks missing real sustained lag that is still impactful — it changes sensitivity rather than filtering transients. Option B is wrong because reducing the period to 30 seconds increases granularity and would make the alarm more sensitive to spikes, worsening false positives. Option D is wrong because a composite alarm combining AuroraReplicaLag with CPUUtilization adds an unrelated condition and does not address the transient-spike problem; it could also suppress real lag alerts when CPU is normal.

86
Multi-Selectmedium

A company is using Amazon CloudWatch Logs to collect logs from multiple EC2 instances. They need to filter logs in real time and send specific log events to a custom application for processing. Which TWO services can they use to achieve this?

Select 2 answers
A.Use Amazon Kinesis Data Analytics to process the log stream.
B.Configure a CloudWatch Logs subscription filter that invokes an AWS Lambda function.
C.Create a CloudWatch Events rule to capture log events and send them to Amazon SQS.
D.Configure a CloudWatch Logs subscription filter that sends data to Amazon Kinesis Data Firehose.
E.Use Amazon S3 event notifications to trigger a Lambda function on new log files.
AnswersB, D

A CloudWatch Logs subscription filter can be configured with a Lambda function as its destination, enabling real-time processing of log events as they arrive. When a log event matches the filter pattern, CloudWatch Logs invokes the Lambda function asynchronously, passing the event payload (base64-encoded and gzipped) for your custom logic to parse and forward. This is a native, low-latency pattern that satisfies the requirement to filter and stream logs in real time, and it is a widely recommended approach for building log-processing pipelines.

Why this answer

Option B is correct because a CloudWatch Logs subscription filter performs real-time filtering of log events and can deliver matching events directly to AWS Lambda, which then runs the custom application logic for processing. Option D is correct because a subscription filter can also stream filtered log events to Amazon Kinesis Data Firehose, which reliably delivers them to a custom destination for processing. Option A is incorrect because Kinesis Data Analytics analyzes streaming data but is not a CloudWatch Logs subscription destination for real-time log filtering.

Option C is incorrect because CloudWatch Events (EventBridge) rules react to AWS service events, not individual CloudWatch Logs log events, and cannot filter log events to SQS. Option E is incorrect because S3 event notifications only trigger on object creation in S3 and do not provide real-time filtering of CloudWatch Logs streams.

Exam trap

DOP-C02 often tests the difference between CloudWatch Logs subscription filters and CloudWatch Events; candidates may confuse the two and select CloudWatch Events for log processing.

87
Multi-Selectmedium

A DevOps engineer needs to set up centralized logging for an application running on multiple EC2 instances across different AWS accounts. The logs must be aggregated in a single S3 bucket and also be analyzed in near real-time. Which TWO services should be used together to achieve this?

Select 2 answers
A.Amazon Simple Queue Service (SQS)
B.Amazon Kinesis Data Firehose
C.AWS CloudTrail
D.Amazon CloudWatch Logs subscription
E.AWS Lambda
AnswersB, D

Amazon Kinesis Data Firehose is a fully managed streaming ingestion service that can receive log events directly from a CloudWatch Logs subscription and deliver them near-real-time to centralized destinations such as Amazon S3, Amazon Redshift, or Amazon OpenSearch Service. It automatically handles buffering, compression, and encryption, which reduces cost and operational overhead. This makes it the ideal component for aggregating logs from multiple AWS accounts into a single, queryable data lake for centralized analysis.

Why this answer

Option B, Amazon Kinesis Data Firehose, is correct because it can ingest streaming log data and reliably deliver it to a single S3 bucket, with optional near real-time transformation and buffering, satisfying both the aggregation and near real-time analysis requirements. Option D, Amazon CloudWatch Logs subscription, is correct because a subscription filter on a CloudWatch log group can stream log events in near real time to Kinesis Data Firehose, enabling centralized collection from EC2 instances across multiple AWS accounts. Together, CloudWatch Logs subscriptions feed Firehose, which delivers the aggregated logs to S3.

Option A, Amazon SQS, is not appropriate because it is a message queue, not a log ingestion or delivery service to S3. Option C, AWS CloudTrail, records API activity rather than application logs, so it does not meet the application logging requirement. Option E, AWS Lambda, is compute for processing events and is not the primary service for aggregating logs into S3.

Exam trap

DOP-C02 often tests the confusion between CloudWatch Logs (application logs) and CloudTrail (API audit logs), and between SQS (decoupling) and Firehose (delivery to S3) — candidates must recognize that only the CloudWatch Logs subscription + Firehose pairing provides both aggregation and near real-time delivery.

88
MCQeasy

A company uses Amazon CloudWatch to monitor its production environment. The DevOps team wants to receive an email notification whenever the average CPU utilization of any EC2 instance exceeds 90% for 5 consecutive minutes. Which steps should be taken to set up this notification?

A.Install the CloudWatch Logs agent on each EC2 instance and configure a metric filter to trigger an SNS notification
B.Create a CloudWatch alarm on CPUUtilization with a threshold of 90% for 5 consecutive periods, and configure an SNS topic to send email
C.Use AWS CloudTrail to monitor CPU utilization and send notifications via SNS
D.Use AWS Config to create a rule that triggers an SNS notification when CPU utilization exceeds 90%
AnswerB

CloudWatch automatically emits the CPUUtilization metric for each EC2 instance, expressed as a percentage, so an alarm can directly evaluate it. Configuring the alarm with a threshold of 90% for five consecutive evaluation periods (5-minute periods, or 1-minute if detailed monitoring is enabled) ensures the alert only fires on sustained load, not transient spikes. The action sends an SNS notification to a topic with an email subscription, which is the canonical AWS approach for metric-based alerting.

Why this answer

Create a CloudWatch alarm on the CPUUtilization metric with a threshold of 90% for 5 consecutive periods (5 minutes assuming 1-minute periods), and configure an SNS topic to send email notifications. Option A is incorrect because the CloudWatch Logs agent collects log data, not metrics, and metric filters are for log analysis, not setting alarms on CPU utilization. Option C is incorrect because AWS CloudTrail logs API calls, not CPU utilization metrics.

Option D is incorrect because AWS Config is for resource configuration auditing and compliance, not for monitoring CPU utilization metrics.

89
MCQmedium

A DevOps team is troubleshooting a slow application. They enabled AWS X-Ray tracing and see that one of the downstream services has a high average response time. However, the traces show that the service itself is fast; the delay is in the network call from the upstream service. Which X-Ray feature should the team use to identify the root cause?

A.Examine the trace map to see the connection between services.
B.Add annotations to the traces for better filtering.
C.View the raw segments of the upstream service.
D.Adjust the sampling rules to capture more traces.
AnswerA

The AWS X-Ray trace map is the correct tool because it visualizes each service as a node and the connections between them as edges, with latency metrics for each edge. This directly reveals whether the identified slowness stems from network communication between services (for example, high I/O wait or retries) rather than from code execution inside a single service, allowing the DevOps team to pinpoint the exact segment of the request path that is underperforming.

Why this answer

The trace map in AWS X-Ray provides a visual representation of the service graph, showing the connections and latency between services. Since the delay is in the network call from the upstream service to the downstream service, the trace map can highlight the specific edge where the high latency occurs, allowing the team to pinpoint whether the issue is due to network congestion, DNS resolution, or a slow HTTP connection. This is the most direct way to identify the root cause of the inter-service communication delay.

Exam trap

The trap here is that candidates might focus on the downstream service's segment (Option C) thinking the delay is inside that service, when the trace map is specifically designed to reveal inter-service communication latency that is not captured by individual segment durations.

How to eliminate wrong answers

Option B is wrong because annotations are key-value pairs added to traces for custom metadata filtering, not for diagnosing network latency between services. Option C is wrong because viewing raw segments of the upstream service would show the service's own processing time and subsegments, but the delay is in the downstream network call, which is captured as a subsegment of the upstream service's trace; however, the trace map is more efficient for visualizing the edge-level latency. Option D is wrong because adjusting sampling rules increases the number of traces captured but does not help identify the root cause of an existing latency issue in the network call.

90
MCQmedium

A DevOps engineer runs the command above to retrieve CPU utilization for an EC2 instance, but gets no data points. The instance is running and has basic monitoring enabled. What is the most likely reason?

A.The dimension name should be 'InstanceId' with a different case.
B.The instance has basic monitoring disabled.
C.The period of 300 seconds is less than the minimum supported period.
D.The IAM user executing the command lacks 'cloudwatch:GetMetricStatistics' permission.
AnswerD

The IAM user must have the cloudwatch:GetMetricStatistics permission to call the CloudWatch API and retrieve metric statistics. If the IAM policy attached to the user does not explicitly allow this action, the request is denied and no CPUUtilization data points are returned. Since the namespace, metric name, dimensions, and period are all valid, the most plausible reason for the lack of data is that the IAM permission is missing.

Why this answer

The command likely fails because the IAM user executing it does not have the 'cloudwatch:GetMetricStatistics' permission. Basic monitoring publishes CPUUtilization metrics every 5 minutes (300 seconds), so a period of 300 seconds is valid. The dimension name 'InstanceId' is correct as shown.

Option A is wrong because the dimension name is case-sensitive and 'InstanceId' is correct. Option B is wrong because basic monitoring is enabled by default and publishes CPUUtilization. Option C is wrong because 300 seconds equals the default 5-minute interval, which is the minimum supported period for basic monitoring.

91
Multi-Selectmedium

A DevOps engineer is designing a monitoring solution for a serverless application using AWS Lambda, Amazon API Gateway, and Amazon DynamoDB. The team needs to monitor for errors and latency. Which TWO actions should the engineer take to implement comprehensive monitoring? (Choose TWO.)

Select 2 answers
A.Enable DynamoDB Accelerator (DAX) to reduce latency.
B.Enable detailed billing metrics for cost analysis.
C.Configure CloudWatch Logs for DynamoDB.
D.Enable AWS X-Ray tracing on API Gateway and Lambda.
E.Set up CloudWatch Alarms on Lambda error count and API Gateway 5XX count.
AnswersD, E

AWS X-Ray traces requests end-to-end from API Gateway to Lambda and downstream services like DynamoDB, capturing segments, sub-segments, and annotations that reveal where latency is consumed and which component is returning errors. By enabling X-Ray on API Gateway and Lambda, the DevOps engineer can inspect trace timelines, service maps, and error rates per operation to pinpoint performance bottlenecks and failed invocations. This provides correlated, request-level observability that CloudWatch aggregate metrics alone cannot offer, making it a correct component of a comprehensive monitoring solution.

Why this answer

AWS X-Ray provides end-to-end tracing for requests as they travel through API Gateway, Lambda, and DynamoDB, enabling the team to identify latency bottlenecks and errors across the entire serverless application. This is essential for comprehensive monitoring of distributed applications, as it captures detailed timing and error data for each component.

Exam trap

The trap here is that candidates may confuse performance optimization tools (like DAX) or cost monitoring (like billing metrics) with actual monitoring solutions, or assume that DynamoDB has native CloudWatch Logs support, when in fact it only emits metrics and requires X-Ray or CloudTrail for detailed request tracing.

92
MCQmedium

A DevOps engineer is troubleshooting an application that runs on Amazon EC2 instances behind an Application Load Balancer. Users report intermittent 503 errors. CloudWatch metrics for the ALB show an increase in 'HTTPCode_ELB_5XX_Count' but the backend 'HealthyHostCount' remains stable. Which action should the engineer take to identify the root cause?

A.Increase the size of the EC2 instances to handle more requests.
B.Enable detailed CloudWatch metrics on the EC2 instances to monitor CPU and memory.
C.Enable and review the ALB access logs stored in Amazon S3 to analyze the HTTP response codes and request patterns.
D.Increase the idle timeout setting on the ALB.
AnswerC

Enabling ALB access logs and storing them in Amazon S3 records every request with fields such as elb_status_code, target_status_code, request_processing_time, target_processing_time, and response_processing_time, as well as the exact request path and user agent. By querying these logs, you can identify the specific HTTP response codes (including the 503s), observe patterns by path or client, and determine whether the timeouts occur in the load balancer or at the target. This is the definitive diagnostic step to isolate the cause of intermittent 503s.

Why this answer

ALB access logs capture detailed information about each request, including HTTP status codes, request processing times, and target response times. By analyzing these logs in Amazon S3, the engineer can identify which specific requests are resulting in 503 errors, examine if there are patterns such as high latency or target unavailability, and determine the root cause. Option A is incorrect because increasing instance size does not address the intermittent 503 errors; if the instances are healthy (HealthyHostCount stable), the issue is likely not capacity but something else like configuration or request handling.

Option B is incorrect because detailed CloudWatch metrics on instances, while useful for performance, would not directly reveal why the ALB is returning 503 errors; the backend might be healthy but returning errors or timing out. Option D is incorrect because increasing the idle timeout would only help if requests are being dropped due to idle connections, but 503 errors typically indicate that the targets are not responding or are returning errors, not idle timeouts.

93
MCQhard

A company runs an Auto Scaling group of EC2 instances that publish custom application metrics to CloudWatch using the PutMetricData API. During a traffic spike, the operations team reports that alarms based on these metrics did not trigger even though application error rates rose sharply. The metrics are published with a one-minute resolution. Which action should a DevOps engineer take to make the alarms respond reliably during spikes?

A.Change the alarm statistic to SampleCount and set the period to 60 seconds.
B.Configure the alarm to treat missing data as breaching and shorten the evaluation period.
C.Publish the custom metrics as high-resolution metrics and configure the alarm with a shorter evaluation period and M-out-of-N datapoints to alarm.
D.Increase the alarm's evaluation period to five minutes so more data points are averaged together.
AnswerC

High-resolution metrics allow one-second granularity, and combining a short evaluation period with M-out-of-N datapoints to alarm makes the alarm evaluate recent error data quickly. This directly addresses the delay and sparsity of custom metrics during rapid spikes, improving alarm responsiveness.

Why this answer

Custom metrics published at standard one-minute resolution can lag behind a fast spike, and a long evaluation window dilutes the signal. Publishing high-resolution metrics and configuring the alarm with a short evaluation period plus M-out-of-N datapoints to alarm lets the alarm react quickly to a genuine error-rate increase.

Exam trap

The trap here is responding to a missed alarm by widening the evaluation period, when that averaging actually makes the spike harder to detect.

94
MCQhard

A Lambda function is unable to write logs to CloudWatch Logs. The IAM policy attached to the function's execution role is shown above. What is the issue?

A.The resource ARN is incorrect; it should include the log stream name.
B.The region in the ARN does not match the Lambda function's region.
C.The action should be 'logs:PutLogEvents' but the resource is too restrictive.
D.The policy is missing the 'logs:CreateLogGroup' and 'logs:CreateLogStream' actions.
AnswerD

Lambda's execution role needs `logs:CreateLogGroup` and `logs:CreateLogStream` before `logs:PutLogEvents` can succeed, since the log group and stream must exist first. Without these actions, the function cannot create its destination, so writes fail regardless of the PutLogEvents permission. This satisfies the stem's requirement for the missing write-enabling actions.

Why this answer

The IAM policy attached to the Lambda execution role is missing the logs:CreateLogGroup and logs:CreateLogStream actions. Without these, Lambda cannot create the log group or log stream, and thus cannot write logs to CloudWatch Logs. The logs:PutLogEvents action alone is insufficient because the log group and stream must exist first.

Exam trap

The trap is focusing on the logs:PutLogEvents action and its resource ARN, while overlooking the prerequisite actions needed to create the log group and stream.

How to eliminate wrong answers

Option A is wrong because the resource ARN for logs:PutLogEvents typically includes the log group and log stream, but the issue is not the ARN format; it's the missing actions. Option B is wrong because the region in the ARN is likely correct; the problem is not a region mismatch. Option C is wrong because while logs:PutLogEvents is necessary, the resource being too restrictive is not the primary issue; the missing CreateLogGroup and CreateLogStream actions prevent any logging.

95
MCQmedium

A DevOps engineer needs to monitor the number of messages in an Amazon SQS queue and trigger an auto scaling action when the queue depth exceeds a threshold. Which combination of services should be used?

A.Amazon CloudWatch Logs and Amazon EC2 Auto Scaling
B.Amazon CloudWatch and Amazon EC2 Auto Scaling
C.Amazon EventBridge and Amazon EC2 Auto Scaling
D.Amazon SQS and AWS Lambda
AnswerB

Amazon CloudWatch natively ingests the `ApproximateNumberOfMessagesVisible` metric that SQS publishes every minute, making queue depth directly measurable. A CloudWatch alarm that transitions to ALARM state based on this metric can be attached to a scaling policy in EC2 Auto Scaling (for step or simple scaling), or you can use a target tracking policy with a custom metric. This is the standard, supported mechanism for scaling compute resources based on queue backlog.

Why this answer

Amazon SQS automatically publishes queue metrics (such as ApproximateNumberOfMessagesVisible) to Amazon CloudWatch. CloudWatch alarms can monitor these metrics and trigger scaling policies on an EC2 Auto Scaling group when the threshold is breached. This is the standard, native integration for queue-depth-based auto scaling.

Exam trap

DOP-C02 often tests the misconception that EventBridge can directly monitor SQS queue depth and trigger scaling, but EventBridge is for event routing, not metric-based alarms; CloudWatch is the correct service for metric monitoring and alarms.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs is for storing and analyzing log data, not for monitoring SQS metrics or triggering scaling actions. Option C is wrong because EventBridge is used for event-driven architectures and routing events, but it does not natively monitor SQS queue depth metrics or directly trigger EC2 Auto Scaling policies. Option D is wrong because SQS alone does not provide monitoring or scaling capabilities, and Lambda is a compute service that could process messages but does not directly trigger EC2 Auto Scaling actions based on queue depth.

96
MCQhard

Refer to the exhibit. A DevOps engineer runs this query to investigate a spike in errors. What is the most likely interpretation?

A.The error rate is increasing sharply in the last 15 minutes.
B.The error rate is decreasing over time.
C.The error rate is stable with no significant change.
D.The query is incorrectly filtering log streams.
AnswerA

The query results show a sharp upward spike in the error count in the most recent time bins, jumping from 1 to 12, which is a clear and significant acceleration in the error rate over the last 15 minutes. This trend is not random fluctuation; the consistent climb in the final bins indicates an emerging incident that requires immediate attention. The slope of the line in the visualization confirms the error rate is increasing, not just an isolated outlier.

Why this answer

The query counts ERROR messages per 5-minute bin for a specific log stream. The output shows a clear increasing trend from 1 to 12 errors over the last 20 minutes, indicating a recent escalation of errors.

97
MCQmedium

A company uses Amazon RDS for MySQL. The database performance has degraded, and the engineer suspects that slow queries are the cause. Which service should be used to identify and analyze the slow queries?

A.Amazon RDS Performance Insights
B.Amazon CloudWatch Metrics
C.Amazon CloudWatch Logs
D.AWS X-Ray
AnswerA

Amazon RDS Performance Insights is the correct choice because it provides a database-focused dashboard that visualizes the DB Load metric in units of Average Active Sessions (AAS), directly correlating temporal load spikes with the specific SQL statements, waits, hosts, and users responsible. This enables you to drill down into individual slow queries and diagnose performance degradation without needing to manually parse slow query logs or instrument application code.

Why this answer

Amazon RDS Performance Insights is the correct service because it provides a database-specific performance schema that visualizes database load and identifies the exact SQL queries causing performance degradation. It integrates directly with RDS for MySQL, offering a dashboard that breaks down wait events, SQL text, and host-level metrics, making it the ideal tool for analyzing slow queries.

Exam trap

The trap here is that candidates may confuse CloudWatch Logs (which can store slow query logs) with a native analysis tool, overlooking that Performance Insights provides immediate, built-in visualization and query-level analysis without requiring custom log parsing.

How to eliminate wrong answers

Option B is wrong because Amazon CloudWatch Metrics provides aggregated performance metrics like CPU utilization and IOPS, but it does not capture individual SQL query text or detailed database wait event analysis needed to identify slow queries. Option C is wrong because Amazon CloudWatch Logs can store MySQL slow query logs if configured, but it requires manual setup and does not provide built-in visualization or analysis of query performance; it is a log storage service, not an analysis tool. Option D is wrong because AWS X-Ray is designed for tracing distributed application requests and debugging microservices, not for analyzing database query performance or slow SQL statements.

98
MCQhard

A company runs a critical application on an Auto Scaling group of EC2 instances behind an Application Load Balancer (ALB). The DevOps team needs to implement a dashboard that shows real-time request latency, error rates, and the number of healthy hosts. Which AWS service should be used to create this dashboard?

A.Amazon QuickSight
B.AWS CloudTrail
C.AWS Config
D.Amazon CloudWatch Dashboards
AnswerD

Amazon CloudWatch Dashboards are the native solution for building customizable operational views that aggregate metrics from AWS services, including EC2 auto scaling groups, load balancers, and custom application metrics. They support real-time graphing with automatic refresh, can combine multiple metrics, alarms, and logs into a single pane of glass, and allow cross-account or cross-region aggregation. With built-in integration for Auto Scaling group metrics like average CPU utilization and healthy host counts, CloudWatch Dashboards are the appropriate choice for a critical application's real-time operational monitoring.

Why this answer

Amazon CloudWatch Dashboards is the correct choice because it provides real-time monitoring and visualization of metrics such as request latency, error rates, and healthy host counts directly from CloudWatch. These metrics are automatically emitted by the Application Load Balancer (e.g., TargetResponseTime, HTTPCode_ELB_5XX_Count, HealthyHostCount) and can be displayed on a customizable dashboard without additional data transformation or querying.

Exam trap

The trap here is that candidates may confuse Amazon QuickSight with CloudWatch Dashboards, assuming QuickSight is the go-to for any dashboarding need, but QuickSight is for business analytics and not designed for real-time infrastructure monitoring with sub-minute latency metrics.

How to eliminate wrong answers

Option A is wrong because Amazon QuickSight is a business intelligence service for interactive dashboards and ad-hoc analysis, not designed for real-time operational monitoring of live infrastructure metrics like ALB latency or healthy hosts. Option B is wrong because AWS CloudTrail records API activity and governance events, not real-time performance metrics or health status of EC2 instances behind a load balancer. Option C is wrong because AWS Config tracks resource configuration changes and compliance, not real-time operational metrics such as request latency or error rates.

99
Multi-Selecthard

A company runs a critical application on Amazon EKS. The operations team needs to monitor the health of the Kubernetes cluster and the applications running on it. Which THREE services can be used together to achieve comprehensive monitoring? (Choose THREE.)

Select 3 answers
A.AWS X-Ray
B.Amazon VPC Flow Logs
C.AWS CloudTrail
D.Amazon CloudWatch Container Insights
E.Amazon Managed Service for Prometheus
AnswersA, D, E

AWS X-Ray traces requests as they flow through a distributed application, providing end-to-end visibility into latencies, errors, and dependencies across microservices running on Amazon EKS. By instrumenting the application with the X-Ray SDK, you can identify the specific service or database call causing performance degradation. X-Ray's service graph and trace timelines reveal exactly where requests are spending time, making it the correct choice for troubleshooting application-level performance issues.

Why this answer

AWS X-Ray (A) is correct because it provides distributed tracing for applications running on EKS, letting the team follow requests across microservices and pinpoint latency or errors in the application layer. Amazon CloudWatch Container Insights (D) is correct because it collects and aggregates container-level metrics and logs (CPU, memory, disk, network) from EKS nodes and pods, giving cluster and workload health visibility. Amazon Managed Service for Prometheus (E) is correct because it is a Prometheus-compatible managed service that scrapes and stores Kubernetes metrics and supports alerting, complementing Container Insights for deeper, PromQL-based monitoring.

Amazon VPC Flow Logs (B) only captures IP traffic metadata at the network interface level and does not provide application or Kubernetes-level health insight. AWS CloudTrail (C) records AWS API activity for auditing and governance, not runtime health or performance monitoring of the cluster or its applications.

Exam trap

DOP-C02 often tests whether candidates confuse network/audit logging services (VPC Flow Logs, CloudTrail) with observability services that actually surface application and container health, causing them to pick the wrong 'monitoring' trio.

100
MCQhard

A company has a microservices architecture with 50 services running on Amazon ECS. The DevOps team wants to collect and analyze logs from all services centrally. They need to query logs across services and set up alerts for error patterns. Which solution is the most scalable and cost-effective?

A.Use AWS CloudTrail to capture all log events and store them in an S3 bucket for analysis
B.Deploy an Amazon Elasticsearch cluster and configure the ECS Fargate agent to send logs directly to Elasticsearch
C.Use the awslogs driver to send logs to Amazon CloudWatch Logs and use CloudWatch Logs Insights for querying and metric filters for alerts
D.Send logs to Amazon S3 and use Amazon Athena for querying, with scheduled queries for alerts
AnswerC

The awslogs driver natively ships ECS container logs to CloudWatch Logs, where Logs Insights queries across all 50 services and metric filters trigger alerts on error patterns. This satisfies centralised querying and alerting without managing extra infrastructure.

Why this answer

Using the awslogs driver to ship ECS container logs to Amazon CloudWatch Logs is the most scalable and cost-effective native solution. CloudWatch Logs Insights provides a query language for searching across log groups, and metric filters can trigger CloudWatch alarms on error patterns. This requires no cluster management, scales automatically with log volume, and integrates natively with ECS task definitions via the logConfiguration block.

Exam trap

DOP-C02 often tests whether candidates default to self-managed Elasticsearch or S3+Athena for log analytics when the native, lower-overhead CloudWatch Logs solution is the intended answer — the trap is over-engineering.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail captures API activity and management events, not application logs from ECS containers — it cannot collect microservice application output. Option B is wrong because deploying and managing a self-hosted Amazon Elasticsearch (now OpenSearch) cluster adds significant operational overhead and cost, and sending logs directly from Fargate tasks bypasses the managed ingestion path, making it less scalable and more expensive than CloudWatch Logs. Option D is wrong because sending logs to S3 and querying with Athena introduces latency (Athena is not real-time), requires schema/partition management, and lacks native alerting — scheduled queries are a poor substitute for CloudWatch metric filters and alarms.

101
MCQeasy

A DevOps engineer is troubleshooting a slow-running Lambda function. The function processes messages from an SQS queue. Which CloudWatch metric should be examined first to determine if the function is experiencing throttling?

A.Invocations
B.ConcurrentExecutions
C.Duration
D.Throttles
AnswerD

The Throttles metric is the definitive CloudWatch counter for how many invocation requests Lambda rejected because the function had no available concurrency capacity. Each time Lambda receives a request while the function or account is over its concurrency limit, it increments Throttles and returns a 429 TooManyRequestsException, prompting SDK retries or event-source retry logic. For a slow-running function, elevated Throttles explains intermittent delays caused by client-side retries, making this the correct metric to inspect when troubleshooting such issues.

Why this answer

The Throttles metric directly indicates when Lambda is rejecting invocation requests due to concurrency limits being reached. Since the question asks specifically about throttling, this is the first metric to examine to confirm whether the function is being rate-limited by AWS.

Exam trap

The trap here is that candidates may confuse high ConcurrentExecutions with throttling, but throttling is a separate metric that directly counts rejected invocations, not the number of concurrent runs.

How to eliminate wrong answers

Option A is wrong because Invocations counts total function invocations, including successful ones, and does not indicate throttling events. Option B is wrong because ConcurrentExecutions shows the number of function instances running at a given time, but it does not directly measure throttling; high concurrency can lead to throttling but the metric itself is not the throttling indicator. Option C is wrong because Duration measures how long the function runs, which can be affected by throttling indirectly but is not a direct measure of throttling events.

102
MCQhard

A DevOps engineer manages a production environment with EC2 instances behind an Application Load Balancer (ALB). The application logs show intermittent 5xx errors from the ALB. The engineer needs to identify whether the errors originate from the targets or the ALB itself. Which CloudWatch metric should be examined to differentiate between these two sources?

A.TargetResponseTime
B.UnhealthyHostCount
C.HTTPCode_Target_5XX_Count
D.RequestCount
AnswerC

HTTPCode_Target_5XX_Count is the correct choice because it represents the number of HTTP responses with a 5xx status code that were returned by registered targets (e.g., EC2 instances) to the Application Load Balancer. This metric excludes 5xx errors generated by the ALB itself, such as 502 Bad Gateway or 503 Service Unavailable from the load balancer when no healthy targets exist. By checking this metric with a Sum statistic, the engineer can directly quantify application-level server errors and distinguish them from load-balancer-level failures.

Why this answer

HTTPCode_Target_5XX_Count, is the correct metric to identify 5xx errors originating from the targets (EC2 instances). The ALB has separate metrics for target errors (HTTPCode_Target_5XX_Count) and ALB errors (HTTPCode_ELB_5XX_Count). Option A (TargetResponseTime) measures latency, not error codes.

Option B (UnhealthyHostCount) counts unhealthy targets based on health checks, not specific HTTP errors. Option D (RequestCount) is total requests, not error codes.

103
MCQeasy

A DevOps engineer is tasked with setting up a centralized logging solution for a multi-account AWS environment. Which service should be used to aggregate logs from multiple accounts?

A.Amazon S3 with cross-region replication
B.AWS CloudTrail with organization trails
C.Amazon CloudWatch Logs with cross-account subscription
D.AWS Config with aggregated compliance rules
AnswerC

Amazon CloudWatch Logs cross-account subscriptions let you forward log events from multiple source accounts to a central destination (via a Kinesis Data Stream or Data Firehose) using subscription filters, enabling near-real-time aggregation of application logs. This is the correct pattern because it directly receives application log streams from existing CloudWatch Logs agents and routes them to a central account for storage and analysis.

Why this answer

Amazon CloudWatch Logs can aggregate logs across accounts using cross-account subscriptions with a central destination (e.g., Kinesis or Lambda). Option C is correct. Option A is incorrect because S3 is a storage service, not for real-time aggregation.

Option B is incorrect as CloudTrail is for API activity, not application logs. Option D is incorrect because AWS Config is for configuration compliance.

104
MCQhard

A DevOps engineer is configuring CloudWatch Logs for a Lambda function that processes streaming data from Kinesis. The function sometimes fails due to memory exhaustion. The engineer wants to ensure that logs from the function are shipped to CloudWatch Logs even when the function fails. Which configuration should be used?

A.Configure a Kinesis Agent on the Lambda execution environment to stream logs to CloudWatch Logs
B.Install the CloudWatch Logs agent on the Lambda function to continuously send logs
C.Enable detailed CloudWatch metrics for the Lambda function
D.Ensure the Lambda function writes logs to stdout or stderr; CloudWatch Logs will automatically capture them
AnswerD

Correct. The Lambda runtime (for both the AWS-provided runtimes and custom runtimes that use the Runtime API) intercepts all output written to stdout and stderr and automatically streams it to CloudWatch Logs under the log group /aws/lambda/<function-name>. This happens regardless of whether the function completes successfully or crashes, so logging to stdout/stderr is the standard, documented way to generate logs. Each invocation gets a unique log stream, and the exact log line format can include request IDs if you also log the Lambda context object, but the capture itself requires no extra configuration beyond the Lambda service execution role's CloudWatch Logs permissions.

Why this answer

Lambda functions automatically send all output written to stdout (via print or console.log) and stderr to CloudWatch Logs, regardless of whether the function succeeds or fails. This is a built-in behavior of the Lambda runtime, so no additional agents or configuration are needed to capture logs from a failed invocation due to memory exhaustion.

Exam trap

The trap here is that candidates may overthink the solution and assume a separate agent or service is required for log shipping in failure scenarios, when in fact Lambda’s native stdout/stderr capture works automatically and reliably even on invocation failure.

How to eliminate wrong answers

Option A is wrong because a Kinesis Agent is designed to run on EC2 instances or on-premises servers to send data to Kinesis, not to stream logs from a Lambda execution environment; Lambda does not support installing or running external agents. Option B is wrong because the CloudWatch Logs agent is intended for EC2 instances or on-premises servers, and cannot be installed inside a Lambda function’s ephemeral execution environment. Option C is wrong because enabling detailed CloudWatch metrics provides performance metrics (e.g., duration, invocations, errors) but does not capture or ship log output from the function.

105
MCQmedium

A DevOps engineer is troubleshooting a slow web application. The application runs on EC2 instances behind an ALB. The engineer notices that the ALB's TargetResponseTime metric shows high p99 values, but the CPU and memory on the EC2 instances are well below thresholds. What is the most likely cause?

A.The Auto Scaling group has too many instances, causing increased network overhead
B.The ALB is routing requests to instances in different Availability Zones, increasing latency
C.The application is waiting on a slow database query or external API call
D.The ALB idle timeout is set too low, causing connections to be dropped
AnswerC

Synchronous dependencies like database queries and external API calls are the classic cause of elevated application latency; the end-user response time becomes the sum of all blocking calls in the request path, and a single slow query (e.g., missing index, lock contention, or throttled API) holds up the entire page. This pattern usually shows as high response times with low CPU and network utilization on the application instances, because threads are parked waiting for the downstream I/O to complete. In an Auto Scaling environment, simply adding instances won't help while the database or API service remains the bottleneck.

Why this answer

High p99 TargetResponseTime on the ALB with low CPU and memory on the EC2 instances indicates that the bottleneck is not compute capacity but rather a dependency external to the application server. The application is likely waiting on a slow database query or external API call, which increases response time without consuming significant local CPU or memory. This is a classic symptom of an I/O-bound or network-bound dependency.

Exam trap

The trap here is that candidates often assume high response times must be caused by compute saturation (CPU/memory) or network issues, but the question deliberately shows low resource utilization to force you to consider external dependencies as the root cause.

How to eliminate wrong answers

Option A is wrong because having too many instances in the Auto Scaling group would reduce per-instance load and likely decrease response times, not increase them; network overhead from more instances is negligible compared to the ALB's connection management. Option B is wrong because ALB inherently routes requests to instances across Availability Zones with minimal latency overhead (typically <1 ms), and cross-AZ data transfer costs are not a significant factor in p99 response time. Option D is wrong because a low ALB idle timeout would cause connections to be dropped prematurely, resulting in client-side errors (e.g., 504 Gateway Timeout) rather than consistently high p99 response times; the metric would show timeouts, not slow completions.

106
Multi-Selecthard

A company runs a critical application on Amazon ECS with Fargate. The application emits structured logs in JSON format. The DevOps team wants to monitor for specific error codes and receive near-real-time alerts. The team also needs to retain logs for 5 years for compliance. Which TWO steps should the team implement?

Select 2 answers
A.Create a CloudWatch Logs metric filter to count occurrences of specific error codes and create an alarm
B.Use Amazon Kinesis Data Analytics to analyze logs in real-time and send alerts
C.Enable AWS CloudTrail to log the application's API calls
D.Stream logs to Amazon S3 via Amazon Kinesis Data Firehose and use S3 event notifications to trigger alerts
E.Configure a CloudWatch Logs retention policy to keep logs for 5 years
AnswersA, E

A CloudWatch Logs metric filter continuously scans new log events as they arrive in the log group and increments a custom metric whenever a pattern such as 'ERROR' or a specific error code appears. This metric can then drive a CloudWatch alarm with an associated SNS topic, providing a near-real-time notification without building any separate ingestion pipeline. Metric filters are the native, low-latency mechanism for monitoring textual patterns in logs.

Why this answer

CloudWatch Logs metric filters can parse JSON-structured logs to count occurrences of specific error codes, and you can create a CloudWatch alarm on that metric to trigger near-real-time notifications via SNS. This is a native, low-latency solution for monitoring specific patterns in ECS Fargate logs without additional infrastructure.

Exam trap

The trap here is that candidates often confuse CloudTrail (which logs AWS API calls) with application-level logging, or they over-engineer the solution with Kinesis Data Analytics or Firehose when CloudWatch native features (metric filters and retention policies) are sufficient and more cost-effective for this use case.

107
MCQhard

A company is using Amazon CloudWatch Synthetics canaries to monitor its web application endpoints. The canaries are deployed in multiple AWS regions. The team wants to aggregate the canary results into a single dashboard in the US East (N. Virginia) region. What is the MOST efficient way to achieve this?

A.Replicate the canaries to US East (N. Virginia) and run them from there.
B.Create a cross-region CloudWatch dashboard and add metrics from each region using metric math.
C.Set up a Lambda function in each region to push canary results to a central S3 bucket, then create a dashboard from S3.
D.Create a CloudWatch Logs Insights query across all regions and visualize results.
AnswerB

CloudWatch dashboards are not region-bound artifacts: each widget can explicitly specify a different source region, and metric math can combine those cross-region metrics within a single expression, for example summing SuccessPercent or averaging Duration across all canary regions. This natively aggregates the existing Synthetics metrics without duplicating canaries, running Lambda functions, or parsing logs. It is the intended, low-operational-overhead mechanism for a consolidated cross-region view and requires no custom infrastructure.

Why this answer

CloudWatch cross-region dashboards allow you to aggregate metrics from multiple regions into a single dashboard without data movement. By using metric math, you can reference metric IDs from different regions directly in the dashboard widget, enabling real-time aggregation of Synthetics canary success/failure rates and latency metrics from all regions into a unified view in US East (N. Virginia).

This approach avoids unnecessary data replication, reduces latency, and minimizes operational overhead.

Exam trap

The trap here is that candidates may assume cross-region aggregation requires data movement (e.g., to S3 or Lambda) or that CloudWatch dashboards are region-scoped, but AWS actually supports cross-region dashboards natively, making option B the most efficient and direct solution.

How to eliminate wrong answers

Option A is wrong because replicating canaries to US East (N. Virginia) would only monitor endpoints from that single region, losing the geographic distribution and failing to aggregate results from the original regions. Option C is wrong because pushing canary results to an S3 bucket and then creating a dashboard from S3 introduces unnecessary complexity, latency, and potential data staleness; CloudWatch Synthetics already stores metrics and logs in CloudWatch, so a cross-region dashboard is more direct and efficient.

Option D is wrong because CloudWatch Logs Insights queries cannot span multiple regions; they are scoped to a single region and log group, making cross-region aggregation impossible without additional tooling.

108
MCQmedium

A company is using AWS CloudFormation to deploy infrastructure. They want to receive notifications when a stack operation fails, including the specific resource that caused the failure. Which approach should they use?

A.Create a CloudWatch alarm on the 'StackFailure' metric.
B.Configure an SNS topic as a notification option in the CloudFormation stack and subscribe to receive stack events.
C.Create an AWS Lambda function that polls the CloudFormation DescribeStackEvents API every minute and sends an email on failure.
D.Enable AWS CloudTrail to log CloudFormation API calls and configure an SNS notification on the trail.
AnswerB

Configuring an SNS topic as a stack notification option delivers every CloudFormation stack event, including ResourceStatusReason for the specific resource that failed, satisfying the requirement to identify the failing resource. Subscribing an email or Lambda endpoint to that topic then surfaces the failure notification automatically.

Why this answer

CloudFormation allows you to specify an SNS topic ARN as a notification option when creating or updating a stack. When a stack operation fails, CloudFormation publishes a notification to that SNS topic, and the notification includes the logical resource ID and the status reason for the failure. This provides real-time, event-driven notifications without requiring polling or additional services.

Exam trap

The trap here is that candidates may confuse CloudWatch metrics or CloudTrail with CloudFormation's native notification capability, assuming that failure events are exposed as metrics or logs rather than through SNS topic subscriptions.

How to eliminate wrong answers

Option A is wrong because CloudFormation does not emit a 'StackFailure' metric to CloudWatch; CloudFormation publishes stack events to SNS topics, not CloudWatch metrics. Option C is wrong because polling the DescribeStackEvents API every minute introduces latency, unnecessary cost, and complexity compared to the native SNS notification mechanism; it also violates the principle of event-driven architecture. Option D is wrong because AWS CloudTrail logs API calls for auditing, but it does not provide real-time notifications on stack operation failures; configuring SNS on a trail only delivers log file delivery notifications, not stack failure events.

109
MCQhard

Refer to the exhibit. A DevOps engineer runs the AWS CLI command to get the average TargetResponseTime for an ALB over a 1-hour period. The output shows only three datapoints. What is the most likely reason?

A.The ALB did not receive any requests during most of the 5-minute periods.
B.The metric TargetResponseTime is not available for Application Load Balancers.
C.The command is missing the 'Statistics' parameter with 'Average'.
D.The period of 300 seconds is too large; a smaller period should be used.
AnswerA

The ALB emits TargetResponseTime only when it actually processes requests. For any 5-minute interval in which no request was routed through the load balancer, CloudWatch receives no samples and therefore publishes no datapoint. The get-metric-statistics output contains a datapoint only for each period with at least one request, so periods with no traffic are simply missing from the result rather than appearing as zero. This exactly explains why fewer datapoints than expected are returned.

Why this answer

The command uses a period of 300 seconds (5 minutes), so over a 1-hour period we expect 12 data points if the metric is consistently reported. However, TargetResponseTime is only emitted when the ALB receives at least one request in that interval. The output shows only three data points, indicating that for the majority of the 5-minute periods, the ALB received no requests, so no metric data was published.

Option A correctly identifies this. Option B is incorrect because TargetResponseTime is a valid metric for Application Load Balancers. Option C is incorrect because the command appears to include the necessary parameters to retrieve the average.

Option D is incorrect because a 300-second period is standard; using a smaller period would not produce more data points if there are no requests.

110
MCQmedium

Refer to the exhibit. A DevOps engineer runs the CloudWatch Logs Insights query and sees a spike in errors at 12:00. Which action would best help identify the root cause?

A.Add a filter for @logStream to see which log stream has the most errors.
B.Increase the limit to 100 to see more results.
C.Query the logs around 12:00 without aggregation to see the actual error messages.
D.Change the bin time to 1m to get more granular data.
AnswerC

Removing the count() and bin() aggregation and running a query that returns the raw @timestamp and @message fields around 12:00 will display the actual error strings, stack traces, and exception types that caused the spike. This is the definitive method for root-cause analysis because it lets you see exactly what the application logged at the moment of failure, such as a TimeoutException, a 500 Internal Server Error, or a specific failed database call. You can then filter further using regular expressions or parse commands to isolate the common pattern.

Why this answer

The query shows a sudden spike at 12:00. To identify the root cause, the engineer should look at the error messages themselves around that time. Filtering by @timestamp and @message to see the actual error messages will help identify the type of error.

Adding a filter for a specific error pattern or grouping by error message would also help.

111
MCQmedium

A company uses AWS CloudTrail to log API activity across multiple accounts in AWS Organizations. The security team wants to receive near-real-time notifications for specific high-risk API calls, such as IAM policy changes or S3 bucket policy modifications. What is the MOST efficient and scalable solution?

A.Deliver CloudTrail logs to an S3 bucket, enable S3 Event Notifications to trigger a Lambda function that filters and publishes to SNS.
B.Create a CloudWatch Events rule that matches the specific API calls and publishes to an SNS topic.
C.Use CloudWatch Logs Insights to query CloudTrail logs and set up a metric filter with an alarm.
D.Enable AWS Config rules to detect changes and trigger an SNS notification.
AnswerA

CloudTrail writes log files to S3 as compressed JSON objects, and S3 Event Notifications fire as soon as each object is created, triggering a Lambda function in near-real-time. The Lambda can decompress the log file, parse individual events, and apply precise filters—such as specific event names, source IPs, or IAM principals—before publishing only the high-risk actions to SNS. This serverless, event-driven pattern scales automatically with account activity and avoids the cost and noise of notifying on every raw CloudTrail event, making it both efficient and cost-effective for high-volume accounts.

Why this answer

It uses S3 Event Notifications to trigger a Lambda function in near-real-time when CloudTrail logs are delivered to S3. The Lambda function can filter for specific high-risk API calls (e.g., IAM policy changes, S3 bucket policy modifications) and publish only relevant events to an SNS topic, providing a scalable and cost-effective solution that avoids polling or complex querying.

Exam trap

The trap here is that candidates often assume CloudWatch Events (EventBridge) is the default choice for real-time CloudTrail monitoring, but they overlook that S3 Event Notifications with Lambda provide a more direct and scalable path for filtering high-volume log data without the overhead of streaming all logs to CloudWatch Logs.

How to eliminate wrong answers

Option B is wrong because CloudWatch Events (now Amazon EventBridge) can match specific API calls from CloudTrail, but it does not support near-real-time notifications for all CloudTrail log entries; it relies on CloudTrail delivering logs to CloudWatch Logs, which can introduce latency and is less efficient for high-volume filtering. Option C is wrong because CloudWatch Logs Insights is a query tool for ad-hoc analysis, not a real-time notification mechanism; metric filters and alarms can trigger notifications but require logs to be streamed to CloudWatch Logs, adding complexity and potential delay. Option D is wrong because AWS Config rules detect configuration changes (e.g., resource modifications) but are not designed for real-time API-level monitoring; they evaluate resources periodically or on configuration changes, which may not capture all high-risk API calls and introduces evaluation delays.

112
MCQmedium

A company uses AWS CloudFormation to deploy infrastructure. The DevOps team needs to receive notifications when stack creation fails. Which approach should be used to automate this monitoring?

A.Create a CloudWatch Events rule that matches CloudFormation 'CREATE_FAILED' stack events and targets an SNS topic.
B.Use AWS Config rules to detect failed stack creations.
C.Enable CloudTrail and create a metric filter for 'CreateStack' API calls.
D.Stream CloudFormation logs to CloudWatch Logs and create a metric filter for 'CREATE_FAILED'.
AnswerA

CloudFormation emits stack status change events to CloudWatch Events (Amazon EventBridge) whenever a stack transitions to a terminal state. A rule can filter for detail-type 'CloudFormation Stack Status Change' and the 'status-detail' of 'CREATE_FAILED', then route the event to an SNS topic. This is the native, event-driven mechanism for receiving notifications about stack creation failures, and it includes the stack name and status in the event payload.

Why this answer

CloudWatch Events (now Amazon EventBridge) can match CloudFormation stack events with the detail-type 'CloudFormation Stack Status Change' and filter for CREATE_FAILED, then target an SNS topic for notification. This is the native, event-driven approach for automating failure alerts without polling or log parsing. It provides near-real-time notification with minimal configuration.

Exam trap

DOP-C02 often tests the misconception that CloudTrail or CloudWatch Logs can directly filter CloudFormation failure statuses; candidates must recognize EventBridge as the native event source for stack status changes.

How to eliminate wrong answers

Option B is wrong because AWS Config rules evaluate resource compliance state, not CloudFormation stack event statuses, and cannot directly detect CREATE_FAILED events. Option C is wrong because CloudTrail logs API calls like CreateStack but does not emit a 'CREATE_FAILED' event; a metric filter on CreateStack would only count API invocations, not failures. Option D is wrong because CloudFormation does not stream stack events to CloudWatch Logs by default; you would need to poll DescribeStackEvents or use EventBridge, making this approach ineffective.

113
MCQhard

A DevOps engineer manages a multi-account AWS Organization. The security team requires that all CloudWatch Logs log groups in every account retain data for at least 400 days and that no developer can shorten that retention. Which combination of actions should the engineer take?

A.Enable CloudWatch Logs data protection and configure a log group policy that blocks retention changes.
B.Set retention on each log group to 400 days using a script, and rely on IAM policies in each account to deny logs:PutRetentionPolicy.
C.Create an AWS Lambda function in each account that runs hourly, detects retention changes, and restores 400 days, and subscribe the function to an Amazon SNS topic.
D.Apply a service control policy in AWS Organizations that denies logs:PutRetentionPolicy and logs:DeleteRetentionPolicy unless the request comes from a designated governance role, and enforce a baseline retention with AWS Config or a CloudFormation StackSet.
AnswerD

SCPs set the maximum permissions for member accounts, so denying retention changes except for a governance role prevents developers from shortening retention. Combined with a StackSet or Config rule that establishes the 400-day baseline, this gives centralized, drift-resistant enforcement across all accounts, including future ones, which matches the requirement.

Why this answer

Service control policies provide preventive guardrails across an AWS Organization, so denying retention policy changes except for a governance role stops developers from shortening retention. A StackSet or AWS Config rule then enforces the 400-day baseline consistently. Reactive scripts, data protection, and per-account IAM policies do not deliver centralized prevention.

Exam trap

The trap here is confusing CloudWatch Logs data protection policies, which mask sensitive data, with retention controls that govern how long log events are stored.

114
MCQhard

A company has a critical application running on Amazon EC2 instances behind an Application Load Balancer. The application is experiencing intermittent latency spikes. The DevOps team has enabled detailed monitoring on the EC2 instances and is using CloudWatch metrics. They notice that CPU utilization and network traffic are normal during the spikes. Which additional diagnostic step should the team take to identify the root cause?

A.Instrument the application with AWS X-Ray to trace requests and identify bottlenecks.
B.Use CloudWatch Container Insights to monitor the performance of the EC2 instances.
C.Enable CloudWatch Synthetics to create canaries that monitor the application endpoints.
D.Run an AWS Trusted Advisor check to identify performance-related recommendations.
AnswerA

Instrumenting the application with AWS X-Ray creates an end-to-end trace for each user request as it traverses your application, generating a service map and detailed segments/subsegments for each downstream call (e.g., web APIs, SQL queries, external HTTP calls). This lets you pinpoint which specific operation contributes the most latency on EC2, including CPU-bound code vs. I/O wait, and correlate trace data with host metrics. Unlike other options, X-Ray provides the distributed tracing needed to identify and reason about internal request bottlenecks in a complex application.

Why this answer

AWS X-Ray provides end-to-end tracing of requests as they travel through the application. Since CPU and network metrics appear normal, the intermittent latency is likely caused by application-level bottlenecks such as slow database queries, external API calls, or inefficient code paths. X-Ray can trace each request end-to-end and identify the specific service or component introducing delay.

Container Insights (B) is designed for monitoring containerized workloads (Amazon ECS/EKS), not EC2 instances directly. CloudWatch Synthetics (C) creates canaries that monitor external endpoint availability and response times, but it does not trace internal request paths. AWS Trusted Advisor (D) offers general best-practice recommendations but is not a diagnostic tool for real-time latency issues.

115
MCQeasy

A DevOps engineer needs to centrally collect and analyze logs from multiple AWS accounts and on-premises servers. Which AWS service should be used to aggregate logs in a single dashboard?

A.Amazon Athena.
B.Amazon S3.
C.Amazon CloudWatch Logs.
D.Amazon Kinesis Data Firehose.
AnswerC

Amazon CloudWatch Logs is the correct choice because it can centrally aggregate logs from multiple AWS accounts, Regions, and on-premises sources via the CloudWatch Logs agent, Kinesis Data Firehose, or subscription filters. It provides native log analytics with Logs Insights, metric filters to create custom metrics, and CloudWatch Dashboards to visualize log-derived data in real time. This directly satisfies the need for centralized collection, analysis, and dashboard visualization without requiring additional services.

Why this answer

Amazon CloudWatch Logs is the correct service for centrally collecting, storing, and analyzing logs from multiple AWS accounts and on-premises servers, with the ability to visualize them in a single dashboard via CloudWatch Logs Insights and CloudWatch dashboards. It supports cross-account log aggregation through subscription filters and centralized log groups.

Exam trap

The trap is that candidates confuse log storage (S3), log query (Athena), and log streaming (Firehose) with the service that actually provides centralized log collection and dashboards — CloudWatch Logs.

How to eliminate wrong answers

Option A is wrong because Amazon Athena is a query service for data in S3 — it can analyze logs but does not natively collect or aggregate them in real time, nor does it provide a dashboard. Option B is wrong because Amazon S3 is object storage; while logs can be archived there, S3 alone does not provide analysis or dashboards. Option D is wrong because Kinesis Data Firehose is a delivery service that streams data to destinations like S3, Redshift, or Splunk — it does not itself provide a dashboard or log analysis.

116
MCQmedium

A DevOps engineer needs to audit changes to IAM policies over the past 90 days. The engineer wants to see who made the change, what the change was, and when it occurred. Which AWS tool should be used?

A.AWS Config
B.Amazon CloudWatch Logs
C.AWS CloudTrail
D.IAM Access Analyzer
AnswerC

AWS CloudTrail is the correct service for auditing IAM policy changes because it records every API call as an event, including the identity of the requesting principal, the timestamp, source IP, request parameters, and response elements. You can view these events directly in the CloudTrail event history or create a trail for long-term storage in S3 and analysis via CloudWatch Logs or Athena, making it the authoritative audit source.

Why this answer

AWS CloudTrail is the correct choice because it records all API calls made to the AWS environment, including IAM policy changes, and stores them as events with details such as the identity of the caller (IAM user or role), the time of the request, and the request parameters. By querying CloudTrail logs over the past 90 days, the DevOps engineer can audit who made the change, what the change was (e.g., the specific IAM policy document modification), and when it occurred.

Exam trap

The trap here is that candidates often confuse AWS Config's ability to track configuration changes with CloudTrail's ability to provide a detailed audit trail of API calls, leading them to choose AWS Config for auditing who made a change, when in fact Config only shows the state change, not the identity of the actor.

How to eliminate wrong answers

Option A is wrong because AWS Config is a configuration management and compliance service that tracks resource configuration changes and evaluates them against rules, but it does not record who made the change or the exact API call details; it focuses on the state of resources, not the audit trail of actions. Option B is wrong because Amazon CloudWatch Logs is used to monitor, store, and access log files from AWS resources and applications, but it does not natively capture IAM API calls; it would require custom logging or integration with CloudTrail to obtain such data. Option D is wrong because IAM Access Analyzer is designed to identify resources shared with external entities and analyze access policies for unintended public or cross-account access, not to provide a historical audit trail of who made changes to IAM policies.

117
MCQhard

A company is using AWS Lambda to process streaming data from Amazon Kinesis. The processing rate is slower than expected, and the engineer needs to monitor the number of records that are failing processing. Which metric should be used to create a CloudWatch alarm?

A.Invocations
B.IteratorAge
C.Errors
D.Throttles
AnswerB

IteratorAge is the correct metric to monitor for Kinesis-triggered Lambda because it directly measures the age of the oldest unprocessed record, reported in milliseconds. When the Lambda consumer can't keep up with the shard's data rate, the iterator age grows, indicating that stream records are sitting unprocessed for longer — a clear sign of a processing bottleneck or backlog. A sustained increase in IteratorAge typically drives alarms for scaling out the Lambda function or increasing the number of shards, making it the definitive indicator of stream processing lag.

Why this answer

The IteratorAge metric measures the age of the last record in the Lambda function's iterator, indicating how far behind real-time the processing is. A high or increasing IteratorAge suggests that records are being retried or stuck due to processing failures, making it the correct metric to monitor for records failing processing in a Kinesis-triggered Lambda.

Exam trap

The trap here is that candidates confuse 'Errors' (Lambda function exceptions) with 'record processing failures' in a Kinesis stream, not realizing that Kinesis retries failed batches internally, so the Lambda may not emit an error metric for each failed record.

How to eliminate wrong answers

Option A (Invocations) is wrong because it counts the total number of function invocations, not failures; a high invocation count could indicate success or failure, but it does not isolate failing records. Option C (Errors) is wrong because it tracks Lambda function errors (e.g., exceptions in code), but Kinesis stream processing failures often result in retries and do not always surface as Lambda errors if the function returns a success after a partial failure. Option D (Throttles) is wrong because it measures when Lambda concurrency limits are exceeded, which is unrelated to record processing failures; throttling would cause slower processing but not directly indicate failed records.

118
MCQmedium

A company uses Amazon CloudWatch Logs to store application logs. A DevOps engineer needs to create a real-time dashboard that displays the count of ERROR-level log entries across all instances. Which approach is the MOST efficient and cost-effective?

A.Create a CloudWatch Logs metric filter for each log group to count ERROR entries, and then create a CloudWatch dashboard
B.Use CloudWatch Logs Insights to run a query that counts ERROR entries across all log groups and add the query to a CloudWatch dashboard
C.Export logs to Amazon S3 and use Amazon Athena to query and visualize in Amazon QuickSight
D.Create a Kinesis Data Firehose delivery stream to stream logs to Amazon OpenSearch Service and build a dashboard in OpenSearch Dashboards
AnswerB

CloudWatch Logs Insights runs a query like `fields @timestamp, @logGroup | filter @message like /ERROR/ | stats count(*) by @logGroup` across all specified log groups in the selected time range, and the exact same query can be added as a dashboard widget using the 'Add to dashboard' option. This gives real-time results without any per-group configuration, and the dashboard automatically reflects the current set of log groups as long as they are included in the query scope. It is the intended native mechanism for interactive multi-group log analysis, making it the correct choice for this scenario.

Why this answer

CloudWatch Logs Insights allows you to run a query across all log groups in real time using a single query, and you can add that query directly to a CloudWatch dashboard. This approach is both efficient (no need to create per-log-group metric filters) and cost-effective (you pay only for the data scanned by the query, not for ongoing metric filter evaluation).

Exam trap

The trap here is that candidates often assume metric filters (Option A) are the only native way to get counts into a dashboard, overlooking that CloudWatch Logs Insights queries can be embedded directly into dashboards for real-time, cross-log-group analysis without the overhead of per-group filters.

How to eliminate wrong answers

Option A is wrong because creating a metric filter for each log group incurs ongoing costs for each filter evaluation, and managing filters across many log groups is inefficient compared to a single cross-group query. Option C is wrong because exporting logs to S3 and using Athena/QuickSight introduces latency (logs are not real-time) and additional costs for S3 storage, Athena queries, and QuickSight subscriptions, making it less efficient and more expensive for a real-time dashboard. Option D is wrong because streaming logs to OpenSearch Service via Kinesis Data Firehose adds complexity, latency, and cost for the delivery stream, OpenSearch cluster, and dashboard, which is overkill for a simple count of ERROR entries.

119
Multi-Selectmedium

A company is using Amazon CloudWatch Logs to store application logs. The security team requires that logs are encrypted at rest using a customer-managed AWS KMS key. Which TWO steps are necessary to achieve this?

Select 2 answers
A.Use the CloudWatch Logs console or API to associate the KMS key with the log group
B.Enable default encryption for CloudWatch Logs in the AWS account settings
C.Update the log group's resource policy to reference the KMS key
D.Associate the KMS key with each log stream individually
E.Create a customer-managed KMS key with appropriate key policy that allows CloudWatch Logs to use the key
AnswersA, E

The correct action is to explicitly associate your customer-managed AWS KMS key with the CloudWatch Logs log group, either through the console or by calling the `AssociateKmsKey` API. This association defines the encryption boundary at the log group level and applies to all existing and future log streams within that group. The key-policy prerequisites must already be in place, but the actual enabling step is this association.

Why this answer

Option E is correct because CloudWatch Logs encryption at rest with a customer-managed key requires first creating a customer-managed KMS key in KMS and attaching a key policy that grants the CloudWatch Logs service principal (logs.<region>.amazonaws.com) the necessary kms:Encrypt, kms:Decrypt, kms:ReEncrypt*, kms:GenerateDataKey*, and kms:Describe* permissions. Option A is correct because, once the key exists, you must explicitly associate it with the specific log group using the CloudWatch Logs console, the AWS CLI (aws logs associate-kms-key --log-group-name <name> --kms-key-id <key-arn>), or the AssociateKmsKey API; encryption is applied at the log group level. Option B is not correct because CloudWatch Logs has no account-level 'default encryption' toggle; encryption must be set per log group.

Option C is not correct because a log group resource policy controls access to log data, not KMS encryption, and does not enable encryption at rest. Option D is not correct because KMS keys are associated with log groups, not individual log streams, so per-stream association is neither possible nor required.

Exam trap

DOP-C02 often tests the misconception that a log group resource policy or an account-wide setting can enable KMS encryption, when in fact you need both a properly permissioned customer-managed key AND a per-log-group association.

120
MCQeasy

A company wants to monitor the number of messages that are published to an Amazon SNS topic. Which CloudWatch metric should be used?

A.SMSMonthToDateSpentUSD
B.PublishSize
C.NumberOfNotificationsDelivered
D.NumberOfMessagesPublished
AnswerD

NumberOfMessagesPublished is the correct Amazon SNS CloudWatch metric for monitoring message volume. It increments for every successful publish operation to an SNS topic, including individual messages sent via the Publish API and each message within a PublishBatch request. This metric provides the exact count of messages accepted by the topic, making it the appropriate choice for tracking publication activity.

Why this answer

The CloudWatch metric 'NumberOfMessagesPublished' tracks the number of messages published to an SNS topic. Option A (SMSMonthToDateSpentUSD) tracks monthly SMS spend, not messages published. Option B (PublishSize) is not a standard SNS metric; the relevant metric is 'PublishSize' for message size, but it's not for counting messages.

Option C (NumberOfNotificationsDelivered) counts messages delivered to subscribers, not published.

121
MCQeasy

A DevOps engineer needs to monitor the memory utilization of an Amazon EC2 instance running a critical application. Which AWS service should be used to collect and track this metric?

A.AWS CloudTrail
B.AWS X-Ray
C.AWS Config
D.Amazon CloudWatch
AnswerD

Amazon CloudWatch is the native monitoring service for AWS, collecting metrics, logs, and events from AWS resources and applications. The CloudWatch Agent, installed on EC2 instances or on-premises servers, can collect memory utilization metrics (such as mem_used_percent) and publish them as custom metrics to CloudWatch, enabling alarms and dashboards. Without the agent, standard EC2 metrics include CPU, disk, and network but not memory, so using CloudWatch with the agent is the appropriate solution.

Why this answer

Amazon CloudWatch is the correct service for monitoring memory utilization of EC2 instances. CloudWatch can collect custom metrics like memory utilization via the CloudWatch Agent. AWS CloudTrail (A) records API calls, not memory metrics.

AWS X-Ray (B) traces application requests, not system metrics. AWS Config (C) records resource configuration changes, not utilization metrics.

122
Multi-Selecthard

A company uses Amazon CloudWatch Logs to centralize logs from multiple EC2 instances running a web application. The DevOps team needs to create a metric filter that parses logs for HTTP status codes (e.g., 4xx and 5xx) and increment a metric. Additionally, they need to create a CloudWatch alarm on the error count. Which of the following are required to achieve this? (Select TWO.)

Select 2 answers
A.Create an IAM role that allows CloudWatch Logs to read the log data and publish metrics.
B.Define the metric filter pattern to match HTTP status codes in the log entries.
C.Create a metric filter in CloudWatch Logs on the log group that contains the application logs.
D.Configure a subscription filter to forward the logs to a Lambda function that creates the metric.
E.Install the CloudWatch Agent on the EC2 instances to send the logs.
AnswersB, C

The metric filter pattern is the core of the extraction process: it defines exactly which log events should be matched and how the metric value is derived, such as counting occurrences of a 4xx or 5xx status code. For example, a pattern like "[*, _, _, status_code]" or a JSON pattern like "{ $.status_code = 4* }" is needed to parse the relevant field correctly from the log entry. If the pattern is not defined, CloudWatch Logs has no way to know which log lines correspond to an HTTP status code or what value to emit for the CloudWatch metric.

Why this answer

The correct answers are B and C. A metric filter must be defined on a log group (option C) to extract metrics, and the filter pattern must match the HTTP status codes in the log entries (option B). Option A is not required because CloudWatch Logs can publish metrics without an additional IAM role; the log group already has sufficient permissions.

Option D is incorrect because subscription filters are for streaming logs to other destinations, not for creating metrics. Option E is not required because the CloudWatch Logs agent or the unified CloudWatch agent can send logs, but the question does not specify which agent; the default agent suffices.

123
MCQhard

A company has a production Amazon EKS cluster with multiple node groups. The DevOps team notices that some pods are frequently restarting due to OOMKilled errors, but the cluster-level metrics (CPU, memory) appear normal. Which CloudWatch Container Insights metric should be analyzed to identify the specific node or pod causing the issue?

A.node_memory_utilization.
B.number_of_running_pods.
C.pod_memory_utilization.
D.pod_cpu_utilization.
AnswerC

Pod memory utilization directly shows memory usage per pod compared to its configured limit, which is the exact condition that triggers the kernel's out-of-memory killer. When a container's working set (container_memory_working_set_bytes) exceeds its memory limit, the OOM killer terminates it with exit code 137, so plotting this metric against the limit surfaces the specific pod(s) at fault and helps correlate with OOMKilled events in kubectl get pods.

Why this answer

OOMKilled errors are caused by individual containers exceeding their memory limits, which is a pod-level (container-level) condition. CloudWatch Container Insights exposes 'pod_memory_utilization' (and 'container_memory_utilization') per pod, allowing you to pinpoint which pod is hitting its memory limit even when node-level memory looks normal. This is the metric that directly correlates with OOMKilled restarts.

Exam trap

DOP-C02 often tests the node-vs-pod metric distinction by making cluster-level metrics look normal, so candidates who pick node_memory_utilization miss that OOMKilled is a per-container cgroup limit event, not a node exhaustion event.

How to eliminate wrong answers

Option A is wrong because node_memory_utilization shows aggregate memory usage per node — if the node has spare memory, this metric looks normal even though a specific pod is being OOMKilled due to its own limit. Option B is wrong because number_of_running_pods is a count metric that tells you how many pods are running, not how much memory any pod is consuming, so it cannot identify the offending pod. Option D is wrong because pod_cpu_utilization measures CPU, not memory — OOMKilled is a memory-limit event, so CPU metrics are irrelevant to diagnosing it.

124
MCQeasy

A DevOps engineer needs to centralize logs from multiple AWS accounts into a single CloudWatch Logs account. Which feature should be used?

A.CloudWatch Logs Insights
B.AWS CloudTrail
C.Amazon Kinesis Data Firehose
D.CloudWatch Logs cross-account subscription
AnswerD

CloudWatch Logs cross-account subscription allows a central (destination) account to create a destination policy and use subscription filters in source accounts to forward log events in real time to a central log group or a Kinesis Data Streams/Firehose pipeline. This feature natively supports aggregating logs from multiple AWS accounts into a single location, providing centralized monitoring and analysis. It directly meets the requirement to centralize logs from multiple AWS accounts, making it the correct choice.

Why this answer

CloudWatch Logs cross-account subscription (Option D) allows you to stream log data from multiple source AWS accounts to a single destination account's CloudWatch Logs log group. This is the native, managed feature designed specifically for centralizing logs across accounts without needing additional infrastructure or data transformation.

Exam trap

The trap here is that candidates often confuse CloudWatch Logs Insights (a query tool) with the actual cross-account log forwarding mechanism, or they incorrectly assume that CloudTrail or Kinesis Data Firehose are the primary services for cross-account log centralization, when in fact the native CloudWatch Logs cross-account subscription is the correct, managed solution.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs Insights is a query and analysis tool for searching and visualizing log data within a single account; it does not provide cross-account log ingestion or forwarding. Option B is wrong because AWS CloudTrail records API activity and can deliver logs to CloudWatch Logs, but it is not a mechanism for centralizing logs from multiple accounts into a single CloudWatch Logs destination. Option C is wrong because Amazon Kinesis Data Firehose is a streaming data delivery service that can load logs into destinations like S3 or Redshift, but it is not designed for direct cross-account CloudWatch Logs subscription; it would require additional setup and does not natively support the cross-account subscription model.

125
MCQeasy

A DevOps engineer needs to monitor the number of messages in an Amazon SQS queue and trigger an Auto Scaling policy to add more EC2 instances when the queue depth exceeds a threshold. Which CloudWatch metric should the alarm use?

A.NumberOfMessagesSent
B.ApproximateNumberOfMessagesNotVisible
C.ApproximateNumberOfMessagesVisible
D.SentMessageSize
AnswerC

ApproximateNumberOfMessagesVisible is the correct CloudWatch metric for monitoring queue depth because it reports the number of messages available in the SQS queue for immediate retrieval by consumers. This value directly reflects the backlog of work waiting to be processed, making it the standard metric used for autoscaling consumers or triggering alarms. It is approximate due to SQS's distributed architecture, but it is polled frequently enough to provide a reliable real-time indicator of queue pressure.

Why this answer

(ApproximateNumberOfMessagesVisible) is the correct metric because it represents the number of messages available to be retrieved from the queue. When this number exceeds a threshold, it indicates that consumer EC2 instances are not keeping up, so an Auto Scaling policy can add more instances to handle the load. Option A (NumberOfMessagesSent) is a count of messages sent, not the current queue depth.

Option B (ApproximateNumberOfMessagesNotVisible) represents messages that are in flight (being processed) and not available for retrieval. Option D (SentMessageSize) is the size of messages sent, not the count. Therefore, C is correct.

126
MCQmedium

A company's DevOps team notices that their Amazon RDS for PostgreSQL instance's CPU utilization spikes to 90% every day at 10:00 AM, causing application latency. They want to be notified when the CPU utilization exceeds 80% for more than 5 minutes to investigate the cause. Which solution should they implement?

A.Use Amazon CloudWatch Logs Insights to query the RDS logs and trigger an SNS notification when CPU utilization is high.
B.Enable AWS Trusted Advisor to automatically create a CloudWatch alarm on the CPU utilization metric.
C.Create an Amazon CloudWatch alarm on the CPUUtilization metric with a period of 5 minutes and a threshold of 80, and set the alarm action to send a notification to an Amazon SNS topic.
D.Create an AWS CloudTrail trail to monitor CPU utilization and trigger an AWS Lambda function to send an email notification.
AnswerC

A CloudWatch alarm on CPUUtilization with a five-minute period and 80% threshold, notifying via Amazon SNS, directly matches the stated requirement to alert when utilisation exceeds 80% for more than five minutes, enabling investigation of the daily 10:00 AM spike.

Why this answer

Amazon CloudWatch natively collects the CPUUtilization metric for RDS, and an alarm with a 5-minute period and an 80% threshold directly matches the requirement to alert when CPU exceeds 80% for more than 5 minutes. Configuring the alarm action to publish to an Amazon SNS topic delivers the notification to the DevOps team, which is the standard, lowest-effort solution.

Exam trap

DOP-C02 often tests the misconception that CloudTrail or Logs Insights can monitor resource metrics, when only CloudWatch metrics and alarms are designed for threshold-based performance alerting.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs Insights queries log data, not metrics, and cannot itself trigger SNS notifications based on CPU utilization thresholds. Option B is wrong because AWS Trusted Advisor provides best-practice checks and does not automatically create CloudWatch alarms on RDS CPU metrics. Option D is wrong because AWS CloudTrail records API activity, not performance metrics like CPU utilization, so it cannot detect or alert on resource saturation.

127
Drag & Dropmedium

Drag and drop the steps to set up an AWS CodeBuild project to build a Docker image and push it to Amazon ECR.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

The correct order to set up an AWS CodeBuild project for building a Docker image and pushing to Amazon ECR is: first create the ECR repository, then write the buildspec.yml file, then create the CodeBuild project with privileged mode enabled, and finally start the build. This sequence ensures the repository exists, the build specification is ready, and the project has the necessary permissions to build and push the Docker image.

128
MCQeasy

A DevOps engineer needs to set up a monitoring solution for an AWS Lambda function that processes messages from an Amazon SQS queue. The engineer wants to be alerted if the function fails to process a message (i.e., the message ends up in the dead-letter queue). Which approach should they use?

A.Create a CloudWatch alarm on the ApproximateNumberOfMessagesVisible metric of the dead-letter queue.
B.Enable CloudTrail to log SQS API calls and create a metric filter for SendMessage to the DLQ.
C.Create a CloudWatch Events rule to monitor the Lambda function errors.
D.Configure the Lambda function's dead-letter queue to send notifications via Amazon SNS.
AnswerA

The ApproximateNumberOfMessagesVisible metric of the dead-letter queue is a native SQS metric that reports the number of messages available for retrieval. When the redrive policy moves messages from the source queue to the DLQ after the maximum receive count is exceeded, this metric immediately increases. Creating a CloudWatch alarm on this metric provides a direct, near-real-time signal that messages are failing processing, and you can trigger an SNS notification or other action from the alarm.

Why this answer

The `ApproximateNumberOfMessagesVisible` metric on the dead-letter queue (DLQ) directly reflects the number of messages that have failed processing and been moved there. By creating a CloudWatch alarm on this metric (e.g., when it exceeds 0 for a period), the engineer receives an alert precisely when messages are failing, without needing to parse logs or rely on indirect indicators.

Exam trap

The trap here is that candidates often confuse monitoring Lambda function errors (Option C) with monitoring DLQ messages, not realizing that a message can end up in the DLQ due to exhaustion of retries (configured in the SQS event source mapping) without the Lambda function itself throwing an error.

How to eliminate wrong answers

Option B is wrong because CloudTrail logs SQS API calls (like SendMessage) but does not provide a real-time metric for DLQ message count; creating a metric filter on SendMessage to the DLQ would require parsing every API call and does not natively aggregate to a simple alarm threshold. Option C is wrong because monitoring Lambda function errors (e.g., via CloudWatch Events or Lambda metrics) captures function invocation failures but does not specifically indicate that a message was sent to the DLQ—messages can fail processing without a Lambda error (e.g., if the function returns an error but the SQS trigger retries and eventually sends to DLQ). Option D is wrong because configuring the Lambda function's DLQ to send notifications via SNS would require custom code or configuration to publish a notification each time a message is moved to the DLQ, which is not a built-in feature of SQS or Lambda; SNS can be used as a target for DLQ messages only if the DLQ itself is an SNS topic, but SQS DLQs are queues, not topics, and SNS does not automatically emit notifications when messages are added to an SQS queue.

129
MCQeasy

An application running on Amazon EC2 instances sends custom metrics to CloudWatch. The team notices that some metrics are not appearing. What is the most likely cause?

A.The custom metric namespace is not pre-registered in CloudWatch.
B.The EC2 instances are in a private subnet without a NAT gateway.
C.The IAM role attached to the EC2 instance lacks permissions to publish metrics.
D.The CloudWatch agent is not installed or configured on the EC2 instances.
AnswerD

Custom metrics are not emitted by default by EC2; they require an explicit mechanism such as the CloudWatch agent or direct PutMetricData calls from the application. If the agent is not installed or its configuration file (common for statsd or collectd) is missing/invalid, no metric data will be sent even though the application is running. This is the most common root cause when an application expects to send custom metrics but nothing appears in CloudWatch, because without the agent’s collection loop, nothing triggers the metric submission.

Why this answer

The most likely cause is that the CloudWatch agent is not installed or configured on the EC2 instances. Custom metrics require the CloudWatch agent to collect and send data to CloudWatch. Option A is incorrect because custom metric namespaces do not need to be pre-registered; they are automatically created when metrics are published.

Option B is incorrect because instances in a private subnet can still publish metrics using a VPC endpoint or a proxy, so a NAT gateway is not strictly required. Option C is incorrect because even with proper IAM permissions, the CloudWatch agent must be installed and running to collect and send custom metrics. Therefore, option D is correct.

130
MCQeasy

A DevOps engineer is troubleshooting an issue where an Amazon RDS for MySQL instance is experiencing high latency. The engineer wants to identify which queries are causing the problem. Which AWS service should be used?

A.Amazon RDS Performance Insights
B.AWS CloudTrail
C.Amazon CloudWatch Metrics
D.VPC Flow Logs
AnswerA

Amazon RDS Performance Insights is the correct choice because it is a dedicated database performance tuning feature that visualizes database load and provides granular, per-query metrics. It breaks down DB load by wait events, SQL statements, and hosts, directly surfacing the exact query text and execution statistics (e.g., rows returned, latency) causing high load. This allows a DevOps engineer to identify problematic queries quickly, unlike network or API-level logs.

Why this answer

Amazon RDS Performance Insights is a database performance tuning tool that visualizes database load (Average Active Sessions) broken down by SQL statement, wait event, user, and host. It directly identifies which queries are consuming the most DB load, making it the correct tool for diagnosing high-latency MySQL workloads. It retains up to 7 days of free performance data and up to 2 years with paid retention.

Exam trap

DOP-C02 often tests the distinction between control-plane logging (CloudTrail), infrastructure metrics (CloudWatch), and database-level query analytics (Performance Insights) — candidates confuse CloudWatch Metrics with query-level diagnostics.

How to eliminate wrong answers

Option B is wrong because AWS CloudTrail records API-level control-plane activity (who called CreateDBInstance, etc.), not SQL query execution or database performance. Option C is wrong because CloudWatch Metrics provides aggregate database metrics like CPUUtilization, FreeableMemory, and ReadIOPS — useful for spotting resource pressure but not for attributing latency to specific queries. Option D is wrong because VPC Flow Logs capture IP-level network metadata (source/dest IP, ports, bytes), not database query behavior.

131
MCQmedium

A company is using Amazon RDS for PostgreSQL and wants to monitor the database for performance issues. They need to capture slow queries and analyze them over time. Which combination of AWS services should they use?

A.Enable RDS Event notifications to send alerts for performance issues.
B.Enable RDS Performance Insights and Enhanced Monitoring.
C.Enable CloudWatch Logs for PostgreSQL and export logs to Amazon S3.
D.Use CloudWatch Metrics to monitor database connections and CPU utilization.
AnswerB

RDS Performance Insights is the correct approach because it provides a database-load dashboard that correlates wait events, SQL statements, and hosts in real time, allowing you to pinpoint the exact queries causing bottlenecks. Enhanced Monitoring complements this by delivering OS-level metrics—including CPU, memory, file I/O, and process lists—from the underlying hypervisor, which helps identify resource saturation that affects query performance. Together, these services give both the query-level and system-level telemetry required to diagnose slow PostgreSQL performance.

Why this answer

RDS Performance Insights provides a database performance schema with detailed wait events and SQL-level metrics to identify slow queries, while Enhanced Monitoring offers OS-level metrics (CPU, memory, disk I/O) at sub-minute granularity. Together, they allow you to capture and analyze slow queries over time without additional log parsing or export overhead.

Exam trap

The trap here is that candidates often confuse general monitoring (CloudWatch Metrics) or log export (CloudWatch Logs to S3) with the specialized, integrated performance analysis tools (Performance Insights and Enhanced Monitoring) that are purpose-built for diagnosing slow queries in RDS.

How to eliminate wrong answers

Option A is wrong because RDS Event notifications only send alerts for instance lifecycle events (e.g., failover, maintenance) and do not capture or analyze slow query performance data. Option C is wrong because exporting PostgreSQL logs to S3 via CloudWatch Logs provides raw log files, but requires additional tooling (e.g., Athena) to parse and analyze slow queries; it is not a native, integrated solution for ongoing performance analysis. Option D is wrong because CloudWatch Metrics for database connections and CPU utilization provide aggregate resource metrics, not the detailed query-level or wait-event data needed to identify and analyze slow queries.

132
MCQeasy

A company uses AWS CloudTrail to log API activity in their AWS account. They need to ensure that any changes to CloudTrail configuration itself are detected and alerted upon in real time. Which service should they use?

A.Use Amazon CloudWatch Events (EventBridge) to create a rule matching the StopLogging or UpdateTrail API calls.
B.Enable AWS Config rules to monitor CloudTrail configuration changes.
C.Use Amazon CloudWatch Logs Insights to query CloudTrail logs for changes.
D.Enable Amazon GuardDuty to detect changes to CloudTrail.
AnswerA

EventBridge rules match CloudTrail management events by API name, so a rule filtering StopLogging and UpdateTrail invokes a target such as SNS within seconds of the call. This delivers the real-time detection of tampering with CloudTrail's own configuration that the scenario requires.

Why this answer

Amazon CloudWatch Events (EventBridge) can monitor CloudTrail API calls in real time by creating a rule that matches specific API calls such as StopLogging or UpdateTrail. When these calls are made, the rule triggers an action (e.g., SNS notification or Lambda function) to alert administrators immediately. This provides the real-time detection required for changes to CloudTrail configuration itself.

Exam trap

The trap here is that candidates often confuse AWS Config (which is for compliance and configuration history) with real-time event-driven alerting, or they think GuardDuty covers all security monitoring, but neither provides the specific real-time API call detection that EventBridge offers.

How to eliminate wrong answers

Option B is wrong because AWS Config rules are designed for continuous compliance assessment and configuration auditing, not real-time event-driven alerting; they evaluate resources periodically or on configuration changes but do not provide instantaneous alerts. Option C is wrong because CloudWatch Logs Insights is a query tool for analyzing historical log data, not a real-time alerting mechanism; it cannot proactively detect changes as they occur. Option D is wrong because Amazon GuardDuty is a threat detection service that focuses on malicious activity and anomalies (e.g., unusual API calls or compromised credentials), not specifically on monitoring CloudTrail configuration changes for compliance or operational awareness.

133
MCQeasy

A company wants to receive a notification when an AWS IAM user creates a new access key. Which AWS service should be used to capture this event and trigger a notification?

A.Amazon GuardDuty
B.AWS CloudTrail with CloudWatch Events
C.Amazon CloudWatch
D.AWS Config
AnswerB

AWS CloudTrail records all API calls made by IAM users as events, and CloudWatch Events (now Amazon EventBridge) can be configured with event patterns to match specific actions. This integration enables real-time notifications by routing matched events to an SNS topic or Lambda function. Thus, CloudTrail provides the audit log while CloudWatch Events performs the filtering and alerting.

Why this answer

AWS CloudTrail captures API activity, including the CreateAccessKey event when an IAM user creates a new access key. By sending these events to Amazon CloudWatch Events (now part of Amazon EventBridge), you can define a rule that triggers a notification via SNS, Lambda, or other targets. This combination provides the real-time event-driven notification the company requires.

Exam trap

The trap here is that candidates often confuse CloudWatch (which handles metrics and logs) with CloudWatch Events (which handles event-driven triggers), leading them to pick Option C, even though CloudWatch alone cannot capture API calls without CloudTrail integration.

How to eliminate wrong answers

Option A is wrong because Amazon GuardDuty is a threat detection service that analyzes VPC Flow Logs, DNS logs, and CloudTrail management events for malicious activity, but it does not directly trigger custom notifications for specific IAM actions like CreateAccessKey. Option C is wrong because Amazon CloudWatch is a monitoring service for metrics, logs, and alarms, but it cannot natively capture API-level events like CreateAccessKey; it relies on CloudTrail or CloudWatch Events to ingest such events. Option D is wrong because AWS Config evaluates resource configurations and compliance rules, but it does not capture real-time API calls or trigger notifications for specific IAM user actions like creating an access key.

134
MCQeasy

A company is using AWS CloudFormation to manage infrastructure. The DevOps team wants to receive notifications when CloudFormation stack creation fails. Which AWS service should be used to capture the stack failure event and send a notification?

A.Amazon SQS
B.Amazon CloudWatch Logs
C.AWS CloudTrail
D.Amazon EventBridge
AnswerD

Amazon EventBridge is the correct service because it is a serverless event bus that natively captures AWS service events, including CloudFormation stack lifecycle changes (CREATE, UPDATE, DELETE), and delivers them to targets like SNS, Lambda, or SQS for immediate action. EventBridge rules filter on resource type, event detail type, and other JSON fields, allowing you to build precise event-driven workflows without custom polling. This makes it the standard, low-latency solution for reacting to CloudFormation events in real time.

Why this answer

Amazon EventBridge can capture CloudFormation stack events (such as CREATE_FAILED) using event rules and route them to targets like Amazon SNS for notifications. Option A is wrong because Amazon SQS is a message queue service and does not directly send notifications; it requires a consumer. Option B is wrong because Amazon CloudWatch Logs stores log data but does not capture CloudFormation events or send notifications.

Option C is wrong because AWS CloudTrail records API calls but is not designed for real-time event-driven notifications; it is better suited for auditing.

135
MCQmedium

A company uses AWS CloudTrail to log all API calls in their AWS account. They need to ensure that any changes to CloudTrail configuration (such as disabling the trail or modifying the log file validation) are immediately detected and trigger an automated response. Which solution should the DevOps engineer implement?

A.Enable Amazon GuardDuty and configure it to monitor CloudTrail logs for suspicious activity.
B.Create an Amazon EventBridge rule that matches CloudTrail API calls like StopLogging or UpdateTrail and triggers an SNS topic.
C.Use AWS Config rules with remediation actions to detect and revert changes to CloudTrail.
D.Use AWS Trusted Advisor to check CloudTrail configuration and send alerts via email.
AnswerB

Amazon EventBridge can natively consume CloudTrail management events and evaluate them against a rule with an event pattern that matches specific API calls such as StopLogging or UpdateTrail. Because the rule is event-driven, it triggers an SNS topic within seconds of the incident, enabling immediate notification to security operations. This approach requires no polling or configuration state evaluation, making it the most direct real-time mechanism for responding to changes to CloudTrail itself.

Why this answer

Amazon EventBridge can match AWS CloudTrail API calls such as StopLogging or UpdateTrail and trigger an SNS topic for immediate notification and automated response. This provides real-time detection of changes to CloudTrail configuration, satisfying the requirement.

Exam trap

The trap is confusing monitoring services like GuardDuty or Config with real-time event-driven detection; candidates may pick Config because it can detect changes, but it lacks the immediate EventBridge-based response required.

How to eliminate wrong answers

Option A is wrong because GuardDuty monitors for malicious activity and anomalies, not specifically for CloudTrail configuration changes, and it does not provide immediate automated response to such changes. Option C is wrong because AWS Config rules evaluate configuration changes but are not real-time and require remediation actions that may take time; they are better for compliance auditing than immediate detection. Option D is wrong because Trusted Advisor provides best-practice checks and email alerts, but it is not real-time and does not trigger automated responses.

136
MCQhard

A company runs a critical application on Amazon ECS with Fargate. The application experiences intermittent slow responses. The DevOps team enabled Container Insights and CloudWatch ServiceLens. However, traces from the application do not appear in ServiceLens. The application uses the AWS X-Ray SDK for tracing. What is the MOST likely cause?

A.The X-Ray daemon is not running in the task definition.
B.The X-Ray SDK cannot send traces to AWS X-Ray from Fargate tasks.
C.ServiceLens does not support Amazon ECS with Fargate launch type.
D.The application is not sending metrics to CloudWatch Container Insights.
AnswerA

The X-Ray daemon must be deployed as a sidecar container in the ECS task definition when using the Fargate launch type. The SDK only emits UDP trace data to localhost:2000; without the daemon listening, those UDP packets are silently dropped and no segments ever reach the X-Ray service. Fargate does not run a host-level X-Ray agent, so the task definition must explicitly include the daemon container and the application must be configured to communicate with it via the daemon's port.

Why this answer

The X-Ray daemon is required to act as a local intermediary that receives trace segments from the X-Ray SDK and forwards them to the AWS X-Ray API. In Amazon ECS with Fargate, the daemon must be explicitly included as a sidecar container in the task definition. Without it, the SDK cannot send traces, which explains why traces are missing from ServiceLens despite the SDK being integrated.

Exam trap

The trap here is that candidates assume the X-Ray SDK can send traces directly to the AWS X-Ray API without a local daemon, but the SDK is designed to offload segment buffering and transmission to the daemon, making it a mandatory component in containerized environments like Fargate.

How to eliminate wrong answers

Option B is wrong because the X-Ray SDK can send traces from Fargate tasks when the X-Ray daemon is properly configured as a sidecar container; there is no inherent limitation preventing trace transmission from Fargate. Option C is wrong because ServiceLens fully supports Amazon ECS with Fargate launch type, including both EC2 and Fargate, as long as the required agents and permissions are in place. Option D is wrong because Container Insights metrics are not required for traces to appear in ServiceLens; ServiceLens aggregates traces from X-Ray and metrics from CloudWatch independently, and missing metrics do not prevent trace visibility.

137
MCQmedium

A DevOps team has set up centralized logging for multiple AWS accounts using Amazon OpenSearch Service. The team uses CloudWatch cross-account observability to collect logs from various accounts into a monitoring account. Recently, logs from one source account stopped appearing in the monitoring account's OpenSearch dashboard. Other source accounts continue to send logs successfully. Which step should the team take to troubleshoot this issue?

A.Verify that the monitoring account's CloudWatch cross-account observability is enabled.
B.Check the source account's CloudWatch Logs subscription filter for the OpenSearch destination.
C.Review the source account's CloudWatch Logs retention policy to confirm logs are not expired.
D.Ensure the IAM role in the source account has the correct trust policy for the monitoring account.
AnswerB

The source account's CloudWatch Logs subscription filter is the mechanism that continuously streams matching log events to a destination such as an OpenSearch ingestion pipeline or a Kinesis stream. If this filter was accidentally removed, disabled, or its destination ARN changed, log forwarding for that account ceases even though the logs remain in CloudWatch Logs. Verifying and if necessary recreating the subscription filter and its destination policy is the correct first step to restore the flow.

Why this answer

The most likely cause of logs from a single source account failing to appear is a misconfigured or broken CloudWatch Logs subscription filter. This filter is responsible for forwarding log events from the source account to the OpenSearch destination in the monitoring account. If the filter is missing, misconfigured, or has been accidentally deleted, logs will not be sent, while other accounts continue to work normally.

Exam trap

The trap here is that candidates confuse the cross-account observability setup (which uses IAM roles and trust policies) with the actual log delivery mechanism (subscription filters), leading them to check the IAM role or the monitoring account configuration instead of the source account's subscription filter.

How to eliminate wrong answers

Option A is wrong because cross-account observability is already working for other source accounts, so the monitoring account's feature is enabled. Option C is wrong because a retention policy would cause logs to stop appearing for all accounts after the retention period, not selectively for one account. Option D is wrong because the IAM role's trust policy is used for cross-account access to CloudWatch metrics and logs, but the actual log delivery to OpenSearch is handled by the subscription filter, not by assuming a role in the source account.

138
MCQmedium

A company uses AWS CloudTrail to log API activity. The security team needs to be alerted when an IAM user creates a new access key. How can this be achieved with minimal overhead?

A.Use CloudWatch Logs Insights to run a query every hour and send results via email.
B.Set up an AWS Config rule to detect when an access key is created.
C.Configure S3 event notifications on the CloudTrail bucket to trigger a Lambda function.
D.Create an Amazon EventBridge rule that matches the 'CreateAccessKey' event and targets an SNS topic.
AnswerD

An EventBridge rule with an event pattern that matches the AWS API-call event for 'CreateAccessKey' (specifically detail.eventSource for iam.amazonaws.com and detail.eventName) synchronously triggers an SNS topic, delivering email or other notifications in near-real-time. This is the natively supported pattern for reacting to CloudTrail API activity, as CloudTrail forwards events to EventBridge, and it is fully serverless with no polling or log-parsing required.

Why this answer

D is correct because Amazon EventBridge can capture real-time CloudTrail API events, such as 'CreateAccessKey', and route them directly to an SNS topic for immediate notification. This approach requires no polling, no custom code, and minimal overhead, as EventBridge handles event matching and delivery natively.

Exam trap

The trap here is that candidates often confuse AWS Config (which evaluates resource state) with EventBridge (which evaluates API events), leading them to choose Option B, but Config cannot react to API calls in real-time and is not designed for event-driven alerting.

How to eliminate wrong answers

Option A is wrong because running a CloudWatch Logs Insights query every hour introduces latency (up to 1 hour delay) and requires manual setup for scheduling and email delivery, which is not minimal overhead. Option B is wrong because AWS Config rules are designed for resource configuration compliance and drift detection, not for real-time API event alerting; they evaluate resources periodically or on configuration changes, not on API calls. Option C is wrong because S3 event notifications on the CloudTrail bucket would trigger on object creation (log file delivery), not on the specific API event itself, and CloudTrail log files are typically delivered in batches with a delay, making this approach less precise and more complex.

139
Multi-Selecteasy

A company is using AWS CloudFormation to deploy a microservices architecture. The operations team wants to receive real-time notifications when any stack operation fails. Which TWO AWS services can be used together to achieve this?

Select 2 answers
A.AWS Lambda
B.Amazon CloudWatch Logs
C.AWS CloudTrail
D.Amazon CloudWatch Events (Amazon EventBridge)
E.Amazon Simple Notification Service (SNS)
AnswersD, E

EventBridge is a serverless event bus that captures AWS service events, including CloudFormation stack status changes. You can create an event rule with a pattern matching CloudFormation events such as CREATE_FAILED or UPDATE_COMPLETE, and route them to targets like SNS, Lambda, or Step Functions. This provides near real-time, event-driven notifications for CloudFormation lifecycle changes, making it the correct service to detect and act on these events.

Why this answer

Amazon CloudWatch Events (now Amazon EventBridge) can capture CloudFormation stack state changes, including failures, and trigger an Amazon SNS topic to send real-time notifications. Therefore, options D (CloudWatch Events/EventBridge) and E (SNS) are correct. Option A is incorrect because AWS Lambda alone does not provide notification capabilities; it would need to be triggered by another service.

Option B is incorrect because CloudWatch Logs stores logs but does not natively trigger notifications for stack failures. Option C is incorrect because AWS CloudTrail records API activity for auditing, not real-time stack event notifications.

140
MCQmedium

A company uses AWS CloudFormation to deploy a three-tier web application. The stack includes an Application Load Balancer (ALB), an Auto Scaling group of EC2 instances, and an Amazon RDS Multi-AZ database. The DevOps team has configured the EC2 instances to send application logs to CloudWatch Logs using the CloudWatch agent. They also set up a CloudWatch alarm on the ALB's 5xx error count. During a recent deployment, the team noticed that the alarm did not trigger even though the application was returning 5xx errors. The team verified that the CloudWatch agent is running on the instances and logs are appearing in CloudWatch Logs. What should the team do to ensure the alarm triggers correctly?

A.Create a metric filter on the EC2 instance's log group to count 5xx errors and create an alarm on that.
B.Change the CloudWatch alarm to use the 'HTTPCode_ELB_5XX_Count' metric instead.
C.Restart the CloudWatch agent on the EC2 instances.
D.Verify that the CloudWatch alarm is using the correct metric 'HTTPCode_Target_5XX_Count' and that the threshold is appropriate.
AnswerD

For a three-tier web application behind an ALB, the correct metric to alarm on for backend errors is HTTPCode_Target_5XX_Count, which counts the number of HTTP 5xx responses sent by registered targets. You must verify that the alarm uses this metric and that the threshold is aligned with your accepted error budget—for example, alarm after 10 errors in 5 minutes—to avoid noisy alerts from transient issues. This metric directly reflects the health of your EC2 instances as the target group, making it the appropriate signal for detecting application failures.

Why this answer

The ALB publishes two separate 5xx metrics: HTTPCode_ELB_5XX_Count (errors generated by the load balancer itself, e.g., 503 due to no healthy targets) and HTTPCode_Target_5XX_Count (errors returned by the backend targets). Since the application is returning 5xx errors from the EC2 instances, the alarm should use HTTPCode_Target_5XX_Count. The team likely set the alarm on HTTPCode_ELB_5XX_Count, which would not trigger for backend errors.

Option A is not the best because creating a log-based metric filter would create a separate alarm and does not fix the existing ALB alarm. Option B is wrong because switching to HTTPCode_ELB_5XX_Count still tracks the wrong metric. Option C is unrelated because the agent is already working and logs are flowing.

Exam trap

Recognize the difference between ALB-generated 5xx errors (HTTPCode_ELB_5XX_Count) and target-generated 5xx errors (HTTPCode_Target_5XX_Count).

141
Multi-Selecthard

A company wants to centralize logging from multiple AWS accounts and regions. The logs should be stored in a central S3 bucket for compliance. Which THREE steps are required to achieve this? (Choose THREE.)

Select 3 answers
A.Create a cross-account subscription in Amazon CloudWatch Logs to stream logs to the central account.
B.Create an S3 bucket in the central account with appropriate bucket policy granting permissions to CloudTrail.
C.Enable CloudTrail in each region where the company operates.
D.Enable AWS CloudTrail in each account and configure it to deliver logs to the central S3 bucket.
E.Turn on CloudTrail data events for all S3 and Lambda resources.
AnswersB, C, D

This bucket policy must explicitly grant the CloudTrail service principal from each source account permission to write objects, typically via s3:PutObject, and also add s3:GetBucketAcl to allow CloudTrail to verify ownership. Without these cross-account permissions, CloudTrail in other accounts cannot deliver logs to the bucket, so this step is a mandatory prerequisite for any centralized log aggregation. The policy should also include a condition that restricts access to specific CloudTrail resources or source accounts to prevent unauthorized writes.

Why this answer

To centralize logs from multiple accounts into a single S3 bucket, the bucket in the central account must have a bucket policy that explicitly grants the necessary permissions (e.g., `s3:PutObject`) to the CloudTrail service principal (`cloudtrail.amazonaws.com`) from each source account. This policy allows CloudTrail in the source accounts to write log files directly into the central bucket, enabling centralized storage for compliance.

Exam trap

The trap here is that candidates often confuse the need for a cross-account CloudWatch Logs subscription (Option A) with the direct S3 delivery mechanism of CloudTrail, or they mistakenly think enabling data events (Option E) is mandatory for centralization, when only management events and proper S3 bucket policy configuration are required.

142
MCQhard

A company runs a critical application on Amazon ECS with Fargate. The DevOps engineer wants to receive alerts when the application's error rate exceeds 5% over a 5-minute period. Which combination of services should be used?

A.Amazon CloudWatch Synthetics to monitor the application endpoint and create an alarm.
B.Amazon CloudWatch Logs Insights to query logs every 5 minutes and trigger a Lambda function.
C.Amazon CloudWatch Logs metric filter to count errors, then a CloudWatch alarm.
D.Amazon CloudWatch Contributor Insights to detect error patterns, then an alarm.
AnswerC

A CloudWatch Logs metric filter parses each incoming log event in real time, incrementing a custom (or extracted) metric whenever a log line matches a pattern (e.g., contains "ERROR" or a JSON field). You can then create a CloudWatch Alarm on that metric with an appropriate threshold, period, and statistic (e.g., Sum of errors over 5 minutes). This approach provides native, serverless, low-latency alerting without additional compute, and it is exactly the intended use case for detecting error rates from logs.

Why this answer

The correct solution is to use Amazon CloudWatch Logs metric filter to count errors from application logs, which creates a custom metric representing the error count or rate. Then, a CloudWatch alarm can be configured on that metric to trigger when the error rate exceeds 5% over a 5-minute period. Option C is correct.

Option A is incorrect because CloudWatch Synthetics is for synthetic monitoring (endpoint health checks), not for analyzing logs for error rates. Option B is incorrect because CloudWatch Logs Insights is an interactive query tool for ad-hoc analysis, not designed for real-time alarming. Option D is incorrect because CloudWatch Contributor Insights analyzes top contributors (e.g., error sources), but does not directly count error rates for alarms.

143
Multi-Selecthard

A company uses Amazon CloudWatch Synthetics canaries to monitor endpoint availability. The canaries are failing intermittently with timeout errors. The DevOps team needs to diagnose the issue. Which THREE aspects should they investigate?

Select 3 answers
A.The canary schedule and frequency.
B.The canary script execution time and memory usage.
C.Network connectivity and routing from the canary's VPC to the target endpoint.
D.CloudWatch Synthetics canary logs for error messages.
E.CloudWatch Logs retention policy for the canary logs.
AnswersB, C, D

CloudWatch Synthetics enforces a maximum execution time per canary run (default 2 minutes, configurable up to 14 minutes). If the script's code, including any synchronous HTTP requests or Puppeteer operations, exceeds this limit, the run is terminated with a timeout error. Similarly, if the memory allocated to the canary's Lambda function is exceeded, the process can be killed, which also surfaces as a timeout or failure. Monitoring these metrics directly identifies whether the script itself is the bottleneck.

Why this answer

Canary scripts have a maximum execution time of 5 minutes (300 seconds) and a memory limit of 1 GB. If the script execution time or memory usage exceeds these limits, the canary will fail with a timeout error. Investigating these metrics in CloudWatch can reveal whether the script is too resource-intensive or slow, causing the intermittent failures.

Exam trap

The trap here is that candidates may confuse canary schedule frequency with execution timeout, thinking that running the canary less often will fix timeout errors, when the root cause is actually script performance or network latency.

144
Multi-Selecthard

A company runs a containerized application on Amazon ECS with Fargate. They want to monitor the application logs and metrics. Which THREE steps should they take to collect and visualize this data? (Choose THREE.)

Select 3 answers
A.Create a CloudWatch Dashboard to display logs and metrics.
B.Use AWS CloudFormation to monitor resource metrics.
C.Configure the ECS task definition to use the awslogs log driver.
D.Use AWS X-Ray to trace requests and collect logs.
E.Enable the CloudWatch agent as a sidecar container in the task definition.
AnswersA, C, E

A CloudWatch Dashboard is a customizable operational view that consolidates CloudWatch Logs insights, metric graphs, and alarm states into a single pane. For an ECS workload, you can combine ECS service-level metrics, container instance metrics, and log-group-based insights to visualize application health and troubleshoot in near real time. This directly satisfies the requirement to display logs and metrics, making it a valid monitoring component.

Why this answer

CloudWatch Dashboards can display both logs and metrics, allowing visualization of application data collected from ECS tasks. Option C is correct because the awslogs log driver sends container logs to CloudWatch Logs directly from the ECS task definition. Option E is correct because the CloudWatch agent can be deployed as a sidecar container to collect custom metrics and logs from the application.

Option B is incorrect because AWS CloudFormation is used for infrastructure provisioning and management, not for monitoring resource metrics. Option D is incorrect because AWS X-Ray is designed for tracing and analyzing requests, not for collecting logs or application metrics.

145
MCQeasy

A company needs to monitor the CPU utilization of its Amazon RDS for PostgreSQL instance. The metric should be available in Amazon CloudWatch with a granularity of 1 minute. Which action should the team take?

A.Install the CloudWatch agent on the RDS instance.
B.Enable Enhanced Monitoring for the RDS instance.
C.No additional configuration is needed; RDS automatically sends metrics to CloudWatch.
D.Enable Performance Insights for the RDS instance.
AnswerC

RDS automatically emits the CPUUtilization metric to CloudWatch in the AWS/RDS namespace for every database instance, with no setup or configuration required. This metric is collected from the hypervisor at the instance level and is available in the CloudWatch console, enabling alarms and dashboards immediately after the database is provisioned. Basic monitoring is included as part of the service, and you can optionally use detailed monitoring or Enhanced Monitoring for faster granularity.

Why this answer

Amazon RDS for PostgreSQL automatically publishes metrics, including CPU utilization, to CloudWatch with a default granularity of 1 minute for standard instances. No additional configuration is required to enable this basic monitoring. The metrics are collected by the RDS hypervisor layer and sent to CloudWatch without needing an agent or extra setup.

Exam trap

The trap here is that candidates often confuse Enhanced Monitoring (which provides OS-level metrics at higher granularity) with basic CloudWatch monitoring, leading them to incorrectly select Option B when the question only requires standard 1-minute CPU utilization metrics.

How to eliminate wrong answers

Option A is wrong because the CloudWatch agent cannot be installed on an RDS instance; RDS is a managed service that does not allow direct OS-level access or agent installation. Option B is wrong because Enhanced Monitoring provides OS-level metrics (e.g., memory, disk I/O) at a granularity of 1 second or more, but it is not required for basic CPU utilization metrics, which are already sent to CloudWatch at 1-minute granularity. Option D is wrong because Performance Insights is a database performance tuning feature that visualizes database load and waits, not a mechanism for sending CPU utilization metrics to CloudWatch.

146
MCQeasy

A DevOps engineer is troubleshooting a slow API response. They suspect that the issue is related to database queries. The application runs on EC2 instances behind an ALB and uses Amazon RDS for MySQL. Which monitoring approach will provide the most granular insight into database query performance?

A.Enable CloudWatch Logs for the RDS instance and query the error log.
B.Install the CloudWatch agent on the EC2 instances to collect database performance counters.
C.Enable Enhanced Monitoring and Performance Insights for the RDS instance.
D.Monitor the ALB's 'TargetResponseTime' metric and correlate with RDS 'CPUUtilization'.
AnswerC

Enhanced Monitoring delivers OS-level metrics for the RDS host at up to one-second granularity, while Performance Insights provides database load per wait event, per SQL, and per host, enabling you to link slow API calls to specific queries and bottlenecks. Together they let you determine whether the issue is CPU saturation, storage I/O, lock waits, or a poorly performing statement, and then drill into the top SQL contributing to database load. This is the targeted diagnostic combination for an RDS-backed API performance problem.

Why this answer

Amazon RDS Performance Insights provides a database load view that breaks down wait events by SQL statement, user, host, and database, which is the most granular way to identify slow queries. Enhanced Monitoring complements this with OS-level metrics at up to one-second granularity. Together they give the DevOps engineer the query-level and host-level detail needed to pinpoint the slow API's database bottleneck.

Exam trap

DOP-C02 often tests the difference between high-level CloudWatch metrics and query-level database tooling, tricking candidates into choosing CloudWatch-based options that cannot reveal per-query wait events.

How to eliminate wrong answers

Option A is wrong because the RDS error log records errors and warnings, not query execution times or wait events, so it cannot reveal which query is slow. Option B is wrong because the CloudWatch agent on EC2 instances collects OS metrics from the application servers, not from the RDS database engine, so it cannot see query-level performance. Option D is wrong because ALB TargetResponseTime and RDS CPUUtilization are high-level metrics that show correlation but not causation — they cannot identify the specific slow query or its wait event.

147
MCQmedium

A company is running a production web application on Auto Scaling EC2 instances behind an ALB. They have enabled detailed CloudWatch metrics on the EC2 instances and enabled CloudTrail. Recently, users reported intermittent 503 errors. The operations team reviews CloudWatch dashboards but sees no spike in CPU or memory. What is the MOST likely cause of the 503 errors?

A.Insufficient CloudTrail logging trail configuration
B.The target group has an insufficient number of healthy instances due to health check failures
C.Detailed monitoring is disabled for the EC2 instances
D.The security group for the ALB is misconfigured
AnswerB

The ALB routes requests only to targets that have successfully passed their configured health checks. If health check failures occur—due to an incorrect health check path, a timeout threshold being too low, or an application dependency failing—the corresponding instances are marked unhealthy and removed from the rotation. When the number of healthy targets drops below the minimum needed (typically zero healthy targets in a target group), the ALB returns HTTP 503 Service Unavailable. This can happen without CPU or memory utilization rising, because health checks validate application-level readiness, not just resource utilization.

Why this answer

ALB returns HTTP 503 when no healthy targets are available in the target group to serve the request. If health checks are failing intermittently — due to application-level issues, misconfigured health check paths, or slow responses — targets are marked unhealthy and removed from rotation, leaving insufficient capacity and producing 503s. Because CPU and memory show no spike, the cause is not resource saturation but target availability.

Exam trap

DOP-C02 often tests the distinction between ALB 503 (no healthy targets) and 502 (bad gateway from a target) — candidates chase resource metrics or security group misconfigurations when the real signal is target health check failures.

How to eliminate wrong answers

Option A is wrong because CloudTrail logs API activity for auditing, not application request handling; insufficient trail configuration cannot cause 503 errors. Option C is wrong because detailed monitoring (1-minute metrics) affects metric granularity, not target health — disabling it would not cause 503s, and the scenario says detailed metrics are enabled. Option D is wrong because a misconfigured ALB security group would typically cause connection timeouts or refused connections (and would affect all traffic consistently), not intermittent 503s from the ALB itself.

148
MCQhard

A media company runs a video transcoding pipeline on AWS. The pipeline uses AWS Step Functions to orchestrate multiple Lambda functions that transcode video files stored in Amazon S3. The company wants to implement a monitoring solution to track the progress of each workflow execution, including which step is currently running, the duration of each step, and any errors. The solution should provide near real-time visibility and allow the team to troubleshoot failed executions quickly. Which solution meets these requirements?

A.Create custom CloudWatch metrics from Lambda functions for each step, and build a CloudWatch dashboard.
B.Use Amazon EventBridge to capture Step Functions execution status changes and build a custom dashboard in CloudWatch.
C.Configure each Lambda function to write logs to CloudWatch Logs with the execution ID, and use CloudWatch Logs Insights to query and visualize.
D.Enable AWS X-Ray tracing on the Step Functions and Lambda functions to get a service map and trace details.
AnswerB

Correct. Amazon EventBridge captures Step Functions execution state changes (e.g., 'ExecutionStarted', 'TaskStateEntered', 'ExecutionFailed') in near real-time. These events can be used to build a CloudWatch dashboard that shows the current step, duration per step, and errors, meeting all requirements without custom instrumentation.

Why this answer

Amazon EventBridge (formerly CloudWatch Events) can capture Step Functions execution state changes (e.g., step started, succeeded, failed). These events can be used to build a custom dashboard in CloudWatch, providing near real-time visibility into workflow progress, step durations, and errors. Option A is incorrect because creating custom metrics from Lambda functions requires additional instrumentation and does not provide workflow-level context easily.

Option C is incorrect because CloudWatch Logs Insights queries are not near real-time; they require searching through logs, and the solution needs real-time visibility. Option D is incorrect because AWS X-Ray provides distributed tracing for individual requests, but it does not offer high-level workflow step tracking with durations and errors aggregated across executions in near real-time.

149
MCQhard

A company is using Amazon RDS for MySQL and needs to monitor the number of slow queries. They have enabled slow query logs. How can they effectively monitor and alert on the number of slow queries per minute?

A.Use RDS Events to send slow query metrics to CloudWatch.
B.Enable RDS Enhanced Monitoring and publish metrics to CloudWatch.
C.Use AWS CloudTrail to monitor SQL queries.
D.Publish slow query logs to CloudWatch Logs, create a metric filter, and set an alarm.
AnswerD

The correct approach is to enable the RDS MySQL slow query log by setting slow_query_log=1 and long_query_time in the DB parameter group, then configure RDS to stream those logs to a CloudWatch Logs log group. Once the log events are flowing, you create a CloudWatch Logs metric filter—patterned to match the slow query log format, such as lines containing 'Query_time'—to count matching events as a custom metric, and then set a CloudWatch alarm to trigger when the count exceeds a threshold.

Why this answer

Slow query logs from RDS MySQL can be published to CloudWatch Logs. Once the logs are in CloudWatch Logs, you can create a metric filter to count the number of slow queries per minute and set a CloudWatch alarm on that metric to alert when the count exceeds a threshold. Option A is incorrect because RDS Events provide notifications about instance events (e.g., failover, maintenance), not slow query metrics.

Option B is incorrect because Enhanced Monitoring provides OS-level metrics (e.g., CPU, memory) but not slow query log data. Option C is incorrect because CloudTrail records API calls made to AWS services, not SQL queries executed within RDS.

150
MCQhard

A DevOps team is implementing a comprehensive logging strategy for a microservices architecture running on Amazon EKS. They need to collect logs from all containers and send them to a centralized log analytics platform. The solution must be agentless and support multi-line log events. Which approach should the team use?

A.Deploy a Fluent Bit DaemonSet on the EKS cluster and configure it to send logs to Amazon CloudWatch Logs.
B.Use the Amazon CloudWatch agent as a sidecar container in each pod to forward logs to CloudWatch Logs.
C.Install the Amazon Kinesis Agent on each EC2 instance and configure it to stream logs to Amazon Kinesis Data Firehose.
D.Deploy a Fluentd DaemonSet on the EKS cluster and configure it to send logs to Amazon S3.
AnswerA

Fluent Bit is a lightweight, high-throughput log processor that runs as a DaemonSet, placing one pod on every cluster node. It automatically discovers and collects container stdout/stderr logs without requiring application-side changes, making it effectively agentless for application teams. It supports multi-line log parsing and its native CloudWatch Logs output plugin streams logs directly to CloudWatch Logs for real-time aggregation. This is the recommended pattern for comprehensive logging on EKS.

Why this answer

Fluent Bit is a lightweight, CNCF-graduated log processor that can be deployed as a DaemonSet on EKS to collect logs from all nodes without requiring sidecar containers. It supports multi-line log events natively via its multiline filter plugin, and it can output directly to Amazon CloudWatch Logs using the cloudwatch_logs output plugin, meeting the agentless requirement since it runs as a Kubernetes DaemonSet rather than as a per-pod sidecar.

Exam trap

The trap here is that candidates often confuse 'agentless' with 'no software at all,' but in Kubernetes, agentless means no sidecar injection per pod; a DaemonSet is considered agentless because it runs as a cluster-level service, not as part of the application deployment.

How to eliminate wrong answers

Option B is wrong because deploying the CloudWatch agent as a sidecar container in each pod is not agentless; it requires modifying every pod definition and increases resource overhead, whereas the requirement specifies an agentless solution. Option C is wrong because the Amazon Kinesis Agent is an EC2-level agent that must be installed on each underlying EC2 instance, which is not agentless and does not integrate with EKS pod-level log collection; it also does not natively support multi-line log events without custom configuration. Option D is wrong because Fluentd is a heavier log collector compared to Fluent Bit, and while it can send logs to Amazon S3, S3 is a storage service, not a centralized log analytics platform; the requirement specifies sending logs to a centralized log analytics platform, which CloudWatch Logs fulfills.

← PreviousPage 2 of 3 · 197 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Monitoring and Logging questions.