DOP-C02 Monitoring and Logging Practice Question
A company runs a multi-region application on Amazon EC2 instances across us-east-1 and eu-west-1. The application uses an Amazon Aurora global database for writes in us-east-1 and reads in eu-west-1. The DevOps team wants to monitor the replication lag between the primary and secondary regions. They have set up a CloudWatch alarm on the AuroraReplicaLag metric in both regions. However, they notice that the alarm in eu-west-1 sometimes triggers false positives when the lag spikes briefly but then recovers. The team wants to reduce false alarms while still being alerted to sustained high lag that could impact read replicas. The team is already using a standard CloudWatch alarm with a period of 1 minute and evaluation periods of 1. What should the team change to reduce false positives?
⚠ Common exam trap
DOP-C02 often tests the difference between changing a threshold (sensitivity) and changing evaluation periods/datapoints (transient filtering), and candidates frequently pick threshold changes or composite alarms when the real fix is sustained-breach evaluation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase the number of evaluation periods to 3, so the alarm triggers only if the lag is high for 3 consecutive minutes.
The false positives occur because a single 1-minute evaluation period triggers on brief, transient lag spikes. Increasing the number of evaluation periods to 3 means the alarm fires only when the lag exceeds the threshold for 3 consecutive 1-minute periods, filtering out short spikes while still catching sustained high lag. This is the standard CloudWatch approach to reduce false positives without losing sensitivity to real issues.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the alarm threshold to a higher value, such as 10 seconds.
Why it's wrong here
Raising the threshold to 10 seconds does not address the root cause of the false positives: brief, transient spikes in replication lag that exceed the current threshold but are not indicative of a sustained problem. With a fixed evaluation period of 1 minute, a single high-lag data point still flips the alarm to ALARM, so a higher threshold only shifts the problem to a different lag value and may cause you to miss genuine replication degradation that stays below the new threshold.
- ✗
Reduce the metric period to 30 seconds to get more granular data.
Why it's wrong here
Reducing the metric period to 30 seconds provides more frequent data points but actually increases the likelihood of false positives. With a 30-second period and the same evaluation period count, the alarm can evaluate on more individual samples per minute, making it even more sensitive to short-lived lag spikes. Moreover, CloudWatch alarms with shorter periods incur higher costs and do not add any smoothing or anomaly detection; they merely increase the resolution of the raw metric, not the alarm's ability to distinguish transient blips from sustained issues.
- ✓
Increase the number of evaluation periods to 3, so the alarm triggers only if the lag is high for 3 consecutive minutes.
Why this is correct
Increasing the number of evaluation periods to 3 with a 1-minute period means the alarm must observe the replication lag exceeding the threshold for three consecutive datapoints (three consecutive minutes) before entering ALARM state. This creates a temporal smoothing effect that filters out brief, self-correcting spikes in AuroraReplicaLag, which are often caused by momentary write bursts or replica catch-up delays. CloudWatch evaluates all three most recent datapoints when determining alarm state, so sustained high lag triggers the alarm, while isolated outliers do not, directly reducing false positives without changing the sensitivity to genuine long-duration lag.
- ✗
Create a composite alarm that triggers when both the AuroraReplicaLag and CPUUtilization metrics are high.
Why it's wrong here
A composite alarm that requires both high AuroraReplicaLag and high CPUUtilization is conceptually flawed because the two metrics are not causally linked in a way that validates replication lag. Replication lag can be high when the replica is idle or when the source has heavy write activity, while CPUUtilization on the replica might be low; conversely, high CPU on the replica can be due to read traffic or other workloads and does not imply replication lag problems. Composite alarms are intended to combine related metrics that jointly define an operational condition (e.g., low free disk space AND high queue length), but here the additional CPU condition would mask real lag-related failures, creating false negatives and leaving the alarm even less reliable.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DOP-C02 question from scratch — 1,298 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.