A company runs a critical application on a fleet of EC2 instances managed by an Auto Scaling group. The application generates logs that are sent to CloudWatch Logs using the CloudWatch agent. Recently, the operations team noticed that some instances are missing logs for certain periods. The CloudWatch agent is configured to batch log events and send them every 5 seconds. The instances have high CPU utilization (90%+) during the missing periods. The DevOps engineer suspects that the agent is being throttled or failing. Which of the following is the MOST likely cause and the BEST course of action?
The CloudWatch agent runs as a separate user-space daemon that periodically reads log files and sends them to the CloudWatch Logs API. When the host's CPU is saturated — especially on T-series instances with exhausted CPU credits — the agent's log collection and flush loop can be delayed or preempted for long enough that it begins dropping buffered events to avoid creating an ever-growing backlog. Increasing instance size or CPU credits gives the agent the scheduling time it needs to reliably process and upload log batches, directly resolving the observed missing periods.
Why this answer
When CPU utilization is sustained at 90%+, the CloudWatch agent competes for CPU and can be starved, causing it to drop or delay log batches. The agent buffers events in memory and on disk; under CPU starvation, the buffer may overflow or the agent may fail to flush within the 5-second interval. Increasing CPU credits (for burstable instances) or moving to a larger instance size gives the agent the resources it needs.
Exam trap
DOP-C02 often tests resource contention — candidates blame network or disk because logs are I/O, but the question explicitly states high CPU, pointing to agent starvation.
How to eliminate wrong answers
Option A is wrong because network saturation would affect all traffic, not just logs, and the symptom is CPU-related (90%+ utilization). Option B is wrong because a 1-day retention policy deletes logs after one day, not during the missing periods, and retention does not cause gaps in delivery. Option D is wrong because disk space issues would produce agent errors about buffer overflow, and the question points to CPU as the stressor.