Courseiva

DOP-C02 Incident and Event Response Practice Question

A DevOps engineer receives a CloudWatch alarm indicating that an EC2 instance's CPU utilization has exceeded 90% for 10 minutes. The instance is part of an Auto Scaling group behind an Application Load Balancer. What is the MOST efficient initial step to troubleshoot the high CPU usage?

⚠ Common exam trap

The trap here is that candidates often jump to scaling actions (Options B or D) or application-layer metrics (Option C) without first checking the instance's foundational health metrics, specifically CPU credit balance for burstable instances, which is the most efficient diagnostic step per AWS Well-Architected best practices.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Review the EC2 instance's CloudWatch metrics for CPU credit balance and network utilization.

Reviewing the EC2 instance's CloudWatch metrics for CPU credit balance and network utilization is the most efficient initial step to diagnose high CPU usage. CPU credit balance is critical for burstable performance instances (e.g., T2/T3), as a depleted credit balance directly causes sustained high CPU utilization. Network utilization metrics can reveal if the high CPU is driven by excessive traffic or a DDoS-like pattern, allowing targeted remediation without unnecessary scaling or configuration changes.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Review the EC2 instance's CloudWatch metrics for CPU credit balance and network utilization.

    Why this is correct

    Reviewing the instance's CPU credit balance is the correct first step because a T-series instance that exhausts its earned credits will be throttled to the baseline CPU, causing slowdowns even when the alarm threshold is breached. The network utilization metric is also essential since high network throughput can drive CPU overhead from interrupt handling and packet processing, which may be the actual root cause. Together, these metrics reveal whether the alarm reflects genuine resource contention or a misconfigured alarm threshold.

  • ✗

    Modify the Auto Scaling group to use a larger instance type.

    Why it's wrong here

    Modifying the Auto Scaling group to use a larger instance type is a reactive capacity change rather than a diagnostic action; it assumes the problem is insufficient vCPUs without verifying whether the instance is actually CPU-constrained. It also involves updating the launch template and potentially recreating instances, which is disruptive and costly, and it may mask a temporary spike that a credit balance review would have identified. This should be a later remediation step, not the first response to an alarm.

  • ✗

    Check the ALB's HTTP 5xx error rate metric for the target group.

    Why it's wrong here

    Checking the ALB's HTTP 5xx error rate for the target group is a valid health indicator but it measures application layer failures, not host CPU utilization. A 5xx can result from application bugs, timeouts, or connection limits and may occur even when CPU utilization is normal, so it cannot explain why the CPU alarm fired. The alarm is specifically about resource consumption on the EC2 instance, making instance-level metrics the relevant diagnostic source.

  • ✗

    Immediately increase the desired capacity of the Auto Scaling group.

    Why it's wrong here

    Immediately raising the desired capacity of the Auto Scaling group is a blind reaction that increases cost and can create a false sense of resolution without identifying the underlying cause. If the problem is CPU credit exhaustion on a single unexpectedly busy instance, adding more instances will not fix the existing instance's throttling unless scaling policies and health checks force a replacement. Proper troubleshooting requires reading the instance's CloudWatch metrics first, because an alarm that goes off while utilization is low may indicate an alarm threshold problem rather than a true capacity shortage.

About these practice questions

Courseiva writes every DOP-C02 question from scratch — 1,298 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.