Courseiva

DOP-C02 Incident and Event Response Practice Question

A company runs a critical application on Amazon ECS with Fargate. The application is deployed across multiple Availability Zones and uses an Application Load Balancer (ALB) as the front-end. During a recent incident, users experienced intermittent connectivity failures. The DevOps team suspects that tasks are being stopped due to resource exhaustion. Which combination of metrics and actions should the team use to diagnose and prevent recurrence?

⚠ Common exam trap

Many exam-takers confuse horizontal scaling (increasing task count) with vertical scaling (increasing task size), assuming that adding more tasks resolves resource exhaustion when the actual issue is insufficient resources per task.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Monitor CPU and memory utilization metrics in CloudWatch; increase the task size (CPU and memory) in the task definition.

CPU and memory utilization metrics in CloudWatch directly indicate resource exhaustion, which is the suspected cause of tasks being stopped. Increasing the task size (CPU and memory) in the task definition provides more resources per task, preventing the OOM killer or CPU throttling from stopping tasks, without changing the number of tasks or scaling logic.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Monitor CPU and memory utilization metrics in CloudWatch; increase the task size (CPU and memory) in the task definition.

    Why this is correct

    Fargate task-level CPU and memory metrics in CloudWatch reveal resource exhaustion causing task stops; raising the task definition's CPU and memory values gives tasks sufficient headroom, directly addressing the suspected exhaustion constraint and preventing recurrence.

  • ✗

    Set up CloudWatch Logs for the application and check for out-of-memory errors; then increase the number of tasks.

    Why it's wrong here

    CloudWatch Logs can show out-of-memory errors, but increasing the number of tasks does not address insufficient resources per task; it may spread load but not fix the root cause if tasks still have inadequate resources.

  • ✗

    Monitor NetworkPacketsIn and NetworkPacketsOut metrics in CloudWatch; increase the number of tasks.

    Why it's wrong here

    NetworkPacketsIn and NetworkPacketsOut measure traffic volume, not CPU or memory consumption, so they cannot reveal the resource exhaustion stopping tasks. They are tempting because network metrics do help diagnose connectivity symptoms, and would be the right choice if the incident stemmed from packet loss or throughput saturation.

  • ✗

    Monitor the ALB error metrics (5xx count) and scale the ECS service based on request count.

    Why it's wrong here

    ALB 5xx counts and request-count scaling track front-end symptoms and demand, not the task-level resource exhaustion that stops Fargate tasks. They are tempting because 5xx errors correlate with user-facing failures, and request-based scaling would be correct if the problem were insufficient capacity under load.

About these practice questions

One of 1,298 original DOP-C02 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.