DOP-C02 Resilient Cloud Solutions Practice Question
A company runs a microservices architecture on Amazon ECS with Fargate. Services communicate via an internal Application Load Balancer (ALB). The operations team notices that occasional traffic spikes cause increased latency and timeouts. The team wants to improve resilience without over-provisioning. Which THREE steps should be taken? (Choose THREE.)
⚠ Common exam trap
Many exam-takers confuse vertical scaling (increasing task resources) with horizontal scaling (adding more tasks), and they may overlook the importance of application-level resilience patterns like graceful shutdowns and service mesh features for traffic management.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable ECS Service Connect for inter-service communication to manage traffic distribution.
B is correct because ECS Service Connect provides built-in traffic management for inter-service communication, including load balancing, retries, and circuit breaking. This helps distribute traffic more evenly during spikes, reducing latency and timeouts without requiring over-provisioning.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the CPU and memory limits in the task definitions.
Why it's wrong here
Increasing the CPU and memory limits in the task definitions only affects the resources available to the container; it does not address latency caused by inefficient application logic, slow downstream dependencies, or network bottlenecks. If the task is not actually CPU- or memory-bound, raising these limits wastes cluster capacity and increases cost without improving response times. The correct approach is to first identify the real bottleneck before allocating more resources.
- ✓
Enable ECS Service Connect for inter-service communication to manage traffic distribution.
Why this is correct
ECS Service Connect provides a resilient service mesh that simplifies inter-service communication by giving each service a stable DNS name and managing traffic distribution at the application layer. It enables fine-grained traffic splitting, automatic retries, and connection draining, which collectively reduce latency and improve reliability between microservices. This is a direct remedy for latency caused by inefficient service-to-service calls, as it optimizes the network path and avoids overloading individual task instances.
- ✓
Configure ECS service auto scaling with a target tracking policy based on ALB request count per target.
Why this is correct
Configuring ECS service auto scaling with a target tracking policy based on ALB request count per target ensures the number of running tasks dynamically adjusts to match actual traffic demand. This prevents service saturation during traffic spikes by adding tasks before response times degrade, and removes tasks when demand drops, maintaining cost efficiency. By directly aligning capacity with request volume, this option mitigates latency caused by under-provisioning under variable load.
- ✓
Implement a graceful shutdown handler in the application to handle SIGTERM.
Why this is correct
A graceful shutdown handler in the application listens for the SIGTERM signal sent by ECS during scale-in or rolling deploys, allowing the process to stop accepting new connections, finish draining in-flight requests, and clean up resources before the container stops. Without this handler, ECS forcibly kills the task mid-request, causing dropped connections, client retries, and noticeable latency spikes during deployments or scale-down events. This control is essential for maintaining service quality and avoiding errors during lifecycle transitions.
- ✗
Use EC2 launch type with Spot Instances to reduce cost.
Why it's wrong here
Using the EC2 launch type with Spot Instances is a cost optimization strategy, not a latency improvement method. Spot Instances can be reclaimed by AWS with a two-minute warning, causing abrupt task termination and potential data loss or request failures, which introduces unpredictable latency and reliability issues. For a latency-sensitive microservices architecture, the instability of Spot Instances can degrade performance, and any cost savings are overshadowed by increased operational complexity and inconsistency.
Go deeper
Related to this question
About these practice questions
This DOP-C02 question is part of Courseiva's 1,298-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.