DOP-C02 Resilient Cloud Solutions Practice Question
A company runs a production e-commerce platform on AWS. The architecture includes an Application Load Balancer (ALB) that distributes traffic to a fleet of Amazon EC2 instances running in an Auto Scaling group across three Availability Zones (AZs). The application stores session state in Amazon ElastiCache for Redis (cluster mode disabled) with a single node. The database is an Amazon Aurora MySQL DB cluster with one writer and two reader instances in different AZs. The platform experiences intermittent slowdowns and occasional timeouts during peak traffic hours. The CloudWatch metrics show that the ALB's TargetResponseTime is elevated, and the Redis CPU utilization is consistently above 80% during these periods. The Auto Scaling group is scaling out, but new instances take several minutes to become healthy. The DevOps team has been asked to improve the resilience and performance of the application with minimal changes to the application code. Which solution should the team implement?
⚠ Common exam trap
The trap here is that candidates focus on scaling the database or load balancer (options A, B, C) instead of recognizing that the single-node Redis cache is the bottleneck and requires horizontal scaling and high availability to resolve both performance and resilience issues.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Migrate ElastiCache for Redis to a cluster mode enabled configuration with multiple shards and enable Multi-AZ with automatic failover. Also, use an ElastiCache replication group with read replicas in different AZs.
The primary bottleneck is the single-node Redis instance (CPU > 80%), which cannot scale horizontally and lacks high availability. Migrating to cluster mode enabled with multiple shards distributes CPU load across shards, while Multi-AZ with automatic failover and read replicas in different AZs provides high availability and read scaling. This directly addresses the elevated ALB TargetResponseTime caused by Redis latency. Note that the application will need to use a Redis Cluster-compatible client; since the stem allows minimal code changes, this solution is still appropriate.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Replace the ALB with a Network Load Balancer (NLB) to reduce latency, and use an Auto Scaling group with a step scaling policy based on Redis CPU utilization.
Why it's wrong here
Replacing the ALB with an NLB only optimizes Layer 4 load balancing and does nothing to relieve CPU pressure on the ElastiCache node, which is the actual bottleneck. Moreover, Auto Scaling step policies cannot use Redis CPU utilization as a native CloudWatch metric because ElastiCache publishes cache metrics (e.g., CPUUtilization) only at the cache-cluster level, and the policy would scale application instances rather than the cache itself, leaving the overworked Redis primary untouched.
- ✗
Increase the instance size of the ElastiCache for Redis node and the size of the Aurora writer instance. Also, increase the cooldown period for the Auto Scaling group to allow new instances to warm up.
Why it's wrong here
Vertically scaling the ElastiCache node and Aurora writer may temporarily mask the CPU saturation, but it does not remove the single point of failure and hits hard instance-size ceilings, especially for Redis with synchronous writes. Increasing the Auto Scaling cooldown period slows down the addition of web/app capacity during traffic spikes, which can worsen latency and undermine the very elasticity needed to handle the production load, so this approach both fails to address root cause and degrades responsiveness.
- ✗
Implement Amazon RDS Proxy in front of the Aurora cluster to reduce database connection overhead, and increase the size of the Redis instance to handle more connections.
Why it's wrong here
RDS Proxy reduces connection churn and overhead for the Aurora database, but the reported bottleneck is Redis CPU saturation, not database connection limits. Increasing the Redis instance size provides more memory and possibly vCPUs, but a single large node still cannot shard data or spread read load, so CPU remains pinned during peak traffic and the cache remains a single point of failure—this option optimizes the wrong tier entirely.
- ✓
Migrate ElastiCache for Redis to a cluster mode enabled configuration with multiple shards and enable Multi-AZ with automatic failover. Also, use an ElastiCache replication group with read replicas in different AZs.
Why this is correct
Migrating to cluster mode enabled with multiple shards horizontally partitions the Redis keyspace across nodes, which directly lowers per-shard CPU utilization and enables linear scaling as traffic grows. Enabling Multi-AZ with automatic failover and placing read replicas in different AZs provides high availability and lets reads be served by replicas, reducing primary node load and cutting failover time from minutes to seconds—this tackles both the immediate CPU bottleneck and the resilience requirement for a production e-commerce platform.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DOP-C02 question from scratch — 1,298 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.