Courseiva

SOA-C02 Reliability and Business Continuity Practice Question

A company runs a critical e-commerce application on Amazon ECS with Fargate launch type, fronted by an Application Load Balancer. The application uses an Amazon ElastiCache for Redis cluster for session state and an Amazon RDS for MySQL Multi-AZ database for persistent data. Recently, during a deployment of a new service version, the application became unresponsive for 15 minutes. The SysOps administrator discovered that the deployment updated the task definition with a new environment variable that pointed to an incorrect ElastiCache endpoint. The ECS service was configured with a rolling update, minimum healthy percent of 50%, and maximum percent of 200%. After the deployment, all tasks failed health checks due to a connection timeout to the wrong Redis endpoint. What is the MOST effective way to prevent this issue in future deployments?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Implement a blue/green deployment strategy using AWS CodeDeploy and test the new task definition before shifting traffic.

Implementing a blue/green deployment with AWS CodeDeploy allows testing the new task definition in a separate target group before shifting traffic. If the new tasks fail health checks (e.g., due to incorrect ElastiCache endpoint), traffic remains on the blue environment, preventing application downtime. Option A is incorrect because a CloudWatch alarm only triggers an alert or rollback after the issue occurs; it does not prevent the deployment from impacting users. Option B is incorrect because updating one task at a time (canary) still exposes tasks to the wrong configuration, and since the minimum healthy percent is 50%, at least half the tasks would fail before detection. Option D is incorrect because the ECS deployment circuit breaker rolls back only after the deployment fails, but it does not prevent the initial impact during the rolling update.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Configure a CloudWatch alarm that triggers an automatic rollback if the error rate exceeds 10%.

    Why it's wrong here

    Configuring a CloudWatch alarm that fires an automatic rollback at a 10% error rate is purely reactive: the bad task definition is already serving real production traffic, and users are already experiencing errors before any rollback action starts. Alarm thresholds also do not evaluate the new task definition's health in advance, and ECS has no native 'alarm-driven rollback' — you would have to build a custom Lambda/Step Functions workflow to interpret the alarm and invoke a deployment rollback, adding complexity and latency. This approach cannot prevent the initial outage caused by the misconfigured variable.

  • ✗

    Update the ECS service to use a canary deployment by updating one task at a time.

    Why it's wrong here

    Updating the ECS service with a canary approach that replaces one task at a time still exposes the faulty task definition to production traffic, and if the image or environment variable is invalid, that single new task will crash or fail health checks while the remaining old tasks continue serving. ECS rolling deployments do not include a pre-traffic validation phase; you only observe the new task's health after it has already been started in the service. If the deployment configuration is too aggressive or the bad task fails slowly, the partial failure can still cause a percentage of requests to be routed to an unhealthy target, resulting in user-facing errors before a rollback or circuit-breaker mechanism activates.

  • ✓

    Implement a blue/green deployment strategy using AWS CodeDeploy and test the new task definition before shifting traffic.

    Why this is correct

    CodeDeploy's blue/green deployment on ECS creates a new 'green' task set running the proposed task definition while the existing 'blue' task set continues to serve full production traffic. This lets you run smoke tests, endpoint checks, or a Lambda-based validation hook against the green task set before shifting any traffic from the blue to the green target group using the ALB's weighted routing. Only after you explicitly confirm the new task definition works can you shift traffic (optionally incrementally), and if issues are detected, you can re-shift back to the original high-fidelity task set with minimal disruption. This pre-validation is exactly what prevents the misconfigured variable from ever affecting end users.

  • ✗

    Enable ECS deployment circuit breaker and set the rollback configuration to automatically roll back failed deployments.

    Why it's wrong here

    Enabling the ECS deployment circuit breaker with automatic rollback triggers only after the service has detected a persistent failure — the bad task definition is already deployed and task health checks have already failed, so the service has experienced real downtime and failed requests during the rollback detection window. The circuit breaker compensates for deployment failures by reverting to the last successful task set, but it does not provide a testing phase or the ability to hold traffic on the new version until validation passes, which is the key missing capability for this scenario. Also, the rollback is binary and can be slow because it depends on health check warning rules and draining time.

About these practice questions

Courseiva writes every SOA-C02 question from scratch — 1,169 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This SOA-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SOA-C02 exam.