CloudFormation Stack Update Rollback Failure: Troubleshooting Steps for DevOps Engineers
An organization uses AWS CloudFormation to manage infrastructure. During an incident, a stack update fails with 'UPDATE_ROLLBACK_FAILED' status. The engineer needs to bring the stack to a consistent state without losing data. What is the BEST approach?
⚠ Common exam trap
Many exam-takers choose manual correction (Option C) thinking they can fix the resource and retry the update, but they overlook that the stack is in a failed rollback state that blocks further updates until the rollback is resolved, making 'ContinueUpdateRollback' the only viable path to a consistent state without data loss.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the 'ContinueUpdateRollback' API to skip the resource that caused the failure.
The 'ContinueUpdateRollback' API is the best approach because it allows the stack to resume the rollback process, skipping the resource that caused the failure, and bringing the stack to a consistent 'UPDATE_ROLLBACK_COMPLETE' state without manual intervention or data loss. This API is specifically designed for the 'UPDATE_ROLLBACK_FAILED' status, enabling you to skip resources that cannot be rolled back (e.g., due to a non-reversible change) while preserving the rest of the stack's state.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use the 'ContinueUpdateRollback' API to skip the resource that caused the failure.
Why this is correct
The `ContinueUpdateRollback` API is the designed recovery action when a CloudFormation stack is stuck in the `UPDATE_ROLLBACK_FAILED` state. By invoking it with the `ResourcesToSkip` parameter, you explicitly instruct CloudFormation to skip the specific resource that caused the rollback failure, allowing the stack to return to a stable `UPDATE_COMPLETE` state. This bypasses the problematic resource without requiring manual intervention. It is the recommended and least disruptive method to recover from a failed stack update.
- ✗
Create a new stack from the same template and migrate resources.
Why it's wrong here
Creating a new stack from the same template and migrating resources is an unnecessarily complex and risky alternative. It requires manually replicating or re-creating all resources, which may involve data migration, downtime, and reconfiguration of dependent systems. CloudFormation's `CreateStack` does not automatically capture the existing stack's state, and migrating an existing stack's live resources is not a trivial operation. Moreover, this approach does not resolve the original stack's failed rollback state, leaving the old stack in a broken condition unless it is also cleaned up.
- ✗
Manually correct the resource configuration that caused the failure, then perform a stack update.
Why it's wrong here
Manually correcting the resource configuration that caused the failure, then performing a stack update, is against CloudFormation's management model. Any manual changes outside CloudFormation create stack drift, meaning the actual resource state no longer matches the template-defined desired state. A subsequent stack update would likely detect mismatches, attempt to 'fix' the resource based on the template, and could cause an unintended replacement or further failure. CloudFormation does not recognize manual repairs as part of its state machine, so the stack remains in `UPDATE_ROLLBACK_FAILED` until you use an API operation like `ContinueUpdateRollback`.
- ✗
Delete the stack and then recreate it from the same template.
Why it's wrong here
Deleting the stack and then recreating it from the same template is an extreme and potentially destructive action. In an `UPDATE_ROLLBACK_FAILED` state, the stack may not delete cleanly because the same resource that prevented rollback could also block stack deletion. Furthermore, if the stack contains stateful resources (e.g., an RDS database) or data-bearing volumes, deleting the stack risks irreversible data loss unless deletion protection or a snapshot strategy is in place. Recreating also loses any existing resource IDs, endpoints, or configurations that dependencies may rely on. CloudFormation's `ContinueUpdateRollback` is specifically designed to recover from such failures without resorting to this disruptive approach.
Go deeper
Related to this question
About these practice questions
This DOP-C02 question is part of Courseiva's 1,298-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.