A company uses Ansible to manage rolling updates of a web server fleet. During a deployment, the playbook fails on one host due to a transient network error, and the rest of the fleet is left in an inconsistent state. Which strategy would best minimize the risk of inconsistency in future rolling updates?
Trap 1: Add retries to each task so transient errors are automatically…
Retries address transient task failures but leave the playbook's linear execution model intact: when a host exhausts retries, remaining hosts still halt mid-rollout. Retries suit flaky API calls or brief network blips within an otherwise healthy run, not fleet-wide consistency during rolling updates, which serial batching with failure isolation handles.
Trap 2: Use a larger serial batch size to complete the rollout faster.
A larger serial batch widens the blast radius: more hosts update concurrently, so a single failure strands a bigger portion of the fleet mid-rollout. Larger batches suit homogeneous, low-risk changes where speed matters and rollback is trivial, not consistency-critical rolling updates needing small, verifiable increments.
Trap 3: Set ignore_errors: yes on all tasks to continue despite failures.
ignore_errors lets the play continue past a failed host, so that host stays on the old release while others advance, producing exactly the inconsistency described. It suits non-critical cleanup or reporting tasks where a failure is tolerable, not deployment steps whose failure must halt the rollout.
- A
Add retries to each task so transient errors are automatically retried.
Why it fails: Retries address transient task failures but leave the playbook's linear execution model intact: when a host exhausts retries, remaining hosts still halt mid-rollout. Retries suit flaky API calls or brief network blips within an otherwise healthy run, not fleet-wide consistency during rolling updates, which serial batching with failure isolation handles.
- B
Use a larger serial batch size to complete the rollout faster.
Why it fails: A larger serial batch widens the blast radius: more hosts update concurrently, so a single failure strands a bigger portion of the fleet mid-rollout. Larger batches suit homogeneous, low-risk changes where speed matters and rollback is trivial, not consistency-critical rolling updates needing small, verifiable increments.
- C
Set ignore_errors: yes on all tasks to continue despite failures.
Why it fails: ignore_errors lets the play continue past a failed host, so that host stays on the old release while others advance, producing exactly the inconsistency described. It suits non-critical cleanup or reporting tasks where a failure is tolerable, not deployment steps whose failure must halt the rollout.
- D
Set max_fail_percentage to 0 in the serial block to abort the rollout on any failure.
Setting max_fail_percentage to 0 makes Ansible abort the entire serial batch the moment any host fails, so remaining hosts never proceed to the next batch. This directly addresses the inconsistent-fleet risk caused by the transient error, halting the rollout before further divergence occurs.