You are designing a rolling update playbook for a 20-node application cluster. The application requires that no more than 25% of the nodes be unavailable at any time. You want Ansible to automatically pause the play if the failure rate within a batch exceeds a threshold, so that you can investigate before continuing. Which play-level keyword should you use?
`max_fail_percentage` is a play-level keyword that aborts the play if the percentage of hosts that fail in a batch exceeds the specified value. In this scenario, setting it to a value that corresponds to the allowed unavailability (e.g., 25) will cause Ansible to stop the rolling update when too many hosts fail, allowing you to investigate before more batches are affected.
Why this answer
The `max_fail_percentage` keyword is designed to abort a play when the failure rate within a batch exceeds a given percentage. By setting it to 25, Ansible will stop the rolling update if more than 25% of the hosts in a batch fail, matching the application's availability requirement. This provides an automatic safety check during the update.
Exam trap
The trap here is confusing `max_fail_percentage` with `any_errors_fatal` or `serial`. `max_fail_percentage` is the only keyword that lets you define a percentage-based failure threshold per batch, while `serial` only controls batch size and `any_errors_fatal` aborts on any single error.