Ansible Serial Keyword for Quorum in Rolling Updates
An operations team is designing a rolling update for a stateful application that requires quorum (minimum 3 out of 5 nodes online). They plan to use Ansible's serial keyword. Which serial value ensures the update proceeds without breaking quorum while still being efficient?
Quick Answer
serial: 2 is correct because it is the largest batch size that still respects the quorum requirement stated in the scenario: with 5 nodes total and a minimum of 3 needed online at all times, taking 2 nodes down for the update leaves exactly 3 running, which is the smallest number that still satisfies quorum. Any larger batch size would drop the online count below 3 and break quorum, which for a stateful application could mean data unavailability or a split-brain condition, not just a performance blip, so this is not just about minimizing downtime in a general sense, it is about respecting a hard constraint the application imposes. At the same time, serial: 2 is more efficient than the safer but slower serial: 1, since it updates the fleet in fewer batches while still never crossing the line the application requires. This is the general pattern worth taking away: when a scenario gives you a specific quorum, minimum-online, or capacity constraint alongside a total node count, the correct serial value is the largest batch size that keeps the number of remaining online nodes at or above that stated minimum - go any higher and you violate the constraint, any lower and you are sacrificing efficiency for no additional safety.
⚠ Common exam trap
Candidates often confuse 'quorum' with 'majority' and incorrectly choose serial: 3, thinking that 3 out of 5 is a majority, but fail to realize that taking down 3 nodes leaves only 2 online, which is below the quorum threshold of 3.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
serial: 2
Setting serial: 2 ensures that only 2 nodes are taken down at a time during the rolling update. With a quorum requirement of 3 out of 5 nodes, taking down 2 nodes leaves 3 online, maintaining quorum. This is the most efficient value that does not risk breaking quorum.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
serial: 2
Why this is correct
Serial 2 updates two nodes at a time, leaving three of five online, which preserves the quorum minimum throughout the rolling update. Larger values risk dropping below three; serial 1 is safe but slower, so 2 balances safety with efficiency.
- ✗
serial: 1
Why it's wrong here
serial: 1 updates nodes one at a time, keeping four online and preserving quorum, but it is the least efficient choice here. It is tempting as the safest possible setting, and would be correct for clusters with no spare capacity, yet the question asks for efficiency alongside quorum safety.
- ✗
serial: 3
Why it's wrong here
serial: 3 updates three nodes at once, leaving only two online, which falls below the three-node quorum minimum. It is tempting because it matches the quorum size, and would work if the requirement were updating exactly a quorum's worth, but the remaining nodes must still satisfy quorum during each batch.
- ✗
serial: 5
Why it's wrong here
serial: 5 updates all five nodes in one batch, taking the whole cluster offline simultaneously and breaking the three-node quorum. It is tempting as the fastest single-pass option, and would suit a stateless fleet where no node interdependence exists, but quorum-dependent stateful clusters cannot tolerate it.
Go deeper
Related to this question
About these practice questions
This EX294 question is part of Courseiva's 392-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
4 more ways this is tested on EX294
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. Which THREE statements correctly describe the behavior of the 'serial' keyword in Ansible? (Choose exactly three.)
hard- ✓ A.It can be set as a percentage of the total hosts.
- ✓ B.It causes the playbook to run on a subset of hosts at a time.
- ✓ C.It can be combined with max_fail_percentage to control failure thresholds.
- D.It guarantees that only one task runs across all hosts at any time.
- E.It applies globally to all plays in the playbook.
Why A: Option A is correct because the serial keyword accepts a percentage value (e.g., serial: 25%) that Ansible interprets as a fraction of the total hosts in the play, batching them accordingly. Option B is correct because serial defines the number of hosts (or batch size) that Ansible targets per play iteration, so the play runs on a subset of hosts at a time rather than all at once. Option C is correct because serial is commonly paired with max_fail_percentage, which aborts the play if failures within a serial batch exceed the given threshold, enabling controlled rolling updates. Option D is incorrect because serial controls host batching, not task concurrency; it does not guarantee only one task runs across all hosts at any time. Option E is incorrect because serial is set at the play level and applies only to that specific play, not globally to all plays in a playbook.
Variation 2. Which TWO of the following are best practices when coordinating rolling updates with Ansible?
hard- ✓ A.Define a 'max_fail_percentage' to abort the update if too many hosts fail.
- ✓ B.Use the 'serial' keyword to update a subset of hosts at a time.
- C.Use 'strategy: free' to allow hosts to run tasks independently.
- D.Use 'gather_facts: no' to speed up the playbook.
- E.Set 'any_errors_fatal: true' to stop the update on the first failure.
Why A: Option B is correct because the 'serial' keyword controls how many hosts are targeted per play iteration, which is the fundamental mechanism for performing a rolling update in Ansible—updating a small batch at a time rather than all hosts at once. Option A is correct because 'max_fail_percentage' works together with 'serial' to abort the entire play if the number of failed hosts in a batch exceeds the defined threshold, preventing a bad rollout from cascading across the fleet. Option C is not a best practice for rolling updates because 'strategy: free' lets each host proceed through tasks independently without waiting for others, which breaks the controlled batch-by-batch ordering that rolling updates require. Option D is not inherently a rolling-update best practice; disabling fact gathering may speed execution but does not coordinate or sequence updates and can break plays that depend on facts. Option E is not appropriate here because 'any_errors_fatal: true' aborts the whole play on the very first host failure, which is more aggressive than the graduated tolerance provided by 'max_fail_percentage' and is not the recommended pairing for controlled rolling updates.
Variation 3. Which TWO options are best practices for coordinating rolling updates with Ansible? (Choose exactly two.)
easy- A.Set ignore_errors: yes to ensure the playbook continues even if some hosts fail.
- ✓ B.Use the serial keyword to update hosts in batches.
- C.Use the default serial setting (all hosts) for simplicity.
- ✓ D.Set max_fail_percentage to limit the number of failed hosts before aborting.
- E.Run all hosts in parallel to minimize total update time.
Why B: Option B is correct because the serial keyword in Ansible controls how many hosts are targeted per play iteration, allowing rolling updates in controlled batches (e.g., serial: 2 or serial: 25%) so that only a subset of hosts is updated at a time while the rest continue serving traffic. Option D is correct because max_fail_percentage defines a failure threshold within a serial batch; if the percentage of failed hosts exceeds that value, Ansible aborts the play, preventing a bad update from cascading across the entire fleet. Option A is not a best practice for rolling updates because ignore_errors: yes masks failures and lets the playbook proceed even when hosts are broken, defeating the safety purpose of batching. Option C is wrong because the default serial value is effectively all hosts in the play, which performs a simultaneous update rather than a rolling one. Option E is also wrong because running all hosts in parallel maximizes blast radius and downtime risk, the opposite of a rolling-update strategy.
Variation 4. A large enterprise manages thousands of servers grouped by data center. They are designing a rolling update that must complete within a maintenance window. Which combination of Ansible strategies best minimizes total update time while maintaining safety?
hard- A.Set serial: 0 to update all hosts simultaneously.
- ✓ B.Set serial to 10% and max_fail_percentage to 25%.
- C.Set forks to 100 and max_fail_percentage to 50.
- D.Set serial to 1 to update one host at a time with max_fail_percentage: 0.
Why B: Setting `serial: 10%` updates hosts in batches of 10% of the inventory, which parallelizes the update across many hosts to minimize total time, while `max_fail_percentage: 25%` provides a safety net by aborting the play if more than 25% of the batch fails, preventing a cascade of failures from taking down the entire data center. This combination balances speed and safety for large-scale rolling updates within a maintenance window.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This EX294 practice question is part of Courseiva's free Red Hat certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the EX294 exam.