Courseiva

CCNA Coordinate rolling updates Questions

32 questions · Coordinate rolling updates · All types, answers revealed

1
MCQhard

You are designing a rolling update playbook for a 20-node application cluster. The application requires that no more than 25% of the nodes be unavailable at any time. You want Ansible to automatically pause the play if the failure rate within a batch exceeds a threshold, so that you can investigate before continuing. Which play-level keyword should you use?

A.`serial`
B.`any_errors_fatal`
C.`ignore_errors`
D.`max_fail_percentage`
AnswerD

`max_fail_percentage` is a play-level keyword that aborts the play if the percentage of hosts that fail in a batch exceeds the specified value. In this scenario, setting it to a value that corresponds to the allowed unavailability (e.g., 25) will cause Ansible to stop the rolling update when too many hosts fail, allowing you to investigate before more batches are affected.

Why this answer

The `max_fail_percentage` keyword is designed to abort a play when the failure rate within a batch exceeds a given percentage. By setting it to 25, Ansible will stop the rolling update if more than 25% of the hosts in a batch fail, matching the application's availability requirement. This provides an automatic safety check during the update.

Exam trap

The trap here is confusing `max_fail_percentage` with `any_errors_fatal` or `serial`. `max_fail_percentage` is the only keyword that lets you define a percentage-based failure threshold per batch, while `serial` only controls batch size and `any_errors_fatal` aborts on any single error.

2
MCQmedium

During a rolling update using an Ansible playbook with serial: 2, one host in the first batch becomes unreachable. The playbook fails with an unreachable host error. How should the administrator proceed to complete the update on the remaining hosts while excluding the problematic host?

A.Use 'ansible-playbook playbook.yml --forks 1' to slow down the update.
B.Use 'ansible-playbook playbook.yml --limit all:!hostname' to exclude the unreachable host.
C.Add 'any_errors_fatal: false' to the playbook and rerun.
D.Rerun the playbook with the same command; it will skip the unreachable host automatically.
AnswerB

The `--limit all:!hostname` pattern applies an inventory exclusion, so Ansible targets every host except the unreachable one. This satisfies the requirement to continue the rolling update on remaining hosts while excluding the problematic host, without editing the playbook's serial setting or inventory file.

Why this answer

The `--limit` flag with the pattern `all:!hostname` uses Ansible's inventory host pattern syntax to exclude a specific host from the playbook run. This allows the administrator to rerun the playbook against all hosts except the unreachable one, completing the rolling update without re-attempting the failed host. The `serial: 2` setting is irrelevant once the host is excluded, as the playbook will only target the remaining reachable hosts.

Exam trap

The trap here is that candidates assume Ansible automatically retries or skips unreachable hosts on subsequent runs, when in fact it will fail again unless the host is explicitly excluded using `--limit` or the connectivity issue is resolved.

How to eliminate wrong answers

Option A is wrong because `--forks 1` reduces the number of parallel connections to 1, which slows down execution but does not exclude the unreachable host; the playbook will still fail when it attempts to connect to that host. Option C is wrong because `any_errors_fatal: false` (the default) does not prevent failure from an unreachable host; unreachable hosts cause a fatal error regardless of this setting, and the playbook will still abort. Option D is wrong because Ansible does not automatically skip unreachable hosts on a rerun; the playbook will fail again on the same host unless it is explicitly excluded or the connectivity issue is resolved.

3
Drag & Dropmedium

Drag and drop the steps to configure a firewall rule using firewalld to allow HTTPS traffic in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Firewalld commands: check zone, add service with --permanent, reload, verify, test.

4
MCQeasy

In OpenShift, a DeploymentConfig uses the RollingUpdate strategy. Which parameter controls the maximum number of pods that can be unavailable during an update?

A.minReadySeconds
B.maxSurge
C.revisionHistoryLimit
D.maxUnavailable
E.progressDeadlineSeconds
AnswerD

maxUnavailable caps how many pods may be taken down simultaneously during a rolling update, directly satisfying the stem's constraint on maximum unavailable pods. It works alongside maxSurge, which governs extra pods created above the desired replica count.

Why this answer

In OpenShift, the RollingUpdate strategy for a DeploymentConfig uses the `maxUnavailable` parameter to specify the maximum number or percentage of pods that can be unavailable during the update process. This ensures that the desired number of pods remain available to serve traffic while the update rolls out, controlling the trade-off between update speed and availability.

Exam trap

The trap here is that candidates often confuse `maxUnavailable` with `maxSurge`, mistakenly thinking that controlling how many extra pods are created is the same as controlling how many can be unavailable, but `maxSurge` limits overshoot while `maxUnavailable` limits undershoot.

How to eliminate wrong answers

Option A is wrong because `minReadySeconds` controls how long a pod must be ready before it is considered available, not the number of unavailable pods during an update. Option B is wrong because `maxSurge` controls the maximum number of pods that can be created above the desired count during an update, not the number that can be unavailable. Option C is wrong because `revisionHistoryLimit` controls how many old ReplicationControllers are retained for rollback, not the update availability threshold.

Option E is wrong because `progressDeadlineSeconds` sets the maximum time for the deployment to make progress before it is considered failed, not the number of unavailable pods.

5
Multi-Selectmedium

A company uses Ansible to perform a rolling update of 10 web servers behind an HAProxy load balancer. The playbook uses the `serial` keyword and includes tasks to disable a host from the load balancer, update the web server package, and re-enable the host. Which TWO best practices should the administrator apply to minimize downtime and ensure a successful rolling update?

Select 2 answers
A.Use `any_errors_fatal: true` to stop the playbook if any host fails.
B.Set `serial: 1` to update one host at a time.
C.Use `throttle: 1` to limit the number of concurrent tasks across all hosts.
D.Ensure the load balancer draining timeout is longer than the maximum expected update time per host.
E.Use `async` and `poll` to run the update tasks in the background while proceeding to the next host immediately.
AnswersB, D

Updating one host at a time minimizes the impact on the load balancer pool and ensures continuous service availability.

Why this answer

Setting `serial: 1` ensures that only one host is updated at a time, which is the safest way to perform a rolling update without overwhelming the load balancer or causing a service outage. This allows the playbook to complete the full update cycle (disable, update, re-enable) for each host before moving to the next, minimizing the number of hosts out of service simultaneously.

Exam trap

The trap here is confusing `serial` with `throttle` or `async`; candidates often think `throttle` or `async` can achieve the same serialization, but only `serial` ensures one host completes the entire update cycle before the next begins, which is essential for minimizing downtime in a rolling update scenario.

6
MCQmedium

You are performing a rolling update on a 5-node application cluster using an Ansible playbook with `serial: 1`. The playbook includes a task that uses the `uri` module to check the application health endpoint after each node is updated. The health check must wait until the application returns HTTP 200 before proceeding to the next node. Which approach ensures that the playbook waits for the health check to succeed before moving to the next host?

A.Use the wait_for module with the host and port parameters to wait for the application to start listening.
B.Use the uri module with the status_code parameter set to 200 and register the result, then use a wait_for condition on the result.
C.Use the pause module to wait for a fixed amount of time after updating each node.
D.Use the uri module with retries and until to repeatedly check the endpoint until it returns HTTP 200.
AnswerD

The uri module can be used with retries and until to poll the health endpoint until it returns the desired status code. This ensures that the playbook waits for the application to become healthy before proceeding to the next host, which is essential for a safe rolling update.

Why this answer

To ensure the playbook waits for the application to become healthy after each node update, you should use the uri module with retries and until. This combination repeatedly polls the health endpoint until it returns HTTP 200 or the retries are exhausted. This approach is reliable and ensures that the application is ready before proceeding to the next node in the rolling update.

Exam trap

The trap here is assuming that the wait_for module can check HTTP responses, but it is designed for files, ports, or conditions, not HTTP status codes.

7
MCQmedium

A team uses Ansible to update a database cluster with one primary and two replicas. The goal is zero downtime. Which update order is the safest?

A.Update replicas first, then the primary.
B.Update in random order.
C.Update all nodes simultaneously.
D.Update the primary first, then replicas.
AnswerA

Updating replicas first keeps the primary serving traffic throughout, so no write outage occurs. Each replica is drained from the load balancer, patched, and rejoined before touching the next. Only after both replicas run the new version is the primary failed over and updated, satisfying the zero-downtime constraint.

Why this answer

Updating replicas first ensures that if the update introduces a regression, it affects only the read-only replicas, which can be quickly rolled back without impacting write availability. Once replicas are confirmed healthy, the primary is updated and a controlled failover (e.g., using `patronictl switchover` or `repmgr standby switchover`) promotes a replica to primary, minimizing downtime to seconds. This order aligns with the principle of reducing blast radius and maintaining quorum in a cluster.

Exam trap

The trap here is that candidates assume updating the primary first is safer because it is the 'source of truth,' but in a clustered environment with zero-downtime requirements, updating replicas first is the standard practice to preserve write availability and allow safe rollback.

How to eliminate wrong answers

Option B is wrong because updating in random order risks updating the primary first, causing a write outage if the update fails, and may break replication consistency if replicas are updated before the primary without a controlled failover. Option C is wrong because updating all nodes simultaneously can cause a complete cluster outage if the update introduces a bug, and it violates the zero-downtime requirement by potentially losing quorum or causing split-brain scenarios. Option D is wrong because updating the primary first forces a failover to a replica that still runs the old version, which may be incompatible with the updated primary's data format or replication protocol, leading to replication lag or cluster instability.

8
MCQhard

A company wants to implement a rolling update for a stateful application where hosts cannot be updated in parallel due to data consistency. They also need to ensure that if any host fails, the entire update is rolled back. Which strategy meets these requirements?

A.Use serial: 2 and any_errors_fatal: yes
B.Use serial: 1 and ignore_errors: yes
C.Use serial: 0 and max_fail_percentage: 0
D.Use serial: 1 and any_errors_fatal: yes
AnswerD

Serial: 1 forces strictly sequential host updates, preventing parallel changes that would corrupt shared stateful data. Any_errors_fatal: yes aborts the entire play on the first failure, triggering rollback rather than leaving remaining hosts partially updated, meeting both the consistency and rollback constraints.

Why this answer

Setting `serial: 1` ensures hosts are updated one at a time, preventing parallel updates that could break data consistency for a stateful application. Adding `any_errors_fatal: yes` causes the entire playbook run to abort immediately if any host fails, which satisfies the requirement for a full rollback on failure.

Exam trap

The trap here is that candidates often confuse `serial: 1` with `serial: 0` or think `max_fail_percentage: 0` alone triggers a rollback, but `any_errors_fatal` is required to abort the entire playbook immediately on any failure, not just stop the current batch.

How to eliminate wrong answers

Option A is wrong because `serial: 2` allows two hosts to be updated in parallel, violating the requirement that hosts cannot be updated simultaneously due to data consistency. Option B is wrong because `ignore_errors: yes` causes Ansible to continue the update even if a host fails, which prevents the required rollback on failure. Option C is wrong because `serial: 0` is invalid (serial must be a positive integer or a percentage string), and `max_fail_percentage: 0` only stops the batch if 0% of hosts fail, which is effectively the same as `any_errors_fatal` but does not trigger a rollback of the entire update—it only halts further batches without reverting completed changes.

9
MCQhard

You are performing a rolling update of a 15-node RHEL cluster with an Ansible playbook that uses `serial: 5`. During the second batch, a host fails to restart its application service, and the task fails. You want the playbook to stop the entire rollout immediately so that no further batches are updated, allowing you to investigate the failure. Which play keyword should you set to achieve this behavior?

A.`max_fail_percentage: 0`
B.`ignore_errors: false` on the service restart task.
C.`serial: 1`
D.`any_errors_fatal: true`
AnswerD

Setting any_errors_fatal to true causes the play to abort immediately when any task fails on any host, stopping the entire rollout across all batches. This matches the requirement to halt further updates as soon as the service restart failure occurs, allowing investigation before continuing.

Why this answer

To stop the entire play immediately when any task fails on any host, the play must set any_errors_fatal to true. This keyword overrides the default behavior where a failed host is removed and other hosts continue, causing the whole run to abort as soon as the failure occurs. It is the correct mechanism for halting a rolling update on the first error.

Exam trap

The trap here is assuming that max_fail_percentage or serial adjustments will stop the rollout immediately, when any_errors_fatal is the keyword that aborts the play on the first task failure across all hosts.

10
MCQhard

You are writing an Ansible playbook to perform a rolling update of a 20-node RHEL application cluster. The application requires that no more than 20 percent of the fleet be offline at any time. You want the playbook to update hosts in batches that respect this constraint and to abort the entire rollout if the failure rate within any batch exceeds 10 percent. Which combination of play keywords should you use?

A.`serial: 4` and `max_fail_percentage: 25`
B.`serial: 4` and `max_fail_percentage: 10`
C.`serial: 20%` and `max_fail_percentage: 10`
D.`serial: 20` and `max_fail_percentage: 10`
AnswerC

Using `serial: 20%` sizes each batch to 20 percent of the 20-node fleet, which is exactly 4 hosts, satisfying the requirement that no more than 20 percent be offline. `max_fail_percentage: 10` then aborts the rollout if more than 10 percent of the hosts in the current batch fail. This directly encodes both constraints in the play.

Why this answer

The requirement is to cap concurrent offline hosts at 20 percent of the fleet while aborting when batch failures exceed 10 percent. With 20 nodes, 20 percent equals 4 hosts, so `serial: 20%` expresses the constraint as a percentage that scales with inventory size. `max_fail_percentage: 10` then aborts the play when more than 10 percent of the hosts in the current batch fail. Together they encode both the availability and failure-tolerance policies directly in the play.

Exam trap

The trap here is confusing percentage-based serial values with absolute host counts, or assuming max_fail_percentage applies to the whole inventory rather than only to the hosts in the current serial batch.

11
MCQeasy

You are performing a rolling update of a 6-node RHEL web server fleet using an Ansible playbook. The playbook currently uses `serial: 1` and takes a long time because each host is updated sequentially. You want to update two hosts at a time to reduce the total rollout duration while still keeping at least four hosts serving traffic. Which change should you make?

A.Change `serial: 1` to `serial: 2`.
B.Add `forks: 2` to the play.
C.Change `serial: 1` to `serial: 4`.
D.Add `throttle: 2` to the play.
AnswerA

Changing serial from 1 to 2 updates two hosts per batch, leaving four hosts online at all times, which satisfies the requirement to keep at least four hosts serving traffic. It halves the number of batches compared to serial of 1, reducing the total rollout time while respecting the availability constraint.

Why this answer

To update two hosts at a time in a six-node fleet while keeping four hosts online, the serial batch size must be two. Changing serial from 1 to 2 directly controls how many hosts Ansible targets per play iteration, halving the number of batches and shortening the rollout. Other keywords such as throttle and forks affect concurrency of task execution but do not define rolling update batch boundaries.

Exam trap

The trap here is confusing concurrency controls like forks or throttle with the serial keyword, which is the actual mechanism that defines rolling update batch size.

12
MCQhard

A platform team uses an Ansible playbook with 'serial: "20%"' to roll out a new configuration to 50 hosts. During the third batch, a task fails on several hosts, and the play aborts with a message about max_fail_percentage. Which statement correctly describes how Ansible determined the batch size and the failure threshold in this run?

A.The batch size was 5 hosts, and the failure threshold was evaluated against the 5 hosts in the current batch.
B.The batch size was 20 hosts, and the failure threshold was evaluated against 50 hosts.
C.The batch size was 10 hosts, and the failure threshold was evaluated against the 10 hosts in the current batch.
D.The batch size was 10 hosts, and the failure threshold was evaluated against all 50 hosts regardless of batch.
AnswerC

serial: "20%" on 50 hosts produces batches of 10. max_fail_percentage is then calculated against the size of the current serial batch, so failures are compared to those 10 hosts. This is exactly how Ansible determines both values during a rolling update.

Why this answer

When serial is a percentage, Ansible multiplies that percentage by the total number of hosts in the play to determine each batch size. For 50 hosts at 20%, each batch contains 10 hosts. max_fail_percentage is then evaluated against the current batch, so the abort decision is based on failures among those 10 hosts, not the full inventory.

Exam trap

The trap here is assuming max_fail_percentage is measured against the full inventory when serial is expressed as a percentage.

13
MCQhard

An Ansible rolling update playbook has 'serial: 1' and 'max_fail_percentage: 0'. During the update of a 5-host group, the first host fails. What is the outcome?

A.The play pauses for manual intervention
B.The play retries the failed host
C.The play aborts immediately
D.The play continues with the remaining 4 hosts
E.The play marks the host as unreachable and continues
AnswerC

With `serial: 1`, each host forms its own batch, so the first host's failure leaves a 100% batch failure rate. Since `max_fail_percentage: 0` permits no failures, Ansible aborts the play immediately, preventing any remaining hosts from being updated.

Why this answer

With 'serial: 1', Ansible updates one host at a time. 'max_fail_percentage: 0' means that if any host fails (0% failure tolerance), the entire playbook run is aborted immediately. When the first host fails, Ansible stops further execution because the failure percentage exceeds the threshold, and no retries or continuation occur.

Exam trap

The trap here is that candidates often assume 'serial: 1' means the play will skip the failed host and continue with the next, but 'max_fail_percentage: 0' overrides that by aborting on any failure.

How to eliminate wrong answers

Option A is wrong because Ansible does not pause for manual intervention by default; it aborts based on max_fail_percentage. Option B is wrong because Ansible does not automatically retry failed hosts unless a retry mechanism (like 'until' or 'retries') is explicitly configured, which is not the case here. Option D is wrong because 'max_fail_percentage: 0' prevents continuation after any failure; the play does not proceed to remaining hosts.

Option E is wrong because marking a host as unreachable is a separate behavior (e.g., when connection fails), but here the host fails during the update, and the failure percentage triggers an abort, not an unreachable status.

14
Multi-Selecthard

Which TWO of the following are best practices when coordinating rolling updates with Ansible?

Select 2 answers
A.Define a 'max_fail_percentage' to abort the update if too many hosts fail.
B.Use the 'serial' keyword to update a subset of hosts at a time.
C.Use 'strategy: free' to allow hosts to run tasks independently.
D.Use 'gather_facts: no' to speed up the playbook.
E.Set 'any_errors_fatal: true' to stop the update on the first failure.
AnswersA, B

Setting `max_fail_percentage` halts the play once failed hosts exceed the threshold, preventing a faulty update from cascading across the remaining batch. This directly satisfies the stem's coordination requirement by bounding blast radius during rolling updates, rather than letting Ansible continue through every host regardless of accumulating failures.

Why this answer

Option B is correct because the 'serial' keyword controls how many hosts are targeted per play iteration, which is the fundamental mechanism for performing a rolling update in Ansible—updating a small batch at a time rather than all hosts at once. Option A is correct because 'max_fail_percentage' works together with 'serial' to abort the entire play if the number of failed hosts in a batch exceeds the defined threshold, preventing a bad rollout from cascading across the fleet. Option C is not a best practice for rolling updates because 'strategy: free' lets each host proceed through tasks independently without waiting for others, which breaks the controlled batch-by-batch ordering that rolling updates require.

Option D is not inherently a rolling-update best practice; disabling fact gathering may speed execution but does not coordinate or sequence updates and can break plays that depend on facts. Option E is not appropriate here because 'any_errors_fatal: true' aborts the whole play on the very first host failure, which is more aggressive than the graduated tolerance provided by 'max_fail_percentage' and is not the recommended pairing for controlled rolling updates.

Exam trap

The trap here is that candidates often confuse 'any_errors_fatal' (which stops on the first failure globally) with 'max_fail_percentage' (which aborts only after a threshold of failures in a batch), leading them to select option E instead of A.

15
MCQeasy

An administrator is rolling out a configuration change to a fleet of 12 application servers. The change must be applied to one server at a time so that the load balancer always has 11 healthy backends. Which playbook directive guarantees this behavior?

A.throttle: 1
B.serial: 1
C.order: sequential
D.forks: 1
AnswerB

serial: 1 instructs Ansible to process exactly one host per batch, so the play runs completely on one server before moving to the next. On a 12-node fleet, that leaves 11 servers untouched and available to the load balancer at all times. This directly satisfies the requirement of updating one server at a time without taking down capacity.

Why this answer

serial defines how many hosts Ansible includes in each rolling batch. Setting serial: 1 ensures the entire play finishes on one host before the next host begins, which keeps 11 of 12 servers available to the load balancer. Other keywords such as throttle and forks affect parallelism or task concurrency, but they do not create per-host rolling batches.

Exam trap

The trap here is confusing task-level concurrency controls like throttle or forks with the play-level batch size set by serial.

16
MCQeasy

An administrator notices that during a rolling update, the playbook seems to hang after updating the first host. The playbook uses serial: 5. What is the most likely cause?

A.The playbook has an infinite loop.
B.One of the hosts in the batch is taking too long to complete its tasks.
C.The SSH control path is exhausted.
D.The max_fail_percentage is set too high.
AnswerB

With `serial: 5`, Ansible waits for every host in the current batch to finish all tasks before starting the next batch. A single slow host therefore blocks the remaining four, making the playbook appear to hang after the first host completes. The constraint is batch-wide synchronisation, not task failure.

Why this answer

When `serial: 5` is set, Ansible processes hosts in batches of five. If one host in the batch takes an unusually long time to complete its tasks (e.g., due to a slow network, a hanging service restart, or a long-running command), the entire batch will appear to hang because Ansible waits for all hosts in the current batch to finish before proceeding to the next batch. This is the most likely cause of the observed behavior during a rolling update.

Exam trap

Red Hat often tests the misconception that `serial` controls parallelism across all hosts (like `forks`), but the trap here is that `serial` batches hosts sequentially, so a single slow host in a batch blocks the entire batch from completing, causing the playbook to appear to hang.

How to eliminate wrong answers

Option A is wrong because an infinite loop would cause the playbook to run indefinitely on a single host, not hang after updating the first host in a batch; the playbook would continue looping on that host without progressing. Option C is wrong because SSH control path exhaustion would typically manifest as SSH connection failures or errors, not a hang after the first host completes; it is a connection pooling issue, not a batch processing delay. Option D is wrong because `max_fail_percentage` controls how many hosts can fail before Ansible aborts the playbook; a high value would allow more failures, not cause a hang, and it does not affect the timing of task completion within a batch.

17
MCQeasy

A team uses Ansible to update a web application across 10 servers with minimal downtime. Which playbook directive achieves one-at-a-time updates?

A.run_once: true
B.delegate_to: localhost
C.serial: 1
D.throttle: 1
E.forks: 10
AnswerC

Setting serial to 1 makes Ansible run the play against a single host per batch, so only one of the ten servers restarts at a time. Remaining hosts keep serving traffic, satisfying the minimal-downtime constraint.

Why this answer

C is correct because the `serial: 1` directive in an Ansible playbook controls the number of hosts that are updated simultaneously. Setting `serial: 1` forces Ansible to execute the playbook on one host at a time, ensuring that the web application is updated sequentially across the 10 servers, which minimizes downtime by keeping the other 9 servers available during each individual update.

Exam trap

The trap here is that candidates confuse `serial` with `forks` or `throttle`, mistakenly thinking that limiting parallel connections (`forks: 1`) or task concurrency (`throttle: 1`) achieves the same sequential host behavior as `serial`, but only `serial` controls the batch size of hosts processed by the playbook.

How to eliminate wrong answers

Option A is wrong because `run_once: true` executes a task on only one host in the batch, not sequentially across all hosts, and is typically used for one-time setup tasks like generating a shared secret. Option B is wrong because `delegate_to: localhost` runs a task on the Ansible control node instead of the target servers, which does not control the order or batch size of host updates. Option D is wrong because `throttle: 1` limits the number of concurrent forks for a specific task but does not enforce sequential host processing across the entire play; it can still allow parallel execution of other tasks.

Option E is wrong because `forks: 10` sets the maximum number of parallel connections Ansible can make, but it does not guarantee one-at-a-time updates; with 10 forks, Ansible could attempt to update all 10 servers simultaneously.

18
MCQmedium

You are performing a rolling update of a 12-node web server fleet managed by Ansible. The playbook uses `serial: 4`. During the second batch, the task `Restart httpd` fails on one host because the service name is misspelled. The playbook aborts with an error. You fix the typo and rerun the playbook. What is the default behavior regarding the hosts that were already updated successfully in the first batch?

A.The playbook re-runs all tasks on every host, including the first batch, because Ansible is idempotent and will simply reapply the configuration.
B.The playbook fails immediately because the inventory still marks the failed host as unreachable, blocking any further execution.
C.The playbook starts over from the first batch and processes all hosts again, applying tasks to hosts that were already updated.
D.The playbook resumes from the failed batch, skipping the first batch entirely, because Ansible tracks successful hosts in a fact cache.
AnswerC

By default, Ansible targets all hosts in the inventory when you rerun a playbook. It does not remember which hosts succeeded in a previous run. Therefore, the first batch will be processed again, and tasks will be reapplied. This is why idempotent tasks and careful use of `serial` with external tracking are important in rolling updates.

Why this answer

Ansible does not persist playbook progress between runs. When you rerun a playbook, it targets the entire inventory by default, so hosts from the first batch are processed again. To avoid re-updating already-successful hosts, you would need to use `--limit` with a list of remaining hosts or implement a custom serial strategy that records completed batches externally.

Exam trap

The trap here is assuming Ansible remembers which hosts succeeded and automatically resumes from the failed batch.

19
MCQeasy

A junior administrator needs to perform a rolling update of 8 web servers where exactly 2 servers are updated at a time. Which play-level keyword and value should be used in the Ansible playbook?

A.throttle: 2
B.max_fail_percentage: 2
C.forks: 2
D.serial: 2
AnswerD

serial: 2 instructs Ansible to process hosts in batches of two during the play. Each batch completes the full task list before the next batch begins, so exactly two servers are updated at a time, which matches the administrator's requirement for this 8-node fleet.

Why this answer

The serial keyword is the standard way to define rolling update batch size in an Ansible play. Setting serial: 2 causes the play to process two hosts at a time through the full task list, ensuring only two of the eight web servers are updated concurrently. This is the direct and correct control for the administrator's requirement.

Exam trap

The trap here is mixing up forks, which controls parallel connections, with serial, which controls rolling batch size.

20
MCQhard

An operations team is designing a rolling update for a stateful application that requires quorum (minimum 3 out of 5 nodes online). They plan to use Ansible's serial keyword. Which serial value ensures the update proceeds without breaking quorum while still being efficient?

A.serial: 2
B.serial: 1
C.serial: 3
D.serial: 5
AnswerA

Serial 2 updates two nodes at a time, leaving three of five online, which preserves the quorum minimum throughout the rolling update. Larger values risk dropping below three; serial 1 is safe but slower, so 2 balances safety with efficiency.

Why this answer

Setting serial: 2 ensures that only 2 nodes are taken down at a time during the rolling update. With a quorum requirement of 3 out of 5 nodes, taking down 2 nodes leaves 3 online, maintaining quorum. This is the most efficient value that does not risk breaking quorum.

Exam trap

The trap here is that candidates may confuse 'quorum' with 'majority' and incorrectly choose serial: 3, thinking that 3 out of 5 is a majority, but fail to realize that taking down 3 nodes leaves only 2 online, which is below the quorum threshold of 3.

How to eliminate wrong answers

Option B is wrong because serial: 1 would take down only 1 node at a time, which is safe but less efficient than serial: 2 since it increases the total update time. Option C is wrong because serial: 3 would take down 3 nodes at once, leaving only 2 online, which breaks the quorum requirement of 3 out of 5 nodes. Option D is wrong because serial: 5 would take down all 5 nodes simultaneously, completely breaking quorum and causing the application to fail.

21
MCQeasy

An administrator wants to update a web server fleet with minimal downtime. They need to update each server one at a time. Which Ansible playbook directive should be used?

A.throttle: 1
B.forks: 1
C.serial: 1
D.max_fail_percentage: 0
AnswerC

The serial directive controls how many hosts Ansible targets per batch. Setting serial: 1 processes each server individually, ensuring only one host is updated at a time, which satisfies the requirement for minimal downtime across the fleet.

Why this answer

The `serial: 1` directive in an Ansible playbook controls the batch size of hosts that are updated simultaneously. Setting it to 1 ensures that only one host is updated at a time, which minimizes downtime by allowing the rest of the fleet to remain available while each server is sequentially updated.

Exam trap

The trap here is that candidates often confuse `serial` with `forks` or `throttle`, mistakenly thinking that limiting parallel task execution (`forks: 1`) or task concurrency (`throttle: 1`) achieves the same one-at-a-time host update behavior, but only `serial` controls the batch size of hosts processed sequentially.

How to eliminate wrong answers

Option A is wrong because `throttle: 1` limits the number of concurrent tasks per host or per play, but it does not control the batch size of hosts being updated; it limits task concurrency, not the sequential update of hosts. Option B is wrong because `forks: 1` sets the number of parallel processes Ansible uses to execute tasks on hosts, but it still allows all hosts in the batch to be processed in parallel; it does not enforce a one-at-a-time update across the entire fleet. Option D is wrong because `max_fail_percentage: 0` defines the maximum percentage of hosts that can fail before the playbook aborts, but it does not control the order or batch size of updates; it is a failure threshold, not a sequencing mechanism.

22
MCQeasy

You are running an Ansible playbook with `serial: 2` to update a fleet of 6 web servers. The playbook includes a task that restarts the web service. After the first batch of 2 hosts is updated, you notice that both hosts are restarted simultaneously. You want to ensure that within each batch, the hosts are updated one at a time to avoid a temporary loss of capacity. Which Ansible keyword should you add to the play to achieve this?

A.`throttle: 1` on the restart task
B.`order: inventory` at the play level
C.`serial: 1` at the play level
D.`strategy: linear` at the play level
AnswerA

The `throttle` keyword limits the number of workers that can execute a task simultaneously. Setting `throttle: 1` on the restart task ensures that only one host in the batch restarts the service at a time, even though both hosts are in the same batch. This prevents simultaneous restarts and maintains capacity.

Why this answer

The `throttle` keyword limits the number of concurrent executions of a task across all hosts. By setting `throttle: 1` on the restart task, you ensure that only one host restarts the web service at a time, even within a batch of 2. This maintains the rolling update's goal of minimizing downtime while keeping the batch size efficient.

Exam trap

The trap here is thinking that `serial: 1` is the only way to serialize updates, or that `strategy: linear` serializes tasks. `serial` controls batch size, while `throttle` controls concurrency of individual tasks within those batches.

23
MCQhard

An OpenShift rolling update is failing because new pods crash immediately. Which parameter automatically triggers a rollback if no progress is made?

A.revisionHistoryLimit
B.maxSurge
C.progressDeadlineSeconds
D.maxUnavailable
E.minReadySeconds
AnswerC

If the deployment does not progress within this time, it is considered failed and rolls back.

Why this answer

The `progressDeadlineSeconds` parameter specifies the maximum duration (in seconds) that a deployment can make no progress before it is considered to have failed. When this deadline is exceeded, the deployment controller automatically triggers a rollback to the previous revision. This is the correct parameter for automatically rolling back a failed rolling update where new pods crash immediately.

Exam trap

The trap here is that candidates confuse `progressDeadlineSeconds` with `minReadySeconds`, thinking that a readiness check alone will trigger a rollback, but `minReadySeconds` only delays availability without initiating a rollback.

How to eliminate wrong answers

Option A is wrong because `revisionHistoryLimit` controls how many old ReplicaSets are retained for rollback, not the timing or automatic rollback trigger. Option B is wrong because `maxSurge` defines the maximum number of pods that can be created above the desired replica count during an update, not a rollback mechanism. Option D is wrong because `maxUnavailable` specifies the maximum number of pods that can be unavailable during the update process, not a progress deadline.

Option E is wrong because `minReadySeconds` determines how long a pod must be ready before it is considered available, but it does not trigger a rollback if no progress is made.

24
Multi-Selectmedium

An administrator needs to update a web application that runs as a Kubernetes Deployment with 5 replicas. The application is stateless, but the update must not cause any downtime. Which TWO strategies ensure zero-downtime rolling updates?

Select 2 answers
A.Omit the liveness probe from the pod spec.
B.Set strategy type to RollingUpdate with maxUnavailable=0 and maxSurge=1.
C.Set maxUnavailable=1 and maxSurge=0.
D.Use the Recreate strategy.
E.Configure a readiness probe that checks the application's health endpoint.
AnswersB, E

maxUnavailable=0 guarantees all five existing pods stay ready during the rollout, while maxSurge=1 allows one extra pod to be created first. New pods must pass readiness before old ones terminate, preserving continuous availability throughout the update.

Why this answer

Option B is correct because a RollingUpdate strategy with maxUnavailable=0 and maxSurge=1 guarantees that no existing pod is terminated until a new pod is fully available, so the Deployment always keeps all 5 replicas serving traffic during the update. Option E is correct because a readiness probe tied to the application's health endpoint ensures Kubernetes only routes traffic to new pods once they are actually ready, preventing requests from hitting uninitialized instances during the rollout. Together, maxUnavailable=0/maxSurge=1 plus a readiness probe are the standard mechanism for zero-downtime updates of a stateless Deployment.

Option A is wrong because omitting a liveness probe does not prevent downtime and removes Kubernetes' ability to restart unhealthy containers. Option C is wrong because maxUnavailable=1 allows a pod to be taken out of service before its replacement is ready, reducing capacity and risking dropped requests. Option D is wrong because the Recreate strategy terminates all existing pods before creating new ones, causing guaranteed downtime.

Exam trap

The trap here is that candidates often confuse `maxUnavailable` and `maxSurge` values, mistakenly thinking that allowing one unavailable pod (maxUnavailable=1) is acceptable for zero-downtime, when in fact it can cause a temporary capacity deficit if the readiness probe is not fast enough.

25
Multi-Selecthard

Which THREE statements correctly describe the behavior of the 'serial' keyword in Ansible? (Choose exactly three.)

Select 3 answers
A.It can be set as a percentage of the total hosts.
B.It causes the playbook to run on a subset of hosts at a time.
C.It can be combined with max_fail_percentage to control failure thresholds.
D.It guarantees that only one task runs across all hosts at any time.
E.It applies globally to all plays in the playbook.
AnswersA, B, C

Ansible accepts serial as an integer or a percentage, so specifying 25% runs the play across a quarter of the matched hosts per batch. This satisfies the stem's requirement for a statement describing serial's percentage-based behaviour, distinct from fixed host counts.

Why this answer

Option A is correct because the serial keyword accepts a percentage value (e.g., serial: 25%) that Ansible interprets as a fraction of the total hosts in the play, batching them accordingly. Option B is correct because serial defines the number of hosts (or batch size) that Ansible targets per play iteration, so the play runs on a subset of hosts at a time rather than all at once. Option C is correct because serial is commonly paired with max_fail_percentage, which aborts the play if failures within a serial batch exceed the given threshold, enabling controlled rolling updates.

Option D is incorrect because serial controls host batching, not task concurrency; it does not guarantee only one task runs across all hosts at any time. Option E is incorrect because serial is set at the play level and applies only to that specific play, not globally to all plays in a playbook.

Exam trap

The trap here is that candidates often confuse 'serial' with a task-level concurrency control or assume it applies globally across all plays, when in fact it is a per-play batch size setting that controls how many hosts execute the entire play simultaneously.

26
MCQmedium

An Ansible Engineer is planning a rolling update for a web application deployed across 10 nodes. The playbook uses the 'delegate_to' directive to manage load balancer health checks. Which of the following best describes the recommended approach to minimize downtime?

A.Use 'serial: 1' and delegate load balancer disable/enable tasks to localhost, ensuring each node is taken out of rotation before updating.
B.Run the update playbook with 'serial: 10' to update all nodes at once, then run a separate playbook to update the load balancer.
C.Run the update on each node manually using 'ansible-playbook --limit' and skip load balancer management to save time.
D.Use 'strategy: free' to allow nodes to update independently without controlling the load balancer.
AnswerA

Using `serial: 1` updates one node at a time, so the remaining nine keep serving traffic. Delegating the load balancer disable and enable tasks to localhost runs them once from the controller rather than on each managed node, satisfying the rolling-update constraint of removing a node from rotation before patching and restoring it afterwards.

Why this answer

Using 'serial: 1' ensures that only one node is updated at a time, and delegating load balancer disable/enable tasks to localhost (or the Ansible control node) allows the playbook to interact with the load balancer API to remove the node from the pool before the update and re-add it after. This minimizes downtime by ensuring traffic is not sent to a node being updated, while other nodes continue serving requests.

Exam trap

The trap here is that candidates may think 'serial: 10' is efficient because it updates all nodes quickly, but they overlook that it causes a full outage, whereas the correct approach prioritizes availability over speed.

How to eliminate wrong answers

Option B is wrong because 'serial: 10' updates all nodes simultaneously, which would cause a complete outage during the update window, defeating the purpose of a rolling update. Option C is wrong because manually running with '--limit' and skipping load balancer management does not automate the process and leaves nodes in the load balancer pool while they are being updated, causing traffic to be sent to an unavailable node and increasing downtime. Option D is wrong because 'strategy: free' allows nodes to run tasks independently without any serialization or load balancer coordination, leading to potential race conditions and no guarantee of minimizing downtime.

27
Multi-Selecteasy

Which TWO options are best practices for coordinating rolling updates with Ansible? (Choose exactly two.)

Select 2 answers
A.Set ignore_errors: yes to ensure the playbook continues even if some hosts fail.
B.Use the serial keyword to update hosts in batches.
C.Use the default serial setting (all hosts) for simplicity.
D.Set max_fail_percentage to limit the number of failed hosts before aborting.
E.Run all hosts in parallel to minimize total update time.
AnswersB, D

The serial keyword partitions the play's host list into batches, so each batch completes before the next begins. This limits blast radius during rolling updates, satisfying the requirement to update hosts incrementally rather than all at once.

Why this answer

Option B is correct because the serial keyword in Ansible controls how many hosts are targeted per play iteration, allowing rolling updates in controlled batches (e.g., serial: 2 or serial: 25%) so that only a subset of hosts is updated at a time while the rest continue serving traffic. Option D is correct because max_fail_percentage defines a failure threshold within a serial batch; if the percentage of failed hosts exceeds that value, Ansible aborts the play, preventing a bad update from cascading across the entire fleet. Option A is not a best practice for rolling updates because ignore_errors: yes masks failures and lets the playbook proceed even when hosts are broken, defeating the safety purpose of batching.

Option C is wrong because the default serial value is effectively all hosts in the play, which performs a simultaneous update rather than a rolling one. Option E is also wrong because running all hosts in parallel maximizes blast radius and downtime risk, the opposite of a rolling-update strategy.

Exam trap

The trap here is that candidates often confuse `ignore_errors` with error handling for rolling updates, not realizing that it bypasses failure detection, whereas `max_fail_percentage` is the correct way to control abort behavior during batch updates.

28
MCQmedium

An Ansible rolling update playbook includes 'max_fail_percentage: 20'. If more than 20% of hosts fail during any batch, what happens?

A.The play pauses and waits for user input
B.The failed hosts are removed from inventory
C.The play retries failed hosts
D.The play aborts immediately
E.The play continues with remaining hosts
AnswerD

max_fail_percentage sets a tolerance threshold; once failed hosts in a batch exceed 20%, Ansible halts the play for all remaining hosts rather than continuing. This aborts the rolling update immediately, preventing further hosts from being updated.

Why this answer

The `max_fail_percentage` parameter in Ansible's rolling update strategy defines the maximum percentage of hosts that can fail in a single batch before the playbook aborts entirely. When the failure rate exceeds this threshold, Ansible stops execution immediately to prevent cascading failures or inconsistent state across the remaining hosts.

Exam trap

The trap here is that candidates often confuse `max_fail_percentage` with `any_errors_fatal` or assume the play will simply skip failed hosts and continue, but Ansible strictly aborts the entire play when the threshold is exceeded to enforce safety limits.

How to eliminate wrong answers

Option A is wrong because `max_fail_percentage` does not pause the play for user input; that behavior is controlled by `serial` with `pause` or `wait_for` tasks, not by failure thresholds. Option B is wrong because failed hosts are not removed from inventory; Ansible does not modify inventory files dynamically based on playbook failures. Option C is wrong because `max_fail_percentage` does not trigger automatic retries; retry behavior is configured separately via `retries` and `until` on individual tasks or via `any_errors_fatal`.

Option E is wrong because the play does not continue with remaining hosts when the failure percentage is exceeded; instead, it aborts immediately to prevent further execution.

29
Multi-Selectmedium

An administrator is designing a rolling update playbook for a 20-node application cluster. The playbook must limit the blast radius of failures and allow the rollout to pause safely if too many hosts fail. Which two Ansible play-level keywords directly control how many hosts are updated at once and whether the play aborts based on failures? (Choose two.)

Select 2 answers
A.forks
B.serial
C.order
D.any_errors_fatal
E.max_fail_percentage
AnswersB, E

serial defines the number or percentage of hosts processed per batch during a rolling update. In this scenario it directly controls how many of the 20 nodes are updated at one time, limiting the blast radius and ensuring the play progresses in controlled groups rather than all at once.

Why this answer

serial controls how many hosts are processed in each rolling batch, directly limiting the blast radius. max_fail_percentage defines the failure threshold per batch that aborts the play, preventing the rollout from continuing when too many hosts fail. Together they provide the two controls the administrator needs for a safe rolling update.

Exam trap

The trap here is confusing parallelism controls like forks with rolling-update controls like serial and max_fail_percentage.

30
MCQhard

In OpenShift, a deployment must gradually shift traffic to new pods during a rolling update. Which default strategy achieves this?

A.Blue-green deployment
B.RollingUpdate
C.Canary deployment
D.Custom strategy
E.Recreate
AnswerB

RollingUpdate is the Deployment strategy that replaces pods incrementally, keeping a configurable maxUnavailable and maxSurge so traffic shifts gradually to new pods while old ones terminate. This satisfies the stem's requirement for a gradual traffic shift during update, unlike Recreate, which stops all pods first.

Why this answer

The RollingUpdate strategy is the default deployment strategy in OpenShift (and Kubernetes) that gradually replaces old pods with new ones while maintaining application availability. It achieves this by incrementally scaling down the old ReplicaSet and scaling up the new one, controlled by parameters like maxSurge and maxUnavailable, ensuring a smooth transition of traffic to new pods.

Exam trap

The trap here is that candidates often confuse the default OpenShift deployment strategy with advanced deployment patterns like blue-green or canary, which are not built-in defaults but require additional configuration, leading them to select those incorrect options.

How to eliminate wrong answers

Option A is wrong because blue-green deployment is not a default strategy in OpenShift; it requires manual configuration or additional tooling (e.g., Routes and Services) to switch traffic between two separate environments. Option C is wrong because canary deployment is not a built-in default strategy; it involves routing a small percentage of traffic to new pods via custom configurations or service mesh features like Istio, not a native OpenShift deployment strategy. Option D is wrong because 'Custom strategy' is not a valid default deployment strategy in OpenShift; the platform only supports RollingUpdate and Recreate as native strategies, with custom behavior achievable via parameters but not as a named default.

Option E is wrong because Recreate is an alternative strategy that terminates all old pods before creating new ones, causing downtime, and is not the default for gradual traffic shifting.

31
MCQmedium

An Ansible playbook sets 'serial: 20%' for rolling updates, but the inventory contains 5 hosts. How many hosts are updated simultaneously?

A.1
B.2
C.3
D.0
E.5
AnswerA

Serial accepts a percentage, which Ansible converts to a host count by rounding up fractional results. Twenty percent of five hosts equals one, so a single host is updated per batch, giving strictly sequential updates across the inventory.

Why this answer

When 'serial: 20%' is set in an Ansible playbook, the percentage is calculated based on the total number of hosts in the inventory. With 5 hosts, 20% of 5 equals 1.0, which is rounded down to 1. Therefore, only 1 host is updated at a time during the rolling update.

Exam trap

The trap here is that candidates often assume percentages are rounded up or that a fractional result like 1.0 would be treated as 2, but Ansible uses floor rounding (truncation) for serial batch sizes, and with exactly 1.0, the result is 1, not 2.

How to eliminate wrong answers

Option B is wrong because 2 would represent 40% of 5 hosts, not 20%. Option C is wrong because 3 would be 60% of the inventory, far exceeding the 20% specification. Option D is wrong because 0 would only occur if the percentage rounded down to zero (e.g., less than 1 host), but 20% of 5 is exactly 1.0, which rounds to 1.

Option E is wrong because 5 would represent 100% of the hosts, which would be a serial value of '100%' or 'serial: 5', not '20%'.

32
MCQhard

You are designing a rolling update for a stateful service where each node must be removed from a load balancer, updated, and re-added before the next node is touched. The update must never take more than one node offline at a time, and if the update fails on a node, the play must stop and leave the remaining nodes untouched. Which playbook configuration achieves this?

A.serial: 1 with ignore_errors: true on the update task so the play can continue to the next node.
B.serial: 1 with any_errors_fatal: true and tasks ordered to drain, update, and re-add within the same play.
C.serial: 2 with max_fail_percentage: 50 so that one failure in a batch is tolerated and the play continues.
D.serial: 1 with run_once: true on the drain and re-add tasks so they execute only once per batch.
AnswerB

serial: 1 processes one host per batch, guaranteeing that only one node is offline at a time. any_errors_fatal: true ensures that a failure on that node stops the entire play, leaving all remaining nodes untouched. Ordering drain, update, and re-add within the same play keeps the node out of rotation only during its own update.

Why this answer

Ensuring only one node is offline at a time and halting on failure requires a batch size of one and play-level fatal error handling. serial: 1 isolates each node, while any_errors_fatal: true stops the play immediately on any failure, leaving the remaining nodes untouched and still in rotation.

Exam trap

The trap here is thinking that ignore_errors or a failure percentage provides safety, when both allow the rollout to continue and risk taking additional nodes offline after a failure.

Ready to test yourself?

Try a timed practice session using only Coordinate rolling updates questions.