Courseiva
Administration →mediumMultiple Choice

NCP-AIO Administration Practice Question

An administrator needs to ensure that all GPU drivers are updated across a heterogeneous cluster without causing downtime. What is the best strategy?

⚠ Common exam trap

Candidates often select disruptive cluster-wide reboots or simultaneous upgrades instead of controlled, node-by-node rolling updates with draining.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Perform rolling updates by draining nodes one by one

Implementing a rolling update strategy allows the cluster to be updated one node at a time while the remaining nodes continue to process training jobs. By draining the target node of all active jobs, upgrading the drivers, and then re-adding the node to the production pool, the administrator maintains service availability. This is the standard practice in production AI operations to minimize disruption and ensure consistent driver versions across the entire fleet.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Simultaneously update all nodes during a maintenance window

    Why it's wrong here

    Simultaneously updating all nodes requires significant downtime, which is unacceptable in production environments. This strategy increases the risk of a widespread failure if an issue occurs with the new driver version, making it impossible to perform a gradual, controlled rollout and verification process across the infrastructure.

  • ✓

    Perform rolling updates by draining nodes one by one

    Why this is correct

    Rolling updates ensure that the cluster remains operational throughout the entire upgrade process. By draining one node at a time, the administrator can safely update drivers without terminating active workloads prematurely. This approach provides a clear path for verification and rollback if the update encounters any unforeseen compatibility issues.

  • ✗

    Use a containerized driver approach to avoid host updates

    Why it's wrong here

    While containerized drivers can simplify deployment, they still rely on the host kernel. If the kernel version changes or requires an update to support new driver features, a host-level upgrade is unavoidable. This method does not eliminate the need for a formal maintenance strategy for the host OS and underlying drivers.

  • ✗

    Only update the driver on the master node

    Why it's wrong here

    Updating only the master node leaves the worker nodes running an older, potentially incompatible driver version. In a distributed training environment, this leads to version mismatch errors between the master and worker nodes, which can cause the entire training cluster to fail when launching collective communication operations.

About these practice questions

One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.