NCP-AIO Workload Management Practice Question
An AI operations team runs long-running training jobs on a Kubernetes cluster with NVIDIA GPU Operator. They observe that after a node is rebooted for maintenance, some pods resume but report CUDA 'unknown error' and the device plugin shows unhealthy GPUs. Which configuration should the administrator review to ensure the driver and device plugin recover cleanly after reboot?
⚠ Common exam trap
The trap here is blaming the application container or its CUDA version for a fault that originates in post-reboot node initialization by the GPU Operator.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The GPU Operator's driver validation and node reboot handling settings, including the driver DaemonSet and readiness gates.
A reboot invalidates the previously loaded driver and device plugin state, so the GPU Operator must reload the driver and re-register devices before pods can use them. Driver validation and readiness gates hold workloads until health checks pass, preventing CUDA errors from reaching applications. Reviewing these settings ensures clean recovery and avoids the unhealthy plugin state observed after maintenance.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The pod's restartPolicy set to Always so containers restart automatically after node recovery.
Why it's wrong here
restartPolicy governs container restarts within a pod, not driver or device plugin recovery on the node. A container can restart successfully yet still fail CUDA calls if the driver is not fully loaded. This option misattributes a node-level initialization problem to pod lifecycle policy and will not correct the unhealthy GPU state.
- ✓
The GPU Operator's driver validation and node reboot handling settings, including the driver DaemonSet and readiness gates.
Why this is correct
After a reboot, the GPU Operator must reload the driver and reinitialize the device plugin before workloads can use GPUs safely. Driver validation and readiness gates prevent pods from starting until the driver and plugin report healthy. Reviewing these settings addresses the CUDA error and unhealthy device plugin state observed after maintenance reboots.
- ✗
The cluster autoscaler settings so replacement nodes are provisioned immediately after reboot.
Why it's wrong here
Autoscaling adds or removes nodes based on demand and does not govern driver initialization on an existing rebooted node. Provisioning a replacement node may mask the problem temporarily but leaves the root cause unresolved. This option does not address why the device plugin reports unhealthy GPUs or why CUDA errors appear after the reboot.
- ✗
The container image's CUDA version to ensure it matches the host driver version.
Why it's wrong here
CUDA version compatibility matters, but the scenario describes a state change caused by a reboot, not a persistent version mismatch that worked before. If the image and driver were compatible before maintenance, version alignment is not the variable that changed. This option distracts from the post-reboot initialization sequence that actually determines device plugin health.
About these practice questions
One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.