NCP-AIO Workload Management Practice Question
An administrator is tuning a Kubernetes cluster that runs GPU Operator. Users report that GPU jobs are sometimes scheduled onto nodes whose drivers are older than the CUDA version the container needs, causing runtime failures. The administrator wants to prevent incompatible placements before pods are bound. (Choose two.)
⚠ Common exam trap
The trap here is expecting the device plugin or container runtime to reconcile CUDA and driver versions, when that compatibility decision belongs to scheduling and admission, not device allocation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Advertise the driver and CUDA versions as node labels through the GPU Operator's node feature discovery, then use nodeAffinity on the pods to require a compatible version.
Preventing incompatible placement requires exposing driver and CUDA versions as schedulable node attributes and then constraining pods against them. Node feature discovery under the GPU Operator publishes version labels, and nodeAffinity lets pods require a compatible generation. A validating admission policy adds a second layer by rejecting pods whose declared CUDA needs exceed the target node's advertised driver, catching mistakes before binding and avoiding the runtime failures users report.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Rely on the device plugin to compare the container's CUDA version against the node driver and refuse to allocate the GPU when they mismatch.
Why it's wrong here
The device plugin allocates GPU devices and advertises counts; it does not parse container images or compare CUDA versions against the host driver. Its responsibility ends at handing out device nodes and health checks. Assuming it performs version negotiation misattributes a scheduling concern to the allocation layer, so mismatched pods still bind and fail later during CUDA initialization.
- ✗
Enable time-slicing with a replica count high enough that every node advertises spare GPU capacity, so incompatible nodes are simply bypassed.
Why it's wrong here
Time-slicing multiplies advertised GPU count but says nothing about driver or CUDA compatibility. A node with an old driver still advertises capacity and will happily accept a pod whose CUDA version exceeds what its driver supports, producing the same runtime failure. Replica count is unrelated to version matching, so this does not prevent incompatible placements at all.
- ✗
Set the container image's CUDA version to match the newest driver in the cluster and rely on the container runtime to downgrade the driver automatically.
Why it's wrong here
Container runtimes do not downgrade or upgrade host GPU drivers based on the image's CUDA version. CUDA within the container depends on the host driver, and a newer CUDA toolkit than the driver supports fails at initialization. Relying on automatic downgrade is a false assumption, so this approach leaves the incompatibility intact and does not prevent the scheduling of mismatched pods.
- ✓
Advertise the driver and CUDA versions as node labels through the GPU Operator's node feature discovery, then use nodeAffinity on the pods to require a compatible version.
Why this is correct
Node feature discovery running under the GPU Operator publishes labels such as nvidia.com/cuda.driver.major and minor versions. Pods can then express a nodeAffinity requirement matching the needed driver generation, so the scheduler only binds them to nodes whose advertised driver satisfies the container's CUDA expectation. This moves incompatibility detection to scheduling time, preventing the runtime failures users currently see.
- ✓
Use a validating admission policy that rejects pods requesting a CUDA version incompatible with the target node's advertised driver labels.
Why this is correct
A validating admission policy can inspect a pod's annotations or image metadata and compare it against the node's advertised driver labels, rejecting requests that would run on incompatible hardware. This enforces compatibility before binding rather than after a crash. Combined with version labels from node feature discovery, it gives a policy-driven guardrail that scales across the cluster without per-pod affinity rules.
About these practice questions
This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.