NCP-AIO Workload Management Practice Question
A data science team submits a PyTorch distributed training job to a Kubernetes cluster with the NVIDIA GPU Operator installed. The job's pods repeatedly fail with a CUDA initialization error, while a simple `nvidia-smi` check inside an interactive pod on the same node succeeds. The administrator confirms the node's driver is healthy and the device plugin is advertising GPUs. Which configuration should the administrator verify first?
⚠ Common exam trap
The trap here is chasing driver or toolkit version mismatches when the real cause is that the workload never requested a GPU, so the device was never injected into the container.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
That each training pod includes a GPU resource request or limit so the device plugin injects the driver libraries and device nodes.
In an operator-managed cluster, GPU access is granted during admission and kubelet setup only when a pod requests an extended GPU resource; the device plugin and admission controller then inject the driver libraries, CUDA binaries, and device nodes. A pod without that request runs with no visible device, which is precisely why CUDA initialization fails while nvidia-smi succeeds in a different pod on the same node. The other options concern version skew, a specific capability, or a legacy runtime setting that would not produce this selective symptom.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
That the node's `/etc/docker/daemon.json` still lists `nvidia` as the default runtime for all containers.
Why it's wrong here
The GPU Operator typically manages the container runtime configuration itself and does not require the default runtime to be nvidia; in fact relying on a default runtime is an older pattern. Misconfiguration there would affect many workloads, not selectively fail one training job while an interactive pod on the same node initializes CUDA successfully.
- ✗
That the pods' securityContext drops the `IPC_LOCK` capability required for pinned host memory.
Why it's wrong here
Dropping IPC_LOCK affects pinned memory allocation used by some collective communication paths, but it produces a specific lock-related error, not a CUDA initialization failure. The device plugin's default behavior does not require this capability for basic CUDA context creation, so this is not the first thing to inspect for the symptom described.
- ✗
That the training pods' containers were built with a CUDA toolkit version newer than the node driver supports.
Why it's wrong here
A toolkit newer than the driver would typically surface as a runtime API error about an unsupported driver version rather than a generic initialization failure, and the operator normally injects a compatible driver library set. While version skew is worth checking, it does not explain why nvidia-smi succeeds on the same node while the training job cannot initialize CUDA.
- ✓
That each training pod includes a GPU resource request or limit so the device plugin injects the driver libraries and device nodes.
Why this is correct
The GPU Operator advertises devices through the device plugin, and the kubelet only mounts the driver libraries, device nodes, and CUDA binaries into a container that requests nvidia.com/gpu. Without that request the container starts with no GPU access, so CUDA initialization fails while nvidia-smi in a GPU-requesting pod on the same node works, matching the observed behavior exactly.
About these practice questions
This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.