NCP-AIO Workload Management Practice Question
An administrator manages a Kubernetes cluster where a training job repeatedly fails with an OutOfMemory error on the GPU even though the pod requests one nvidia.com/gpu. DCGM metrics show that another pod on the same node is consuming GPU memory concurrently. GPU sharing via time-slicing is enabled cluster-wide. Which action should the administrator take to prevent this cross-pod interference while preserving the ability to share GPUs among trusted inference workloads?
⚠ Common exam trap
The trap here is thinking that increasing GPU requests or adding quotas creates memory isolation, when only dedicated devices or MIG instances actually partition GPU memory.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Disable time-slicing on the node pool hosting training jobs and use MIG or dedicated GPUs for those workloads
Time-slicing multiplexes a GPU across pods without memory partitioning, so one workload can starve another and trigger OutOfMemory errors. The reliable fix is to separate untrusted or memory-intensive training workloads onto dedicated GPUs or MIG instances, while leaving time-slicing for inference workloads where memory contention is acceptable and controlled.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Apply a ResourceQuota limiting nvidia.com/gpu to one per namespace
Why it's wrong here
A ResourceQuota restricts aggregate namespace consumption but does not isolate memory between pods already scheduled onto the same physical GPU. Two pods in different namespaces could still share a time-sliced GPU and contend for memory. The quota also would not prevent the scheduler from placing the training pod alongside an inference pod, so the interference remains.
- ✗
Set the CUDA_VISIBLE_DEVICES environment variable manually in the training pod spec
Why it's wrong here
Manually setting CUDA_VISIBLE_DEVICES bypasses the device plugin's allocation logic and can conflict with the kubelet's injected value. It also does not create memory isolation; another pod on the same time-sliced GPU can still consume memory. This approach risks scheduling errors and does not address the root cause of cross-pod GPU memory contention.
- ✓
Disable time-slicing on the node pool hosting training jobs and use MIG or dedicated GPUs for those workloads
Why this is correct
Time-slicing provides no memory isolation, so concurrent pods can exhaust GPU memory and cause OutOfMemory failures. Removing training jobs from time-sliced nodes and giving them dedicated GPUs or MIG instances ensures exclusive memory access while still allowing time-slicing to remain available for trusted inference workloads on separate node pools.
- ✗
Increase the pod's nvidia.com/gpu request to two GPUs
Why it's wrong here
Requesting a second GPU does not prevent another pod from sharing the first GPU when time-slicing is enabled, because the device plugin still advertises replicas of each physical device. The training pod may receive two virtual GPU replicas that map to shared physical devices, so the concurrent memory consumption problem persists rather than being resolved.
About these practice questions
This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.