NCP-AIO Workload Management Practice Question
A data science team submits a PyTorch training job to a Kubernetes cluster managed by Run:ai. The job requests two GPUs but only one is allocated, and the second worker hangs waiting for a peer. Which Run:ai capability should the administrator verify is configured so the distributed job receives all requested GPUs atomically?
⚠ Common exam trap
The trap here is assuming that priority or affinity alone can prevent partial placement, when only all-or-nothing gang scheduling guarantees every rank starts together.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Gang scheduling, which ensures all pods in a distributed job are scheduled together or not at all.
Distributed training depends on every rank being present before collective operations begin; a single missing worker causes the rest to block. Run:ai's gang scheduling is the mechanism that admits the entire workload as a unit, so all requested GPUs are granted together. Confirming gang scheduling is enabled and applied to the job resolves the partial allocation that produced the hang.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
A higher priority class assigned to the training job so it preempts other workloads.
Why it's wrong here
Priority and preemption affect which workloads run first, not whether all pods of a single job are scheduled together. A high-priority job can still be partially placed if free GPUs are fragmented across nodes. Preemption may eventually free capacity, but it does not guarantee the all-or-nothing allocation that eliminates the peer-wait hang.
- ✓
Gang scheduling, which ensures all pods in a distributed job are scheduled together or not at all.
Why this is correct
Distributed training jobs require all workers to start together; partial allocation causes hangs because ranks wait for peers that never launch. Run:ai's gang scheduling treats the workload as an atomic unit, allocating all requested GPUs or leaving the job pending. Verifying this configuration addresses the symptom of one GPU allocated and a stalled peer directly.
- ✗
Node affinity rules that pin each worker to a specific GPU node by hostname.
Why it's wrong here
Node affinity constrains where pods can run but does not guarantee that all workers are admitted simultaneously. With affinity alone, the scheduler can still place one worker while another remains Pending due to insufficient GPUs elsewhere. This option does not provide the atomic allocation needed to prevent the deadlock described in the scenario.
- ✗
Enabling time-slicing on the GPU device plugin to increase the apparent GPU count.
Why it's wrong here
Time-slicing advertises more logical GPUs per device, but it does not synchronize scheduling of the pods in a distributed job. Workers could still start at different times and hang waiting for peers. This option increases sharing capacity rather than enforcing atomic allocation, so it does not resolve the described deadlock.
Visual reference
About these practice questions
This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.