NCP-AIO Workload Management Practice Question
An AI operations team runs a shared Kubernetes cluster with the NVIDIA GPU Operator and several namespaces owned by different groups. A group reports that its training pods are stuck Pending with an event indicating insufficient nvidia.com/gpu, yet cluster-wide dashboards show many GPUs idle. Investigation reveals the idle GPUs belong to nodes labeled for another group, and the affected namespace has a node affinity rule pinning its pods to a specific GPU generation that is fully consumed. Which action best resolves the Pending pods while respecting multi-tenant boundaries?
⚠ Common exam trap
The trap here is treating idle cluster-wide GPUs as available capacity, when node affinity can make those GPUs ineligible for the pending pods.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Provision additional GPUs of the required generation for that tenant or rebalance existing capacity to that node pool.
The pods are constrained by a node affinity rule targeting a specific GPU generation whose nodes are fully allocated. Idle GPUs on other nodes cannot satisfy that rule. The correct fix is to add capacity of the required generation or rebalance matching cards into that pool, which resolves the shortage while preserving the tenant isolation the affinity rule enforces.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Disable the device plugin on the other tenants' nodes so their GPUs become available to the affected namespace.
Why it's wrong here
Disabling the device plugin would make those GPUs unschedulable for everyone, including their rightful tenants, and would disrupt running workloads. It does not help the affected pods because those nodes are still excluded by the affinity rule. This destructive action harms other tenants and fails to address the generation mismatch causing the Pending state.
- ✗
Lower the GPU request in the pod spec so the scheduler accepts a partial device from an idle node.
Why it's wrong here
Kubernetes does not support fractional nvidia.com/gpu requests without a sharing mode such as MIG or time-slicing. Changing the request to a fraction would be invalid or ineffective, and it would not move the pods onto the idle nodes because the affinity rule still restricts placement. This misdiagnoses both the resource model and the scheduling constraint.
- ✗
Remove the node affinity rule so the pods can schedule onto any idle GPU node in the cluster.
Why it's wrong here
Deleting the affinity rule would let the pods land on other groups' nodes, violating the tenant boundaries the cluster was designed to enforce. It also ignores why the rule existed, which is likely to keep workloads on compatible hardware. The pods might start, but the change undermines isolation and could cause cross-tenant interference, so it is not the best resolution.
- ✓
Provision additional GPUs of the required generation for that tenant or rebalance existing capacity to that node pool.
Why this is correct
The pods are Pending because the specific GPU generation they require is exhausted, not because the cluster lacks GPUs entirely. Adding capacity to that node pool, or moving idle cards of the right generation into it, satisfies the affinity constraint while keeping tenants separated. This addresses the real bottleneck without weakening isolation rules.
Visual reference
About these practice questions
This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.