NCP-AIO Workload Management Practice Question
An AI operations engineer manages a shared Kubernetes cluster running NVIDIA GPU Operator. Several teams report that their inference pods remain in a Pending state with the event '0/8 nodes are available: 8 Insufficient nvidia.com/gpu.' The administrator verifies that nvidia-smi on all nodes shows idle GPUs. Which action should the administrator take first to resolve the scheduling failure?
⚠ Common exam trap
The trap here is assuming that idle GPUs visible in nvidia-smi automatically become schedulable Kubernetes resources without the device plugin advertising them.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Inspect the node allocatable GPU count and any taints or labels that prevent scheduling on the idle nodes.
The Pending event shows the scheduler believes no node has a free nvidia.com/gpu resource, even though nvidia-smi reports idle devices. That gap is almost always caused by the device plugin not advertising GPUs, or by taints, labels, or affinity rules excluding the nodes. Inspecting allocatable counts and scheduling constraints identifies the actual blocker before any disruptive remediation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Reinstall the NVIDIA GPU Operator and restart the kubelet on all worker nodes.
Why it's wrong here
Reinstalling the GPU Operator and restarting kubelet is a disruptive, last-resort action that does not address the reported event. The scheduler message indicates a resource accounting mismatch, not a broken driver stack. Idle GPUs shown by nvidia-smi confirm the hardware and driver are functional, so this action would cause unnecessary downtime and likely leave the root cause intact.
- ✓
Inspect the node allocatable GPU count and any taints or labels that prevent scheduling on the idle nodes.
Why this is correct
The event 'Insufficient nvidia.com/gpu' means the scheduler sees zero or fewer allocatable GPUs than requested on every candidate node. This can result from missing device plugin advertisements, taints, node selectors, or nodes excluded by affinity. Checking allocatable resources and taints directly identifies why idle hardware is invisible or ineligible to the scheduler.
- ✗
Add a nodeSelector for kubernetes.io/os=linux to the pod template.
Why it's wrong here
A linux OS node selector is already satisfied by typical GPU worker nodes and does not explain why the scheduler reports insufficient GPUs on all eight nodes. This change would not alter GPU resource accounting and could even further restrict eligible nodes. It misdiagnoses a resource advertisement problem as an OS labeling problem.
- ✗
Increase the GPU memory limit in the pod specification so the scheduler can fit the workload.
Why it's wrong here
GPU memory is not a schedulable extended resource in Kubernetes; the scheduler only tracks the integer count of nvidia.com/gpu devices. Raising a memory limit does not change the 'Insufficient nvidia.com/gpu' event and cannot make a pod fit. This option confuses VRAM capacity with device count scheduling, so it will not resolve the Pending state.
Visual reference
About these practice questions
Courseiva writes every NCP-AIO question from scratch — 309 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.