Courseiva
Administration →mediumMultiple Choice

NCP-AIO Administration Practice Question

An administrator supports a shared Kubernetes cluster running NVIDIA GPU Operator. Data scientists report that their inference pods remain in Pending state, yet the GPU Operator pods and node feature discovery pods are healthy, and the GPU nodes show no hardware alarms. The administrator confirms that the cluster has a mixture of MIG-capable A100 nodes and non-MIG T4 nodes. Which immediate administrative action is most appropriate to diagnose the scheduling failure?

⚠ Common exam trap

The trap here is assuming that healthy GPU Operator pods guarantee that GPU resources are correctly advertised and schedulable, when in fact a MIG strategy mismatch can leave pods Pending without any operator-level failure.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Inspect the pod events with kubectl describe pod and verify the MIG strategy configured in the device plugin, because a mismatch between the node's MIG configuration and the requested resource can prevent GPU resource advertisement.

The pod events are the fastest way to see whether the scheduler cannot find a suitable node due to resource requests. In a mixed MIG and non-MIG environment, the device plugin advertises different resource names, and a mismatch between the requested GPU resource and what the node exposes leaves pods Pending. Verifying the MIG strategy and node labels directly identifies that mismatch.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Inspect the pod events with kubectl describe pod and verify the MIG strategy configured in the device plugin, because a mismatch between the node's MIG configuration and the requested resource can prevent GPU resource advertisement.

    Why this is correct

    Pod events reveal whether the scheduler rejected the pod due to missing nvidia.com/gpu or nvidia.com/mig-* resources. If the device plugin advertises MIG profiles but the workload requests a full GPU (or vice versa), the pod stays Pending. Checking the MIG strategy and node labels directly addresses this common scheduling mismatch.

  • ✗

    Restart the kubelet on all worker nodes to force the device plugin to re-register GPU resources, because a stale kubelet cache is the most frequent cause of Pending workloads.

    Why it's wrong here

    Restarting kubelet is disruptive and unnecessary here. The GPU Operator and node feature discovery pods are healthy, indicating that kubelet is functioning. A stale cache would typically show node NotReady or plugin registration errors, neither of which is reported. This action risks evicting running workloads without addressing the actual scheduling constraint.

  • ✗

    Delete the pending pods and recreate them with a higher priority class, because the scheduler may be preempting them in favor of system pods.

    Why it's wrong here

    Priority preemption would show events indicating preemption, and system pods typically have lower or equal priority. Recreating pods with higher priority does not resolve a missing or mismatched GPU resource advertisement. This action may cause unintended preemption of other workloads and does not diagnose the root cause.

  • ✗

    Scale the GPU Operator controller manager to zero replicas and back to one, because the operator may have lost track of node labels after a recent upgrade.

    Why it's wrong here

    The operator's controller manager manages lifecycle components, not pod scheduling decisions. Scaling it down and up could interrupt reconciliation and is not a diagnostic step. Since operator pods are healthy, the issue is more likely a resource request mismatch or node selector, not operator state.

About these practice questions

Courseiva writes every NCP-AIO question from scratch — 309 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.