Courseiva
Workload Management →easyMultiple Choice

NCP-AIO Workload Management Practice Question

A platform team runs an on-premises Kubernetes cluster for AI inference. Several teams submit pods that request the same GPU device, and the scheduler places more pods onto a node than there are available GPUs, causing OOM errors on the device. The administrator wants the Kubernetes scheduler itself to prevent overcommitting GPUs without any custom admission controller. Which action should the administrator take?

⚠ Common exam trap

The trap here is assuming that Kubernetes natively understands GPUs as limited resources; without the device plugin exposing them as extended resources, the scheduler treats GPU nodes as ordinary nodes and will happily overcommit devices.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Install the NVIDIA device plugin so GPUs are advertised as schedulable extended resources and add a matching resource request to each pod spec.

Advertising GPUs as extended resources through the NVIDIA device plugin is what lets the default scheduler treat each device as countable, finite node capacity. Pods that declare a GPU request are then placed only where an unallocated device exists, and no custom admission controller is required because the scheduler performs the arithmetic itself. Quotas, taints, and autoscalers influence eligibility or CPU and memory sizing, not device-level accounting.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Install the NVIDIA device plugin so GPUs are advertised as schedulable extended resources and add a matching resource request to each pod spec.

    Why this is correct

    The NVIDIA k8s-device-plugin registers each GPU as an extended resource such as nvidia.com/gpu, which makes the scheduler count devices as finite node capacity. When a pod requests one GPU, the scheduler subtracts it from allocatable capacity, so no node can be overcommitted. This directly solves the reported overplacement without writing any admission logic, because scheduling math handles accounting natively.

  • ✗

    Set a node taint on each GPU node and add a toleration to pods, so only pods that explicitly opt in are placed there.

    Why it's wrong here

    Taints and tolerations control which pods are eligible for a node, not how many pods of that eligible set fit. Every opted-in pod can still be scheduled to the same node until CPU or memory limits are hit, so the GPU itself remains overcommitted. This changes admission eligibility rather than device accounting, leaving the original OOM behaviour intact.

  • ✗

    Enable the Kubernetes Vertical Pod Autoscaler in recommendation mode on the GPU namespaces to detect device pressure and reschedule pods.

    Why it's wrong here

    Vertical Pod Autoscaler adjusts CPU and memory requests based on observed usage and can evict and restart pods, but it does not model GPUs as schedulable units and does not know how many devices a node physically exposes. It reacts after eviction rather than preventing the overplacement in the first place, so device OOM errors would continue during the learning window.

  • ✗

    Create a ResourceQuota in each team namespace limiting the count of pods that may run, so fewer pods land on GPU nodes.

    Why it's wrong here

    A ResourceQuota caps aggregate pod counts or resource totals per namespace, but a pod count limit is unrelated to how many GPUs a node exposes. Two pods on a single-GPU node still overcommit the device if neither pod requests the GPU resource. The quota also does not inform the scheduler about device inventory, so placement remains blind to GPU availability.

About these practice questions

One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.