Courseiva
Workload Management →hardMultiple Choice

NCP-AIO Workload Management Practice Question

An AI operations team runs mixed training and inference workloads on a Kubernetes cluster managed with the NVIDIA GPU Operator. Inference pods frequently arrive in bursts and must start within seconds, while long-running training jobs occupy most MIG-capable A100 GPUs for days. Administrators want burst inference pods to obtain GPU capacity immediately without preempting or restarting the training jobs, and they want the cluster to reclaim those resources automatically when the burst ends. Which approach best satisfies these requirements?

⚠ Common exam trap

The trap here is assuming that priority and preemption are the natural answer for burst workloads, when the requirement that training jobs never restart makes preemption disqualifying.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Create a node pool reserved for inference, and configure a cluster autoscaler with scale-to-zero so burst pods trigger new GPU nodes that are removed when idle.

The scenario requires two properties simultaneously: burst inference capacity that appears without disturbing running training jobs, and automatic reclamation when the burst ends. Only a dedicated autoscaled node pool satisfies both, because new GPU nodes are provisioned on demand for inference pods while training pods remain scheduled and running on their existing nodes. When demand drops, scale-to-zero removes the extra nodes, freeing GPU resources without any preemption or job restarts.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Create a node pool reserved for inference, and configure a cluster autoscaler with scale-to-zero so burst pods trigger new GPU nodes that are removed when idle.

    Why this is correct

    A dedicated inference node pool with scale-to-zero autoscaling lets burst pods trigger immediate node provisioning while training jobs keep running untouched on their own nodes. When the burst subsides, the autoscaler drains and removes the idle inference nodes, reclaiming GPU capacity and cost automatically. This directly matches both requirements: no preemption of training and automatic reclamation after the burst.

  • ✗

    Configure the inference Deployment with a topologySpreadConstraint across all GPU nodes so its replicas distribute evenly and find free capacity.

    Why it's wrong here

    Topology spread constraints only influence where pods are placed among nodes that already have allocatable GPU capacity; they do not create capacity. With training jobs occupying nearly all A100 GPUs, spreading inference replicas simply leaves them Pending, because no node has a free device. This adds scheduling complexity without addressing the real constraint, which is the absence of unallocated GPU hardware during bursts.

  • ✗

    Enable the GPU Operator's time-slicing configuration so inference and training containers share the same physical GPUs through a shared device plugin.

    Why it's wrong here

    Time-slicing lets multiple containers share a GPU by interleaving execution, but it does not guarantee inference pods start within seconds, since they still compete for the same device cycles as training kernels. Sharing a device with a heavy training workload introduces latency jitter and can slow both workloads. It also does not provide the automatic reclamation behavior requested, since the shared GPUs remain allocated to the node's training pods.

  • ✗

    Deploy inference pods with a higher PriorityClass and enable preemption so the scheduler evicts lower-priority training pods when capacity is exhausted.

    Why it's wrong here

    Priority-based preemption does free capacity quickly, but it precisely violates the stated requirement that training jobs must not be preempted or restarted. Evicted training pods lose their running state and must restart from a checkpoint, which for multi-day jobs can waste hours of GPU time. This solves the latency problem by sacrificing the long-running workloads the team explicitly wants protected, so it is the wrong design here.

About these practice questions

Courseiva writes every NCP-AIO question from scratch — 309 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.