Courseiva
Workload Management →mediumMultiple Choice

NCP-AIO Workload Management Practice Question

An AI platform team runs GPU-accelerated inference pods on a Kubernetes cluster with the NVIDIA GPU Operator. During peak load, high-priority latency-sensitive inference pods are frequently preempted by large batch training jobs that were submitted later. The team wants the scheduler to guarantee that inference pods always win placement and preemption decisions against training pods without manually cordoning nodes. Which mechanism should the administrator configure?

⚠ Common exam trap

The trap here is assuming that GPU sharing features such as time-slicing or MIG establish workload precedence, when only PriorityClass plus scheduler preemption actually reorders and evicts competing pods.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Assign the inference pods a higher PriorityClass value and reference it in their pod spec so the kube-scheduler preempts lower-priority training pods when resources are scarce.

Kubernetes resolves GPU contention through the scheduler's priority and preemption logic, not through device sharing or quota limits. Giving inference pods a higher PriorityClass value makes the kube-scheduler place them first and evict lower-priority training pods when capacity is exhausted, which delivers the guaranteed preference the platform team requires.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable the GPU Operator's time-slicing feature with a replica count greater than one so inference and training pods share each physical GPU concurrently.

    Why it's wrong here

    Time-slicing increases the number of logical GPU slots so several containers can share one device, but it does not establish any ordering or preemption rules between pods. Under contention, training and inference pods would simply compete for the same shared slices, so latency-sensitive inference would still be starved rather than guaranteed placement.

  • ✗

    Create a separate namespace with a ResourceQuota that limits training pods to a small number of GPUs and place all inference pods in the default namespace.

    Why it's wrong here

    A ResourceQuota caps aggregate consumption inside a namespace, but it does not cause the kube-scheduler to preempt a running training pod or to prefer an inference pod during scheduling. Training jobs admitted before the quota is reached would keep their GPUs, so the latency-sensitive pods could still be blocked.

  • ✓

    Assign the inference pods a higher PriorityClass value and reference it in their pod spec so the kube-scheduler preempts lower-priority training pods when resources are scarce.

    Why this is correct

    PriorityClass is the native Kubernetes scheduling mechanism that orders pending pods; a higher integer value makes the kube-scheduler prefer those pods and preempt lower-priority workloads that already occupy GPUs. Referencing the class in the pod spec ensures inference pods win placement and preemption decisions, which is exactly what the team needs.

  • ✗

    Add a nodeSelector to the inference pods that pins them to nodes labeled with the NVIDIA GPU product name, so they only run on nodes not used by training.

    Why it's wrong here

    A nodeSelector only filters which nodes a pod is eligible for; it neither raises scheduling precedence nor triggers preemption of pods already running on those nodes. If training pods occupy every GPU node, the pinned inference pods remain Pending, so this approach cannot guarantee that inference wins under contention.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

Courseiva writes every NCP-AIO question from scratch — 309 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.