Courseiva
Workload Management →hardMultiple Choice

NCP-AIO Workload Management Practice Question

A cluster runs mixed workloads: latency-sensitive inference services and best-effort batch jobs. Administrators observe that batch jobs occasionally occupy all GPUs, causing inference requests to queue and breach service level objectives. They want inference pods to be admitted immediately while allowing batch work to use remaining capacity and be preempted when needed. Which approach should they implement?

⚠ Common exam trap

The trap here is equating resource sharing or quota limits with prioritization, when only priority and preemption can guarantee immediate admission.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Define priority classes and a preemption policy so inference pods have higher priority than batch jobs.

The requirement is immediate admission for inference and preemptible use of leftover capacity by batch work. Priority classes with preemption give inference pods the ability to displace lower-priority batch pods when GPUs are scarce, while batch jobs still consume idle capacity during quiet periods. Affinity, quotas, and time-slicing each constrain or share resources but cannot prioritize or preempt running workloads.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable time-slicing on all GPUs so batch and inference pods share devices.

    Why it's wrong here

    Time-slicing lets multiple pods share a GPU by interleaving contexts, which improves utilization but provides no priority or preemption. A batch pod can still occupy the device and degrade inference latency, and there is no mechanism to evict it. Sharing memory also risks contention, so this approach does not protect latency-sensitive services.

  • ✗

    Set a ResourceQuota on the batch namespace limiting total GPU requests.

    Why it's wrong here

    A ResourceQuota caps how many GPUs the batch namespace may request, which prevents it from consuming the entire cluster. However, it does not give inference pods priority or the ability to preempt. If inference demand spikes beyond reserved capacity, requests still queue, and idle quota held by batch jobs is not reclaimed, so service objectives remain at risk.

  • ✗

    Apply node affinity so inference pods only run on a dedicated subset of nodes.

    Why it's wrong here

    Node affinity constrains placement but does not create preemption or priority. If the dedicated nodes are full, inference pods still wait, and batch jobs cannot use those nodes even when idle. This wastes capacity and fails to guarantee immediate admission, since affinity alone cannot displace running workloads or arbitrate between competing pods.

  • ✓

    Define priority classes and a preemption policy so inference pods have higher priority than batch jobs.

    Why this is correct

    Kubernetes priority classes let higher-priority pods preempt lower-priority ones when resources are scarce. Assigning inference a high priority and batch a low priority ensures inference is admitted immediately, while batch jobs yield their GPUs when inference arrives. This matches the requirement that batch work uses leftover capacity and is preempted rather than blocking critical services.

About these practice questions

Courseiva writes every NCP-AIO question from scratch — 309 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.