NCP-AIO Workload Management Practice Question
An administrator supports a shared inference cluster where a single A100 GPU must serve several small models concurrently. They configure the NVIDIA device plugin with a time-slicing configuration that advertises multiple replicas of the same physical device. After deployment, users report that one noisy model starves the others and latency spikes unpredictably. Which statement best explains the observed behaviour?
⚠ Common exam trap
The trap here is equating multiple advertised GPU replicas with multiple isolated GPUs; time-sliced replicas share one device's memory and compute engines and therefore cannot guarantee fair or predictable service levels.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Time-slicing interleaves work on one physical GPU with no memory or fault isolation, so a heavy workload can dominate device time and memory bandwidth.
Time-slicing creates multiple schedulable replicas of one physical GPU, which improves packing density but provides no hardware-level isolation. All replicas contend for the same memory, caches, and execution bandwidth, so a dominant workload can starve lighter ones and produce erratic tail latency. Where strict isolation or predictable performance is required, hardware partitioning such as Multi-Instance GPU is the appropriate mechanism instead of replica multiplexing.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Multi-Instance GPU mode must be enabled alongside time-slicing, because the two modes are designed to be combined on the same device.
Why it's wrong here
MIG and time-slicing are mutually exclusive strategies for a given GPU configuration: MIG carves the device into isolated hardware instances, while time-slicing multiplexes the whole device. They are not layered together. Enabling MIG would change the device into fixed partitions and remove the replica behaviour entirely, so this is not the reason for the noisy-neighbour symptom.
- ✗
The device plugin failed to register replicas, so the scheduler placed all pods on one logical device instead of distributing them.
Why it's wrong here
If replica registration had failed, the pods requesting replicas would remain Pending rather than run and contend. The fact that all models execute but interfere with each other indicates replicas were advertised and scheduled correctly. The symptom is resource contention inside a shared device, not a scheduling or registration failure, so this explanation does not fit the evidence.
- ✓
Time-slicing interleaves work on one physical GPU with no memory or fault isolation, so a heavy workload can dominate device time and memory bandwidth.
Why this is correct
Time-slicing exposes several logical replicas of one physical GPU and lets the driver rotate execution among them, but the replicas share the same memory space and compute engine. There is no partitioning of memory capacity, cache, or bandwidth, so a demanding model can monopolize execution slots and evict or slow others. That matches the reported starvation and unpredictable latency exactly.
- ✗
The pods lack a RuntimeClass reference, so the NVIDIA container runtime never applies the configured replica strategy at container start.
Why it's wrong here
The replica count is defined in the device plugin configuration on the node, and the plugin advertises that count regardless of pod-level RuntimeClass usage. A missing RuntimeClass would typically prevent GPU access altogether rather than cause uneven sharing. Since the models are running and competing, the replica strategy is clearly in effect, so this cause is inconsistent with the observation.
About these practice questions
This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.