Courseiva

NCP-AIO · topic practice

Workload Management practice questions

Workload Management on the NCP-AIO exam covers scheduling, isolating, and sharing GPU resources across containers and Kubernetes pods on NVIDIA systems. Questions present operational symptoms — low GPU utilization, OOM crashes, or jobs stuck Pending — and ask you to identify the misconfiguration in device plugins, MIG, time-slicing, or PriorityClass settings.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Workload Management

What the exam tests

What to know about Workload Management

Be able to map symptoms to the correct GPU-sharing mechanism and scheduling configuration. The single most important thing: know which sharing mode (MIG, time-slicing, or exclusive) a workload needs, and verify device plugin, resource requests, and PriorityClass align before blaming hardware.

Configuring NVIDIA device plugin, MIG, and time-slicing so multiple containers share one physical GPU

Diagnosing low GPU utilization versus high CPU load in Kubernetes training jobs

Tracing OOM errors to workload management causes such as memory limits or oversubscription

Interpreting PriorityClass and scheduling behavior when GPU jobs fail to start

Watch out for

Common Workload Management exam traps

  • ▸Assuming OOM always means insufficient VRAM, ignoring container memory limits or concurrent processes on the same GPU
  • ▸Confusing MIG partitioning with time-slicing, and expecting isolation guarantees that the chosen mode does not provide
  • ▸Treating Pending GPU pods as hardware failures instead of checking PriorityClass, node selectors, and device plugin capacity

Practice set

Workload Management questions

20 questions · select your answer, then reveal the explanation

An AI researcher is running a multi-node training job on a DGX SuperPOD using Kubernetes. The job frequently fails due to GPU memory fragmentation. Which workload management strategy best mitigates this issue?

Which THREE factors should be considered when defining GPU resource requests for a containerized AI workload to prevent scheduling failures?

An administrator is optimizing a Kubernetes cluster using NVIDIA GPU Operator. Which TWO configurations must be correctly implemented to ensure that GPU-bound workloads are scheduled efficiently across nodes with heterogeneous GPU types?

When a job fails due to an OOM (Out of Memory) error on the GPU, which action should a workload manager take to best improve job success rates for subsequent attempts?

Which THREE factors are essential when planning a cluster-wide policy for GPU workload preemption?

Refer to the exhibit. The pod is failing to schedule in a cluster with available A100 GPUs. Which factor is most likely preventing the scheduler from placing the pod?

Exhibit

apiVersion: v1
kind: Pod
metadata:
  name: gpu-workload
spec:
  containers:
  - name: cuda-container
    image: nvidia/cuda:11.0-base
    resources:
      limits:
        nvidia.com/gpu: 2
  nodeSelector:
    gpu-type: tesla-a100

When implementing a job queue for LLM training that requires periodic checkpointing to distributed storage, which strategy best minimizes the impact of checkpointing on training throughput?

Refer to the exhibit. If a high-priority training job is submitted, what is the expected outcome of this policy in a restricted workload management system?

Exhibit

{
  "policy": "deny",
  "resources": ["gpu-cluster-a"],
  "max_gpu_usage": "10%",
  "priority_level": "low"
}

An AI researcher reports that a multi-node training job is frequently preempted on a Kubernetes cluster using the NVIDIA GPU Operator. Which component is responsible for orchestrating the preemption logic based on pod priority and GPU resource availability?

An AI Ops engineer is optimizing a large-scale training cluster. Which TWO actions would effectively improve GPU utilization in a multi-tenant environment?

Which parameter in the NVIDIA Device Plugin configuration is used to enable the creation of multiple GPU instances via time-slicing?

Why would an AI Ops engineer choose to use 'drain' mode on a GPU node before performing maintenance?

An organization is running multi-tenant AI training workloads on a Kubernetes cluster with NVIDIA GPU Operator. Which mechanism should an administrator use to ensure that a high-priority training job preempts low-priority development pods when GPU resources are exhausted?

Refer to the exhibit. The pod remains in a 'Pending' state. After verifying that the GPU Operator is installed and nodes have available capacity, what is the most likely cause of this scheduling failure?

Exhibit

apiVersion: v1
kind: Pod
metadata:
  name: gpu-workload
spec:
  containers:
  - name: cuda-container
    image: nvcr.io/nvidia/k8s/cuda-sample:nbody
    resources:
      limits:
        nvidia.com/gpu: 1
  nodeSelector:
    nvidia.com/gpu.present: "true"

Which THREE factors are primary drivers for choosing between 'exclusive' GPU mode and 'shared' GPU mode using Time-Slicing in a production AI environment?

A platform engineer is deploying a multi-node large language model training job on a Slurm cluster managed by NVIDIA Base Command Manager. The job repeatedly fails with a 'node not responding' error, and the scheduler log shows that the job was allocated nodes that had been drained for maintenance but were not yet returned to service. Which Slurm configuration should the engineer verify to ensure the scheduler does not assign jobs to nodes in a drained state?

An AI operations team runs mixed inference and training pods on a shared NVIDIA-accelerated Kubernetes cluster. They want the cluster scheduler to pack inference pods onto GPUs that already host other inference pods, while keeping training pods isolated on dedicated GPUs. Which Kubernetes mechanism should the administrator configure to achieve this placement behavior?

A research team runs a mixture of long-running interactive Jupyter notebook sessions and short-lived batch inference jobs on the same Kubernetes cluster managed by the NVIDIA GPU Operator. Interactive sessions frequently hold GPU memory for hours while idle, starving the batch jobs. The administrator must let batch jobs preempt idle interactive pods without killing sessions that are actively computing. Which Kubernetes capability should be configured?

An AI operations team is using NVIDIA Run:ai to manage a shared GPU cluster. They want to ensure that a high-priority inference job can preempt lower-priority training jobs when GPU resources are scarce. The team has defined priorities in the Run:ai UI but observes that preemption is not occurring. Which Run:ai feature must be enabled to allow preemption based on priority?

A platform team runs mixed training and inference workloads on a Kubernetes cluster with NVIDIA GPU Operator and MIG-enabled A100 GPUs. They want to improve GPU utilization while keeping tenant isolation for latency-sensitive inference. Which TWO configurations should they implement? (Choose two.)

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Workload Management sessions

Start a Workload Management only practice session

Every question in these sessions is drawn from the Workload Management domain — nothing else.

Related practice questions

Related NCP-AIO topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the NCP-AIO exam test about Workload Management?
Be able to map symptoms to the correct GPU-sharing mechanism and scheduling configuration. The single most important thing: know which sharing mode (MIG, time-slicing, or exclusive) a workload needs, and verify device plugin, resource requests, and PriorityClass align before blaming hardware.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Workload Management questions in a focused session?
Yes — the session launcher on this page draws every question from the Workload Management domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other NCP-AIO topics?
Use the topic links above to move to related areas, or go back to the NCP-AIO question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the NCP-AIO exam covers. They are not copied from any real exam or dump site.