Be able to map symptoms to the correct GPU-sharing mechanism and scheduling configuration. The single most important thing: know which sharing mode (MIG, time-slicing, or exclusive) a workload needs, and verify device plugin, resource requests, and PriorityClass align before blaming hardware.
Start practicing
Workload Management — choose a session length
Free · No account required
Domain overview
Workload Management on the NCP-AIO exam covers scheduling, isolating, and sharing GPU resources across containers and Kubernetes pods on NVIDIA systems. Questions present operational symptoms — low GPU utilization, OOM crashes, or jobs stuck Pending — and ask you to identify the misconfiguration in device plugins, MIG, time-slicing, or PriorityClass settings.
Exam objectives
Configuring NVIDIA device plugin, MIG, and time-slicing so multiple containers share one physical GPU
Diagnosing low GPU utilization versus high CPU load in Kubernetes training jobs
Tracing OOM errors to workload management causes such as memory limits or oversubscription
Interpreting PriorityClass and scheduling behavior when GPU jobs fail to start
Assuming OOM always means insufficient VRAM, ignoring container memory limits or concurrent processes on the same GPU
Confusing MIG partitioning with time-slicing, and expecting isolation guarantees that the chosen mode does not provide
Treating Pending GPU pods as hardware failures instead of checking PriorityClass, node selectors, and device plugin capacity
Click any question to see the full explanation and answer options, or start a focused practice session above.
Which TWO methods are effective for enforcing GPU resource isolation in a multi-tenant NVIDIA Kubernetes environment?
2When managing large-scale model training jobs, what is the primary purpose of using a Job Scheduler like Slurm or Kubernetes Batch?
3A production inference service experiences intermittent latency spikes. The service is deployed on shared infrastructure. Which tool would best help an AI Ops engineer identify if GPU resource contention is the cause?
4Why is it important to use a persistent storage volume for model checkpoints in a distributed training job?
5In a multi-node training scenario, what is the significance of the NVIDIA Collective Communications Library (NCCL) in workload management?
6Which TWO of the following are benefits of using containerized GPU workloads compared to bare-metal deployment?
7An AI ops engineer notices that a specific training workload is experiencing high 'wait' times for GPU resources despite the cluster having available idle GPUs. What is the most likely cause?
8Which mechanism does the NVIDIA Device Plugin use to communicate GPU availability to the Kubernetes Kubelet?
9What is the primary function of an 'InitContainer' in an NVIDIA GPU-enabled pod deployment?
10An organization is deploying large-scale LLM training workloads on an NVIDIA DGX SuperPOD. The data science team reports that training jobs are frequently preempted by higher-priority batch jobs, leading to significant checkpointing overhead. Which Workload Manager configuration strategy best minimizes resource fragmentation and improves overall cluster utilization while maintaining SLA requirements?
11When managing workloads on NVIDIA DGX systems, what is the primary role of the NVIDIA device plugin in the Kubernetes ecosystem?
12A cluster administrator notices that GPU utilization is low despite high queue volume. After analyzing the logs, they identify that many pods are failing because they cannot access the necessary CUDA libraries. What is the most likely cause, and which component should be verified?
13When designing a workload management strategy for multi-tenant AI training, what is the most effective way to ensure isolation between different tenants using the same physical GPU nodes?
14A data engineering team is deploying a distributed data processing workload. Which THREE metrics are most important to monitor in the Workload Manager to ensure optimal GPU throughput and identify potential bottlenecks?
15Which component in the NVIDIA AI ecosystem is responsible for monitoring and reporting GPU telemetry data, such as power usage, temperature, and utilization, to Prometheus?
16Which of the following is the most appropriate workload management technique for a bursty AI inference workload that requires low latency but does not need full GPU power for every request?
17In an NVIDIA-accelerated Kubernetes environment, why is it critical to configure a 'RuntimeClass' for GPU-enabled pods?
18An organization is migrating their on-premises AI training to a hybrid cloud environment. Which component is most important to maintain consistent workload management across both the on-premises DGX systems and cloud-based GPU nodes?
19An AI researcher is running a large-scale training job on an NVIDIA DGX system using Kubernetes. They observe that GPU utilization is consistently low despite high CPU load. Which workload management configuration is most likely to resolve this bottleneck by optimizing data pipeline throughput?
20Which TWO strategies should an administrator implement to ensure fair resource scheduling in a multi-tenant NVIDIA cluster using Kubernetes and the NVIDIA device plugin?
21A DevOps engineer needs to monitor GPU health metrics in real-time for workload management. Which tool provides the most granular visibility into GPU utilization and power consumption for individual containers?
22An organization is migrating AI workloads to a private cloud. Which feature is essential for ensuring that GPU resources are dynamically reclaimed and reallocated to different departments without manual intervention?
23Which THREE factors must be considered when sizing a persistent storage solution for multi-node distributed training checkpoints?
24What is the primary benefit of using NVIDIA GPU Operator in a Kubernetes cluster for workload management?
25When running multi-instance GPU (MIG) workloads, what is the main advantage of assigning specific MIG profiles to different Kubernetes namespaces?
26Which approach is most effective for scaling an inference workload that experiences sudden, unpredictable spikes in request volume?
27When debugging a workload that consistently crashes with 'Out of Memory' (OOM) errors despite sufficient GPU VRAM, what is the most likely cause related to workload management?
28Which scheduling strategy is recommended to maximize the efficiency of long-running training jobs on preemptible instances?
29Refer to the exhibit. A cluster administrator notices that GPU jobs with this PriorityClass are failing to start even when empty GPUs are available. What is the most likely cause?
30Which NVIDIA technology allows for partitioning a single physical GPU into multiple independent instances, each with dedicated compute and memory resources for smaller workloads?
31When managing GPU resources in a shared cluster, which configuration best prevents 'noisy neighbor' scenarios where one GPU task consumes all available memory bandwidth?
32Refer to the exhibit. The cluster administrator observes near-capacity memory utilization across three GPUs. What is the most likely consequence if an additional pod is scheduled to these GPUs without memory partitioning?
33Which strategy is most effective for managing heterogeneous GPU clusters containing both older architectures (e.g., V100) and newer architectures (e.g., H100)?
34Which THREE factors should be considered when estimating GPU memory requirements for a Large Language Model (LLM) fine-tuning job?
35An administrator is optimizing multi-GPU utilization. Which TWO of the following configurations allow multiple containers to share a single physical GPU on a supported NVIDIA architecture?
36Which component of the NVIDIA GPU Operator is responsible for monitoring GPU health and reporting telemetry data to the Kubernetes control plane?
37Refer to the exhibit. An administrator is attempting to deploy a job to a namespace with a ResourceQuota defined. What is the cause of this error?
38When deploying large-scale distributed training jobs, why is it recommended to use the NVIDIA Network Operator in conjunction with the GPU Operator?
39An AI operations team runs mixed training and inference workloads on a Kubernetes cluster managed with the NVIDIA GPU Operator. Inference pods frequently arrive in bursts and must start within seconds, while long-running training jobs occupy most MIG-capable A100 GPUs for days. Administrators want burst inference pods to obtain GPU capacity immediately without preempting or restarting the training jobs, and they want the cluster to reclaim those resources automatically when the burst ends. Which approach best satisfies these requirements?
40A Kubernetes cluster running the NVIDIA GPU Operator is shared by an inference team and a research team. The research team's training pods repeatedly evict the inference pods from GPUs, causing latency spikes in production. The administrator wants to guarantee that inference pods always get GPU access first. Which Kubernetes scheduling mechanism should be configured?
41An administrator must run a batch inference job that requires exactly two NVIDIA GPUs on a Kubernetes cluster managed by the NVIDIA GPU Operator. Which pod specification field should be used to request those GPUs?
42A platform team operates a Kubernetes cluster where several teams submit GPU training jobs. The administrator needs to enforce per-namespace limits on the number of GPUs that can be consumed and prevent a single namespace from monopolizing all GPU capacity. Which TWO Kubernetes resources should be configured to achieve this? (Choose two.)
43An administrator runs a Kubernetes cluster with the NVIDIA GPU Operator. A data science team wants to run several small inference containers that each use only a fraction of a GPU's compute and memory, but the cluster currently assigns whole GPUs per pod. Which approach allows multiple containers to share a single physical GPU with memory isolation?
44A platform team runs an NVIDIA AI cluster with the GPU Operator deployed. Users submit jobs directly with kubectl and frequently request whole GPUs even when their notebooks only need a fraction of one. The team wants Kubernetes itself to admit and queue jobs based on GPU demand without users changing their manifests. Which component should the team deploy to meet this requirement?
45A researcher submits a distributed training job that spans four pods, each needing one GPU, and the pods must start together or not at all. The administrator wants Kubernetes to schedule all four pods only when four GPUs are simultaneously available. Which workload management construct should be used?
46An administrator manages a shared NVIDIA cluster where several teams run inference services. One team's pods are being evicted repeatedly, and DCGM metrics show the node's GPUs are healthy but memory on the devices is nearly exhausted. The team insists their model fits. Which action should the administrator take FIRST to identify the cause?
47An AI operations engineer manages a shared Kubernetes cluster running NVIDIA GPU Operator. Several teams report that their inference pods remain in a Pending state with the event '0/8 nodes are available: 8 Insufficient nvidia.com/gpu.' The administrator verifies that nvidia-smi on all nodes shows idle GPUs. Which action should the administrator take first to resolve the scheduling failure?
48A platform engineer is preparing an NVIDIA-accelerated Kubernetes cluster for a new team that will submit PyTorch training jobs. The team wants jobs to request GPUs without hardcoding device indices. Which Kubernetes resource should the engineer ensure is installed and healthy so pods can request nvidia.com/gpu resources?
49A data engineering team runs nightly batch inference on a Kubernetes cluster with NVIDIA GPUs. Jobs sometimes fail because two pods are scheduled onto the same physical GPU and one exhausts framebuffer memory. The team wants each pod to receive an isolated slice of a GPU with dedicated memory. Which NVIDIA feature should they enable?
50An AI platform team runs GPU-accelerated inference pods on a Kubernetes cluster with the NVIDIA GPU Operator. During peak load, high-priority latency-sensitive inference pods are frequently preempted by large batch training jobs that were submitted later. The team wants the scheduler to guarantee that inference pods always win placement and preemption decisions against training pods without manually cordoning nodes. Which mechanism should the administrator configure?
51An AI operations engineer manages a Kubernetes cluster where the NVIDIA GPU Operator's device plugin exposes GPUs as schedulable resources. A data science team submits a batch inference job that requests one GPU but does not specify a node selector or tolerations. The job stays in Pending while other GPU nodes remain idle because they carry a taint the GPU Operator applied to reserve them for a specific workload class. Which approach is the most appropriate for the engineer to make the job schedulable without disrupting the reserved nodes?
52An MLOps engineer needs to guarantee that a latency-sensitive inference Deployment always has GPU capacity available, even when a large training Job is submitted to the same namespace. The cluster uses the NVIDIA GPU Operator and nodes have four A100 GPUs each. Which approach reliably reserves GPU capacity for the inference Deployment?
53An administrator manages a Kubernetes cluster where a training job repeatedly fails with an OutOfMemory error on the GPU even though the pod requests one nvidia.com/gpu. DCGM metrics show that another pod on the same node is consuming GPU memory concurrently. GPU sharing via time-slicing is enabled cluster-wide. Which action should the administrator take to prevent this cross-pod interference while preserving the ability to share GPUs among trusted inference workloads?
54A platform team runs an NVIDIA GPU Operator-managed Kubernetes cluster shared by two research groups. Group A's pods request `nvidia.com/gpu: 1` and are scheduled correctly, but Group B's pods that omit any GPU resource request are also landing on GPU nodes and consuming host memory and CPU, degrading Group A's jobs. The team wants Group B's non-GPU pods to stop consuming capacity on the GPU node pool without changing Group B's manifests. Which action should the administrator take?
55A platform team runs a Kubernetes cluster where the NVIDIA GPU Operator is installed and time-slicing is configured with a ConfigMap that advertises four replicas per physical GPU. A data scientist submits a PyTorch training pod requesting nvidia.com/gpu: 1. The pod stays Pending indefinitely, and the scheduler event reads 'Insufficient nvidia.com/gpu'. The node's GPUs are otherwise idle and healthy. Which action most directly resolves the pending state?
56A cluster runs mixed workloads: latency-sensitive inference services and best-effort batch jobs. Administrators observe that batch jobs occasionally occupy all GPUs, causing inference requests to queue and breach service level objectives. They want inference pods to be admitted immediately while allowing batch work to use remaining capacity and be preempted when needed. Which approach should they implement?
57A platform team runs mixed training and inference workloads on a Kubernetes cluster with the NVIDIA GPU Operator. Inference pods are latency-sensitive and must not be preempted, while training pods can be interrupted and restarted. The team wants training jobs to yield GPUs to inference jobs when capacity is scarce, without manual intervention. Which Kubernetes mechanism should the team configure to achieve this behavior?
58An administrator is using NVIDIA Base Command Manager to manage a cluster with a mix of GPU and CPU nodes. They need to ensure that a newly added GPU node is correctly recognized and that jobs can be scheduled on it. Which TWO actions must be performed to integrate the new node into the Base Command Manager cluster? (Choose two.)
59A platform team runs an on-premises Kubernetes cluster for AI inference. Several teams submit pods that request the same GPU device, and the scheduler places more pods onto a node than there are available GPUs, causing OOM errors on the device. The administrator wants the Kubernetes scheduler itself to prevent overcommitting GPUs without any custom admission controller. Which action should the administrator take?
60A research group submits a distributed PyTorch training job spanning eight GPUs across two nodes. The job completes but produces a model with accuracy far below the single-node baseline, and logs show that several ranks started training before their peers had initialized the process group. The administrator must ensure that all ranks are launched together and that a failed rank terminates the whole job. Which combination of Kubernetes mechanisms should be used?
61A platform team is preparing a Kubernetes cluster for AI workloads and wants the GPU device plugin, driver containers, and monitoring components deployed and kept in sync automatically on every GPU node. Which component should be installed to achieve this?
62An MLOps engineer manages a Kubernetes cluster where the NVIDIA GPU Operator runs the MIG manager. Several inference pods must each receive an isolated, fixed slice of a single A100, and the team wants the slices to survive node reboots without manual reconfiguration. Which combination of settings should the engineer apply?
63A research team is submitting many short-lived experiment jobs to an NVIDIA-accelerated Kubernetes cluster. The operations team wants to reduce GPU idle time and improve overall utilization without modifying the training code. Which TWO approaches should the operations team implement? (Choose two.)
64A cloud operations team is using NVIDIA AI Enterprise with Kubernetes to deploy inference workloads. They want to ensure that GPU resources are allocated to pods only when explicitly requested, and that pods without GPU requests do not consume GPU resources. Which Kubernetes feature should they use to enforce this behavior?
65An AI operations engineer is troubleshooting a Kubernetes cluster where several GPU training pods fail to start with a device plugin allocation error, even though the nodes report healthy GPUs. The engineer suspects the pods are requesting more GPU resources than a single physical card can provide without a sharing mechanism. Which TWO configurations would legitimately allow multiple pods to consume a single physical GPU on these nodes? (Choose two.)
66An administrator supports a shared inference cluster where a single A100 GPU must serve several small models concurrently. They configure the NVIDIA device plugin with a time-slicing configuration that advertises multiple replicas of the same physical device. After deployment, users report that one noisy model starves the others and latency spikes unpredictably. Which statement best explains the observed behaviour?
67A platform engineer manages an NVIDIA-accelerated Kubernetes cluster running the NVIDIA GPU Operator on nodes with A100 GPUs. Several data-science teams submit training jobs, and the engineer must ensure each team's pods receive a full physical GPU exclusively, with no two pods sharing the same device. Which scheduling configuration should the engineer apply to the pod specification to guarantee exclusive whole-GPU allocation?
68An administrator is configuring a Kubernetes cluster where some nodes have A100 GPUs and others have H100 GPUs. A training job requires specific GPU memory capacity and CUDA compute capability. Which mechanism should the administrator use to ensure the job is only scheduled onto nodes with the correct GPU model?
69An administrator manages a Kubernetes cluster where the NVIDIA GPU Operator has deployed the device plugin and MIG Manager. A tenant wants to run several small inference services that each need only a fraction of a GPU, isolated from other tenants' memory and fault domains. The administrator decides to use Multi-Instance GPU mode. Which TWO actions must be performed to make MIG-backed GPU resources schedulable to those pods? (Choose two.)
70An AI operations team runs a shared Kubernetes cluster with the NVIDIA GPU Operator and several namespaces owned by different groups. A group reports that its training pods are stuck Pending with an event indicating insufficient nvidia.com/gpu, yet cluster-wide dashboards show many GPUs idle. Investigation reveals the idle GPUs belong to nodes labeled for another group, and the affected namespace has a node affinity rule pinning its pods to a specific GPU generation that is fully consumed. Which action best resolves the Pending pods while respecting multi-tenant boundaries?
71A data science team submits a PyTorch training job to a Kubernetes cluster managed by Run:ai. The job requests two GPUs but only one is allocated, and the second worker hangs waiting for a peer. Which Run:ai capability should the administrator verify is configured so the distributed job receives all requested GPUs atomically?
72An administrator manages a cluster where inference services and batch training share the same GPU nodes. During business hours, inference pods must be scheduled promptly, while training jobs can wait. The administrator wants preemption so that a pending high-priority inference pod can evict a lower-priority training pod when no GPU is free, with evicted training resuming later. Which configuration achieves this?
73An ML platform team runs an NVIDIA GPU Operator-managed cluster and wants to allow multiple pods to share a single A100 GPU so that small inference services can co-reside without each consuming a whole device. The team needs a time-slicing configuration that applies to all GPU nodes in the cluster. Which approach should the administrator take?
74A data science team submits a PyTorch training job to a Kubernetes cluster where the NVIDIA GPU Operator is installed. The pod stays in Pending state, and kubectl describe shows the message '0/6 nodes are available: 6 Insufficient nvidia.com/gpu.' The administrator confirms the nodes have healthy GPUs and the device plugin pods are Running. What is the most likely cause?
75An AI operations engineer manages a Kubernetes cluster running the NVIDIA GPU Operator. A team wants its long-running inference deployment to be automatically rescheduled if the GPU on a node develops an uncorrectable error that the device plugin or health checks detect. The team also wants the node to stop accepting new GPU pods until the issue is resolved. Which combination of behaviors should the engineer rely on to meet these requirements?
76A data science team submits a PyTorch distributed training job to a Kubernetes cluster with the NVIDIA GPU Operator installed. The job's pods repeatedly fail with a CUDA initialization error, while a simple `nvidia-smi` check inside an interactive pod on the same node succeeds. The administrator confirms the node's driver is healthy and the device plugin is advertising GPUs. Which configuration should the administrator verify first?
77A research team runs a multi-node distributed training job spanning eight GPU nodes. Jobs frequently begin execution before all worker pods are running, and the collective initialization hangs until the operator manually scales the job down and up. The administrator wants the scheduler to admit the job only when all of its pods can be placed together. Which mechanism should be used?
78An AI operations team runs long-running training jobs on a Kubernetes cluster with NVIDIA GPU Operator. They observe that after a node is rebooted for maintenance, some pods resume but report CUDA 'unknown error' and the device plugin shows unhealthy GPUs. Which configuration should the administrator review to ensure the driver and device plugin recover cleanly after reboot?
79A research organization runs an NVIDIA DGX SuperPOD with a Kubernetes cluster managed by the NVIDIA GPU Operator and Network Operator. A distributed training job using PyTorch DDP across 32 nodes stalls at initialization, and the administrator suspects the collective communication library is not selecting the high-speed fabric. Which configuration should the administrator verify first to ensure NCCL uses the correct network interface and topology?
80A financial services firm must prove to auditors that an AI training job ran on hardware located only in its Frankfurt data center and that no pod could ever be scheduled onto GPUs in other regions. The cluster spans three regions with nodes labeled topology.kubernetes.io/region. Which approach most directly enforces this placement requirement?
81An administrator is tuning a Kubernetes cluster that runs many small inference pods on NVIDIA GPUs. Utilization is low because each pod reserves a full GPU while using only a fraction of its memory and compute. The administrator wants to share GPUs across pods while preserving memory-level isolation between processes. Which TWO configurations achieve this? (Choose two.)
82An administrator is tuning a Kubernetes cluster that runs GPU Operator. Users report that GPU jobs are sometimes scheduled onto nodes whose drivers are older than the CUDA version the container needs, causing runtime failures. The administrator wants to prevent incompatible placements before pods are bound. (Choose two.)
83An administrator supports a multi-tenant cluster where several teams share GPUs. Leadership requires that each team's batch jobs receive a fair share of GPU time and that one team cannot monopolize devices by submitting thousands of low-priority pods. Jobs are submitted through a Kubernetes-native batch scheduler that supports queueing. Which approach best enforces fair-share GPU allocation across teams?
84An AI operations team manages a shared Kubernetes cluster where a nightly batch training workload requests nvidia.com/gpu resources and occasionally consumes all GPU memory on a node, causing a co-located interactive notebook pod to fail with CUDA out-of-memory errors. The team wants the interactive notebook to be isolated from the batch workload's memory usage without adding new hardware. Which action best achieves this on supported data center GPUs?
85An AI operations engineer manages a shared Kubernetes cluster where several teams submit GPU jobs. The engineer must prevent any single namespace from consuming all GPU capacity and must also ensure that jobs from one team cannot starve others during peak periods. Which combination of Kubernetes and NVIDIA GPU Operator features should the engineer implement?
86An administrator notices that GPU utilization on a training cluster hovers around 25 percent even though many jobs are queued. Investigation shows that each job requests a full GPU, but the models are small and alternate between short data-loading phases and brief compute bursts. The administrator wants to increase effective GPU utilization without changing model code. Which action should be taken first?
87A research team submits a multi-node training job using a `Job` with eight pods, each requesting one GPU. The cluster has eight GPU nodes, each with one A100. The administrator observes that all eight pods are spread one per node and the job runs, but throughput is far below expectations and NCCL logs show repeated fallback from GPUDirect RDMA to socket transport. Which action most directly addresses the root cause?
Be able to map symptoms to the correct GPU-sharing mechanism and scheduling configuration. The single most important thing: know which sharing mode (MIG, time-slicing, or exclusive) a workload needs, and verify device plugin, resource requests, and PriorityClass align before blaming hardware.
The Courseiva NCP-AIO question bank contains 87 questions in the Workload Management domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Workload Management domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included