Be able to deploy the NVIDIA GPU Operator on Kubernetes and confirm drivers, container toolkit, and device plugin are healthy, then validate multi-node NCCL/RDMA performance and read nvidia-smi output. The key skill is tracing a symptom to the correct deployment layer before changing anything.
Start practicing
Installation and Deployment — choose a session length
Free · No account required
Domain overview
This domain covers standing up and validating NVIDIA-accelerated infrastructure: installing the GPU Operator on Kubernetes, provisioning drivers and the container toolkit, configuring multi-node fabrics like NCCL and GPUDirect RDMA, and diagnosing pods stuck Pending or GPUs underperforming after deployment. Questions are scenario-based, pairing a symptom with the deployment layer or tool responsible.
Exam objectives
NVIDIA GPU Operator deployment on Kubernetes, including automatic driver, container toolkit, and device plugin installation on GPU nodes
Multi-node interconnect configuration and validation using NCCL, InfiniBand, and GPUDirect RDMA for distributed training
Real-time GPU monitoring and management with nvidia-smi for power, temperature, utilization, and memory statistics
Diagnosing GPU pod scheduling failures, including Pending states caused by missing device plugin, taints, or resource requests
Blaming application code or model configuration for slow inter-node throughput instead of checking the network fabric, NCCL settings, and RDMA path
Assuming the GPU Operator installs everything by default without verifying driver, toolkit, and device plugin DaemonSets are actually running
Debugging a Pending GPU pod as a hardware fault when the real cause is an unschedulable node, missing nvidia.com/gpu resource, or unready device plugin
Click any question to see the full explanation and answer options, or start a focused practice session above.
An administrator is deploying NVIDIA AI Enterprise on a bare-metal cluster. The workload requires full GPU isolation with minimal latency. Which configuration should the administrator select to achieve this goal?
2Refer to the exhibit. An administrator notices poor performance in an AI training job. What is the most likely cause based on the CLI output?
3Which component is strictly necessary for managing NVIDIA AI Enterprise licensing across a distributed cluster of nodes?
4When deploying NVIDIA AI Enterprise, why is the selection of the correct CUDA version in the container image critical during the installation phase?
5Refer to the exhibit. The administrator has deployed the GPU Operator, but the node does not show GPU resources. What is the most likely cause?
6During an NVIDIA AI Enterprise deployment, you are asked to configure the 'NVIDIA Container Toolkit'. What is its primary function?
7Which of the following describes the purpose of the NVIDIA GPU Operator's 'Driver Container'?
8In an air-gapped environment, what must an administrator do to ensure the GPU Operator correctly installs the necessary software components?
9Which NVIDIA technology enables the partitioning of a single physical GPU into multiple independent instances for use by different virtual machines or containers?
10You are troubleshooting a node where the GPU is detected, but the application fails to utilize it. Which log source would provide the most relevant information?
11Which NVIDIA tool allows you to verify that the GPU and its driver are properly installed and functioning on a Linux system?
12An AI engineer is deploying a large language model on an NVIDIA DGX system. The deployment fails with an error indicating an insufficient NVIDIA driver version for the required CUDA toolkit. Which action should the engineer take to resolve the dependency mismatch?
13Refer to the exhibit. An AI engineer observes that a model training job is running slower than expected. Based on the output, what is the primary cause of the performance degradation?
14When installing the NVIDIA Container Toolkit to enable GPU acceleration in Docker, which file must be modified or verified to ensure the container runtime can access the NVIDIA runtime?
15Which utility is primarily used to monitor and manage NVIDIA GPU power, temperature, and usage statistics in real-time on a Linux-based deployment?
16What is the primary function of the NVIDIA Persistence Daemon in an AI deployment?
17When configuring a multi-tenant environment on an NVIDIA DGX system, how are MIG (Multi-Instance GPU) instances best provisioned?
18An administrator is deploying NVIDIA AI Enterprise on a bare-metal cluster. Which component is mandatory to provide the necessary abstraction layer for containerized GPU resources?
19When configuring a node for NVIDIA AI Enterprise in a Kubernetes environment, what is the primary function of the NVIDIA Container Toolkit?
20Which file format is commonly used to define the configuration and state for the NVIDIA GPU Operator within a Kubernetes environment?
21Which NVIDIA software component is responsible for providing the necessary CUDA libraries to containerized applications?
22What is the primary benefit of using an NVIDIA-Certified System for AI Enterprise deployments?
23What is the primary role of a private container registry in an NVIDIA AI Enterprise deployment?
24When installing the NVIDIA GPU Operator, which namespace is typically used to ensure proper isolation and role-based access control?
25When deploying NVIDIA containers using the NVIDIA Container Toolkit, what is the primary function of the 'nvidia-container-runtime'?
26Which tool is the industry standard for monitoring and managing NVIDIA data center GPUs in a large-scale cluster deployment?
27Which TWO of the following are prerequisites for installing the NVIDIA Container Toolkit on a Linux host?
28When configuring the NVIDIA Device Plugin for Kubernetes, what is the purpose of the 'time-slicing' configuration?
29Which action must be performed after updating the NVIDIA driver on a Linux host to ensure that all active GPU containers recognize the new driver version?
30Which THREE of the following are benefits of using Multi-Instance GPU (MIG) technology in a Kubernetes environment?
31When installing NVIDIA drivers via a package manager, what is the importance of the 'dkms' package?
32Which configuration file is typically modified to enable the NVIDIA Device Plugin in a Kubernetes cluster?
33What is the primary role of the NVIDIA Data Center GPU Manager (DCGM) Exporter in a cloud-native monitoring stack?
34An administrator is deploying NVIDIA AI Enterprise on a bare-metal cluster. Which component must be installed first to ensure proper communication between the Kubernetes scheduler and the underlying GPU hardware?
35An administrator is preparing to deploy an NVIDIA AI Enterprise solution on an OpenShift cluster. Which TWO steps must be completed to ensure the NVIDIA drivers are loaded correctly on the worker nodes?
36When deploying the NVIDIA GPU Operator in a restricted-access environment (air-gapped), which THREE requirements must be addressed to ensure a successful installation?
37An administrator is planning to monitor GPU utilization across a large cluster. Which component should be deployed to collect metrics that are compatible with Prometheus?
38Which action is required when updating the NVIDIA driver on a node managed by the GPU Operator to ensure that running workloads are not interrupted abruptly?
39Which mechanism does the NVIDIA GPU Operator use to ensure that the driver installed on a worker node matches the specific architecture of the installed GPU hardware?
40When troubleshooting an NVIDIA GPU Operator installation, which TWO locations should an administrator check to identify why the driver installation pod is failing?
41Which component is responsible for exposing the GPU as a schedulable resource in a Kubernetes cluster?
42An administrator observes that despite the GPU Operator being installed, the pods cannot access the GPU. What is the most likely cause if the NVIDIA container runtime is properly configured?
43Which security configuration is necessary when deploying NVIDIA GPUs in a multi-tenant environment to prevent unauthorized access between containers?
44Which TWO of the following steps are essential when deploying the NVIDIA GPU Operator on a Kubernetes cluster to ensure that GPU resources are discoverable by the scheduler?
45Which component in the NVIDIA AI Enterprise stack is responsible for providing the necessary user-space libraries and binaries to run GPU-accelerated applications inside containers?
46During the deployment of an AI model training workload on a multi-node cluster, the administrator notices that inter-node communication is significantly slower than expected. Which deployment aspect should be investigated first?
47Refer to the exhibit. An administrator attempts to deploy a GPU-based pod, but it remains in the 'Pending' state. What is the most likely cause based on the error log?
48When deploying NVIDIA AI Enterprise, why is the use of the NVIDIA NGC Catalog recommended over public container repositories?
49An administrator is tasked with deploying a multi-node training job using the NVIDIA GPU Operator. Which configuration must be present to ensure that pods are scheduled on nodes with identical GPU architectures to prevent performance degradation?
50Which command is used to verify that the NVIDIA GPU Operator has successfully installed the necessary components on a Kubernetes node?
51An AI platform engineer is preparing a fleet of NVIDIA DGX H100 systems for production workloads using the NVIDIA Base Command Manager (BCM). During initial bare-metal provisioning via PXE boot, the provisioning server successfully hands out IP addresses, but nodes consistently fail during the OS image deployment phase, throwing a kernel panic related to missing storage drivers. Which deployment step must be verified or corrected to ensure successful hardware-specific image deployment?
52A platform team is deploying the NVIDIA GPU Operator on a Kubernetes cluster to manage GPU nodes. They want the Operator to automatically install the NVIDIA driver, the container toolkit, and the device plugin on each GPU node. Which component of the GPU Operator is responsible for installing the NVIDIA driver on the host?
53A platform engineer is preparing a Kubernetes cluster to run AI training jobs that require GPU access. The cluster nodes have NVIDIA GPUs, and the engineer wants the GPU Operator to manage the driver lifecycle. Which component must be installed on the host nodes to allow the GPU Operator to load kernel modules and create device nodes?
54An administrator is preparing a Kubernetes cluster for AI workloads and needs to ensure that the NVIDIA GPU Operator can be installed. The cluster nodes have NVIDIA GPUs, and the administrator wants to verify that the nodes are ready. Which command should the administrator run to check if the NVIDIA driver is already loaded on a node?
55A financial services company is deploying NVIDIA AI Enterprise on a Kubernetes cluster with strict security policies. They need to ensure that GPU workloads are isolated and that the NVIDIA GPU Operator components are deployed with least privilege. Which feature of the NVIDIA GPU Operator allows administrators to define granular permissions for its components?
56An AI operations engineer is preparing a Kubernetes cluster to run GPU-accelerated inference workloads using the NVIDIA GPU Operator. The cluster nodes already have NVIDIA data center GPUs installed, and the engineer wants to avoid installing the driver manually on each node. Which component of the GPU Operator is responsible for automatically deploying the NVIDIA driver on worker nodes?
57A platform engineer is preparing a bare-metal Kubernetes cluster to run GPU-accelerated AI workloads using the NVIDIA GPU Operator. The cluster nodes have NVIDIA Ampere GPUs and run Ubuntu 22.04 with containerd as the container runtime. The engineer wants to avoid installing any NVIDIA drivers or CUDA components directly on the host. Which GPU Operator configuration should be used to achieve this?
58A company is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing NVIDIA A100 GPUs. They need to enable vGPU functionality for virtual machines running AI workloads. Which configuration step is required on the ESXi host to allow vGPU assignment to VMs?
59A platform engineer is preparing a bare-metal Ubuntu 22.04 server with four A100 GPUs for an NVIDIA AI Enterprise deployment. The GPUs are not yet visible to the operating system tooling. Which command should the engineer run to confirm the driver loaded successfully and that all four GPUs are enumerated with their current driver version?
60A platform engineer is preparing an Ubuntu 22.04 server that will host GPU-accelerated inference containers managed by containerd (not Docker). The team wants the NVIDIA Container Toolkit to expose GPUs to those containers. After installing the toolkit packages, which action must the engineer take so that containerd actually invokes the NVIDIA runtime for GPU workloads?
61An AI operations team is deploying the NVIDIA GPU Operator on a Kubernetes cluster that uses containerd as the container runtime. The cluster nodes have NVIDIA GPUs, and the team wants to ensure that GPU workloads can request GPU resources. After installing the operator, they notice that pods requesting 'nvidia.com/gpu' remain in Pending state. Which component of the GPU Operator is most likely misconfigured or missing?
62A system administrator is installing the NVIDIA Container Toolkit on a standalone Ubuntu server to run GPU-accelerated containers. After installation, they want to verify that the toolkit is correctly configured. Which command should they run to test GPU access from a container?
63An AI operations team is installing the NVIDIA GPU Operator on a Kubernetes cluster that uses a custom containerd configuration. They need to ensure that the GPU Operator can properly manage the container runtime. Which action should they take before installing the GPU Operator?
64An AI operations team is validating a new Kubernetes cluster before installing the NVIDIA GPU Operator with the driver managed by the Operator itself. The nodes run a supported Linux distribution with GPUs physically installed. Which two conditions must be satisfied for the Operator's driver container to build and load the kernel module successfully? (Choose two.)
65An AI infrastructure team is deploying NVIDIA AI Enterprise on a Kubernetes cluster using the NVIDIA GPU Operator. They need to ensure that the GPU Operator can successfully manage GPUs and that workloads can consume GPU resources. Which two components does the GPU Operator deploy to enable GPU scheduling and container GPU access? (Choose two.)
66A cloud operations engineer is deploying the NVIDIA GPU Operator on a managed Kubernetes service where the worker nodes already have the NVIDIA data center driver installed by the cloud provider. The team wants the Operator to manage only the device plugin, container toolkit, and monitoring components. Which Helm value should the engineer set during installation?
67A cloud architect is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing NVIDIA A100 GPUs. The architect must enable GPU virtualization using NVIDIA vGPU. Which two components are required to support vGPU on the ESXi hosts? (Choose two.)
68A financial services company is deploying NVIDIA AI Enterprise in an air-gapped data center. They need to install the NVIDIA GPU Operator on their Kubernetes cluster without internet access. Which additional step must they take to ensure a successful installation?
69An AI infrastructure team is preparing an air-gapped data center to install the NVIDIA GPU Operator. They have mirrored all required container images into a private registry. During installation, the Operator's pods fail with ImagePullBackOff because the components still reference images under nvcr.io. Which configuration is required to make the Operator and its managed components pull from the private registry?
70A DevOps engineer is deploying the NVIDIA GPU Operator on a Kubernetes cluster that uses containerd as the container runtime. The engineer notices that the Operator's validation pod fails with an error indicating that the NVIDIA container runtime is not configured. Which action should the engineer take to resolve this?
71A platform team is preparing a bare-metal Kubernetes cluster to run AI workloads. They want the GPU Operator to install and manage the NVIDIA driver automatically on each node. Which prerequisite must be satisfied on every worker node before the GPU Operator can succeed?
72A media company is deploying an inference service on a Kubernetes cluster with the NVIDIA GPU Operator installed. The service pods remain in Pending with the message that no nodes have the requested nvidia.com/gpu resource, even though the GPUs are healthy and the driver loads correctly on every node. Which troubleshooting step should the engineer perform first?
73An administrator is installing the NVIDIA GPU Operator on a Kubernetes cluster. They want to verify that the GPU Operator's components are running correctly after installation. Which command should they use to check the status of the GPU Operator pods?
74An AI operations team is deploying NVIDIA AI Enterprise on a bare-metal Kubernetes cluster with DGX A100 systems. They need to enable GPUDirect Storage to accelerate data loading from a local NVMe array. Which component must be installed and configured on the DGX nodes to support GPUDirect Storage?
75An administrator is deploying the NVIDIA GPU Operator into an existing Kubernetes cluster where the NVIDIA driver is already installed and maintained by the node image. The team wants the Operator to manage only the container runtime, device plugin, and monitoring components. Which configuration should be applied to the GPU Operator's ClusterPolicy?
76An administrator is deploying the NVIDIA GPU Operator on a Kubernetes cluster in an air-gapped environment. The cluster nodes have no internet access, and all container images must be pulled from a private registry. Which two actions are required to ensure a successful deployment? (Choose two.)
77A system administrator is installing NVIDIA AI Enterprise on a Kubernetes cluster that will use Multi-Instance GPU (MIG) on A100 GPUs. The administrator wants to ensure that MIG instances are properly exposed as schedulable resources. Which action must be taken after enabling MIG mode on the GPUs?
78A platform engineer must validate a new NVIDIA GPU Operator deployment on a Kubernetes cluster before handing it to data scientists. Which two checks confirm that the Operator has correctly exposed GPU resources to the cluster scheduler? (Choose two.)
79A systems administrator is installing the NVIDIA Container Toolkit on a stand-alone server running Ubuntu 22.04 to enable Docker containers to access NVIDIA GPUs. After installation, they run a test container and find that it cannot see the GPU. Which step is most likely missing?
80A platform team is installing the NVIDIA GPU Operator on a Kubernetes cluster that runs a mix of GPU and non-GPU nodes. They want the operator to manage the driver lifecycle only on nodes that actually have NVIDIA GPUs, without requiring manual taints on non-GPU nodes. Which configuration should they apply to the ClusterPolicy to achieve this?
81An AI operations engineer is deploying a multi-node Kubernetes cluster with NVIDIA A100 GPUs for distributed training. The engineer wants to ensure that GPUs are correctly discovered and that workloads can request GPU resources. After installing the NVIDIA GPU Operator, the engineer notices that the GPU nodes are not advertising any 'nvidia.com/gpu' resources. Which component should the engineer verify first to resolve this issue?
82An administrator is preparing to install the NVIDIA GPU Operator on a new Kubernetes cluster. The cluster uses containerd as the container runtime. Which prerequisite must be satisfied on each GPU node before the operator can successfully deploy the driver container?
83An AI operations engineer is preparing a Kubernetes cluster to run GPU-accelerated workloads using the NVIDIA GPU Operator. The cluster nodes already have NVIDIA data center GPUs installed and the NVIDIA driver is pre-installed on the host. The engineer wants to use the GPU Operator to manage the container toolkit, device plugin, and monitoring components but must avoid the Operator managing or upgrading the driver. Which configuration should be applied to the GPU Operator deployment?
84A healthcare company is deploying NVIDIA AI Enterprise on a Kubernetes cluster to run medical imaging AI models. The cluster administrator needs to verify that the NVIDIA GPU Operator is installed and functioning correctly. Which command should the administrator use to check the status of the GPU Operator pods?
85An administrator is deploying NVIDIA AI Enterprise on a Kubernetes cluster and must decide how GPU workloads should request accelerators. The environment has a mix of full-GPU training jobs and inference services that share a single A100. Which approach correctly allows a pod to consume a specific MIG-backed slice rather than the whole device?
86A cloud operations team is deploying the NVIDIA GPU Operator in an environment where the Kubernetes control plane cannot reach the public internet, but worker nodes can access an internal HTTP registry that mirrors required images. The team wants to avoid manual image pulls on each node. Which two configurations should they implement to enable a successful air-gapped installation? (Choose two.)
87During an air-gapped installation of the NVIDIA GPU Operator, the administrator must make all required images available to the cluster. Which component is responsible for pulling the Operator's operand images from the private registry?
88A financial services company is deploying NVIDIA AI Enterprise on a VMware vSphere cluster with NVIDIA A100 GPUs. The security team requires that GPU workloads be isolated at the hardware level, with separate memory and fault domains, to meet regulatory compliance. The company also wants to maximize GPU utilization by running multiple workloads concurrently. Which NVIDIA feature should be enabled to meet these requirements?
89A Kubernetes cluster administrator is installing the NVIDIA GPU Operator and wants to ensure that GPU workloads are scheduled only on nodes with healthy GPUs. The administrator plans to use the operator's built-in health checks. Which component is responsible for monitoring GPU health and marking nodes as unschedulable when a GPU fails?
90An administrator is deploying the NVIDIA GPU Operator on a Kubernetes cluster using Helm. The cluster nodes have NVIDIA GPUs and the administrator wants to ensure that the GPU Operator can automatically label nodes with GPU properties and install the device plugin. Which prerequisite must be met on the cluster nodes before installing the GPU Operator?
91A research lab is deploying NVIDIA AI Enterprise on an air-gapped Kubernetes cluster. The cluster has no internet access, and all software must be installed from a local registry. The administrator plans to use the NVIDIA GPU Operator. Which two actions must be performed to ensure a successful deployment in this environment? (Choose two.)
92An AI operations team is deploying NVIDIA AI Enterprise on a Kubernetes cluster using the NVIDIA GPU Operator. They need to ensure that GPU metrics such as utilization, memory usage, and temperature are collected and exposed to Prometheus for monitoring. Which component of the GPU Operator is responsible for this?
Be able to deploy the NVIDIA GPU Operator on Kubernetes and confirm drivers, container toolkit, and device plugin are healthy, then validate multi-node NCCL/RDMA performance and read nvidia-smi output. The key skill is tracing a symptom to the correct deployment layer before changing anything.
The Courseiva NCP-AIO question bank contains 92 questions in the Installation and Deployment domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Installation and Deployment domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included