NCP-AIO · domain
Installation and Deployment
This domain covers standing up and validating NVIDIA-accelerated infrastructure: installing the GPU Operator on Kubernetes, provisioning drivers and the container toolkit, configuring multi-node fabrics like NCCL and GPUDirect RDMA, and diagnosing pods stuck Pending or GPUs underperforming after deployment. Questions are scenario-based, pairing a symptom with the deployment layer or tool responsible.
Focused practice
Practice Installation and Deployment questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Installation and Deployment
Be able to deploy the NVIDIA GPU Operator on Kubernetes and confirm drivers, container toolkit, and device plugin are healthy, then validate multi-node NCCL/RDMA performance and read nvidia-smi output. The key skill is tracing a symptom to the correct deployment layer before changing anything.
NVIDIA GPU Operator deployment on Kubernetes, including automatic driver, container toolkit, and device plugin installation on GPU nodes
Multi-node interconnect configuration and validation using NCCL, InfiniBand, and GPUDirect RDMA for distributed training
Real-time GPU monitoring and management with nvidia-smi for power, temperature, utilization, and memory statistics
Diagnosing GPU pod scheduling failures, including Pending states caused by missing device plugin, taints, or resource requests
Watch out for
Common Installation and Deployment exam traps
- ▸Blaming application code or model configuration for slow inter-node throughput instead of checking the network fabric, NCCL settings, and RDMA path
- ▸Assuming the GPU Operator installs everything by default without verifying driver, toolkit, and device plugin DaemonSets are actually running
- ▸Debugging a Pending GPU pod as a hardware fault when the real cause is an unschedulable node, missing nvidia.com/gpu resource, or unready device plugin
Question index
All Installation and Deployment questions (92)
Click any question to see the full explanation, or start a practice session above.
A media company is deploying an inference service on a Kubernetes cluster with the NVIDIA GPU Operator installed. The service pods remain in Pending with the message that no nodes have the requested nvidia.com/gpu resource, even though the GPUs are healthy and the driver loads correctly on every node. Which troubleshooting step should the engineer perform first?
Medium2A DevOps engineer is deploying the NVIDIA GPU Operator on a Kubernetes cluster that uses containerd as the container runtime. The engineer notices that the Operator's validation pod fails with an error indicating that the NVIDIA container runtime is not configured. Which action should the engineer take to resolve this?
Easy3When installing the NVIDIA Container Toolkit to enable GPU acceleration in Docker, which file must be modified or verified to ensure the container runtime can access the NVIDIA runtime?
Easy4A systems administrator is installing the NVIDIA Container Toolkit on a stand-alone server running Ubuntu 22.04 to enable Docker containers to access NVIDIA GPUs. After installation, they run a test container and find that it cannot see the GPU. Which step is most likely missing?
Easy5A platform engineer is preparing an Ubuntu 22.04 server that will host GPU-accelerated inference containers managed by containerd (not Docker). The team wants the NVIDIA Container Toolkit to expose GPUs to those containers. After installing the toolkit packages, which action must the engineer take so that containerd actually invokes the NVIDIA runtime for GPU workloads?
Medium6You are troubleshooting a node where the GPU is detected, but the application fails to utilize it. Which log source would provide the most relevant information?
Medium7In an air-gapped environment, what must an administrator do to ensure the GPU Operator correctly installs the necessary software components?
Hard8Which TWO of the following are prerequisites for installing the NVIDIA Container Toolkit on a Linux host?
Medium9An AI operations team is deploying the NVIDIA GPU Operator on a Kubernetes cluster that uses containerd as the container runtime. The cluster nodes have NVIDIA GPUs, and the team wants to ensure that GPU workloads can request GPU resources. After installing the operator, they notice that pods requesting 'nvidia.com/gpu' remain in Pending state. Which component of the GPU Operator is most likely misconfigured or missing?
Medium10When troubleshooting an NVIDIA GPU Operator installation, which TWO locations should an administrator check to identify why the driver installation pod is failing?
Hard11Which configuration file is typically modified to enable the NVIDIA Device Plugin in a Kubernetes cluster?
Medium12An AI infrastructure team is preparing an air-gapped data center to install the NVIDIA GPU Operator. They have mirrored all required container images into a private registry. During installation, the Operator's pods fail with ImagePullBackOff because the components still reference images under nvcr.io. Which configuration is required to make the Operator and its managed components pull from the private registry?
Hard13A platform engineer is preparing a bare-metal Ubuntu 22.04 server with four A100 GPUs for an NVIDIA AI Enterprise deployment. The GPUs are not yet visible to the operating system tooling. Which command should the engineer run to confirm the driver loaded successfully and that all four GPUs are enumerated with their current driver version?
Easy14A financial services company is deploying NVIDIA AI Enterprise on a VMware vSphere cluster with NVIDIA A100 GPUs. The security team requires that GPU workloads be isolated at the hardware level, with separate memory and fault domains, to meet regulatory compliance. The company also wants to maximize GPU utilization by running multiple workloads concurrently. Which NVIDIA feature should be enabled to meet these requirements?
Hard15Which NVIDIA tool allows you to verify that the GPU and its driver are properly installed and functioning on a Linux system?
Easy16Which file format is commonly used to define the configuration and state for the NVIDIA GPU Operator within a Kubernetes environment?
Medium17When deploying NVIDIA AI Enterprise, why is the use of the NVIDIA NGC Catalog recommended over public container repositories?
Medium18Which tool is the industry standard for monitoring and managing NVIDIA data center GPUs in a large-scale cluster deployment?
Medium19Which action is required when updating the NVIDIA driver on a node managed by the GPU Operator to ensure that running workloads are not interrupted abruptly?
Medium20An AI operations team is validating a new Kubernetes cluster before installing the NVIDIA GPU Operator with the driver managed by the Operator itself. The nodes run a supported Linux distribution with GPUs physically installed. Which two conditions must be satisfied for the Operator's driver container to build and load the kernel module successfully? (Choose two.)
Hard21What is the primary role of the NVIDIA Data Center GPU Manager (DCGM) Exporter in a cloud-native monitoring stack?
Easy22An AI operations engineer is preparing a Kubernetes cluster to run GPU-accelerated workloads using the NVIDIA GPU Operator. The cluster nodes already have NVIDIA data center GPUs installed and the NVIDIA driver is pre-installed on the host. The engineer wants to use the GPU Operator to manage the container toolkit, device plugin, and monitoring components but must avoid the Operator managing or upgrading the driver. Which configuration should be applied to the GPU Operator deployment?
Medium23Which command is used to verify that the NVIDIA GPU Operator has successfully installed the necessary components on a Kubernetes node?
Easy24A Kubernetes cluster administrator is installing the NVIDIA GPU Operator and wants to ensure that GPU workloads are scheduled only on nodes with healthy GPUs. The administrator plans to use the operator's built-in health checks. Which component is responsible for monitoring GPU health and marking nodes as unschedulable when a GPU fails?
Medium25Which of the following describes the purpose of the NVIDIA GPU Operator's 'Driver Container'?
Medium26When installing NVIDIA drivers via a package manager, what is the importance of the 'dkms' package?
Easy27When installing the NVIDIA GPU Operator, which namespace is typically used to ensure proper isolation and role-based access control?
Easy28Which action must be performed after updating the NVIDIA driver on a Linux host to ensure that all active GPU containers recognize the new driver version?
Medium29A healthcare company is deploying NVIDIA AI Enterprise on a Kubernetes cluster to run medical imaging AI models. The cluster administrator needs to verify that the NVIDIA GPU Operator is installed and functioning correctly. Which command should the administrator use to check the status of the GPU Operator pods?
Easy30A platform engineer must validate a new NVIDIA GPU Operator deployment on a Kubernetes cluster before handing it to data scientists. Which two checks confirm that the Operator has correctly exposed GPU resources to the cluster scheduler? (Choose two.)
Hard31When deploying the NVIDIA GPU Operator in a restricted-access environment (air-gapped), which THREE requirements must be addressed to ensure a successful installation?
Hard32Which NVIDIA technology enables the partitioning of a single physical GPU into multiple independent instances for use by different virtual machines or containers?
Medium33Which NVIDIA software component is responsible for providing the necessary CUDA libraries to containerized applications?
Easy34An AI engineer is deploying a large language model on an NVIDIA DGX system. The deployment fails with an error indicating an insufficient NVIDIA driver version for the required CUDA toolkit. Which action should the engineer take to resolve the dependency mismatch?
Medium35Which THREE of the following are benefits of using Multi-Instance GPU (MIG) technology in a Kubernetes environment?
Hard36An AI operations team is deploying NVIDIA AI Enterprise on a bare-metal Kubernetes cluster with DGX A100 systems. They need to enable GPUDirect Storage to accelerate data loading from a local NVMe array. Which component must be installed and configured on the DGX nodes to support GPUDirect Storage?
Hard37Which mechanism does the NVIDIA GPU Operator use to ensure that the driver installed on a worker node matches the specific architecture of the installed GPU hardware?
Medium38An administrator is deploying NVIDIA AI Enterprise on a bare-metal cluster. Which component is mandatory to provide the necessary abstraction layer for containerized GPU resources?
Medium39A system administrator is installing the NVIDIA Container Toolkit on a standalone Ubuntu server to run GPU-accelerated containers. After installation, they want to verify that the toolkit is correctly configured. Which command should they run to test GPU access from a container?
Easy40During the deployment of an AI model training workload on a multi-node cluster, the administrator notices that inter-node communication is significantly slower than expected. Which deployment aspect should be investigated first?
Medium41A platform engineer is preparing a bare-metal Kubernetes cluster to run GPU-accelerated AI workloads using the NVIDIA GPU Operator. The cluster nodes have NVIDIA Ampere GPUs and run Ubuntu 22.04 with containerd as the container runtime. The engineer wants to avoid installing any NVIDIA drivers or CUDA components directly on the host. Which GPU Operator configuration should be used to achieve this?
Medium42A financial services company is deploying NVIDIA AI Enterprise in an air-gapped data center. They need to install the NVIDIA GPU Operator on their Kubernetes cluster without internet access. Which additional step must they take to ensure a successful installation?
Medium43An administrator is deploying NVIDIA AI Enterprise on a Kubernetes cluster and must decide how GPU workloads should request accelerators. The environment has a mix of full-GPU training jobs and inference services that share a single A100. Which approach correctly allows a pod to consume a specific MIG-backed slice rather than the whole device?
Hard44Which component is responsible for exposing the GPU as a schedulable resource in a Kubernetes cluster?
Medium45A research lab is deploying NVIDIA AI Enterprise on an air-gapped Kubernetes cluster. The cluster has no internet access, and all software must be installed from a local registry. The administrator plans to use the NVIDIA GPU Operator. Which two actions must be performed to ensure a successful deployment in this environment? (Choose two.)
Hard46An administrator is preparing a Kubernetes cluster for AI workloads and needs to ensure that the NVIDIA GPU Operator can be installed. The cluster nodes have NVIDIA GPUs, and the administrator wants to verify that the nodes are ready. Which command should the administrator run to check if the NVIDIA driver is already loaded on a node?
Easy47A system administrator is installing NVIDIA AI Enterprise on a Kubernetes cluster that will use Multi-Instance GPU (MIG) on A100 GPUs. The administrator wants to ensure that MIG instances are properly exposed as schedulable resources. Which action must be taken after enabling MIG mode on the GPUs?
Medium48An administrator is deploying the NVIDIA GPU Operator on a Kubernetes cluster using Helm. The cluster nodes have NVIDIA GPUs and the administrator wants to ensure that the GPU Operator can automatically label nodes with GPU properties and install the device plugin. Which prerequisite must be met on the cluster nodes before installing the GPU Operator?
Easy49An administrator is installing the NVIDIA GPU Operator on a Kubernetes cluster. They want to verify that the GPU Operator's components are running correctly after installation. Which command should they use to check the status of the GPU Operator pods?
Easy50A cloud operations engineer is deploying the NVIDIA GPU Operator on a managed Kubernetes service where the worker nodes already have the NVIDIA data center driver installed by the cloud provider. The team wants the Operator to manage only the device plugin, container toolkit, and monitoring components. Which Helm value should the engineer set during installation?
Easy51A financial services company is deploying NVIDIA AI Enterprise on a Kubernetes cluster with strict security policies. They need to ensure that GPU workloads are isolated and that the NVIDIA GPU Operator components are deployed with least privilege. Which feature of the NVIDIA GPU Operator allows administrators to define granular permissions for its components?
Hard52An administrator is deploying NVIDIA AI Enterprise on a bare-metal cluster. The workload requires full GPU isolation with minimal latency. Which configuration should the administrator select to achieve this goal?
Medium53Refer to the exhibit. An administrator notices poor performance in an AI training job. What is the most likely cause based on the CLI output?
Medium54During an NVIDIA AI Enterprise deployment, you are asked to configure the 'NVIDIA Container Toolkit'. What is its primary function?
Easy55Refer to the exhibit. An AI engineer observes that a model training job is running slower than expected. Based on the output, what is the primary cause of the performance degradation?
Medium56When deploying NVIDIA containers using the NVIDIA Container Toolkit, what is the primary function of the 'nvidia-container-runtime'?
Easy57An AI operations team is installing the NVIDIA GPU Operator on a Kubernetes cluster that uses a custom containerd configuration. They need to ensure that the GPU Operator can properly manage the container runtime. Which action should they take before installing the GPU Operator?
Medium58When configuring the NVIDIA Device Plugin for Kubernetes, what is the purpose of the 'time-slicing' configuration?
Medium59During an air-gapped installation of the NVIDIA GPU Operator, the administrator must make all required images available to the cluster. Which component is responsible for pulling the Operator's operand images from the private registry?
Medium60An AI platform engineer is preparing a fleet of NVIDIA DGX H100 systems for production workloads using the NVIDIA Base Command Manager (BCM). During initial bare-metal provisioning via PXE boot, the provisioning server successfully hands out IP addresses, but nodes consistently fail during the OS image deployment phase, throwing a kernel panic related to missing storage drivers. Which deployment step must be verified or corrected to ensure successful hardware-specific image deployment?
Medium61Which component is strictly necessary for managing NVIDIA AI Enterprise licensing across a distributed cluster of nodes?
Easy62Refer to the exhibit. The administrator has deployed the GPU Operator, but the node does not show GPU resources. What is the most likely cause?
Medium63A cloud operations team is deploying the NVIDIA GPU Operator in an environment where the Kubernetes control plane cannot reach the public internet, but worker nodes can access an internal HTTP registry that mirrors required images. The team wants to avoid manual image pulls on each node. Which two configurations should they implement to enable a successful air-gapped installation? (Choose two.)
Hard64An administrator observes that despite the GPU Operator being installed, the pods cannot access the GPU. What is the most likely cause if the NVIDIA container runtime is properly configured?
Hard65What is the primary role of a private container registry in an NVIDIA AI Enterprise deployment?
Medium66A company is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing NVIDIA A100 GPUs. They need to enable vGPU functionality for virtual machines running AI workloads. Which configuration step is required on the ESXi host to allow vGPU assignment to VMs?
Hard67A platform team is deploying the NVIDIA GPU Operator on a Kubernetes cluster to manage GPU nodes. They want the Operator to automatically install the NVIDIA driver, the container toolkit, and the device plugin on each GPU node. Which component of the GPU Operator is responsible for installing the NVIDIA driver on the host?
Medium68Which utility is primarily used to monitor and manage NVIDIA GPU power, temperature, and usage statistics in real-time on a Linux-based deployment?
Easy69Refer to the exhibit. An administrator attempts to deploy a GPU-based pod, but it remains in the 'Pending' state. What is the most likely cause based on the error log?
Medium70An administrator is preparing to install the NVIDIA GPU Operator on a new Kubernetes cluster. The cluster uses containerd as the container runtime. Which prerequisite must be satisfied on each GPU node before the operator can successfully deploy the driver container?
Easy71An administrator is preparing to deploy an NVIDIA AI Enterprise solution on an OpenShift cluster. Which TWO steps must be completed to ensure the NVIDIA drivers are loaded correctly on the worker nodes?
Medium72An AI operations engineer is preparing a Kubernetes cluster to run GPU-accelerated inference workloads using the NVIDIA GPU Operator. The cluster nodes already have NVIDIA data center GPUs installed, and the engineer wants to avoid installing the driver manually on each node. Which component of the GPU Operator is responsible for automatically deploying the NVIDIA driver on worker nodes?
Easy73A platform engineer is preparing a Kubernetes cluster to run AI training jobs that require GPU access. The cluster nodes have NVIDIA GPUs, and the engineer wants the GPU Operator to manage the driver lifecycle. Which component must be installed on the host nodes to allow the GPU Operator to load kernel modules and create device nodes?
Medium74An administrator is planning to monitor GPU utilization across a large cluster. Which component should be deployed to collect metrics that are compatible with Prometheus?
Medium75What is the primary benefit of using an NVIDIA-Certified System for AI Enterprise deployments?
Medium76When deploying NVIDIA AI Enterprise, why is the selection of the correct CUDA version in the container image critical during the installation phase?
Hard77An AI operations engineer is deploying a multi-node Kubernetes cluster with NVIDIA A100 GPUs for distributed training. The engineer wants to ensure that GPUs are correctly discovered and that workloads can request GPU resources. After installing the NVIDIA GPU Operator, the engineer notices that the GPU nodes are not advertising any 'nvidia.com/gpu' resources. Which component should the engineer verify first to resolve this issue?
Hard78Which TWO of the following steps are essential when deploying the NVIDIA GPU Operator on a Kubernetes cluster to ensure that GPU resources are discoverable by the scheduler?
Medium79Which security configuration is necessary when deploying NVIDIA GPUs in a multi-tenant environment to prevent unauthorized access between containers?
Medium80What is the primary function of the NVIDIA Persistence Daemon in an AI deployment?
Medium81An AI operations team is deploying NVIDIA AI Enterprise on a Kubernetes cluster using the NVIDIA GPU Operator. They need to ensure that GPU metrics such as utilization, memory usage, and temperature are collected and exposed to Prometheus for monitoring. Which component of the GPU Operator is responsible for this?
Medium82An administrator is deploying the NVIDIA GPU Operator on a Kubernetes cluster in an air-gapped environment. The cluster nodes have no internet access, and all container images must be pulled from a private registry. Which two actions are required to ensure a successful deployment? (Choose two.)
Medium83When configuring a node for NVIDIA AI Enterprise in a Kubernetes environment, what is the primary function of the NVIDIA Container Toolkit?
Medium84A platform team is preparing a bare-metal Kubernetes cluster to run AI workloads. They want the GPU Operator to install and manage the NVIDIA driver automatically on each node. Which prerequisite must be satisfied on every worker node before the GPU Operator can succeed?
Easy85A cloud architect is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing NVIDIA A100 GPUs. The architect must enable GPU virtualization using NVIDIA vGPU. Which two components are required to support vGPU on the ESXi hosts? (Choose two.)
Hard86When configuring a multi-tenant environment on an NVIDIA DGX system, how are MIG (Multi-Instance GPU) instances best provisioned?
Hard87A platform team is installing the NVIDIA GPU Operator on a Kubernetes cluster that runs a mix of GPU and non-GPU nodes. They want the operator to manage the driver lifecycle only on nodes that actually have NVIDIA GPUs, without requiring manual taints on non-GPU nodes. Which configuration should they apply to the ClusterPolicy to achieve this?
Medium88An administrator is tasked with deploying a multi-node training job using the NVIDIA GPU Operator. Which configuration must be present to ensure that pods are scheduled on nodes with identical GPU architectures to prevent performance degradation?
Hard89An AI infrastructure team is deploying NVIDIA AI Enterprise on a Kubernetes cluster using the NVIDIA GPU Operator. They need to ensure that the GPU Operator can successfully manage GPUs and that workloads can consume GPU resources. Which two components does the GPU Operator deploy to enable GPU scheduling and container GPU access? (Choose two.)
Hard90An administrator is deploying the NVIDIA GPU Operator into an existing Kubernetes cluster where the NVIDIA driver is already installed and maintained by the node image. The team wants the Operator to manage only the container runtime, device plugin, and monitoring components. Which configuration should be applied to the GPU Operator's ClusterPolicy?
Medium91An administrator is deploying NVIDIA AI Enterprise on a bare-metal cluster. Which component must be installed first to ensure proper communication between the Kubernetes scheduler and the underlying GPU hardware?
Medium92Which component in the NVIDIA AI Enterprise stack is responsible for providing the necessary user-space libraries and binaries to run GPU-accelerated applications inside containers?
EasyOther domains
All NCP-AIO exam domains
Frequently asked questions
- What does the Installation and Deployment domain cover on the NCP-AIO exam?
- Be able to deploy the NVIDIA GPU Operator on Kubernetes and confirm drivers, container toolkit, and device plugin are healthy, then validate multi-node NCCL/RDMA performance and read nvidia-smi output. The key skill is tracing a symptom to the correct deployment layer before changing anything.
- How many questions are in this domain?
- This page lists all 92 Installation and Deployment questions in the NCP-AIO question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Installation and Deployment questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.