Be able to deploy GPU-enabled containers on Kubernetes and configure MIG or vGPU for isolated sharing. The single most important thing: know which NVIDIA technology provides strict hardware isolation versus mere time-slicing, and verify the full container GPU stack is installed end to end.
Start practicing
Administration — choose a session length
Free · No account required
Domain overview
The Administration domain of the NCP-AIO exam covers deploying, securing, and operating NVIDIA GPU infrastructure for AI workloads across containers, Kubernetes, and virtualized environments. It tests practical configuration of GPU sharing, container runtimes, cluster performance tuning, and security policy application, requiring administrators to reason about the immediate operational effects of their choices.
Exam objectives
Configuring NVIDIA Container Toolkit and device plugin for GPU access in Kubernetes pods
Applying Multi-Instance GPU (MIG) and vGPU for hardware-isolated GPU sharing across tenants
Using NVIDIA Base Command Manager to monitor and tune training cluster performance
Interpreting security policy effects on containerized AI workloads at runtime
Assuming GPU sharing always provides isolation; only MIG or vGPU enforce strict hardware partitioning, while time-slicing does not.
Forgetting that the NVIDIA device plugin and container runtime hooks must both be installed before pods can request GPUs.
Treating Base Command as only a scheduler; it also handles node health, fabric monitoring, and performance consistency.
Click any question to see the full explanation and answer options, or start a focused practice session above.
An administrator is configuring a multi-tenant NVIDIA AI Enterprise environment. Which mechanism is most effective for ensuring hardware-level isolation between concurrent training jobs on a single A100 GPU?
2An administrator is optimizing a cluster for AI model training using NVIDIA Base Command. Which TWO tasks are critical for ensuring consistent performance across the training nodes?
3Refer to the exhibit. An administrator applies this security policy to a container runtime environment. What is the immediate effect on containerized AI applications within this scope?
4When monitoring GPU health in an enterprise cluster, which command provides the most comprehensive snapshot of real-time power, temperature, and memory utilization?
5An administrator notices that GPU utilization is high, but throughput in an AI training job remains low. What is the most likely bottleneck?
6Refer to the exhibit. An administrator observes this in the logs during a multi-GPU training job. What is the performance implication of this setting?
7Which NVIDIA technology enables a GPU to be shared among multiple virtual machines or containers while maintaining strict hardware isolation?
8An administrator needs to ensure that all GPU drivers are updated across a heterogeneous cluster without causing downtime. What is the best strategy?
9Which THREE components are required for a container to successfully leverage NVIDIA GPUs on a Kubernetes cluster?
10An administrator is configuring a new cluster and wants to ensure that telemetry data from GPUs is collected in a centralized manner. Which tool is best suited for this requirement?
11An administrator wants to prevent unauthorized users from accessing sensitive model weights stored in GPU memory. Which security feature should be implemented to ensure hardware-level isolation of the memory space?
12Which administrative practice ensures that a cluster is prepared for the arrival of new NVIDIA GPU hardware with minimal downtime?
13An administrator manages an NVIDIA AI Enterprise cluster running multiple Kubernetes nodes, each with several A100 GPUs. After upgrading the NVIDIA GPU Operator to a newer version, the administrator notices that pods requesting GPUs remain in a Pending state, and the node's allocatable GPU count is reported as zero. Which command should the administrator run first to diagnose the issue?
14An administrator manages an NVIDIA AI Enterprise cluster and needs to enforce GPU resource quotas across multiple Kubernetes namespaces. Which NVIDIA component should be configured to enforce these quotas?
15An administrator is deploying NVIDIA AI Enterprise on a Kubernetes cluster and wants to verify that the GPU Operator has successfully installed all required components. Which command should the administrator use to check the status of the GPU Operator pods?
16An administrator manages an NVIDIA AI Enterprise cluster using NVIDIA Run:ai. A data science team complains that their submitted training job has been stuck in a Pending state for over an hour, even though the Run:ai scheduler shows free GPUs in the cluster. The administrator verifies that the job requests 2 GPUs and the node pool has 4 idle GPUs. Which Run:ai administrative configuration is the most likely cause of the job remaining Pending?
17An administrator is configuring NVIDIA Base Command Manager to manage a cluster of DGX nodes. They want to ensure that when a node's GPU temperature exceeds a defined threshold, the node is automatically drained and an alert is sent to the operations team. Which combination of Base Command Manager features should the administrator configure to achieve this?
18An administrator is preparing a bare-metal GPU server for AI workloads and needs to verify that the NVIDIA driver and CUDA toolkit are properly installed. The server has an NVIDIA A100 GPU. Which command should the administrator run to display the GPU model, driver version, and CUDA version?
19An administrator is responsible for maintaining an NVIDIA AI Enterprise cluster and needs to ensure high availability of GPU resources for critical inference workloads. Which two practices should the administrator implement? (Choose two.)
20An administrator is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing several NVIDIA A100 GPUs. They need to enable GPU sharing across multiple virtual machines to maximize utilization. Which vSphere feature should they configure?
21An administrator is preparing a multi-node NVIDIA DGX H100 cluster for a distributed training job using NVIDIA Base Command. The cluster nodes have InfiniBand adapters, but the job's inter-node throughput is far below expectations. The administrator runs `ibstat` and sees that the ports are in the INIT state rather than ACTIVE. Which action should the administrator take first?
22An administrator needs to collect GPU telemetry from an NVIDIA AI Enterprise cluster and store it in a time-series database for long-term analysis. Which component should be deployed to export GPU metrics in Prometheus format?
23An administrator manages an NVIDIA AI Enterprise deployment on Kubernetes and needs to ensure that GPU telemetry is exported to an existing Prometheus instance. The administrator deploys the NVIDIA DCGM Exporter but sees no GPU metrics in Prometheus. Which configuration should the administrator verify first?
24An administrator supports a shared Kubernetes cluster running NVIDIA GPU Operator. Data scientists report that their inference pods remain in Pending state, yet the GPU Operator pods and node feature discovery pods are healthy, and the GPU nodes show no hardware alarms. The administrator confirms that the cluster has a mixture of MIG-capable A100 nodes and non-MIG T4 nodes. Which immediate administrative action is most appropriate to diagnose the scheduling failure?
25An administrator is setting up an NVIDIA AI Enterprise cluster and wants to verify that the NVIDIA GPU Operator has successfully deployed all required components on a worker node. Which command should the administrator use to list the GPU Operator pods running on that node?
26An AI operations engineer is preparing a DGX H100 system for a multi-node training workload. The engineer runs `nvidia-smi topo -m` and notices that GPU4 and GPU5 report a connection type of SYS, while all other GPU pairs show NV18. What is the most likely cause of this topology anomaly?
27An administrator is setting up NVIDIA Base Command Manager to provision and manage a new AI cluster. They need to ensure that the cluster can automatically discover and configure new GPU nodes. Which component is responsible for node discovery and initial configuration?
28An administrator is preparing a cluster for a new large language model training job that will use NVIDIA Magnum IO GPUDirect Storage to stream training data directly from a parallel file system to GPU memory. The administrator must verify that the environment supports GPUDirect Storage before the job starts. (Choose two.)
29An administrator is responsible for monitoring a large-scale AI cluster with hundreds of NVIDIA GPUs. They need to collect telemetry data such as GPU utilization, temperature, and power consumption from all nodes and store it centrally for analysis and alerting. Which NVIDIA tool should they use to collect and export GPU metrics to a monitoring system like Prometheus?
30An AI operations team is using NVIDIA DCGM (Data Center GPU Manager) to monitor a cluster of A100 GPUs. They want to set up proactive health checks to detect and mitigate GPU issues before they cause job failures. Which two DCGM features should they configure? (Choose two.)
31An administrator is troubleshooting a multi-node NVIDIA GPU training job that intermittently hangs during the all-reduce phase. The job uses NCCL over InfiniBand. Logs show that some ranks time out while others complete. The administrator suspects a network fabric issue. Which action should the administrator take first to isolate whether the problem is in the InfiniBand fabric or in the NCCL configuration?
32An administrator is responsible for a large NVIDIA DGX SuperPOD used for multi-node training. They need to ensure that GPU telemetry and health metrics are collected centrally and can trigger alerts when GPUs exceed temperature thresholds. Which component of NVIDIA Base Command Manager (BCM) should they configure to achieve this?
33An AI operations team is deploying NVIDIA Base Command Manager to manage a cluster of DGX nodes. They want to ensure that only authorized users can submit jobs and that all job submissions are audited. Which combination of Base Command Manager features should the administrator configure to meet these requirements?
34A company runs multiple AI workloads on a shared Kubernetes cluster with NVIDIA GPUs. The administrator needs to enforce that only pods with a specific label can consume GPU resources, while other pods are denied. Which Kubernetes admission control mechanism should be used to implement this policy?
35A site reliability engineer is troubleshooting a DGX A100 node that intermittently drops out of the cluster during large NCCL all-reduce jobs. `nvidia-smi` shows all eight A100 GPUs healthy, but DCGM reports XID errors 74 and 79 on one GPU during the failures. The engineer needs to determine the most likely cause and the correct administrative action. Which combination best describes the cause and the appropriate first step?
36An administrator is responsible for maintaining a fleet of NVIDIA-certified servers running AI workloads. They need to quickly identify which servers have GPUs that are overheating and may throttle performance. Which NVIDIA tool should the administrator use to monitor GPU temperature across the fleet in real time?
37An administrator is managing an NVIDIA DGX SuperPOD used for large-scale AI training. The cluster uses a Slurm workload manager. The administrator needs to ensure that jobs are scheduled only on nodes with healthy GPUs and that failed GPUs are automatically drained from the pool. Which integration should be configured to achieve this?
38An administrator is using NVIDIA Base Command Manager to provision a new GPU cluster. They need to ensure that the compute nodes are configured with the correct GPU driver and CUDA toolkit versions. Which Base Command Manager feature should they use to automate this?
39An administrator is responsible for an NVIDIA AI Enterprise deployment on Kubernetes. The security team requires that all GPU-accelerated pods run with the least privilege necessary and that GPU device nodes are not exposed to pods that do not request them. Which combination of configurations should the administrator implement to meet these requirements?
40An AI operations team needs to monitor GPU health and utilization across a fleet of DGX nodes from a single dashboard. They want per-GPU metrics such as power, temperature, utilization, and ECC errors, and they want to retain historical data for capacity planning. Which NVIDIA tool is purpose-built to collect and expose these GPU telemetry metrics for centralized monitoring?
41An administrator is configuring NVIDIA GPUDirect Storage (GDS) on a cluster to accelerate data loading for AI training jobs. The cluster uses Mellanox InfiniBand adapters and NVMe storage. Which two actions are required to enable GDS and ensure optimal performance? (Choose two.)
42An administrator is managing a Kubernetes cluster with NVIDIA GPU Operator. A new policy requires that all GPU workloads run with specific environment variables set, such as NVIDIA_VISIBLE_DEVICES and NVIDIA_DRIVER_CAPABILITIES. The administrator wants to enforce these variables automatically for any pod that requests a GPU, without modifying each pod specification manually. Which approach should the administrator use?
43An administrator is troubleshooting a performance degradation in a multi-node NVIDIA NCCL-based training job. The job spans four DGX nodes connected via InfiniBand. The administrator suspects that NCCL is not using the optimal network path. Which action should the administrator take to verify and enforce the use of GPUDirect RDMA for inter-node communication?
44An administrator is configuring an NVIDIA AI Enterprise cluster to run multi-tenant inference workloads on Kubernetes. The administrator must ensure that GPU resources are isolated and that tenants cannot access each other's GPU memory. Which two actions should the administrator take? (Choose two.)
Be able to deploy GPU-enabled containers on Kubernetes and configure MIG or vGPU for isolated sharing. The single most important thing: know which NVIDIA technology provides strict hardware isolation versus mere time-slicing, and verify the full container GPU stack is installed end to end.
The Courseiva NCP-AIO question bank contains 44 questions in the Administration domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Administration domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included