NCP-AIO · domain
Administration
The Administration domain of the NCP-AIO exam covers deploying, securing, and operating NVIDIA GPU infrastructure for AI workloads across containers, Kubernetes, and virtualized environments. It tests practical configuration of GPU sharing, container runtimes, cluster performance tuning, and security policy application, requiring administrators to reason about the immediate operational effects of their choices.
Focused practice
Practice Administration questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Administration
Be able to deploy GPU-enabled containers on Kubernetes and configure MIG or vGPU for isolated sharing. The single most important thing: know which NVIDIA technology provides strict hardware isolation versus mere time-slicing, and verify the full container GPU stack is installed end to end.
Configuring NVIDIA Container Toolkit and device plugin for GPU access in Kubernetes pods
Applying Multi-Instance GPU (MIG) and vGPU for hardware-isolated GPU sharing across tenants
Using NVIDIA Base Command Manager to monitor and tune training cluster performance
Interpreting security policy effects on containerized AI workloads at runtime
Watch out for
Common Administration exam traps
- ▸Assuming GPU sharing always provides isolation; only MIG or vGPU enforce strict hardware partitioning, while time-slicing does not.
- ▸Forgetting that the NVIDIA device plugin and container runtime hooks must both be installed before pods can request GPUs.
- ▸Treating Base Command as only a scheduler; it also handles node health, fabric monitoring, and performance consistency.
Question index
All Administration questions (44)
Click any question to see the full explanation, or start a practice session above.
An administrator is preparing a cluster for a new large language model training job that will use NVIDIA Magnum IO GPUDirect Storage to stream training data directly from a parallel file system to GPU memory. The administrator must verify that the environment supports GPUDirect Storage before the job starts. (Choose two.)
Medium2An administrator is setting up NVIDIA Base Command Manager to provision and manage a new AI cluster. They need to ensure that the cluster can automatically discover and configure new GPU nodes. Which component is responsible for node discovery and initial configuration?
Easy3An administrator is setting up an NVIDIA AI Enterprise cluster and wants to verify that the NVIDIA GPU Operator has successfully deployed all required components on a worker node. Which command should the administrator use to list the GPU Operator pods running on that node?
Easy4An administrator is using NVIDIA Base Command Manager to provision a new GPU cluster. They need to ensure that the compute nodes are configured with the correct GPU driver and CUDA toolkit versions. Which Base Command Manager feature should they use to automate this?
Medium5An administrator needs to collect GPU telemetry from an NVIDIA AI Enterprise cluster and store it in a time-series database for long-term analysis. Which component should be deployed to export GPU metrics in Prometheus format?
Medium6An administrator is configuring an NVIDIA AI Enterprise cluster to run multi-tenant inference workloads on Kubernetes. The administrator must ensure that GPU resources are isolated and that tenants cannot access each other's GPU memory. Which two actions should the administrator take? (Choose two.)
Hard7Refer to the exhibit. An administrator applies this security policy to a container runtime environment. What is the immediate effect on containerized AI applications within this scope?
Hard8An administrator is responsible for maintaining a fleet of NVIDIA-certified servers running AI workloads. They need to quickly identify which servers have GPUs that are overheating and may throttle performance. Which NVIDIA tool should the administrator use to monitor GPU temperature across the fleet in real time?
Easy9An administrator is optimizing a cluster for AI model training using NVIDIA Base Command. Which TWO tasks are critical for ensuring consistent performance across the training nodes?
Medium10An administrator manages an NVIDIA AI Enterprise cluster running multiple Kubernetes nodes, each with several A100 GPUs. After upgrading the NVIDIA GPU Operator to a newer version, the administrator notices that pods requesting GPUs remain in a Pending state, and the node's allocatable GPU count is reported as zero. Which command should the administrator run first to diagnose the issue?
Medium11A site reliability engineer is troubleshooting a DGX A100 node that intermittently drops out of the cluster during large NCCL all-reduce jobs. `nvidia-smi` shows all eight A100 GPUs healthy, but DCGM reports XID errors 74 and 79 on one GPU during the failures. The engineer needs to determine the most likely cause and the correct administrative action. Which combination best describes the cause and the appropriate first step?
Hard12An administrator is troubleshooting a performance degradation in a multi-node NVIDIA NCCL-based training job. The job spans four DGX nodes connected via InfiniBand. The administrator suspects that NCCL is not using the optimal network path. Which action should the administrator take to verify and enforce the use of GPUDirect RDMA for inter-node communication?
Hard13An administrator is configuring a new cluster and wants to ensure that telemetry data from GPUs is collected in a centralized manner. Which tool is best suited for this requirement?
Medium14An AI operations team is deploying NVIDIA Base Command Manager to manage a cluster of DGX nodes. They want to ensure that only authorized users can submit jobs and that all job submissions are audited. Which combination of Base Command Manager features should the administrator configure to meet these requirements?
Medium15An AI operations team needs to monitor GPU health and utilization across a fleet of DGX nodes from a single dashboard. They want per-GPU metrics such as power, temperature, utilization, and ECC errors, and they want to retain historical data for capacity planning. Which NVIDIA tool is purpose-built to collect and expose these GPU telemetry metrics for centralized monitoring?
Easy16An administrator is preparing a multi-node NVIDIA DGX H100 cluster for a distributed training job using NVIDIA Base Command. The cluster nodes have InfiniBand adapters, but the job's inter-node throughput is far below expectations. The administrator runs `ibstat` and sees that the ports are in the INIT state rather than ACTIVE. Which action should the administrator take first?
Medium17An administrator is configuring a multi-tenant NVIDIA AI Enterprise environment. Which mechanism is most effective for ensuring hardware-level isolation between concurrent training jobs on a single A100 GPU?
Medium18An administrator supports a shared Kubernetes cluster running NVIDIA GPU Operator. Data scientists report that their inference pods remain in Pending state, yet the GPU Operator pods and node feature discovery pods are healthy, and the GPU nodes show no hardware alarms. The administrator confirms that the cluster has a mixture of MIG-capable A100 nodes and non-MIG T4 nodes. Which immediate administrative action is most appropriate to diagnose the scheduling failure?
Medium19An administrator is responsible for a large NVIDIA DGX SuperPOD used for multi-node training. They need to ensure that GPU telemetry and health metrics are collected centrally and can trigger alerts when GPUs exceed temperature thresholds. Which component of NVIDIA Base Command Manager (BCM) should they configure to achieve this?
Hard20An administrator manages an NVIDIA AI Enterprise deployment on Kubernetes and needs to ensure that GPU telemetry is exported to an existing Prometheus instance. The administrator deploys the NVIDIA DCGM Exporter but sees no GPU metrics in Prometheus. Which configuration should the administrator verify first?
Hard21An administrator wants to prevent unauthorized users from accessing sensitive model weights stored in GPU memory. Which security feature should be implemented to ensure hardware-level isolation of the memory space?
Medium22An AI operations team is using NVIDIA DCGM (Data Center GPU Manager) to monitor a cluster of A100 GPUs. They want to set up proactive health checks to detect and mitigate GPU issues before they cause job failures. Which two DCGM features should they configure? (Choose two.)
Medium23An administrator is responsible for an NVIDIA AI Enterprise deployment on Kubernetes. The security team requires that all GPU-accelerated pods run with the least privilege necessary and that GPU device nodes are not exposed to pods that do not request them. Which combination of configurations should the administrator implement to meet these requirements?
Hard24An administrator is configuring NVIDIA GPUDirect Storage (GDS) on a cluster to accelerate data loading for AI training jobs. The cluster uses Mellanox InfiniBand adapters and NVMe storage. Which two actions are required to enable GDS and ensure optimal performance? (Choose two.)
Hard25A company runs multiple AI workloads on a shared Kubernetes cluster with NVIDIA GPUs. The administrator needs to enforce that only pods with a specific label can consume GPU resources, while other pods are denied. Which Kubernetes admission control mechanism should be used to implement this policy?
Medium26An administrator is managing an NVIDIA DGX SuperPOD used for large-scale AI training. The cluster uses a Slurm workload manager. The administrator needs to ensure that jobs are scheduled only on nodes with healthy GPUs and that failed GPUs are automatically drained from the pool. Which integration should be configured to achieve this?
Medium27When monitoring GPU health in an enterprise cluster, which command provides the most comprehensive snapshot of real-time power, temperature, and memory utilization?
Easy28An administrator is configuring NVIDIA Base Command Manager to manage a cluster of DGX nodes. They want to ensure that when a node's GPU temperature exceeds a defined threshold, the node is automatically drained and an alert is sent to the operations team. Which combination of Base Command Manager features should the administrator configure to achieve this?
Hard29An administrator is deploying NVIDIA AI Enterprise on a Kubernetes cluster and wants to verify that the GPU Operator has successfully installed all required components. Which command should the administrator use to check the status of the GPU Operator pods?
Easy30An AI operations engineer is preparing a DGX H100 system for a multi-node training workload. The engineer runs `nvidia-smi topo -m` and notices that GPU4 and GPU5 report a connection type of SYS, while all other GPU pairs show NV18. What is the most likely cause of this topology anomaly?
Medium31Which NVIDIA technology enables a GPU to be shared among multiple virtual machines or containers while maintaining strict hardware isolation?
Easy32An administrator is troubleshooting a multi-node NVIDIA GPU training job that intermittently hangs during the all-reduce phase. The job uses NCCL over InfiniBand. Logs show that some ranks time out while others complete. The administrator suspects a network fabric issue. Which action should the administrator take first to isolate whether the problem is in the InfiniBand fabric or in the NCCL configuration?
Medium33Which THREE components are required for a container to successfully leverage NVIDIA GPUs on a Kubernetes cluster?
Hard34An administrator is managing a Kubernetes cluster with NVIDIA GPU Operator. A new policy requires that all GPU workloads run with specific environment variables set, such as NVIDIA_VISIBLE_DEVICES and NVIDIA_DRIVER_CAPABILITIES. The administrator wants to enforce these variables automatically for any pod that requests a GPU, without modifying each pod specification manually. Which approach should the administrator use?
Medium35Which administrative practice ensures that a cluster is prepared for the arrival of new NVIDIA GPU hardware with minimal downtime?
Medium36An administrator is responsible for maintaining an NVIDIA AI Enterprise cluster and needs to ensure high availability of GPU resources for critical inference workloads. Which two practices should the administrator implement? (Choose two.)
Medium37An administrator manages an NVIDIA AI Enterprise cluster using NVIDIA Run:ai. A data science team complains that their submitted training job has been stuck in a Pending state for over an hour, even though the Run:ai scheduler shows free GPUs in the cluster. The administrator verifies that the job requests 2 GPUs and the node pool has 4 idle GPUs. Which Run:ai administrative configuration is the most likely cause of the job remaining Pending?
Medium38An administrator is responsible for monitoring a large-scale AI cluster with hundreds of NVIDIA GPUs. They need to collect telemetry data such as GPU utilization, temperature, and power consumption from all nodes and store it centrally for analysis and alerting. Which NVIDIA tool should they use to collect and export GPU metrics to a monitoring system like Prometheus?
Easy39An administrator manages an NVIDIA AI Enterprise cluster and needs to enforce GPU resource quotas across multiple Kubernetes namespaces. Which NVIDIA component should be configured to enforce these quotas?
Medium40An administrator is preparing a bare-metal GPU server for AI workloads and needs to verify that the NVIDIA driver and CUDA toolkit are properly installed. The server has an NVIDIA A100 GPU. Which command should the administrator run to display the GPU model, driver version, and CUDA version?
Medium41An administrator needs to ensure that all GPU drivers are updated across a heterogeneous cluster without causing downtime. What is the best strategy?
Medium42Refer to the exhibit. An administrator observes this in the logs during a multi-GPU training job. What is the performance implication of this setting?
Hard43An administrator notices that GPU utilization is high, but throughput in an AI training job remains low. What is the most likely bottleneck?
Medium44An administrator is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing several NVIDIA A100 GPUs. They need to enable GPU sharing across multiple virtual machines to maximize utilization. Which vSphere feature should they configure?
EasyOther domains
All NCP-AIO exam domains
Frequently asked questions
- What does the Administration domain cover on the NCP-AIO exam?
- Be able to deploy GPU-enabled containers on Kubernetes and configure MIG or vGPU for isolated sharing. The single most important thing: know which NVIDIA technology provides strict hardware isolation versus mere time-slicing, and verify the full container GPU stack is installed end to end.
- How many questions are in this domain?
- This page lists all 44 Administration questions in the NCP-AIO question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Administration questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.