Courseiva

NCP-AIO · domain

Administration

The Administration domain of the NCP-AIO exam covers deploying, securing, and operating NVIDIA GPU infrastructure for AI workloads across containers, Kubernetes, and virtualized environments. It tests practical configuration of GPU sharing, container runtimes, cluster performance tuning, and security policy application, requiring administrators to reason about the immediate operational effects of their choices.

44 questions9 easy24 medium11 hard

Focused practice

Practice Administration questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Administration

Be able to deploy GPU-enabled containers on Kubernetes and configure MIG or vGPU for isolated sharing. The single most important thing: know which NVIDIA technology provides strict hardware isolation versus mere time-slicing, and verify the full container GPU stack is installed end to end.

Configuring NVIDIA Container Toolkit and device plugin for GPU access in Kubernetes pods

Applying Multi-Instance GPU (MIG) and vGPU for hardware-isolated GPU sharing across tenants

Using NVIDIA Base Command Manager to monitor and tune training cluster performance

Interpreting security policy effects on containerized AI workloads at runtime

Watch out for

Common Administration exam traps

  • ▸Assuming GPU sharing always provides isolation; only MIG or vGPU enforce strict hardware partitioning, while time-slicing does not.
  • ▸Forgetting that the NVIDIA device plugin and container runtime hooks must both be installed before pods can request GPUs.
  • ▸Treating Base Command as only a scheduler; it also handles node health, fabric monitoring, and performance consistency.

Question index

All Administration questions (44)

Click any question to see the full explanation, or start a practice session above.

1

An administrator is preparing a cluster for a new large language model training job that will use NVIDIA Magnum IO GPUDirect Storage to stream training data directly from a parallel file system to GPU memory. The administrator must verify that the environment supports GPUDirect Storage before the job starts. (Choose two.)

Medium
2

An administrator is setting up NVIDIA Base Command Manager to provision and manage a new AI cluster. They need to ensure that the cluster can automatically discover and configure new GPU nodes. Which component is responsible for node discovery and initial configuration?

Easy
3

An administrator is setting up an NVIDIA AI Enterprise cluster and wants to verify that the NVIDIA GPU Operator has successfully deployed all required components on a worker node. Which command should the administrator use to list the GPU Operator pods running on that node?

Easy
4

An administrator is using NVIDIA Base Command Manager to provision a new GPU cluster. They need to ensure that the compute nodes are configured with the correct GPU driver and CUDA toolkit versions. Which Base Command Manager feature should they use to automate this?

Medium
5

An administrator needs to collect GPU telemetry from an NVIDIA AI Enterprise cluster and store it in a time-series database for long-term analysis. Which component should be deployed to export GPU metrics in Prometheus format?

Medium
6

An administrator is configuring an NVIDIA AI Enterprise cluster to run multi-tenant inference workloads on Kubernetes. The administrator must ensure that GPU resources are isolated and that tenants cannot access each other's GPU memory. Which two actions should the administrator take? (Choose two.)

Hard
7

Refer to the exhibit. An administrator applies this security policy to a container runtime environment. What is the immediate effect on containerized AI applications within this scope?

Hard
8

An administrator is responsible for maintaining a fleet of NVIDIA-certified servers running AI workloads. They need to quickly identify which servers have GPUs that are overheating and may throttle performance. Which NVIDIA tool should the administrator use to monitor GPU temperature across the fleet in real time?

Easy
9

An administrator is optimizing a cluster for AI model training using NVIDIA Base Command. Which TWO tasks are critical for ensuring consistent performance across the training nodes?

Medium
10

An administrator manages an NVIDIA AI Enterprise cluster running multiple Kubernetes nodes, each with several A100 GPUs. After upgrading the NVIDIA GPU Operator to a newer version, the administrator notices that pods requesting GPUs remain in a Pending state, and the node's allocatable GPU count is reported as zero. Which command should the administrator run first to diagnose the issue?

Medium
11

A site reliability engineer is troubleshooting a DGX A100 node that intermittently drops out of the cluster during large NCCL all-reduce jobs. `nvidia-smi` shows all eight A100 GPUs healthy, but DCGM reports XID errors 74 and 79 on one GPU during the failures. The engineer needs to determine the most likely cause and the correct administrative action. Which combination best describes the cause and the appropriate first step?

Hard
12

An administrator is troubleshooting a performance degradation in a multi-node NVIDIA NCCL-based training job. The job spans four DGX nodes connected via InfiniBand. The administrator suspects that NCCL is not using the optimal network path. Which action should the administrator take to verify and enforce the use of GPUDirect RDMA for inter-node communication?

Hard
13

An administrator is configuring a new cluster and wants to ensure that telemetry data from GPUs is collected in a centralized manner. Which tool is best suited for this requirement?

Medium
14

An AI operations team is deploying NVIDIA Base Command Manager to manage a cluster of DGX nodes. They want to ensure that only authorized users can submit jobs and that all job submissions are audited. Which combination of Base Command Manager features should the administrator configure to meet these requirements?

Medium
15

An AI operations team needs to monitor GPU health and utilization across a fleet of DGX nodes from a single dashboard. They want per-GPU metrics such as power, temperature, utilization, and ECC errors, and they want to retain historical data for capacity planning. Which NVIDIA tool is purpose-built to collect and expose these GPU telemetry metrics for centralized monitoring?

Easy
16

An administrator is preparing a multi-node NVIDIA DGX H100 cluster for a distributed training job using NVIDIA Base Command. The cluster nodes have InfiniBand adapters, but the job's inter-node throughput is far below expectations. The administrator runs `ibstat` and sees that the ports are in the INIT state rather than ACTIVE. Which action should the administrator take first?

Medium
17

An administrator is configuring a multi-tenant NVIDIA AI Enterprise environment. Which mechanism is most effective for ensuring hardware-level isolation between concurrent training jobs on a single A100 GPU?

Medium
18

An administrator supports a shared Kubernetes cluster running NVIDIA GPU Operator. Data scientists report that their inference pods remain in Pending state, yet the GPU Operator pods and node feature discovery pods are healthy, and the GPU nodes show no hardware alarms. The administrator confirms that the cluster has a mixture of MIG-capable A100 nodes and non-MIG T4 nodes. Which immediate administrative action is most appropriate to diagnose the scheduling failure?

Medium
19

An administrator is responsible for a large NVIDIA DGX SuperPOD used for multi-node training. They need to ensure that GPU telemetry and health metrics are collected centrally and can trigger alerts when GPUs exceed temperature thresholds. Which component of NVIDIA Base Command Manager (BCM) should they configure to achieve this?

Hard
20

An administrator manages an NVIDIA AI Enterprise deployment on Kubernetes and needs to ensure that GPU telemetry is exported to an existing Prometheus instance. The administrator deploys the NVIDIA DCGM Exporter but sees no GPU metrics in Prometheus. Which configuration should the administrator verify first?

Hard
21

An administrator wants to prevent unauthorized users from accessing sensitive model weights stored in GPU memory. Which security feature should be implemented to ensure hardware-level isolation of the memory space?

Medium
22

An AI operations team is using NVIDIA DCGM (Data Center GPU Manager) to monitor a cluster of A100 GPUs. They want to set up proactive health checks to detect and mitigate GPU issues before they cause job failures. Which two DCGM features should they configure? (Choose two.)

Medium
23

An administrator is responsible for an NVIDIA AI Enterprise deployment on Kubernetes. The security team requires that all GPU-accelerated pods run with the least privilege necessary and that GPU device nodes are not exposed to pods that do not request them. Which combination of configurations should the administrator implement to meet these requirements?

Hard
24

An administrator is configuring NVIDIA GPUDirect Storage (GDS) on a cluster to accelerate data loading for AI training jobs. The cluster uses Mellanox InfiniBand adapters and NVMe storage. Which two actions are required to enable GDS and ensure optimal performance? (Choose two.)

Hard
25

A company runs multiple AI workloads on a shared Kubernetes cluster with NVIDIA GPUs. The administrator needs to enforce that only pods with a specific label can consume GPU resources, while other pods are denied. Which Kubernetes admission control mechanism should be used to implement this policy?

Medium
26

An administrator is managing an NVIDIA DGX SuperPOD used for large-scale AI training. The cluster uses a Slurm workload manager. The administrator needs to ensure that jobs are scheduled only on nodes with healthy GPUs and that failed GPUs are automatically drained from the pool. Which integration should be configured to achieve this?

Medium
27

When monitoring GPU health in an enterprise cluster, which command provides the most comprehensive snapshot of real-time power, temperature, and memory utilization?

Easy
28

An administrator is configuring NVIDIA Base Command Manager to manage a cluster of DGX nodes. They want to ensure that when a node's GPU temperature exceeds a defined threshold, the node is automatically drained and an alert is sent to the operations team. Which combination of Base Command Manager features should the administrator configure to achieve this?

Hard
29

An administrator is deploying NVIDIA AI Enterprise on a Kubernetes cluster and wants to verify that the GPU Operator has successfully installed all required components. Which command should the administrator use to check the status of the GPU Operator pods?

Easy
30

An AI operations engineer is preparing a DGX H100 system for a multi-node training workload. The engineer runs `nvidia-smi topo -m` and notices that GPU4 and GPU5 report a connection type of SYS, while all other GPU pairs show NV18. What is the most likely cause of this topology anomaly?

Medium
31

Which NVIDIA technology enables a GPU to be shared among multiple virtual machines or containers while maintaining strict hardware isolation?

Easy
32

An administrator is troubleshooting a multi-node NVIDIA GPU training job that intermittently hangs during the all-reduce phase. The job uses NCCL over InfiniBand. Logs show that some ranks time out while others complete. The administrator suspects a network fabric issue. Which action should the administrator take first to isolate whether the problem is in the InfiniBand fabric or in the NCCL configuration?

Medium
33

Which THREE components are required for a container to successfully leverage NVIDIA GPUs on a Kubernetes cluster?

Hard
34

An administrator is managing a Kubernetes cluster with NVIDIA GPU Operator. A new policy requires that all GPU workloads run with specific environment variables set, such as NVIDIA_VISIBLE_DEVICES and NVIDIA_DRIVER_CAPABILITIES. The administrator wants to enforce these variables automatically for any pod that requests a GPU, without modifying each pod specification manually. Which approach should the administrator use?

Medium
35

Which administrative practice ensures that a cluster is prepared for the arrival of new NVIDIA GPU hardware with minimal downtime?

Medium
36

An administrator is responsible for maintaining an NVIDIA AI Enterprise cluster and needs to ensure high availability of GPU resources for critical inference workloads. Which two practices should the administrator implement? (Choose two.)

Medium
37

An administrator manages an NVIDIA AI Enterprise cluster using NVIDIA Run:ai. A data science team complains that their submitted training job has been stuck in a Pending state for over an hour, even though the Run:ai scheduler shows free GPUs in the cluster. The administrator verifies that the job requests 2 GPUs and the node pool has 4 idle GPUs. Which Run:ai administrative configuration is the most likely cause of the job remaining Pending?

Medium
38

An administrator is responsible for monitoring a large-scale AI cluster with hundreds of NVIDIA GPUs. They need to collect telemetry data such as GPU utilization, temperature, and power consumption from all nodes and store it centrally for analysis and alerting. Which NVIDIA tool should they use to collect and export GPU metrics to a monitoring system like Prometheus?

Easy
39

An administrator manages an NVIDIA AI Enterprise cluster and needs to enforce GPU resource quotas across multiple Kubernetes namespaces. Which NVIDIA component should be configured to enforce these quotas?

Medium
40

An administrator is preparing a bare-metal GPU server for AI workloads and needs to verify that the NVIDIA driver and CUDA toolkit are properly installed. The server has an NVIDIA A100 GPU. Which command should the administrator run to display the GPU model, driver version, and CUDA version?

Medium
41

An administrator needs to ensure that all GPU drivers are updated across a heterogeneous cluster without causing downtime. What is the best strategy?

Medium
42

Refer to the exhibit. An administrator observes this in the logs during a multi-GPU training job. What is the performance implication of this setting?

Hard
43

An administrator notices that GPU utilization is high, but throughput in an AI training job remains low. What is the most likely bottleneck?

Medium
44

An administrator is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing several NVIDIA A100 GPUs. They need to enable GPU sharing across multiple virtual machines to maximize utilization. Which vSphere feature should they configure?

Easy

Frequently asked questions

What does the Administration domain cover on the NCP-AIO exam?
Be able to deploy GPU-enabled containers on Kubernetes and configure MIG or vGPU for isolated sharing. The single most important thing: know which NVIDIA technology provides strict hardware isolation versus mere time-slicing, and verify the full container GPU stack is installed end to end.
How many questions are in this domain?
This page lists all 44 Administration questions in the NCP-AIO question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Administration questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
nvidia-ncp-aio NVIDIA-NCP-AIO ncp aio administration Practice Questions