Courseiva

CCNA Administration Questions

44 questions · Administration · All types, answers revealed

1
Multi-Selectmedium

An administrator is preparing a cluster for a new large language model training job that will use NVIDIA Magnum IO GPUDirect Storage to stream training data directly from a parallel file system to GPU memory. The administrator must verify that the environment supports GPUDirect Storage before the job starts. (Choose two.)

Select 2 answers
A.Disable IOMMU in the BIOS on all GPU nodes to allow direct memory access between storage devices and GPUs.
B.Ensure that the parallel file system client supports the cuFile API and that the file system is mounted with options that allow direct I/O and peer-to-peer memory access.
C.Configure the cluster to use RDMA over Converged Ethernet (RoCE) exclusively for all storage traffic, because GPUDirect Storage requires RDMA.
D.Confirm that the NVIDIA driver and CUDA toolkit versions installed on the nodes meet the minimum requirements for the GPUDirect Storage release being used, and that the nvidia-fs kernel module is loaded.
E.Enable NVIDIA Multi-Instance GPU (MIG) on every GPU to provide isolated memory partitions for storage transfers.
AnswersB, D

The file system client must support cuFile and be mounted appropriately for direct memory access. Without cuFile support, applications cannot use the GPUDirect Storage path. Mount options and client capabilities determine whether the storage stack can perform peer-to-peer transfers to GPU memory.

Why this answer

GPUDirect Storage requires compatible drivers, the nvidia-fs kernel module, and a file system client that supports the cuFile API with suitable mount options. These two checks confirm that the software stack and storage client can perform direct memory transfers. MIG, exclusive RoCE, and disabling IOMMU are not prerequisites and may be counterproductive.

Exam trap

The trap here is assuming that GPUDirect Storage depends on MIG or a specific network transport, when it actually relies on driver, kernel module, and file system client support.

2
MCQeasy

An administrator is setting up NVIDIA Base Command Manager to provision and manage a new AI cluster. They need to ensure that the cluster can automatically discover and configure new GPU nodes. Which component is responsible for node discovery and initial configuration?

A.NVIDIA Fleet Command
B.NVIDIA GPU Operator
C.NVIDIA Morpheus
D.NVIDIA Base Command Manager head node
AnswerD

The head node in Base Command Manager is the central management server that handles node discovery, provisioning, and configuration. It runs services like DHCP, TFTP, and the provisioning engine that automatically detect new nodes when they PXE boot. It then applies the appropriate image and configuration. This is the core component for automating cluster setup, making it the correct answer for node discovery and initial configuration.

Why this answer

In NVIDIA Base Command Manager, the head node is the central management server that performs node discovery and initial configuration. It provides the provisioning services that allow new GPU nodes to be automatically detected and configured. The GPU Operator, Fleet Command, and Morpheus serve different purposes and are not involved in the initial provisioning of a Base Command Manager cluster.

Exam trap

The trap here is confusing Base Command Manager with Kubernetes-focused tools like the GPU Operator, which handle different layers of the stack.

3
MCQeasy

An administrator is setting up an NVIDIA AI Enterprise cluster and wants to verify that the NVIDIA GPU Operator has successfully deployed all required components on a worker node. Which command should the administrator use to list the GPU Operator pods running on that node?

A.nvidia-smi -q -d COMPUTE
B.kubectl describe node <node-name> | grep nvidia
C.kubectl get pods -n gpu-operator --field-selector spec.nodeName=<node-name>
D.helm list -n gpu-operator
AnswerC

The GPU Operator deploys its components into the gpu-operator namespace by default, and the field selector spec.nodeName filters pods to a specific node. This command directly lists the operator pods on the target node, allowing the administrator to confirm that components such as the driver daemonset, container toolkit, and device plugin are running. It is the precise and efficient way to verify operator deployment on a worker node.

Why this answer

The NVIDIA GPU Operator runs its components as pods in the gpu-operator namespace. To confirm deployment on a specific worker node, filtering pods by spec.nodeName in that namespace lists exactly the operator pods scheduled there. This provides direct evidence that the driver, container toolkit, device plugin, and related daemonsets are running on the node.

Exam trap

The trap here is confusing Helm release status with actual pod health, when verifying operator deployment on a node requires inspecting pods filtered by node name.

4
MCQmedium

An administrator is using NVIDIA Base Command Manager to provision a new GPU cluster. They need to ensure that the compute nodes are configured with the correct GPU driver and CUDA toolkit versions. Which Base Command Manager feature should they use to automate this?

A.Ansible playbooks executed manually
B.Node groups with software images
C.PXE boot with custom kickstart scripts
D.Cron-based package updates
AnswerB

Base Command Manager allows administrators to define node groups and assign software images that include specific GPU driver and CUDA toolkit versions. During provisioning, nodes in the group automatically receive the defined software stack, ensuring consistency. This automation reduces manual configuration errors and ensures all compute nodes have the correct versions for AI workloads.

Why this answer

Base Command Manager uses node groups to categorize nodes with similar roles and software requirements. By assigning a software image to a node group, administrators define the exact GPU driver and CUDA toolkit versions. During provisioning, Base Command Manager applies the image, ensuring consistency across compute nodes.

This is the intended feature for automating software configuration in a Base Command Manager cluster.

Exam trap

The trap here is assuming that generic automation tools like Ansible or PXE are sufficient, but Base Command Manager's integrated software image and node group features are specifically designed for this scenario.

5
MCQmedium

An administrator needs to collect GPU telemetry from an NVIDIA AI Enterprise cluster and store it in a time-series database for long-term analysis. Which component should be deployed to export GPU metrics in Prometheus format?

A.NVIDIA Container Toolkit
B.NVIDIA Base Command Manager
C.NVIDIA DCGM Exporter
D.NVIDIA GPU Operator
AnswerC

DCGM Exporter is designed to collect GPU metrics using DCGM and expose them in Prometheus format. It runs as a container and provides metrics such as GPU utilization, memory usage, temperature, and power. These metrics can be scraped by Prometheus and stored in a time-series database for analysis. This is the standard solution for exporting GPU telemetry in Kubernetes environments.

Why this answer

DCGM Exporter is the component that collects GPU metrics via DCGM and exposes them as Prometheus metrics. It is typically deployed as a daemonset or pod and can be scraped by Prometheus for long-term storage and analysis. Other NVIDIA components manage deployment or container runtime integration but do not provide metric export functionality.

Exam trap

The trap here is assuming that the GPU Operator or Base Command Manager exports metrics, when metric export is specifically the role of DCGM Exporter.

6
Multi-Selecthard

An administrator is configuring an NVIDIA AI Enterprise cluster to run multi-tenant inference workloads on Kubernetes. The administrator must ensure that GPU resources are isolated and that tenants cannot access each other's GPU memory. Which two actions should the administrator take? (Choose two.)

Select 2 answers
A.Use Kubernetes ResourceQuota and LimitRange objects to restrict GPU memory per namespace.
B.Configure NVIDIA vGPU with different vGPU profiles for each tenant.
C.Deploy the NVIDIA GPU Operator with the device plugin configured to advertise MIG resources.
D.Enable Multi-Instance GPU (MIG) mode on supported GPUs and assign separate MIG instances to each tenant.
E.Enable time-slicing of GPUs so that multiple tenants share the same GPU context.
AnswersC, D

For Kubernetes to schedule workloads onto MIG instances, the NVIDIA device plugin must advertise each MIG slice as a schedulable resource, which the GPU Operator configures via its MIG strategy setting. Without this, MIG instances exist on the GPU but are invisible to the scheduler. Enabling the device plugin to expose MIG resources is therefore a necessary action to make tenant isolation effective.

Why this answer

Hardware-level isolation for multi-tenant inference on Kubernetes is achieved by enabling MIG on supported GPUs and having the GPU Operator's device plugin advertise MIG instances as schedulable resources. Together these actions partition the GPU into isolated slices and make those slices available to tenants without memory sharing, satisfying the isolation requirement.

Exam trap

The trap here is assuming that Kubernetes quotas or time-slicing provide GPU memory isolation, when only MIG creates hardware-partitioned instances with dedicated memory.

7
MCQhard

Refer to the exhibit. An administrator applies this security policy to a container runtime environment. What is the immediate effect on containerized AI applications within this scope?

A.The container gains elevated privileges to access the host GPU
B.The container is unable to detect or communicate with the NVIDIA GPU
C.The container can use the GPU but cannot perform memory mapping
D.The container experiences increased latency for GPU operations
AnswerB

The NVIDIA driver requires access to specific device nodes under /dev/nvidia* to function. By denying access to these files, the runtime prevents the container's CUDA libraries from establishing a connection to the GPU driver, rendering the GPU invisible to the application code executing inside the container environment.

Why this answer

The policy explicitly denies access to the character device files associated with the NVIDIA GPU. Without access to /dev/nvidia0, /dev/nvidiactl, and /dev/nvidia-uvm, the CUDA runtime cannot interact with the GPU hardware. Consequently, any attempt to initialize a CUDA device will fail, causing the application to crash.

This policy is a common restrictive measure in high-security environments where GPU access must be strictly managed or audited.

Exam trap

Candidates often assume that security policies only affect network traffic or file system access, overlooking that blocking character device nodes directly breaks the CUDA runtime's ability to initialize hardware.

8
MCQeasy

An administrator is responsible for maintaining a fleet of NVIDIA-certified servers running AI workloads. They need to quickly identify which servers have GPUs that are overheating and may throttle performance. Which NVIDIA tool should the administrator use to monitor GPU temperature across the fleet in real time?

A.NVIDIA Nsight Systems
B.NVIDIA Data Center GPU Manager (DCGM)
C.NVIDIA TensorRT
D.NVIDIA CUDA Toolkit
AnswerB

DCGM is designed for data center GPU monitoring and management at scale. It provides real-time temperature, power, and health metrics for each GPU and can be integrated with monitoring systems. For fleet-wide temperature monitoring and throttle detection, DCGM is the appropriate tool in this scenario.

Why this answer

DCGM is the NVIDIA tool purpose-built for data center GPU monitoring. It exposes temperature, power, utilization, and health metrics through APIs and integrations with popular monitoring platforms. For an administrator needing real-time temperature visibility across many servers, DCGM is the correct choice.

Exam trap

The trap here is confusing performance profiling or development tools like Nsight Systems or CUDA Toolkit with fleet monitoring tools, which have different purposes and do not provide centralized temperature telemetry.

9
Multi-Selectmedium

An administrator is optimizing a cluster for AI model training using NVIDIA Base Command. Which TWO tasks are critical for ensuring consistent performance across the training nodes?

Select 2 answers
A.Standardizing NVIDIA driver and NCCL library versions across all nodes
B.Increasing the size of the swap partition on each node
C.Enabling GPUDirect RDMA to bypass system memory for GPU-to-GPU data transfer
D.Disabling the NVIDIA persistence mode to save power
E.Reducing the number of CUDA streams per training job
AnswersA, C

Distributed training performance is highly sensitive to the communication stack. Inconsistent NCCL versions can lead to suboptimal collective operation performance, while mismatched drivers can cause compatibility issues with the underlying hardware interconnects, ultimately leading to performance degradation and difficult-to-debug failures during large-scale model training cluster operations.

Why this answer

Consistency in high-performance AI clusters depends on hardware synchronization and resource availability. Ensuring that all nodes run identical driver and firmware versions prevents subtle performance regressions during multi-node training. Furthermore, configuring GPUDirect RDMA is essential for reducing latency in inter-GPU communication, which is a major bottleneck in distributed training jobs.

These steps are fundamental to maintaining high utilization and predictable training times across large-scale NVIDIA accelerated infrastructure.

Exam trap

Candidates often focus only on software frameworks while ignoring crucial low-level networking optimizations like GPUDirect RDMA and driver version synchronization across multi-node setups.

10
MCQmedium

An administrator manages an NVIDIA AI Enterprise cluster running multiple Kubernetes nodes, each with several A100 GPUs. After upgrading the NVIDIA GPU Operator to a newer version, the administrator notices that pods requesting GPUs remain in a Pending state, and the node's allocatable GPU count is reported as zero. Which command should the administrator run first to diagnose the issue?

A.kubectl describe node <node-name>
B.kubectl get pods --all-namespaces -o wide
C.nvidia-smi -q
D.kubectl logs -n gpu-operator <gpu-operator-pod>
AnswerA

This command shows detailed node status, including conditions, capacity, and allocatable resources. If the GPU Operator's device plugin is not functioning, the node will report zero allocatable nvidia.com/gpu resources, and events may indicate plugin registration failures. It directly reveals whether the node recognizes the GPUs as schedulable resources, making it the essential first diagnostic step.

Why this answer

The node's allocatable GPU count is reported as zero, indicating that the kubelet is not receiving GPU resource advertisements from the NVIDIA device plugin. Describing the node reveals capacity, allocatable resources, and relevant events, such as device plugin registration failures. This directly identifies whether the GPU Operator's device plugin is functioning, making it the correct first step.

Exam trap

The trap here is assuming that checking the GPU Operator logs or running nvidia-smi will directly explain why Kubernetes reports zero allocatable GPUs, when the node's resource status is the authoritative source.

11
MCQhard

A site reliability engineer is troubleshooting a DGX A100 node that intermittently drops out of the cluster during large NCCL all-reduce jobs. `nvidia-smi` shows all eight A100 GPUs healthy, but DCGM reports XID errors 74 and 79 on one GPU during the failures. The engineer needs to determine the most likely cause and the correct administrative action. Which combination best describes the cause and the appropriate first step?

A.The errors are benign and expected during large all-reduce jobs; suppress them by setting the DCGM health check to ignore XID 74 and 79.
B.The errors indicate a thermal shutdown; immediately lower the GPU clock with `nvidia-smi -lgc` and rerun the job.
C.The errors are caused by an outdated NCCL version; upgrade NCCL to the latest release and rerun the job without further hardware checks.
D.The errors point to a GPU falling off the bus or an NVLink error; run `nvidia-smi -q` and DCGM diagnostics on the suspect GPU, then consider reseating or replacing the GPU board.
AnswerD

XID 74 and 79 are associated with GPU falling off the bus and NVLink errors respectively. The correct administrative response is to collect detailed GPU and NVLink state with `nvidia-smi -q`, run DCGM diagnostics to isolate the failing GPU or link, and then perform hardware remediation such as reseating the GPU board or replacing it if diagnostics confirm a persistent fault.

Why this answer

XID 74 and 79 are driver-reported errors indicating a GPU has fallen off the bus or an NVLink error has occurred. These are hardware and link-level faults, so the right first step is to gather detailed GPU and NVLink diagnostics with `nvidia-smi -q` and DCGM, then remediate the hardware. Ignoring or masking these errors risks job failures and data corruption, while treating them as thermal or software issues misses the actual fault domain.

Exam trap

The trap here is treating XID 74 and 79 as generic performance or thermal warnings, when they specifically signal GPU bus loss and NVLink faults that require hardware-level diagnosis.

12
MCQhard

An administrator is troubleshooting a performance degradation in a multi-node NVIDIA NCCL-based training job. The job spans four DGX nodes connected via InfiniBand. The administrator suspects that NCCL is not using the optimal network path. Which action should the administrator take to verify and enforce the use of GPUDirect RDMA for inter-node communication?

A.Configure the InfiniBand fabric to use RoCEv2 instead of InfiniBand native protocol to improve compatibility with GPUDirect RDMA.
B.Set NCCL_NET_GDR_LEVEL=0 to force NCCL to use GPUDirect RDMA for all inter-node communication.
C.Set the environment variable NCCL_DEBUG=INFO and inspect the logs for messages indicating that GPUDirect RDMA is enabled.
D.Run nvidia-smi topo -m on each node to check the GPU-to-NIC affinity and ensure that GPUs and InfiniBand adapters are connected via PCIe switches.
AnswerC

Setting NCCL_DEBUG=INFO generates detailed logs that show which transports NCCL uses, including whether GPUDirect RDMA is active. The logs will indicate if the network plugin is using RDMA and if there are any fallbacks to slower paths. This is the standard method to verify NCCL's behavior and diagnose performance issues related to network communication.

Why this answer

To verify and enforce GPUDirect RDMA usage, the administrator should enable NCCL debug logging with NCCL_DEBUG=INFO and examine the logs for transport information. This will show if GDR is active and if any fallbacks occur. Checking topology with nvidia-smi topo -m is useful but does not confirm active GDR.

Setting NCCL_NET_GDR_LEVEL=0 would disable GDR, and switching to RoCEv2 is irrelevant.

Exam trap

The trap here is misunderstanding NCCL_NET_GDR_LEVEL: a value of 0 disables GPUDirect RDMA, while higher values enable it based on topology distance.

13
MCQmedium

An administrator is configuring a new cluster and wants to ensure that telemetry data from GPUs is collected in a centralized manner. Which tool is best suited for this requirement?

A.nvidia-smi log file rotation
B.NVIDIA DCGM Exporter
C.Manual polling of /proc/driver/nvidia
D.NVIDIA Nsight Systems
AnswerB

The DCGM Exporter is purpose-built for this task. It collects detailed telemetry, such as power, temperature, and usage, directly from the GPUs via DCGM and makes it available to monitoring platforms. This provides a clean, automated, and scalable architecture for centralized tracking of GPU performance metrics across the cluster.

Why this answer

NVIDIA DCGM Exporter is the industry-standard tool for collecting GPU telemetry data in Kubernetes environments. It aggregates metrics from the Data Center GPU Manager and exposes them in a format compatible with Prometheus, allowing for centralized monitoring, visualization, and alerting. This visibility is essential for understanding cluster performance, identifying anomalies, and optimizing resource utilization in large-scale AI infrastructure deployments, making it the preferred choice for enterprise monitoring solutions.

Exam trap

Test-takers sometimes confuse standard Kubernetes metrics servers with specialized GPU telemetry collectors like the NVIDIA DCGM Exporter.

14
MCQmedium

An AI operations team is deploying NVIDIA Base Command Manager to manage a cluster of DGX nodes. They want to ensure that only authorized users can submit jobs and that all job submissions are audited. Which combination of Base Command Manager features should the administrator configure to meet these requirements?

A.Deploy a separate Kubernetes cluster with RBAC policies and use Base Command Manager only for node provisioning.
B.Configure IPMI access controls on each DGX node and enable the Base Command Manager telemetry collector.
C.Enable SELinux enforcing mode on all nodes and configure Base Command Manager to use SSH key-based authentication only.
D.Integrate Base Command Manager with an LDAP or Active Directory identity provider and enable the audit logging feature.
AnswerD

Base Command Manager supports integration with enterprise identity providers such as LDAP or Active Directory for authentication and authorization. Enabling audit logging captures job submission events and user actions. Together, these features enforce authorized access and provide the required audit trail for job submissions in this scenario.

Why this answer

Base Command Manager centralizes cluster administration, including user authentication and job scheduling. Integrating with LDAP or Active Directory ensures only authorized users can authenticate and submit jobs, while audit logging records those submissions for compliance. This combination directly addresses both the access control and auditing requirements in the scenario.

Exam trap

The trap here is assuming that infrastructure hardening measures like SELinux or IPMI controls provide user-level job authorization and audit trails, when those functions require identity integration and audit logging in Base Command Manager.

15
MCQeasy

An AI operations team needs to monitor GPU health and utilization across a fleet of DGX nodes from a single dashboard. They want per-GPU metrics such as power, temperature, utilization, and ECC errors, and they want to retain historical data for capacity planning. Which NVIDIA tool is purpose-built to collect and expose these GPU telemetry metrics for centralized monitoring?

A.NVIDIA DCGM (Data Center GPU Manager) with the DCGM exporter for Prometheus.
B.NVIDIA Base Command Manager's job scheduler logs.
C.NVIDIA Nsight Systems for profiling GPU kernels and collecting timeline traces.
D.NVIDIA CUDA Toolkit's `nvidia-smi` command run manually on each node.
AnswerA

DCGM is the NVIDIA tool designed for data center GPU monitoring and management. It collects health, utilization, power, temperature, and ECC metrics, and the DCGM exporter exposes them to Prometheus for centralized dashboards and long-term retention. This directly matches the requirement for fleet-wide GPU telemetry and historical capacity planning data.

Why this answer

DCGM is NVIDIA's purpose-built data center GPU monitoring and management tool. It gathers health and performance metrics including power, temperature, utilization, and ECC errors, and the DCGM exporter makes them available to Prometheus for centralized dashboards and historical retention. This aligns exactly with the team's need for fleet-wide GPU telemetry and capacity planning data.

Exam trap

The trap here is confusing a profiler or a manual command-line utility with a continuous telemetry pipeline, when only DCGM with its exporter provides centralized, historical GPU health metrics.

16
MCQmedium

An administrator is preparing a multi-node NVIDIA DGX H100 cluster for a distributed training job using NVIDIA Base Command. The cluster nodes have InfiniBand adapters, but the job's inter-node throughput is far below expectations. The administrator runs `ibstat` and sees that the ports are in the INIT state rather than ACTIVE. Which action should the administrator take first?

A.Enable GPUDirect Storage on each node to bypass the CPU for inter-node communication.
B.Reinstall the NVIDIA GPU driver on all nodes to reset the InfiniBand firmware.
C.Disable ECC memory on the GPUs to increase available bandwidth for inter-node traffic.
D.Verify that the InfiniBand subnet manager is running and that the IPoIB interfaces are configured correctly.
AnswerD

An InfiniBand port remains in INIT when it has not been initialized by a subnet manager, so checking that the subnet manager service is active on the fabric and that IPoIB (or RDMA) interfaces are up is the correct first diagnostic step. Without subnet manager configuration, ports never transition to ACTIVE and inter-node bandwidth collapses regardless of GPU or driver health.

Why this answer

InfiniBand ports stuck in INIT have not been configured by a subnet manager, which is the authoritative service that initializes and activates fabric links. Checking the subnet manager and IPoIB configuration directly addresses why ports never reach ACTIVE. Without an active subnet manager, distributed training traffic cannot flow at expected speeds, so this is the correct first action.

Exam trap

The trap here is assuming that low inter-node throughput is always a GPU or driver problem, when an InfiniBand port in INIT points to a fabric-layer subnet manager or link configuration issue.

17
MCQmedium

An administrator is configuring a multi-tenant NVIDIA AI Enterprise environment. Which mechanism is most effective for ensuring hardware-level isolation between concurrent training jobs on a single A100 GPU?

A.Kubernetes namespaces with resource quotas
B.NVIDIA Driver process scheduling priority
C.Multi-Instance GPU (MIG) partitioning
D.Docker container CPU pinning
AnswerC

MIG hardware partitions ensure that each workload receives a dedicated set of compute units and memory buffers. By physically isolating the GPU resources, you guarantee deterministic performance for each tenant, which is necessary when running sensitive or high-throughput AI training models concurrently on a single hardware accelerator.

Why this answer

NVIDIA Multi-Instance GPU (MIG) allows a single physical GPU to be partitioned into multiple isolated instances, each with dedicated memory and compute cores. This is critical in multi-tenant AI environments because it prevents noisy neighbor issues, ensuring one workload's memory usage or compute demand does not degrade the performance of another. Proper resource partitioning is essential for maintaining strict SLAs and security boundaries within shared infrastructure.

Exam trap

Candidates often suggest software-level container limits or time-slicing when the question specifically asks for hardware-level isolation using partitioning features like MIG.

18
MCQmedium

An administrator supports a shared Kubernetes cluster running NVIDIA GPU Operator. Data scientists report that their inference pods remain in Pending state, yet the GPU Operator pods and node feature discovery pods are healthy, and the GPU nodes show no hardware alarms. The administrator confirms that the cluster has a mixture of MIG-capable A100 nodes and non-MIG T4 nodes. Which immediate administrative action is most appropriate to diagnose the scheduling failure?

A.Inspect the pod events with kubectl describe pod and verify the MIG strategy configured in the device plugin, because a mismatch between the node's MIG configuration and the requested resource can prevent GPU resource advertisement.
B.Restart the kubelet on all worker nodes to force the device plugin to re-register GPU resources, because a stale kubelet cache is the most frequent cause of Pending workloads.
C.Delete the pending pods and recreate them with a higher priority class, because the scheduler may be preempting them in favor of system pods.
D.Scale the GPU Operator controller manager to zero replicas and back to one, because the operator may have lost track of node labels after a recent upgrade.
AnswerA

Pod events reveal whether the scheduler rejected the pod due to missing nvidia.com/gpu or nvidia.com/mig-* resources. If the device plugin advertises MIG profiles but the workload requests a full GPU (or vice versa), the pod stays Pending. Checking the MIG strategy and node labels directly addresses this common scheduling mismatch.

Why this answer

The pod events are the fastest way to see whether the scheduler cannot find a suitable node due to resource requests. In a mixed MIG and non-MIG environment, the device plugin advertises different resource names, and a mismatch between the requested GPU resource and what the node exposes leaves pods Pending. Verifying the MIG strategy and node labels directly identifies that mismatch.

Exam trap

The trap here is assuming that healthy GPU Operator pods guarantee that GPU resources are correctly advertised and schedulable, when in fact a MIG strategy mismatch can leave pods Pending without any operator-level failure.

19
MCQhard

An administrator is responsible for a large NVIDIA DGX SuperPOD used for multi-node training. They need to ensure that GPU telemetry and health metrics are collected centrally and can trigger alerts when GPUs exceed temperature thresholds. Which component of NVIDIA Base Command Manager (BCM) should they configure to achieve this?

A.NVIDIA Base Command Manager's built-in 'cluster monitoring' dashboard
B.NVIDIA Management Library (NVML) on each node
C.NVIDIA Container Toolkit on each compute node
D.NVIDIA Data Center GPU Manager (DCGM) integrated with BCM
AnswerD

DCGM is designed for centralized GPU telemetry, health monitoring, and policy enforcement in data center environments. BCM integrates with DCGM to collect metrics from all nodes and can forward them to monitoring systems like Prometheus for alerting. Configuring DCGM within BCM provides the required centralized collection and threshold-based alerts for temperature and other metrics.

Why this answer

NVIDIA DCGM is the standard tool for centralized GPU telemetry and health monitoring in data centers. When integrated with Base Command Manager, it collects metrics from all nodes and can be configured with alerting rules for temperature and other thresholds. This provides the required centralized visibility and automated alerts for the DGX SuperPOD.

Exam trap

The trap here is confusing a monitoring dashboard with the actual telemetry collection component; the dashboard displays data but does not collect or alert on its own.

20
MCQhard

An administrator manages an NVIDIA AI Enterprise deployment on Kubernetes and needs to ensure that GPU telemetry is exported to an existing Prometheus instance. The administrator deploys the NVIDIA DCGM Exporter but sees no GPU metrics in Prometheus. Which configuration should the administrator verify first?

A.The NVIDIA GPU Operator's device plugin is configured to advertise GPUs with the correct resource name.
B.The GPU driver version installed on the nodes matches the version bundled in the DCGM Exporter container.
C.The cluster's CNI plugin supports multicast so that Prometheus can discover the exporter.
D.The DCGM Exporter ServiceMonitor or PodMonitor selector labels match the Prometheus operator's serviceMonitorSelector.
AnswerD

When Prometheus is managed by the Prometheus Operator, it only scrapes targets whose ServiceMonitor or PodMonitor labels match the operator's serviceMonitorSelector or podMonitorSelector. If the DCGM Exporter's monitor labels do not match, Prometheus silently ignores the endpoint even though the exporter is running. Verifying selector label alignment is the correct first step because it is the most common cause of missing metrics in operator-managed Prometheus.

Why this answer

In Prometheus Operator deployments, scrape targets are selected by label matching between the ServiceMonitor or PodMonitor and the Prometheus custom resource. If the DCGM Exporter's monitor labels do not align with the operator's selector, Prometheus never scrapes it, producing exactly the symptom of no GPU metrics. Verifying label alignment is therefore the correct first check.

Exam trap

The trap here is assuming that a running DCGM Exporter automatically gets scraped, when operator-managed Prometheus only scrapes targets whose monitor labels match its selector.

21
MCQmedium

An administrator wants to prevent unauthorized users from accessing sensitive model weights stored in GPU memory. Which security feature should be implemented to ensure hardware-level isolation of the memory space?

A.Enable standard Linux filesystem permissions.
B.Implement NVIDIA Confidential Computing.
C.Use a simple SSH firewall rule.
D.Update the NVIDIA CUDA shared library path.
AnswerB

NVIDIA Confidential Computing utilizes hardware-based Trusted Execution Environments (TEEs) to encrypt data while it resides in GPU memory. This prevents unauthorized access from other processes, the kernel, or the hypervisor, ensuring that sensitive model weights remain secure even if the software environment is considered untrusted or potentially compromised.

Why this answer

Hardware-level memory isolation via Confidential Computing technologies, such as NVIDIA Confidential Computing (CC), protects data in use by encrypting memory contents. For administrators handling sensitive IP, this provides a root-of-trust that persists even if the OS or hypervisor is compromised. This is a critical administrative control for ensuring compliance and data sovereignty in multi-tenant cloud environments where infrastructure is shared among different departments or organizations.

Exam trap

Candidates often confuse software-level encryption or standard disk encryption with hardware-level memory isolation. They incorrectly select general security tools instead of the specific NVIDIA Confidential Computing framework required for GPU-level memory protection.

22
Multi-Selectmedium

An AI operations team is using NVIDIA DCGM (Data Center GPU Manager) to monitor a cluster of A100 GPUs. They want to set up proactive health checks to detect and mitigate GPU issues before they cause job failures. Which two DCGM features should they configure? (Choose two.)

Select 2 answers
A.DCGM policy management for automated remediation
B.DCGM profiling metrics for real-time utilization
C.DCGM group configuration for multi-node synchronization
D.DCGM health checks with periodic diagnostics
E.DCGM API integration with Prometheus for alerting
AnswersA, D

DCGM policy management allows administrators to define policies that trigger actions when health violations occur, such as resetting a GPU or cordoning a node. This enables automated mitigation, reducing downtime. It works in conjunction with health checks. Configuring policies ensures that detected issues are handled promptly, which is essential for proactive health management in a large GPU cluster.

Why this answer

DCGM health checks with periodic diagnostics and DCGM policy management for automated remediation are the two features that directly enable proactive health monitoring and mitigation. Health checks detect issues, and policies define actions to take when issues are found. Together, they form a proactive health management system.

Other options are either monitoring metrics or management features that do not directly address proactive health.

Exam trap

The trap here is selecting monitoring metrics like profiling as health checks, when they only provide performance data, not issue detection and remediation.

23
MCQhard

An administrator is responsible for an NVIDIA AI Enterprise deployment on Kubernetes. The security team requires that all GPU-accelerated pods run with the least privilege necessary and that GPU device nodes are not exposed to pods that do not request them. Which combination of configurations should the administrator implement to meet these requirements?

A.Deploy the NVIDIA GPU Operator with the device plugin and configure a PodSecurityPolicy or OPA Gatekeeper policy that requires pods to request nvidia.com/gpu and forbids privileged mode.
B.Use the NVIDIA GPU Operator with the device plugin and configure pod security contexts to drop all capabilities and add only the NVIDIA_VISIBLE_DEVICES environment variable.
C.Install the NVIDIA k8s-device-plugin standalone and set the --pass-device-specs flag to true, then allow all pods to run as root.
D.Enable the NVIDIA device plugin with the --fail-on-init-error=false flag and set privileged: true in the pod security context.
AnswerA

The GPU Operator's device plugin exposes GPUs as schedulable resources (nvidia.com/gpu). Enforcing a policy that requires pods to request this resource ensures that only pods explicitly asking for GPUs receive device nodes. Forbidding privileged mode enforces least privilege. Together, these configurations ensure GPU device nodes are not exposed to pods that do not request them and that pods run with minimal privileges.

Why this answer

Using the GPU Operator with the device plugin, combined with a policy that mandates explicit nvidia.com/gpu requests and prohibits privileged containers, ensures that GPU device nodes are only injected into pods that request them and that pods run with minimal privileges. This satisfies both the least privilege and isolation requirements.

Exam trap

The trap here is believing that setting the NVIDIA_VISIBLE_DEVICES environment variable alone is sufficient for GPU access in Kubernetes; in reality, the device plugin must allocate the resource via a resource request, and without it the device nodes are not mounted.

24
Multi-Selecthard

An administrator is configuring NVIDIA GPUDirect Storage (GDS) on a cluster to accelerate data loading for AI training jobs. The cluster uses Mellanox InfiniBand adapters and NVMe storage. Which two actions are required to enable GDS and ensure optimal performance? (Choose two.)

Select 2 answers
A.Ensure that the storage devices and network adapters support RDMA and are properly configured for peer-to-peer communication.
B.Set the environment variable NVIDIA_GDS_ENABLE=1 on all nodes.
C.Configure the GPU nodes to use the NVIDIA Container Toolkit for all training containers.
D.Disable IOMMU in the BIOS to allow direct memory access between devices.
E.Install the NVIDIA GPUDirect Storage kernel module and user-space libraries on all GPU nodes.
AnswersA, E

GDS relies on RDMA and peer-to-peer (P2P) communication to transfer data directly between storage and GPU memory without CPU involvement. The storage devices (NVMe over Fabrics) and network adapters (InfiniBand) must support RDMA, and the system must be configured to allow P2P DMA. This includes enabling PCIe Access Control Services (ACS) and IOMMU settings that permit direct peer-to-peer transfers. Without this, GDS cannot achieve direct data paths.

Why this answer

Enabling GPUDirect Storage requires installing the GDS kernel module and user-space libraries on GPU nodes, and ensuring that storage and network hardware support RDMA and peer-to-peer communication. These two actions create the necessary software and hardware foundation for direct data transfers between storage and GPU memory. Other options are either incorrect or not specific to GDS.

Exam trap

The trap here is assuming that GDS can be enabled by a simple environment variable or that container toolkit configuration is sufficient, when it actually requires low-level driver and hardware configuration.

25
MCQmedium

A company runs multiple AI workloads on a shared Kubernetes cluster with NVIDIA GPUs. The administrator needs to enforce that only pods with a specific label can consume GPU resources, while other pods are denied. Which Kubernetes admission control mechanism should be used to implement this policy?

A.PodSecurityPolicy
B.ValidatingAdmissionWebhook
C.ResourceQuota
D.LimitRange
AnswerB

A ValidatingAdmissionWebhook can intercept pod creation requests and evaluate custom logic, such as checking for a specific label before allowing the pod to request GPU resources. This provides the flexibility to enforce label-based policies on extended resources. By deploying a webhook that inspects pod labels and resource requests, the administrator can deny non-compliant pods, meeting the requirement.

Why this answer

A ValidatingAdmissionWebhook allows custom admission logic to be applied to pod creation. By configuring a webhook that checks for a specific label before permitting GPU resource requests, the administrator can enforce label-based access control. This is the correct mechanism because it can inspect pod metadata and resource requests and reject non-compliant pods.

Exam trap

The trap here is assuming that ResourceQuota can enforce per-pod label conditions, when it only limits aggregate namespace consumption.

26
MCQmedium

An administrator is managing an NVIDIA DGX SuperPOD used for large-scale AI training. The cluster uses a Slurm workload manager. The administrator needs to ensure that jobs are scheduled only on nodes with healthy GPUs and that failed GPUs are automatically drained from the pool. Which integration should be configured to achieve this?

A.NVIDIA Base Command Manager with Slurm integration
B.NVIDIA Fleet Command with Slurm edge scheduling
C.NVIDIA Data Center GPU Manager (DCGM) with the Slurm health check plugin
D.NVIDIA Container Toolkit with Slurm's --gres flag
AnswerC

DCGM provides comprehensive GPU health monitoring, and its integration with Slurm via the health check plugin allows automatic detection of unhealthy GPUs. When DCGM identifies a failed GPU, the plugin can drain the node from Slurm, preventing new jobs from being scheduled on it. This directly addresses the requirement to schedule only on healthy GPUs and automatically remove failed ones.

Why this answer

DCGM with the Slurm health check plugin is the correct integration because it continuously monitors GPU health and can automatically drain nodes with failed GPUs from Slurm, ensuring jobs run only on healthy hardware. Other options lack the necessary health monitoring and automatic remediation capabilities.

Exam trap

The trap here is assuming that Base Command Manager or the Container Toolkit alone can handle GPU health monitoring; they manage deployment and resource allocation, but DCGM is specifically designed for health checks and integration with schedulers like Slurm.

27
MCQeasy

When monitoring GPU health in an enterprise cluster, which command provides the most comprehensive snapshot of real-time power, temperature, and memory utilization?

A.nvcc --version
B.nvidia-smi
C.nvidia-bug-report.sh
D.cat /proc/driver/nvidia/gpus/*/information
AnswerB

The nvidia-smi tool is specifically designed to provide a comprehensive view of all NVIDIA GPUs installed in the system. It displays critical health metrics like power usage, temperature, memory usage, and compute utilization in a real-time format, serving as the primary diagnostic tool for AI infrastructure administrators.

Why this answer

The 'nvidia-smi' utility is the industry-standard tool for GPU management and monitoring. It provides a tabular overview of the system, including per-GPU power draw, thermal status, memory allocation, and active process lists. For administrators, this is the first point of contact for diagnosing performance bottlenecks or thermal throttling issues in a data center, making it the essential command for routine maintenance and operational health monitoring of NVIDIA hardware.

Exam trap

Candidates often look for complex monitoring software or cloud-native dashboards, forgetting that the native 'nvidia-smi' tool is the fundamental, built-in utility for real-time hardware diagnostics.

28
MCQhard

An administrator is configuring NVIDIA Base Command Manager to manage a cluster of DGX nodes. They want to ensure that when a node's GPU temperature exceeds a defined threshold, the node is automatically drained and an alert is sent to the operations team. Which combination of Base Command Manager features should the administrator configure to achieve this?

A.Health checks with a custom script that triggers a node drain via the Base Command Manager API, and an alert rule that sends a notification.
B.NVIDIA GPU Operator with node problem detector, and a Kubernetes taint that evicts pods.
C.Prometheus Alertmanager with a webhook that drains the node, and Grafana dashboards for visualization.
D.Base Command Manager job scheduler with a preemption policy that kills jobs when temperature is high.
AnswerA

Base Command Manager health checks can run custom scripts periodically. A script can query GPU temperature and, if over threshold, call the Base Command Manager API to drain the node. An alert rule then sends a notification. This provides automated remediation and alerting as required.

Why this answer

Base Command Manager provides health checks that can execute custom scripts on nodes. A script can monitor GPU temperature and invoke the Base Command Manager API to drain the node, while an alert rule notifies the team. This leverages native features for automated remediation and alerting, unlike external monitoring stacks or Kubernetes-specific tools.

Exam trap

The trap here is confusing Base Command Manager's native health and alerting features with external monitoring tools like Prometheus or Kubernetes node problem detector.

29
MCQeasy

An administrator is deploying NVIDIA AI Enterprise on a Kubernetes cluster and wants to verify that the GPU Operator has successfully installed all required components. Which command should the administrator use to check the status of the GPU Operator pods?

A.nvidia-smi -q
B.kubectl get nodes -o wide
C.dcgmi discovery -l
D.kubectl get pods -n gpu-operator
AnswerD

The GPU Operator deploys its components into the gpu-operator namespace by default. Running kubectl get pods -n gpu-operator lists all pods in that namespace, allowing the administrator to see if the operator, driver, container toolkit, device plugin, and other components are running. This is the standard way to verify the operator's deployment status.

Why this answer

The GPU Operator installs its components as pods in the gpu-operator namespace. Checking pod status with kubectl get pods -n gpu-operator is the correct way to verify that all required components are running. Other commands provide GPU-level or node-level information but do not show the status of the operator's pods.

Exam trap

The trap here is confusing host-level GPU queries with Kubernetes pod status checks; nvidia-smi and dcgmi do not show operator pods.

30
MCQmedium

An AI operations engineer is preparing a DGX H100 system for a multi-node training workload. The engineer runs `nvidia-smi topo -m` and notices that GPU4 and GPU5 report a connection type of SYS, while all other GPU pairs show NV18. What is the most likely cause of this topology anomaly?

A.The NVSwitch fabric is operating in a fallback mode that only activates when peer-to-peer traffic exceeds a threshold.
B.GPU4 and GPU5 are configured in MIG mode, which forces inter-GPU communication over the PCIe bus.
C.GPU4 and GPU5 are connected through the PCIe host bridge rather than the NVSwitch fabric, possibly due to a degraded or missing NVLink connection.
D.GPU4 and GPU5 are reserved for display output, so the driver automatically disables their NVLink interfaces.
AnswerC

SYS indicates that communication between those two GPUs traverses the system's PCIe and CPU interconnect instead of the high-speed NVLink/NVSwitch fabric. On a DGX H100, all GPU pairs should normally show NV18 via NVSwitch, so this points to a degraded or missing NVLink path between GPU4 and GPU5 that must be investigated.

Why this answer

A healthy DGX H100 shows NV18 (NVLink 4.0, 18 links) for every GPU pair because all GPUs are connected through NVSwitch. A SYS entry for a specific pair means traffic between those GPUs falls back to the PCIe/CPU path, which severely reduces bandwidth and increases latency. This typically signals a hardware or cabling fault on the NVLink side that should be resolved before running distributed training.

Exam trap

The trap here is assuming that SYS is a normal topology state for some GPU pairs on an NVSwitch-equipped system, when it actually indicates a degraded or absent NVLink path.

31
MCQeasy

Which NVIDIA technology enables a GPU to be shared among multiple virtual machines or containers while maintaining strict hardware isolation?

A.NVIDIA CUDA streams
B.NVIDIA Multi-Instance GPU (MIG)
C.NVIDIA Collective Communications Library (NCCL)
D.NVIDIA GPUDirect RDMA
AnswerB

MIG allows a physical GPU to be partitioned into multiple isolated instances at the hardware level. Each instance has its own dedicated memory and compute resources, providing the isolation necessary for multi-tenant environments where security and performance guarantees are required for concurrent, independent workloads on a single piece of hardware.

Why this answer

NVIDIA Multi-Instance GPU (MIG) is the core technology that enables hardware-level partitioning on supported GPUs. By creating dedicated instances that have their own memory and compute resources, MIG ensures that workloads remain isolated, providing predictable quality of service and security in multi-tenant environments. This is a foundational technology for maximizing GPU utilization in cloud and enterprise data center environments where mixed workloads are common.

Exam trap

Test-takers sometimes confuse software-based time-slicing with true hardware isolation technologies like MIG when multi-tenant security is required.

32
MCQmedium

An administrator is troubleshooting a multi-node NVIDIA GPU training job that intermittently hangs during the all-reduce phase. The job uses NCCL over InfiniBand. Logs show that some ranks time out while others complete. The administrator suspects a network fabric issue. Which action should the administrator take first to isolate whether the problem is in the InfiniBand fabric or in the NCCL configuration?

A.Disable InfiniBand and force NCCL to use TCP sockets instead, because TCP is more reliable for collective operations.
B.Immediately replace all InfiniBand cables and transceivers on the affected nodes, because intermittent hangs during all-reduce almost always indicate physical link errors.
C.Run the NCCL tests (nccl-tests) with the same topology and environment variables, and enable NCCL debug logging to capture detailed transport and topology information.
D.Increase the NCCL timeout value in the training script and restart the job, because the hangs are likely due to slow network convergence.
AnswerC

Running nccl-tests with debug logging reproduces the collective communication pattern outside the training job. It reveals which transport (InfiniBand, RoCE, or sockets) is used, whether topology detection fails, and where timeouts occur. This isolates NCCL configuration issues from fabric problems before deeper network diagnostics.

Why this answer

Using nccl-tests with debug logging is a controlled way to reproduce the collective communication and observe transport selection, topology detection, and timeouts. It helps determine whether the issue is in NCCL configuration or the InfiniBand fabric. Replacing hardware, forcing TCP, or increasing timeouts are premature and do not isolate the root cause.

Exam trap

The trap here is jumping to hardware replacement or timeout adjustments before verifying whether NCCL is correctly configured and using the expected transport.

33
Multi-Selecthard

Which THREE components are required for a container to successfully leverage NVIDIA GPUs on a Kubernetes cluster?

Select 3 answers
A.NVIDIA Container Toolkit
B.Kubernetes Device Plugin for NVIDIA
C.NVIDIA Data Center GPU Manager (DCGM)
D.NVIDIA Driver installed on the host node
E.Pre-installed CUDA Toolkit inside the container
AnswersA, B, D

The NVIDIA Container Toolkit is essential for enabling the container runtime to interact with the host's NVIDIA drivers. It performs the necessary steps to inject the GPU devices into the container namespace and ensures that the required libraries are mapped correctly so the application can communicate with the GPU hardware.

Why this answer

Successful GPU integration in Kubernetes relies on the interaction between the container engine, the NVIDIA driver, and the device plugin. The driver provides the kernel-level interface, the container runtime handles the low-level injection, and the device plugin advertises the hardware to the scheduler. These three parts must be correctly configured and compatible to ensure that containers can schedule and utilize GPUs effectively without manual configuration or persistent connectivity issues.

Exam trap

Candidates often forget the role of the NVIDIA Container Toolkit, incorrectly assuming that the driver and device plugin alone are sufficient to bridge the container runtime to the GPU.

34
MCQmedium

An administrator is managing a Kubernetes cluster with NVIDIA GPU Operator. A new policy requires that all GPU workloads run with specific environment variables set, such as NVIDIA_VISIBLE_DEVICES and NVIDIA_DRIVER_CAPABILITIES. The administrator wants to enforce these variables automatically for any pod that requests a GPU, without modifying each pod specification manually. Which approach should the administrator use?

A.Use a ValidatingAdmissionWebhook to reject pods that do not include the required environment variables.
B.Configure the NVIDIA GPU Operator to set default environment variables in the device plugin daemonset.
C.Create a MutatingAdmissionWebhook that injects the required environment variables into pods that request nvidia.com/gpu resources.
D.Set the environment variables in the NVIDIA Container Runtime configuration on each node.
AnswerC

A MutatingAdmissionWebhook can intercept pod creation requests and automatically modify the pod specification to add environment variables when the pod requests GPU resources. This enforces the policy cluster-wide without requiring manual changes to each pod. It is a standard Kubernetes mechanism for such automation and integrates well with the GPU Operator's resource requests.

Why this answer

A MutatingAdmissionWebhook is the correct Kubernetes-native way to automatically inject environment variables into pods based on their resource requests. It can inspect incoming pod specs and add the required variables when a GPU resource is requested, ensuring consistent configuration without manual intervention. Validating webhooks only reject, and runtime configuration is not pod-specific.

Exam trap

The trap here is confusing validating and mutating admission webhooks, or thinking that node-level runtime configuration can enforce pod-specific environment variables.

35
MCQmedium

Which administrative practice ensures that a cluster is prepared for the arrival of new NVIDIA GPU hardware with minimal downtime?

A.Wait for the hardware to arrive before checking driver support.
B.Update the OS and NVIDIA drivers across the cluster.
C.Manually install every CUDA library on all nodes.
D.Modify the BIOS settings to ignore PCIe errors.
AnswerB

Proactively updating the OS and drivers to versions that support the upcoming GPU hardware ensures that when the physical installation occurs, the software stack is ready. This minimizes the time spent troubleshooting driver incompatibilities and allows for a smooth, plug-and-play integration of the new hardware into the production cluster.

Why this answer

Maintaining an up-to-date driver and software stack via a robust package management system is the key to minimizing downtime during hardware upgrades. Administrators must ensure that the kernel and driver versions are compatible with the new hardware before it arrives. This proactive preparation is vital for maximizing cluster uptime and ensuring that researchers can immediately begin using the new hardware for their AI experiments without configuration delays.

Exam trap

Candidates often assume physical hardware installation alone or container runtime updates are sufficient, overlooking the critical requirement to synchronize the underlying host OS kernel and NVIDIA driver stack.

36
Multi-Selectmedium

An administrator is responsible for maintaining an NVIDIA AI Enterprise cluster and needs to ensure high availability of GPU resources for critical inference workloads. Which two practices should the administrator implement? (Choose two.)

Select 2 answers
A.Enable GPU time-slicing to allow multiple pods to share a single GPU
B.Schedule all inference pods on a single node with multiple GPUs to reduce network latency
C.Use NVIDIA MIG to partition a GPU into isolated instances for each workload
D.Implement pod disruption budgets to maintain a minimum number of available replicas
E.Configure node affinity and anti-affinity rules to spread pods across multiple nodes
AnswersD, E

Pod disruption budgets (PDBs) ensure that a specified minimum number of pods remain available during voluntary disruptions, such as node drains or upgrades. This helps maintain service availability for critical inference workloads. By defining a PDB, the administrator can prevent simultaneous eviction of too many replicas, thus supporting high availability.

Why this answer

High availability for inference workloads requires spreading pods across multiple nodes to avoid single points of failure and using pod disruption budgets to maintain minimum replica counts during disruptions. Node anti-affinity ensures distribution, while PDBs protect against voluntary evictions. Time-slicing, MIG, and single-node concentration do not provide fault tolerance against node or GPU failures.

Exam trap

The trap here is equating resource sharing or isolation features like MIG or time-slicing with high availability, when they do not protect against hardware failure.

37
MCQmedium

An administrator manages an NVIDIA AI Enterprise cluster using NVIDIA Run:ai. A data science team complains that their submitted training job has been stuck in a Pending state for over an hour, even though the Run:ai scheduler shows free GPUs in the cluster. The administrator verifies that the job requests 2 GPUs and the node pool has 4 idle GPUs. Which Run:ai administrative configuration is the most likely cause of the job remaining Pending?

A.The job's container image does not include the NVIDIA Container Toolkit, so the scheduler cannot start the job.
B.The project's GPU quota is exhausted by other running workloads, so the scheduler cannot allocate the requested GPUs.
C.The NVIDIA GPU Operator is not installed on the cluster, so the scheduler cannot see GPU resources.
D.The node pool is configured with a node affinity that excludes the nodes with idle GPUs, so the scheduler cannot place the job.
AnswerB

Run:ai projects enforce a GPU quota per project. Even if the cluster has idle GPUs, a job in a project that has already consumed its quota will remain Pending until quota is freed. This matches the scenario where the cluster shows free GPUs but the job cannot start because the project-level quota is the limiting factor.

Why this answer

Run:ai uses projects to enforce GPU quotas. When a project's quota is fully consumed, new jobs remain Pending even if the cluster has idle GPUs. The administrator should check the project's quota usage and either increase the quota or free resources.

Node affinity, missing GPU Operator, or missing Container Toolkit would produce different symptoms.

Exam trap

The trap here is assuming that free GPUs in the cluster automatically mean a job can be scheduled, ignoring project-level quota enforcement in Run:ai.

38
MCQeasy

An administrator is responsible for monitoring a large-scale AI cluster with hundreds of NVIDIA GPUs. They need to collect telemetry data such as GPU utilization, temperature, and power consumption from all nodes and store it centrally for analysis and alerting. Which NVIDIA tool should they use to collect and export GPU metrics to a monitoring system like Prometheus?

A.NVIDIA Nsight Systems
B.NVIDIA DCGM Exporter
C.nvidia-smi
D.NVIDIA Base Command Manager
AnswerB

NVIDIA DCGM Exporter is a Prometheus exporter that collects GPU metrics using NVIDIA Data Center GPU Manager (DCGM) and exposes them in a format that Prometheus can scrape. It provides a wide range of metrics, including utilization, temperature, power, and memory usage, making it ideal for centralized monitoring of large GPU clusters.

Why this answer

For centralized monitoring of GPU metrics in a large cluster, NVIDIA DCGM Exporter is the appropriate tool. It leverages NVIDIA Data Center GPU Manager (DCGM) to collect a comprehensive set of metrics from each GPU and exposes them via an HTTP endpoint that Prometheus can scrape. This allows administrators to aggregate metrics across all nodes, create dashboards in Grafana, and set up alerts based on thresholds.

Exam trap

The trap here is confusing nvidia-smi with a scalable monitoring solution, when it is only a per-node command-line tool without native Prometheus integration.

39
MCQmedium

An administrator manages an NVIDIA AI Enterprise cluster and needs to enforce GPU resource quotas across multiple Kubernetes namespaces. Which NVIDIA component should be configured to enforce these quotas?

A.Kubernetes ResourceQuota with extended resources
B.NVIDIA DCGM Exporter
C.NVIDIA Base Command Manager
D.NVIDIA GPU Operator
AnswerA

Kubernetes ResourceQuota objects can limit the aggregate quantity of extended resources, such as nvidia.com/gpu, that can be requested within a namespace. By defining a ResourceQuota that specifies a maximum for nvidia.com/gpu, the administrator can enforce GPU quotas across namespaces. This is the native Kubernetes mechanism for quota enforcement and works in conjunction with the NVIDIA device plugin that advertises GPUs as extended resources.

Why this answer

Kubernetes ResourceQuota is the native mechanism to limit aggregate resource consumption per namespace, including extended resources like nvidia.com/gpu. When the NVIDIA device plugin advertises GPUs, administrators can define a ResourceQuota specifying a maximum for nvidia.com/gpu, thereby enforcing GPU quotas. Other NVIDIA tools focus on deployment, management, or monitoring and do not provide quota enforcement.

Exam trap

The trap here is assuming that NVIDIA GPU Operator or Base Command Manager enforces GPU quotas, when quota enforcement is actually a Kubernetes-native function.

40
MCQmedium

An administrator is preparing a bare-metal GPU server for AI workloads and needs to verify that the NVIDIA driver and CUDA toolkit are properly installed. The server has an NVIDIA A100 GPU. Which command should the administrator run to display the GPU model, driver version, and CUDA version?

A.nvidia-debugdump --list
B.nvcc --version
C.lspci | grep -i nvidia
D.nvidia-smi
AnswerD

The nvidia-smi command queries the NVIDIA driver and displays detailed information about all installed GPUs, including the GPU model, driver version, CUDA version, and current utilization. It is the standard tool for verifying driver and CUDA compatibility on a GPU-enabled system, making it the correct choice for this scenario.

Why this answer

The nvidia-smi utility is the primary interface for monitoring and managing NVIDIA GPUs. It provides a summary of each GPU's model, driver version, CUDA version, temperature, power usage, and memory utilization. For an administrator validating a new server, running nvidia-smi immediately confirms that the driver is loaded and the GPU is recognized, and it shows the CUDA version supported by the driver, which is essential for AI workload compatibility.

Exam trap

The trap here is assuming that nvcc --version reports the driver version, when it actually only reports the CUDA compiler version.

41
MCQmedium

An administrator needs to ensure that all GPU drivers are updated across a heterogeneous cluster without causing downtime. What is the best strategy?

A.Simultaneously update all nodes during a maintenance window
B.Perform rolling updates by draining nodes one by one
C.Use a containerized driver approach to avoid host updates
D.Only update the driver on the master node
AnswerB

Rolling updates ensure that the cluster remains operational throughout the entire upgrade process. By draining one node at a time, the administrator can safely update drivers without terminating active workloads prematurely. This approach provides a clear path for verification and rollback if the update encounters any unforeseen compatibility issues.

Why this answer

Implementing a rolling update strategy allows the cluster to be updated one node at a time while the remaining nodes continue to process training jobs. By draining the target node of all active jobs, upgrading the drivers, and then re-adding the node to the production pool, the administrator maintains service availability. This is the standard practice in production AI operations to minimize disruption and ensure consistent driver versions across the entire fleet.

Exam trap

Candidates often select disruptive cluster-wide reboots or simultaneous upgrades instead of controlled, node-by-node rolling updates with draining.

42
MCQhard

Refer to the exhibit. An administrator observes this in the logs during a multi-GPU training job. What is the performance implication of this setting?

A.The training job will have lower memory consumption
B.GPU-to-GPU data transfer performance is significantly reduced
C.The job will be more stable across nodes
D.The system will utilize less power during training
AnswerB

By disabling P2P, the system is forced to move data through the PCIe bus or system memory, which is significantly slower than using NVLink or direct P2P access. This causes a major bottleneck in collective operations like AllReduce, which are fundamental to the scalability of distributed AI training jobs.

Why this answer

Disabling Peer-to-Peer (P2P) communication forces NCCL to use system memory (often via the network or PCIe bus) to facilitate data movement between GPUs. This bypasses the high-bandwidth NVLink interconnect, resulting in significantly increased latency and lower throughput for collective operations. This configuration is typically used only for debugging or when hardware incompatibility prevents direct P2P access, as it severely hinders the performance of multi-GPU, multi-node AI training workloads.

Exam trap

Candidates often assume that if a training job runs, it is running optimally, failing to realize that disabling P2P forces traffic over the slower PCIe bus instead of high-speed NVLink.

43
MCQmedium

An administrator notices that GPU utilization is high, but throughput in an AI training job remains low. What is the most likely bottleneck?

A.Insufficient CUDA cores on the GPU
B.Data pipeline or storage I/O bottleneck
C.Incompatible NVIDIA driver version
D.Excessive usage of GPU registers
AnswerB

High GPU utilization accompanied by low throughput indicates the GPU is spending time waiting for data to arrive from the CPU or storage. This starvation effect is a classic symptom of an inefficient data loader or slow storage subsystem, which limits the overall throughput despite the GPU's apparent activity.

Why this answer

When GPU utilization is high but throughput is low, the GPU is likely stalling while waiting for data. This is typically a sign of an input/output (I/O) bottleneck, where the data pipeline (e.g., loading images from disk or network) cannot keep up with the GPU's processing speed. Ensuring the data preprocessing pipeline is sufficiently parallelized and optimized is crucial for maximizing GPU utilization and maintaining high training throughput in AI workloads.

Exam trap

Candidates often assume that high GPU utilization is a sign of a healthy, efficient training job, failing to recognize that it can actually indicate the GPU is idling while waiting for data.

44
MCQeasy

An administrator is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing several NVIDIA A100 GPUs. They need to enable GPU sharing across multiple virtual machines to maximize utilization. Which vSphere feature should they configure?

A.vSphere DRS
B.SR-IOV
C.NVIDIA vGPU
D.DirectPath I/O
AnswerC

NVIDIA vGPU allows a single physical GPU to be partitioned into multiple virtual GPUs, each assigned to a different VM. This enables sharing and maximizes utilization. It requires vSphere with NVIDIA vGPU software and compatible GPUs. This is the correct feature for GPU sharing in vSphere.

Why this answer

NVIDIA vGPU is the vSphere feature that partitions a physical GPU into multiple virtual GPUs, allowing multiple VMs to share the same physical GPU. This maximizes GPU utilization and is designed for AI workloads. DirectPath I/O provides exclusive access, while DRS and SR-IOV do not enable GPU sharing.

Exam trap

The trap here is confusing GPU sharing with GPU passthrough, where DirectPath I/O gives exclusive access to one VM, not sharing.

Ready to test yourself?

Try a timed practice session using only Administration questions.