Courseiva

NVIDIA Certified Professional: AI Operations (NCP-AIO) — Questions 226–300

309 questions total · 5pages · All types, answers revealed

Page 3

Page 4 of 5

Page 5
226
MCQmedium

A platform team runs a Kubernetes cluster where the NVIDIA GPU Operator is installed and time-slicing is configured with a ConfigMap that advertises four replicas per physical GPU. A data scientist submits a PyTorch training pod requesting nvidia.com/gpu: 1. The pod stays Pending indefinitely, and the scheduler event reads 'Insufficient nvidia.com/gpu'. The node's GPUs are otherwise idle and healthy. Which action most directly resolves the pending state?

A.Restart the nvidia-device-plugin pod so it re-reads the time-slicing ConfigMap and republishes the replicated resource capacity to the kubelet.
B.Add a nodeSelector that pins the pod to a specific GPU node, because the scheduler cannot match GPU requests without an explicit node affinity rule.
C.Verify the time-slicing ConfigMap is labeled for the device plugin and that the node's GPU capacity now reports the multiplied replica count, then requeue the pod.
D.Change the pod's resource request to nvidia.com/gpu.shared: 1 so it matches the resource name that time-slicing exposes.
AnswerC

Time-slicing only takes effect when the ConfigMap carries the label the device plugin watches and the plugin successfully reloads it. Confirming the node's allocatable nvidia.com/gpu reflects replicas times physical GPUs proves the configuration propagated. If capacity is still one per GPU, the ConfigMap is unlabeled or malformed, and correcting that plus requeuing the pod restores schedulability.

Why this answer

Time-slicing is enabled by a ConfigMap that the NVIDIA device plugin watches; without the expected label the plugin ignores it and keeps advertising one GPU per device. Confirming the ConfigMap is labeled and that the node's allocatable nvidia.com/gpu equals replicas multiplied by physical GPUs verifies the change propagated. Once capacity is correctly republished, the pending pod can be scheduled without altering its resource request.

Exam trap

The trap here is assuming time-slicing creates a distinct resource name such as nvidia.com/gpu.shared, when it actually multiplies the existing nvidia.com/gpu count.

227
MCQhard

A system administrator is troubleshooting a 'CUDA error: invalid device ordinal' when launching a job on a multi-GPU system. What is the most likely cause?

A.The GPU driver is corrupted and requires a reinstall.
B.CUDA_VISIBLE_DEVICES is set to an out-of-range index.
C.The GPU memory is full.
D.The InfiniBand fabric is misconfigured.
AnswerB

If the environment variable CUDA_VISIBLE_DEVICES specifies an index that does not exist on the host, the CUDA runtime will throw an 'invalid device ordinal' error. This is a configuration error where the system believes it has fewer GPUs than the application is trying to access via the environment variable.

Why this answer

The 'invalid device ordinal' error typically indicates that the application is attempting to access a GPU index (e.g., GPU 4) that does not exist or is not visible to the process. This is often caused by environment variables like CUDA_VISIBLE_DEVICES being configured incorrectly, mapping the application to non-existent hardware. Ensuring the logical-to-physical GPU mapping is accurate is vital for correct job execution on shared infrastructure.

Exam trap

Candidates frequently assume this error implies a faulty physical GPU or driver crash, overlooking the simpler and more common configuration error where environment variables restrict access to non-existent indices.

228
Multi-Selecthard

Which THREE components are required for a container to successfully leverage NVIDIA GPUs on a Kubernetes cluster?

Select 3 answers
A.NVIDIA Container Toolkit
B.Kubernetes Device Plugin for NVIDIA
C.NVIDIA Data Center GPU Manager (DCGM)
D.NVIDIA Driver installed on the host node
E.Pre-installed CUDA Toolkit inside the container
AnswersA, B, D

The NVIDIA Container Toolkit is essential for enabling the container runtime to interact with the host's NVIDIA drivers. It performs the necessary steps to inject the GPU devices into the container namespace and ensures that the required libraries are mapped correctly so the application can communicate with the GPU hardware.

Why this answer

Successful GPU integration in Kubernetes relies on the interaction between the container engine, the NVIDIA driver, and the device plugin. The driver provides the kernel-level interface, the container runtime handles the low-level injection, and the device plugin advertises the hardware to the scheduler. These three parts must be correctly configured and compatible to ensure that containers can schedule and utilize GPUs effectively without manual configuration or persistent connectivity issues.

Exam trap

Candidates often forget the role of the NVIDIA Container Toolkit, incorrectly assuming that the driver and device plugin alone are sufficient to bridge the container runtime to the GPU.

229
MCQhard

A company is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing NVIDIA A100 GPUs. They need to enable vGPU functionality for virtual machines running AI workloads. Which configuration step is required on the ESXi host to allow vGPU assignment to VMs?

A.Install the NVIDIA vGPU Manager on the ESXi host and set the graphics type to shared or direct.
B.Install the NVIDIA GPU Operator on the ESXi host and create a custom resource for vGPU.
C.Enable SR-IOV on the ESXi host and configure virtual functions for each VM.
D.Install the NVIDIA Container Toolkit on the ESXi host and configure the Docker daemon.
AnswerA

To enable vGPU on ESXi, you must install the NVIDIA vGPU Manager VIB on each host and configure the GPU's graphics type. For vGPU, the graphics type is typically set to 'shared' to allow multiple VMs to share the GPU, or 'direct' for passthrough. This is a required step before vGPU profiles can be assigned to VMs.

Why this answer

Enabling vGPU on ESXi requires the NVIDIA vGPU Manager VIB and setting the GPU graphics type appropriately. This allows the hypervisor to present virtual GPUs to VMs. Alternatives like the Container Toolkit or GPU Operator are not applicable to ESXi host-level configuration, and SR-IOV is not the mechanism used for vGPU in this context.

Exam trap

The trap here is assuming that Kubernetes-centric tools like the GPU Operator or Container Toolkit are used for vSphere vGPU setup, when actually the hypervisor-specific vGPU Manager is required.

230
Multi-Selectmedium

An AI operations engineer is troubleshooting a Kubernetes cluster where several GPU training pods fail to start with a device plugin allocation error, even though the nodes report healthy GPUs. The engineer suspects the pods are requesting more GPU resources than a single physical card can provide without a sharing mechanism. Which TWO configurations would legitimately allow multiple pods to consume a single physical GPU on these nodes? (Choose two.)

Select 2 answers
A.Enable NVIDIA Multi-Instance GPU (MIG) on supported GPUs and expose the MIG instances as schedulable resources.
B.Configure time-slicing in the device plugin so several pods share a GPU through interleaved execution.
C.Add a node selector that allows multiple pods to bind to the same nvidia.com/gpu resource slot.
D.Create a ResourceQuota that counts GPU requests as fractional values such as 0.5 per pod.
E.Raise the GPU Operator's driver version so the device plugin reports extra virtual GPUs per card.
AnswersA, B

MIG partitions a supported GPU into isolated instances, each with dedicated memory and compute slices. When the GPU Operator exposes these instances through the device plugin, each MIG instance is advertised as its own resource, so multiple pods can run concurrently on one physical card with hardware-level isolation. This is a supported way to share a single GPU across pods.

Why this answer

Both MIG and time-slicing change how the device plugin advertises and allocates a physical GPU, enabling concurrent consumption by more than one pod. MIG offers hardware-partitioned, isolated instances, while time-slicing interleaves workloads on the whole card. Each is a supported sharing mode configured through the GPU Operator, unlike driver upgrades, node selectors, or quota edits, which do not alter device advertisement.

Exam trap

The trap here is believing that raising a driver version or adjusting scheduling hints can create shareable GPUs, when only explicit device plugin sharing modes change resource advertisement.

231
MCQmedium

An administrator is managing a Kubernetes cluster with NVIDIA GPU Operator. A new policy requires that all GPU workloads run with specific environment variables set, such as NVIDIA_VISIBLE_DEVICES and NVIDIA_DRIVER_CAPABILITIES. The administrator wants to enforce these variables automatically for any pod that requests a GPU, without modifying each pod specification manually. Which approach should the administrator use?

A.Use a ValidatingAdmissionWebhook to reject pods that do not include the required environment variables.
B.Configure the NVIDIA GPU Operator to set default environment variables in the device plugin daemonset.
C.Create a MutatingAdmissionWebhook that injects the required environment variables into pods that request nvidia.com/gpu resources.
D.Set the environment variables in the NVIDIA Container Runtime configuration on each node.
AnswerC

A MutatingAdmissionWebhook can intercept pod creation requests and automatically modify the pod specification to add environment variables when the pod requests GPU resources. This enforces the policy cluster-wide without requiring manual changes to each pod. It is a standard Kubernetes mechanism for such automation and integrates well with the GPU Operator's resource requests.

Why this answer

A MutatingAdmissionWebhook is the correct Kubernetes-native way to automatically inject environment variables into pods based on their resource requests. It can inspect incoming pod specs and add the required variables when a GPU resource is requested, ensuring consistent configuration without manual intervention. Validating webhooks only reject, and runtime configuration is not pod-specific.

Exam trap

The trap here is confusing validating and mutating admission webhooks, or thinking that node-level runtime configuration can enforce pod-specific environment variables.

232
MCQmedium

A platform team is deploying the NVIDIA GPU Operator on a Kubernetes cluster to manage GPU nodes. They want the Operator to automatically install the NVIDIA driver, the container toolkit, and the device plugin on each GPU node. Which component of the GPU Operator is responsible for installing the NVIDIA driver on the host?

A.NVIDIA Container Toolkit
B.NVIDIA Driver Container
C.NVIDIA Device Plugin
D.NVIDIA GPU Operator Validator
AnswerB

The NVIDIA Driver Container is a containerized version of the NVIDIA driver that the GPU Operator deploys as a DaemonSet on GPU nodes. It installs the driver into the host's filesystem, enabling driver management without pre-installing the driver on the host OS. This is exactly what the team needs for automatic driver installation.

Why this answer

The GPU Operator automates the deployment of all NVIDIA software components on Kubernetes. The Driver Container is specifically designed to install the NVIDIA driver on the host. It runs as a DaemonSet and installs the driver into the host's filesystem, eliminating the need for manual driver installation.

The other components handle container runtime integration, resource advertisement, and validation.

Exam trap

The trap here is assuming that the NVIDIA Container Toolkit installs the driver, when it actually only enables containers to access GPUs and requires the driver to be present.

233
Multi-Selecthard

An AI researcher is debugging a multi-node training job using NCCL. Which TWO actions should they take to diagnose potential network-related performance degradation?

Select 2 answers
A.Set NCCL_DEBUG=INFO to monitor communication patterns and connection issues.
B.Increase the number of threads in the PyTorch DataLoader.
C.Run nccl-tests to establish a baseline for collective performance.
D.Switch from NCCL to MPI for all communication operations.
E.Lower the precision of the model to float16.
AnswersA, C

Setting NCCL_DEBUG to INFO provides detailed logs about how NCCL discovers the network topology and establishes peer-to-peer connections between nodes. This is essential for identifying misconfigured interconnects or slow paths that could be hindering collective operation performance during distributed training cycles across multiple GPU instances.

Why this answer

Debugging multi-node communication requires verifying both software configurations and physical network health. NCCL_DEBUG settings provide granular insight into connection establishment and topology detection, while NCCL tests provide baseline performance metrics. Identifying these issues early is essential for scaling models across large clusters, as network overhead can quickly become the dominant factor in training time during distributed synchronized gradient descent.

Exam trap

Candidates often attempt to debug the code logic or model hyperparameters, ignoring the specialized diagnostic tools (NCCL_DEBUG, nccl-tests) specifically designed to isolate communication and network-level issues in distributed training.

234
MCQeasy

Which utility is primarily used to monitor and manage NVIDIA GPU power, temperature, and usage statistics in real-time on a Linux-based deployment?

A.top
B.nvidia-smi
C.htop
D.dmesg
AnswerB

This is the official utility for managing NVIDIA GPU devices. It provides comprehensive real-time data on GPU usage, memory allocation, power consumption, and thermal status. It also allows administrators to set performance limits and reset devices, which is essential for maintaining a healthy and performant AI cluster.

Why this answer

nvidia-smi (NVIDIA System Management Interface) is the command-line utility used to interact with the NVIDIA driver for monitoring and managing GPU devices. It is the fundamental tool for AI ops teams to verify that GPUs are properly detected and performing within expected thermal and power envelopes during model training or inference tasks on DGX or workstation systems.

Exam trap

Candidates frequently confuse nvidia-smi with higher-level management tools or cloud-native exporters like DCGM, overlooking its role as the foundational host-level CLI utility.

235
MCQmedium

Refer to the exhibit. An administrator attempts to deploy a GPU-based pod, but it remains in the 'Pending' state. What is the most likely cause based on the error log?

A.The pod is requesting more CPU cores than the node can provide.
B.The NVIDIA device plugin is not correctly advertising GPU resources.
C.The node is in a 'NotReady' state due to high disk latency.
D.The pod security policy prohibits the use of GPU-accelerated containers.
AnswerB

The device plugin is the component responsible for telling the Kubernetes scheduler how many GPUs are available on a specific node. If this process fails or is not running, the scheduler will not see any allocatable GPU resources, resulting in the 'Insufficient' error even if physical GPUs are present.

Why this answer

The error 'Insufficient nvidia.com/gpu' indicates that the Kubernetes scheduler cannot find a node with available GPU resources that match the pod's request. This typically happens when the device plugin is not correctly reporting resource availability to the scheduler, or the cluster is over-provisioned. Resolving this requires ensuring the GPU Operator is running and that the device plugin has successfully registered the GPUs with the Kubelet on the worker nodes.

Exam trap

Candidates often blame pod resource limits or application errors. They fail to recognize the 'Insufficient' resource error as a direct signal that the device plugin is not communicating with the Kubelet.

236
Multi-Selecthard

An AI operations engineer is optimizing a PyTorch training job on an NVIDIA DGX A100 system. The job uses a data loader with multiple workers, but the engineer observes that GPU utilization fluctuates between 40% and 60%, and `nvidia-smi dmon` shows periods of zero GPU utilization. The engineer suspects that the data input pipeline is the bottleneck. Which TWO actions should the engineer take to improve GPU utilization? (Choose two.)

Select 2 answers
A.Increase the number of DataLoader workers to better overlap data loading with GPU computation.
B.Decrease the batch size to reduce memory usage and allow more frequent updates.
C.Set the CUDA_LAUNCH_BLOCKING environment variable to 1 to ensure synchronous kernel execution.
D.Enable mixed precision training to reduce memory footprint and increase throughput.
E.Enable pinned memory in the DataLoader to speed up host-to-device transfers.
AnswersA, E

Increasing DataLoader workers allows more parallel data preprocessing, reducing the time the GPU waits for data. This directly addresses the bottleneck by improving overlap between CPU data loading and GPU computation, leading to higher and more consistent GPU utilization. It is a standard optimization for input-bound training jobs.

Why this answer

The observed GPU utilization fluctuations and zero-utilization periods indicate the GPU is often waiting for data. Increasing DataLoader workers parallelizes data preprocessing, and enabling pinned memory accelerates host-to-device transfers. Together, these reduce data loading latency and improve overlap with GPU computation, raising utilization.

Other options either do not address the bottleneck or harm performance.

Exam trap

The trap here is focusing on GPU-side optimizations like mixed precision when the bottleneck is actually the CPU data pipeline, which requires adjusting DataLoader parameters.

237
Multi-Selectmedium

A team trains a model inside an NGC PyTorch container on a DGX H100 node. Training starts, but after a few minutes the process dies and `dmesg` shows `Xid 79: GPU has fallen off the bus` on one GPU. The team needs to determine whether the fault is hardware or software before opening an RMA. Which two actions should they take to gather useful evidence? (Choose two.)

Select 2 answers
A.Immediately reboot the node and re-run the training job to see whether it fails again.
B.Run `nvidia-smi -q` and `nvidia-smi -q -d ECC,PAGE_RETIREMENT,ROW_REMAPPER` and capture the output.
C.Reset the GPU with `nvidia-smi --gpu-reset` and continue using the node if training succeeds afterward.
D.Collect `nvidia-bug-report.sh` output and the relevant `dmesg`/`journalctl -k` window around the failure.
E.Delete and recreate the NGC container, then pull the image again from the registry.
AnswersB, D

These queries report ECC errors, pending row remaps, and page retirement state for each GPU. On an H100, Xid 79 is frequently linked to a row-remapping or ECC escalation event, so capturing these counters before any reset preserves the evidence needed to distinguish a hardware fault from a transient software issue. The output is also requested by NVIDIA support when validating an RMA case.

Why this answer

Xid 79 indicates the GPU dropped off the PCIe bus, which can stem from a hardware fault, a degraded link, or a driver escalation after ECC or row-remap events. Before resetting anything, engineers should query ECC and row-remapper state and bundle the full bug report with the matching kernel log. These two sources together show whether the GPU had pending remaps or link errors, enabling an accurate hardware-versus-software determination and a defensible RMA decision.

Exam trap

The trap here is assuming that rebooting or resetting the GPU is a safe first step, when it actually erases the volatile counters and logs that distinguish a failing GPU from a software or fabric issue.

238
MCQeasy

An administrator is preparing to install the NVIDIA GPU Operator on a new Kubernetes cluster. The cluster uses containerd as the container runtime. Which prerequisite must be satisfied on each GPU node before the operator can successfully deploy the driver container?

A.The Linux kernel headers for the running kernel must be available on the node.
B.The node must have the `nvidia.com/gpu.present` label applied manually by the administrator.
C.The NVIDIA Container Toolkit must be installed manually on each node.
D.The node must have the NVIDIA driver preinstalled at the exact version specified by the operator.
AnswerA

The driver container compiles the NVIDIA kernel modules against the running kernel, so the matching kernel headers and build tools must be present. Without them, the driver build fails and the GPU cannot be initialized. This is a fundamental prerequisite for any node where the operator will install the driver via container.

Why this answer

The driver container builds NVIDIA kernel modules on the host, which requires kernel headers and a compiler toolchain matching the running kernel. Without these, the container cannot compile the modules and the driver fails to load. This prerequisite is essential for any node where the operator manages the driver, and it is often overlooked during cluster preparation.

Exam trap

The trap here is assuming the NVIDIA Container Toolkit or driver must be preinstalled, when the operator handles those components and instead requires kernel headers for module compilation.

239
MCQhard

A team is running a multi-GPU training job on an NVIDIA DGX A100 system using NCCL for inter-GPU communication. Training throughput is much lower than expected, and the NCCL logs show repeated 'NCCL WARN Call to ibv_reg_mr failed' errors. The job uses a container with host networking. Which action should the AI operations engineer take to resolve the issue?

A.Reduce the batch size per GPU to lower memory pressure and avoid the need for large memory registrations.
B.Increase the NCCL_IB_TIMEOUT environment variable to allow more time for memory registration.
C.Disable InfiniBand by setting NCCL_IB_DISABLE=1 to force NCCL to use TCP sockets.
D.Verify and increase the container's locked memory limit (ulimit -l) and ensure the NVIDIA driver and OFED stack are compatible.
AnswerD

The 'ibv_reg_mr failed' error typically occurs when the process cannot pin enough memory for RDMA, often because the locked memory limit is too low or the OFED stack is incompatible with the NVIDIA driver. In containers, the default ulimit -l may be insufficient. Raising the limit and validating driver/OFED compatibility directly addresses the registration failure, resolving the NCCL warnings and restoring throughput.

Why this answer

The NCCL warning 'ibv_reg_mr failed' points to an inability to register memory regions for InfiniBand, commonly caused by a low locked memory limit or mismatched OFED and NVIDIA drivers. In containerized environments, the default ulimit -l is often too low. Increasing the locked memory limit and ensuring driver compatibility directly fixes the registration failure, restoring proper NCCL communication and training throughput.

Exam trap

The trap here is treating the NCCL warning as a timeout or bandwidth issue and adjusting unrelated variables instead of addressing the memory registration failure.

240
MCQeasy

A data science team submits a PyTorch distributed training job to a Kubernetes cluster with the NVIDIA GPU Operator installed. The job's pods repeatedly fail with a CUDA initialization error, while a simple `nvidia-smi` check inside an interactive pod on the same node succeeds. The administrator confirms the node's driver is healthy and the device plugin is advertising GPUs. Which configuration should the administrator verify first?

A.That the node's `/etc/docker/daemon.json` still lists `nvidia` as the default runtime for all containers.
B.That the pods' securityContext drops the `IPC_LOCK` capability required for pinned host memory.
C.That the training pods' containers were built with a CUDA toolkit version newer than the node driver supports.
D.That each training pod includes a GPU resource request or limit so the device plugin injects the driver libraries and device nodes.
AnswerD

The GPU Operator advertises devices through the device plugin, and the kubelet only mounts the driver libraries, device nodes, and CUDA binaries into a container that requests nvidia.com/gpu. Without that request the container starts with no GPU access, so CUDA initialization fails while nvidia-smi in a GPU-requesting pod on the same node works, matching the observed behavior exactly.

Why this answer

In an operator-managed cluster, GPU access is granted during admission and kubelet setup only when a pod requests an extended GPU resource; the device plugin and admission controller then inject the driver libraries, CUDA binaries, and device nodes. A pod without that request runs with no visible device, which is precisely why CUDA initialization fails while nvidia-smi succeeds in a different pod on the same node. The other options concern version skew, a specific capability, or a legacy runtime setting that would not produce this selective symptom.

Exam trap

The trap here is chasing driver or toolkit version mismatches when the real cause is that the workload never requested a GPU, so the device was never injected into the container.

241
MCQhard

A research organization runs an NVIDIA DGX SuperPOD with a Kubernetes cluster managed by the NVIDIA GPU Operator and Network Operator. A distributed training job using PyTorch DDP across 32 nodes stalls at initialization, and the administrator suspects the collective communication library is not selecting the high-speed fabric. Which configuration should the administrator verify first to ensure NCCL uses the correct network interface and topology?

A.Increase the pod's nvidia.com/gpu limit to 8 so each node exposes all GPUs to the training process.
B.Set the CUDA_VISIBLE_DEVICES variable to list all GPUs and restart the training job.
C.Enable the NVIDIA MIG feature on all nodes so each rank gets an isolated GPU slice for communication.
D.Confirm that NCCL_IB_DISABLE is set to 0, NCCL_SOCKET_IFNAME matches the high-speed fabric interface, and NCCL_TOPO_FILE or the topology-aware plugin is loaded on each node.
AnswerD

NCCL relies on these environment variables and topology data to select InfiniBand or RoCE interfaces and to build the correct ring or tree topology across nodes. If NCCL_IB_DISABLE is set to 1, or NCCL_SOCKET_IFNAME points at the management interface, NCCL falls back to slower paths and initialization can stall. Verifying these values is the primary diagnostic step.

Why this answer

NCCL chooses transports and interfaces based on environment variables and detected topology. When NCCL_IB_DISABLE is set incorrectly or NCCL_SOCKET_IFNAME points to the wrong interface, collectives fall back to TCP over the management network or fail to connect, causing distributed jobs to hang at initialization. Verifying these variables and the topology file on every node is the correct first diagnostic step.

Exam trap

The trap here is assuming that adding GPUs per pod or adjusting CUDA_VISIBLE_DEVICES will fix a distributed hang, when the real cause is NCCL's network interface and fabric selection.

242
MCQmedium

Which administrative practice ensures that a cluster is prepared for the arrival of new NVIDIA GPU hardware with minimal downtime?

A.Wait for the hardware to arrive before checking driver support.
B.Update the OS and NVIDIA drivers across the cluster.
C.Manually install every CUDA library on all nodes.
D.Modify the BIOS settings to ignore PCIe errors.
AnswerB

Proactively updating the OS and drivers to versions that support the upcoming GPU hardware ensures that when the physical installation occurs, the software stack is ready. This minimizes the time spent troubleshooting driver incompatibilities and allows for a smooth, plug-and-play integration of the new hardware into the production cluster.

Why this answer

Maintaining an up-to-date driver and software stack via a robust package management system is the key to minimizing downtime during hardware upgrades. Administrators must ensure that the kernel and driver versions are compatible with the new hardware before it arrives. This proactive preparation is vital for maximizing cluster uptime and ensuring that researchers can immediately begin using the new hardware for their AI experiments without configuration delays.

Exam trap

Candidates often assume physical hardware installation alone or container runtime updates are sufficient, overlooking the critical requirement to synchronize the underlying host OS kernel and NVIDIA driver stack.

243
Multi-Selectmedium

An administrator is preparing to deploy an NVIDIA AI Enterprise solution on an OpenShift cluster. Which TWO steps must be completed to ensure the NVIDIA drivers are loaded correctly on the worker nodes?

Select 2 answers
A.Deploy the NVIDIA GPU Operator via the OperatorHub.
B.Manually install NVIDIA drivers on every host OS.
C.Configure the Node Feature Discovery (NFD) operator.
D.Install the NVIDIA Triton Inference Server.
E.Set the Kubernetes scheduler to 'AlwaysPull'.
AnswersA, C

The OperatorHub provides the standard interface for managing lifecycle updates and dependencies within OpenShift. Deploying the GPU Operator from here automatically pulls in required dependencies like NFD, ensuring that the environment is correctly primed to handle the proprietary NVIDIA kernel modules and device plugin communication protocols.

Why this answer

On OpenShift, the NVIDIA GPU Operator leverages the Node Feature Discovery (NFD) and the Machine Config Operator to manage kernel modules. Ensuring these are configured correctly allows the driver container to load the proprietary NVIDIA kernel modules successfully. Proper node labeling is also required so the operator can target the correct hardware, ensuring that the driver is injected into the node's operating system environment during the boot process.

Exam trap

Candidates often overlook Node Feature Discovery (NFD), mistakenly assuming the GPU Operator alone can automatically detect hardware features on OpenShift worker nodes without NFD labeling.

244
Multi-Selectmedium

An administrator is responsible for maintaining an NVIDIA AI Enterprise cluster and needs to ensure high availability of GPU resources for critical inference workloads. Which two practices should the administrator implement? (Choose two.)

Select 2 answers
A.Enable GPU time-slicing to allow multiple pods to share a single GPU
B.Schedule all inference pods on a single node with multiple GPUs to reduce network latency
C.Use NVIDIA MIG to partition a GPU into isolated instances for each workload
D.Implement pod disruption budgets to maintain a minimum number of available replicas
E.Configure node affinity and anti-affinity rules to spread pods across multiple nodes
AnswersD, E

Pod disruption budgets (PDBs) ensure that a specified minimum number of pods remain available during voluntary disruptions, such as node drains or upgrades. This helps maintain service availability for critical inference workloads. By defining a PDB, the administrator can prevent simultaneous eviction of too many replicas, thus supporting high availability.

Why this answer

High availability for inference workloads requires spreading pods across multiple nodes to avoid single points of failure and using pod disruption budgets to maintain minimum replica counts during disruptions. Node anti-affinity ensures distribution, while PDBs protect against voluntary evictions. Time-slicing, MIG, and single-node concentration do not provide fault tolerance against node or GPU failures.

Exam trap

The trap here is equating resource sharing or isolation features like MIG or time-slicing with high availability, when they do not protect against hardware failure.

245
MCQhard

Which THREE factors should be considered when estimating GPU memory requirements for a Large Language Model (LLM) fine-tuning job?

A.The memory size of the model parameters.
B.The precision used for optimizer states.
C.The activation memory during the forward pass.
D.The total number of CPU threads used.
E.The network latency between nodes.
AnswerA, B, C

Model weights occupy a significant portion of GPU memory. For large models, this is often the baseline requirement. When determining the cluster footprint, the total size of these parameters must be considered, especially if using techniques like sharding, which distribute these weights across multiple GPUs in a cluster.

Why this answer

When sizing LLM workloads, you must account for the model weights, optimizer states, and gradient buffers. Additionally, activations consume significant memory during the forward and backward passes. Understanding these components is essential for AI Ops, as incorrect sizing leads to OOM crashes early in the training process, wasting significant compute time and delaying model delivery in production environments.

Exam trap

Candidates frequently focus only on model weights, ignoring the significant memory overhead consumed by activation buffers during the forward pass, which often causes OOM errors in large training jobs.

246
MCQmedium

Refer to the exhibit. During a multi-node training job, communication between nodes fails. What is the most likely cause of this error?

A.The GPU driver version is incompatible with the installed NCCL library.
B.The NCCL_SOCKET_IFNAME environment variable is misconfigured for the cluster network.
C.The model weights are too large to be synchronized across the network.
D.The CUDA visibility is restricted to only one GPU per node.
AnswerB

NCCL uses the NCCL_SOCKET_IFNAME variable to identify the specific network interface for inter-node communication. If this variable is unset or points to an incorrect interface, nodes cannot discover each other, leading to connection refused errors. Correctly specifying the high-speed interconnect is essential for successful multi-node operations.

Why this answer

The error indicates a network connectivity issue during the NCCL initialization phase, specifically failing to route traffic between nodes. NCCL relies on correct interface configuration for inter-node communication. Misconfigured network interfaces or firewall rules between nodes are common failure points in distributed training.

Troubleshooting these network layer issues is vital for ensuring high-performance communication across GPU clusters during large-scale model training.

Exam trap

Test-takers commonly blame the deep learning framework or code implementation rather than checking low-level cluster networking environment variables required for multi-node NCCL communication.

247
MCQmedium

Refer to the exhibit. An AI administrator investigates why a GPU node is performing significantly slower than expected. Based on the output, what is the most likely cause?

A.The GPU driver is outdated and needs a patch.
B.The server's cooling system is failing or airflow is restricted.
C.The power supply unit is malfunctioning.
D.The workload exceeds the GPU memory capacity.
AnswerB

The 'HW Thermal Slowdown' status directly confirms that the GPU has reached an internal temperature threshold and is actively reducing performance to lower heat generation. This indicates a physical cooling deficiency, likely caused by obstructed air intakes, failing server fans, or an inadequate data center ambient temperature environment.

Why this answer

The exhibit shows 'HW Thermal Slowdown' and 'HW Slowdown' are active. This indicates the GPU hardware is actively reducing its clock frequency to prevent damage due to excessive heat. This is a critical performance issue that requires immediate attention to the data center cooling or physical airflow within the server chassis.

Ensuring adequate thermal management is fundamental to maintaining consistent compute performance during heavy training loads.

Exam trap

Candidates often blame software bottlenecks or outdated drivers when unexpected performance drops occur, failing to check hardware thermal and clock throttling status.

248
MCQmedium

An AI operations engineer manages a shared Kubernetes cluster running NVIDIA GPU Operator. Several teams report that their inference pods remain in a Pending state with the event '0/8 nodes are available: 8 Insufficient nvidia.com/gpu.' The administrator verifies that nvidia-smi on all nodes shows idle GPUs. Which action should the administrator take first to resolve the scheduling failure?

A.Reinstall the NVIDIA GPU Operator and restart the kubelet on all worker nodes.
B.Inspect the node allocatable GPU count and any taints or labels that prevent scheduling on the idle nodes.
C.Add a nodeSelector for kubernetes.io/os=linux to the pod template.
D.Increase the GPU memory limit in the pod specification so the scheduler can fit the workload.
AnswerB

The event 'Insufficient nvidia.com/gpu' means the scheduler sees zero or fewer allocatable GPUs than requested on every candidate node. This can result from missing device plugin advertisements, taints, node selectors, or nodes excluded by affinity. Checking allocatable resources and taints directly identifies why idle hardware is invisible or ineligible to the scheduler.

Why this answer

The Pending event shows the scheduler believes no node has a free nvidia.com/gpu resource, even though nvidia-smi reports idle devices. That gap is almost always caused by the device plugin not advertising GPUs, or by taints, labels, or affinity rules excluding the nodes. Inspecting allocatable counts and scheduling constraints identifies the actual blocker before any disruptive remediation.

Exam trap

The trap here is assuming that idle GPUs visible in nvidia-smi automatically become schedulable Kubernetes resources without the device plugin advertising them.

249
MCQeasy

An AI operations engineer is preparing a Kubernetes cluster to run GPU-accelerated inference workloads using the NVIDIA GPU Operator. The cluster nodes already have NVIDIA data center GPUs installed, and the engineer wants to avoid installing the driver manually on each node. Which component of the GPU Operator is responsible for automatically deploying the NVIDIA driver on worker nodes?

A.NVIDIA Device Plugin
B.NVIDIA GPU Operator's driver container
C.NVIDIA DCGM Exporter
D.NVIDIA Container Toolkit
AnswerB

The GPU Operator includes a driver container that runs as a DaemonSet on GPU nodes. It compiles and loads the NVIDIA kernel driver inside a container, eliminating the need for manual host driver installation. This is the intended mechanism for automated driver deployment in Kubernetes environments, ensuring consistent driver versions across nodes without direct host modifications.

Why this answer

The driver container within the GPU Operator automates the deployment and management of NVIDIA drivers on Kubernetes nodes. It runs as a DaemonSet and ensures the correct driver version is loaded without manual intervention. This simplifies operations and maintains consistency across the cluster, which is essential for AI workloads that depend on specific driver capabilities.

Exam trap

The trap here is confusing the NVIDIA Container Toolkit with the driver container; the toolkit exposes GPUs to containers but does not install the kernel driver.

250
MCQmedium

An administrator manages an NVIDIA AI Enterprise cluster using NVIDIA Run:ai. A data science team complains that their submitted training job has been stuck in a Pending state for over an hour, even though the Run:ai scheduler shows free GPUs in the cluster. The administrator verifies that the job requests 2 GPUs and the node pool has 4 idle GPUs. Which Run:ai administrative configuration is the most likely cause of the job remaining Pending?

A.The job's container image does not include the NVIDIA Container Toolkit, so the scheduler cannot start the job.
B.The project's GPU quota is exhausted by other running workloads, so the scheduler cannot allocate the requested GPUs.
C.The NVIDIA GPU Operator is not installed on the cluster, so the scheduler cannot see GPU resources.
D.The node pool is configured with a node affinity that excludes the nodes with idle GPUs, so the scheduler cannot place the job.
AnswerB

Run:ai projects enforce a GPU quota per project. Even if the cluster has idle GPUs, a job in a project that has already consumed its quota will remain Pending until quota is freed. This matches the scenario where the cluster shows free GPUs but the job cannot start because the project-level quota is the limiting factor.

Why this answer

Run:ai uses projects to enforce GPU quotas. When a project's quota is fully consumed, new jobs remain Pending even if the cluster has idle GPUs. The administrator should check the project's quota usage and either increase the quota or free resources.

Node affinity, missing GPU Operator, or missing Container Toolkit would produce different symptoms.

Exam trap

The trap here is assuming that free GPUs in the cluster automatically mean a job can be scheduled, ignoring project-level quota enforcement in Run:ai.

251
MCQmedium

A platform engineer is preparing a Kubernetes cluster to run AI training jobs that require GPU access. The cluster nodes have NVIDIA GPUs, and the engineer wants the GPU Operator to manage the driver lifecycle. Which component must be installed on the host nodes to allow the GPU Operator to load kernel modules and create device nodes?

A.NVIDIA Kubernetes Device Plugin
B.NVIDIA Container Toolkit
C.NVIDIA GPU Operator Driver Container
D.NVIDIA GPU Driver
AnswerC

The GPU Operator deploys a driver container that runs on each node to install the NVIDIA driver, load kernel modules, and create device nodes. This container manages the driver lifecycle, ensuring the correct version is used. It is the component that enables the GPU Operator to handle driver installation and updates automatically, which is essential for this scenario.

Why this answer

The GPU Operator uses a driver container to manage the NVIDIA driver on host nodes. This container loads kernel modules, creates device nodes, and ensures the driver version matches the operator's requirements. The NVIDIA Container Toolkit and Device Plugin are also deployed by the operator but serve different purposes: the toolkit enables container GPU access, and the device plugin advertises GPUs to Kubernetes.

The driver container is specifically responsible for the driver lifecycle, making it the correct component.

Exam trap

The trap here is confusing the driver container with the NVIDIA Container Toolkit or Device Plugin, assuming that any GPU-related component can manage the driver.

252
MCQmedium

An administrator is planning to monitor GPU utilization across a large cluster. Which component should be deployed to collect metrics that are compatible with Prometheus?

A.NVIDIA GPU Operator
B.NVIDIA DCGM Exporter
C.NVIDIA Triton Inference Server
D.NVIDIA CUDA Toolkit
AnswerB

The DCGM Exporter is purpose-built to extract hardware metrics and expose them in a format that Prometheus can scrape. This is the correct choice for integrating GPU telemetry into existing monitoring pipelines, enabling administrators to visualize performance data and effectively manage GPU resources within their Kubernetes infrastructure clusters.

Why this answer

NVIDIA DCGM (Data Center GPU Manager) is the industry-standard tool for collecting GPU health and telemetry data. By deploying the DCGM Exporter, the administrator enables the conversion of hardware telemetry into a Prometheus-friendly format. This integration is vital for observability, allowing teams to set up alerts and dashboards to track GPU usage, power consumption, and memory allocation across the entire fleet of accelerated computing nodes.

Exam trap

Candidates frequently confuse general Kubernetes metrics servers with GPU-specific telemetry tools, choosing standard kube-state-metrics instead of the dedicated NVIDIA DCGM Exporter.

253
MCQmedium

What is the primary benefit of using an NVIDIA-Certified System for AI Enterprise deployments?

A.Guaranteed 100% network uptime
B.Validated compatibility and performance
C.Automatic software license renewal
D.Free access to all AI models
AnswerB

NVIDIA-Certified Systems are tested against a strict validation suite to ensure that hardware, firmware, and software drivers work seamlessly together. This reduces the risk of deployment failure and performance degradation, providing a predictable and supported environment for AI Enterprise software workloads across diverse enterprise data center environments.

Why this answer

NVIDIA-Certified Systems have passed rigorous testing to ensure optimal performance and compatibility with the NVIDIA AI Enterprise software stack. This certification minimizes the risk of hardware-software conflicts, performance bottlenecks, and stability issues during production. For AI Ops professionals, this translates to faster deployment times, lower maintenance costs, and high confidence in the reliability of the underlying infrastructure for mission-critical deep learning and inference workloads.

Exam trap

Candidates often confuse 'NVIDIA-Certified' with 'any GPU-enabled server.' They incorrectly assume any hardware with an NVIDIA GPU provides the same level of validated performance, software stack compatibility, and production-grade support.

254
MCQhard

When deploying NVIDIA AI Enterprise, why is the selection of the correct CUDA version in the container image critical during the installation phase?

A.It determines the speed of the GPU's memory bus.
B.The CUDA version must be compatible with the host driver version.
C.It enables the use of the NVIDIA License System.
D.It is required for the installation of the GPU Operator.
AnswerB

NVIDIA drivers follow a backward-compatibility model where the driver must support the CUDA version used by the application. Using a container with a newer CUDA version than the driver supports will cause the application to fail to initialize, as it cannot properly map the required kernel functions.

Why this answer

The CUDA version dictates which APIs and features are available to the AI application. Because the driver on the host must support the CUDA version used by the container (backward compatibility), mismatching these leads to runtime failures. This is a crucial AI Ops consideration as it directly affects the stability of the entire stack, ensuring that the software environment aligns with the underlying hardware capabilities for maximum performance and reliability.

Exam trap

Candidates often assume that the container image includes its own driver, failing to realize that the host driver must be compatible with the CUDA version installed inside the container.

255
Multi-Selectmedium

Which TWO of the following are benefits of using containerized GPU workloads compared to bare-metal deployment?

Select 2 answers
A.Easier management of heterogeneous software dependencies.
B.Significant increase in raw GPU compute performance.
C.Improved portability across development and production environments.
D.Automatic elimination of GPU driver compatibility issues.
E.Direct access to the GPU firmware for kernel customization.
AnswersA, C

Containers allow each workload to bundle its own specific versions of libraries like CUDA and PyTorch. This avoids conflicts on the host system, where different training jobs might otherwise require incompatible driver versions or shared library dependencies, enabling more efficient sharing of the same underlying physical node.

Why this answer

Containerization provides portability, version control, and dependency isolation, which are essential for managing complex AI stacks. By packaging the environment, developers ensure that models run consistently across development, testing, and production clusters. This significantly reduces 'environment drift' and simplifies the management of different CUDA, cuDNN, and framework versions required by various research projects within the same shared infrastructure cluster.

Exam trap

Candidates often select 'increased performance' as a benefit. Containerization adds a slight abstraction layer and does not inherently increase raw GPU compute performance compared to bare-metal execution.

256
MCQhard

An inference service runs a 70B parameter model with TensorRT-LLM on a single H100 using in-flight batching. Operators report that time-to-first-token is acceptable, but inter-token latency degrades sharply once concurrent request count exceeds a certain point, and GPU memory utilization sits near 98 percent. Which change most directly addresses the inter-token latency degradation?

A.Disable in-flight batching so each request is processed to completion before the next begins.
B.Increase the maximum batch size so more requests are processed per decode step.
C.Enable FP8 quantization for the KV cache and reduce the maximum batch size in the TensorRT-LLM build configuration.
D.Switch the service to a round-robin load balancer across two replicas of the same model on one GPU.
AnswerC

Near-saturated memory with growing concurrency means the KV cache is crowding out the workspace and forcing the scheduler to admit requests it cannot serve efficiently. FP8 KV cache roughly halves cache footprint, and lowering the maximum batch size keeps the runtime from over-admitting sequences. Together they reduce per-step memory pressure and shorten decode iterations, directly improving inter-token latency under high concurrency.

Why this answer

When GPU memory is nearly exhausted, the runtime cannot hold enough KV cache plus workspace to serve all admitted sequences efficiently, so each decode step stretches and inter-token latency rises with concurrency. Shrinking the KV cache with FP8 quantization and lowering the maximum batch size reduces per-step memory and compute demand, restoring shorter decode iterations. These changes target the resource constraint that actually causes the latency curve to bend upward.

Exam trap

The trap here is treating higher concurrency as a throughput problem to solve by enlarging the batch, when the observed memory saturation means the batch ceiling is already too high for the available KV cache and workspace.

257
MCQhard

An AI operations team runs a shared Kubernetes cluster with the NVIDIA GPU Operator and several namespaces owned by different groups. A group reports that its training pods are stuck Pending with an event indicating insufficient nvidia.com/gpu, yet cluster-wide dashboards show many GPUs idle. Investigation reveals the idle GPUs belong to nodes labeled for another group, and the affected namespace has a node affinity rule pinning its pods to a specific GPU generation that is fully consumed. Which action best resolves the Pending pods while respecting multi-tenant boundaries?

A.Disable the device plugin on the other tenants' nodes so their GPUs become available to the affected namespace.
B.Lower the GPU request in the pod spec so the scheduler accepts a partial device from an idle node.
C.Remove the node affinity rule so the pods can schedule onto any idle GPU node in the cluster.
D.Provision additional GPUs of the required generation for that tenant or rebalance existing capacity to that node pool.
AnswerD

The pods are Pending because the specific GPU generation they require is exhausted, not because the cluster lacks GPUs entirely. Adding capacity to that node pool, or moving idle cards of the right generation into it, satisfies the affinity constraint while keeping tenants separated. This addresses the real bottleneck without weakening isolation rules.

Why this answer

The pods are constrained by a node affinity rule targeting a specific GPU generation whose nodes are fully allocated. Idle GPUs on other nodes cannot satisfy that rule. The correct fix is to add capacity of the required generation or rebalance matching cards into that pool, which resolves the shortage while preserving the tenant isolation the affinity rule enforces.

Exam trap

The trap here is treating idle cluster-wide GPUs as available capacity, when node affinity can make those GPUs ineligible for the pending pods.

258
MCQeasy

An administrator is responsible for monitoring a large-scale AI cluster with hundreds of NVIDIA GPUs. They need to collect telemetry data such as GPU utilization, temperature, and power consumption from all nodes and store it centrally for analysis and alerting. Which NVIDIA tool should they use to collect and export GPU metrics to a monitoring system like Prometheus?

A.NVIDIA Nsight Systems
B.NVIDIA DCGM Exporter
C.nvidia-smi
D.NVIDIA Base Command Manager
AnswerB

NVIDIA DCGM Exporter is a Prometheus exporter that collects GPU metrics using NVIDIA Data Center GPU Manager (DCGM) and exposes them in a format that Prometheus can scrape. It provides a wide range of metrics, including utilization, temperature, power, and memory usage, making it ideal for centralized monitoring of large GPU clusters.

Why this answer

For centralized monitoring of GPU metrics in a large cluster, NVIDIA DCGM Exporter is the appropriate tool. It leverages NVIDIA Data Center GPU Manager (DCGM) to collect a comprehensive set of metrics from each GPU and exposes them via an HTTP endpoint that Prometheus can scrape. This allows administrators to aggregate metrics across all nodes, create dashboards in Grafana, and set up alerts based on thresholds.

Exam trap

The trap here is confusing nvidia-smi with a scalable monitoring solution, when it is only a per-node command-line tool without native Prometheus integration.

259
MCQhard

An AI operations engineer is deploying a multi-node Kubernetes cluster with NVIDIA A100 GPUs for distributed training. The engineer wants to ensure that GPUs are correctly discovered and that workloads can request GPU resources. After installing the NVIDIA GPU Operator, the engineer notices that the GPU nodes are not advertising any 'nvidia.com/gpu' resources. Which component should the engineer verify first to resolve this issue?

A.NVIDIA DCGM Exporter
B.NVIDIA GPU Operator's driver container
C.NVIDIA Container Toolkit
D.NVIDIA Device Plugin
AnswerD

The NVIDIA Device Plugin is the component that discovers GPUs and advertises them as schedulable resources (e.g., nvidia.com/gpu) to the Kubernetes API server. If GPUs are not appearing as resources, the device plugin is the first component to check. It could be failing to start, unable to communicate with the kubelet, or missing driver dependencies.

Why this answer

When GPU resources are not advertised, the NVIDIA Device Plugin is the primary component to investigate. It runs as a DaemonSet on GPU nodes and registers GPUs with the kubelet. Issues such as pod failures, missing driver libraries, or misconfigured kubelet settings can prevent resource advertisement.

Checking its logs and status will reveal the root cause.

Exam trap

The trap here is focusing on the driver or toolkit first; however, the device plugin is the component that directly advertises GPU resources to Kubernetes.

260
MCQmedium

Which TWO of the following steps are essential when deploying the NVIDIA GPU Operator on a Kubernetes cluster to ensure that GPU resources are discoverable by the scheduler?

A.Enable the node feature discovery (NFD) service in the operator configuration.
B.Manually install the NVIDIA Container Toolkit on every worker node prior to operator deployment.
C.Configure the GPU device plugin to register NVIDIA-specific resources with the Kubelet.
D.Modify the Kubernetes API server manifest to include the NVIDIA-specific admission controller.
E.Disable the default Kubernetes scheduler to allow the NVIDIA scheduler plugin to take over.
AnswerA, C

NFD is essential for labeling nodes based on hardware features, such as GPU architecture. The GPU Operator uses these labels to identify nodes suitable for GPU-accelerated workloads. Without NFD, the cluster cannot dynamically detect and categorize the GPU hardware capabilities required to schedule pods accurately across the node pool.

Why this answer

The NVIDIA GPU Operator automates the lifecycle of NVIDIA software components, including the driver, toolkit, and device plugin. By configuring the operator to deploy the GPU device plugin and the Node Feature Discovery (NFD) service, the cluster gains the ability to identify GPU hardware and advertise it as an allocatable resource. Without these, the Kubernetes scheduler cannot place pods requiring GPU resources, leading to 'Pending' status for those workloads.

Exam trap

Candidates often assume that installing the GPU driver is sufficient. They miss that the Kubernetes scheduler requires explicit registration via NFD and the device plugin to actually 'see' the GPU as a resource.

261
MCQmedium

An administrator manages an NVIDIA AI Enterprise cluster and needs to enforce GPU resource quotas across multiple Kubernetes namespaces. Which NVIDIA component should be configured to enforce these quotas?

A.Kubernetes ResourceQuota with extended resources
B.NVIDIA DCGM Exporter
C.NVIDIA Base Command Manager
D.NVIDIA GPU Operator
AnswerA

Kubernetes ResourceQuota objects can limit the aggregate quantity of extended resources, such as nvidia.com/gpu, that can be requested within a namespace. By defining a ResourceQuota that specifies a maximum for nvidia.com/gpu, the administrator can enforce GPU quotas across namespaces. This is the native Kubernetes mechanism for quota enforcement and works in conjunction with the NVIDIA device plugin that advertises GPUs as extended resources.

Why this answer

Kubernetes ResourceQuota is the native mechanism to limit aggregate resource consumption per namespace, including extended resources like nvidia.com/gpu. When the NVIDIA device plugin advertises GPUs, administrators can define a ResourceQuota specifying a maximum for nvidia.com/gpu, thereby enforcing GPU quotas. Other NVIDIA tools focus on deployment, management, or monitoring and do not provide quota enforcement.

Exam trap

The trap here is assuming that NVIDIA GPU Operator or Base Command Manager enforces GPU quotas, when quota enforcement is actually a Kubernetes-native function.

262
MCQhard

An AI operations engineer is investigating intermittent failures in a long-running distributed training job. The job occasionally aborts with a collective timeout error, but no GPU errors, ECC events, or fabric link flaps appear in logs. Which action should the engineer take first to identify the root cause?

A.Restart the job with a reduced global batch size to decrease the time each collective step requires.
B.Enable per-rank NCCL debug logging and correlate the last completed collective across all ranks to find the straggler.
C.Lower the NCCL timeout value so failures surface faster and can be correlated with system events.
D.Switch the collective algorithm from ring to tree to reduce sensitivity to a single slow rank.
AnswerB

Collective timeouts occur when one or more ranks fail to arrive at a collective. Per-rank debug logs show the last operation each rank completed, so comparing them identifies the rank that stalled and the operation it was executing. This directly localizes the fault without disrupting the run, making it the correct first step.

Why this answer

A collective timeout with no hardware errors indicates one rank is arriving late or not at all. Enabling per-rank NCCL debug logging and comparing the last completed collective across ranks isolates the straggler and the operation it was executing, providing direct evidence of the fault without altering the job configuration or losing diagnostic information.

Exam trap

The trap here is treating a collective timeout as a network or GPU hardware fault, when the absence of ECC and link-flap events points instead to a single rank stalling in software or host resources.

263
MCQmedium

When managing GPU resources in a shared cluster, which configuration best prevents 'noisy neighbor' scenarios where one GPU task consumes all available memory bandwidth?

A.Increasing the CUDA_VISIBLE_DEVICES environment variable.
B.Implementing NVIDIA MIG and strictly defining resource limits.
C.Setting a higher priority class for inference pods.
D.Configuring the NVIDIA DCGM Exporter alerts.
AnswerB

MIG partitions the GPU at the hardware level, providing dedicated memory and compute paths. When combined with Kubernetes resource limits, it ensures that workloads are physically constrained to their slice, preventing a single process from monopolizing memory bandwidth or compute cycles, thus effectively eliminating the noisy neighbor problem.

Why this answer

Using a combination of NVIDIA MIG and Kubernetes resource quotas provides the strongest isolation. By enforcing hardware-level partitioning via MIG, you ensure that memory bandwidth is strictly bounded for each slice. This is vital for AI Ops because it prevents high-throughput training jobs from starving small inference requests, ensuring stable latency across the entire cluster environment and improving total system reliability.

Exam trap

Candidates often choose software-level container limits or basic Kubernetes namespaces alone, forgetting that true bandwidth isolation requires hardware-level partitioning like NVIDIA MIG to prevent high-throughput tasks from monopolizing shared memory paths.

264
MCQeasy

When managing large-scale model training jobs, what is the primary purpose of using a Job Scheduler like Slurm or Kubernetes Batch?

A.To increase the clock speed of individual GPU cores.
B.To automate resource allocation, job queuing, and throughput optimization.
C.To provide a direct IDE interface for writing model code.
D.To replace the need for containerization technologies.
AnswerB

Schedulers are essential for maximizing cluster utilization by managing queues and resource mapping. They automate the lifecycle of compute tasks, ensuring that jobs are placed on nodes with the required hardware specifications, thereby optimizing throughput and ensuring that expensive GPU resources are consistently kept productive.

Why this answer

Job schedulers serve as the orchestration layer to queue, prioritize, and allocate computational resources based on policy. In AI operations, they ensure that high-priority training runs receive the necessary GPU throughput while lower-priority jobs wait. By automating the allocation process, schedulers prevent resource idleness and manage contention, allowing data scientists to focus on model development rather than manual infrastructure management or resource conflict resolution during heavy cluster usage.

Exam trap

Candidates often focus on the model training code itself rather than the orchestration layer, failing to realize that job schedulers are essential for preventing resource contention in large clusters.

265
MCQmedium

An engineer is troubleshooting a CUDA program that terminates unexpectedly. Which tool should be used to detect memory leaks and race conditions in the CUDA kernel code?

A.nvidia-smi
B.NVIDIA Compute Sanitizer
C.nsys (Nsight Systems)
D.nvprof
AnswerB

The NVIDIA Compute Sanitizer is a functional correctness checking tool for CUDA kernels. It can detect memory access errors, race conditions, and various other issues that lead to unexpected program termination, making it the correct choice for debugging problematic kernel code during development or deployment.

Why this answer

Debugging parallel code is inherently difficult due to the non-deterministic nature of race conditions and memory access patterns. The NVIDIA Compute Sanitizer is the essential diagnostic tool for identifying these issues, such as out-of-bounds memory access or race conditions, before they lead to erratic, intermittent crashes in production. Using it early in the development lifecycle saves massive amounts of time by ensuring code correctness and robustness.

Exam trap

Test-takers often recommend standard CPU debuggers or basic logging flags, overlooking the specialized memory and race condition diagnostic requirements unique to parallel CUDA kernel code execution.

266
MCQeasy

Which NVIDIA technology allows for partitioning a single physical GPU into multiple independent instances, each with dedicated compute and memory resources for smaller workloads?

A.CUDA Streams
B.Multi-Instance GPU (MIG)
C.NVIDIA Collective Communications Library (NCCL)
D.NVIDIA Container Runtime
AnswerB

MIG enables the hardware-level partitioning of GPUs. By creating distinct instances with dedicated compute, memory, and cache, it provides fault isolation and guaranteed performance, which is vital for modern multi-tenant AI clusters that need to support varying workload sizes efficiently without resource contention or cross-tenant interference.

Why this answer

Multi-Instance GPU (MIG) is a feature in NVIDIA A100 and H100 architectures that enables the partitioning of a physical GPU into up to seven separate instances. This is essential for AI Ops to maximize resource utilization by running multiple inference tasks on a single GPU without interference, ensuring Quality of Service for each instance while preventing a single process from consuming all hardware resources.

Exam trap

Candidates often confuse MIG with general virtualization or container-based GPU sharing, failing to realize MIG is a hardware-level partitioning technology specific to Ampere and newer NVIDIA architectures for isolated, secure workloads.

267
MCQmedium

An operations team runs a multi-node NCCL all-reduce training job across four DGX nodes connected via InfiniBand. Training throughput is far below the expected linear scaling, and `nvidia-smi` shows GPU utilization oscillating between 20% and 40%. The network fabric is healthy and the GPUs are not thermally throttled. Which diagnostic step is MOST appropriate to identify the bottleneck?

A.Reinstall the CUDA toolkit on all nodes to ensure matching driver and runtime versions.
B.Increase the training batch size by 4x to raise arithmetic intensity and re-measure GPU utilization.
C.Enable ECC memory scrubbing on all GPUs and monitor for corrected errors during training.
D.Run `nccl-tests` (all_reduce_perf) with varying message sizes and collect NCCL debug logs by setting NCCL_DEBUG=INFO to inspect topology and algorithm selection.
AnswerD

Low, oscillating GPU utilization in a multi-node all-reduce job typically points to communication stalls. `nccl-tests` with varying message sizes reveals the achieved bus bandwidth per size, and NCCL_DEBUG=INFO exposes the detected topology, ring/tree algorithm choice, and channel count. This directly isolates whether the collective is the bottleneck and why, making it the correct diagnostic action for this scenario.

Why this answer

The oscillating, low GPU utilization across multiple nodes during an all-reduce strongly suggests the collective communication is stalling. Running `nccl-tests` with multiple message sizes measures achieved bandwidth and highlights which sizes are inefficient, while NCCL_DEBUG=INFO reveals the detected topology, algorithm, and channels. Together these pinpoint whether the bottleneck is the fabric, the ring/tree algorithm, or an unexpected topology detection.

Exam trap

The trap here is assuming low GPU utilization always means a compute problem, when in multi-node all-reduce jobs the bottleneck is often the collective communication path.

268
MCQmedium

Which security configuration is necessary when deploying NVIDIA GPUs in a multi-tenant environment to prevent unauthorized access between containers?

A.Increase the Kubernetes pod memory limit.
B.Enable MIG (Multi-Instance GPU).
C.Use a privileged container.
D.Install the latest version of CUDA.
AnswerB

MIG provides hardware-level partitioning, ensuring that individual GPU instances have isolated compute and memory. This is the recommended security practice for multi-tenancy, preventing cross-tenant resource leakage and providing strong isolation that software-based approaches cannot match, which is critical for compliance and security in shared infrastructure deployments.

Why this answer

In multi-tenant environments, using MIG (Multi-Instance GPU) is the standard method for hardware-level isolation. MIG allows a single physical GPU to be partitioned into multiple isolated instances, each with its own dedicated compute and memory. This ensures that one tenant's workload cannot interfere with or access the resources of another, providing a robust security boundary that software-only isolation cannot reliably guarantee in accelerated computing environments.

Exam trap

Candidates often suggest software-based container isolation (like namespaces or cgroups). While these provide process boundaries, they do not provide the hardware-level memory and compute isolation required for secure GPU multi-tenancy.

269
MCQmedium

What is the primary function of an 'InitContainer' in an NVIDIA GPU-enabled pod deployment?

A.To serve as a secondary GPU compute thread during training.
B.To ensure prerequisites are satisfied before the main application starts.
C.To permanently store the model weights after training.
D.To bypass GPU scheduling constraints and force-load the job.
AnswerB

InitContainers provide a controlled sequence of operations. By running before the main application container, they allow for critical setup tasks like verifying library installations, checking driver compatibility, or prepping datasets, which ensures that the training job starts in a valid, predictable state, minimizing late-stage failures.

Why this answer

InitContainers are often used to ensure that environment-specific requirements—such as configuring GPU drivers, validating libraries, or mounting persistent volumes—are met before the main training application starts. In AI workloads, this is critical for ensuring that dependencies are correctly loaded or data is staged, preventing runtime failures that would occur if the main training process attempted to execute in an incomplete environment. This pattern enhances the robustness of automated AI workflows.

Exam trap

Candidates mistakenly think InitContainers handle the main application training logic or continuously monitor runtime performance throughout the pod lifecycle.

270
MCQmedium

An administrator is preparing a bare-metal GPU server for AI workloads and needs to verify that the NVIDIA driver and CUDA toolkit are properly installed. The server has an NVIDIA A100 GPU. Which command should the administrator run to display the GPU model, driver version, and CUDA version?

A.nvidia-debugdump --list
B.nvcc --version
C.lspci | grep -i nvidia
D.nvidia-smi
AnswerD

The nvidia-smi command queries the NVIDIA driver and displays detailed information about all installed GPUs, including the GPU model, driver version, CUDA version, and current utilization. It is the standard tool for verifying driver and CUDA compatibility on a GPU-enabled system, making it the correct choice for this scenario.

Why this answer

The nvidia-smi utility is the primary interface for monitoring and managing NVIDIA GPUs. It provides a summary of each GPU's model, driver version, CUDA version, temperature, power usage, and memory utilization. For an administrator validating a new server, running nvidia-smi immediately confirms that the driver is loaded and the GPU is recognized, and it shows the CUDA version supported by the driver, which is essential for AI workload compatibility.

Exam trap

The trap here is assuming that nvcc --version reports the driver version, when it actually only reports the CUDA compiler version.

271
MCQmedium

An AI operations engineer is troubleshooting an inference service on an NVIDIA A100 GPU that shows intermittent stalls. The monitoring dashboard reports GPU utilization at 100%, but request throughput is far below the validated baseline. Running nvidia-smi dmon reveals the SM utilization is high while memory controller utilization is low. Which action should the engineer take first to identify the bottleneck?

A.Switch the inference precision from FP16 to FP32 to improve numerical stability.
B.Enable MPS (Multi-Process Service) to allow concurrent kernel execution from multiple processes.
C.Profile the inference process with Nsight Systems to inspect kernel execution and identify serialization gaps.
D.Increase the batch size in the inference server configuration to improve GPU occupancy.
AnswerC

Nsight Systems captures a timeline of CPU and GPU activity, revealing kernel serialization, gaps, and synchronization stalls that inflate utilization without producing throughput. Because memory controller utilization is low, the bottleneck is likely execution serialization or CPU-side latency, which Nsight Systems can pinpoint. This is the correct first step to diagnose the cause before making configuration changes.

Why this answer

The high SM utilization with low memory controller utilization suggests the GPU is busy but not doing useful work, often due to kernel serialization or CPU-GPU synchronization stalls. Profiling with Nsight Systems provides the timeline needed to see gaps and serialization. Only after identifying the specific stall should configuration changes be considered, making profiling the correct first step.

Exam trap

The trap here is assuming that 100% GPU utilization always means the GPU is efficiently processing work, when it can indicate stalls or serialization.

272
MCQmedium

What is the primary function of the NVIDIA Persistence Daemon in an AI deployment?

A.It automatically updates the NVIDIA driver when a new version is released on the web.
B.It keeps the GPU driver initialized to reduce latency during application startup.
C.It monitors GPU health and automatically initiates a reboot if an error is detected.
D.It manages the network traffic between the GPU and the storage backend.
AnswerB

By maintaining the driver's state in memory, the Persistence Daemon eliminates the overhead associated with the driver unloading and reloading process. This ensures that GPU resources are always ready for immediate use, which is essential for performance-sensitive AI applications that require rapid task execution.

Why this answer

The Persistence Daemon ensures that the NVIDIA driver remains loaded even when no applications are using the GPU. This prevents the driver from unloading and then reloading when a job starts, which significantly reduces the startup latency of AI models. It is a critical configuration for high-performance environments where frequent job scheduling would otherwise incur unnecessary overhead from repeated driver and device initialization.

Exam trap

Candidates often think the persistence daemon is used for saving machine learning model checkpoints, confusing storage persistence with driver state persistence.

273
MCQeasy

A financial services firm must prove to auditors that an AI training job ran on hardware located only in its Frankfurt data center and that no pod could ever be scheduled onto GPUs in other regions. The cluster spans three regions with nodes labeled topology.kubernetes.io/region. Which approach most directly enforces this placement requirement?

A.Define a nodeAffinity rule in the pod spec requiring topology.kubernetes.io/region to equal eu-central-1, and mark it requiredDuringSchedulingIgnoredDuringExecution.
B.Add a preferredDuringSchedulingIgnoredDuringExecution nodeAffinity rule that favors the Frankfurt region with a high weight.
C.Set the pod's restartPolicy to Never and add a toleration for the region-specific taint on Frankfurt nodes.
D.Create a PodDisruptionBudget for the training job that limits voluntary evictions to zero during the run.
AnswerA

A required node affinity rule is a hard constraint: the kube-scheduler will only place the pod on nodes whose labels match, and it leaves the pod Pending rather than scheduling it elsewhere. Pinning the region label guarantees the training job runs exclusively in the Frankfurt nodes, which is the enforceable evidence auditors need.

Why this answer

Hard placement guarantees come from required node affinity, which makes matching labels a precondition for scheduling. Because the scheduler will not place the pod on any node whose region label differs, the job cannot leave the Frankfurt data center, giving the firm an enforceable, auditable control over GPU location.

Exam trap

The trap here is confusing soft preferred affinity, which the scheduler may ignore under pressure, with required affinity, which is a hard constraint that keeps the pod Pending instead of violating the rule.

274
MCQmedium

An AI operations team is deploying NVIDIA AI Enterprise on a Kubernetes cluster using the NVIDIA GPU Operator. They need to ensure that GPU metrics such as utilization, memory usage, and temperature are collected and exposed to Prometheus for monitoring. Which component of the GPU Operator is responsible for this?

A.NVIDIA GPU Operator Validator
B.NVIDIA GPU Device Plugin
C.NVIDIA Container Toolkit
D.NVIDIA DCGM Exporter
AnswerD

The NVIDIA DCGM Exporter is a component deployed by the GPU Operator that collects GPU telemetry using NVIDIA Data Center GPU Manager (DCGM) and exposes it as Prometheus metrics. It provides metrics on utilization, memory, temperature, power, and more. This is the correct component for integrating GPU monitoring with Prometheus in a Kubernetes environment managed by the GPU Operator.

Why this answer

The NVIDIA DCGM Exporter is the component of the GPU Operator that gathers GPU metrics via DCGM and exposes them in Prometheus format. It is specifically designed for monitoring GPU health and performance in Kubernetes. Other components like the Container Toolkit, Device Plugin, and Validator serve different purposes and do not provide metrics collection for Prometheus.

Exam trap

The trap here is confusing the device plugin, which handles scheduling, with the DCGM Exporter, which handles monitoring metrics.

275
MCQmedium

An administrator needs to ensure that all GPU drivers are updated across a heterogeneous cluster without causing downtime. What is the best strategy?

A.Simultaneously update all nodes during a maintenance window
B.Perform rolling updates by draining nodes one by one
C.Use a containerized driver approach to avoid host updates
D.Only update the driver on the master node
AnswerB

Rolling updates ensure that the cluster remains operational throughout the entire upgrade process. By draining one node at a time, the administrator can safely update drivers without terminating active workloads prematurely. This approach provides a clear path for verification and rollback if the update encounters any unforeseen compatibility issues.

Why this answer

Implementing a rolling update strategy allows the cluster to be updated one node at a time while the remaining nodes continue to process training jobs. By draining the target node of all active jobs, upgrading the drivers, and then re-adding the node to the production pool, the administrator maintains service availability. This is the standard practice in production AI operations to minimize disruption and ensure consistent driver versions across the entire fleet.

Exam trap

Candidates often select disruptive cluster-wide reboots or simultaneous upgrades instead of controlled, node-by-node rolling updates with draining.

276
Multi-Selectmedium

An administrator is deploying the NVIDIA GPU Operator on a Kubernetes cluster in an air-gapped environment. The cluster nodes have no internet access, and all container images must be pulled from a private registry. Which two actions are required to ensure a successful deployment? (Choose two.)

Select 2 answers
A.Ensure that the private registry supports manifest lists for multi-architecture images.
B.Configure the GPU Operator to use the private registry by setting the image repository in the Helm chart values.
C.Mirror all required NVIDIA GPU Operator images to the private registry.
D.Disable the driver container and install drivers manually on each node.
E.Set the environment variable 'AIRGAP=true' in the operator's deployment.
AnswersB, C

The GPU Operator Helm chart allows specifying a custom image registry and repository for all components. By setting the appropriate values, the Operator will pull images from the private registry instead of NVIDIA's public registry. This configuration is essential to direct the Operator to the mirrored images, enabling successful deployment in an air-gapped setup.

Why this answer

Air-gapped deployments require that all container images are available in a private registry, and the GPU Operator must be configured to pull from that registry. Mirroring the images ensures availability, and setting the image repository in the Helm values directs the Operator to the correct location. These two actions are essential for a successful deployment without internet access.

Exam trap

The trap here is assuming that a special air-gap flag exists in the Operator, when in reality air-gapped support is achieved by mirroring images and configuring the registry settings.

277
MCQhard

Refer to the exhibit. An administrator observes this in the logs during a multi-GPU training job. What is the performance implication of this setting?

A.The training job will have lower memory consumption
B.GPU-to-GPU data transfer performance is significantly reduced
C.The job will be more stable across nodes
D.The system will utilize less power during training
AnswerB

By disabling P2P, the system is forced to move data through the PCIe bus or system memory, which is significantly slower than using NVLink or direct P2P access. This causes a major bottleneck in collective operations like AllReduce, which are fundamental to the scalability of distributed AI training jobs.

Why this answer

Disabling Peer-to-Peer (P2P) communication forces NCCL to use system memory (often via the network or PCIe bus) to facilitate data movement between GPUs. This bypasses the high-bandwidth NVLink interconnect, resulting in significantly increased latency and lower throughput for collective operations. This configuration is typically used only for debugging or when hardware incompatibility prevents direct P2P access, as it severely hinders the performance of multi-GPU, multi-node AI training workloads.

Exam trap

Candidates often assume that if a training job runs, it is running optimally, failing to realize that disabling P2P forces traffic over the slower PCIe bus instead of high-speed NVLink.

278
MCQmedium

When configuring a node for NVIDIA AI Enterprise in a Kubernetes environment, what is the primary function of the NVIDIA Container Toolkit?

A.Provisioning virtual machines
B.Exposing GPUs to containers
C.Optimizing neural network layers
D.Managing Kubernetes network policies
AnswerB

The NVIDIA Container Toolkit provides the necessary runtime libraries and hooks that allow container orchestrators to inject GPU resources into a container. By modifying the container runtime specification, it ensures that device nodes and necessary drivers are mapped into the container's namespace during startup.

Why this answer

The NVIDIA Container Toolkit allows container engines, such as Docker or containerd, to interface with physical GPUs. It enables the exposure of GPUs inside containers, ensuring that applications can access CUDA libraries and hardware acceleration. Without this toolkit, the container runtime cannot bridge the gap between the host's GPU hardware and the containerized process, effectively rendering the GPU unusable for AI workloads.

Exam trap

Candidates often confuse the Container Toolkit with the GPU Operator itself. They assume the Toolkit manages the entire cluster lifecycle rather than its primary, specific purpose of exposing host GPUs to containers.

279
MCQmedium

Which strategy is most effective for managing heterogeneous GPU clusters containing both older architectures (e.g., V100) and newer architectures (e.g., H100)?

A.Enabling uniform scheduling across all nodes.
B.Using Node Selectors and Taints/Tolerations.
C.Configuring the NVIDIA Device Plugin to hide older GPUs.
D.Deploying a single unified container image for all jobs.
AnswerB

Node Selectors and Taints/Tolerations are the standard mechanisms for enforcing hardware compatibility in Kubernetes. They allow operators to isolate workloads to specific GPU architectures, ensuring that code requiring newer hardware features is only placed on nodes capable of supporting them, while older jobs can safely run on legacy hardware.

Why this answer

Using Kubernetes Node Selectors and Taints/Tolerations allows administrators to route specific workloads to compatible hardware. This is essential because newer GPUs support features like MIG or specific precision formats that older hardware lacks. By properly tagging nodes based on architecture, AI Ops prevents 'Incompatible Device' errors and ensures that high-performance training jobs are scheduled on hardware that meets their specific computational requirements, optimizing cluster-wide performance and efficiency.

Exam trap

Many candidates incorrectly suggest using only automated scheduling policies like priority classes, forgetting that heterogeneous hardware requires explicit node-level constraints to prevent incompatible jobs from failing at runtime.

280
MCQeasy

A platform team is preparing a bare-metal Kubernetes cluster to run AI workloads. They want the GPU Operator to install and manage the NVIDIA driver automatically on each node. Which prerequisite must be satisfied on every worker node before the GPU Operator can succeed?

A.The node must have the NVIDIA Container Toolkit preinstalled and configured.
B.The node must be running the NVIDIA Data Center GPU Manager (DCGM) as a systemd service before joining the cluster.
C.The node must have a supported Linux kernel with matching kernel headers available for the driver container to build against.
D.The node must have the NVIDIA vGPU Manager installed so the physical GPU can be partitioned before driver installation.
AnswerC

The GPU Operator's driver container compiles the NVIDIA kernel module against the running kernel, so the matching kernel headers or development package must be present on the host. Without them the driver build fails during installation. This is a documented prerequisite for nodes where the Operator manages the driver rather than relying on a preinstalled one.

Why this answer

The driver container builds the NVIDIA kernel module at runtime against the host kernel, so matching kernel headers or development packages must be present on each node. The GPU Operator manages the container toolkit, DCGM, and device plugin itself, so those do not need to be preinstalled. vGPU Manager applies to virtualized deployments, not bare-metal driver automation.

Exam trap

The trap here is assuming the GPU Operator installs the driver without needing host-level build dependencies such as matching kernel headers.

281
MCQmedium

A platform team runs an NVIDIA AI cluster with the GPU Operator deployed. Users submit jobs directly with kubectl and frequently request whole GPUs even when their notebooks only need a fraction of one. The team wants Kubernetes itself to admit and queue jobs based on GPU demand without users changing their manifests. Which component should the team deploy to meet this requirement?

A.NVIDIA MIG Manager configured for every GPU
B.NVIDIA GPU Operator with the device plugin enabled
C.NVIDIA DCGM Exporter with Prometheus alerting
D.NVIDIA KAI Scheduler
AnswerD

KAI Scheduler is NVIDIA's Kubernetes-native scheduler for AI workloads. It plugs into the cluster as a secondary scheduler and performs gang scheduling, queueing, and GPU fraction accounting, so jobs that request partial GPUs or exceed current capacity wait in a queue instead of being rejected. Because admission and queueing happen inside the scheduler, users keep submitting standard manifests and do not need to learn new tooling.

Why this answer

The requirement is for Kubernetes to make admission and queueing decisions based on GPU demand while users submit normal manifests. KAI Scheduler is the NVIDIA component that provides queue-based, gang-aware scheduling for AI workloads and understands fractional GPU requests, so jobs wait rather than fail. The GPU Operator, MIG Manager, and DCGM Exporter each address installation, partitioning, or observability rather than scheduling admission.

Exam trap

The trap here is assuming the GPU Operator itself performs scheduling or admission, when it only installs and manages the GPU software stack.

282
MCQmedium

An organization is migrating AI workloads to a private cloud. Which feature is essential for ensuring that GPU resources are dynamically reclaimed and reallocated to different departments without manual intervention?

A.Manual pod scheduling with affinity rules.
B.Static GPU reservation per user.
C.Automated cluster autoscaling and priority-based scheduling.
D.Hard-coding node IP addresses in the configuration files.
AnswerC

Automated autoscaling and priority-based scheduling allow the platform to dynamically adjust capacity based on real-time demand. High-priority workloads can preempt lower-priority tasks, ensuring critical research proceeds, while idle resources are automatically reclaimed and made available to other users, maximizing the ROI of the NVIDIA hardware investment.

Why this answer

Dynamic resource scheduling, often implemented via Kubernetes schedulers or custom job orchestrators, is essential for multi-tenant environments. By utilizing features like auto-scaling, preemption, and resource quotas, the system can automatically reclaim idle GPUs from one department and reallocate them to another. This automation maximizes hardware utility and ensures that expensive GPU infrastructure is never left sitting idle due to administrative delays.

Exam trap

Test-takers frequently suggest static allocation schemes or manual administrative interventions, missing the core requirement for automated cluster autoscaling and priority-based scheduling to dynamically reclaim resources.

283
Multi-Selecthard

A cloud architect is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing NVIDIA A100 GPUs. The architect must enable GPU virtualization using NVIDIA vGPU. Which two components are required to support vGPU on the ESXi hosts? (Choose two.)

Select 2 answers
A.NVIDIA GPU Operator for Kubernetes
B.NVIDIA CUDA Toolkit installed on the ESXi host
C.NVIDIA Container Toolkit installed on the ESXi host
D.NVIDIA vGPU guest driver installed in each VM
E.NVIDIA vGPU Manager for VMware ESXi
AnswersD, E

Each VM that uses a vGPU must have the NVIDIA vGPU guest driver installed. This driver communicates with the vGPU Manager on the host to access the GPU. Without the guest driver, the VM cannot utilize the virtual GPU. It is a required component for vGPU functionality.

Why this answer

To enable NVIDIA vGPU on VMware ESXi, the host must have the NVIDIA vGPU Manager VIB installed, which allows the hypervisor to partition the physical GPU. Each VM that uses a vGPU must also have the NVIDIA vGPU guest driver installed. These two components are essential; other tools like the GPU Operator or Container Toolkit are not used on ESXi.

Exam trap

The trap here is confusing Kubernetes-focused tools like the GPU Operator with hypervisor-level components required for vGPU on ESXi.

284
MCQmedium

A DevOps engineer needs to monitor GPU health metrics in real-time for workload management. Which tool provides the most granular visibility into GPU utilization and power consumption for individual containers?

A.Kubernetes Horizontal Pod Autoscaler (HPA) with CPU metrics.
B.NVIDIA DCGM Exporter.
C.Standard Linux 'top' command.
D.Docker stats command.
AnswerB

The DCGM Exporter provides comprehensive, hardware-specific metrics that are essential for deep visibility into GPU performance. It enables the capture of precise utilization data per container, which allows administrators to make data-driven decisions regarding resource allocation, capacity planning, and the optimization of AI training and inference workloads.

Why this answer

The NVIDIA DCGM (Data Center GPU Manager) Exporter is the industry-standard tool for collecting fine-grained GPU telemetry. By integrating with Prometheus and Grafana, it allows administrators to track metrics like power usage, temperature, and utilization at the container level. This granular data is vital for proactive workload management, allowing for autoscaling based on actual hardware demand rather than generic CPU metrics.

Exam trap

Candidates frequently choose generic Kubernetes metrics or standard Prometheus exporters. These tools often lack the specific hooks needed to read deep GPU hardware registers like power draw and memory bandwidth.

285
MCQhard

When configuring a multi-tenant environment on an NVIDIA DGX system, how are MIG (Multi-Instance GPU) instances best provisioned?

A.Through the BIOS settings, creating physical partitions on the GPU die.
B.Using the nvidia-smi command to define the desired MIG profiles.
C.By editing the /etc/nvidia/mig-config.json file directly.
D.Via the OS kernel boot parameters in the GRUB configuration.
AnswerB

nvidia-smi is the correct utility for creating and managing MIG profiles. By applying these profiles, the GPU is partitioned into distinct instances. This provides hardware-level isolation, which is critical for security and performance when multiple users share the same physical GPU infrastructure for different training tasks.

Why this answer

MIG allows a single GPU to be partitioned into multiple isolated instances, each with dedicated compute and memory resources. Configuring MIG via the nvidia-smi tool at the OS level ensures that these partitions are visible to the container runtime. This is crucial for AI operations, as it enables safe, secure multi-tenancy where different users can run jobs simultaneously without resource contention or cross-tenant interference.

Exam trap

Candidates frequently assume MIG is configured via Kubernetes manifests or YAML files. While Kubernetes manages the pods, the actual hardware partitioning must be defined at the OS level via nvidia-smi.

286
MCQmedium

A platform team is installing the NVIDIA GPU Operator on a Kubernetes cluster that runs a mix of GPU and non-GPU nodes. They want the operator to manage the driver lifecycle only on nodes that actually have NVIDIA GPUs, without requiring manual taints on non-GPU nodes. Which configuration should they apply to the ClusterPolicy to achieve this?

A.Set `driver.enabled: false` and rely on preinstalled drivers on all nodes.
B.Set `driver.rdma.enabled: true` to limit driver operations to GPU nodes.
C.Set `driver.nodeSelector` in the ClusterPolicy to match a label such as `nvidia.com/gpu.present: "true"`.
D.Set `driver.enabled: true` and `driver.useNvidiaDriverRoot: true` in the ClusterPolicy.
AnswerC

The `driver.nodeSelector` field in the ClusterPolicy restricts where the driver daemonset is scheduled. By selecting only nodes labeled with `nvidia.com/gpu.present: "true"`, the operator manages drivers exclusively on GPU nodes. This avoids manual taints on non-GPU nodes and matches the desired automatic scoping to GPU hardware.

Why this answer

The GPU Operator uses node selectors to determine where components are deployed. Configuring `driver.nodeSelector` with a label that identifies GPU-equipped nodes ensures driver management is scoped to those nodes only. This avoids the need for manual taints on non-GPU nodes and aligns with the operator's declarative model for heterogeneous clusters.

Exam trap

The trap here is confusing driver configuration fields such as `driver.enabled` or `driver.rdma.enabled` with node placement controls, when only `driver.nodeSelector` governs which nodes are targeted.

287
MCQeasy

A data science team submits a PyTorch training job to a Kubernetes cluster where the NVIDIA GPU Operator is installed. The pod stays in Pending state, and kubectl describe shows the message '0/6 nodes are available: 6 Insufficient nvidia.com/gpu.' The administrator confirms the nodes have healthy GPUs and the device plugin pods are Running. What is the most likely cause?

A.The pod requested more nvidia.com/gpu resources than any single node can advertise
B.The container image does not include the NVIDIA CUDA base layer
C.The pod lacks a nodeSelector matching the GPU node label nvidia.com/gpu.present
D.The GPU Operator's driver container has not yet built the kernel module on the nodes
AnswerA

Extended resources such as nvidia.com/gpu are integer-quantized per node and cannot be oversubscribed across nodes. If the pod requests a count exceeding what any single node advertises, the scheduler reports insufficient nvidia.com/gpu on every node and the pod stays Pending. Reducing the request to fit one node resolves the condition.

Why this answer

Extended resources like nvidia.com/gpu are advertised per node and cannot be aggregated across nodes, so a request exceeding any single node's advertised count leaves the pod unschedulable with an insufficient-resource message. Verifying the per-node GPU count and lowering the pod's request to fit within one node restores scheduling.

Exam trap

The trap here is reading 'Insufficient nvidia.com/gpu' as a driver or image problem, when it actually means the requested GPU count exceeds what any individual node advertises.

288
MCQhard

An AI operations team is running a large language model inference service on NVIDIA H100 GPUs using NVIDIA Triton Inference Server. They observe that the first inference request after a period of inactivity takes significantly longer than subsequent requests. The model is loaded and ready, but the GPU shows low utilization during the first request. Which optimization should the team implement to reduce this latency spike?

A.Configure Triton's model warmup to run dummy inference requests during model loading.
B.Enable Triton's dynamic batching with a large maximum batch size.
C.Set the Triton `--pinned-memory-pool-byte-size` to a larger value.
D.Increase the number of model instances per GPU to allow more concurrent executions.
AnswerA

Triton's model warmup feature executes a specified number of inference requests when the model is loaded, ensuring that CUDA kernels are compiled, memory allocations are made, and the GPU is initialized. This eliminates the cold-start penalty for the first real request. Setting warmup with representative input shapes and batch sizes directly reduces the latency spike after periods of inactivity.

Why this answer

The first inference after inactivity is slow because CUDA kernels and memory allocations are not yet initialized on the GPU. Triton's model warmup runs dummy requests at load time to trigger this initialization, so the first real request executes at normal speed. Other options target throughput or memory pooling but do not eliminate the cold-start penalty.

Exam trap

The trap here is confusing cold-start latency with throughput optimization, leading to batching or instance scaling instead of pre-warming the model.

289
MCQhard

An administrator is tasked with deploying a multi-node training job using the NVIDIA GPU Operator. Which configuration must be present to ensure that pods are scheduled on nodes with identical GPU architectures to prevent performance degradation?

A.Enable the 'auto-scaling' feature in the NVIDIA GPU Operator.
B.Use node affinity labels based on NFD-provided hardware information.
C.Increase the timeout values for the NCCL collective communication operations.
D.Set the 'nvidia.com/gpu' resource limit to zero on all nodes except the master node.
AnswerB

Node Feature Discovery (NFD) labels nodes with their GPU model and architecture. By using these labels in the training pod's affinity configuration, the administrator forces the scheduler to select only nodes that match the desired hardware profile, ensuring consistent performance for distributed training jobs that rely on identical hardware capabilities.

Why this answer

To maintain high-performance, synchronized training, all participating nodes should share the same GPU architecture and interconnect type (e.g., NVLink or InfiniBand). Using Kubernetes node affinity or anti-affinity rules combined with the hardware labels automatically generated by the Node Feature Discovery (NFD) service ensures that the scheduler places the job exclusively on compatible nodes. This prevents the training from falling back to slower, sub-optimal communication paths or heterogeneous hardware modes.

Exam trap

Candidates often rely on default scheduling, forgetting that Kubernetes is unaware of GPU architecture differences. They fail to use NFD labels, leading to heterogeneous nodes that cause significant performance degradation.

290
MCQhard

An MLOps engineer needs to guarantee that a latency-sensitive inference Deployment always has GPU capacity available, even when a large training Job is submitted to the same namespace. The cluster uses the NVIDIA GPU Operator and nodes have four A100 GPUs each. Which approach reliably reserves GPU capacity for the inference Deployment?

A.Create a separate node pool and use nodeSelector or nodeAffinity to pin the inference Deployment to nodes tainted for inference only
B.Set a higher PriorityClass on the inference Deployment so it preempts training pods when needed
C.Enable the NVIDIA MPS control daemon and configure each inference pod with a shared memory fraction
D.Apply a ResourceQuota on the namespace that limits nvidia.com/gpu to the number of inference replicas
AnswerA

Dedicating a tainted node pool and pinning the inference Deployment with nodeSelector or nodeAffinity guarantees that training pods cannot consume those GPUs. Taints repel pods that lack the matching toleration, so the reserved capacity is genuinely protected regardless of how much training demand arrives, which is the only option that provides a hard reservation.

Why this answer

Only a dedicated, tainted node pool combined with nodeSelector or nodeAffinity creates a hard capacity reservation that training workloads cannot violate. Priority-based preemption and MPS improve scheduling behavior or sharing but do not prevent a co-located training Job from consuming the GPUs first, and a namespace ResourceQuota limits aggregate requests without reserving specific capacity for the inference Deployment.

Exam trap

The trap here is believing that a higher PriorityClass reserves GPU capacity, when it only enables eviction after a scheduling failure, not proactive reservation.

291
MCQmedium

A Kubernetes cluster running the NVIDIA GPU Operator is shared by an inference team and a research team. The research team's training pods repeatedly evict the inference pods from GPUs, causing latency spikes in production. The administrator wants to guarantee that inference pods always get GPU access first. Which Kubernetes scheduling mechanism should be configured?

A.Enable the NVIDIA MIG feature on all GPUs and dedicate a MIG instance to each inference pod.
B.Assign a higher PriorityClass to the inference pods and enable preemption on the scheduler.
C.Configure a taint on the GPU nodes and add the corresponding toleration only to the inference pods.
D.Create a PodDisruptionBudget for the inference deployment and set maxUnavailable to zero.
AnswerB

PriorityClass with preemption allows higher-priority inference pods to be scheduled and to evict lower-priority training pods when GPU resources are scarce. This directly protects production inference latency by ensuring inference workloads are admitted first, which matches the requirement to guarantee GPU access for the inference team.

Why this answer

Priority and preemption are the native Kubernetes controls that decide which pods win when GPU resources are contended. Giving inference pods a higher PriorityClass lets the scheduler admit them first and evict lower-priority training pods when necessary, which directly satisfies the requirement that production inference always obtains GPU access.

Exam trap

The trap here is assuming that taints, tolerations, or PodDisruptionBudgets control scheduler contention, when in fact only PriorityClass with preemption orders competing pods for scarce GPU resources.

292
MCQhard

Refer to the exhibit. The system has two GPUs. What is the most likely cause of the observed performance discrepancy?

A.The GPUs have different amounts of physical memory.
B.Only one GPU is receiving the workload due to improper data distribution.
C.The system is limited by the PCIe bus bandwidth.
D.The model is too small to be parallelized across two GPUs.
AnswerB

The discrepancy between high GPU utilization on one device and low utilization on the other strongly suggests that only one GPU is performing the compute-intensive training loop. This is typical when the DataParallel or DistributedDataParallel wrapper is not correctly configured across all available devices.

Why this answer

The exhibit shows one GPU heavily utilized while the second is idling despite similar memory consumption. This indicates a data parallelism imbalance, where one process is doing the bulk of the work. This is a common issue in multi-GPU setups where workload distribution is not correctly handled, leading to massive inefficiencies where expensive hardware is under-utilized, significantly increasing the time required for model training or inference tasks.

Exam trap

Candidates often guess hardware failure or driver mismatch, when the most common issue is a simple failure to properly initialize or distribute workloads across both available GPUs in the application code.

293
Multi-Selecthard

An AI infrastructure team is deploying NVIDIA AI Enterprise on a Kubernetes cluster using the NVIDIA GPU Operator. They need to ensure that the GPU Operator can successfully manage GPUs and that workloads can consume GPU resources. Which two components does the GPU Operator deploy to enable GPU scheduling and container GPU access? (Choose two.)

Select 2 answers
A.NVIDIA Device Plugin
B.NVIDIA Container Toolkit
C.NVIDIA Persistence Daemon
D.NVIDIA MIG Manager
E.NVIDIA GPU Driver
AnswersA, B

The NVIDIA Device Plugin is deployed by the GPU Operator to advertise GPU resources to the Kubernetes API server. It allows the scheduler to allocate GPUs to pods that request them. Without it, Kubernetes would not know about the GPUs, and pods would not be scheduled with GPU resources. This component is essential for GPU scheduling.

Why this answer

The NVIDIA GPU Operator deploys the NVIDIA Device Plugin to enable Kubernetes to schedule pods with GPU resources, and the NVIDIA Container Toolkit to allow containers to access GPUs. The device plugin registers GPUs with the Kubernetes API, while the container toolkit configures the container runtime. Together, they provide the necessary integration for GPU-accelerated workloads.

Other components like the driver and MIG manager are important but not the specific enablers for scheduling and container access.

Exam trap

The trap here is assuming that the GPU driver alone is sufficient for GPU scheduling and container access, overlooking the need for the device plugin and container toolkit.

294
MCQhard

A team is deploying a large language model for inference using NVIDIA TensorRT-LLM on an H100 GPU. They observe that the first inference request takes several seconds, while subsequent requests are fast. They want to reduce this initial latency. Which technique should they implement?

A.Increase the GPU's power limit to boost clock speeds during the first request.
B.Use TensorRT-LLM's built-in paged KV cache and enable continuous batching.
C.Precompile the TensorRT engine and load it at server startup, then perform a warm-up inference.
D.Reduce the model's precision to INT4 to decrease computation time.
AnswerC

The first request latency includes engine deserialization, CUDA context creation, and kernel loading. Precompiling the engine and loading it during startup, followed by a warm-up inference, ensures that these one-time costs are paid before actual requests arrive. This directly reduces the first-request latency for users.

Why this answer

The first inference request incurs one-time costs such as TensorRT engine deserialization, CUDA context setup, and kernel compilation/loading. By precompiling the engine and loading it at startup, and then running a warm-up inference, these costs are moved to server initialization. Subsequent requests then benefit from a fully initialized environment, reducing the observed initial latency.

Exam trap

The trap here is confusing steady-state optimizations like continuous batching or precision reduction with cold-start latency, which is caused by initialization overhead.

295
MCQeasy

A data science team submits a PyTorch training job to a Kubernetes cluster managed by Run:ai. The job requests two GPUs but only one is allocated, and the second worker hangs waiting for a peer. Which Run:ai capability should the administrator verify is configured so the distributed job receives all requested GPUs atomically?

A.A higher priority class assigned to the training job so it preempts other workloads.
B.Gang scheduling, which ensures all pods in a distributed job are scheduled together or not at all.
C.Node affinity rules that pin each worker to a specific GPU node by hostname.
D.Enabling time-slicing on the GPU device plugin to increase the apparent GPU count.
AnswerB

Distributed training jobs require all workers to start together; partial allocation causes hangs because ranks wait for peers that never launch. Run:ai's gang scheduling treats the workload as an atomic unit, allocating all requested GPUs or leaving the job pending. Verifying this configuration addresses the symptom of one GPU allocated and a stalled peer directly.

Why this answer

Distributed training depends on every rank being present before collective operations begin; a single missing worker causes the rest to block. Run:ai's gang scheduling is the mechanism that admits the entire workload as a unit, so all requested GPUs are granted together. Confirming gang scheduling is enabled and applied to the job resolves the partial allocation that produced the hang.

Exam trap

The trap here is assuming that priority or affinity alone can prevent partial placement, when only all-or-nothing gang scheduling guarantees every rank starts together.

296
MCQhard

An administrator wants to ensure that a training process is limited to a single GPU on a multi-GPU node. Which environment variable should be set?

A.NCCL_DEBUG=INFO
B.CUDA_VISIBLE_DEVICES=0
C.NVIDIA_DRIVER_CAPABILITIES=compute
D.OMP_NUM_THREADS=1
AnswerB

CUDA_VISIBLE_DEVICES is the standard environment variable used to mask specific GPUs from a process. Setting it to a specific index restricts the application to use only that hardware device, which is the standard method for isolating jobs in a multi-GPU system.

Why this answer

Controlling GPU visibility is a fundamental skill for resource management in multi-tenant environments. By using CUDA_VISIBLE_DEVICES, an administrator can restrict a process to a specific device, preventing multiple jobs from competing for the same GPU. This isolation is crucial for maintaining performance stability and ensuring that individual jobs receive consistent, predictable access to compute resources without interference from other concurrent tasks.

Exam trap

Students often mistakenly select command-line flags or code-level device placement arguments instead of the standard operating system environment variable required to restrict GPU visibility globally.

297
MCQmedium

An AI engineer observes that a model training job on an NVIDIA DGX system is underutilizing the GPU. The training loop shows frequent "CPU bottleneck" warnings in the logs. Which action should the engineer take first to optimize throughput?

A.Upgrade to a higher-end GPU model to handle the processing load.
B.Increase the batch size significantly to fill the GPU memory.
C.Implement NVIDIA DALI to offload preprocessing tasks from the CPU to the GPU.
D.Reduce the number of training epochs to lower CPU overhead.
AnswerC

NVIDIA DALI is specifically designed to accelerate data preprocessing pipelines by moving them from the CPU to the GPU. This eliminates the bottleneck by ensuring that data augmentation and transformation tasks occur at the same high speed as the training process, maximizing overall system hardware utilization.

Why this answer

Identifying CPU bottlenecks is critical because data pipelines often struggle to keep up with GPU compute speed. By optimizing data preprocessing, specifically increasing the number of workers in the DataLoader or using NVIDIA DALI, the engineer ensures the GPU remains saturated with data. This optimization directly impacts total training time and infrastructure cost efficiency, ensuring that high-performance hardware is not left idling while waiting for I/O operations.

Exam trap

Candidates often suggest upgrading the GPU or increasing the batch size, which exacerbates the CPU bottleneck rather than solving the underlying data ingestion starvation occurring at the preprocessing layer.

298
Multi-Selecthard

An AI researcher is deploying a multi-node training job using NCCL on an InfiniBand network. The job is suffering from intermittent latency spikes. Which TWO steps should the engineer perform to troubleshoot the network configuration?

Select 2 answers
A.Check IB link status and error counters using 'ibstat'.
B.Reset the GPU thermal throttling threshold.
C.Update the OS kernel to the latest version.
D.Verify NCCL_IB_HCA and NCCL_IB_GID_INDEX settings.
E.Increase the system swap space to 512GB.
AnswersA, D

Monitoring InfiniBand link counters is essential for detecting physical layer issues like CRC errors or link retrains. If the physical link is unstable, NCCL collective operations will experience significant latency, causing training to stall or fail. Identifying these errors early isolates the issue to the fabric cabling or switches.

Why this answer

Troubleshooting high-performance networking in distributed training requires checking for physical layer stability and protocol-level misconfigurations. Verifying InfiniBand counters helps identify packet drops or link flapping, while ensuring NCCL environment variables are correctly set for the specific topology avoids suboptimal routing. These steps ensure that the inter-GPU communication remains within the high-bandwidth low-latency envelope required for scaling deep learning workloads effectively.

Exam trap

Candidates often try to debug the training application code or model synchronization logic, failing to check the physical InfiniBand layer and NCCL environment variables that govern high-speed network communication.

299
MCQmedium

An administrator notices that GPU utilization is high, but throughput in an AI training job remains low. What is the most likely bottleneck?

A.Insufficient CUDA cores on the GPU
B.Data pipeline or storage I/O bottleneck
C.Incompatible NVIDIA driver version
D.Excessive usage of GPU registers
AnswerB

High GPU utilization accompanied by low throughput indicates the GPU is spending time waiting for data to arrive from the CPU or storage. This starvation effect is a classic symptom of an inefficient data loader or slow storage subsystem, which limits the overall throughput despite the GPU's apparent activity.

Why this answer

When GPU utilization is high but throughput is low, the GPU is likely stalling while waiting for data. This is typically a sign of an input/output (I/O) bottleneck, where the data pipeline (e.g., loading images from disk or network) cannot keep up with the GPU's processing speed. Ensuring the data preprocessing pipeline is sufficiently parallelized and optimized is crucial for maximizing GPU utilization and maintaining high training throughput in AI workloads.

Exam trap

Candidates often assume that high GPU utilization is a sign of a healthy, efficient training job, failing to recognize that it can actually indicate the GPU is idling while waiting for data.

300
MCQeasy

An administrator is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing several NVIDIA A100 GPUs. They need to enable GPU sharing across multiple virtual machines to maximize utilization. Which vSphere feature should they configure?

A.vSphere DRS
B.SR-IOV
C.NVIDIA vGPU
D.DirectPath I/O
AnswerC

NVIDIA vGPU allows a single physical GPU to be partitioned into multiple virtual GPUs, each assigned to a different VM. This enables sharing and maximizes utilization. It requires vSphere with NVIDIA vGPU software and compatible GPUs. This is the correct feature for GPU sharing in vSphere.

Why this answer

NVIDIA vGPU is the vSphere feature that partitions a physical GPU into multiple virtual GPUs, allowing multiple VMs to share the same physical GPU. This maximizes GPU utilization and is designed for AI workloads. DirectPath I/O provides exclusive access, while DRS and SR-IOV do not enable GPU sharing.

Exam trap

The trap here is confusing GPU sharing with GPU passthrough, where DirectPath I/O gives exclusive access to one VM, not sharing.

Page 3

Page 4 of 5

Page 5

All pages