Courseiva

NVIDIA Certified Professional: AI Operations (NCP-AIO) — Questions 151–225

309 questions total · 5pages · All types, answers revealed

Page 2

Page 3 of 5

Page 4
151
MCQeasy

A monitoring system reports that a DGX node's GPUs are running at reduced clocks during a long training job, and `nvidia-smi -q -d PERFORMANCE` shows the throttle reason as 'SW Power Cap'. The job's power draw is at the configured limit. Which action is MOST appropriate to restore higher clocks?

A.Reduce the batch size so the GPU consumes less power and stays below the cap.
B.Raise the GPU power limit with `nvidia-smi -pl` to a value within the GPU's supported maximum and re-run the job.
C.Enable persistence mode with `nvidia-smi -pm 1` to keep the driver loaded and improve clock stability.
D.Lock the GPU clocks to their maximum with `nvidia-smi -lgc` to override the throttle.
AnswerB

A 'SW Power Cap' throttle reason means the GPU is hitting its configured power limit, so clocks are reduced to stay within budget. If the hardware and cooling support it, raising the power limit to a validated value within the GPU's supported range allows higher sustained clocks, directly addressing the throttle cause. This is the correct, targeted remediation for the reported condition.

Why this answer

The throttle reason 'SW Power Cap' indicates the GPU is constrained by its configured power limit. Raising the power limit to a supported value within the GPU's validated maximum allows the device to sustain higher clocks, provided cooling and power delivery can support it. This directly removes the cause of the reduced clocks rather than working around it.

Exam trap

The trap here is confusing a power-cap throttle with a thermal or clock-lock issue and attempting to override clocks instead of adjusting the power budget.

152
MCQmedium

An ML platform team runs an NVIDIA GPU Operator-managed cluster and wants to allow multiple pods to share a single A100 GPU so that small inference services can co-reside without each consuming a whole device. The team needs a time-slicing configuration that applies to all GPU nodes in the cluster. Which approach should the administrator take?

A.Add a tolerations entry for nvidia.com/gpu to each pod and increase the kubelet's --max-pods flag on GPU nodes.
B.Set the environment variable NVIDIA_VISIBLE_DEVICES=all on each pod and let the runtime divide GPU time among containers.
C.Create a ConfigMap containing the time-slicing configuration and reference it in the ClusterPolicy so the GPU Operator propagates the device plugin config across nodes.
D.Install the NVIDIA MIG Manager and set the nvidia.com/mig.config label to all-1g.5gb on each node to subdivide the GPUs.
AnswerC

The GPU Operator supports time-slicing by reading a ConfigMap referenced in the ClusterPolicy's device plugin configuration. When applied, the operator propagates the config to all GPU nodes and restarts the device plugin, causing each physical GPU to advertise a multiplied replica count so multiple pods can share one device. This is the supported cluster-wide method.

Why this answer

Time-slicing in the NVIDIA GPU Operator is driven by a device-plugin configuration that the operator distributes to every GPU node through the ClusterPolicy. The referenced ConfigMap declares a replica count, and the device plugin then advertises that many virtual GPUs per physical device, allowing the scheduler to place multiple pods on one GPU. This is the supported, cluster-wide mechanism for enabling time-sliced sharing.

Exam trap

The trap here is confusing per-pod GPU visibility variables or MIG partitioning with time-slicing, which actually requires a device-plugin ConfigMap referenced by the ClusterPolicy.

153
MCQmedium

An organization is migrating their on-premises AI training to a hybrid cloud environment. Which component is most important to maintain consistent workload management across both the on-premises DGX systems and cloud-based GPU nodes?

A.Deploying identical physical GPU hardware across all sites.
B.Using a unified Kubernetes orchestration layer with consistent GPU Operator versions.
C.Hard-coding all GPU resource requests to match the smallest cloud GPU instance.
D.Implementing separate workload managers for cloud and on-premises sites.
AnswerB

Maintaining a unified orchestration layer allows for uniform resource management and scheduling logic across environments. By standardizing the GPU Operator and the Kubernetes API, teams can move workloads seamlessly without rewriting manifests or changing operational workflows, which is the primary challenge in managing hybrid GPU infrastructure.

Why this answer

A unified Kubernetes control plane, managed by an orchestrator like NVIDIA Base Command or a managed Kubernetes service (e.g., GKE or EKS with NVIDIA GPU support), provides a consistent API. This allows developers to use the same manifest files and CI/CD pipelines regardless of whether the physical hardware is in a local datacenter or in the cloud. It ensures that GPU scheduling, resource requests, and monitoring tools remain identical, simplifying the operations lifecycle.

Exam trap

Candidates often focus on data synchronization or network latency, missing that the primary operational hurdle in hybrid environments is inconsistent software versions across the GPU Operator and drivers.

154
MCQhard

A production inference service on NVIDIA GPUs reports that p99 latency spikes every few minutes while p50 remains stable. Metrics show GPU memory utilization near the limit and periodic `cudaMalloc` calls in the application logs. The model and batch size are fixed. Which change is MOST likely to eliminate the latency spikes?

A.Enable CUDA lazy module loading to defer kernel module initialization until first use.
B.Increase the number of model instances per GPU to smooth out request distribution.
C.Lower the GPU clock to reduce power draw and stabilize thermals during inference.
D.Pre-allocate the inference workspace and reuse CUDA memory pools instead of calling cudaMalloc during request handling.
AnswerD

Periodic `cudaMalloc` calls during request handling force synchronous memory allocation and can trigger driver-level operations that stall the GPU pipeline, producing p99 spikes while p50 stays flat. Pre-allocating workspaces and using a memory pool (or CUDA graphs with static allocations) removes allocation from the hot path, which directly addresses the observed pattern and is the correct remediation here.

Why this answer

The correlation between p99 spikes and periodic `cudaMalloc` calls, combined with stable p50, points to allocation occurring inside the request path. Synchronous allocation can serialize GPU work and introduce jitter. Pre-allocating workspaces and using a memory pool or CUDA graphs with static buffers removes the allocator from the hot path, eliminating the periodic stalls while leaving steady-state latency unchanged.

Exam trap

The trap here is attributing p99 spikes to concurrency or thermals, when the log evidence ties them to allocator calls in the request path.

155
MCQmedium

A researcher submits a distributed training job that spans four pods, each needing one GPU, and the pods must start together or not at all. The administrator wants Kubernetes to schedule all four pods only when four GPUs are simultaneously available. Which workload management construct should be used?

A.A Kubernetes Job with parallelism set to 4 and completions set to 4.
B.A StatefulSet with podManagementPolicy set to Parallel.
C.A DaemonSet that places one training pod on each GPU node in the cluster.
D.A PodGroup managed by a scheduler that supports gang scheduling, such as Volcano or the scheduler-plugins coscheduling plugin.
AnswerD

Gang scheduling ensures that all pods in a PodGroup are scheduled together or none are, preventing partial starts that waste GPUs and stall distributed training. This directly satisfies the requirement that the four GPU pods begin only when four GPUs are simultaneously available.

Why this answer

Distributed training needs gang scheduling so that either every worker pod is placed or none is. A PodGroup interpreted by a gang-aware scheduler like Volcano or the coscheduling plugin holds the pods until the full set of GPUs is available, avoiding deadlocks and wasted accelerator time.

Exam trap

The trap here is assuming that a standard Job or StatefulSet provides atomic, all-or-nothing scheduling, when in fact only gang-scheduling constructs enforce that guarantee.

156
MCQhard

A research team submits a multi-node training job using a `Job` with eight pods, each requesting one GPU. The cluster has eight GPU nodes, each with one A100. The administrator observes that all eight pods are spread one per node and the job runs, but throughput is far below expectations and NCCL logs show repeated fallback from GPUDirect RDMA to socket transport. Which action most directly addresses the root cause?

A.Set `NCCL_P2P_DISABLE=1` in the job's environment to force peer-to-peer transfers over PCIe.
B.Deploy the NVIDIA Network Operator to install and configure the RDMA stack and GPUDirect RDMA on the GPU nodes.
C.Add a node affinity rule forcing all eight pods onto a single node to enable NVLink communication.
D.Increase the number of replicas in the Job so more GPUs participate in the collective.
AnswerB

NCCL falling back to socket transport indicates the RDMA path is unavailable, typically because the high-speed fabric drivers, RDMA devices, and GPUDirect RDMA support are not provisioned on the nodes. The Network Operator deploys and configures those components alongside the GPU Operator, restoring the RDMA transport that NCCL prefers for multi-node collectives.

Why this answer

The symptom points to NCCL abandoning RDMA and using the slower socket transport for inter-node collectives. That path depends on the high-speed fabric drivers, RDMA devices, and GPUDirect RDMA support being installed and configured on every GPU node, which is precisely what the NVIDIA Network Operator provisions in tandem with the GPU Operator. Changing rank counts, disabling peer-to-peer, or collapsing topology onto one node does not restore the missing RDMA capability.

Exam trap

The trap here is tuning NCCL environment variables or job topology when the actual gap is that the RDMA and GPUDirect components were never deployed on the nodes.

157
MCQhard

A financial services company is deploying NVIDIA AI Enterprise on a Kubernetes cluster with strict security policies. They need to ensure that GPU workloads are isolated and that the NVIDIA GPU Operator components are deployed with least privilege. Which feature of the NVIDIA GPU Operator allows administrators to define granular permissions for its components?

A.Security Context Constraints
B.PodSecurityPolicy
C.Role-Based Access Control (RBAC)
D.Network Policies
AnswerC

RBAC in Kubernetes allows administrators to define roles and role bindings that specify which actions are permitted on which resources. The GPU Operator uses RBAC to grant its components the minimum necessary permissions. This enables least-privilege access and is essential for strict security environments.

Why this answer

RBAC is the Kubernetes mechanism for defining granular permissions. The NVIDIA GPU Operator deploys components with specific ServiceAccounts and RBAC roles that grant only the permissions needed to perform their functions. This aligns with least-privilege principles and is critical for security-sensitive deployments.

Other options are either deprecated, network-focused, or platform-specific.

Exam trap

The trap here is assuming that PodSecurityPolicy or Network Policies can restrict API permissions; only RBAC governs what actions components can perform on Kubernetes resources.

158
MCQhard

An administrator manages an NVIDIA AI Enterprise deployment on Kubernetes and needs to ensure that GPU telemetry is exported to an existing Prometheus instance. The administrator deploys the NVIDIA DCGM Exporter but sees no GPU metrics in Prometheus. Which configuration should the administrator verify first?

A.The NVIDIA GPU Operator's device plugin is configured to advertise GPUs with the correct resource name.
B.The GPU driver version installed on the nodes matches the version bundled in the DCGM Exporter container.
C.The cluster's CNI plugin supports multicast so that Prometheus can discover the exporter.
D.The DCGM Exporter ServiceMonitor or PodMonitor selector labels match the Prometheus operator's serviceMonitorSelector.
AnswerD

When Prometheus is managed by the Prometheus Operator, it only scrapes targets whose ServiceMonitor or PodMonitor labels match the operator's serviceMonitorSelector or podMonitorSelector. If the DCGM Exporter's monitor labels do not match, Prometheus silently ignores the endpoint even though the exporter is running. Verifying selector label alignment is the correct first step because it is the most common cause of missing metrics in operator-managed Prometheus.

Why this answer

In Prometheus Operator deployments, scrape targets are selected by label matching between the ServiceMonitor or PodMonitor and the Prometheus custom resource. If the DCGM Exporter's monitor labels do not align with the operator's selector, Prometheus never scrapes it, producing exactly the symptom of no GPU metrics. Verifying label alignment is therefore the correct first check.

Exam trap

The trap here is assuming that a running DCGM Exporter automatically gets scraped, when operator-managed Prometheus only scrapes targets whose monitor labels match its selector.

159
MCQmedium

An administrator wants to prevent unauthorized users from accessing sensitive model weights stored in GPU memory. Which security feature should be implemented to ensure hardware-level isolation of the memory space?

A.Enable standard Linux filesystem permissions.
B.Implement NVIDIA Confidential Computing.
C.Use a simple SSH firewall rule.
D.Update the NVIDIA CUDA shared library path.
AnswerB

NVIDIA Confidential Computing utilizes hardware-based Trusted Execution Environments (TEEs) to encrypt data while it resides in GPU memory. This prevents unauthorized access from other processes, the kernel, or the hypervisor, ensuring that sensitive model weights remain secure even if the software environment is considered untrusted or potentially compromised.

Why this answer

Hardware-level memory isolation via Confidential Computing technologies, such as NVIDIA Confidential Computing (CC), protects data in use by encrypting memory contents. For administrators handling sensitive IP, this provides a root-of-trust that persists even if the OS or hypervisor is compromised. This is a critical administrative control for ensuring compliance and data sovereignty in multi-tenant cloud environments where infrastructure is shared among different departments or organizations.

Exam trap

Candidates often confuse software-level encryption or standard disk encryption with hardware-level memory isolation. They incorrectly select general security tools instead of the specific NVIDIA Confidential Computing framework required for GPU-level memory protection.

160
MCQmedium

An AI operations engineer is troubleshooting a multi-node NCCL training job on an NVIDIA DGX SuperPOD. The job runs but scales poorly: inter-node bandwidth is roughly half of the expected 200 Gb/s per GPU, while intra-node NVLink traffic is at full rate. Running `nvidia-smi topo -m` shows that GPUs in each node are connected to the NICs through the PCIe switch, but the job sets `NCCL_SOCKET_IFNAME` to the management interface. Which action is the most appropriate to resolve the bottleneck?

A.Enable GPUDirect RDMA by allowing the container access to the host IB verbs devices and the `/dev/infiniband` tree, and set `NCCL_IB_HCA` to the compute fabric adapters.
B.Increase `NCCL_BUFFSIZE` to 16 MB and set `NCCL_NTHREADS` to 8 to raise channel throughput.
C.Set `NCCL_P2P_DISABLE=1` so that all inter-node traffic is routed through the CPUs, avoiding PCIe switch contention.
D.Bind the job to a single NUMA node using `numactl --cpunodebind=0 --membind=0` to reduce cross-socket traffic.
AnswerA

The symptom of full intra-node NVLink but degraded inter-node throughput points to the job falling back to TCP over the management interface. GPUDirect RDMA lets the NIC read/write GPU memory directly over the InfiniBand/RoCE compute fabric, bypassing host memory copies. Exposing `/dev/infiniband` to the container and pinning `NCCL_IB_HCA` restores the intended high-speed path.

Why this answer

When intra-node NVLink performs at full rate but inter-node bandwidth is roughly half, the job is usually not using the compute fabric. GPUDirect RDMA bypasses host memory copies by letting the HCA access GPU memory directly, and `NCCL_IB_HCA` ensures NCCL selects the correct adapters. Exposing `/dev/infiniband` to the container is required for RDMA verbs.

This restores the expected inter-node bandwidth.

Exam trap

The trap here is assuming that NCCL tuning parameters such as buffer size or thread count can compensate for a transport that is falling back to TCP over the management network.

161
MCQmedium

An AI engineer observes that a training job on an NVIDIA DGX H100 system is experiencing significant performance degradation. The GPU utilization is high, but the throughput remains low. Which tool should be used first to identify if the bottleneck is related to data loading or PCIe bandwidth saturation?

A.NVIDIA-SMI
B.NVIDIA DCGM
C.NVIDIA Nsight Systems
D.NVIDIA Nsight Compute
AnswerC

Nsight Systems provides a system-wide view of CPU and GPU activities, allowing engineers to correlate data transfer operations with kernel execution. This visibility is essential for identifying stalls, data starvation, or PCIe bus congestion, which are the most common causes of low throughput despite high GPU utilization.

Why this answer

NVIDIA Nsight Systems is the primary tool for analyzing system-wide performance, including CPU-GPU interactions and data transfer bottlenecks. Identifying whether a bottleneck occurs in the data pipeline or within the hardware interconnects is critical for optimizing training speed. By visualizing timelines, engineers can pinpoint if the GPU is starving for data or if the PCIe bus is congested, enabling targeted remediation for large-scale distributed training clusters.

Exam trap

Candidates often select NVIDIA-SMI or Nsight Compute, failing to realize that Nsight Systems is the correct tool for identifying system-wide bottlenecks between the CPU, data pipeline, and GPU.

162
Multi-Selectmedium

An AI operations team is using NVIDIA DCGM (Data Center GPU Manager) to monitor a cluster of A100 GPUs. They want to set up proactive health checks to detect and mitigate GPU issues before they cause job failures. Which two DCGM features should they configure? (Choose two.)

Select 2 answers
A.DCGM policy management for automated remediation
B.DCGM profiling metrics for real-time utilization
C.DCGM group configuration for multi-node synchronization
D.DCGM health checks with periodic diagnostics
E.DCGM API integration with Prometheus for alerting
AnswersA, D

DCGM policy management allows administrators to define policies that trigger actions when health violations occur, such as resetting a GPU or cordoning a node. This enables automated mitigation, reducing downtime. It works in conjunction with health checks. Configuring policies ensures that detected issues are handled promptly, which is essential for proactive health management in a large GPU cluster.

Why this answer

DCGM health checks with periodic diagnostics and DCGM policy management for automated remediation are the two features that directly enable proactive health monitoring and mitigation. Health checks detect issues, and policies define actions to take when issues are found. Together, they form a proactive health management system.

Other options are either monitoring metrics or management features that do not directly address proactive health.

Exam trap

The trap here is selecting monitoring metrics like profiling as health checks, when they only provide performance data, not issue detection and remediation.

163
MCQhard

An administrator is responsible for an NVIDIA AI Enterprise deployment on Kubernetes. The security team requires that all GPU-accelerated pods run with the least privilege necessary and that GPU device nodes are not exposed to pods that do not request them. Which combination of configurations should the administrator implement to meet these requirements?

A.Deploy the NVIDIA GPU Operator with the device plugin and configure a PodSecurityPolicy or OPA Gatekeeper policy that requires pods to request nvidia.com/gpu and forbids privileged mode.
B.Use the NVIDIA GPU Operator with the device plugin and configure pod security contexts to drop all capabilities and add only the NVIDIA_VISIBLE_DEVICES environment variable.
C.Install the NVIDIA k8s-device-plugin standalone and set the --pass-device-specs flag to true, then allow all pods to run as root.
D.Enable the NVIDIA device plugin with the --fail-on-init-error=false flag and set privileged: true in the pod security context.
AnswerA

The GPU Operator's device plugin exposes GPUs as schedulable resources (nvidia.com/gpu). Enforcing a policy that requires pods to request this resource ensures that only pods explicitly asking for GPUs receive device nodes. Forbidding privileged mode enforces least privilege. Together, these configurations ensure GPU device nodes are not exposed to pods that do not request them and that pods run with minimal privileges.

Why this answer

Using the GPU Operator with the device plugin, combined with a policy that mandates explicit nvidia.com/gpu requests and prohibits privileged containers, ensures that GPU device nodes are only injected into pods that request them and that pods run with minimal privileges. This satisfies both the least privilege and isolation requirements.

Exam trap

The trap here is believing that setting the NVIDIA_VISIBLE_DEVICES environment variable alone is sufficient for GPU access in Kubernetes; in reality, the device plugin must allocate the resource via a resource request, and without it the device nodes are not mounted.

164
MCQmedium

Which mechanism does the NVIDIA Device Plugin use to communicate GPU availability to the Kubernetes Kubelet?

A.It writes to a configuration file that the Kubelet watches.
B.It uses a gRPC-based device plugin API.
C.It queries the API server directly via REST calls.
D.It relies on the container runtime to inspect hardware.
AnswerB

The Kubernetes Device Plugin framework is built on gRPC. The NVIDIA plugin implements this API to provide the Kubelet with the necessary information about GPU resources. This standard interface allows Kubernetes to treat different hardware devices consistently while maintaining the flexibility needed for NVIDIA-specific hardware management.

Why this answer

The NVIDIA Device Plugin functions as a gRPC service that registers itself with the Kubelet upon startup. It performs active monitoring of the GPU nodes and reports discovered devices, including their count and health, to the Kubelet. This information is subsequently relayed to the Kubernetes API server, allowing the scheduler to make placement decisions based on real-time GPU availability and ensuring efficient workload distribution in AI clusters.

Exam trap

Candidates often guess standard REST APIs or custom webhooks instead of the gRPC-based device plugin API used by Kubernetes Kubelet.

165
MCQmedium

An administrator is deploying NVIDIA AI Enterprise on a bare-metal cluster. The workload requires full GPU isolation with minimal latency. Which configuration should the administrator select to achieve this goal?

A.Deploy using NVIDIA vGPU on a KVM hypervisor.
B.Configure the system using NVIDIA License System (NLS) in disconnected mode.
C.Implement bare-metal installation with NVIDIA GPUDirect RDMA enabled.
D.Utilize containerized GPU passthrough with a standard Docker runtime.
AnswerC

GPUDirect RDMA allows direct memory access between the GPU and third-party devices such as NICs, effectively bypassing the host CPU. This architecture eliminates unnecessary data copies and context switches, providing the lowest possible latency for high-speed AI data pipelines and distributed training environments on physical infrastructure.

Why this answer

GPU Direct and bare-metal deployments are essential for latency-sensitive workloads. By avoiding hypervisor overhead, the administrator ensures direct path access to the GPU memory and interconnects. This configuration is critical in high-performance computing environments where jitter and interrupt latency can degrade model training performance.

Selecting the right deployment mode is a foundational step in AI Operations to ensure optimal hardware utilization and predictable execution times for deep learning models.

Exam trap

Candidates often incorrectly choose virtualization or container-only solutions, failing to recognize that 'bare-metal' and 'GPUDirect RDMA' are the specific, non-negotiable requirements for minimizing latency in high-performance computing environments.

166
MCQeasy

A data scientist reports that a Jupyter notebook running on a DGX station cannot allocate GPU memory, even though other users' jobs are running fine. The notebook kernel was started before a system administrator updated the NVIDIA driver and rebooted the node. Which action should the data scientist take to resolve the issue?

A.Run `nvidia-smi --gpu-reset` to reset the GPU and clear any stale contexts.
B.Set the environment variable `CUDA_VISIBLE_DEVICES=0` to force the notebook to use a specific GPU.
C.Restart the Jupyter kernel to reinitialize CUDA and pick up the new driver.
D.Reinstall the NVIDIA driver using the `.run` installer without rebooting.
AnswerC

When the NVIDIA driver is updated and the system reboots, any existing processes that had already initialized CUDA hold references to the old driver. The Jupyter kernel is such a process. Restarting the kernel terminates the old process and starts a new one that loads the updated driver, allowing CUDA calls to succeed and GPU memory to be allocated.

Why this answer

After a driver update and reboot, any process that started before the update holds an outdated CUDA context. The Jupyter kernel is such a process, so it cannot use the new driver. Restarting the kernel creates a fresh process that loads the updated driver and can allocate GPU memory normally, resolving the issue without affecting other users.

Exam trap

The trap here is assuming that a GPU reset or environment variable change is needed, when simply restarting the long-running process that predates the driver update is sufficient.

167
MCQmedium

Refer to the exhibit. An administrator notices poor performance in an AI training job. What is the most likely cause based on the CLI output?

A.The GPU is overheating due to a fan failure.
B.The power limit is set too low for the current workload.
C.The GPU driver is corrupted or out of date.
D.The workload is waiting for CPU memory allocation.
AnswerB

The 'Sw Power Cap' status confirms that the GPU is limited by the current software configuration. To improve performance, the administrator should evaluate the power policy settings to determine if the wattage limit can be safely increased to allow the GPU to reach its maximum boost clock frequency.

Why this answer

The output indicates that the GPU is currently throttling due to 'Sw Power Cap'. This means the software-defined power limit is restricting the GPU performance to stay within a specific wattage budget. This is common in densely packed servers or cloud environments where power infrastructure is shared.

Monitoring power usage is vital because it directly impacts clock speeds, which in turn bottleneck training throughput and lengthen the time required for model convergence.

Exam trap

Candidates often mistake power capping errors for hardware faults or thermal overheating issues, ignoring the explicit 'Sw Power Cap' message in the telemetry output.

168
MCQmedium

An operations team is troubleshooting a multi-node NCCL training job on an NVIDIA DGX SuperPOD that stalls at initialization and never begins gradient exchange. Running `nccl-tests` with `all_reduce_perf` on the same nodes fails identically, but single-node `all_reduce_perf` succeeds. Which action should the team take first to isolate the fault?

A.Set NCCL_DEBUG=INFO and NCCL_DEBUG_SUBSYS=INIT,COLL,NET on the launcher and inspect which transport (NET/IB or P2P) is selected and where the hang occurs.
B.Disable GPUDirect RDMA by setting NCCL_NET_GDR_LEVEL=0 across all nodes to force host-memory staging.
C.Increase NCCL_BUFFSIZE to 16 MB on every rank and rerun the job to rule out buffer exhaustion.
D.Reduce the number of ranks per node so each process has a dedicated GPU, then rerun the multi-node test.
AnswerA

NCCL_DEBUG=INFO with NCCL_DEBUG_SUBSYS=INIT,COLL,NET surfaces transport selection, ring/tree topology construction, and per-rank connection progress. Because single-node all_reduce_perf passes but multi-node fails, the fault is almost certainly in the inter-node path, and this logging pinpoints whether NCCL fell back to a broken interface or is stuck establishing connections.

Why this answer

Because single-node all_reduce_perf succeeds while multi-node fails, the failure is in inter-node communication setup rather than GPU-local collectives. Enabling NCCL_DEBUG=INFO with INIT, COLL, and NET subsystems exposes which transport NCCL selects, how rings and trees are built, and where connection establishment stalls, giving the team actionable evidence before any configuration change.

Exam trap

The trap here is assuming a collective hang is a buffer-size or GPU-memory problem, when a job that never starts exchanging gradients is actually stuck in NCCL transport and topology initialization.

169
Multi-Selecthard

An administrator is configuring NVIDIA GPUDirect Storage (GDS) on a cluster to accelerate data loading for AI training jobs. The cluster uses Mellanox InfiniBand adapters and NVMe storage. Which two actions are required to enable GDS and ensure optimal performance? (Choose two.)

Select 2 answers
A.Ensure that the storage devices and network adapters support RDMA and are properly configured for peer-to-peer communication.
B.Set the environment variable NVIDIA_GDS_ENABLE=1 on all nodes.
C.Configure the GPU nodes to use the NVIDIA Container Toolkit for all training containers.
D.Disable IOMMU in the BIOS to allow direct memory access between devices.
E.Install the NVIDIA GPUDirect Storage kernel module and user-space libraries on all GPU nodes.
AnswersA, E

GDS relies on RDMA and peer-to-peer (P2P) communication to transfer data directly between storage and GPU memory without CPU involvement. The storage devices (NVMe over Fabrics) and network adapters (InfiniBand) must support RDMA, and the system must be configured to allow P2P DMA. This includes enabling PCIe Access Control Services (ACS) and IOMMU settings that permit direct peer-to-peer transfers. Without this, GDS cannot achieve direct data paths.

Why this answer

Enabling GPUDirect Storage requires installing the GDS kernel module and user-space libraries on GPU nodes, and ensuring that storage and network hardware support RDMA and peer-to-peer communication. These two actions create the necessary software and hardware foundation for direct data transfers between storage and GPU memory. Other options are either incorrect or not specific to GDS.

Exam trap

The trap here is assuming that GDS can be enabled by a simple environment variable or that container toolkit configuration is sufficient, when it actually requires low-level driver and hardware configuration.

170
MCQhard

An AI operations engineer is troubleshooting a multi-GPU training job that fails intermittently with a NCCL error: 'unhandled system error'. The job runs on a DGX-1 with eight V100 GPUs connected via NVLink. Which step should the engineer take first to resolve the issue?

A.Disable NVLink by setting NCCL_P2P_DISABLE=1 to force PCIe communication.
B.Reinstall the NVIDIA driver and CUDA toolkit on all nodes.
C.Reduce the number of GPUs used in the job to four to see if the error disappears.
D.Set NCCL_DEBUG=INFO and reproduce the failure to capture detailed logs.
AnswerD

NCCL debug logs provide detailed information about the communication setup, including which transports are used, any fallbacks, and specific error codes. This is the most direct way to diagnose the cause of an 'unhandled system error', which can stem from hardware, driver, or configuration issues. Capturing logs during failure is the essential first step before attempting fixes.

Why this answer

The first step in troubleshooting any NCCL error should be to gather detailed logs using NCCL_DEBUG=INFO. This provides visibility into the communication path and error specifics, enabling targeted fixes. Other options either mask the issue, involve unnecessary reinstallation, or reduce resources without diagnosis.

Exam trap

The trap here is jumping to hardware workarounds or reinstallations without first collecting diagnostic information that NCCL can provide.

171
MCQmedium

When designing a workload management strategy for multi-tenant AI training, what is the most effective way to ensure isolation between different tenants using the same physical GPU nodes?

A.Set strict Kubernetes memory limits on every pod to prevent memory leakage.
B.Implement NVIDIA MIG to partition hardware at the compute and memory level.
C.Use Kubernetes namespaces to logically separate the tenants.
D.Enable GPU sharing through the NVIDIA Container Toolkit's time-slicing configuration.
AnswerB

MIG provides spatial and temporal hardware partitioning, ensuring that individual workloads have their own dedicated hardware paths. This prevents performance degradation caused by noisy neighbors and provides security by isolating memory and compute resources, which is critical for multi-tenant clusters that require predictable performance and isolation.

Why this answer

Using NVIDIA Multi-Instance GPU (MIG) is the most robust method for hardware-level isolation. MIG partitions a single physical GPU into multiple independent instances, each with its own memory and compute resources. This allows multiple tenants to run workloads simultaneously on the same hardware without interfering with each other's performance, ensuring predictable and secure resource allocation for multi-tenant AI environments.

Exam trap

Candidates often suggest software-level resource limits (like cgroups) for GPU isolation, which do not provide the same level of hardware-enforced memory and compute partitioning as NVIDIA MIG.

172
MCQhard

An operations engineer is investigating a sudden drop in throughput for a multi-GPU training job on an NVIDIA DGX A100. The job uses PyTorch with DDP. Logs show that one GPU is consistently at 100% utilization while others are below 50%. Which tool and approach should be used to identify the bottleneck?

A.Use nvidia-smi to monitor GPU utilization and memory, then rebalance the workload by reducing the batch size on the overutilized GPU.
B.Use NVIDIA Nsight Systems to profile the job and examine the CUDA API and kernel timeline to identify serialization or synchronization issues.
C.Use NVIDIA Nsight Compute to profile each GPU's kernels and compare execution times to find the slowest kernel.
D.Use nvidia-smi topo -m to check the GPU topology and then adjust the NCCL communication settings to use a different ring order.
AnswerB

Nsight Systems provides a detailed timeline of CPU and GPU activity, showing kernel executions, memory copies, and synchronization points. It can reveal if one GPU is waiting on data or if there is an imbalance in computation. This is the correct tool to pinpoint the bottleneck in a multi-GPU job, as it captures cross-GPU interactions and can highlight stragglers.

Why this answer

Nsight Systems is designed for system-level profiling, capturing CPU and GPU activities across multiple devices. It can show if one GPU is performing more work or if others are idle waiting for synchronization. This makes it the right tool to identify the bottleneck in a DDP job where one GPU is overutilized.

Profiling before making changes is essential.

Exam trap

The trap here is choosing a profiling tool that is too granular, like Nsight Compute, or a monitoring tool that lacks detail, like nvidia-smi, instead of the system-level profiler needed for this multi-GPU imbalance.

173
MCQeasy

During an NVIDIA AI Enterprise deployment, you are asked to configure the 'NVIDIA Container Toolkit'. What is its primary function?

A.To provide high-level APIs for neural network training.
B.To allow the container runtime to interact with the host's GPU.
C.To act as a package manager for AI model weight files.
D.To automatically optimize neural network hyper-parameters.
AnswerB

The Container Toolkit provides the necessary hooks and libraries to map GPU resources into the container namespace. This allows the application running inside the container to make calls to the GPU hardware, which would otherwise be inaccessible due to the isolation boundaries of the container runtime environment.

Why this answer

The Container Toolkit enables containers to access the GPU by exposing the necessary drivers, libraries, and device files from the host into the container runtime. It acts as the critical bridge between the hardware-level drivers and the containerized applications. Understanding this is foundational for AI Ops, as it ensures that containerized AI models can actually utilize the GPU hardware for training or inference tasks without manual configuration.

Exam trap

Candidates often describe the Toolkit as a library installer for the container, rather than its primary role as the runtime interface enabling GPU resource access.

174
MCQmedium

Refer to the exhibit. An AI engineer observes that a model training job is running slower than expected. Based on the output, what is the primary cause of the performance degradation?

A.The GPU is overheating and the thermal management system has engaged to prevent hardware damage.
B.The GPU is currently idle and the driver has shifted the card to an energy-saving state.
C.An administrator has set a software power limit that is lower than the GPU's maximum performance threshold.
D.The GPU is experiencing a PCIe bus error, forcing the system to reduce clock speeds to maintain stability.
AnswerC

The 'SW Power Cap' indicator explicitly confirms that an administrative policy or software command has capped the power usage of the GPU. This forces the GPU to maintain lower clock speeds, directly impacting the throughput of the training job by limiting the available computational resources.

Why this answer

The 'SW Power Cap' throttle reason indicates that the power limit is configured below the card's maximum design capacity, causing the GPU to downclock to P12 state to stay within the power envelope. This is crucial for AI operations because it signals that the hardware is being artificially throttled, necessitating a review of power policies to ensure optimal training performance in high-compute scenarios.

Exam trap

Candidates frequently mistake power throttling for thermal throttling. They look at temperature metrics instead of examining the specific 'Power Cap' status flag, which indicates an administrative software limit is active.

175
MCQeasy

An AI operations team is troubleshooting a training job that crashes with a segmentation fault after several hours. The job uses multiple GPUs and NCCL for communication. System logs show no errors, but dmesg reveals repeated 'NVRM: Xid' errors. Which action should be taken first to diagnose the issue?

A.Restart the NVIDIA driver with rmmod and modprobe commands.
B.Check the Xid error code in the NVIDIA documentation to identify the specific GPU fault.
C.Enable NCCL_DEBUG=INFO and re-run the job to capture detailed NCCL logs.
D.Run nvidia-smi -q to check GPU temperature and power usage.
AnswerB

Xid errors are reported by the NVIDIA driver and indicate GPU hardware or driver issues. Each Xid code corresponds to a specific fault, such as a corrupted memory access or a fallen off the bus error. Looking up the code in NVIDIA's documentation provides the exact cause and recommended actions. This is the most direct first step to diagnose the segmentation fault linked to GPU errors.

Why this answer

Xid errors are logged by the NVIDIA driver and each code maps to a specific GPU fault. Interpreting the Xid code is the fastest way to understand whether the segmentation fault is due to a hardware issue, driver bug, or application error. This guides subsequent troubleshooting steps, such as replacing hardware or updating the driver.

Exam trap

The trap here is jumping to application-level debugging tools like NCCL_DEBUG when the system logs already point to a GPU-level fault via Xid errors.

176
MCQmedium

A machine learning engineer is optimizing a recommendation model for inference on an NVIDIA T4 GPU. The model uses dynamic input shapes, and profiling shows that kernel launch overhead is a significant contributor to latency. Which optimization technique should be applied to reduce this overhead?

A.Convert the model to TensorRT with dynamic shapes and enable CUDA graphs during inference.
B.Use mixed precision (FP16) to reduce the number of kernels executed.
C.Enable NVIDIA MPS (Multi-Process Service) to allow concurrent kernel execution.
D.Increase the batch size to amortize kernel launch overhead across more samples.
AnswerA

CUDA graphs capture a sequence of kernel launches and replay them with a single launch, drastically reducing launch overhead. TensorRT supports dynamic shapes and can build engines that use CUDA graphs. This is ideal for models with variable input shapes where launch overhead is high. The combination reduces CPU overhead and improves latency.

Why this answer

CUDA graphs reduce kernel launch overhead by capturing a sequence of operations and replaying them as a single graph. TensorRT with dynamic shapes can leverage CUDA graphs to maintain flexibility while minimizing overhead. This directly addresses the profiling finding, making it the most effective optimization for the described scenario.

Exam trap

The trap here is assuming that mixed precision or larger batch sizes reduce launch overhead, but they target different bottlenecks and do not minimize the number of CPU-GPU interactions.

177
MCQeasy

When deploying NVIDIA containers using the NVIDIA Container Toolkit, what is the primary function of the 'nvidia-container-runtime'?

A.It automatically recompiles the application code for the specific GPU architecture found on the host.
B.It manages the lifecycle of the GPU driver installation on the host operating system.
C.It exposes the host's NVIDIA GPUs and driver libraries to the container environment.
D.It monitors the temperature and power consumption of the GPUs during container execution.
AnswerC

This runtime transparently mounts the host GPU devices and user-mode NVIDIA libraries into the container. It modifies the container's OCI specification during the execution phase, ensuring that the containerized process can communicate with the physical GPU drivers installed on the host operating system.

Why this answer

The nvidia-container-runtime is a crucial component that allows Docker containers to interface with host GPUs. By modifying the container runtime specification, it ensures that the necessary device nodes and NVIDIA driver libraries are injected into the container namespace at startup. This enables seamless hardware acceleration for AI applications without requiring users to manually manage drivers or complex device path configurations inside their container images.

Exam trap

Candidates frequently confuse the container runtime's role with orchestration components like the device plugin, failing to realize the runtime directly injects driver libraries and device nodes into the container namespace.

178
MCQhard

When debugging a workload that consistently crashes with 'Out of Memory' (OOM) errors despite sufficient GPU VRAM, what is the most likely cause related to workload management?

A.The CUDA driver version is incompatible with the installed GPU.
B.The container memory limit is lower than the data processing requirements.
C.The GPU device plugin is not correctly reporting free memory.
D.The training job is using mixed-precision training (FP16).
AnswerB

Workloads often require substantial system memory for preprocessing before moving data to the GPU. If the container memory limit is exceeded, the orchestrator will terminate the pod with an OOM error, regardless of how much GPU VRAM is available. This is a common oversight when configuring resource limits.

Why this answer

OOM errors can occur due to host-side system memory exhaustion if the workload manages large datasets in system RAM before loading them into the GPU. If the container memory limit is set too low for the data processing pipeline, the entire container will be terminated. This highlights the need to correctly balance both GPU VRAM and system memory limits in the container specification.

Exam trap

Candidates mistakenly assume that OOM errors on GPU nodes are always caused by insufficient VRAM, completely ignoring container system memory limits during data preprocessing.

179
MCQmedium

An AI operations team is installing the NVIDIA GPU Operator on a Kubernetes cluster that uses a custom containerd configuration. They need to ensure that the GPU Operator can properly manage the container runtime. Which action should they take before installing the GPU Operator?

A.Label the nodes with the appropriate container runtime version.
B.Set the default runtime in containerd to nvidia-container-runtime.
C.Disable the containerd systemd service and let the GPU Operator start its own runtime.
D.Ensure that the containerd configuration does not already include conflicting NVIDIA runtime settings.
AnswerD

The GPU Operator manages the NVIDIA Container Toolkit and runtime configuration. If containerd already has manually added NVIDIA runtime settings, these can conflict with the operator's configuration, leading to failures. It is best practice to remove any existing NVIDIA runtime configuration before installation so the operator can set it up cleanly.

Why this answer

Before installing the GPU Operator, any pre-existing NVIDIA runtime configuration in containerd should be removed to avoid conflicts. The operator manages the runtime configuration itself, so manual settings can interfere. Other options like changing the default runtime or disabling containerd are incorrect because the operator integrates with the existing runtime rather than replacing or requiring manual overrides.

Exam trap

The trap here is assuming the GPU Operator requires manual runtime configuration, when in fact it manages the runtime and conflicts can arise from pre-existing settings.

180
MCQmedium

A production inference service on an NVIDIA A100 GPU experiences a gradual increase in latency over several hours, eventually requiring a pod restart. GPU memory utilization climbs steadily, but the model and batch size are fixed. Which action should an AI operations engineer take first to diagnose the root cause?

A.Use PyTorch's torch.cuda.memory_summary() or TensorFlow's memory profiler to capture allocation snapshots and compare over time.
B.Run nvidia-smi --query-gpu=memory.used --format=csv -l 1 to log memory usage over time and correlate with request rate.
C.Enable CUDA memory leak detection with compute-sanitizer --tool memcheck on the running inference process.
D.Increase the GPU memory limit in the Kubernetes pod spec to prevent the pod from being killed.
AnswerA

Framework-level memory profilers show exactly which tensors or operations are allocating memory and whether those allocations are freed. By taking snapshots at intervals, an engineer can see if memory is retained across inference calls, indicating a leak in the application code or framework caching allocator. This directly identifies the source without disrupting the service.

Why this answer

The steady memory increase with fixed workload strongly suggests a memory leak in the inference application or framework. Framework-specific memory profilers provide allocation-level visibility, allowing engineers to identify unreleased tensors or cached allocations. Aggregate GPU monitoring or sanitizer tools lack the necessary granularity.

Increasing limits merely postpones failure.

Exam trap

The trap here is assuming that nvidia-smi memory monitoring is sufficient to diagnose leaks, when it only shows aggregate usage without per-process or allocation detail.

181
MCQhard

An administrator manages a cluster where inference services and batch training share the same GPU nodes. During business hours, inference pods must be scheduled promptly, while training jobs can wait. The administrator wants preemption so that a pending high-priority inference pod can evict a lower-priority training pod when no GPU is free, with evicted training resuming later. Which configuration achieves this?

A.Set the training pods' terminationGracePeriodSeconds to zero and add a PodDisruptionBudget so the scheduler can remove them immediately when capacity is needed.
B.Apply node affinity rules that pin inference pods to dedicated nodes and training pods to separate nodes, then let each workload schedule only within its own pool.
C.Create a PriorityClass with a high value for inference and a lower one for training, reference them in the pod specs, and enable preemption in the scheduler so lower-priority pods are evicted when needed.
D.Define a ResourceQuota on the training namespace capping GPU requests so that free capacity always remains available for inference pods.
AnswerC

Kubernetes priority and preemption let a pending pod with a higher PriorityClass displace running pods of lower priority when resources are scarce. Assigning a high class to inference and a lower class to training makes the scheduler evict a training pod to admit the inference pod. Combined with a group-aware training controller, the evicted job can be requeued and resume from checkpoint.

Why this answer

Priority and preemption are the native scheduling features designed for exactly this pattern: a higher-priority pending pod triggers eviction of lower-priority running pods to free capacity. Defining distinct PriorityClasses for inference and training and referencing them in pod specs enables the scheduler to make that decision. Because the training workload is managed by a group-aware controller, the evicted job can be requeued and restarted from its last checkpoint.

Exam trap

The trap here is confusing admission-time controls such as quotas and affinity with runtime preemption; only priority classes cause a pending pod to evict a running one, while quotas and affinity merely shape where or whether pods are admitted.

182
MCQeasy

An AI operations engineer is troubleshooting a model inference service deployed with NVIDIA Triton Inference Server on a GPU. The service occasionally returns incorrect predictions, and the engineer suspects that the input data is not being preprocessed correctly. The model expects input tensors in FP32 format, but the client is sending FP16 data. Which action should the engineer take to resolve the issue?

A.Enable dynamic batching in Triton to automatically convert FP16 inputs to FP32.
B.Set the model to use FP16 precision to match the client's input data type.
C.Increase the instance count to handle the load and reduce the chance of data corruption.
D.Configure the Triton model's input to specify the correct data type (FP32) and ensure the client sends matching data.
AnswerD

Triton validates input data types against the model's configuration. If the model expects FP32 but receives FP16, it may either reject the request or misinterpret the data, leading to incorrect predictions. Setting the correct data type in the model configuration and aligning the client ensures proper preprocessing and accurate inference.

Why this answer

The root cause is a mismatch between the client's data type (FP16) and the model's expected input type (FP32). Triton requires that input tensors match the model's configuration. By setting the correct data type in the model configuration and ensuring the client sends FP32 data, the engineer ensures proper data handling and restores prediction accuracy.

Exam trap

The trap here is assuming that Triton automatically handles data type conversions, when in fact it enforces strict type matching based on the model configuration.

183
Multi-Selectmedium

An administrator is tuning a Kubernetes cluster that runs GPU Operator. Users report that GPU jobs are sometimes scheduled onto nodes whose drivers are older than the CUDA version the container needs, causing runtime failures. The administrator wants to prevent incompatible placements before pods are bound. (Choose two.)

Select 2 answers
A.Rely on the device plugin to compare the container's CUDA version against the node driver and refuse to allocate the GPU when they mismatch.
B.Enable time-slicing with a replica count high enough that every node advertises spare GPU capacity, so incompatible nodes are simply bypassed.
C.Set the container image's CUDA version to match the newest driver in the cluster and rely on the container runtime to downgrade the driver automatically.
D.Advertise the driver and CUDA versions as node labels through the GPU Operator's node feature discovery, then use nodeAffinity on the pods to require a compatible version.
E.Use a validating admission policy that rejects pods requesting a CUDA version incompatible with the target node's advertised driver labels.
AnswersD, E

Node feature discovery running under the GPU Operator publishes labels such as nvidia.com/cuda.driver.major and minor versions. Pods can then express a nodeAffinity requirement matching the needed driver generation, so the scheduler only binds them to nodes whose advertised driver satisfies the container's CUDA expectation. This moves incompatibility detection to scheduling time, preventing the runtime failures users currently see.

Why this answer

Preventing incompatible placement requires exposing driver and CUDA versions as schedulable node attributes and then constraining pods against them. Node feature discovery under the GPU Operator publishes version labels, and nodeAffinity lets pods require a compatible generation. A validating admission policy adds a second layer by rejecting pods whose declared CUDA needs exceed the target node's advertised driver, catching mistakes before binding and avoiding the runtime failures users report.

Exam trap

The trap here is expecting the device plugin or container runtime to reconcile CUDA and driver versions, when that compatibility decision belongs to scheduling and admission, not device allocation.

184
MCQmedium

When configuring the NVIDIA Device Plugin for Kubernetes, what is the purpose of the 'time-slicing' configuration?

A.To schedule tasks based on the specific time of day for load balancing.
B.To increase the GPU's clock frequency during high-demand periods.
C.To allow multiple pods to share a single physical GPU through context switching.
D.To restrict access to the GPU based on container process priority.
AnswerC

Time-slicing enables a physical GPU to be oversubscribed by allowing multiple pods to execute in turns on the same hardware. This increases the utilization of the GPU in scenarios where the individual pods do not require the full dedicated throughput of the entire device at all times.

Why this answer

Time-slicing is a technique that allows multiple Kubernetes pods to share a single GPU by rapidly switching context between them. In scenarios where full GPU isolation is not required, this increases resource utilization. It is a vital deployment strategy for optimizing cost and efficiency in shared environments, allowing administrators to balance the workload across limited GPU hardware without needing more expensive virtual machine-based partitioning solutions.

Exam trap

Candidates confuse time-slicing with hardware partitioning (MIG). Time-slicing is software-based context switching, whereas MIG provides true hardware-level isolation of GPU resources.

185
MCQmedium

A company runs multiple AI workloads on a shared Kubernetes cluster with NVIDIA GPUs. The administrator needs to enforce that only pods with a specific label can consume GPU resources, while other pods are denied. Which Kubernetes admission control mechanism should be used to implement this policy?

A.PodSecurityPolicy
B.ValidatingAdmissionWebhook
C.ResourceQuota
D.LimitRange
AnswerB

A ValidatingAdmissionWebhook can intercept pod creation requests and evaluate custom logic, such as checking for a specific label before allowing the pod to request GPU resources. This provides the flexibility to enforce label-based policies on extended resources. By deploying a webhook that inspects pod labels and resource requests, the administrator can deny non-compliant pods, meeting the requirement.

Why this answer

A ValidatingAdmissionWebhook allows custom admission logic to be applied to pod creation. By configuring a webhook that checks for a specific label before permitting GPU resource requests, the administrator can enforce label-based access control. This is the correct mechanism because it can inspect pod metadata and resource requests and reject non-compliant pods.

Exam trap

The trap here is assuming that ResourceQuota can enforce per-pod label conditions, when it only limits aggregate namespace consumption.

186
MCQmedium

During a training job, the system reports "NCCL WARN" regarding a slow network path. What is the most likely culprit for this performance bottleneck in a multi-node InfiniBand environment?

A.The GPU clock speed is set too low.
B.InfiniBand link speed is negotiated at a lower rate than expected.
C.The batch size is too small.
D.The CPU is running at maximum capacity.
AnswerB

If an InfiniBand link negotiates at a lower rate (e.g., SDR instead of HDR), the network throughput will be severely throttled. This mismatch is a classic cause of "slow path" NCCL warnings, as the collective communication operations are unable to achieve the expected bandwidth required for efficient multi-node training.

Why this answer

In high-performance clusters, the network fabric is often the bottleneck. InfiniBand performance relies on proper subnet manager configuration and accurate link-speed reporting. Identifying "slow paths" using tools like ibdiagnet helps pinpoint physical or configuration issues in the interconnect.

Addressing these issues is vital because, in distributed training, the speed of the cluster is limited by the slowest link in the communication path, severely impacting the overall training efficiency.

Exam trap

Candidates often blame the GPU driver or the training code itself, overlooking that InfiniBand interconnects are physical network layers that can negotiate down to lower speeds due to cabling issues.

187
MCQmedium

A platform team runs mixed training and inference workloads on a Kubernetes cluster with the NVIDIA GPU Operator. Inference pods are latency-sensitive and must not be preempted, while training pods can be interrupted and restarted. The team wants training jobs to yield GPUs to inference jobs when capacity is scarce, without manual intervention. Which Kubernetes mechanism should the team configure to achieve this behavior?

A.A PodDisruptionBudget on the training pods that guarantees a minimum number of running replicas.
B.PriorityClass with preemptionPolicy set to PreemptLowerPriority on the inference pods.
C.ResourceQuota on the training namespace limiting total GPU requests below cluster capacity.
D.A node affinity rule on the inference pods that targets nodes with the highest available GPU memory.
AnswerB

PriorityClass assigns a numeric priority, and when a high-priority pod cannot schedule, the scheduler may evict lower-priority pods on a node to make room. Setting the inference pods to a higher priority with preemption enabled lets them displace training pods, which can restart. This directly implements automatic yielding of GPUs to latency-sensitive inference workloads without manual operator action.

Why this answer

Priority and preemption are the scheduler features designed for exactly this pattern. Assigning inference pods a higher-priority class with preemption enabled lets the scheduler evict lower-priority training pods when a node lacks free GPUs. Because training pods are restartable, the disruption is acceptable, and inference keeps its latency guarantees without manual intervention.

Exam trap

The trap here is confusing PodDisruptionBudget, which limits voluntary evictions, with preemption, which actively evicts lower-priority pods to place a higher-priority one.

188
Multi-Selectmedium

An administrator is using NVIDIA Base Command Manager to manage a cluster with a mix of GPU and CPU nodes. They need to ensure that a newly added GPU node is correctly recognized and that jobs can be scheduled on it. Which TWO actions must be performed to integrate the new node into the Base Command Manager cluster? (Choose two.)

Select 2 answers
A.Install the Base Command Manager agent on the new node and register it with the head node.
B.Define the node in the Base Command Manager cluster configuration and assign it to a node group or category.
C.Add the node's hostname and IP address to the /etc/hosts file on all cluster nodes.
D.Configure a static IP address for the node and add it to the cluster's DNS server.
E.Manually install the NVIDIA GPU driver on the node without using Base Command Manager's provisioning tools.
AnswersA, B

The Base Command Manager agent is required on each node to communicate with the head node, report status, and receive commands. Installing and registering the agent allows the head node to recognize the new node, manage its resources, and include it in the cluster's scheduling pool. Without the agent, the node remains unmanaged and cannot be utilized for jobs.

Why this answer

Integrating a new node into a Base Command Manager cluster requires installing and registering the Base Command Manager agent so the head node can manage it, and defining the node in the cluster configuration with an appropriate node group assignment so it inherits policies and becomes schedulable. These two actions together ensure the node is recognized, provisioned correctly, and available for job scheduling.

Exam trap

The trap here is assuming that manual network configuration or driver installation is necessary, when Base Command Manager handles these through its own provisioning and management framework.

189
MCQmedium

During an air-gapped installation of the NVIDIA GPU Operator, the administrator must make all required images available to the cluster. Which component is responsible for pulling the Operator's operand images from the private registry?

A.The NVIDIA Container Toolkit must be pointed at the private registry through its config.toml file.
B.The Operator's Helm chart values must include an imagePullSecret for the private registry on every namespace.
C.The kubelet on each node must be configured with a mirror registry that rewrites all container image pulls.
D.The GPU Operator's ClusterPolicy must reference the private registry so operand pods use images from that registry.
AnswerD

In an air-gapped environment, the ClusterPolicy is configured with the private registry path and image repository settings for each operand, such as driver, toolkit, device plugin, and DCGM exporter. This ensures the Operator deploys pods that pull from the internal registry rather than the public NVIDIA registry, which is unreachable.

Why this answer

Air-gapped deployments require the ClusterPolicy to specify the private registry and repository paths for each operand image. This directs the Operator to deploy pods that pull driver, toolkit, device plugin, and monitoring images from the internal registry. Kubelet mirrors, toolkit configuration, and image pull secrets address adjacent concerns but do not select the operand image sources.

Exam trap

The trap here is assuming a kubelet mirror or pull secret is sufficient to redirect operand images, when the ClusterPolicy must explicitly reference the private registry.

190
MCQeasy

Which component in the NVIDIA AI ecosystem is responsible for monitoring and reporting GPU telemetry data, such as power usage, temperature, and utilization, to Prometheus?

A.The NVIDIA Device Plugin.
B.The NVIDIA DCGM Exporter.
C.The Kubernetes Kubelet.
D.The NVIDIA Container Runtime.
AnswerB

The DCGM Exporter is purpose-built to interface with the underlying NVIDIA driver and DCGM to gather detailed telemetry. It exposes these metrics as Prometheus-formatted data, which is essential for monitoring the health and performance of GPU workloads within a cluster and making data-driven infrastructure management decisions.

Why this answer

The NVIDIA DCGM Exporter is the standard component designed to collect GPU metrics via the Data Center GPU Manager (DCGM) and export them into a format that Prometheus can scrape. This is vital for workload management because it provides the real-time observability required to trigger auto-scaling or identify jobs that are failing to utilize allocated GPU resources effectively across the cluster.

Exam trap

Candidates often confuse the exporter with the DCGM agent itself. The agent collects data, but the exporter is the specific component that translates it for Prometheus scraping.

191
MCQmedium

A distributed training job on a multi-GPU node is exhibiting poor scaling efficiency: each GPU shows high utilization, but overall throughput increases by only 15% when doubling the number of GPUs. The job uses NCCL for communication. Which diagnostic step is most appropriate to identify the bottleneck?

A.Enable NVIDIA MPS (Multi-Process Service) to allow multiple processes to share the GPUs more efficiently.
B.Check the GPU clock speeds with `nvidia-smi -q -d CLOCK` to ensure they are running at maximum frequency.
C.Increase the batch size per GPU to reduce the frequency of gradient synchronization across GPUs.
D.Run `nvidia-smi topo -m` to inspect the GPU interconnect topology and verify that NCCL is using the fastest available path between GPUs.
AnswerD

Poor scaling with high GPU utilization often indicates communication overhead. `nvidia-smi topo -m` reveals whether GPUs are connected via NVLink or PCIe and whether data traverses slower paths like QPI/UPI. If NCCL cannot leverage NVLink, all-reduce operations become the bottleneck. This command is the standard first step to diagnose interconnect issues and ensure NCCL selects the optimal topology-aware route, directly addressing the scaling inefficiency.

Why this answer

Poor scaling efficiency in distributed training often results from communication bottlenecks. The `nvidia-smi topo -m` command provides a matrix of GPU interconnect paths, helping verify whether NCCL can use NVLink or is forced over slower PCIe/QPI links. If the topology shows suboptimal paths, adjusting NCCL environment variables or hardware configuration can improve scaling.

Thus, inspecting the topology is the most direct diagnostic step.

Exam trap

The trap here is assuming that high GPU utilization guarantees optimal scaling, overlooking that inter-GPU communication overhead can dominate even when GPUs appear busy.

192
MCQmedium

An AI engineer is optimizing a real-time inference pipeline on an NVIDIA A100 GPU. The model uses dynamic input shapes, and profiling shows that the GPU spends significant time on memory copies between host and device. Which optimization should be implemented to reduce this overhead?

A.Enable Unified Memory to automatically manage data migration.
B.Increase the batch size to amortize the cost of memory copies.
C.Use pinned (page-locked) host memory for input and output buffers.
D.Use CUDA streams to overlap memory copies with computation.
AnswerC

Pinned memory allows the GPU to directly access host memory via DMA, avoiding the overhead of staging through pageable memory. This reduces the time spent on host-to-device and device-to-host copies, which is critical for real-time inference with dynamic shapes where copies are frequent. It is a standard optimization to improve transfer efficiency and lower latency.

Why this answer

The overhead from host-device memory copies is best reduced by using pinned memory, which enables faster DMA transfers. This is especially important for real-time inference with dynamic shapes where copies are frequent. While CUDA streams can overlap transfers with compute, they are more effective when combined with pinned memory.

Exam trap

The trap here is assuming that Unified Memory or CUDA streams alone will solve the copy overhead, when the fundamental issue is pageable memory causing slow transfers.

193
MCQmedium

An AI platform engineer is preparing a fleet of NVIDIA DGX H100 systems for production workloads using the NVIDIA Base Command Manager (BCM). During initial bare-metal provisioning via PXE boot, the provisioning server successfully hands out IP addresses, but nodes consistently fail during the OS image deployment phase, throwing a kernel panic related to missing storage drivers. Which deployment step must be verified or corrected to ensure successful hardware-specific image deployment?

A.Reconfigure the DHCP server scope options to extend lease times and include custom vendor-class identifiers for the provisioning daemon.
B.Update the system BIOS Unified Extensible Firmware Interface boot order to prioritize internal redundant array of independent disks storage over network interfaces.
C.Verify that the target operating system image profile assigned in Base Command Manager includes the necessary storage controller modules and hardware support packages for the specific server model.
D.Modify the dynamic port forwarding settings on the top-of-rack management switches to allow uninterrupted trivial file transfer protocol block size extensions.
AnswerC

A kernel panic citing missing storage drivers means the assigned OS image profile lacks the storage controller modules and hardware support packages required by that server model. Correcting the image profile in Base Command Manager lets the deployed kernel detect the disks.

Why this answer

NVIDIA Base Command Manager relies heavily on matching the hardware profile of complex servers like DGX H100 with the correct custom software image and specialized storage drivers. If the deployment image lacks the required kernel modules for modern high-performance NVMe controllers, PXE booting fails at the kernel load stage. Ensuring the software image profile incorporates the exact hardware-specific drivers resolves this mismatch, allowing enterprise-grade bare-metal automation to complete seamlessly across the GPU cluster.

Exam trap

Candidates often assume standard Linux server golden images work universally across NVIDIA DGX hardware without customizing kernel modules for specialized high-performance storage and networking controllers.

194
Multi-Selecthard

Which THREE factors must be considered when sizing a persistent storage solution for multi-node distributed training checkpoints?

Select 3 answers
A.Aggregate write bandwidth to avoid stalling the training loop.
B.Total storage capacity to store multiple checkpoint versions.
C.The number of CPU cores on the storage controller.
D.High availability of the storage backend to prevent data loss.
E.The number of users accessing the storage simultaneously.
AnswersA, B, D

Distributed training generates massive amounts of state data simultaneously. If the storage system cannot handle the aggregate write load, the compute nodes will stall, waiting for I/O completion. Sufficient write bandwidth is therefore non-negotiable to maintain the high performance required for large-scale GPU training clusters.

Why this answer

Checkpointing large models requires significant I/O throughput to avoid blocking the training loop, sufficient capacity for multiple historical versions, and high availability to ensure data integrity. By addressing these three factors—throughput, capacity, and reliability—organizations can ensure that the training process remains robust against hardware failures without incurring unnecessary performance penalties during the frequent write cycles required for large-scale model training.

Exam trap

Candidates often overlook aggregate throughput, focusing only on total capacity. In distributed training, having enough space is useless if the write speed is too slow, causing the GPUs to idle.

195
MCQmedium

An administrator is managing an NVIDIA DGX SuperPOD used for large-scale AI training. The cluster uses a Slurm workload manager. The administrator needs to ensure that jobs are scheduled only on nodes with healthy GPUs and that failed GPUs are automatically drained from the pool. Which integration should be configured to achieve this?

A.NVIDIA Base Command Manager with Slurm integration
B.NVIDIA Fleet Command with Slurm edge scheduling
C.NVIDIA Data Center GPU Manager (DCGM) with the Slurm health check plugin
D.NVIDIA Container Toolkit with Slurm's --gres flag
AnswerC

DCGM provides comprehensive GPU health monitoring, and its integration with Slurm via the health check plugin allows automatic detection of unhealthy GPUs. When DCGM identifies a failed GPU, the plugin can drain the node from Slurm, preventing new jobs from being scheduled on it. This directly addresses the requirement to schedule only on healthy GPUs and automatically remove failed ones.

Why this answer

DCGM with the Slurm health check plugin is the correct integration because it continuously monitors GPU health and can automatically drain nodes with failed GPUs from Slurm, ensuring jobs run only on healthy hardware. Other options lack the necessary health monitoring and automatic remediation capabilities.

Exam trap

The trap here is assuming that Base Command Manager or the Container Toolkit alone can handle GPU health monitoring; they manage deployment and resource allocation, but DCGM is specifically designed for health checks and integration with schedulers like Slurm.

196
MCQmedium

An AI operations engineer is investigating an inference service running on NVIDIA Triton Inference Server. Clients report sporadic 500 errors under peak load. The server logs show occasional 'Failed to allocate memory' messages, and `nvidia-smi` shows VRAM nearly full. The service uses dynamic batching with a maximum batch size of 64 and multiple model instances per GPU. Which change is the most appropriate first step to stabilize the service?

A.Reduce the number of model instances per GPU and lower the maximum dynamic batch size to match the available VRAM headroom.
B.Switch the service to use the CPU execution provider for overflow requests when GPU memory is exhausted.
C.Enable `--strict-model-config=false` so Triton can auto-generate the model configuration and allocate memory more efficiently.
D.Increase the GPU memory pool by setting `CUDA_MPS_ACTIVE_THREAD_PERCENTAGE` to 100.
AnswerA

The allocation failures under peak load with near-full VRAM indicate that the configured instances and maximum batch size exceed the memory budget when concurrent requests spike. Reducing instances and capping the batch size lowers the peak memory footprint per inference, restoring headroom so the allocator can satisfy every request instead of failing.

Why this answer

Triton allocates memory per model instance and per active batch execution. When instances and the maximum dynamic batch size are sized for average load, peak concurrency can exceed VRAM and cause allocation failures. Reducing instances and capping batch size brings the peak footprint within budget, letting the scheduler queue requests instead of failing them, which stabilizes error rates under load.

Exam trap

The trap here is assuming that runtime flags or MPS settings can expand available VRAM, when the real issue is a peak memory footprint larger than the device can satisfy.

197
MCQmedium

A production inference service running on NVIDIA GPUs exhibits periodic latency spikes every few minutes, correlating with CPU-side stalls and low GPU utilization during those intervals. Profiling with Nsight Systems shows large gaps between kernel launches and frequent cudaMalloc/cudaFree calls. Which action best addresses the root cause?

A.Switch the model to use TensorRT with FP16 precision to reduce kernel execution time.
B.Enable CUDA Graph capture for the inference sequence and reuse the captured graph for repeated executions.
C.Increase the number of worker threads on the CPU to better overlap preprocessing with GPU execution.
D.Enable Multi-Process Service (MPS) to allow concurrent kernel execution from multiple processes.
AnswerB

CUDA Graphs capture a sequence of kernel launches and memory operations into a single graph, then replay it with minimal CPU overhead. This eliminates repeated cudaMalloc/cudaFree and launch gaps, smoothing latency spikes. It directly targets the CPU-side stalls and low GPU utilization observed, making it the appropriate optimization for this scenario.

Why this answer

The profile shows CPU-side stalls and frequent cudaMalloc/cudaFree, indicating launch and allocation overhead. CUDA Graphs capture the entire sequence and replay it with minimal CPU involvement, eliminating both the allocation churn and the launch gaps. This directly reduces latency spikes and improves GPU utilization, making it the most effective remedy for the described symptoms.

Exam trap

The trap here is assuming that reducing kernel execution time or adding CPU threads will fix latency spikes caused by launch overhead, rather than addressing the orchestration inefficiency itself.

198
MCQeasy

Which component is strictly necessary for managing NVIDIA AI Enterprise licensing across a distributed cluster of nodes?

A.NVIDIA CUDA Toolkit.
B.NVIDIA License System (NLS) Instance.
C.NVIDIA Container Toolkit.
D.NVIDIA Deep Learning GPU Training System (DIGITS).
AnswerB

The NLS instance acts as the centralized point for distributing and validating licenses to nodes in the cluster. It ensures that the enterprise software features are appropriately licensed, providing the necessary reporting and validation required by the AI Enterprise software subscription models for large-scale distributed computing environments.

Why this answer

The NVIDIA License System (NLS) is the centralized authority for managing product entitlements. In a distributed environment, nodes must reach this server to validate their software keys. This component is crucial because it decouples the license management from the individual worker nodes, allowing for flexible scaling and centralized compliance reporting, which are essential in enterprise AI infrastructure where hardware resources are frequently provisioned and decommissioned.

Exam trap

Candidates mistakenly believe that individual offline license keys stored locally on each worker node are sufficient for managing enterprise-wide distributed NVIDIA AI Enterprise deployments.

199
MCQmedium

Refer to the exhibit. The administrator has deployed the GPU Operator, but the node does not show GPU resources. What is the most likely cause?

A.The node is not part of a valid Kubernetes namespace.
B.The NFD service is not installed or configured correctly.
C.The GPU firmware is out of date.
D.The cluster is using an unsupported CPU architecture.
AnswerB

Node Feature Discovery is responsible for scanning the hardware and applying labels such as 'nvidia.com/gpu.present'. Without these labels, the GPU Operator will not know which nodes to target for driver installation, causing the node to remain without detected GPU resources in the Kubernetes API server.

Why this answer

The lack of GPU resources in the node description indicates that the Node Feature Discovery (NFD) or the GPU Operator has not correctly identified or labeled the GPU hardware. This is a common deployment issue where the hardware-to-software handshake fails due to missing labels. Identifying this early is key to ensuring that the Kubernetes scheduler can correctly place AI workloads on nodes equipped with the necessary compute resources.

Exam trap

Candidates often focus on the GPU Operator itself, forgetting that the GPU Operator relies on NFD to detect and label the hardware so the Kubernetes scheduler knows where to place workloads.

200
Multi-Selecthard

A cloud operations team is deploying the NVIDIA GPU Operator in an environment where the Kubernetes control plane cannot reach the public internet, but worker nodes can access an internal HTTP registry that mirrors required images. The team wants to avoid manual image pulls on each node. Which two configurations should they implement to enable a successful air-gapped installation? (Choose two.)

Select 2 answers
A.Use a private image registry and configure image pull secrets in the `gpu-operator` namespace for authentication.
B.Configure the ClusterPolicy with `operator.defaultRuntime: containerd` and set `driver.repository` to the internal registry path.
C.Disable the Node Feature Discovery (NFD) component to reduce the number of images that need to be mirrored.
D.Mirror all GPU Operator component images to the internal registry and update the ClusterPolicy to reference that registry for each component.
E.Set `driver.enabled: false` to avoid pulling the driver image from the public registry.
AnswersA, D

A private registry often requires authentication. Creating an image pull secret in the operator's namespace and referencing it in the ClusterPolicy or service accounts allows nodes to pull images securely. This is a standard requirement for air-gapped installations where the internal registry is not publicly accessible, ensuring components can authenticate and retrieve images.

Why this answer

Air-gapped installations require that all container images used by the GPU Operator are available in a reachable registry. Mirroring every component image and updating the ClusterPolicy to reference the internal registry ensures pods can start without internet access. Additionally, if the registry requires authentication, an image pull secret must be configured in the operator's namespace to allow secure pulls.

Exam trap

The trap here is thinking that simply disabling a component or setting a repository path is enough for air-gapped operation, when all images must be mirrored and authentication configured if needed.

201
MCQeasy

When monitoring GPU health in an enterprise cluster, which command provides the most comprehensive snapshot of real-time power, temperature, and memory utilization?

A.nvcc --version
B.nvidia-smi
C.nvidia-bug-report.sh
D.cat /proc/driver/nvidia/gpus/*/information
AnswerB

The nvidia-smi tool is specifically designed to provide a comprehensive view of all NVIDIA GPUs installed in the system. It displays critical health metrics like power usage, temperature, memory usage, and compute utilization in a real-time format, serving as the primary diagnostic tool for AI infrastructure administrators.

Why this answer

The 'nvidia-smi' utility is the industry-standard tool for GPU management and monitoring. It provides a tabular overview of the system, including per-GPU power draw, thermal status, memory allocation, and active process lists. For administrators, this is the first point of contact for diagnosing performance bottlenecks or thermal throttling issues in a data center, making it the essential command for routine maintenance and operational health monitoring of NVIDIA hardware.

Exam trap

Candidates often look for complex monitoring software or cloud-native dashboards, forgetting that the native 'nvidia-smi' tool is the fundamental, built-in utility for real-time hardware diagnostics.

202
MCQhard

A team is profiling a distributed training job using NVIDIA NCCL for inter-GPU communication on a DGX A100 system. They observe that all-reduce operations are taking longer than expected, and the NCCL debug logs show frequent 'NVLS' (NVLink SHARP) errors. Which action should be taken to resolve the issue?

A.Increase the NCCL buffer size by setting NCCL_BUFFSIZE to a larger value to reduce the number of messages.
B.Disable NVLink SHARP by setting NCCL_NVLS_ENABLE=0 in the environment and restart the job.
C.Update the GPU driver and CUDA toolkit to the latest versions to fix NVLS bugs.
D.Switch the NCCL algorithm to 'Tree' by setting NCCL_ALGO=Tree to avoid NVLS usage.
AnswerB

NVLS (NVLink SHARP) is an in-network reduction feature that can accelerate all-reduce operations, but it requires compatible hardware and software. If errors occur, disabling it forces NCCL to fall back to standard ring or tree algorithms, which are reliable. This resolves the immediate errors and restores communication performance, though it may not achieve peak efficiency. It is a valid troubleshooting step when NVLS is unstable.

Why this answer

NVLS errors indicate that the NVLink SHARP acceleration is failing, likely due to hardware or software incompatibility. Disabling NVLS via the NCCL_NVLS_ENABLE environment variable forces NCCL to use standard algorithms, which are robust and should eliminate the errors. Other options either do not target NVLS or are less direct and potentially disruptive.

Exam trap

The trap here is assuming that increasing buffer size or changing algorithms will fix NVLS errors, when the correct approach is to disable the failing feature.

203
MCQhard

An administrator manages a Kubernetes cluster where a training job repeatedly fails with an OutOfMemory error on the GPU even though the pod requests one nvidia.com/gpu. DCGM metrics show that another pod on the same node is consuming GPU memory concurrently. GPU sharing via time-slicing is enabled cluster-wide. Which action should the administrator take to prevent this cross-pod interference while preserving the ability to share GPUs among trusted inference workloads?

A.Apply a ResourceQuota limiting nvidia.com/gpu to one per namespace
B.Set the CUDA_VISIBLE_DEVICES environment variable manually in the training pod spec
C.Disable time-slicing on the node pool hosting training jobs and use MIG or dedicated GPUs for those workloads
D.Increase the pod's nvidia.com/gpu request to two GPUs
AnswerC

Time-slicing provides no memory isolation, so concurrent pods can exhaust GPU memory and cause OutOfMemory failures. Removing training jobs from time-sliced nodes and giving them dedicated GPUs or MIG instances ensures exclusive memory access while still allowing time-slicing to remain available for trusted inference workloads on separate node pools.

Why this answer

Time-slicing multiplexes a GPU across pods without memory partitioning, so one workload can starve another and trigger OutOfMemory errors. The reliable fix is to separate untrusted or memory-intensive training workloads onto dedicated GPUs or MIG instances, while leaving time-slicing for inference workloads where memory contention is acceptable and controlled.

Exam trap

The trap here is thinking that increasing GPU requests or adding quotas creates memory isolation, when only dedicated devices or MIG instances actually partition GPU memory.

204
MCQhard

An administrator is configuring NVIDIA Base Command Manager to manage a cluster of DGX nodes. They want to ensure that when a node's GPU temperature exceeds a defined threshold, the node is automatically drained and an alert is sent to the operations team. Which combination of Base Command Manager features should the administrator configure to achieve this?

A.Health checks with a custom script that triggers a node drain via the Base Command Manager API, and an alert rule that sends a notification.
B.NVIDIA GPU Operator with node problem detector, and a Kubernetes taint that evicts pods.
C.Prometheus Alertmanager with a webhook that drains the node, and Grafana dashboards for visualization.
D.Base Command Manager job scheduler with a preemption policy that kills jobs when temperature is high.
AnswerA

Base Command Manager health checks can run custom scripts periodically. A script can query GPU temperature and, if over threshold, call the Base Command Manager API to drain the node. An alert rule then sends a notification. This provides automated remediation and alerting as required.

Why this answer

Base Command Manager provides health checks that can execute custom scripts on nodes. A script can monitor GPU temperature and invoke the Base Command Manager API to drain the node, while an alert rule notifies the team. This leverages native features for automated remediation and alerting, unlike external monitoring stacks or Kubernetes-specific tools.

Exam trap

The trap here is confusing Base Command Manager's native health and alerting features with external monitoring tools like Prometheus or Kubernetes node problem detector.

205
MCQhard

An administrator observes that despite the GPU Operator being installed, the pods cannot access the GPU. What is the most likely cause if the NVIDIA container runtime is properly configured?

A.The Kubernetes API server is down.
B.The host NVIDIA driver is not loaded.
C.The pod has too little memory requested.
D.The container image is missing the CUDA library.
AnswerB

Even with the correct runtime, the container requires the underlying host driver to communicate with the hardware. If the driver is not installed or the kernel module is not loaded, the runtime will fail to map the GPU device, leading to a situation where the GPU is inaccessible.

Why this answer

If the runtime is configured, the most common remaining failure is the lack of the correct NVIDIA drivers on the host node. If the kernel modules are not loaded or the driver version is incompatible with the installed GPU, the runtime will be unable to successfully inject the device nodes. This is a common installation oversight where the operator might be deployed but the driver installation task failed or was skipped.

Exam trap

Candidates often assume the issue is a Kubernetes misconfiguration or a pod error. They fail to check the underlying host kernel, where driver loading issues are the most frequent root cause.

206
Multi-Selecthard

An AI operations engineer is investigating a training job that exhibits poor scaling efficiency when moving from 8 to 32 GPUs on a DGX SuperPOD. Profiling indicates that the communication time in NCCL all-reduce operations is disproportionately high. Which two actions should be taken to improve scaling efficiency? (Choose two.)

Select 2 answers
A.Check that the InfiniBand fabric is not experiencing congestion or errors by monitoring performance counters.
B.Ensure that the job is using the correct NCCL topology by setting NCCL_TOPO_FILE to a custom topology XML.
C.Increase the NCCL buffer size by setting NCCL_BUFFSIZE to a larger value.
D.Enable NCCL tree algorithm for all-reduce by setting NCCL_ALGO=Tree.
E.Verify that NCCL is using the highest-bandwidth network interface, such as InfiniBand, by checking NCCL_IB_HCA.
AnswersA, E

High communication time can be caused by network congestion or errors on the InfiniBand fabric. Monitoring counters such as port errors, link down events, or congestion can reveal underlying issues. If the fabric is congested, all-reduce operations will slow down, impacting scaling. This is a fundamental check before tuning NCCL parameters. It directly addresses potential network-level bottlenecks.

Why this answer

Poor scaling in NCCL all-reduce often stems from network misconfiguration or fabric issues. Verifying that NCCL uses the correct high-bandwidth HCAs and checking the InfiniBand fabric for congestion or errors are essential first steps. These actions ensure that communication occurs over the fastest available path and that the network is healthy, directly improving scaling efficiency.

Exam trap

The trap here is focusing on NCCL algorithm or buffer tuning before confirming that the underlying network hardware and fabric are correctly utilized and healthy.

207
MCQhard

An AI ops engineer notices that a specific training workload is experiencing high 'wait' times for GPU resources despite the cluster having available idle GPUs. What is the most likely cause?

A.The GPU driver is too new for the current Kubernetes version.
B.The pod has unmet affinity or toleration requirements.
C.The system clock is unsynchronized between nodes.
D.The container image is too large to pull efficiently.
AnswerB

If a pod requires a specific node label or has a taint that the available nodes do not satisfy, it will remain in a pending state. This scenario is a common cause for pods failing to schedule, even when the underlying hardware resources appear idle to the cluster administrator.

Why this answer

High wait times despite idle resources often point to scheduling constraints, such as mismatched node affinity, taints, or tolerations. The scheduler may be unable to place the pod because the available GPUs do not meet specific requirements (e.g., specific memory requirements, architecture, or interconnect features). Troubleshooting requires examining the Kubernetes scheduler logs or pod event descriptions to identify why the pending pod cannot be bound to the available nodes.

Exam trap

Test-takers frequently assume idle GPUs mean hardware failures, overlooking scheduler constraints such as unmet node affinities, taints, or tolerations.

208
MCQmedium

An AI operations engineer manages a Kubernetes cluster running the NVIDIA GPU Operator. A team wants its long-running inference deployment to be automatically rescheduled if the GPU on a node develops an uncorrectable error that the device plugin or health checks detect. The team also wants the node to stop accepting new GPU pods until the issue is resolved. Which combination of behaviors should the engineer rely on to meet these requirements?

A.A Horizontal Pod Autoscaler scales up replicas so healthy copies absorb traffic while the faulty node remains in service.
B.The GPU Operator's health checks mark the GPU unhealthy, the device plugin stops advertising it, and Kubernetes taints or cordons the affected node so pods are rescheduled and no new GPU pods land there.
C.A liveness probe on the inference container restarts the pod on the same node when the GPU error causes a request failure.
D.A PodDisruptionBudget on the inference deployment forces the scheduler to migrate pods off the node when the GPU fails.
AnswerB

The GPU Operator runs health checks that can detect uncorrectable GPU errors and signal the device plugin to stop advertising the faulty device. The operator can also apply taints or cordon the node. Existing pods become unschedulable and are recreated elsewhere by their controller, while new GPU pods are kept off the node, matching both stated requirements.

Why this answer

Meeting both requirements needs health-driven device withdrawal plus node-level scheduling exclusion. The GPU Operator's health checks detect uncorrectable errors, the device plugin stops advertising the faulty GPU, and the node is tainted or cordoned. Controllers then recreate pods on healthy nodes, and the taint keeps new GPU pods away until the fault is cleared.

Exam trap

The trap here is expecting pod-level constructs like liveness probes or disruption budgets to handle hardware faults, when GPU error isolation is driven by device health checks and node taints.

209
MCQmedium

What is the primary role of a private container registry in an NVIDIA AI Enterprise deployment?

A.Managing hardware firmware updates
B.Storing and distributing container images
C.Monitoring real-time GPU thermals
D.Providing GPU driver licensing keys
AnswerB

A private registry acts as a local source of truth for container images. In production environments, it is essential for security and stability, allowing administrators to control exactly which software versions are deployed to the cluster while reducing reliance on external registries that may have latency or availability issues.

Why this answer

A private registry provides a secure, reliable, and high-speed source for container images within the enterprise firewall. By caching NVIDIA-provided containers locally, organizations avoid dependency on public repositories, which is critical for security compliance and offline operational readiness. It ensures that the cluster has consistent, immutable versions of software available, preventing drift and ensuring that deployments remain stable and reproducible across the entire production environment.

Exam trap

Candidates often assume a registry is for model versioning or training data storage. In the context of NVIDIA AI Enterprise, its primary function is strictly container image management and distribution.

210
MCQeasy

An AI operations engineer manages a Kubernetes cluster where the NVIDIA GPU Operator's device plugin exposes GPUs as schedulable resources. A data science team submits a batch inference job that requests one GPU but does not specify a node selector or tolerations. The job stays in Pending while other GPU nodes remain idle because they carry a taint the GPU Operator applied to reserve them for a specific workload class. Which approach is the most appropriate for the engineer to make the job schedulable without disrupting the reserved nodes?

A.Add an appropriate toleration and node selector to the job so it can target the reserved GPU nodes.
B.Delete and recreate the device plugin daemonset so GPUs are re-advertised to the scheduler.
C.Increase the cluster's GPU resource quota in the namespace so the scheduler can allocate a card.
D.Remove the node taint from all GPU nodes so the job can be scheduled anywhere.
AnswerA

A taint on a node only repels pods that do not tolerate it. Adding the matching toleration and a node selector that identifies the reserved GPU nodes lets this job schedule there while every other untolerated workload stays off those nodes. The reservation remains intact and the job becomes runnable, which directly resolves the Pending state.

Why this answer

A node taint repels pods unless they carry a matching toleration. The reserved GPU nodes are intentionally tainted, so a job that neither tolerates the taint nor selects an untainted node cannot schedule. Adding the correct toleration plus a node selector places the job on the reserved nodes while preserving the reservation for other classes of work.

Exam trap

The trap here is assuming a Pending GPU pod is caused by missing GPU resources or quota, when the actual cause is an un-tolerated node taint.

211
Multi-Selectmedium

A research team is submitting many short-lived experiment jobs to an NVIDIA-accelerated Kubernetes cluster. The operations team wants to reduce GPU idle time and improve overall utilization without modifying the training code. Which TWO approaches should the operations team implement? (Choose two.)

Select 2 answers
A.Pin each experiment job to a dedicated physical GPU using node affinity
B.Use a gang-scheduling or batch scheduler with backfill so queued jobs start as soon as GPUs free up
C.Disable the NVIDIA device plugin and mount GPUs directly via hostPath
D.Increase the pod's nvidia.com/gpu request to reserve more GPU memory
E.Enable GPU sharing with time-slicing so multiple experiment pods can occupy the same GPU
AnswersB, E

Batch schedulers with backfill can start smaller jobs ahead of larger queued jobs when resources are available, keeping GPUs busy between experiment runs. This reduces idle gaps and improves throughput for many short-lived jobs, complementing GPU sharing without requiring modifications to the training scripts.

Why this answer

Short-lived experiment jobs create idle GPU gaps. Time-slicing allows multiple such pods to share a device, and a batch scheduler with backfill keeps the queue moving so GPUs are assigned as soon as capacity frees. Both measures raise utilization without touching training code, which is exactly what the operations team needs.

Exam trap

The trap here is assuming that dedicating a GPU per short job or requesting more GPUs improves utilization, when in fact both reduce concurrency and leave GPUs idle.

212
MCQmedium

If a training job on a multi-node cluster shows a significant performance drop during checkpointing, what is the most likely bottleneck?

A.Network latency between compute nodes.
B.Shared filesystem I/O throughput.
C.GPU memory allocation overhead.
D.The model weight update frequency.
AnswerB

Checkpointing requires writing large amounts of data to disk. If multiple nodes attempt to write to a shared filesystem simultaneously, the aggregate I/O demand can exceed the filesystem's bandwidth, causing the training process to hang while waiting for the write operation to complete.

Why this answer

Checkpointing involves writing large model states to non-volatile storage. If this process is not handled asynchronously or if the underlying storage system is undersized, the training process will stall. Identifying this bottleneck allows engineers to implement optimized I/O strategies, such as using distributed file systems or NVMe-based local storage, to minimize the time spent in the checkpoint phase and maximize overall cluster efficiency.

Exam trap

Candidates often assume checkpoint performance drops are caused by insufficient GPU memory or slow CPU speeds, overlooking the massive I/O bottleneck created by writing large state files simultaneously.

213
MCQeasy

Which component of the NVIDIA AI Enterprise stack is primarily responsible for ensuring the long-term stability and compatibility of drivers and libraries across heterogeneous hardware configurations?

A.NVIDIA CUDA Toolkit
B.NVIDIA AI Enterprise
C.NVIDIA Triton Inference Server
D.NVIDIA Collective Communications Library
AnswerB

NVIDIA AI Enterprise is a cloud-native software suite that provides validated, secure, and supported software. It is specifically designed to ensure that the entire stack, from drivers to containers, is compatible and stable, allowing organizations to maintain production consistency across diverse data center hardware environments over long periods.

Why this answer

NVIDIA AI Enterprise provides a validated and supported software stack that ensures consistency across different hardware generations. Stability is the foundation of AI operations; without a supported, validated stack, minor version mismatches between kernels and libraries can lead to silent failures or system crashes. This support model is essential for enterprises to avoid the 'dependency hell' that often plagues research-grade environments when scaling to production operations.

Exam trap

Candidates often name individual components like CUDA or NCCL, ignoring that NVIDIA AI Enterprise is the overarching suite specifically designed for validated support and version compatibility.

214
MCQmedium

An AI researcher is running a training job using mixed precision (FP16/BF16). The loss function is diverging unexpectedly. What is the most likely culprit?

A.The GPU driver is too old for FP16.
B.The learning rate is too high for mixed precision.
C.The lack of loss scaling for FP16.
D.The batch size is too large.
AnswerC

FP16 has a limited dynamic range. During backpropagation, many small gradient values can be rounded to zero (underflow). Loss scaling multiplies the loss before backpropagation, pushing the gradients into a representable range in FP16, then unscaling them before applying weight updates. This is critical for preventing divergence in mixed precision.

Why this answer

Mixed precision training uses lower precision formats to speed up math and reduce memory usage, but this can lead to numerical instability. Divergence in loss is a classic symptom of 'underflow' or 'overflow' issues occurring in FP16, where small gradients or large weight updates exceed the representable range. Implementing loss scaling is the standard industry technique to preserve the precision of gradients and ensure the training process remains stable.

Exam trap

Candidates often assume divergence is caused by a poor learning rate or bad initialization, overlooking the precision limitations inherent in standard FP16 training.

215
MCQhard

Refer to the exhibit. A cluster administrator notices that GPU jobs with this PriorityClass are failing to start even when empty GPUs are available. What is the most likely cause?

A.The PriorityClass value is too low to trigger preemption.
B.Insufficient nodes satisfy the 'minAvailable' requirement for the pod group.
C.The NVIDIA Device Plugin requires a restart to recognize the PriorityClass.
D.The PriorityClass name violates Kubernetes naming conventions.
AnswerB

In high-performance computing, Volcano requires that all pods within a group start simultaneously. If the requested number of GPUs cannot be satisfied across available nodes due to locality or capacity constraints, the job remains pending, even if individual nodes have free GPUs, to prevent incomplete distributed training runs.

Why this answer

When using the Volcano scheduler with NVIDIA GPUs, priority classes alone are insufficient if the Gang Scheduling policy is not met. If the job requests multiple pods that cannot be satisfied simultaneously, Volcano will leave the pods in a pending state to avoid resource fragmentation. This ensures that massive training jobs do not consume isolated resources that would otherwise result in deadlocks during collective communication synchronization.

Exam trap

Candidates often ignore the 'minAvailable' parameter in Volcano, assuming that as long as individual GPUs are free, the job will eventually start, ignoring the gang scheduling requirement.

216
MCQhard

An operations team observes that a distributed training job using NVIDIA Collective Communications Library (NCCL) across eight nodes occasionally hangs during the all-reduce phase. Logs show no errors, and the hang resolves only after a node is manually restarted. Which action is MOST appropriate to diagnose the intermittent hang?

A.Disable InfiniBand and fall back to TCP sockets by setting NCCL_IB_DISABLE=1.
B.Increase the NCCL buffer size with NCCL_BUFFSIZE to reduce the number of messages exchanged.
C.Set NCCL_ALGO=RING to force a single algorithm and eliminate variability.
D.Enable NCCL debug logging with NCCL_DEBUG=INFO and NCCL_DEBUG_SUBSYS=COLL, and set a NCCL watchdog timeout to capture the stalled collective.
AnswerD

Intermittent hangs without errors require visibility into which collective and rank stalled. NCCL_DEBUG=INFO with NCCL_DEBUG_SUBSYS=COLL prints per-collective progress and rank participation, while a watchdog timeout can trigger a dump when progress stops. This combination captures the state at the moment of the hang, making it the correct diagnostic action for this scenario.

Why this answer

Intermittent NCCL hangs without errors are best diagnosed by capturing state at the moment of the stall. Enabling NCCL debug logging for the collective subsystem and configuring a watchdog timeout produces logs that show which rank and collective stopped progressing. This evidence is necessary to distinguish a slow rank, a fabric problem, or a software defect from a simple performance issue.

Exam trap

The trap here is trying to work around the hang by changing transport or algorithm rather than capturing diagnostic data at the point of failure.

217
MCQhard

An AI operations team runs long-running training jobs on a Kubernetes cluster with NVIDIA GPU Operator. They observe that after a node is rebooted for maintenance, some pods resume but report CUDA 'unknown error' and the device plugin shows unhealthy GPUs. Which configuration should the administrator review to ensure the driver and device plugin recover cleanly after reboot?

A.The pod's restartPolicy set to Always so containers restart automatically after node recovery.
B.The GPU Operator's driver validation and node reboot handling settings, including the driver DaemonSet and readiness gates.
C.The cluster autoscaler settings so replacement nodes are provisioned immediately after reboot.
D.The container image's CUDA version to ensure it matches the host driver version.
AnswerB

After a reboot, the GPU Operator must reload the driver and reinitialize the device plugin before workloads can use GPUs safely. Driver validation and readiness gates prevent pods from starting until the driver and plugin report healthy. Reviewing these settings addresses the CUDA error and unhealthy device plugin state observed after maintenance reboots.

Why this answer

A reboot invalidates the previously loaded driver and device plugin state, so the GPU Operator must reload the driver and re-register devices before pods can use them. Driver validation and readiness gates hold workloads until health checks pass, preventing CUDA errors from reaching applications. Reviewing these settings ensures clean recovery and avoids the unhealthy plugin state observed after maintenance.

Exam trap

The trap here is blaming the application container or its CUDA version for a fault that originates in post-reboot node initialization by the GPU Operator.

218
MCQmedium

An AI operations team is using NVIDIA DCGM to monitor a cluster of GPUs. They want to set up alerts for when GPUs are running at high temperatures for extended periods. Which DCGM feature should they use?

A.DCGM's built-in email alerting system configured via dcgm.conf.
B.DCGM diagnostics with the '-r' flag to run a comprehensive test and report temperature violations.
C.NVIDIA-smi's '--query-gpu=temperature.gpu' with a cron job to check and send alerts.
D.DCGM health checks with custom thresholds and alerting via DCGM exporter to Prometheus.
AnswerD

DCGM provides health checks that can be configured with thresholds for temperature and other metrics. The DCGM exporter can expose these metrics to Prometheus, which supports alerting rules based on sustained conditions. This combination allows for proactive monitoring and alerting on high temperatures over time. It is the standard approach for cluster-wide GPU monitoring.

Why this answer

DCGM health checks allow defining thresholds for GPU metrics, including temperature. When combined with the DCGM exporter and Prometheus, teams can create alerting rules that trigger on sustained high temperatures. This is the recommended method for cluster-wide monitoring.

Other options either misuse DCGM diagnostics, assume non-existent features, or use less scalable methods.

Exam trap

The trap here is confusing DCGM diagnostics (point-in-time testing) with continuous monitoring and alerting, which requires integration with a metrics platform.

219
MCQmedium

Which of the following is the most appropriate workload management technique for a bursty AI inference workload that requires low latency but does not need full GPU power for every request?

A.Assigning one full physical GPU to every single inference request.
B.Using MIG or fractional GPU sharing to maximize hardware throughput.
C.Configuring the GPU to run in 'Maximum Performance' mode at all times.
D.Implementing a strict FIFO queue for all incoming inference requests.
AnswerB

MIG and fractional sharing allow multiple inference workloads to run concurrently on a single GPU. This effectively increases the number of available 'virtual' GPUs, maximizing utilization and ensuring that small, bursty requests can be handled with minimal latency, which is essential for cost-effective, high-scale inference services.

Why this answer

NVIDIA Multi-Instance GPU (MIG) combined with time-slicing or fractional GPU sharing is ideal for inference workloads. By partitioning GPUs, the system can handle many smaller, latency-sensitive requests efficiently. This approach balances the need for high-performance hardware with the requirement for multi-tenancy, ensuring that no single inference request monopolizes the entire GPU while maintaining the fast response times required by the application.

Exam trap

Candidates often mistake general load balancing or horizontal scaling for GPU-level resource management, failing to recognize that MIG is the specific hardware-level solution for partitioning GPUs for multi-tenant inference.

220
MCQeasy

An administrator is deploying NVIDIA AI Enterprise on a Kubernetes cluster and wants to verify that the GPU Operator has successfully installed all required components. Which command should the administrator use to check the status of the GPU Operator pods?

A.nvidia-smi -q
B.kubectl get nodes -o wide
C.dcgmi discovery -l
D.kubectl get pods -n gpu-operator
AnswerD

The GPU Operator deploys its components into the gpu-operator namespace by default. Running kubectl get pods -n gpu-operator lists all pods in that namespace, allowing the administrator to see if the operator, driver, container toolkit, device plugin, and other components are running. This is the standard way to verify the operator's deployment status.

Why this answer

The GPU Operator installs its components as pods in the gpu-operator namespace. Checking pod status with kubectl get pods -n gpu-operator is the correct way to verify that all required components are running. Other commands provide GPU-level or node-level information but do not show the status of the operator's pods.

Exam trap

The trap here is confusing host-level GPU queries with Kubernetes pod status checks; nvidia-smi and dcgmi do not show operator pods.

221
MCQeasy

An AI operations engineer notices that a training job on an NVIDIA A100 GPU is running slower than expected. Running nvidia-smi shows that the GPU is in 'P0' state but the 'SM Clock' is significantly lower than the maximum boost clock. The job is not memory-bound. Which action should the engineer take first to diagnose the issue?

A.Reinstall the NVIDIA driver to ensure the latest version.
B.Increase the batch size to improve GPU utilization.
C.Switch the job to use mixed precision to reduce compute load.
D.Check the GPU's power consumption and whether it is hitting the power limit.
AnswerD

Low SM clock despite P0 state often indicates power or thermal throttling. Checking power consumption against the configured limit helps determine if the GPU is throttling due to power constraints. This is a common cause of reduced clocks and is a logical first diagnostic step before changing application settings.

Why this answer

When a GPU is in P0 state but SM clock is low, power or thermal throttling is a likely cause. Checking power consumption and limits is a quick, non-invasive diagnostic step. If the GPU is hitting its power limit, the engineer can then adjust power settings or optimize the workload to reduce power draw.

The other actions do not directly diagnose the throttling issue.

Exam trap

The trap here is jumping to application-level optimizations like batch size or precision changes before verifying whether the GPU is simply power-limited or thermally throttled.

222
MCQmedium

An AI operations engineer is preparing a DGX H100 system for a multi-node training workload. The engineer runs `nvidia-smi topo -m` and notices that GPU4 and GPU5 report a connection type of SYS, while all other GPU pairs show NV18. What is the most likely cause of this topology anomaly?

A.The NVSwitch fabric is operating in a fallback mode that only activates when peer-to-peer traffic exceeds a threshold.
B.GPU4 and GPU5 are configured in MIG mode, which forces inter-GPU communication over the PCIe bus.
C.GPU4 and GPU5 are connected through the PCIe host bridge rather than the NVSwitch fabric, possibly due to a degraded or missing NVLink connection.
D.GPU4 and GPU5 are reserved for display output, so the driver automatically disables their NVLink interfaces.
AnswerC

SYS indicates that communication between those two GPUs traverses the system's PCIe and CPU interconnect instead of the high-speed NVLink/NVSwitch fabric. On a DGX H100, all GPU pairs should normally show NV18 via NVSwitch, so this points to a degraded or missing NVLink path between GPU4 and GPU5 that must be investigated.

Why this answer

A healthy DGX H100 shows NV18 (NVLink 4.0, 18 links) for every GPU pair because all GPUs are connected through NVSwitch. A SYS entry for a specific pair means traffic between those GPUs falls back to the PCIe/CPU path, which severely reduces bandwidth and increases latency. This typically signals a hardware or cabling fault on the NVLink side that should be resolved before running distributed training.

Exam trap

The trap here is assuming that SYS is a normal topology state for some GPU pairs on an NVSwitch-equipped system, when it actually indicates a degraded or absent NVLink path.

223
MCQhard

A team runs a multi-node NCCL all-reduce training job on four DGX H100 nodes connected by an InfiniBand fabric. Scaling efficiency is poor: throughput barely improves beyond two nodes, and `nvidia-smi` shows NIC transmit counters on each GPU's assigned HCA are far below the PCIe link capacity while GPU compute utilization sits at ~55%. The fabric manager logs report all links as up with no symbol errors. Which action should the administrator take first to diagnose the interconnect bottleneck?

A.Switch the job to use the NCCL tree algorithm via NCCL_ALGO=TREE because ring all-reduce cannot scale across four nodes.
B.Enable GPUDirect RDMA by exporting NCCL_NET_GDR_LEVEL=SYS and rerunning the job to confirm whether host-memory staging was the cause.
C.Run `nccl-tests` all_reduce_perf with the same process count and message sizes, then compare its reported bus bandwidth against the theoretical peak for the fabric.
D.Increase the NCCL_BUFFSIZE environment variable to its maximum so each all-reduce chunk transfers more data per operation.
AnswerC

Running nccl-tests all_reduce_perf reproduces the collective in isolation and reports algorithmic and bus bandwidth, so a gap versus theoretical fabric peak localizes the loss to the communication path rather than the model code. Because every link is error-free and HCA counters are low, the fault is likely in topology, ring/tree selection, or placement, and this microbenchmark exposes exactly that before any configuration is changed.

Why this answer

When every fabric link is error-free but HCA counters sit well below link bandwidth, the bottleneck is in how the collective is scheduled or placed, not in raw link health. Reproducing the all-reduce with nccl-tests and comparing achieved bus bandwidth to the fabric's theoretical peak isolates the communication path and reveals whether the gap comes from topology detection, ring construction, or process placement before any tuning is applied.

Exam trap

The trap here is assuming that low NIC counters plus healthy links prove the network is fine, when it actually points to the collective not saturating the fabric at all.

224
MCQeasy

Which NVIDIA technology enables a GPU to be shared among multiple virtual machines or containers while maintaining strict hardware isolation?

A.NVIDIA CUDA streams
B.NVIDIA Multi-Instance GPU (MIG)
C.NVIDIA Collective Communications Library (NCCL)
D.NVIDIA GPUDirect RDMA
AnswerB

MIG allows a physical GPU to be partitioned into multiple isolated instances at the hardware level. Each instance has its own dedicated memory and compute resources, providing the isolation necessary for multi-tenant environments where security and performance guarantees are required for concurrent, independent workloads on a single piece of hardware.

Why this answer

NVIDIA Multi-Instance GPU (MIG) is the core technology that enables hardware-level partitioning on supported GPUs. By creating dedicated instances that have their own memory and compute resources, MIG ensures that workloads remain isolated, providing predictable quality of service and security in multi-tenant environments. This is a foundational technology for maximizing GPU utilization in cloud and enterprise data center environments where mixed workloads are common.

Exam trap

Test-takers sometimes confuse software-based time-slicing with true hardware isolation technologies like MIG when multi-tenant security is required.

225
MCQmedium

An administrator is troubleshooting a multi-node NVIDIA GPU training job that intermittently hangs during the all-reduce phase. The job uses NCCL over InfiniBand. Logs show that some ranks time out while others complete. The administrator suspects a network fabric issue. Which action should the administrator take first to isolate whether the problem is in the InfiniBand fabric or in the NCCL configuration?

A.Disable InfiniBand and force NCCL to use TCP sockets instead, because TCP is more reliable for collective operations.
B.Immediately replace all InfiniBand cables and transceivers on the affected nodes, because intermittent hangs during all-reduce almost always indicate physical link errors.
C.Run the NCCL tests (nccl-tests) with the same topology and environment variables, and enable NCCL debug logging to capture detailed transport and topology information.
D.Increase the NCCL timeout value in the training script and restart the job, because the hangs are likely due to slow network convergence.
AnswerC

Running nccl-tests with debug logging reproduces the collective communication pattern outside the training job. It reveals which transport (InfiniBand, RoCE, or sockets) is used, whether topology detection fails, and where timeouts occur. This isolates NCCL configuration issues from fabric problems before deeper network diagnostics.

Why this answer

Using nccl-tests with debug logging is a controlled way to reproduce the collective communication and observe transport selection, topology detection, and timeouts. It helps determine whether the issue is in NCCL configuration or the InfiniBand fabric. Replacing hardware, forcing TCP, or increasing timeouts are premature and do not isolate the root cause.

Exam trap

The trap here is jumping to hardware replacement or timeout adjustments before verifying whether NCCL is correctly configured and using the expected transport.

Page 2

Page 3 of 5

Page 4

All pages