Courseiva

NVIDIA Certified Professional: AI Operations (NCP-AIO) — Questions 1–75

309 questions total · 5pages · All types, answers revealed

Page 1 of 5

Page 2
1
MCQmedium

During a multi-GPU training job, you notice that one GPU consistently reports lower utilization and longer communication times compared to others. What is the most likely reason for this performance imbalance?

A.The GPU is defective and needs replacement.
B.The GPU is connected via a slower interconnect.
C.The training framework is not using CUDA streams.
D.The ambient server temperature is too high.
AnswerB

If one GPU in a cluster lacks the high-speed NVLink connection enjoyed by the others, it will be limited by the bandwidth of the slower interface (like PCIe). During AllReduce operations, the entire cluster must wait for this 'straggler' to complete, leading to lower total utilization and increased latency.

Why this answer

In multi-GPU systems, uneven load distribution often stems from mismatched NVLink topologies or PCIe lane configurations. If one GPU is connected via a slower PCIe link rather than a high-speed NVLink interconnect, it becomes the bottleneck in collective communication operations like AllReduce. Ensuring symmetric connectivity across all GPUs is essential for predictable performance and preventing the 'straggler' effect in distributed deep learning training.

Exam trap

Candidates often attribute performance imbalances to software bugs or uneven data batching, failing to account for physical hardware topology issues, such as mismatched PCIe lanes or NVLink connectivity.

2
MCQmedium

An AI researcher is using Nsight Systems to profile an application. They notice a large gap in the timeline where neither the CPU nor the GPU is performing significant work. What does this gap most likely represent?

A.The GPU is busy performing background garbage collection.
B.The application is performing blocking synchronization.
C.The profiler has encountered a buffer overflow.
D.The system is entering a power-saving state.
AnswerB

Gaps in a timeline often signify synchronization barriers where the host is waiting for a device to finish, or vice versa. This effectively pauses the entire execution pipeline. By using non-blocking CUDA streams and events, these synchronization gaps can be removed, allowing the system to maintain continuous compute throughput.

Why this answer

Large gaps in profiling timelines often indicate synchronization points where the CPU is waiting for the GPU to finish a task, or vice versa, due to blocking operations. These 'stalls' are common in poorly optimized code that issues many synchronous calls. Identifying these gaps allows the developer to re-structure the code to use asynchronous execution, thereby overlapping compute and data movement to improve performance.

Exam trap

Candidates often assume that a timeline gap means the hardware is broken or under-powered, overlooking software-level synchronization barriers where the CPU and GPU are simply waiting on each other.

3
Multi-Selectmedium

An administrator is preparing a cluster for a new large language model training job that will use NVIDIA Magnum IO GPUDirect Storage to stream training data directly from a parallel file system to GPU memory. The administrator must verify that the environment supports GPUDirect Storage before the job starts. (Choose two.)

Select 2 answers
A.Disable IOMMU in the BIOS on all GPU nodes to allow direct memory access between storage devices and GPUs.
B.Ensure that the parallel file system client supports the cuFile API and that the file system is mounted with options that allow direct I/O and peer-to-peer memory access.
C.Configure the cluster to use RDMA over Converged Ethernet (RoCE) exclusively for all storage traffic, because GPUDirect Storage requires RDMA.
D.Confirm that the NVIDIA driver and CUDA toolkit versions installed on the nodes meet the minimum requirements for the GPUDirect Storage release being used, and that the nvidia-fs kernel module is loaded.
E.Enable NVIDIA Multi-Instance GPU (MIG) on every GPU to provide isolated memory partitions for storage transfers.
AnswersB, D

The file system client must support cuFile and be mounted appropriately for direct memory access. Without cuFile support, applications cannot use the GPUDirect Storage path. Mount options and client capabilities determine whether the storage stack can perform peer-to-peer transfers to GPU memory.

Why this answer

GPUDirect Storage requires compatible drivers, the nvidia-fs kernel module, and a file system client that supports the cuFile API with suitable mount options. These two checks confirm that the software stack and storage client can perform direct memory transfers. MIG, exclusive RoCE, and disabling IOMMU are not prerequisites and may be counterproductive.

Exam trap

The trap here is assuming that GPUDirect Storage depends on MIG or a specific network transport, when it actually relies on driver, kernel module, and file system client support.

4
MCQmedium

A media company is deploying an inference service on a Kubernetes cluster with the NVIDIA GPU Operator installed. The service pods remain in Pending with the message that no nodes have the requested nvidia.com/gpu resource, even though the GPUs are healthy and the driver loads correctly on every node. Which troubleshooting step should the engineer perform first?

A.Reinstall the NVIDIA driver on every node to refresh the kernel module.
B.Increase the kubelet's pod density limit so the plugin can be scheduled.
C.Check whether the NVIDIA device plugin pods are running and have registered the GPUs with kubelet.
D.Enable Multi-Instance GPU mode on all GPUs so each partition advertises capacity.
AnswerC

The scheduler only sees nvidia.com/gpu capacity after the device plugin registers each GPU with kubelet through the plugin socket. If the device plugin pods are crashing or not scheduled, nodes show no GPU resource and pods stay Pending. Verifying plugin pod status and logs is the most direct first step when the driver itself is healthy.

Why this answer

Kubernetes learns about GPUs through the device plugin framework: the NVIDIA device plugin advertises nvidia.com/gpu for each visible GPU by registering with kubelet. When the driver is healthy but pods report no such resource, the plugin is the missing link. Checking its pod status and logs quickly reveals scheduling failures, crashes, or socket registration errors before deeper investigation.

Exam trap

The trap here is jumping to driver or hardware remediation when the symptom, a healthy driver with no advertised nvidia.com/gpu resource, points squarely at the device plugin registration path.

5
MCQeasy

An administrator is setting up NVIDIA Base Command Manager to provision and manage a new AI cluster. They need to ensure that the cluster can automatically discover and configure new GPU nodes. Which component is responsible for node discovery and initial configuration?

A.NVIDIA Fleet Command
B.NVIDIA GPU Operator
C.NVIDIA Morpheus
D.NVIDIA Base Command Manager head node
AnswerD

The head node in Base Command Manager is the central management server that handles node discovery, provisioning, and configuration. It runs services like DHCP, TFTP, and the provisioning engine that automatically detect new nodes when they PXE boot. It then applies the appropriate image and configuration. This is the core component for automating cluster setup, making it the correct answer for node discovery and initial configuration.

Why this answer

In NVIDIA Base Command Manager, the head node is the central management server that performs node discovery and initial configuration. It provides the provisioning services that allow new GPU nodes to be automatically detected and configured. The GPU Operator, Fleet Command, and Morpheus serve different purposes and are not involved in the initial provisioning of a Base Command Manager cluster.

Exam trap

The trap here is confusing Base Command Manager with Kubernetes-focused tools like the GPU Operator, which handle different layers of the stack.

6
MCQeasy

A DevOps engineer is deploying the NVIDIA GPU Operator on a Kubernetes cluster that uses containerd as the container runtime. The engineer notices that the Operator's validation pod fails with an error indicating that the NVIDIA container runtime is not configured. Which action should the engineer take to resolve this?

A.Manually install the NVIDIA Container Toolkit on all nodes and set the default runtime to nvidia.
B.Switch the cluster to use Docker as the container runtime, as the GPU Operator only supports Docker.
C.Ensure that the GPU Operator's container-toolkit daemonset is enabled and has the necessary permissions to modify the container runtime configuration.
D.Disable the validation pod in the GPU Operator's Helm chart to suppress the error.
AnswerC

The GPU Operator deploys a container-toolkit daemonset that installs and configures the NVIDIA Container Toolkit. If this daemonset is disabled or lacks permissions (e.g., privileged access), it cannot modify the containerd configuration. Enabling it and ensuring proper RBAC and security context allows the Operator to set up the runtime, resolving the validation error.

Why this answer

The GPU Operator uses a container-toolkit daemonset to install and configure the NVIDIA Container Toolkit on each node, including modifying the containerd configuration. If this daemonset is disabled or lacks privileges, the runtime won't be set up, causing validation failures. Ensuring the daemonset is enabled and has proper permissions resolves the issue.

Exam trap

The trap here is thinking manual installation or switching runtimes is needed, when the GPU Operator is designed to handle runtime configuration automatically.

7
MCQmedium

What is the most accurate way to verify that a training job is utilizing Tensor Cores?

A.Check the GPU power consumption in nvidia-smi.
B.Monitor the SM Tensor utilization metric via Nsight Compute.
C.Check the memory bandwidth in DCGM.
D.Verify the CUDA driver version is above 500.0.
AnswerB

Nsight Compute provides specific hardware counters for SM Tensor utilization. This allows an engineer to directly verify if the kernels are issuing the specific matrix multiply-accumulate instructions that run on Tensor Cores, confirming that the hardware acceleration is active for the workload.

Why this answer

Tensor Cores are specialized hardware units for mixed-precision matrix operations. Monitoring the SM occupancy and specific instruction usage via Nsight Compute is the definitive way to confirm their engagement. Understanding how to verify this is essential for engineers to ensure that their optimization efforts, such as mixed-precision training, are actually resulting in the intended hardware utilization and performance improvements.

Exam trap

Candidates often assume that high GPU utilization automatically implies Tensor Core usage. They fail to distinguish between general CUDA core compute and specialized matrix operations performed by Tensor Cores.

8
MCQhard

A research team runs a multi-node distributed training job spanning eight GPU nodes. Jobs frequently begin execution before all worker pods are running, and the collective initialization hangs until the operator manually scales the job down and up. The administrator wants the scheduler to admit the job only when all of its pods can be placed together. Which mechanism should be used?

A.Increase the kube-scheduler's default backoff period so pods that fail to schedule retry less aggressively and wait for peers to become ready.
B.Configure gang scheduling through a scheduler plugin so the job's pods are placed atomically only when the full group can be accommodated.
C.Set a high PriorityClass on every worker pod so the scheduler treats the group as important and places members as soon as resources appear.
D.Add a pod anti-affinity rule requiring each worker to run on a distinct node so that placement spreads evenly across the eight GPU nodes.
AnswerB

Gang scheduling holds the entire job group until every member can be placed on available resources, then admits them together. This directly prevents the partial-start condition that stalls collective initialization. Using a scheduler plugin that supports gang or coscheduling semantics gives the atomic placement guarantee the team needs, eliminating the manual scale-down and scale-up workaround.

Why this answer

Collective initialization requires all ranks to be present before computation proceeds, so partial startup leads to hangs. Gang scheduling solves this by treating the job group as a single scheduling unit and admitting it only when every member fits. Scheduler plugins that implement coscheduling or gang semantics provide this all-or-nothing placement, which is why they are standard in large distributed training environments.

Exam trap

The trap here is reaching for priority or anti-affinity as a fix for partial startup; those affect ordering and node distribution but never make pod admission atomic, so the collective can still begin with missing ranks.

9
MCQeasy

When installing the NVIDIA Container Toolkit to enable GPU acceleration in Docker, which file must be modified or verified to ensure the container runtime can access the NVIDIA runtime?

A./etc/fstab
B./etc/docker/daemon.json
C./etc/environment
D./boot/grub/grub.cfg
AnswerB

This is the primary configuration file for the Docker daemon. Including the nvidia-container-runtime in this JSON configuration is mandatory to register the runtime with Docker. This allows the engine to recognize the --gpus flag, enabling GPU hardware pass-through for accelerated workloads inside containers.

Why this answer

The daemon configuration file, typically located at /etc/docker/daemon.json, must be updated to include the NVIDIA runtime. This integration is the foundational step for AI containerization, as it allows the Docker engine to map host GPU resources into the container namespace. Without this configuration, containers will fail to detect GPUs, rendering them incapable of running accelerated AI applications.

Exam trap

Candidates often look for environment variables or shell scripts, missing that the persistent configuration for the Docker runtime resides in the /etc/docker/daemon.json file.

10
MCQeasy

A systems administrator is installing the NVIDIA Container Toolkit on a stand-alone server running Ubuntu 22.04 to enable Docker containers to access NVIDIA GPUs. After installation, they run a test container and find that it cannot see the GPU. Which step is most likely missing?

A.Rebooting the server after installing the toolkit.
B.Adding the user to the 'docker' group to run containers with GPU support.
C.Configuring the Docker daemon to use the NVIDIA container runtime as the default runtime.
D.Installing the NVIDIA GPU driver on the host.
AnswerC

The NVIDIA Container Toolkit requires the Docker daemon to be configured to use the 'nvidia' runtime. This is typically done by adding 'nvidia' to the 'runtimes' section in '/etc/docker/daemon.json' and optionally setting it as the default. Without this configuration, Docker will not use the NVIDIA runtime, and containers will not have GPU access, even if the toolkit is installed.

Why this answer

After installing the NVIDIA Container Toolkit, the Docker daemon must be configured to use the NVIDIA container runtime. This involves editing '/etc/docker/daemon.json' to add the runtime and restarting Docker. Without this configuration, containers will not have GPU access, even though the toolkit is installed.

Exam trap

The trap here is thinking that installing the toolkit alone is sufficient, when Docker must also be told to use the NVIDIA runtime.

11
MCQeasy

An AI operations engineer is deploying a model on an NVIDIA GPU and notices that inference latency is higher than expected. The model uses a batch size of 1, and profiling shows that the GPU is idle between kernel launches. Which optimization technique should the engineer use to reduce latency by overlapping data transfer with computation?

A.Enable NVIDIA MPS to allow multiple processes to share the GPU.
B.Switch to a lower-precision data type such as FP16 to reduce memory transfer size.
C.Increase the batch size to improve GPU utilization and reduce per-inference latency.
D.Use CUDA streams to overlap host-to-device memory copies with kernel execution.
AnswerD

CUDA streams allow asynchronous execution of memory copies and kernels, enabling overlap of data transfer with computation. For a batch size of 1, where kernels are short, this overlap can hide transfer latency and reduce overall inference time. The engineer should use separate streams for data transfer and compute, and synchronize appropriately, to achieve the desired latency reduction.

Why this answer

The idle time between kernel launches indicates that data transfers and kernel executions are serialized. Using CUDA streams to overlap host-to-device copies with kernel execution can hide transfer latency and reduce overall inference time. This is the correct technique for overlapping data transfer with computation, especially with small batch sizes where transfer overhead is significant.

Exam trap

The trap here is assuming that increasing batch size or using lower precision will reduce latency, when the issue is serialization of transfers and kernels.

12
MCQmedium

A platform engineer is preparing an Ubuntu 22.04 server that will host GPU-accelerated inference containers managed by containerd (not Docker). The team wants the NVIDIA Container Toolkit to expose GPUs to those containers. After installing the toolkit packages, which action must the engineer take so that containerd actually invokes the NVIDIA runtime for GPU workloads?

A.Install the nvidia-container-runtime package and symlink it as /usr/bin/runc on the host.
B.Add the user to the video group and grant read-write access to /dev/nvidiactl.
C.Set the environment variable NVIDIA_VISIBLE_DEVICES=all in the host shell profile and reboot the node.
D.Run nvidia-ctk runtime configure --runtime=containerd and restart the containerd service.
AnswerD

The nvidia-ctk runtime configure command edits the containerd configuration (typically /etc/containerd/config.toml) to register the NVIDIA runtime and set it as the default, and containerd must then be restarted to load the change. Without this registration, containerd keeps using its stock runc runtime and GPU devices never appear inside containers, even though the toolkit binaries are installed.

Why this answer

Registering the NVIDIA runtime with the container engine is the essential post-install step. The nvidia-ctk runtime configure command writes the runtime entry and default-runtime setting into containerd's config.toml, and containerd must be restarted to apply it. Merely installing packages or setting container environment variables does not change which runtime containerd uses, so GPU devices remain invisible to containers until the engine configuration is updated and reloaded.

Exam trap

The trap here is assuming that installing the NVIDIA Container Toolkit packages is sufficient and that containers automatically gain GPU access without reconfiguring the container engine's runtime.

13
MCQhard

A production inference service using TensorRT is showing lower than expected throughput. Profiling shows that the model is spending significant time in "host-to-device" transfers. What is the best optimization?

A.Increase the GPU clock speed via nvidia-smi.
B.Use pinned memory for data transfers.
C.Switch to FP64 precision for higher accuracy.
D.Implement multi-threaded data preprocessing on the GPU.
AnswerB

Pinned memory (page-locked memory) allows for faster transfer rates between the CPU and GPU because it enables the GPU to perform direct memory access without the CPU needing to copy data to a temporary buffer. This significantly reduces the overhead of host-to-device transfers in high-throughput inference pipelines.

Why this answer

Data movement is a primary latency bottleneck in deep learning inference. By using pinned memory (page-locked memory) for host-side buffers, the system can enable faster direct memory access (DMA) transfers between the CPU and the GPU. This minimizes the time spent in data copy operations, directly reducing latency and increasing total throughput, which is essential for meeting strict Service Level Agreements (SLAs) in production AI deployments.

Exam trap

Candidates often suggest optimizing the model architecture or increasing batch size, ignoring the fundamental I/O bottleneck caused by standard pageable memory transfers between the CPU and GPU host-device boundary.

14
Multi-Selecthard

Which TWO of the following actions should be taken to optimize GPU memory usage when encountering Out-of-Memory (OOM) errors during model training?

Select 2 answers
A.Implement gradient checkpointing to trade compute for memory.
B.Increase the number of CPU threads in the data loader.
C.Use mixed-precision training (FP16/BF16) to reduce weight storage.
D.Disable the use of NCCL to reduce inter-node memory overhead.
E.Switch from the Adam optimizer to Stochastic Gradient Descent (SGD).
AnswersA, C

Gradient checkpointing stores only a subset of activations and recomputes the rest during the backward pass. This significantly reduces the memory footprint of the activation graph, allowing for larger models or batch sizes to fit in memory at the cost of additional compute cycles during training.

Why this answer

OOM errors are common in deep learning when the model size, activation memory, and batch size exceed VRAM capacity. Implementing gradient checkpointing and mixed-precision training are standard industry practices to manage memory pressure. Mastery of these techniques is essential for AI operations engineers to maintain system stability and enable the training of large-scale models without requiring immediate hardware upgrades.

Exam trap

Candidates often select 'increasing batch size' or 'adding more hardware' as solutions. These actually exacerbate OOM errors or ignore the requirement to optimize memory within existing hardware constraints.

15
MCQhard

An administrator manages a shared NVIDIA cluster where several teams run inference services. One team's pods are being evicted repeatedly, and DCGM metrics show the node's GPUs are healthy but memory on the devices is nearly exhausted. The team insists their model fits. Which action should the administrator take FIRST to identify the cause?

A.Increase the pod's memory limit in the deployment manifest and restart it.
B.Replace the affected GPUs and re-run the job on fresh hardware.
C.Lower the DCGM sampling interval so metrics capture the spike more precisely.
D.Inspect the pod specs and running processes for GPU memory held by other containers on the same node.
AnswerD

When device memory is nearly exhausted but hardware is healthy, the most likely cause is co-tenancy: other containers on the same node are holding GPU memory. Because the device plugin allocates whole GPUs by default, multiple pods can land on the same device only when sharing is explicitly enabled, so checking pod specs and active processes reveals whether a neighbor is consuming the memory the team expects to have.

Why this answer

Healthy GPUs with near-exhausted device memory in a shared cluster usually indicate that another container on the same node is holding memory on the same device. The fastest way to confirm this is to inspect pod specifications and running processes to see which workloads share the GPU. Replacing hardware, changing host memory limits, or tuning telemetry do not address the allocation conflict and delay the correct diagnosis.

Exam trap

The trap here is treating GPU memory exhaustion as a hardware fault rather than as contention between co-located workloads.

16
MCQmedium

Which scheduling strategy is recommended to maximize the efficiency of long-running training jobs on preemptible instances?

A.Avoiding preemptible instances for all training jobs.
B.Implementing granular checkpointing and automated job resumption.
C.Locking the job to a specific node using node affinity.
D.Increasing the priority of the job to the maximum level.
AnswerB

Granular checkpointing minimizes the amount of lost progress during a preemption event. Coupled with automated job resumption, the system can quickly restart the task on a new node from the last checkpoint. This allows for safe usage of low-cost preemptible instances, balancing economic efficiency with the need for persistent progress.

Why this answer

Long-running jobs on preemptible (or spot) instances require frequent, efficient checkpointing to survive node reclamation. By combining a robust checkpointing schedule with smart job resubmission logic that monitors for preemption signals, organizations can take advantage of low-cost instances while minimizing the loss of progress. This approach allows for significant cost savings in non-critical training without sacrificing the overall reliability of the research project.

Exam trap

Test-takers often rely solely on high-availability cluster setups or cheaper instance pricing without establishing application-level mechanisms to preserve state when preemption inevitably occurs.

17
MCQmedium

A team is deploying a large language model for inference using NVIDIA Triton Inference Server on a GPU. They observe that the first inference request has high latency compared to subsequent requests. What is the most likely cause and the appropriate optimization?

A.The GPU is thermal throttling on the first request; set a higher power limit using nvidia-smi -pl to stabilize performance.
B.The model is not using TensorRT; converting it to a TensorRT engine will eliminate first-request latency.
C.The model uses dynamic batching; disable dynamic batching to ensure the first request is processed immediately.
D.The first request triggers model loading and CUDA context initialization; enable model warmup in Triton to pre-load and initialize the model.
AnswerD

Triton's model warmup feature allows the server to run dummy inferences during initialization, loading the model into GPU memory and initializing CUDA contexts. This moves the overhead from the first real request to server startup, reducing first-request latency. It is a standard practice for latency-sensitive deployments. The warmup can be configured with sample inputs to cover typical shapes.

Why this answer

The first inference request often incurs overhead from loading the model into GPU memory, compiling kernels, and initializing CUDA contexts. Triton's model warmup feature performs dummy inferences at startup, effectively pre-warming the model. This shifts the initialization cost away from the first real request, resulting in consistent low latency for all requests.

Exam trap

The trap here is attributing first-request latency to the inference engine or batching strategy, when it is actually caused by lazy initialization that warmup can mitigate.

18
Multi-Selecthard

An administrator manages a Kubernetes cluster where the NVIDIA GPU Operator has deployed the device plugin and MIG Manager. A tenant wants to run several small inference services that each need only a fraction of a GPU, isolated from other tenants' memory and fault domains. The administrator decides to use Multi-Instance GPU mode. Which TWO actions must be performed to make MIG-backed GPU resources schedulable to those pods? (Choose two.)

Select 2 answers
A.Install the NVIDIA Network Operator and enable RDMA device advertisement on the node
B.Enable MIG mode on the physical GPU and define a MIG profile/geometry via the MIG Manager configuration
C.Request the MIG device in the pod spec using the extended resource name advertised by the device plugin, such as nvidia.com/mig-1g.5gb
D.Create a RuntimeClass that points to the nvidia-container-runtime and reference it from each inference pod
E.Apply a ResourceQuota that limits nvidia.com/gpu to zero in the tenant namespace
AnswersB, C

MIG mode must be enabled on the GPU, and the MIG Manager in the GPU Operator applies a geometry configuration that carves the device into instances with dedicated memory and fault isolation. Without an applied profile, no MIG instances exist and the node advertises no MIG resources, so scheduling cannot succeed regardless of pod specification.

Why this answer

MIG-backed scheduling requires two coordinated steps: enabling MIG mode and applying a geometry through the MIG Manager so instances exist, and having pods request the specific profile resource name that the device plugin advertises, such as nvidia.com/mig-1g.5gb. Together these produce isolated GPU slices with dedicated memory and fault domains for each tenant's inference service.

Exam trap

The trap here is treating a RuntimeClass or ResourceQuota as the mechanism that creates MIG slices, when the essential pair is MIG geometry configuration plus requesting the advertised profile-specific extended resource.

19
MCQeasy

A cloud operations team is using NVIDIA AI Enterprise with Kubernetes to deploy inference workloads. They want to ensure that GPU resources are allocated to pods only when explicitly requested, and that pods without GPU requests do not consume GPU resources. Which Kubernetes feature should they use to enforce this behavior?

A.Use node selectors to schedule pods only on GPU nodes.
B.Define resource requests and limits for 'nvidia.com/gpu' in the pod specification.
C.Enable the GPU Operator's 'device plugin' to automatically inject GPU requests into all pods.
D.Use a mutating admission webhook to add GPU requests to pods based on their namespace.
AnswerB

In Kubernetes, GPU resources are requested using the 'nvidia.com/gpu' resource name in the pod's resource requests and limits. When a pod specifies this, the scheduler allocates a GPU to it; pods without this request are not allocated GPUs, even if scheduled on GPU nodes. This enforces explicit GPU allocation and prevents pods from consuming GPU resources unintentionally, satisfying the team's requirement.

Why this answer

Kubernetes requires pods to explicitly request GPU resources using the 'nvidia.com/gpu' resource name in their resource specifications. The scheduler then allocates GPUs only to those pods, ensuring that pods without such requests do not consume GPU resources even if they run on GPU nodes. This mechanism enforces the desired explicit allocation policy and is the standard way to manage GPU resources in Kubernetes.

Exam trap

The trap here is confusing node-level scheduling constraints with resource-level allocation; node selectors only control placement, not whether a GPU is actually assigned.

20
MCQeasy

An administrator is setting up an NVIDIA AI Enterprise cluster and wants to verify that the NVIDIA GPU Operator has successfully deployed all required components on a worker node. Which command should the administrator use to list the GPU Operator pods running on that node?

A.nvidia-smi -q -d COMPUTE
B.kubectl describe node <node-name> | grep nvidia
C.kubectl get pods -n gpu-operator --field-selector spec.nodeName=<node-name>
D.helm list -n gpu-operator
AnswerC

The GPU Operator deploys its components into the gpu-operator namespace by default, and the field selector spec.nodeName filters pods to a specific node. This command directly lists the operator pods on the target node, allowing the administrator to confirm that components such as the driver daemonset, container toolkit, and device plugin are running. It is the precise and efficient way to verify operator deployment on a worker node.

Why this answer

The NVIDIA GPU Operator runs its components as pods in the gpu-operator namespace. To confirm deployment on a specific worker node, filtering pods by spec.nodeName in that namespace lists exactly the operator pods scheduled there. This provides direct evidence that the driver, container toolkit, device plugin, and related daemonsets are running on the node.

Exam trap

The trap here is confusing Helm release status with actual pod health, when verifying operator deployment on a node requires inspecting pods filtered by node name.

21
MCQmedium

In a multi-node training scenario, what is the significance of the NVIDIA Collective Communications Library (NCCL) in workload management?

A.It handles the automatic scaling of pods in the cluster.
B.It provides a mechanism to optimize inter-node data exchange.
C.It replaces the need for high-speed network cabling.
D.It monitors the temperature of the GPUs during training.
AnswerB

NCCL is purpose-built to accelerate collective operations across distributed GPUs. It intelligently utilizes hardware interconnects such as NVLink and InfiniBand to reduce communication latency, which is the primary bottleneck in distributed training, ensuring that gradient synchronization does not throttle the overall training throughput of the job.

Why this answer

NCCL is a critical communication primitive for multi-GPU, multi-node training. It optimizes the collective operations (like AllReduce) required for synchronizing gradient updates across nodes. Effective workload management requires ensuring that nodes are configured for high-bandwidth, low-latency interconnects like NVLink and InfiniBand, which NCCL utilizes to minimize synchronization overhead.

This optimization is essential for scaling training jobs to large clusters without performance degradation caused by network bottlenecks.

Exam trap

Candidates often confuse NCCL with general network protocols, failing to recognize its specific role in optimizing collective operations like AllReduce for multi-node GPU synchronization.

22
MCQmedium

You are troubleshooting a node where the GPU is detected, but the application fails to utilize it. Which log source would provide the most relevant information?

A.The BIOS system event log.
B.The NVIDIA container runtime logs.
C.The cluster's physical network switch logs.
D.The local NTP synchronization logs.
AnswerB

The container runtime logs show the interaction between the runtime and the GPU drivers during container instantiation. If the runtime fails to inject the necessary libraries or access the GPU device, these logs will capture the error, which is the most likely cause when a GPU is physically detected.

Why this answer

The NVIDIA container runtime logs and the application-level logs are the most important sources. If the GPU is visible to the system but not the application, the issue is likely a driver/runtime mismatch or a library path configuration. Checking these logs allows an administrator to isolate whether the fault is in the container orchestration layer or the application's software environment, which is vital for rapid resolution in production environments.

Exam trap

Candidates often suggest checking the kernel logs or application code itself, overlooking that the NVIDIA container runtime is the specific layer responsible for bridging the GPU to the containerized application.

23
MCQmedium

An AI Operations engineer is managing a multi-node training job using NVIDIA NCCL. The logs indicate frequent 'NCCL WARN' messages related to 'net_ib_init' failures. What is the most likely cause of this issue?

A.Incompatible NCCL versions across the cluster nodes.
B.Misconfigured InfiniBand subnet manager or mismatched firmware.
C.Insufficient system RAM on the worker nodes.
D.Incorrect GPU driver version installed on the login node.
AnswerB

NCCL requires a correctly configured InfiniBand fabric to initialize efficiently. 'net_ib_init' errors almost exclusively result from either improper Subnet Manager configuration, outdated HCA firmware, or physical layer issues that prevent the NCCL communicator from establishing the required high-speed RDMA connections between GPUs in the cluster.

Why this answer

NCCL relies on InfiniBand (IB) for high-speed inter-GPU communication in multi-node setups. Failures in 'net_ib_init' suggest that the network configuration or driver stack is misaligned between nodes. Identifying and resolving these connectivity issues is critical because NCCL is the backbone of distributed training; any failure here leads to synchronized hangs, data corruption, or severe training slowdowns, undermining the benefits of cluster-level resource parallelism.

Exam trap

Candidates often assume 'net_ib_init' failures are application-level coding bugs within the training script, rather than recognizing them as infrastructure-level configuration issues within the InfiniBand fabric.

24
MCQhard

In an air-gapped environment, what must an administrator do to ensure the GPU Operator correctly installs the necessary software components?

A.Enable the 'offline-mode' flag in the Kubernetes API.
B.Configure the GPU Operator to use a local image registry.
C.Manually copy the VIB files to every node.
D.Disable the validation of container image signatures.
AnswerB

To succeed in air-gapped environments, the GPU Operator must be instructed to pull its images from an internally accessible registry. This involves mapping image paths and ensuring that all dependencies are hosted locally, which allows the deployment to proceed without needing external network egress to public servers.

Why this answer

In an air-gapped environment, the cluster cannot reach external repositories (e.g., NGC or Docker Hub). The administrator must pre-populate a local container registry with all required images and update the Operator configuration to point to this local registry. This is a common enterprise task for secure environments that prevents deployment failures caused by connection timeouts or authentication issues when accessing public NVIDIA software mirrors.

Exam trap

Candidates mistakenly think they can rely on standard internet-based NGC pulls by adjusting firewall rules, ignoring the true definition of an air-gapped network.

25
Multi-Selecthard

An AI operations team is troubleshooting a distributed training job on an NVIDIA DGX SuperPOD that uses NCCL for inter-GPU communication. The job intermittently hangs during the all-reduce phase. Which two actions should be taken to diagnose and resolve the issue? (Choose two.)

Select 2 answers
A.Disable the use of InfiniBand and force NCCL to use TCP sockets for communication.
B.Set NCCL_DEBUG=INFO and NCCL_DEBUG_SUBSYS=INIT,COLL to capture detailed NCCL initialization and collective logs.
C.Verify that all nodes have consistent NCCL versions and that the network interfaces used for communication are up and have sufficient bandwidth.
D.Restart the training job with a smaller number of GPUs to see if the hang persists.
E.Increase the batch size to reduce the frequency of all-reduce operations.
AnswersB, C

Enabling NCCL debug logging provides insights into the communication setup and collective operations. It can reveal misconfigurations, such as incorrect network interface selection or topology issues, that cause hangs. This is a standard first step in diagnosing NCCL-related problems, as it shows the chosen algorithm and any errors during initialization or execution.

Why this answer

Intermittent hangs in NCCL all-reduce often stem from configuration or network issues. Enabling detailed NCCL logging helps identify the exact failure point, while verifying consistent NCCL versions and network interface health addresses common root causes. These two actions together provide both diagnostic information and a path to resolution without degrading performance.

Exam trap

The trap here is thinking that reducing batch size or GPU count will solve the hang, but those are workarounds that do not diagnose or fix the underlying communication problem.

26
MCQhard

A production inference service running on NVIDIA T4 GPUs shows that GPU utilization is consistently below 20% while request latency is high. Profiling with Nsight Systems reveals that the model execution time is short but there are frequent gaps between kernels. Which optimization should be applied first to improve GPU utilization?

A.Increase the batch size in the inference server configuration.
B.Switch from FP32 to FP16 precision for the model weights.
C.Increase the number of concurrent model instances on each GPU.
D.Enable CUDA graphs to capture and replay the inference sequence.
AnswerA

Increasing batch size allows more requests to be processed per kernel launch, reducing the relative overhead of kernel launch gaps and improving GPU utilization. With small batches, the GPU sits idle between kernels. Larger batches keep the GPU busy and amortize launch overhead, directly addressing the observed gaps and low utilization. This is a standard first optimization for underutilized inference GPUs.

Why this answer

The low GPU utilization and gaps between kernels indicate that the GPU is not receiving enough work per launch. Increasing the batch size allows more data to be processed per kernel, filling the gaps and raising utilization. This is the most direct and effective first step before considering more complex techniques like CUDA graphs or precision changes.

Exam trap

The trap here is focusing on kernel launch overhead as the primary cause, when the real issue is insufficient work per launch due to small batch sizes.

27
MCQmedium

Which TWO of the following are prerequisites for installing the NVIDIA Container Toolkit on a Linux host?

A.A compatible NVIDIA driver already installed on the host.
B.The latest version of the CUDA Toolkit installed in every container.
C.A container runtime such as Docker or containerd.
D.An active subscription to NVIDIA AI Enterprise.
E.A pre-configured Kubernetes cluster with Helm.
AnswerA, C

The NVIDIA driver acts as the kernel-mode component that communicates with the hardware. The container toolkit is essentially a wrapper that relies on this driver to provide GPU access to containers. Without a working driver, the toolkit has no hardware interface to pass through to containers.

Why this answer

To successfully deploy the NVIDIA Container Toolkit, the host must have a functional NVIDIA driver and a container runtime like Docker or containerd. These prerequisites ensure that the runtime has a target to interface with and can correctly map host-side GPU resources into the container namespace. Without these core components, the toolkit cannot bridge the gap between physical hardware and isolated containerized processes.

Exam trap

Candidates select guest OS configurations or specific AI frameworks, forgetting that container toolkits strictly depend on low-level host drivers and runtimes.

28
MCQmedium

An administrator is using NVIDIA Base Command Manager to provision a new GPU cluster. They need to ensure that the compute nodes are configured with the correct GPU driver and CUDA toolkit versions. Which Base Command Manager feature should they use to automate this?

A.Ansible playbooks executed manually
B.Node groups with software images
C.PXE boot with custom kickstart scripts
D.Cron-based package updates
AnswerB

Base Command Manager allows administrators to define node groups and assign software images that include specific GPU driver and CUDA toolkit versions. During provisioning, nodes in the group automatically receive the defined software stack, ensuring consistency. This automation reduces manual configuration errors and ensures all compute nodes have the correct versions for AI workloads.

Why this answer

Base Command Manager uses node groups to categorize nodes with similar roles and software requirements. By assigning a software image to a node group, administrators define the exact GPU driver and CUDA toolkit versions. During provisioning, Base Command Manager applies the image, ensuring consistency across compute nodes.

This is the intended feature for automating software configuration in a Base Command Manager cluster.

Exam trap

The trap here is assuming that generic automation tools like Ansible or PXE are sufficient, but Base Command Manager's integrated software image and node group features are specifically designed for this scenario.

29
MCQeasy

A data scientist reports that a Jupyter notebook running on a GPU-enabled server is extremely slow when training a small neural network, even though nvidia-smi shows the GPU is idle. The notebook uses TensorFlow. Which is the most likely cause?

A.TensorFlow is not configured to use the GPU; it is running on the CPU.
B.The GPU is being used by another process, causing contention.
C.The Jupyter notebook kernel needs to be restarted to detect the GPU.
D.The neural network is too small to benefit from GPU acceleration, so TensorFlow automatically uses the CPU.
AnswerA

If the GPU is idle while training, TensorFlow is likely defaulting to CPU execution. This can happen if the GPU is not visible to TensorFlow due to missing CUDA libraries, incorrect environment variables, or a CPU-only TensorFlow installation. Checking tf.config.list_physical_devices('GPU') would confirm. This is a common oversight in notebook environments where the kernel may not have GPU access.

Why this answer

The idle GPU during training strongly suggests TensorFlow is not using it. Common causes include missing GPU support in TensorFlow, incorrect CUDA/cuDNN versions, or the process not having access to the GPU. Verifying TensorFlow's device configuration is the first step.

Other options are less likely given the evidence.

Exam trap

The trap here is assuming the GPU is too busy or the model too small, when the real issue is that TensorFlow is not configured to use the GPU at all.

30
MCQmedium

An AI operations team is deploying the NVIDIA GPU Operator on a Kubernetes cluster that uses containerd as the container runtime. The cluster nodes have NVIDIA GPUs, and the team wants to ensure that GPU workloads can request GPU resources. After installing the operator, they notice that pods requesting 'nvidia.com/gpu' remain in Pending state. Which component of the GPU Operator is most likely misconfigured or missing?

A.The NVIDIA device plugin, which advertises GPU resources to the Kubernetes API server.
B.The Kubernetes scheduler configuration, which may not be aware of GPU resources.
C.The NVIDIA Container Toolkit, which enables containers to access GPUs.
D.The GPU Operator's driver container, which loads the NVIDIA kernel modules.
AnswerA

The NVIDIA device plugin is responsible for discovering GPUs on each node and advertising them as schedulable resources like 'nvidia.com/gpu'. If it is not running or misconfigured, the Kubernetes scheduler will not see any GPU resources, causing pods that request them to remain Pending. This is the most direct cause for the described symptom.

Why this answer

The NVIDIA device plugin is a DaemonSet that runs on each node and registers GPUs as extended resources. When it is missing or misconfigured, the Kubernetes API server has no knowledge of available GPUs, so any pod requesting 'nvidia.com/gpu' cannot be scheduled. Ensuring the device plugin is healthy and running is essential for GPU scheduling.

Exam trap

The trap here is confusing the role of the NVIDIA Container Toolkit with that of the device plugin; the toolkit enables GPU access at runtime, but the device plugin is what makes GPUs visible to the scheduler.

31
MCQmedium

Which approach is most effective for scaling an inference workload that experiences sudden, unpredictable spikes in request volume?

A.Setting a fixed number of replicas to match peak expected load.
B.Using a load balancer to redirect traffic to an idle cluster.
C.Implementing HPA with custom metrics provided by DCGM Exporter.
D.Increasing the memory allocation for each inference container.
AnswerC

HPA using custom GPU metrics allows for precise scaling triggered by actual hardware usage. As request volume spikes, GPU utilization increases, and the HPA automatically triggers the deployment of additional inference pods. This ensures that the system scales only when needed, maintaining optimal performance while minimizing resource waste during quiet periods.

Why this answer

Predictive or reactive autoscaling based on custom GPU metrics (like utilization) allows the cluster to adjust replica counts in real-time. By monitoring the GPU load and scaling the inference pods accordingly, the system maintains low latency during spikes while keeping costs low during idle periods. This responsiveness is vital for production AI services where performance SLAs are tied directly to user experience.

Exam trap

Candidates frequently choose CPU-based metrics for autoscaling, not realizing that GPU-bound workloads often saturate the GPU while CPU usage remains low, rendering standard HPA configurations ineffective for scaling.

32
MCQmedium

A production inference service experiences intermittent latency spikes. The service is deployed on shared infrastructure. Which tool would best help an AI Ops engineer identify if GPU resource contention is the cause?

A.Standard Kubernetes 'kubectl top' command.
B.NVIDIA DCGM Exporter with Prometheus/Grafana.
C.The Linux 'top' command on the worker node.
D.A simple network latency test tool like 'ping'.
AnswerB

DCGM collects granular, real-time metrics directly from the GPU hardware. By exporting these to Prometheus and visualizing them in Grafana, engineers can identify spikes in GPU activity or bandwidth saturation that correlate with the application's latency, providing a clear path to identifying the source of resource contention.

Why this answer

NVIDIA DCGM (Data Center GPU Manager) provides detailed telemetry data, including GPU utilization, memory bandwidth, and power usage per process. By correlating latency spikes with DCGM metrics, engineers can pinpoint if another process on the same GPU is competing for compute resources or memory bandwidth. This visibility is essential for performance tuning and workload placement, allowing engineers to verify if resource isolation policies are functioning as expected.

Exam trap

Candidates might suggest looking at CPU logs or standard Kubernetes pod metrics. These metrics are often blind to GPU-specific resource contention, such as memory bus saturation or shared SM usage.

33
Multi-Selecthard

When troubleshooting an NVIDIA GPU Operator installation, which TWO locations should an administrator check to identify why the driver installation pod is failing?

Select 2 answers
A.The output of 'kubectl describe pod <driver-pod-name>'.
B.The contents of the /etc/kubernetes/manifests folder.
C.The logs of the driver pod using 'kubectl logs'.
D.The system-wide /var/log/syslog file on the control plane.
E.The NVIDIA license server status page.
AnswersA, C

Describing the pod reveals critical event information such as scheduling errors, image pull failures, or readiness probe failures. This is the first step in diagnosing why a pod failed to reach a 'Running' state, providing clues about potential resource constraints or registry authentication issues preventing deployment.

Why this answer

Checking the pod's logs and describe output is the standard diagnostic path. The 'describe' output shows events like ImagePullBackOff or scheduling failures, while the logs provide the specific error message from the driver installation script itself, such as kernel header mismatches or network timeouts. These two sources provide the necessary visibility to pinpoint whether the failure is infrastructure-related, configuration-based, or due to a missing environmental dependency.

Exam trap

Candidates frequently choose 'kubectl get events' or 'dmesg' on the host. While helpful, the question specifically asks for the two standard Kubernetes-native diagnostic locations for a failing pod.

34
MCQmedium

Which configuration file is typically modified to enable the NVIDIA Device Plugin in a Kubernetes cluster?

A./etc/docker/daemon.json
B.A Kubernetes DaemonSet manifest (YAML).
C./etc/nvidia/nvidia-container-runtime.json
D./etc/kubernetes/kubelet.conf
AnswerB

The NVIDIA Device Plugin is deployed as a DaemonSet in Kubernetes. The configuration for the plugin, including resource settings and feature flags, is defined within the manifest YAML file, which is then applied via kubectl to the cluster to manage the discovery and allocation of GPUs.

Why this answer

The NVIDIA Device Plugin is typically deployed as a DaemonSet using a YAML manifest. Modifying this manifest allows administrators to customize settings like the MIG strategy, time-slicing, and resource allocation policies. Correct configuration of this file is essential for ensuring that the Kubernetes scheduler correctly recognizes and assigns GPU resources to pods, which is the cornerstone of effective AI cluster management in modern cloud-native environments.

Exam trap

Candidates often confuse the NVIDIA Device Plugin with the NVIDIA Container Runtime, mistakenly believing configuration changes happen in a Docker daemon file instead of a Kubernetes DaemonSet manifest.

35
MCQmedium

When troubleshooting a NCCL collective communication timeout in a distributed training environment, which component should be the primary focus of initial investigation?

A.The GPU driver version on the master node.
B.The network interface card (NIC) topology and interconnect health.
C.The model weight initialization strategy.
D.The local disk I/O latency for checkpoint saving.
AnswerB

NCCL relies heavily on high-speed interconnects like InfiniBand or RoCE. Timeout errors typically indicate that packets are being dropped or delayed beyond the threshold defined in the NCCL configuration, pointing directly to network-related bottlenecks, faulty hardware, or incorrect fabric configuration between the participating nodes.

Why this answer

NCCL timeouts are frequently caused by network congestion, MTU mismatches, or faulty interconnect cables between nodes. By verifying the physical and logical network path, engineers can isolate whether the issue is at the application layer or the fabric layer. This focus is critical because distributed training is highly sensitive to latency and packet loss across the GPU cluster interconnect.

Exam trap

Candidates often assume NCCL timeouts are purely application bugs in the deep learning framework, leading them to unnecessarily debug model code instead of checking network hardware.

36
MCQeasy

An AI operations engineer is validating a new NVIDIA-certified server before putting it into production. The job runs correctly but the team wants to confirm that the GPUs are operating at the expected clocks and not being limited by power or thermal constraints. Which command provides the most direct evidence of the current power and thermal limits and any throttling reasons?

A.`nvidia-smi --query-gpu=name,driver_version --format=csv`
B.`nvidia-smi topo -m`
C.`nvidia-smi --gpu-reset -i 0`
D.`nvidia-smi -q -d PERFORMANCE`
AnswerD

The performance query section reports the current clocks, the active performance state, and the specific throttle reasons such as power cap, thermal limit, or hardware slowdown. This directly answers whether the GPU is being limited and why, making it the most appropriate single command for this validation step.

Why this answer

The performance query reports clocks, performance state, and explicit throttle reason flags such as power cap, thermal slowdown, and hardware limit. That combination gives direct evidence of whether a GPU is constrained and by what. Other commands either report inventory or topology and cannot answer the throttling question.

Exam trap

The trap here is choosing a command that lists GPU inventory or topology, which looks authoritative but contains no clock, power, or throttle information.

37
MCQeasy

When managing workloads on NVIDIA DGX systems, what is the primary role of the NVIDIA device plugin in the Kubernetes ecosystem?

A.To compile CUDA code for specific GPU architectures during the job scheduling process.
B.To monitor the health status of GPUs and advertise them to the Kubernetes scheduler.
C.To automatically optimize neural network hyper-parameters for faster training throughput.
D.To replace the need for the NVIDIA Container Toolkit in the container runtime environment.
AnswerB

The device plugin is the bridge that allows Kubernetes to 'see' the GPUs. It queries the NVIDIA driver for device information, reports healthy devices to the API server, and manages the lifecycle of GPU assignments, ensuring the scheduler only places GPU-dependent pods on nodes that can support them.

Why this answer

The NVIDIA device plugin acts as an interface between Kubernetes and the NVIDIA drivers, enabling the scheduler to advertise and allocate GPUs as first-class resources. It is fundamental to GPU workload management because it informs the kubelet about the availability, health, and count of GPUs on the node, ensuring that pods requesting GPUs are placed on nodes with sufficient capacity.

Exam trap

Candidates often confuse the device plugin with the container runtime. The plugin only handles hardware advertisement and discovery, not the injection of CUDA libraries or driver communication into the container.

38
MCQhard

An AI infrastructure team is preparing an air-gapped data center to install the NVIDIA GPU Operator. They have mirrored all required container images into a private registry. During installation, the Operator's pods fail with ImagePullBackOff because the components still reference images under nvcr.io. Which configuration is required to make the Operator and its managed components pull from the private registry?

A.Enable the Operator's auto-mirroring feature so it copies images from nvcr.io at runtime.
B.Set the chart values for the image repository and tag for each component, and provide an image pull secret in the Operator namespace.
C.Configure a DNS override so nvcr.io resolves to the private registry's IP address.
D.Add nvcr.io to the cluster's imagePullPolicy as Never.
AnswerB

In an air-gapped installation, every component image reference must be overridden to point at the private registry, including the Operator itself and the driver, toolkit, device plugin, and DCGM exporter images. The pull secret must exist in the namespace so kubelet can authenticate. Without these overrides the manifests still name nvcr.io and pulls fail.

Why this answer

Air-gapped installs require pre-mirroring every image and rewriting all image references in the Helm values to the private registry, including the Operator and each managed component. Authentication to the private registry is supplied through an image pull secret placed in the Operator namespace so the kubelet can pull. DNS tricks, pull policy changes, or imagined auto-mirroring do not redirect image references and leave the components failing.

Exam trap

The trap here is believing that a network-level redirection such as DNS or a pull policy change can substitute for rewriting the actual image repository references and supplying registry credentials.

39
MCQeasy

A platform engineer is preparing a bare-metal Ubuntu 22.04 server with four A100 GPUs for an NVIDIA AI Enterprise deployment. The GPUs are not yet visible to the operating system tooling. Which command should the engineer run to confirm the driver loaded successfully and that all four GPUs are enumerated with their current driver version?

A.dcgmi discovery -l
B.nvidia-container-cli --info
C.nvidia-smi
D.lspci -d 10de:
AnswerC

nvidia-smi queries the loaded NVIDIA kernel driver and prints every enumerated GPU together with the driver version and CUDA version. On a freshly prepared AI Enterprise node it is the canonical first check that the driver bound to all four A100 devices before any container runtime or GPU Operator work begins.

Why this answer

The NVIDIA kernel driver is the foundation of every AI Enterprise workload, and nvidia-smi is the standard utility that both triggers a driver query and renders the full GPU inventory with driver and CUDA versions. Administrative toolkits such as DCGM or the container toolkit sit above the driver and only report meaningful data once the driver has bound successfully to all adapters.

Exam trap

The trap here is assuming that seeing the GPUs in a PCI listing proves the driver is installed, when PCI enumeration and driver binding are independent states.

40
MCQmedium

A platform team runs an NVIDIA GPU Operator-managed Kubernetes cluster shared by two research groups. Group A's pods request `nvidia.com/gpu: 1` and are scheduled correctly, but Group B's pods that omit any GPU resource request are also landing on GPU nodes and consuming host memory and CPU, degrading Group A's jobs. The team wants Group B's non-GPU pods to stop consuming capacity on the GPU node pool without changing Group B's manifests. Which action should the administrator take?

A.Set `NVIDIA_DRIVER_CAPABILITIES=compute,utility` on the GPU Operator DaemonSet so non-GPU pods cannot attach to the device.
B.Create a `ResourceQuota` in each namespace that caps `requests.nvidia.com/gpu` at zero for Group B.
C.Enable the `nvidia` device plugin's `--fail-on-init-error=false` flag so unresolvable GPU requests fall back to CPU scheduling.
D.Apply a taint such as `nvidia.com/gpu=present:NoSchedule` to the GPU nodes and add a matching toleration to Group A's pod templates.
AnswerD

Tainting GPU nodes with a dedicated key and NoSchedule effect prevents any pod lacking a matching toleration from being placed there, so Group B's unmodified workloads are repelled while Group A's templates are explicitly admitted. This is the standard scheduler-level control for dedicating a node pool, and it requires no edits to Group B's manifests, matching the stated constraint exactly.

Why this answer

Dedicating a node pool to GPU consumers is done at the scheduler layer: a taint on the GPU nodes repels every pod that does not carry a matching toleration. Because Group B's manifests cannot be changed, repelling rather than attracting is the only workable direction, and adding tolerations to Group A's templates is a one-time change under the platform team's control. Quotas, driver flags, and plugin options all operate after or outside scheduling and cannot keep non-GPU pods off those nodes.

Exam trap

The trap here is assuming that a ResourceQuota or device-plugin setting controls node placement, when only taints, tolerations, and node affinity influence the scheduler's decisions.

41
MCQhard

A financial services company is deploying NVIDIA AI Enterprise on a VMware vSphere cluster with NVIDIA A100 GPUs. The security team requires that GPU workloads be isolated at the hardware level, with separate memory and fault domains, to meet regulatory compliance. The company also wants to maximize GPU utilization by running multiple workloads concurrently. Which NVIDIA feature should be enabled to meet these requirements?

A.NVIDIA NVLink with SHARP
B.NVIDIA vGPU with time-sliced scheduling
C.NVIDIA GPUDirect Storage
D.NVIDIA Multi-Instance GPU (MIG)
AnswerD

MIG partitions an A100 GPU into up to seven independent instances, each with dedicated memory, cache, and compute resources. This provides hardware-level isolation and fault domain separation, satisfying regulatory requirements for workload isolation. It also allows multiple workloads to run concurrently on a single GPU, maximizing utilization. MIG is the correct choice for this scenario because it uniquely combines isolation and concurrency on A100 hardware.

Why this answer

NVIDIA Multi-Instance GPU (MIG) is the only feature that provides hardware-level partitioning of an A100 GPU into isolated instances with dedicated memory, cache, and compute resources. This meets the need for fault domain separation and regulatory compliance while enabling concurrent execution of multiple workloads, thereby maximizing utilization. Other options either share resources without isolation or address different concerns like data transfer or inter-GPU communication.

Exam trap

The trap here is confusing time-sliced vGPU or other GPU sharing technologies with true hardware partitioning, which only MIG provides on A100 GPUs.

42
MCQhard

An AI operations team runs mixed training and inference workloads on a Kubernetes cluster managed with the NVIDIA GPU Operator. Inference pods frequently arrive in bursts and must start within seconds, while long-running training jobs occupy most MIG-capable A100 GPUs for days. Administrators want burst inference pods to obtain GPU capacity immediately without preempting or restarting the training jobs, and they want the cluster to reclaim those resources automatically when the burst ends. Which approach best satisfies these requirements?

A.Create a node pool reserved for inference, and configure a cluster autoscaler with scale-to-zero so burst pods trigger new GPU nodes that are removed when idle.
B.Configure the inference Deployment with a topologySpreadConstraint across all GPU nodes so its replicas distribute evenly and find free capacity.
C.Enable the GPU Operator's time-slicing configuration so inference and training containers share the same physical GPUs through a shared device plugin.
D.Deploy inference pods with a higher PriorityClass and enable preemption so the scheduler evicts lower-priority training pods when capacity is exhausted.
AnswerA

A dedicated inference node pool with scale-to-zero autoscaling lets burst pods trigger immediate node provisioning while training jobs keep running untouched on their own nodes. When the burst subsides, the autoscaler drains and removes the idle inference nodes, reclaiming GPU capacity and cost automatically. This directly matches both requirements: no preemption of training and automatic reclamation after the burst.

Why this answer

The scenario requires two properties simultaneously: burst inference capacity that appears without disturbing running training jobs, and automatic reclamation when the burst ends. Only a dedicated autoscaled node pool satisfies both, because new GPU nodes are provisioned on demand for inference pods while training pods remain scheduled and running on their existing nodes. When demand drops, scale-to-zero removes the extra nodes, freeing GPU resources without any preemption or job restarts.

Exam trap

The trap here is assuming that priority and preemption are the natural answer for burst workloads, when the requirement that training jobs never restart makes preemption disqualifying.

43
MCQeasy

Which NVIDIA tool allows you to verify that the GPU and its driver are properly installed and functioning on a Linux system?

A.nvcc --version
B.nvidia-smi
C.lspci | grep nvidia
D.docker run --gpus all
AnswerB

This command is the primary tool for verifying that the NVIDIA driver is loaded and communicating correctly with the GPU. It provides essential diagnostic information, including device names, driver versions, and current memory usage, which are the fundamental metrics for confirming that a GPU installation was successful.

Why this answer

The 'nvidia-smi' (System Management Interface) tool is the standard utility for interacting with the NVIDIA driver. It provides a real-time status of GPU utilization, temperature, memory usage, and driver versions. Being proficient with this tool is essential for an AI Ops professional to quickly validate hardware health, confirm driver installation, and identify if a GPU is accessible by the host OS after a fresh installation or reboot.

Exam trap

Candidates often confuse nvidia-smi with higher-level management tools like NVIDIA AI Enterprise or DCGM. They overlook that nvidia-smi is the fundamental, low-level command for basic driver and hardware verification.

44
MCQmedium

An administrator is optimizing a large model training job to reduce checkpointing time to storage. Which strategy is most effective for minimizing the impact on training throughput?

A.Use faster NVMe drives for checkpoint storage.
B.Implement asynchronous checkpointing.
C.Increase the checkpoint frequency.
D.Compress the model weights before saving.
AnswerB

Asynchronous checkpointing allows the training process to save the model state in the background without blocking the main training loop. By offloading the serialization and I/O to a separate thread, the GPU remains free to continue compute-intensive operations, significantly increasing overall training throughput and reducing total job wall-clock time.

Why this answer

Checkpointing large models involves writing gigabytes of data to storage, which can pause training for significant periods. Using asynchronous checkpointing or offloading the save process to a background thread allows the GPU to continue training while the data is written to persistent storage. This eliminates the I/O wait time, ensuring that the heavy computational resources are not wasted during the periodic saving of model states.

Exam trap

Candidates often suggest faster storage hardware. While helpful, it does not solve the fundamental issue of the GPU being blocked by synchronous I/O operations during the checkpointing process.

45
MCQhard

A multi-tenant NVIDIA GPU-accelerated Kubernetes cluster utilizing NVIDIA AI Enterprise experiences intermittent out-of-memory errors on Triton Inference Server pods despite adequate node memory reservation. Which monitoring and troubleshooting action correctly isolates the root cause?

A.Analyze cAdvisor container memory metrics using kubectl top pods to evaluate resident set size growth over sustained inference workloads.
B.Review the Kubernetes cluster autoscaler logs to determine if node scaling events triggered unexpected pod evictions and restarts.
C.Inspect DCGM-Exporter metrics for DCGM_FI_DEV_FB_FREE and DCGM_FI_DEV_GPU_TEMP alongside Triton server request concurrency logs.
D.Execute systemctl status containerd on the worker node to verify container runtime stability and daemon responsiveness.
AnswerC

Tracking free frame buffer metrics and GPU temperatures via DCGM-Exporter reveals exact memory headroom and potential throttling conditions during peak concurrency. Correlating these metrics with inference request logs confirms whether dynamic batching parameters exceeded available device memory.

Why this answer

Inspect NVIDIA Data Center GPU Manager metrics via Prometheus and DCGM-Exporter to capture real-time device memory consumption patterns. This practice is critical because standard Kubernetes container metrics fail to track internal GPU frame buffer allocations and fragmentation specific to deep learning inference engines.

Exam trap

Many administrators mistakenly rely solely on standard cAdvisor memory metrics exposed by Kubernetes, completely missing GPU-specific memory exhaustion occurring inside the device driver context.

46
MCQmedium

An AI administrator is tasked with monitoring GPU utilization in a multi-user cluster. Which tool provides the most granular real-time visibility into process-level GPU memory usage and compute utilization?

A.The system 'top' utility.
B.The NVIDIA 'nvidia-smi' tool.
C.The 'df' command.
D.The system network logs.
AnswerB

NVIDIA-SMI is the standard interface for querying the status of NVIDIA GPU devices. It displays real-time statistics including per-process memory consumption, duty cycle, and power usage. This granularity is essential for identifying which specific user processes are saturating the GPU or causing resource conflicts in a multi-user environment.

Why this answer

The 'nvidia-smi' utility is the foundational tool for monitoring NVIDIA GPU hardware. It provides direct, real-time feedback on memory usage, compute utilization, and process-level diagnostics. For cluster-wide administration, it is the primary command-line tool used to identify exactly which processes are consuming resources, allowing administrators to troubleshoot contention and manage workloads effectively across the available GPU hardware resources.

Exam trap

Candidates often suggest high-level monitoring dashboards or cloud-native tools. While useful, they lack the immediate, process-level granularity required for troubleshooting specific resource contention on a local DGX node.

47
MCQmedium

Which file format is commonly used to define the configuration and state for the NVIDIA GPU Operator within a Kubernetes environment?

A.JSON
B.YAML
C.TOML
D.XML
AnswerB

YAML is the native format for Kubernetes configurations. The GPU Operator relies on YAML files to define the ClusterPolicy and other custom resources. This structure allows administrators to define the desired state of their GPU infrastructure, enabling automated reconciliation and lifecycle management by the operator.

Why this answer

Kubernetes uses YAML to define declarative states for resources, including Custom Resource Definitions (CRDs) used by operators. The GPU Operator leverages YAML manifests to configure settings such as driver version, toolkit installation, and monitoring agents. Understanding this format is vital for AI Ops professionals, as it allows for version-controlled infrastructure-as-code practices, ensuring deployments are reproducible and auditable across various environments.

Exam trap

Candidates might assume binary or proprietary configuration formats are used, forgetting that Kubernetes operators rely heavily on standard declarative YAML files for managing Custom Resource Definitions.

48
MCQhard

Refer to the exhibit. An administrator notices that 'user_a' is consistently hitting resource limits despite having sufficient total system GPU memory. Based on the policy JSON, what is the cause?

A.The system has reached the max_concurrent_jobs limit.
B.The user is exceeding the 16GB VRAM limit.
C.The NVIDIA driver is blocking the user's access.
D.The GPU memory is fragmented across the nodes.
AnswerB

The JSON clearly defines a resource quota of '16GB'. When the user's training or inference job requests memory beyond this allocated limit, the enforcement mechanism blocks further allocation. This is a deliberate configuration to ensure fair resource distribution among multiple users in a shared GPU cluster environment.

Why this answer

The policy specifies a hard resource quota of '16GB' for the user. Even if the total system memory is higher, the scheduler enforces this limit per user. When the user's workload attempts to allocate more than 16GB of VRAM, the job will fail or throttle.

This is a common method for preventing single users from starving the entire cluster of shared resources.

Exam trap

Candidates often look at the total cluster capacity rather than the specific user-level quota. They assume the user can access all available VRAM, ignoring the scheduler's hard limit policy.

49
MCQhard

An inference model running on Triton Inference Server is reporting high latency for requests. The model uses a fixed-size batching strategy. What is the most effective way to optimize throughput while maintaining latency targets?

A.Increase the number of instances for the model.
B.Configure dynamic batching with a maximum delay.
C.Disable all logging to reduce CPU overhead.
D.Force the model to run on the CPU.
AnswerB

Dynamic batching allows the server to aggregate multiple requests into a single inference call. Setting a maximum delay ensures that the server waits just long enough to improve throughput without exceeding the latency budget. This is the standard method to optimize GPU inference workloads on Triton Inference Server.

Why this answer

Dynamic batching is a powerful feature in Triton Inference Server that groups individual requests together to saturate the GPU's compute capability. By configuring the 'max_queue_delay_microseconds', the system waits briefly to aggregate requests, significantly increasing throughput. This optimization is crucial for balancing the trade-off between individual request latency and overall system efficiency, ensuring that the GPU is not performing trivial computations for tiny batches.

Exam trap

Candidates often choose static batch resizing or manual request throttling, which fails to automatically adapt to fluctuating incoming request rates and traffic spikes.

50
MCQmedium

An administrator needs to collect GPU telemetry from an NVIDIA AI Enterprise cluster and store it in a time-series database for long-term analysis. Which component should be deployed to export GPU metrics in Prometheus format?

A.NVIDIA Container Toolkit
B.NVIDIA Base Command Manager
C.NVIDIA DCGM Exporter
D.NVIDIA GPU Operator
AnswerC

DCGM Exporter is designed to collect GPU metrics using DCGM and expose them in Prometheus format. It runs as a container and provides metrics such as GPU utilization, memory usage, temperature, and power. These metrics can be scraped by Prometheus and stored in a time-series database for analysis. This is the standard solution for exporting GPU telemetry in Kubernetes environments.

Why this answer

DCGM Exporter is the component that collects GPU metrics via DCGM and exposes them as Prometheus metrics. It is typically deployed as a daemonset or pod and can be scraped by Prometheus for long-term storage and analysis. Other NVIDIA components manage deployment or container runtime integration but do not provide metric export functionality.

Exam trap

The trap here is assuming that the GPU Operator or Base Command Manager exports metrics, when metric export is specifically the role of DCGM Exporter.

51
MCQmedium

When deploying NVIDIA AI Enterprise, why is the use of the NVIDIA NGC Catalog recommended over public container repositories?

A.NGC images are the only way to bypass the need for a valid NVIDIA license.
B.NGC images are pre-configured to automatically perform hardware diagnostics.
C.NGC images are curated and optimized for NVIDIA hardware compatibility.
D.NGC images can be deployed without any need for the NVIDIA Container Toolkit.
AnswerC

NGC images are specifically curated to ensure that all required CUDA libraries and drivers are compatible. This optimization guarantees that the software stack works as intended with NVIDIA GPUs, preventing the 'DLL hell' or library version conflicts often encountered when assembling container environments from generic public images and manual installs.

Why this answer

The NGC Catalog provides verified, performance-optimized, and security-scanned container images specifically tailored for the NVIDIA AI Enterprise stack. These images contain all necessary dependencies and are guaranteed to be compatible with supported GPU drivers. Using these images reduces the risk of deployment failures caused by library version mismatches, missing dependencies, or unoptimized software versions that are common in generic, public repositories, thereby ensuring a reliable production-grade AI environment.

Exam trap

Candidates mistakenly believe that public repositories contain identical images to NGC. They ignore the performance tuning and security scanning inherent in curated NGC images that prevent runtime bottlenecks.

52
Multi-Selecthard

An administrator is configuring an NVIDIA AI Enterprise cluster to run multi-tenant inference workloads on Kubernetes. The administrator must ensure that GPU resources are isolated and that tenants cannot access each other's GPU memory. Which two actions should the administrator take? (Choose two.)

Select 2 answers
A.Use Kubernetes ResourceQuota and LimitRange objects to restrict GPU memory per namespace.
B.Configure NVIDIA vGPU with different vGPU profiles for each tenant.
C.Deploy the NVIDIA GPU Operator with the device plugin configured to advertise MIG resources.
D.Enable Multi-Instance GPU (MIG) mode on supported GPUs and assign separate MIG instances to each tenant.
E.Enable time-slicing of GPUs so that multiple tenants share the same GPU context.
AnswersC, D

For Kubernetes to schedule workloads onto MIG instances, the NVIDIA device plugin must advertise each MIG slice as a schedulable resource, which the GPU Operator configures via its MIG strategy setting. Without this, MIG instances exist on the GPU but are invisible to the scheduler. Enabling the device plugin to expose MIG resources is therefore a necessary action to make tenant isolation effective.

Why this answer

Hardware-level isolation for multi-tenant inference on Kubernetes is achieved by enabling MIG on supported GPUs and having the GPU Operator's device plugin advertise MIG instances as schedulable resources. Together these actions partition the GPU into isolated slices and make those slices available to tenants without memory sharing, satisfying the isolation requirement.

Exam trap

The trap here is assuming that Kubernetes quotas or time-slicing provide GPU memory isolation, when only MIG creates hardware-partitioned instances with dedicated memory.

53
MCQmedium

An organization is deploying large-scale LLM training workloads on an NVIDIA DGX SuperPOD. The data science team reports that training jobs are frequently preempted by higher-priority batch jobs, leading to significant checkpointing overhead. Which Workload Manager configuration strategy best minimizes resource fragmentation and improves overall cluster utilization while maintaining SLA requirements?

A.Increase the default job preemption grace period to allow all jobs to complete.
B.Disable multi-instance GPU (MIG) to allow all jobs to access full GPU memory.
C.Implement gang scheduling combined with strict job priority and preemption policies.
D.Manually partition the cluster into static zones for different research teams.
AnswerC

Gang scheduling ensures that all requested resources for a distributed job are acquired simultaneously, preventing deadlocks where multiple jobs wait for partial allocations. Combining this with clear priority levels allows the scheduler to preempt lower-priority tasks efficiently, maximizing cluster utility while protecting SLA-bound high-priority distributed training workloads.

Why this answer

Implementing gang scheduling with preemption thresholds ensures that resources are allocated atomically to distributed training jobs, preventing partial allocations that lead to starvation. By configuring preemption grace periods and job priorities, the scheduler can effectively balance urgent tasks against long-running training runs. This approach is critical in NVIDIA environments to ensure that high-bandwidth inter-node communication remains optimized during heavy compute cycles, preventing inefficient resource usage.

Exam trap

Candidates often suggest simple priority queues without gang scheduling. This causes 'fragmentation' where a job starts with only half its required GPUs, leading to a deadlock or inefficient training.

54
MCQeasy

Why is it important to use a persistent storage volume for model checkpoints in a distributed training job?

A.To increase the write speed of the training data.
B.To ensure model checkpoints survive pod failures.
C.To provide a high-speed cache for real-time inference.
D.To hide the model from unauthorized cluster users.
AnswerB

Pod failures are common in orchestrated environments due to node errors or preemptions. By storing checkpoints on persistent storage, the state is decoupled from the lifecycle of the individual pod, allowing a new pod instance to pick up exactly where the previous process stopped during training.

Why this answer

Distributed training jobs are prone to node failures or preemptions in shared environments. Checkpointing allows the training state to be saved periodically. By storing these checkpoints on persistent volumes, the job can resume from the last saved state rather than restarting from scratch if a failure occurs.

This minimizes wasted compute time and ensures that long-running training tasks can eventually complete, protecting the investment in expensive cluster resources.

Exam trap

Candidates often think persistent storage is for performance or speed. In reality, it is purely for fault tolerance and state persistence during inevitable pod crashes or preemptions.

55
MCQmedium

Which tool is the industry standard for monitoring and managing NVIDIA data center GPUs in a large-scale cluster deployment?

A.nvidia-smi
B.NVIDIA DCGM
C.CUDA Debugger
D.NVIDIA Triton Inference Server
AnswerB

DCGM is the enterprise-grade solution for managing and monitoring GPUs. It supports cluster-wide health monitoring, performance profiling, and configurable diagnostic tests. It is the core component for NVIDIA AI operations, providing the telemetry needed for dashboarding and automated resource management at scale.

Why this answer

The NVIDIA Data Center GPU Manager (DCGM) provides a comprehensive set of APIs and tools for managing and monitoring GPUs in production environments. It is essential for AI operations because it allows for real-time health checks, diagnostic testing, and policy-based management of GPU resources across a distributed cluster, ensuring high availability and identifying performance bottlenecks before they impact critical training or inference workflows.

Exam trap

Candidates often select generic monitoring tools like Prometheus or Grafana. While these visualize data, they are not the primary NVIDIA-specific engine used to manage and diagnose GPU hardware health.

56
MCQhard

Refer to the exhibit. The cluster administrator observes near-capacity memory utilization across three GPUs. What is the most likely consequence if an additional pod is scheduled to these GPUs without memory partitioning?

A.The scheduler will automatically enable MIG.
B.The job will experience OOM errors and fail.
C.The scheduler will use CPU swap memory.
D.The job will be throttled but complete successfully.
AnswerB

With memory utilization already exceeding 95% on all GPUs, there is insufficient headroom to spawn a new process. The operating system or the NVIDIA driver will trigger an Out-of-Memory (OOM) kill signal, resulting in the immediate failure of the new container and potential instability for existing tasks.

Why this answer

Scheduling an additional pod onto these GPUs will lead to Out-of-Memory (OOM) errors. Because the GPUs are already at near-maximum capacity, the driver will likely kill the incoming process or potentially destabilize existing tasks. This highlights the importance of workload monitoring and resource planning, as AI Ops must ensure that requests match actual hardware capacity to prevent catastrophic job failures in production environments.

Exam trap

Candidates might guess that the system will simply slow down or swap to system RAM, forgetting that GPU memory is non-swappable and will cause immediate process termination via OOM error.

57
MCQmedium

An AI platform team runs GPU-accelerated inference pods on a Kubernetes cluster with the NVIDIA GPU Operator. During peak load, high-priority latency-sensitive inference pods are frequently preempted by large batch training jobs that were submitted later. The team wants the scheduler to guarantee that inference pods always win placement and preemption decisions against training pods without manually cordoning nodes. Which mechanism should the administrator configure?

A.Enable the GPU Operator's time-slicing feature with a replica count greater than one so inference and training pods share each physical GPU concurrently.
B.Create a separate namespace with a ResourceQuota that limits training pods to a small number of GPUs and place all inference pods in the default namespace.
C.Assign the inference pods a higher PriorityClass value and reference it in their pod spec so the kube-scheduler preempts lower-priority training pods when resources are scarce.
D.Add a nodeSelector to the inference pods that pins them to nodes labeled with the NVIDIA GPU product name, so they only run on nodes not used by training.
AnswerC

PriorityClass is the native Kubernetes scheduling mechanism that orders pending pods; a higher integer value makes the kube-scheduler prefer those pods and preempt lower-priority workloads that already occupy GPUs. Referencing the class in the pod spec ensures inference pods win placement and preemption decisions, which is exactly what the team needs.

Why this answer

Kubernetes resolves GPU contention through the scheduler's priority and preemption logic, not through device sharing or quota limits. Giving inference pods a higher PriorityClass value makes the kube-scheduler place them first and evict lower-priority training pods when capacity is exhausted, which delivers the guaranteed preference the platform team requires.

Exam trap

The trap here is assuming that GPU sharing features such as time-slicing or MIG establish workload precedence, when only PriorityClass plus scheduler preemption actually reorders and evicts competing pods.

58
MCQhard

In an NVIDIA-accelerated Kubernetes environment, why is it critical to configure a 'RuntimeClass' for GPU-enabled pods?

A.To allow the Kubernetes scheduler to place pods on nodes based on GPU availability.
B.To define the specific NVIDIA driver version required by the application.
C.To ensure the container runtime correctly injects GPU drivers and libraries into the pod.
D.To increase the priority of GPU workloads over standard CPU-only workloads.
AnswerC

The RuntimeClass enables the node to switch to the NVIDIA-specific container runtime when needed. This runtime is responsible for the crucial task of injecting the required driver and library files, enabling the container to recognize and interact with the GPU hardware without requiring manual configuration in every single pod.

Why this answer

The RuntimeClass ensures that the container is started with the correct NVIDIA container runtime (e.g., nvidia-container-runtime). This runtime is responsible for performing the necessary hardware mapping, injecting the CUDA libraries, and setting the environment variables required for the container to actually utilize the GPU. Without this configuration, the pod might be scheduled successfully but will fail at runtime because it cannot communicate with the NVIDIA driver.

Exam trap

Candidates often assume that requesting a GPU resource in the pod spec is sufficient, forgetting that the RuntimeClass is the mechanism that actually injects the necessary NVIDIA libraries.

59
MCQmedium

Which action is required when updating the NVIDIA driver on a node managed by the GPU Operator to ensure that running workloads are not interrupted abruptly?

A.Manually stop all running pods.
B.Use the operator to perform a rolling update.
C.Reboot the entire cluster simultaneously.
D.Delete the node object from the cluster.
AnswerB

The GPU Operator automates the rolling update process, which includes cordoning and draining nodes to move workloads before applying driver upgrades. This ensures that the maintenance happens without unexpected service outages, fulfilling the requirement to manage infrastructure updates safely while maintaining high availability for the dependent AI workloads.

Why this answer

The GPU Operator supports seamless driver upgrades by cordoning and draining nodes. When an update is initiated, the operator gracefully moves workloads to other available nodes before updating the driver. This process prevents application crashes, ensures data integrity, and maintains cluster stability during maintenance windows, which is a key responsibility for AI operations professionals managing production-grade, long-running AI training or inference tasks on shared GPU resources.

Exam trap

Candidates often suggest manually stopping pods or deleting the node, forgetting that the GPU Operator provides automated rolling update capabilities that handle pod draining and cordoning safely.

60
Multi-Selecthard

An AI operations team is validating a new Kubernetes cluster before installing the NVIDIA GPU Operator with the driver managed by the Operator itself. The nodes run a supported Linux distribution with GPUs physically installed. Which two conditions must be satisfied for the Operator's driver container to build and load the kernel module successfully? (Choose two.)

Select 2 answers
A.Each node must have at least one Multi-Instance GPU profile pre-created.
B.The kubelet must be configured with the NVIDIA device plugin endpoint flag.
C.Secure Boot must be disabled or the module must be signed with an enrolled key.
D.The cluster must have the NVIDIA Network Operator installed first.
E.The node's kernel headers for the running kernel must be available to the driver container.
AnswersC, E

When UEFI Secure Boot is enabled, the kernel refuses to load unsigned modules. The NVIDIA driver container either needs Secure Boot turned off or must sign the generated module with a Machine Owner Key that is enrolled in the firmware. Without one of these, module insertion is rejected and the GPU stays unusable even though the build completed.

Why this answer

For the Operator to manage drivers, the driver container must compile the NVIDIA kernel module against the exact running kernel, which requires matching kernel headers and build dependencies on the node. In addition, UEFI Secure Boot blocks unsigned modules, so it must be disabled or the module signed with an enrolled key. These two conditions gate whether the module can be built and inserted; without them the node cannot advertise GPU capacity to the scheduler.

Exam trap

The trap here is treating unrelated cluster add-ons, such as the Network Operator or MIG profiles, as prerequisites for the driver container when the real gating factors are kernel headers and module signing.

61
MCQmedium

Refer to the exhibit. The training job fails with a CUDA OOM error. Given the memory profile, which optimization strategy provides the most immediate relief while maintaining model performance?

A.Increase the batch size to 256.
B.Enable Automatic Mixed Precision (AMP).
C.Disable all CUDA kernels.
D.Replace the GPU with a higher-clocked model.
AnswerB

Mixed precision reduces memory usage by using FP16 or BF16 for most operations, which requires less memory than FP32. This approach allows the training process to utilize significantly less VRAM, providing the necessary headroom to avoid the current OOM error while keeping the model training pipeline functional.

Why this answer

Switching to Mixed Precision (FP16/BF16) effectively halves the memory requirement for model activations and gradients. This is a standard optimization strategy for Transformer models when GPU memory capacity is the primary constraint. By reducing the memory footprint of floating-point operations, researchers can fit larger batch sizes or deeper architectures into the same GPU memory, significantly improving training efficiency and throughput without sacrificing convergence stability.

Exam trap

Candidates often choose complex model architecture changes like gradient checkpointing or model parallelism, which are harder to implement than enabling AMP for immediate memory relief.

62
MCQeasy

What is the primary role of the NVIDIA Data Center GPU Manager (DCGM) Exporter in a cloud-native monitoring stack?

A.To act as a load balancer for GPU-accelerated traffic.
B.To provide a Prometheus-compatible endpoint for GPU metrics.
C.To manage the deployment of containerized AI models.
D.To encrypt data transmission between the GPU and the CPU.
AnswerB

The DCGM Exporter collects raw telemetry data from the DCGM API and exposes it as a scrapeable endpoint for Prometheus. This enables administrators to visualize GPU health, track usage trends, and set alerts for thresholds like high temperature or memory usage within a Grafana dashboard.

Why this answer

The DCGM Exporter is the bridge between NVIDIA hardware metrics and monitoring systems like Prometheus. It gathers real-time telemetry from DCGM and exposes it in a format that Prometheus can scrape. This integration is vital for AI operations because it provides observability into GPU utilization, memory usage, and thermal health, enabling automated alerts and performance dashboards that are essential for maintaining a production-grade AI cluster.

Exam trap

Candidates confuse the DCGM Exporter with the underlying metrics collector tool itself, missing its specific role in formatting data for Prometheus scraping.

63
MCQhard

Refer to the exhibit. An administrator applies this security policy to a container runtime environment. What is the immediate effect on containerized AI applications within this scope?

A.The container gains elevated privileges to access the host GPU
B.The container is unable to detect or communicate with the NVIDIA GPU
C.The container can use the GPU but cannot perform memory mapping
D.The container experiences increased latency for GPU operations
AnswerB

The NVIDIA driver requires access to specific device nodes under /dev/nvidia* to function. By denying access to these files, the runtime prevents the container's CUDA libraries from establishing a connection to the GPU driver, rendering the GPU invisible to the application code executing inside the container environment.

Why this answer

The policy explicitly denies access to the character device files associated with the NVIDIA GPU. Without access to /dev/nvidia0, /dev/nvidiactl, and /dev/nvidia-uvm, the CUDA runtime cannot interact with the GPU hardware. Consequently, any attempt to initialize a CUDA device will fail, causing the application to crash.

This policy is a common restrictive measure in high-security environments where GPU access must be strictly managed or audited.

Exam trap

Candidates often assume that security policies only affect network traffic or file system access, overlooking that blocking character device nodes directly breaks the CUDA runtime's ability to initialize hardware.

64
MCQeasy

An administrator must run a batch inference job that requires exactly two NVIDIA GPUs on a Kubernetes cluster managed by the NVIDIA GPU Operator. Which pod specification field should be used to request those GPUs?

A.resources.requests with the key nvidia.com/gpu set to 2 and no limit
B.annotations with the key nvidia.com/gpu.count set to 2
C.nodeSelector with the key nvidia.com/gpu.count set to 2
D.resources.limits with the key nvidia.com/gpu set to 2
AnswerD

The NVIDIA device plugin advertises GPUs as the extended resource nvidia.com/gpu, and extended resources must be requested through resources.limits. Setting the limit to 2 causes the scheduler to place the pod on a node with two allocatable GPUs and injects those devices into the container, which meets the exact requirement.

Why this answer

GPUs exposed by the NVIDIA device plugin appear to Kubernetes as the extended resource nvidia.com/gpu. Extended resources must be declared in resources.limits, and setting that limit to 2 makes the scheduler reserve two GPUs on a suitable node and pass the devices into the container.

Exam trap

The trap here is treating nvidia.com/gpu like CPU or memory and specifying it only under resources.requests, which the API server rejects for extended resources.

65
MCQmedium

An AI operations engineer is preparing a Kubernetes cluster to run GPU-accelerated workloads using the NVIDIA GPU Operator. The cluster nodes already have NVIDIA data center GPUs installed and the NVIDIA driver is pre-installed on the host. The engineer wants to use the GPU Operator to manage the container toolkit, device plugin, and monitoring components but must avoid the Operator managing or upgrading the driver. Which configuration should be applied to the GPU Operator deployment?

A.Use the NVIDIA GPU Operator with the --set toolkit.enabled=false option to prevent driver installation.
B.Set the driver.enabled parameter to false in the GPU Operator's Helm chart values.
C.Install the GPU Operator with the --set operator.driverVersion=latest flag to pin the driver version.
D.Deploy the GPU Operator but remove the nvidia-driver-daemonset after installation.
AnswerB

Setting driver.enabled=false instructs the GPU Operator to skip deploying the driver container and instead rely on the pre-installed host driver. This is the supported method for clusters where the driver is managed externally, such as via the node's package manager. The Operator will still deploy the container toolkit, device plugin, and DCGM exporter, allowing full GPU scheduling and monitoring without touching the driver.

Why this answer

When the NVIDIA driver is already present on host nodes and should remain externally managed, the GPU Operator must be configured to not deploy its own driver container. The driver.enabled=false Helm value achieves this by skipping the driver daemonset while still deploying other components like the container toolkit, device plugin, and DCGM exporter. This preserves the existing driver and avoids conflicts.

Exam trap

The trap here is assuming that specifying a driver version or disabling the toolkit will prevent driver installation, when only the driver.enabled flag controls driver management.

66
MCQeasy

A platform engineer is preparing an NVIDIA-accelerated Kubernetes cluster for a new team that will submit PyTorch training jobs. The team wants jobs to request GPUs without hardcoding device indices. Which Kubernetes resource should the engineer ensure is installed and healthy so pods can request nvidia.com/gpu resources?

A.NVIDIA Container Toolkit installed on each worker node
B.NVIDIA DCGM Exporter deployed as a DaemonSet
C.NVIDIA Network Operator with RDMA shared device plugin
D.NVIDIA GPU Operator with the device plugin component enabled
AnswerD

The NVIDIA device plugin, deployed by the GPU Operator, registers nvidia.com/gpu as an extended resource and advertises the count of available GPUs on each node. Without it, pods cannot request GPUs by resource name. Ensuring the operator and its device plugin are healthy is the correct step to let PyTorch jobs request GPUs abstractly rather than by device index.

Why this answer

Kubernetes learns about specialized hardware through device plugins that advertise extended resources. The NVIDIA device plugin, managed by the GPU Operator, publishes nvidia.com/gpu counts so the scheduler can allocate GPUs to pods. With it healthy, PyTorch jobs can declare a GPU request and receive an assigned device without the user specifying a physical index.

Exam trap

The trap here is equating GPU driver or container runtime installation with resource advertisement, when only the device plugin registers nvidia.com/gpu with the Kubernetes scheduler.

67
MCQeasy

Which command is used to verify that the NVIDIA GPU Operator has successfully installed the necessary components on a Kubernetes node?

A.nvidia-smi check-components
B.kubectl get pods -n gpu-operator
C.docker inspect nvidia-gpu-operator
D.kube-config verify --gpu
AnswerB

The GPU Operator runs in a specific namespace. Listing the pods in this namespace allows an administrator to see the status of the daemonsets, such as the driver installer and device plugin. All pods being in a 'Running' or 'Completed' state confirms a successful deployment of the operator components.

Why this answer

The 'kubectl get pods -n gpu-operator' command is the primary method to check the status of the operator and its managed components, such as the device plugin and driver daemonsets. This step is essential because it confirms that the operator's control loop has successfully completed the deployment, ensuring that the GPU software stack is active and ready to handle incoming AI workload scheduling requests.

Exam trap

Candidates often confuse cluster-wide resource inspection commands like 'kubectl get nodes' or generic pod queries with operator-specific status checks, forgetting to target the dedicated namespace where the GPU Operator components reside.

68
MCQhard

An administrator notices that a specific containerized training job reports high 'GPU Duty Cycle' but low 'Memory Bandwidth Utilization'. What does this pattern indicate about the workload?

A.The model is suffering from excessive CPU-to-GPU data transfer overhead.
B.The workload is compute-bound, performing heavy arithmetic operations.
C.The system is experiencing PCIe lane bandwidth saturation.
D.The batch size is too small to saturate the GPU compute units.
AnswerB

When the GPU is constantly busy (high duty cycle) but not demanding high amounts of data from VRAM (low bandwidth), it indicates that the kernels are performing a large number of computations relative to the amount of data read, characterizing a compute-bound operation.

Why this answer

High duty cycle combined with low memory bandwidth suggests the model is compute-bound rather than memory-bound. This usually occurs with models that have very high arithmetic intensity, such as small models with many layers. Recognizing this helps in selecting the appropriate hardware, such as focusing on TFLOPS capability rather than HBM bandwidth, to optimize the training speed of the specific neural network architecture.

Exam trap

Candidates often confuse low memory bandwidth with memory-bound workloads, wrongly assuming the GPU lacks sufficient VRAM capacity, whereas it actually indicates that the processor is saturated with intense arithmetic computations.

69
MCQmedium

A Kubernetes cluster administrator is installing the NVIDIA GPU Operator and wants to ensure that GPU workloads are scheduled only on nodes with healthy GPUs. The administrator plans to use the operator's built-in health checks. Which component is responsible for monitoring GPU health and marking nodes as unschedulable when a GPU fails?

A.NVIDIA DCGM Exporter
B.NVIDIA GPU Feature Discovery
C.NVIDIA Device Plugin
D.NVIDIA GPU Operator's health check component (part of the operator's node validation)
AnswerD

The GPU Operator includes a health check mechanism that periodically validates GPU health. When a GPU is found unhealthy, the operator applies a taint to the node, preventing new GPU workloads from being scheduled there. This integrates with Kubernetes scheduling to avoid placing workloads on failing hardware.

Why this answer

The GPU Operator's health check component continuously monitors GPU status and taints nodes with unhealthy GPUs. This prevents Kubernetes from scheduling new GPU workloads on those nodes, improving reliability. The device plugin handles resource advertisement, while DCGM Exporter provides metrics; neither automatically taints nodes based on health.

Exam trap

The trap here is attributing health-based node tainting to the device plugin or DCGM Exporter, when it is actually the operator's dedicated health check component that performs this action.

70
Multi-Selecthard

A data engineering team is deploying a distributed data processing workload. Which THREE metrics are most important to monitor in the Workload Manager to ensure optimal GPU throughput and identify potential bottlenecks?

Select 3 answers
A.GPU Duty Cycle (Active utilization percentage).
B.GPU Memory Bandwidth Utilization.
C.Container CPU usage percentage.
D.PCIe Throughput (Data transfer rates).
E.Pod restart count in the namespace.
AnswersA, B, D

The duty cycle measures the percentage of time the GPU is performing actual computation. If this number is consistently low, it indicates that the GPU is waiting for data or that the workload is CPU-bound, which is a vital indicator for troubleshooting underperforming AI training or processing jobs.

Why this answer

Monitoring these metrics allows administrators to detect inefficiencies in data loading or GPU utilization. GPU Duty Cycle tracks how often the GPU is actively processing data. Memory Bandwidth Utilization highlights if the bottleneck is in data transfer.

Finally, PCIe throughput indicates whether the bus between the CPU and GPU is saturated. Together, these metrics provide a complete picture of why a workload might be underperforming in a high-performance environment.

Exam trap

Candidates often select general CPU or memory metrics instead of specialized GPU metrics like GPU Duty Cycle, Memory Bandwidth, and PCIe Throughput when diagnosing GPU workload bottlenecks in Workload Manager.

71
MCQeasy

A data engineering team runs nightly batch inference on a Kubernetes cluster with NVIDIA GPUs. Jobs sometimes fail because two pods are scheduled onto the same physical GPU and one exhausts framebuffer memory. The team wants each pod to receive an isolated slice of a GPU with dedicated memory. Which NVIDIA feature should they enable?

A.Multi-Instance GPU (MIG)
B.NVIDIA Collective Communications Library (NCCL)
C.GPUDirect Storage
D.NVIDIA Container Toolkit
AnswerA

MIG partitions a single physical GPU into multiple independent instances, each with its own dedicated memory and compute slices. Because the memory is hardware-partitioned, one instance cannot consume another's framebuffer, which directly prevents the failure the team is seeing. The device plugin can then advertise each MIG instance as a schedulable resource so pods receive isolated slices.

Why this answer

The team needs hardware-level isolation with dedicated memory per workload. Multi-Instance GPU partitions a physical GPU into independent instances, each with its own memory and compute, and the device plugin can expose those instances as discrete resources. GPUDirect Storage, NCCL, and the Container Toolkit each address I/O paths, communication, or containerization rather than partitioning a GPU among tenants.

Exam trap

The trap here is confusing containerization or communication tooling, which makes GPUs usable, with partitioning features that actually isolate GPU memory.

72
MCQmedium

What is the primary benefit of using NVIDIA GPU Operator in a Kubernetes cluster for workload management?

A.It enables automatic model fine-tuning for specific frameworks.
B.It provides a declarative way to manage GPU drivers and software components.
C.It optimizes the neural network architecture automatically.
D.It eliminates the need for containerization in AI workflows.
AnswerB

The GPU Operator uses Kubernetes custom resources to declaratively manage NVIDIA software. This ensures that driver versions and plugin configurations are consistent across nodes, eliminating manual setup errors. This automation is crucial for scalability, as it allows administrators to manage thousands of GPUs with the same consistency as a single node.

Why this answer

The NVIDIA GPU Operator automates the lifecycle management of NVIDIA software components, including drivers, the container toolkit, and device plugins. By standardizing these installations across all cluster nodes, it ensures a consistent and stable environment for AI workloads, reducing configuration drift and manual operational overhead that often plagues large-scale distributed infrastructure deployments.

Exam trap

Test-takers often confuse the NVIDIA GPU Operator with general container runtimes or standard Kubernetes schedulers, overlooking its specific role in declaratively managing drivers and software components.

73
MCQmedium

Which of the following describes the purpose of the NVIDIA GPU Operator's 'Driver Container'?

A.To store the persistent state of trained neural networks.
B.To provide a platform-agnostic way to deploy NVIDIA drivers.
C.To manage the licensing of the AI Enterprise suite.
D.To act as a gateway for remote GPU access.
AnswerB

The Driver Container encapsulates the driver installation logic, making it consistent across different node operating systems. It handles the nuances of kernel headers and source code compilation, allowing the GPU Operator to manage drivers as standard Kubernetes workloads rather than requiring manual installation on each node.

Why this answer

The Driver Container is a critical component that builds or pulls the correct driver for the specific host OS and kernel. It automates the complex process of driver installation, ensuring compatibility across heterogeneous node environments. By containerizing the driver, the Operator simplifies maintenance and upgrades, reducing the risk of configuration drift and ensuring that nodes always have a functional driver compatible with the latest AI software releases.

Exam trap

Candidates often think the driver container installs the driver directly onto the host OS. In reality, it packages the driver to be portable and compatible across different kernel versions without manual compilation.

74
MCQeasy

An administrator is responsible for maintaining a fleet of NVIDIA-certified servers running AI workloads. They need to quickly identify which servers have GPUs that are overheating and may throttle performance. Which NVIDIA tool should the administrator use to monitor GPU temperature across the fleet in real time?

A.NVIDIA Nsight Systems
B.NVIDIA Data Center GPU Manager (DCGM)
C.NVIDIA TensorRT
D.NVIDIA CUDA Toolkit
AnswerB

DCGM is designed for data center GPU monitoring and management at scale. It provides real-time temperature, power, and health metrics for each GPU and can be integrated with monitoring systems. For fleet-wide temperature monitoring and throttle detection, DCGM is the appropriate tool in this scenario.

Why this answer

DCGM is the NVIDIA tool purpose-built for data center GPU monitoring. It exposes temperature, power, utilization, and health metrics through APIs and integrations with popular monitoring platforms. For an administrator needing real-time temperature visibility across many servers, DCGM is the correct choice.

Exam trap

The trap here is confusing performance profiling or development tools like Nsight Systems or CUDA Toolkit with fleet monitoring tools, which have different purposes and do not provide centralized temperature telemetry.

75
MCQeasy

When installing NVIDIA drivers via a package manager, what is the importance of the 'dkms' package?

A.It provides the graphical user interface for managing GPU settings.
B.It ensures the driver module is rebuilt automatically during kernel upgrades.
C.It improves the performance of the CUDA compiler during the build process.
D.It manages the firmware updates for the GPU hardware.
AnswerB

DKMS is designed to maintain kernel modules across kernel updates. By automatically recompiling the NVIDIA kernel module against the new kernel headers, it ensures that the GPU driver remains compatible and loaded correctly after the OS performs a system update or upgrade.

Why this answer

DKMS (Dynamic Kernel Module Support) is critical because it automates the recompilation of the NVIDIA kernel module whenever a new Linux kernel is installed. In AI production environments, kernel updates are common for security and stability. DKMS ensures that the GPU remains functional after an update without requiring manual driver intervention, preventing unplanned downtime for AI workloads and simplifying system administration tasks.

Exam trap

Candidates often confuse DKMS with standard package managers or assume manual recompilation is sufficient, forgetting that system updates happen automatically and break modules without DKMS.

Page 1 of 5

Page 2

All pages