Courseiva

NVIDIA Certified Professional: AI Operations (NCP-AIO) — Questions 301–309

309 questions total · 5pages · All types, answers revealed

Page 4

Page 5 of 5

301
MCQmedium

If a GPU job consistently fails with 'Out of Memory' despite the model size being significantly smaller than the total VRAM, what is the most likely cause?

A.The GPU clock speed is too low.
B.Memory fragmentation in the CUDA context.
C.An incompatible version of NCCL library.
D.The system bus width is insufficient.
AnswerB

Memory fragmentation occurs when the available memory is split into small, non-contiguous blocks. When the model requests a large contiguous block of memory, the allocator fails to find one, resulting in an OOM error even if the sum of all free memory is actually greater than the request.

Why this answer

Memory fragmentation occurs when frequent allocations and deallocations leave holes in the memory address space. Even if total free memory seems sufficient, large contiguous blocks cannot be allocated. Monitoring fragmentation patterns is essential for AI engineers to optimize memory management, such as using memory pools or persistent buffers, to ensure stable and predictable training performance in long-running jobs.

Exam trap

Candidates frequently assume an OOM error always means total available memory is exhausted, overlooking how memory fragmentation prevents large contiguous allocations.

302
MCQmedium

An administrator is deploying the NVIDIA GPU Operator into an existing Kubernetes cluster where the NVIDIA driver is already installed and maintained by the node image. The team wants the Operator to manage only the container runtime, device plugin, and monitoring components. Which configuration should be applied to the GPU Operator's ClusterPolicy?

A.Set devicePlugin.enabled to false so the Operator does not advertise GPU resources.
B.Set driver.enabled to false so the Operator does not deploy the driver container.
C.Set dcgmExporter.enabled to false so the Operator does not conflict with the existing driver.
D.Set toolkit.enabled to false so the Operator does not modify the container runtime configuration.
AnswerB

Setting driver.enabled to false tells the GPU Operator to skip driver management and rely on the preinstalled host driver. The Operator then proceeds to deploy the container toolkit, device plugin, DCGM exporter, and other managed components. This is the documented way to integrate the Operator with nodes whose drivers are maintained outside the Operator's lifecycle.

Why this answer

To use the GPU Operator with drivers already managed by the node image, the ClusterPolicy must disable driver management so the Operator skips the driver container. The container toolkit, device plugin, and DCGM exporter remain enabled so the Operator still configures runtime integration, advertises GPU resources, and exposes telemetry as the team intends.

Exam trap

The trap here is confusing driver management with runtime or device plugin management and disabling a component the scenario actually wants the Operator to control.

303
MCQmedium

An administrator is deploying NVIDIA AI Enterprise on a bare-metal cluster. Which component must be installed first to ensure proper communication between the Kubernetes scheduler and the underlying GPU hardware?

A.NVIDIA Triton Inference Server
B.NVIDIA GPU Operator
C.NVIDIA NeMo Framework
D.NVIDIA Base Command Manager
AnswerB

The GPU Operator automates the installation of the NVIDIA driver, the Kubernetes device plugin, the DCGM monitoring agent, and other necessary components. Establishing this layer first ensures the cluster is GPU-aware and that the scheduler can effectively identify and assign physical resources to incoming application workloads.

Why this answer

The NVIDIA GPU Operator is essential for automating the management of all NVIDIA software components in Kubernetes. By installing it first, the administrator ensures that the device plugin, monitoring tools, and drivers are correctly configured. This foundation is critical for scheduling GPU-accelerated pods, as the Kubernetes scheduler requires the device plugin to advertise available GPU resources to the cluster's API server, enabling seamless workload orchestration across the infrastructure.

Exam trap

Candidates often think monitoring tools or device plugins must be installed individually first, overlooking that the GPU Operator automates and manages all underlying subcomponents.

304
MCQeasy

An operations engineer notices that an inference container on an A100 is intermittently returning stale predictions after a model update. The container mounts the model directory from a host path, and the update process replaces files in place. Which change most reliably prevents the stale predictions?

A.Increase the container's shared memory size so the model can be cached entirely in RAM.
B.Reduce the inference batch size so each prediction reads the model from disk again.
C.Enable read-only mounts for the model directory to block writes from the container.
D.Use a versioned, immutable model artifact and restart or roll the inference container so it loads the new version atomically.
AnswerD

Replacing files in place while a process has them mapped or cached can leave the running server serving old weights or a mix of old and new tensors. Publishing each model as an immutable, versioned artifact and restarting the server ensures the process loads a consistent snapshot. This removes the race between file replacement and model loading, eliminating stale or partially updated predictions.

Why this answer

When model files are overwritten in place, a running inference process keeps using the weights it already loaded, so clients see stale results until the process restarts. Treating each model as an immutable, versioned artifact and performing a container restart or rolling deployment guarantees an atomic switch to a consistent set of weights. This removes the timing race between the update and the load, which is the actual source of the stale predictions.

Exam trap

The trap here is assuming that read-only mounts or larger caches make model updates safe, when the real issue is that in-place file replacement never notifies a process that already loaded the weights.

305
MCQhard

Refer to the exhibit. An administrator is attempting to deploy a job to a namespace with a ResourceQuota defined. What is the cause of this error?

A.The GPU driver version is incompatible with the quota controller.
B.The pod manifest is missing the required GPU resource limits.
C.The cluster is out of available GPU capacity.
D.The NVIDIA Device Plugin is not running in the namespace.
AnswerB

The error message explicitly states that limits for 'nvidia.com/gpu' must be specified. This is a common requirement in environments where quotas are implemented to ensure fair scheduling. Without these limits, the admission controller rejects the pod because it cannot account for the GPU usage against the namespace quota.

Why this answer

The error indicates that the namespace has a ResourceQuota enforcing that all pods must specify GPU limits, but the submitted pod manifest lacks a 'resources.limits.nvidia.com/gpu' entry. In multi-tenant environments, ResourceQuotas are essential for preventing a single user from consuming the entire GPU capacity. The manifest must include a valid GPU limit to satisfy the namespace policy, ensuring the cluster remains balanced across different organizational teams.

Exam trap

Candidates often assume ResourceQuota errors stem from cluster-wide node exhaustion, missing the fact that namespace-level policies explicitly require explicit GPU resource limits in the pod manifest.

306
MCQeasy

Which component in the NVIDIA AI Enterprise stack is responsible for providing the necessary user-space libraries and binaries to run GPU-accelerated applications inside containers?

A.NVIDIA GPU Operator
B.NVIDIA Container Toolkit
C.NVIDIA License System
D.NVIDIA Unified Fabric Manager
AnswerB

The NVIDIA Container Toolkit is specifically designed to provide the libraries and binaries required for containerized applications to perform GPU acceleration. It includes the runtime hook that makes the GPU visible to the container environment, ensuring the application can utilize the hardware for compute or graphics tasks.

Why this answer

The NVIDIA Container Toolkit is essential because it allows the container runtime to interact with the host's NVIDIA drivers. It provides the necessary libraries and the container runtime wrapper to expose GPUs inside the container environment. Without this toolkit, applications inside a container cannot access the GPU hardware, even if the host has the correct drivers installed, making it a critical deployment component.

Exam trap

Candidates often confuse the NVIDIA driver itself with the Container Toolkit. They assume drivers alone allow containers to access GPU hardware, ignoring the necessary user-space library mapping provided by the toolkit.

307
Multi-Selecthard

A platform team operates a Kubernetes cluster where several teams submit GPU training jobs. The administrator needs to enforce per-namespace limits on the number of GPUs that can be consumed and prevent a single namespace from monopolizing all GPU capacity. Which TWO Kubernetes resources should be configured to achieve this? (Choose two.)

Select 2 answers
A.A PriorityClass that assigns a low priority value to all training pods in the namespace.
B.A PodSecurityPolicy that denies privileged containers in the namespace.
C.A LimitRange that sets a default and maximum nvidia.com/gpu value for containers in the namespace.
D.A NetworkPolicy that restricts traffic between pods in different namespaces.
E.A ResourceQuota that specifies nvidia.com/gpu in its hard limits for each namespace.
AnswersC, E

LimitRange applies defaults and bounds to individual containers. Setting a maximum nvidia.com/gpu stops a single container from requesting an excessive number of GPUs, and the default ensures pods that omit a GPU request still receive a defined value, complementing the namespace-wide cap enforced by ResourceQuota.

Why this answer

ResourceQuota enforces an aggregate ceiling on nvidia.com/gpu per namespace, while LimitRange constrains and defaults the per-container GPU request. Together they bound total namespace consumption and prevent any single container from grabbing an outsized share, which is exactly the governance the platform team needs.

Exam trap

The trap here is believing that PriorityClass or NetworkPolicy can cap GPU consumption, when only quota and limit-range objects act on resource quantities at admission time.

308
MCQhard

Refer to the exhibit. The training job is showing intermittent "thermal throttling" warnings. Which configuration change is the most appropriate adjustment?

A.Set "compute_mode" to "exclusive_process".
B.Disable ECC mode to lower the power consumption.
C.Lower the "power_limit_watts" value.
D.Set "persistence" to "disabled".
AnswerC

Reducing the power limit directly lowers the heat generated by the GPU. While this may slightly decrease the maximum performance, it prevents the GPU from reaching the thermal trip point, thereby eliminating throttling and ensuring a consistent, albeit slightly lower, performance profile during long-duration training jobs.

Why this answer

Thermal throttling occurs when the GPU reaches its maximum operating temperature and lowers its clock speed to prevent physical damage. While power limits can be adjusted to reduce heat, simply lowering the limit may negatively impact training performance. Adjusting the power limit to a value that balances thermal overhead with compute requirements is a necessary trade-off to ensure stable, consistent performance during long training runs without hardware damage.

Exam trap

Test-takers frequently assume that lowering the power limit will eliminate thermal throttling without side effects, forgetting that overly restricted wattage can severely degrade training compute performance and extend runtime.

309
MCQmedium

A distributed training job using PyTorch DDP across eight GPUs on one DGX A100 node shows GPU utilization oscillating between 20 and 40 percent, while `nvidia-smi dmon` shows low SM activity but sustained high memory-controller utilization. The data loader reads from a local NVMe RAID array and applies heavy CPU augmentation. Which action best improves GPU utilization?

A.Enable NCCL gradient compression to reduce inter-GPU communication volume.
B.Reduce the per-GPU batch size so each step completes faster and utilization appears smoother.
C.Increase the number of DataLoader worker processes and enable pinned memory with non-blocking host-to-device transfers.
D.Switch the job from DDP to a single-process data-parallel loop that iterates over all GPUs in Python.
AnswerC

Low SM activity combined with heavy memory-controller traffic points to input pipeline stalls: the GPUs are idle waiting for batches. More worker processes parallelize CPU augmentation, while pinning host buffers and using non-blocking copies lets the copy engine overlap transfer with compute. This directly shortens the gap between kernel bursts and raises sustained SM occupancy without changing model math.

Why this answer

Oscillating SM utilization with high memory-controller traffic is the classic signature of an input pipeline that cannot keep pace with compute. Adding DataLoader workers parallelizes augmentation across CPU cores, and pinned memory plus non-blocking transfers let DMA copies overlap with kernels. This keeps the GPUs fed continuously, converting idle gaps into productive compute and raising sustained utilization without altering the model or optimizer.

Exam trap

The trap here is blaming inter-GPU communication because the job is distributed, when the profile shows idle SMs rather than busy NVLink, which points to starvation from the host input pipeline instead.

Page 4

Page 5 of 5

All pages