Courseiva

NVIDIA Certified Professional: AI Operations (NCP-AIO) — Questions 76–150

309 questions total · 5pages · All types, answers revealed

Page 1

Page 2 of 5

Page 3
76
Multi-Selectmedium

An administrator is optimizing a cluster for AI model training using NVIDIA Base Command. Which TWO tasks are critical for ensuring consistent performance across the training nodes?

Select 2 answers
A.Standardizing NVIDIA driver and NCCL library versions across all nodes
B.Increasing the size of the swap partition on each node
C.Enabling GPUDirect RDMA to bypass system memory for GPU-to-GPU data transfer
D.Disabling the NVIDIA persistence mode to save power
E.Reducing the number of CUDA streams per training job
AnswersA, C

Distributed training performance is highly sensitive to the communication stack. Inconsistent NCCL versions can lead to suboptimal collective operation performance, while mismatched drivers can cause compatibility issues with the underlying hardware interconnects, ultimately leading to performance degradation and difficult-to-debug failures during large-scale model training cluster operations.

Why this answer

Consistency in high-performance AI clusters depends on hardware synchronization and resource availability. Ensuring that all nodes run identical driver and firmware versions prevents subtle performance regressions during multi-node training. Furthermore, configuring GPUDirect RDMA is essential for reducing latency in inter-GPU communication, which is a major bottleneck in distributed training jobs.

These steps are fundamental to maintaining high utilization and predictable training times across large-scale NVIDIA accelerated infrastructure.

Exam trap

Candidates often focus only on software frameworks while ignoring crucial low-level networking optimizations like GPUDirect RDMA and driver version synchronization across multi-node setups.

77
Multi-Selecthard

An AI operations engineer is tuning a real-time inference service on NVIDIA A100 GPUs. Profiling with Nsight Systems shows that the GPU is idle for long periods while waiting for input data, and that host-to-device memory copies are frequent and small. The service uses a fixed batch size of 1 and a custom data loader. Which two changes are most likely to improve GPU utilization and reduce inference latency? (Choose two.)

Select 2 answers
A.Increase the number of CUDA streams used for memory copies and inference kernels.
B.Increase the inference batch size and implement dynamic batching to group requests.
C.Use `cudaMemcpy` instead of `cudaMemcpyAsync` for all host-to-device transfers.
D.Pin the inference process to a single CPU core to reduce context switching.
E.Enable NVIDIA TensorRT with FP16 precision and optimize the model for the target GPU.
AnswersB, E

Larger batches amortize kernel launch and memory copy overhead across more work, keeping the GPU busier. Dynamic batching groups multiple incoming requests into a single inference pass, improving throughput and reducing per-request latency when the service is under load. This directly addresses the idle GPU periods caused by small, frequent transfers and fixed batch size of one.

Why this answer

The GPU idles because it waits on small, frequent data transfers and processes one sample at a time. Increasing batch size with dynamic batching and optimizing the model with TensorRT in FP16 reduce per-inference overhead and better utilize Tensor Cores. Together they increase work per transfer and speed up computation, directly improving utilization and latency.

Exam trap

The trap here is focusing on CPU-side or stream-level tweaks while overlooking that the core inefficiency is the tiny batch size and unoptimized model execution.

78
MCQmedium

An AI engineer observes that a model training job on an NVIDIA DGX system is underutilizing the GPU. The monitoring logs show high CPU wait times and low GPU duty cycles. Which action should the engineer take first to resolve the bottleneck?

A.Increase the GPU clock frequency via NVML.
B.Enable GPUDirect Storage on the local drive.
C.Optimize the data preprocessing pipeline using NVIDIA DALI.
D.Increase the batch size to maximize memory usage.
AnswerC

NVIDIA DALI offloads data augmentation and preprocessing tasks from the CPU to the GPU. By moving these compute-intensive preprocessing steps onto the hardware acceleration engines, the CPU is relieved of the burden, allowing it to prepare batches faster and keeping the GPU fully saturated with training data.

Why this answer

High CPU wait times alongside low GPU utilization suggest an I/O or data preprocessing bottleneck. The CPU cannot feed data to the GPU fast enough, forcing the GPU to idle while waiting for the next batch. Optimizing the data pipeline, such as increasing prefetch buffers or using NVIDIA DALI to move image processing to the GPU, directly addresses the starvation of the compute resources, ensuring efficient utilization.

Exam trap

Candidates often try to optimize the GPU architecture or hyperparameters, missing the fact that the CPU data preprocessing pipeline is the actual bottleneck.

79
Multi-Selectmedium

An administrator is optimizing multi-GPU utilization. Which TWO of the following configurations allow multiple containers to share a single physical GPU on a supported NVIDIA architecture?

Select 2 answers
A.Enabling NVIDIA Time-Slicing in the GPU Operator configuration.
B.Configuring Multi-Instance GPU (MIG) profiles.
C.Increasing the CUDA_VISIBLE_DEVICES environment variable.
D.Setting a higher GPU limit in the Kubernetes manifest.
E.Deploying the NVIDIA Network Operator.
AnswersA, B

Time-Slicing allows multiple pods to share a GPU by switching context at rapid intervals. This is a software-based approach that enables oversubscription, allowing smaller workloads to execute concurrently on the same hardware, which is highly beneficial for development environments or low-throughput inference tasks that don't need full GPU power.

Why this answer

Multi-Instance GPU (MIG) and Time-Slicing are the primary methods for sharing physical GPU resources. MIG provides hardware-level isolation, while Time-Slicing provides software-based multiplexing. Understanding these options is essential for AI operations, as it allows administrators to maximize hardware ROI by supporting smaller workloads that do not require an entire A100 or H100 GPU, effectively increasing the density of the training or inference environment.

Exam trap

Candidates frequently confuse software-level multiplexing options like Time-Slicing with hypervisor features or mistake them for hardware-partitioning mechanisms like MIG, failing to recognize that both are valid sharing methods.

80
MCQhard

An administrator supports a multi-tenant cluster where several teams share GPUs. Leadership requires that each team's batch jobs receive a fair share of GPU time and that one team cannot monopolize devices by submitting thousands of low-priority pods. Jobs are submitted through a Kubernetes-native batch scheduler that supports queueing. Which approach best enforces fair-share GPU allocation across teams?

A.Use node taints and tolerations to dedicate specific GPU nodes to each team permanently
B.Assign each team a distinct Kubernetes namespace and rely on the default kube-scheduler for GPU placement
C.Define per-team queues with weighted fair-share policies and GPU-aware scheduling in the batch scheduler
D.Apply identical PriorityClass values to all team pods so the scheduler treats them equally
AnswerC

A GPU-aware batch scheduler with per-team queues and weighted fair-share policies allocates GPU capacity proportionally across tenants and prevents any single team from monopolizing devices through sheer volume. It coordinates gang scheduling and quotas at the queue level, directly delivering the fairness and anti-monopolization requirement for shared GPU clusters.

Why this answer

Weighted fair-share queues in a GPU-aware batch scheduler allocate device time proportionally among teams and enforce queue-level limits, preventing a single tenant from monopolizing GPUs by volume. The default scheduler, equal priorities, or static node partitioning cannot provide dynamic proportional sharing or gang scheduling for distributed jobs.

Exam trap

The trap here is equating namespace isolation or equal priorities with fairness, when fair share actually requires a scheduler that tracks per-queue usage and applies weighted allocation across teams.

81
MCQeasy

When installing the NVIDIA GPU Operator, which namespace is typically used to ensure proper isolation and role-based access control?

A.default
B.gpu-operator
C.kube-system
D.public-apps
AnswerB

Creating a dedicated namespace for the GPU Operator is standard industry practice. It provides logical isolation and allows administrators to apply specific RBAC policies to the operator's components, ensuring that the critical system-level software is segregated from general tenant workloads and other cluster services for better security posture.

Why this answer

Using a dedicated namespace like 'gpu-operator' is a best practice in Kubernetes. It isolates the operator's resources, permissions, and lifecycle from other cluster services. This separation allows for granular security policies, making it easier to manage access and ensuring that only authorized personnel can modify the operator's configuration, which is vital for maintaining the security and integrity of the GPU-enabled infrastructure in a production environment.

Exam trap

Candidates often install the Operator in the default namespace. This creates security risks and makes it difficult to apply specific RBAC policies or manage the lifecycle of the operator independently of applications.

82
MCQmedium

When deploying large-scale distributed training jobs, why is it recommended to use the NVIDIA Network Operator in conjunction with the GPU Operator?

A.To increase the number of supported GPU containers.
B.To optimize RDMA throughput for distributed training.
C.To provide automatic GPU driver updates.
D.To enable multi-tenancy on GPU nodes.
AnswerB

Distributed training relies on efficient GPU-to-GPU communication across nodes. The Network Operator manages the configuration of RDMA-enabled NICs, allowing GPUs to communicate directly with minimal CPU intervention. This dramatically increases throughput and reduces latency, which is essential for the performance of large-scale distributed machine learning training jobs.

Why this answer

The Network Operator automates the configuration of high-speed interconnects like InfiniBand or RoCE, which are crucial for distributed training across multiple nodes. By aligning network resources with GPU placement, the operator minimizes latency and maximizes throughput. This ensures that the GPU compute capacity is not bottlenecked by slow network communication, which is essential for scaling complex AI model training across large clusters effectively.

Exam trap

Candidates often assume that standard Kubernetes networking is sufficient for distributed AI training, overlooking the critical requirement for low-latency, high-throughput RDMA interconnects.

83
MCQhard

Which of the following is the primary indicator of PCIe bus saturation when profiling a training job on an NVIDIA DGX system?

A.High SM occupancy but low memory bandwidth usage.
B.Low GPU duty cycle and high PCIe throughput utilization.
C.High GPU temperature and low clock speeds.
D.Frequent ECC errors in the GPU memory logs.
AnswerB

If the GPU duty cycle is low, it means the GPU is waiting for data. If the PCIe throughput utilization is simultaneously high, it indicates that the bus is fully saturated with data transfers, confirming that the GPU is stalled because the data cannot arrive fast enough.

Why this answer

PCIe bus saturation occurs when the GPU's requirement for data exceeds the transfer capacity of the bus. This leads to GPU starvation, where the compute cores idle while waiting for data. Recognizing the specific correlation between low GPU utilization and high PCIe throughput is essential for identifying I/O-bound bottlenecks and optimizing data pipelines for large-scale training jobs.

Exam trap

Candidates often confuse PCIe saturation with general GPU memory bandwidth bottlenecks. They mistakenly assume high GPU utilization is required to identify a bus issue, ignoring that starvation results in low utilization.

84
MCQmedium

Which action must be performed after updating the NVIDIA driver on a Linux host to ensure that all active GPU containers recognize the new driver version?

A.Rebuild all container images with the latest CUDA toolkit.
B.Restart the containerized applications.
C.Upgrade the Docker daemon version.
D.Re-run the nvidia-container-toolkit installation script.
AnswerB

Active containers hold references to the old driver files that were mapped during their initialization. To ensure that the containers are using the updated driver libraries from the host, the processes must be terminated and relaunched so the container runtime can remount the latest versions.

Why this answer

Containers utilize the host-level NVIDIA drivers injected via the container toolkit. When the host driver is updated, active containers will still be referencing the old driver libraries in their environment until they are restarted. Restarting the containers is the mandatory step to force them to bind against the newly installed libraries, ensuring compatibility and preventing runtime errors associated with outdated driver interfaces.

Exam trap

Candidates assume updating the host driver is instantly inherited by running containers, forgetting that active containers retain bindings to old libraries until restarted.

85
MCQmedium

Refer to the exhibit. What is the most effective way to resolve this specific throttling condition?

A.Update the GPU firmware to the latest release.
B.Reinstall the NVIDIA display driver.
C.Increase the power limit for the GPU using nvidia-smi.
D.Replace the power distribution unit (PDU) in the rack.
AnswerC

The 'Sw Power Cap' throttle reason explicitly states that the software power limit is currently restricting the GPU performance. Increasing the power limit via the nvidia-smi tool allows the driver to permit higher clock speeds, effectively removing this constraint and enabling the GPU to perform at full capacity.

Why this answer

The exhibit indicates that the GPU is hitting a software-defined power cap. This is a common situation when a node is configured with a restricted power limit to manage rack-level power consumption. Adjusting the power limit using nvidia-smi ensures the GPU can access the necessary power to reach higher clock frequencies, which is essential for maximizing performance in compute-intensive deep learning tasks.

Exam trap

Candidates often assume the issue is a software bug or driver failure. They overlook the possibility that the GPU is simply hitting a configured power limit enforced by the administrator.

86
MCQmedium

A data scientist reports that a PyTorch training job on an NVIDIA V100 GPU is running slower than expected. The job uses a data loader with num_workers=4. Monitoring shows GPU utilization is around 50%, and CPU usage is high. Which action should an AI operations engineer recommend to improve GPU utilization?

A.Pin memory in the DataLoader to speed up host-to-device transfers.
B.Enable mixed precision training to speed up computations.
C.Increase num_workers in the DataLoader to overlap data preprocessing with GPU computation.
D.Increase the batch size to better utilize the GPU.
AnswerC

When CPU usage is high and GPU utilization is low, the data loading pipeline cannot keep up with the GPU. Increasing num_workers allows more parallel data preprocessing processes, reducing the time the GPU waits for data. This directly addresses the bottleneck and can significantly improve GPU utilization, provided CPU resources are sufficient.

Why this answer

The combination of high CPU usage and low GPU utilization points to a data loading bottleneck where the GPU is starved for data. Increasing the number of DataLoader workers enables more parallel data preprocessing, allowing the GPU to be fed continuously. Other options do not address the CPU-bound pipeline and may not improve the situation.

Exam trap

The trap here is assuming that GPU-side optimizations like mixed precision or larger batches will help, when the real issue is insufficient data preprocessing throughput.

87
Multi-Selecthard

An AI operations team is troubleshooting a distributed training job on a cluster of NVIDIA DGX A100 systems connected via InfiniBand. The job runs but achieves only 40% of expected scaling efficiency. The team suspects communication bottlenecks. Which two actions should they take to confirm and address the issue? (Choose two.)

Select 2 answers
A.Verify that all nodes have identical NCCL versions and environment variables, and that InfiniBand fabric manager is running.
B.Run NCCL_DEBUG=INFO and inspect the logs for warnings about falling back to slower transports or retries.
C.Switch the job to use Ethernet instead of InfiniBand to simplify troubleshooting.
D.Use nvidia-smi to monitor GPU utilization and memory bandwidth on each node during training.
E.Increase the global batch size proportionally to the number of GPUs to improve compute-to-communication ratio.
AnswersA, B

Inconsistent NCCL versions or missing environment variables (e.g., NCCL_IB_HCA, NCCL_SOCKET_IFNAME) can cause suboptimal transport selection or failures. The InfiniBand fabric manager ensures proper link configuration. Ensuring uniformity across nodes is crucial for optimal collective communication performance and is a common fix for scaling inefficiencies.

Why this answer

Scaling inefficiency in distributed training often stems from communication overhead. NCCL debug logs directly show transport selection and errors, while ensuring consistent NCCL versions and InfiniBand fabric health addresses common misconfigurations. Together, these actions confirm and resolve bottlenecks.

The other options either do not diagnose the network or would degrade performance.

Exam trap

The trap here is focusing on GPU-level metrics or hyperparameter changes instead of directly inspecting NCCL communication behavior and InfiniBand fabric health.

88
MCQeasy

A platform engineer manages an NVIDIA-accelerated Kubernetes cluster running the NVIDIA GPU Operator on nodes with A100 GPUs. Several data-science teams submit training jobs, and the engineer must ensure each team's pods receive a full physical GPU exclusively, with no two pods sharing the same device. Which scheduling configuration should the engineer apply to the pod specification to guarantee exclusive whole-GPU allocation?

A.Enable MIG by setting the nvidia.com/mig.config label on the node to all-balanced and request nvidia.com/mig-1g.5gb in the pod.
B.Configure the pod to mount the hostPath /dev/nvidia0 and set privileged: true so the container can access the first GPU directly.
C.Set a nodeSelector for nvidia.com/gpu.product=A100 and rely on the scheduler to place one pod per node automatically.
D.Set resources.limits to nvidia.com/gpu: 1 and resources.requests to nvidia.com/gpu: 1, letting the NVIDIA device plugin allocate an exclusive GPU.
AnswerD

Requesting and limiting nvidia.com/gpu to 1 makes the NVIDIA device plugin treat the GPU as an indivisible integer device, so the kubelet advertises whole GPUs and the scheduler binds exactly one physical A100 exclusively to the pod. Because the extended resource is integer-only, no sharing occurs, which satisfies the exclusivity requirement without extra configuration.

Why this answer

Declaring both a request and a limit for nvidia.com/gpu equal to 1 is the canonical way to obtain an exclusive whole GPU in Kubernetes. The NVIDIA device plugin advertises GPUs as integer extended resources, so the scheduler reserves one entire device for the pod. Because extended resources cannot be fractional, no co-scheduling or sharing occurs, which precisely meets the exclusive allocation requirement.

Exam trap

The trap here is assuming that a nodeSelector or hostPath device mount can enforce exclusivity, when only the nvidia.com/gpu extended resource request actually reserves a whole physical GPU through the device plugin.

89
MCQmedium

An AI researcher is running a large-scale training job on an NVIDIA DGX system using Kubernetes. They observe that GPU utilization is consistently low despite high CPU load. Which workload management configuration is most likely to resolve this bottleneck by optimizing data pipeline throughput?

A.Increase the GPU memory limit in the pod specification.
B.Enable NVIDIA Multi-Instance GPU (MIG) for this specific job.
C.Optimize the data loader by increasing worker threads and implementing buffered prefetching.
D.Lower the batch size to reduce the memory footprint on the GPU.
AnswerC

Optimizing the data loader allows the CPU to prepare future training batches while the current batch is processing on the GPU. By increasing worker threads and utilizing prefetching, the system effectively hides I/O latency, ensuring the GPU cores are saturated with data and significantly improving overall model training efficiency.

Why this answer

Low GPU utilization often stems from data starvation where the GPU waits for the CPU to preprocess or fetch data from storage. Implementing data prefetching and increasing the number of workers in the data loader ensures the GPU remains fed with batches. This is critical in AI operations to maximize compute ROI and reduce total training time for expensive cluster resources.

Exam trap

Candidates frequently try to resolve low GPU utilization by scaling up the GPU instances or adjusting cluster scheduling policies, failing to recognize that the bottleneck is actually CPU-bound data preprocessing.

90
MCQmedium

Refer to the exhibit. The training job fails with an OOM error. Which optimization strategy will most effectively resolve this while maintaining model convergence?

A.Increase the learning rate to compensate for smaller batches.
B.Implement Gradient Accumulation to simulate a larger batch size.
C.Disable mixed-precision training to reduce memory overhead.
D.Clear the GPU cache using torch.cuda.empty_cache() after every iteration.
AnswerB

Gradient accumulation allows you to simulate a large batch size by breaking it into smaller chunks that fit in GPU memory. You perform multiple forward/backward passes and accumulate the gradients, updating weights only after reaching the target batch size, effectively bypassing the physical memory limitation.

Why this answer

Memory management is central to deep learning stability. When a model exceeds physical VRAM, gradient accumulation allows for larger effective batch sizes without increasing memory footprint. By accumulating gradients over multiple small steps and performing a single weight update, the model effectively sees a larger batch size, maintaining the convergence characteristics of the original design while fitting within the strict physical memory constraints of the GPU hardware.

Exam trap

Candidates frequently try to resolve OOM errors by simply reducing the batch size without realizing it can negatively impact model convergence and training accuracy.

91
MCQhard

An administrator is configuring a Kubernetes cluster where some nodes have A100 GPUs and others have H100 GPUs. A training job requires specific GPU memory capacity and CUDA compute capability. Which mechanism should the administrator use to ensure the job is only scheduled onto nodes with the correct GPU model?

A.Apply a taint to nodes without the required GPU and rely on the default scheduler
B.Increase the nvidia.com/gpu resource request to match the GPU memory size
C.Set a nodeSelector matching the GPU model label applied by the GPU Operator's node feature discovery
D.Use a ResourceQuota that limits GPU requests per namespace
AnswerC

Node Feature Discovery, deployed with the GPU Operator, labels nodes with GPU model information such as nvidia.com/gpu.product. A nodeSelector referencing that label constrains the scheduler to nodes with the required GPU model, ensuring the job lands only on A100 or H100 nodes as appropriate without hardcoding node names.

Why this answer

Node Feature Discovery, part of the GPU Operator stack, labels nodes with GPU product details such as nvidia.com/gpu.product. A nodeSelector or affinity rule referencing that label restricts scheduling to nodes with the required GPU model, which is the precise way to satisfy memory and compute capability requirements without manual node naming.

Exam trap

The trap here is treating nvidia.com/gpu as a model or memory selector, when it is only a device count and cannot distinguish between A100 and H100 hardware.

92
MCQeasy

A platform team is preparing a Kubernetes cluster for AI workloads and wants the GPU device plugin, driver containers, and monitoring components deployed and kept in sync automatically on every GPU node. Which component should be installed to achieve this?

A.NVIDIA Network Operator
B.NVIDIA DCGM standalone on each node
C.NVIDIA Container Toolkit installed manually on each node
D.NVIDIA GPU Operator
AnswerD

The NVIDIA GPU Operator uses the Operator pattern to deploy and manage the GPU driver, container runtime hooks, device plugin, DCGM exporter, and related components as DaemonSets. It continuously reconciles node state, so new GPU nodes are automatically provisioned with the full software stack, which directly matches the requirement for automatic, synchronized deployment.

Why this answer

The NVIDIA GPU Operator is purpose-built to manage the entire GPU software stack on Kubernetes nodes through automated reconciliation. It deploys the driver, container toolkit configuration, device plugin, and DCGM-based monitoring, so GPU resources are advertised and maintained without manual per-node work. The Network Operator addresses networking, while DCGM and the Container Toolkit are individual pieces the Operator already orchestrates.

Exam trap

The trap here is confusing the NVIDIA Container Toolkit, which only wires the runtime for GPU access, with the GPU Operator that manages the whole stack automatically.

93
MCQmedium

An AI operations engineer manages a shared Kubernetes cluster where several teams submit GPU jobs. The engineer must prevent any single namespace from consuming all GPU capacity and must also ensure that jobs from one team cannot starve others during peak periods. Which combination of Kubernetes and NVIDIA GPU Operator features should the engineer implement?

A.Define a ResourceQuota on nvidia.com/gpu per namespace and use the GPU Operator's device plugin to advertise capacity so the quota can enforce limits.
B.Enable the MIG Manager and assign a fixed mig profile to every namespace through a LimitRange.
C.Apply a NetworkPolicy that restricts pod-to-pod traffic between namespaces and rely on it to limit GPU consumption.
D.Set a node taint for nvidia.com/gpu and add tolerations only to the highest-priority team's pods.
AnswerA

ResourceQuota can cap the total nvidia.com/gpu requests within a namespace when the device plugin advertises GPUs as extended resources. This prevents one namespace from monopolizing cluster GPU capacity. Combined with the operator's device plugin, quotas become enforceable at admission time, ensuring fair sharing across teams without manual intervention.

Why this answer

ResourceQuota is the Kubernetes mechanism that caps aggregate resource consumption per namespace, and it works for GPU extended resources once the NVIDIA device plugin advertises them. By setting a quota on nvidia.com/gpu, the administrator prevents any single namespace from consuming all GPUs, ensuring other teams retain capacity. This directly addresses both the monopolization and starvation concerns.

Exam trap

The trap here is treating NetworkPolicy, taints, or LimitRange as tools for GPU capacity fairness, when only ResourceQuota enforces an aggregate per-namespace ceiling on nvidia.com/gpu.

94
Multi-Selectmedium

Which TWO of the following actions are recommended for optimizing NVIDIA GPU utilization during a high-concurrency inference deployment?

Select 2 answers
A.Disable ECC memory on all GPUs to increase raw clock speed.
B.Implement NVIDIA Multi-Process Service (MPS) to allow concurrent kernel execution.
C.Convert models to TensorRT format to leverage layer fusion and precision tuning.
D.Switch the GPU power mode to maximum performance via NVML.
E.Increase the batch size to the maximum allowed by GPU memory.
AnswersB, C

MPS allows multiple processes to share GPU resources more effectively by enabling concurrent kernel execution. This is particularly useful for small inference models where a single process cannot saturate the GPU, leading to higher overall utilization and better performance during high-concurrency periods.

Why this answer

Optimizing inference requires maximizing throughput while maintaining low latency. Using TensorRT for model compilation and Multi-Process Service (MPS) for resource sharing are standard industry practices. These methods ensure that compute resources are efficiently allocated across multiple streams, preventing underutilization of the GPU's tensor cores.

Mastering these tools is essential for AI Operations engineers managing production-grade inference pipelines that must scale effectively under varying loads.

Exam trap

Candidates often suggest increasing the batch size or adding more GPUs. These do not address the efficiency of the individual GPU's execution streams or the model's runtime format optimization.

95
MCQeasy

A healthcare company is deploying NVIDIA AI Enterprise on a Kubernetes cluster to run medical imaging AI models. The cluster administrator needs to verify that the NVIDIA GPU Operator is installed and functioning correctly. Which command should the administrator use to check the status of the GPU Operator pods?

A.kubectl describe daemonset nvidia-device-plugin -n kube-system
B.kubectl logs -n gpu-operator nvidia-driver-daemonset
C.kubectl get nodes -o wide
D.kubectl get pods -n gpu-operator
AnswerD

The GPU Operator is typically deployed in the gpu-operator namespace. Running kubectl get pods -n gpu-operator lists all pods managed by the Operator, allowing the administrator to verify that components like the driver, container toolkit, and device plugin are running.

Why this answer

The NVIDIA GPU Operator deploys its components into the gpu-operator namespace. To verify that the Operator is installed and functioning, the administrator should list the pods in that namespace. This provides a quick overview of all related pods, such as the driver, container toolkit, and device plugin, and their current status.

Exam trap

The trap here is assuming that GPU Operator components reside in kube-system or that checking nodes alone is sufficient; the Operator uses its own namespace, and pod status is the definitive check.

96
Multi-Selecthard

An AI operations engineer is investigating a training job on an NVIDIA DGX system that intermittently fails with 'uncorrectable ECC error' on a GPU. The job is using NCCL for multi-GPU communication. The engineer needs to identify the appropriate immediate actions to diagnose and mitigate the issue. (Choose two.)

Select 2 answers
A.Increase the NCCL timeout to allow the job to recover from the ECC error.
B.Drain the affected GPU from the scheduler to prevent new jobs from being assigned.
C.Restart the training job immediately to see if the error recurs.
D.Use nvidia-smi --gpu-reset to reset the GPU and clear the error state.
E.Run nvidia-smi -q to check the ECC error counts and identify the affected GPU.
AnswersB, E

Draining the GPU prevents additional jobs from being scheduled on a potentially failing device, reducing the risk of further job failures and data corruption. This is a standard operational mitigation for hardware errors. It allows the administrator to perform maintenance or replacement without impacting other workloads, and it is an appropriate immediate action after identifying the affected GPU.

Why this answer

The immediate actions should include checking ECC error counts with nvidia-smi -q to identify the affected GPU, and draining that GPU from the scheduler to prevent further job failures. These steps diagnose the issue and mitigate risk while preserving the ability to perform maintenance. Restarting, resetting, or adjusting NCCL timeouts do not address the underlying hardware error and may delay proper resolution.

Exam trap

The trap here is treating an uncorrectable ECC error as a software or communication issue and attempting to fix it with job restarts or NCCL tuning.

97
Multi-Selecthard

A platform engineer must validate a new NVIDIA GPU Operator deployment on a Kubernetes cluster before handing it to data scientists. Which two checks confirm that the Operator has correctly exposed GPU resources to the cluster scheduler? (Choose two.)

Select 2 answers
A.Confirm that the GPU Operator's driver DaemonSet pods are scheduled on all nodes, including CPU-only nodes.
B.Run a CUDA-enabled pod that requests nvidia.com/gpu and verify it reaches Running state and reports the expected device.
C.Confirm that nodes advertise the nvidia.com/gpu resource and that its allocatable count matches the physical GPU count.
D.Check that the NVIDIA driver container image tag matches the CUDA toolkit version installed in the workload image.
E.Verify that the container runtime on each node has been switched from containerd to Docker with the nvidia runtime as default.
AnswersB, C

A scheduled pod that requests the GPU resource exercises the full path: scheduler admission, device plugin allocation, container runtime injection, and driver access inside the container. If the pod runs and nvidia-smi or a CUDA sample reports the device, the end-to-end GPU enablement is verified, not just the resource advertisement.

Why this answer

GPU exposure is confirmed by two complementary signals: the node advertises the nvidia.com/gpu extended resource with an allocatable count matching the physical GPUs, and a test pod requesting that resource schedules successfully and can access the device. Together they validate device plugin registration and end-to-end runtime injection. Driver image tags, node runtime swaps, and DaemonSet placement on CPU-only nodes are irrelevant to this validation.

Exam trap

The trap here is treating driver or runtime version matching as proof of GPU exposure instead of checking the advertised extended resource and an actual scheduled GPU pod.

98
MCQeasy

Which component of the NVIDIA GPU Operator is responsible for monitoring GPU health and reporting telemetry data to the Kubernetes control plane?

A.NVIDIA Device Plugin.
B.NVIDIA DCGM Exporter.
C.NVIDIA Container Toolkit.
D.NVIDIA Node Feature Discovery.
AnswerB

The DCGM Exporter collects detailed GPU metrics and exposes them in a Prometheus-compatible format. This allows administrators to monitor GPU health, performance, and power consumption. It is the core monitoring component that bridges the gap between hardware telemetry and the observability stack in a containerized AI infrastructure.

Why this answer

The DCGM Exporter is the primary tool within the NVIDIA GPU Operator ecosystem for collecting metrics. It interacts with the Data Center GPU Manager (DCGM) to gather telemetry such as utilization, temperature, and memory health. This is vital for AI Operations because it provides the observability necessary to trigger auto-scaling, identify failing hardware, and ensure that training jobs are performing optimally within the cluster environment.

Exam trap

Candidates often confuse management components like the GPU Operator or device plugin with the specialized telemetry collection agent, mixing up scheduling logic with metrics gathering.

99
Multi-Selecthard

Which TWO strategies should an administrator implement to ensure fair resource scheduling in a multi-tenant NVIDIA cluster using Kubernetes and the NVIDIA device plugin?

Select 2 answers
A.Configure Kubernetes ResourceQuotas to limit the total number of GPUs per namespace.
B.Implement static partitioning for all GPUs in the cluster.
C.Define Kubernetes PriorityClasses to influence job preemption policies.
D.Use the NVIDIA Triton Inference Server to manage all batch processing.
E.Disable the NVIDIA device plugin to allow direct node access.
AnswersA, C

ResourceQuotas provide a hard ceiling on the aggregate compute resources a namespace can consume. This prevents a single tenant from launching an excessive number of pods that might exhaust cluster GPU capacity, ensuring that other tenants maintain access to sufficient hardware for their own operational and development needs.

Why this answer

Fair scheduling in multi-tenant environments requires both hard limits to prevent resource monopolization and priority-based mechanisms to ensure critical jobs progress. By combining ResourceQuotas for capacity management and PriorityClasses for job scheduling, admins can prevent a single user from starving the cluster while maintaining high performance for latency-sensitive inference or urgent training tasks during peak usage windows.

Exam trap

Candidates often select only software scheduling flags or generic pod limits, missing the necessary combination of namespace-level capacity caps and explicit job priority classes required for multi-tenant fairness.

100
MCQhard

A team runs multi-node training with NCCL over InfiniBand on a cluster of DGX systems. Jobs scale well to four nodes but throughput drops sharply at eight nodes, and `nccl-tests` all-reduce bandwidth falls well below line rate at that size. The fabric uses a fat-tree topology with adaptive routing enabled. Which investigation is most likely to reveal the cause?

A.Check whether NCCL is selecting the correct InfiniBand HCAs and whether the ring or tree algorithm choice matches the fabric's oversubscription ratio.
B.Move the collective operations from NCCL to a TCP socket backend over the management network.
C.Increase the batch size per GPU so that communication is amortized over more compute.
D.Disable adaptive routing on the InfiniBand fabric to force deterministic paths.
AnswerA

Beyond a certain node count, NCCL's default topology detection can pick a ring that traverses oversubscribed uplinks, or it can select the wrong HCA when multiple adapters exist per node. Verifying `NCCL_IB_HCA` and forcing the appropriate algorithm or topology file aligns communication with the physical fabric. This directly explains why bandwidth collapses only at larger scale while small jobs look healthy.

Why this answer

NCCL chooses rings and trees based on detected topology, and on multi-node jobs this can cross oversubscribed uplinks or use a suboptimal adapter when several are present. When scaling stops at a specific node count, the cause is usually that the communication pattern no longer matches the fabric's capacity. Confirming HCA selection and steering the algorithm or topology to respect the oversubscription ratio restores bandwidth and explains why smaller jobs appeared fine.

Exam trap

The trap here is assuming the interconnect is broken or that adaptive routing is at fault, when the collapse at a specific node count usually reflects NCCL's topology and algorithm choices crossing oversubscribed links.

101
Multi-Selecthard

When deploying the NVIDIA GPU Operator in a restricted-access environment (air-gapped), which THREE requirements must be addressed to ensure a successful installation?

Select 3 answers
A.Mirror all required images to a private registry.
B.Upgrade the host BIOS to the latest version.
C.Configure the operator to use the private registry.
D.Provide local access to kernel headers.
E.Install the NVIDIA Triton Inference Server first.
AnswersA, C, D

Without internet access, the operator cannot fetch images from public sources like NGC. Mirroring is mandatory to ensure the container runtime can pull the images locally. The operator deployment manifest must also be updated to reference the internal registry URL instead of the default public registry paths.

Why this answer

In air-gapped environments, the inability to reach public registries is the primary failure point. Administrators must mirror all required container images to a local private registry, configure the operator to point to these local locations, and ensure that all necessary kernel headers are available locally for the driver build process. These steps ensure that the GPU Operator can complete its setup without external internet dependencies.

Exam trap

Candidates often forget the requirement for local kernel headers. Even with mirrored images, the driver build process will fail if it cannot access the necessary kernel headers locally.

102
MCQhard

An administrator runs a Kubernetes cluster with the NVIDIA GPU Operator. A data science team wants to run several small inference containers that each use only a fraction of a GPU's compute and memory, but the cluster currently assigns whole GPUs per pod. Which approach allows multiple containers to share a single physical GPU with memory isolation?

A.Enable Multi-Instance GPU mode on supported GPUs and configure the device plugin to advertise MIG instances as schedulable resources.
B.Reduce the container's nvidia.com/gpu request to a fractional value such as 0.5.
C.Set the device plugin's time-slicing configuration so that each GPU is advertised multiple times.
D.Configure the MPS control daemon to allow concurrent kernel execution across containers.
AnswerA

MIG partitions a supported GPU into hardware-isolated instances with dedicated compute and memory slices. Advertising those instances through the device plugin lets each inference container receive its own MIG device, providing true memory and fault isolation while allowing multiple containers to share one physical GPU, which matches the requirement.

Why this answer

MIG is the only mechanism here that divides a physical GPU into hardware-isolated instances with separate memory and compute slices. By enabling MIG and having the device plugin advertise each instance, the scheduler can place multiple inference containers on one GPU while keeping their memory and faults isolated.

Exam trap

The trap here is conflating time-slicing or MPS, which share a GPU without memory isolation, with MIG, which provides true hardware-level memory partitioning.

103
MCQmedium

A cluster administrator notices that GPU utilization is low despite high queue volume. After analyzing the logs, they identify that many pods are failing because they cannot access the necessary CUDA libraries. What is the most likely cause, and which component should be verified?

A.Verify the NVIDIA Device Plugin status, as it manages library injection.
B.Check the NVIDIA Container Runtime configuration for proper runtime class mapping.
C.Review Kubernetes Resource Quotas, as they restrict the total amount of GPU memory.
D.Increase the GPU memory limit to ensure enough memory for loading heavy CUDA libraries.
AnswerB

The NVIDIA Container Runtime must be configured to inject the appropriate drivers and libraries into the container. If this mapping is missing, applications will fail to load CUDA libraries, leading to runtime errors. Checking the container runtime ensures that the environment is set up for correct GPU access.

Why this answer

The most likely cause is a misconfiguration of the NVIDIA Container Runtime, which is responsible for injecting the necessary NVIDIA libraries into the container namespace. If the container runtime is not properly configured, the container cannot interact with the GPU, even if the scheduler has successfully placed it. Verifying the container runtime configuration is essential to ensure that GPU-enabled containers can access host-level drivers and CUDA stacks.

Exam trap

Candidates often blame the GPU driver version first. While important, the runtime mapping is the most common configuration error that prevents the container from 'seeing' the host's GPU capabilities.

104
MCQmedium

An administrator manages an NVIDIA AI Enterprise cluster running multiple Kubernetes nodes, each with several A100 GPUs. After upgrading the NVIDIA GPU Operator to a newer version, the administrator notices that pods requesting GPUs remain in a Pending state, and the node's allocatable GPU count is reported as zero. Which command should the administrator run first to diagnose the issue?

A.kubectl describe node <node-name>
B.kubectl get pods --all-namespaces -o wide
C.nvidia-smi -q
D.kubectl logs -n gpu-operator <gpu-operator-pod>
AnswerA

This command shows detailed node status, including conditions, capacity, and allocatable resources. If the GPU Operator's device plugin is not functioning, the node will report zero allocatable nvidia.com/gpu resources, and events may indicate plugin registration failures. It directly reveals whether the node recognizes the GPUs as schedulable resources, making it the essential first diagnostic step.

Why this answer

The node's allocatable GPU count is reported as zero, indicating that the kubelet is not receiving GPU resource advertisements from the NVIDIA device plugin. Describing the node reveals capacity, allocatable resources, and relevant events, such as device plugin registration failures. This directly identifies whether the GPU Operator's device plugin is functioning, making it the correct first step.

Exam trap

The trap here is assuming that checking the GPU Operator logs or running nvidia-smi will directly explain why Kubernetes reports zero allocatable GPUs, when the node's resource status is the authoritative source.

105
MCQmedium

Which NVIDIA technology enables the partitioning of a single physical GPU into multiple independent instances for use by different virtual machines or containers?

A.NVIDIA GPUDirect Storage.
B.NVIDIA Multi-Instance GPU (MIG).
C.NVIDIA NVLink.
D.NVIDIA vGPU Profiles.
AnswerB

MIG enables hardware-level partitioning of the GPU, allowing each instance to have its own compute cores, memory, and cache. This provides robust isolation and performance guarantees for multiple applications or users, ensuring that one workload does not adversely impact the performance of another co-located on the same device.

Why this answer

NVIDIA Multi-Instance GPU (MIG) technology is the correct answer. It allows a single GPU to be securely partitioned at the hardware level, providing guaranteed QoS and isolation for different workloads. This is essential for maximizing GPU utilization in enterprise AI, as it enables the co-location of small inference tasks alongside larger training workloads on a single piece of high-end hardware.

Exam trap

Candidates often confuse MIG with vGPU or Time-Slicing, failing to specify that MIG is the unique hardware-level partitioning technology for NVIDIA GPUs.

106
MCQmedium

An AI engineer needs to monitor GPU utilization across a large cluster of nodes in real-time. Which NVIDIA tool is the most appropriate for this high-level observability task?

A.nvidia-smi
B.NVIDIA DCGM (Data Center GPU Manager)
C.CUDA Profiler (nsys)
D.nvcc
AnswerB

DCGM is the enterprise-grade tool for managing and monitoring NVIDIA GPU clusters. It provides the necessary APIs to export metrics to monitoring systems like Prometheus and Grafana, allowing for real-time observability across large fleets of GPUs, which is critical for maintaining high availability in production AI environments.

Why this answer

Observability at scale requires tools that can aggregate metrics across multiple nodes. NVIDIA DCGM (Data Center GPU Manager) is specifically architected for this purpose, providing a comprehensive API and service to monitor GPU health, utilization, and power consumption across entire clusters. This is essential for operations teams managing large-scale AI infrastructure to proactively detect performance issues and optimize resource utilization across the fleet.

Exam trap

Candidates often confuse local single-GPU utilities like nvidia-smi with cluster-wide observability tools required for managing multi-node infrastructures.

107
MCQeasy

Which NVIDIA software component is responsible for providing the necessary CUDA libraries to containerized applications?

A.NVIDIA Driver
B.NVIDIA Container Toolkit
C.NVIDIA vGPU Manager
D.NVIDIA Triton
AnswerB

The NVIDIA Container Toolkit provides the runtime hooks and libraries that allow containers to access the GPU and CUDA acceleration. It enables the container engine to identify, map, and utilize the host's NVIDIA hardware, ensuring that deep learning frameworks can call CUDA primitives directly from within the container.

Why this answer

The NVIDIA Container Toolkit is the critical bridge. It provides the necessary libraries and container runtime hooks that allow processes inside a container to access the host's GPU and CUDA environment. This setup is fundamental for AI Enterprise, as it enables portability of AI applications while maintaining high-performance access to physical GPU hardware, regardless of the underlying host OS distribution or container runtime used.

Exam trap

Candidates often confuse the Container Toolkit with the NVIDIA driver itself, failing to recognize that the Toolkit provides the bridge for containers to consume host-side drivers.

108
MCQmedium

An AI engineer is deploying a large language model on an NVIDIA DGX system. The deployment fails with an error indicating an insufficient NVIDIA driver version for the required CUDA toolkit. Which action should the engineer take to resolve the dependency mismatch?

A.Reinstall the CUDA toolkit using a generic installer without checking the NVIDIA driver version compatibility matrix.
B.Downgrade the OS kernel to a legacy version to force compatibility with an older CUDA toolkit.
C.Upgrade the NVIDIA driver to a version verified as compatible with the required CUDA toolkit.
D.Modify the LD_LIBRARY_PATH environment variable to prioritize older CUDA libraries found on the system.
AnswerC

Matching the driver version to the CUDA toolkit requirements is the standard procedure for fixing driver-level mismatches. This ensures that the GPU hardware can successfully communicate with the user-space libraries, allowing the AI software stack to utilize the full range of CUDA capabilities.

Why this answer

Drivers and CUDA versions maintain a strict compatibility matrix. Installing the latest driver version supported by the specific DGX OS is essential to ensure the CUDA runtime can interface correctly with the GPU hardware. This task is critical in production environments because mismatched drivers lead to kernel panics or silent performance degradation, directly impacting the availability of AI workloads running on the cluster infrastructure.

Exam trap

Candidates often suggest reinstalling the CUDA toolkit or the entire container runtime. The root issue is the driver-to-CUDA compatibility matrix, which requires updating the driver to match the toolkit.

109
MCQeasy

An AI operations engineer notices that a real-time inference service on an NVIDIA T4 GPU has highly variable latency, with occasional spikes to over 100 ms. The service uses TensorRT and runs in a Docker container. Which action should the engineer take to reduce latency variability?

A.Enable dynamic batching in the inference server configuration.
B.Increase the number of CPU cores allocated to the container.
C.Set the GPU to persistence mode using nvidia-smi -pm 1.
D.Lock the GPU clocks to a fixed frequency using nvidia-smi -lgc.
AnswerD

Locking GPU clocks prevents the GPU from dynamically adjusting frequencies based on power and thermal headroom, which can cause latency variability. A fixed clock ensures consistent performance, reducing spikes. This is especially effective for latency-sensitive inference where predictable execution time is critical. It trades off some power efficiency for stability.

Why this answer

Latency variability in GPU inference often results from dynamic clock adjustments as the GPU balances power and thermal constraints. Locking clocks to a stable frequency eliminates this variability, providing consistent execution times. Other options either add latency (dynamic batching), address different issues (persistence mode), or do not target GPU clock behavior.

Exam trap

The trap here is assuming that persistence mode or CPU allocation affects runtime latency, when the primary cause of variability is often GPU clock throttling.

110
MCQhard

A site reliability engineer is troubleshooting a DGX A100 node that intermittently drops out of the cluster during large NCCL all-reduce jobs. `nvidia-smi` shows all eight A100 GPUs healthy, but DCGM reports XID errors 74 and 79 on one GPU during the failures. The engineer needs to determine the most likely cause and the correct administrative action. Which combination best describes the cause and the appropriate first step?

A.The errors are benign and expected during large all-reduce jobs; suppress them by setting the DCGM health check to ignore XID 74 and 79.
B.The errors indicate a thermal shutdown; immediately lower the GPU clock with `nvidia-smi -lgc` and rerun the job.
C.The errors are caused by an outdated NCCL version; upgrade NCCL to the latest release and rerun the job without further hardware checks.
D.The errors point to a GPU falling off the bus or an NVLink error; run `nvidia-smi -q` and DCGM diagnostics on the suspect GPU, then consider reseating or replacing the GPU board.
AnswerD

XID 74 and 79 are associated with GPU falling off the bus and NVLink errors respectively. The correct administrative response is to collect detailed GPU and NVLink state with `nvidia-smi -q`, run DCGM diagnostics to isolate the failing GPU or link, and then perform hardware remediation such as reseating the GPU board or replacing it if diagnostics confirm a persistent fault.

Why this answer

XID 74 and 79 are driver-reported errors indicating a GPU has fallen off the bus or an NVLink error has occurred. These are hardware and link-level faults, so the right first step is to gather detailed GPU and NVLink diagnostics with `nvidia-smi -q` and DCGM, then remediate the hardware. Ignoring or masking these errors risks job failures and data corruption, while treating them as thermal or software issues misses the actual fault domain.

Exam trap

The trap here is treating XID 74 and 79 as generic performance or thermal warnings, when they specifically signal GPU bus loss and NVLink faults that require hardware-level diagnosis.

111
MCQhard

An MLOps engineer manages a Kubernetes cluster where the NVIDIA GPU Operator runs the MIG manager. Several inference pods must each receive an isolated, fixed slice of a single A100, and the team wants the slices to survive node reboots without manual reconfiguration. Which combination of settings should the engineer apply?

A.Set the MIG manager's config to 'all-disabled' and rely on the device plugin to carve profiles dynamically per pod request.
B.Disable the MIG manager entirely and create MIG instances manually with nvidia-smi on each node, then label the nodes so pods schedule there.
C.Enable time-slicing with four replicas and set the MIG manager strategy to 'single', because replicas provide the same isolation as MIG instances.
D.Configure a MIG config labeled for the MIG manager with named profiles and set the device plugin strategy to 'mixed', so whole GPUs and MIG instances are both advertised.
AnswerD

A labeled MIG config tells the MIG manager which geometry to apply and persist, so instances are recreated after reboots without manual work. The mixed strategy lets the device plugin advertise both full GPUs and MIG instances, allowing the inference pods to request a specific MIG resource such as nvidia.com/mig-1g.5gb. This yields fixed, isolated slices with hardware-level memory and fault separation.

Why this answer

Persistent MIG slices require the MIG manager to read a labeled configuration that defines the desired geometry, which it reapplies after reboots and driver reloads. Setting the device plugin strategy to mixed lets both whole GPUs and the carved MIG instances be advertised under distinct resource names. Pods can then request a specific profile such as nvidia.com/mig-1g.5gb, obtaining fixed, hardware-isolated slices that meet the isolation and persistence requirements.

Exam trap

The trap here is believing time-slicing delivers the same hardware isolation as MIG, when it only shares one engine among workloads with no memory or fault boundaries.

112
MCQhard

Which THREE of the following are benefits of using Multi-Instance GPU (MIG) technology in a Kubernetes environment?

A.Hardware-level isolation between GPU workloads.
B.Increased total GPU count available to the OS kernel.
C.Deterministic Quality of Service (QoS) for different workloads.
D.Automatic translation of CUDA code to work on different GPU architectures.
E.Optimized utilization of expensive GPU hardware resources.
AnswerA, C, E

MIG provides true hardware-level partitioning, which prevents processes in one instance from accessing the memory or compute resources of another instance. This is far more robust than software-based time-slicing, ensuring that different tenants or workloads remain entirely isolated from one another.

Why this answer

MIG allows for hardware-level isolation, ensuring that one workload's memory usage or compute spikes do not impact others on the same physical GPU. This provides deterministic Quality of Service (QoS), which is critical for multi-tenant AI environments. By partitioning resources, administrators can optimize GPU allocation, allowing smaller, less intensive tasks to run on a fraction of the GPU, thereby significantly increasing overall cluster efficiency.

Exam trap

Candidates often mistake MIG for a software-based virtualization tool rather than a hardware-level partitioning technology, leading them to select incorrect benefits related to dynamic software-defined resource sharing instead of isolation.

113
MCQhard

An AI operations team manages a shared Kubernetes cluster where a nightly batch training workload requests nvidia.com/gpu resources and occasionally consumes all GPU memory on a node, causing a co-located interactive notebook pod to fail with CUDA out-of-memory errors. The team wants the interactive notebook to be isolated from the batch workload's memory usage without adding new hardware. Which action best achieves this on supported data center GPUs?

A.Enable time-slicing with a replica count of four so the notebook and batch workloads alternate on the same GPU in round-robin fashion.
B.Set a memory limit on the batch container using the standard Kubernetes resources.limits.memory field to cap its GPU framebuffer usage.
C.Configure Multi-Instance GPU profiles so the notebook and batch workloads run on separate MIG instances with dedicated memory partitions on the same physical GPU.
D.Add a higher PriorityClass to the notebook pod so the kube-scheduler preempts the batch workload whenever memory pressure occurs.
AnswerC

MIG slices a supported GPU into hardware-isolated instances, each with its own dedicated memory and compute resources. Placing the notebook and batch workload on separate instances prevents the batch job from consuming the notebook's memory, resolving the out-of-memory failures without procuring additional GPUs.

Why this answer

The conflict is shared GPU memory, not scheduling order. MIG is the only listed mechanism that gives each workload a dedicated, hardware-isolated memory partition on the same physical device, letting the notebook and batch job coexist without the batch job starving the notebook's framebuffer, and it requires no extra hardware.

Exam trap

The trap here is assuming that time-slicing or container memory limits isolate GPU memory, when neither partitions the device framebuffer and only MIG provides hardware-level memory separation.

114
MCQhard

An AI operations team is deploying NVIDIA AI Enterprise on a bare-metal Kubernetes cluster with DGX A100 systems. They need to enable GPUDirect Storage to accelerate data loading from a local NVMe array. Which component must be installed and configured on the DGX nodes to support GPUDirect Storage?

A.NVIDIA Peer-to-Peer (P2P) over PCIe with IOMMU disabled
B.NVIDIA GPU Operator with the RDMA shared device plugin enabled
C.NVIDIA Container Toolkit with the 'nvidia-container-runtime' configured for privileged mode
D.NVIDIA Magnum IO GPUDirect Storage kernel module and user-space libraries
AnswerD

GPUDirect Storage requires the nvidia-fs kernel module and CUDA libraries that enable direct memory access between storage and GPU memory. These are part of Magnum IO GPUDirect Storage. Installing and configuring them on the DGX nodes allows applications to bypass the CPU and system memory, reducing latency and increasing throughput for data-intensive AI workloads.

Why this answer

GPUDirect Storage is enabled by installing the NVIDIA Magnum IO GPUDirect Storage components, which include the nvidia-fs kernel module and CUDA libraries. These allow direct DMA transfers between NVMe storage and GPU memory, bypassing the CPU. On DGX systems, these components are typically part of the DGX software stack and must be properly configured.

Exam trap

The trap here is confusing GPUDirect Storage with other NVIDIA technologies like RDMA or P2P, which address different data paths.

115
MCQmedium

Which mechanism does the NVIDIA GPU Operator use to ensure that the driver installed on a worker node matches the specific architecture of the installed GPU hardware?

A.NVIDIA System Management Interface (nvidia-smi).
B.Node Feature Discovery (NFD).
C.The container runtime's default configuration.
D.Hardcoding the driver version in the deployment manifest.
AnswerB

NFD labels nodes based on hardware attributes like GPU model and architecture. The GPU Operator uses these labels to match the node to the appropriate driver image. This automated mechanism is essential for scaling deployments across diverse hardware configurations without requiring manual node configuration for every single machine.

Why this answer

The GPU Operator uses Node Feature Discovery (NFD) to identify the specific GPU hardware present on each node. By labeling nodes with these hardware characteristics, the operator can ensure that the correct driver and kernel modules are selected and deployed. This mapping is vital in heterogeneous clusters, where different nodes might require different driver builds, preventing installation errors and ensuring optimal performance across the entire fleet.

Exam trap

Candidates often guess that the GPU Operator performs the hardware detection itself. It actually relies on the Node Feature Discovery (NFD) to identify and label hardware for the operator.

116
MCQeasy

A platform team runs an on-premises Kubernetes cluster for AI inference. Several teams submit pods that request the same GPU device, and the scheduler places more pods onto a node than there are available GPUs, causing OOM errors on the device. The administrator wants the Kubernetes scheduler itself to prevent overcommitting GPUs without any custom admission controller. Which action should the administrator take?

A.Install the NVIDIA device plugin so GPUs are advertised as schedulable extended resources and add a matching resource request to each pod spec.
B.Set a node taint on each GPU node and add a toleration to pods, so only pods that explicitly opt in are placed there.
C.Enable the Kubernetes Vertical Pod Autoscaler in recommendation mode on the GPU namespaces to detect device pressure and reschedule pods.
D.Create a ResourceQuota in each team namespace limiting the count of pods that may run, so fewer pods land on GPU nodes.
AnswerA

The NVIDIA k8s-device-plugin registers each GPU as an extended resource such as nvidia.com/gpu, which makes the scheduler count devices as finite node capacity. When a pod requests one GPU, the scheduler subtracts it from allocatable capacity, so no node can be overcommitted. This directly solves the reported overplacement without writing any admission logic, because scheduling math handles accounting natively.

Why this answer

Advertising GPUs as extended resources through the NVIDIA device plugin is what lets the default scheduler treat each device as countable, finite node capacity. Pods that declare a GPU request are then placed only where an unallocated device exists, and no custom admission controller is required because the scheduler performs the arithmetic itself. Quotas, taints, and autoscalers influence eligibility or CPU and memory sizing, not device-level accounting.

Exam trap

The trap here is assuming that Kubernetes natively understands GPUs as limited resources; without the device plugin exposing them as extended resources, the scheduler treats GPU nodes as ordinary nodes and will happily overcommit devices.

117
MCQmedium

An administrator is deploying NVIDIA AI Enterprise on a bare-metal cluster. Which component is mandatory to provide the necessary abstraction layer for containerized GPU resources?

A.NVIDIA Triton Inference Server
B.NVIDIA GPU Operator
C.NVIDIA Base Command
D.NVIDIA CUDA Toolkit
AnswerB

The GPU Operator utilizes the Kubernetes operator pattern to automate the installation and management of NVIDIA drivers, the NVIDIA Container Toolkit, and device plugins. This provides the critical abstraction layer required to expose physical GPUs as schedulable resources within a container orchestrator environment seamlessly.

Why this answer

The NVIDIA GPU Operator is essential for automating the management of all NVIDIA software components needed to provision Kubernetes with GPU support. It manages the lifecycle of drivers, container runtimes, and monitoring tools. By leveraging the Operator, administrators ensure consistent configuration across the cluster, preventing drift and ensuring that CUDA workloads have the correct dependencies to execute reliably on bare-metal infrastructure.

Exam trap

Candidates often name individual components like the driver or toolkit. The GPU Operator is the mandatory overarching framework that automates the deployment and management of these individual components.

118
MCQeasy

A data scientist reports that a Jupyter notebook running on a NVIDIA GPU server is extremely slow when executing a deep learning model, even though `nvidia-smi` shows the GPU is idle. The notebook uses TensorFlow. Which action should be taken first to diagnose the issue?

A.Increase the notebook's memory allocation by setting `TF_GPU_ALLOCATOR=cuda_malloc_async`.
B.Check that TensorFlow is configured to use the GPU by running `tf.config.list_physical_devices('GPU')`.
C.Restart the Jupyter kernel and clear the GPU memory with `nvidia-smi --gpu-reset`.
D.Run `nvidia-smi -q -d PERFORMANCE` to check if the GPU is throttled due to thermal or power issues.
AnswerB

If the GPU is idle while the notebook runs slowly, TensorFlow may be defaulting to CPU execution. Verifying that TensorFlow detects the GPU is the first diagnostic step. The function `tf.config.list_physical_devices('GPU')` returns available GPUs; if empty, TensorFlow isn't using the GPU. This directly addresses the symptom of an idle GPU during a supposedly GPU-accelerated workload. It is a simple, non-invasive check that can quickly identify misconfiguration.

Why this answer

An idle GPU during a slow TensorFlow workload strongly suggests that TensorFlow is not using the GPU. The quickest way to confirm is to check if TensorFlow detects the GPU via `tf.config.list_physical_devices('GPU')`. If the list is empty, the issue is likely missing CUDA/cuDNN libraries or a CPU-only TensorFlow installation.

This diagnostic step is non-disruptive and directly targets the root cause.

Exam trap

The trap here is assuming that a slow notebook with an idle GPU indicates a hardware problem, when it often means the framework isn't configured to use the GPU.

119
MCQeasy

A system administrator is installing the NVIDIA Container Toolkit on a standalone Ubuntu server to run GPU-accelerated containers. After installation, they want to verify that the toolkit is correctly configured. Which command should they run to test GPU access from a container?

A.docker run --rm --gpus all nvidia/cuda:11.0-base nvidia-smi
B.nvidia-smi
C.systemctl status nvidia-container-runtime
D.nvidia-container-cli --version
AnswerA

This command runs a CUDA container with all GPUs exposed and executes nvidia-smi inside the container. If the toolkit is configured correctly, the container will have access to the GPUs and nvidia-smi will display them. This directly tests the container runtime's ability to pass through GPU devices. It is the standard method to verify NVIDIA Container Toolkit functionality.

Why this answer

To verify that the NVIDIA Container Toolkit is correctly configured, the administrator should run a container with GPU access and execute nvidia-smi inside it. The command docker run --rm --gpus all nvidia/cuda:11.0-base nvidia-smi uses the --gpus all flag to expose all GPUs, and the nvidia-smi output confirms that the container can see and use the GPUs. This tests the entire stack from Docker to the toolkit.

Other commands only check installation or host-level GPU status.

Exam trap

The trap here is assuming that running nvidia-smi on the host or checking the toolkit version is sufficient to verify container GPU access, when a container-based test is required.

120
MCQhard

An administrator is troubleshooting a performance degradation in a multi-node NVIDIA NCCL-based training job. The job spans four DGX nodes connected via InfiniBand. The administrator suspects that NCCL is not using the optimal network path. Which action should the administrator take to verify and enforce the use of GPUDirect RDMA for inter-node communication?

A.Configure the InfiniBand fabric to use RoCEv2 instead of InfiniBand native protocol to improve compatibility with GPUDirect RDMA.
B.Set NCCL_NET_GDR_LEVEL=0 to force NCCL to use GPUDirect RDMA for all inter-node communication.
C.Set the environment variable NCCL_DEBUG=INFO and inspect the logs for messages indicating that GPUDirect RDMA is enabled.
D.Run nvidia-smi topo -m on each node to check the GPU-to-NIC affinity and ensure that GPUs and InfiniBand adapters are connected via PCIe switches.
AnswerC

Setting NCCL_DEBUG=INFO generates detailed logs that show which transports NCCL uses, including whether GPUDirect RDMA is active. The logs will indicate if the network plugin is using RDMA and if there are any fallbacks to slower paths. This is the standard method to verify NCCL's behavior and diagnose performance issues related to network communication.

Why this answer

To verify and enforce GPUDirect RDMA usage, the administrator should enable NCCL debug logging with NCCL_DEBUG=INFO and examine the logs for transport information. This will show if GDR is active and if any fallbacks occur. Checking topology with nvidia-smi topo -m is useful but does not confirm active GDR.

Setting NCCL_NET_GDR_LEVEL=0 would disable GDR, and switching to RoCEv2 is irrelevant.

Exam trap

The trap here is misunderstanding NCCL_NET_GDR_LEVEL: a value of 0 disables GPUDirect RDMA, while higher values enable it based on topology distance.

121
MCQmedium

During the deployment of an AI model training workload on a multi-node cluster, the administrator notices that inter-node communication is significantly slower than expected. Which deployment aspect should be investigated first?

A.Verify that the GPU memory usage is below 50% on all nodes.
B.Check the NCCL_DEBUG and network interface configuration for the training job.
C.Restart the Kubernetes API server to refresh the node connection states.
D.Update the NVIDIA driver to the latest gaming-optimized release.
AnswerB

NCCL communication relies heavily on the correct identification of high-speed network interfaces. If the job is defaulting to an Ethernet interface instead of InfiniBand, training will be severely throttled. Setting NCCL_DEBUG allows administrators to identify which interfaces are being selected and if the desired fabric is actually being used.

Why this answer

NCCL (NVIDIA Collective Communications Library) is the primary engine for inter-node communication in distributed training. Misconfiguration of the network interface or the underlying fabric provider (e.g., InfiniBand or RoCE) will lead to significant performance bottlenecks. Investigating the NCCL configuration and the network topology ensures that the GPUs are utilizing the highest bandwidth available, which is vital for preventing training jobs from stalling during gradient synchronization across multiple nodes.

Exam trap

Candidates often try to troubleshoot the model code or the application logic first, ignoring the communication layer (NCCL) and network fabric which are the primary culprits for inter-node training latency.

122
MCQmedium

A platform engineer is preparing a bare-metal Kubernetes cluster to run GPU-accelerated AI workloads using the NVIDIA GPU Operator. The cluster nodes have NVIDIA Ampere GPUs and run Ubuntu 22.04 with containerd as the container runtime. The engineer wants to avoid installing any NVIDIA drivers or CUDA components directly on the host. Which GPU Operator configuration should be used to achieve this?

A.Set driver.enabled=false and rely on pre-installed host drivers.
B.Deploy the GPU Operator with the default configuration, which includes the driver container.
C.Use the operator's 'driver' Helm chart with 'driver.enabled=true' but set 'driver.usePrecompiled=true'.
D.Install the NVIDIA Container Toolkit manually and disable the operator's driver management.
AnswerB

The default GPU Operator deployment includes a driver container that compiles and loads the NVIDIA kernel modules on the host without requiring a pre-installed driver. This satisfies the requirement to avoid manual host driver installation while still enabling GPU access for workloads, as the Operator manages the driver lifecycle automatically.

Why this answer

The NVIDIA GPU Operator is designed to manage the full stack of GPU software, including the driver, by running a driver container on each node. In a default deployment, the driver container builds and loads the kernel modules, so no manual host driver installation is needed. This aligns with the requirement to keep the host clean and let the Operator handle everything.

Exam trap

The trap here is assuming that the GPU Operator always requires pre-installed host drivers, when in fact its default mode deploys a driver container to manage them.

123
MCQmedium

An administrator is configuring a new cluster and wants to ensure that telemetry data from GPUs is collected in a centralized manner. Which tool is best suited for this requirement?

A.nvidia-smi log file rotation
B.NVIDIA DCGM Exporter
C.Manual polling of /proc/driver/nvidia
D.NVIDIA Nsight Systems
AnswerB

The DCGM Exporter is purpose-built for this task. It collects detailed telemetry, such as power, temperature, and usage, directly from the GPUs via DCGM and makes it available to monitoring platforms. This provides a clean, automated, and scalable architecture for centralized tracking of GPU performance metrics across the cluster.

Why this answer

NVIDIA DCGM Exporter is the industry-standard tool for collecting GPU telemetry data in Kubernetes environments. It aggregates metrics from the Data Center GPU Manager and exposes them in a format compatible with Prometheus, allowing for centralized monitoring, visualization, and alerting. This visibility is essential for understanding cluster performance, identifying anomalies, and optimizing resource utilization in large-scale AI infrastructure deployments, making it the preferred choice for enterprise monitoring solutions.

Exam trap

Test-takers sometimes confuse standard Kubernetes metrics servers with specialized GPU telemetry collectors like the NVIDIA DCGM Exporter.

124
MCQmedium

An operations engineer is troubleshooting a distributed training job that uses NVIDIA Magnum IO GPUDirect Storage to read training data directly from a local NVMe SSD into GPU memory. The job reports lower than expected I/O bandwidth. `nvidia-smi` shows normal GPU utilization, and the NVMe drive's throughput is well below its peak. Which factor is most likely limiting GPUDirect Storage performance in this scenario?

A.The training job is using too many CPU threads for data preprocessing.
B.The GPU's ECC memory is enabled, reducing available bandwidth.
C.The GPU is not connected to the NVMe drive through a supported PCIe topology.
D.The NVMe drive is formatted with a file system that does not support GPUDirect Storage.
AnswerC

GPUDirect Storage requires a direct data path between the storage device and the GPU, typically over PCIe with peer-to-peer support. If the NVMe drive and GPU are behind different PCIe switches or root complexes without proper peer-to-peer capabilities, data must bounce through host memory, reducing bandwidth. This is a common limitation in multi-socket servers where the GPU and NVMe are on different NUMA nodes.

Why this answer

GPUDirect Storage achieves maximum bandwidth only when the GPU and NVMe drive can communicate directly over PCIe with peer-to-peer support. In many servers, the GPU and drive are on different PCIe root complexes or NUMA nodes, forcing data through host memory and halving effective bandwidth. Checking the PCIe topology with tools like `nvidia-smi topo -m` or `lspci` reveals whether a direct path exists.

Other factors like CPU threads or file system are less likely to cause the specific bandwidth shortfall.

Exam trap

The trap here is assuming that any NVMe drive can achieve full GPUDirect Storage bandwidth regardless of PCIe topology, when peer-to-peer support and NUMA locality are critical.

125
MCQmedium

A financial services company is deploying NVIDIA AI Enterprise in an air-gapped data center. They need to install the NVIDIA GPU Operator on their Kubernetes cluster without internet access. Which additional step must they take to ensure a successful installation?

A.Disable the GPU Operator's driver container and manually install the driver on each node.
B.Use a Kubernetes cluster that supports dynamic volume provisioning for GPU drivers.
C.Enable the GPU Operator's built-in proxy to fetch images from the internet.
D.Configure the GPU Operator to use a private container registry with mirrored images.
AnswerD

In an air-gapped environment, the GPU Operator cannot pull images from public registries. The administrator must mirror all required images, including the operator, driver, device plugin, and container toolkit, to a private registry accessible within the network. The GPU Operator must be configured to pull from this registry by setting the appropriate values in the Helm chart. This step is essential for successful installation.

Why this answer

For an air-gapped installation, the critical step is to make all required container images available within the isolated network. This is done by mirroring the images to a private registry and configuring the GPU Operator to use that registry. The operator's Helm chart provides values to specify the registry and image repository.

Without this, the operator cannot pull the necessary components. Other options either do not address the air-gap limitation or propose unnecessary manual steps.

Exam trap

The trap here is thinking that the GPU Operator has a built-in mechanism to fetch images without internet, or that manual driver installation is required, when the actual solution is to use a private registry with mirrored images.

126
MCQhard

A research group submits a distributed PyTorch training job spanning eight GPUs across two nodes. The job completes but produces a model with accuracy far below the single-node baseline, and logs show that several ranks started training before their peers had initialized the process group. The administrator must ensure that all ranks are launched together and that a failed rank terminates the whole job. Which combination of Kubernetes mechanisms should be used?

A.Schedule the job on a single node with eight GPUs and set the CUDA_VISIBLE_DEVICES variable per container to isolate each rank.
B.Create eight separate Deployments, one per rank, and pass the rank index through an environment variable in each Deployment manifest.
C.Deploy the job as a Kubernetes Job with parallelism set to eight and completions set to one, relying on the default pod startup ordering.
D.Use a gang-scheduling mechanism such as Volcano or the Kubeflow Training Operator with PyTorchJob and configure the rendezvous endpoint so all worker replicas are admitted together.
AnswerD

Gang scheduling admits all replicas atomically, so no rank starts until every peer is schedulable, and the PyTorchJob controller injects the master address, rank, and world size into each pod while restarting or failing the group consistently. This directly fixes the initialization race and enforces all-or-nothing execution.

Why this answer

Distributed training needs collective admission and coordinated rank metadata. A gang scheduler plus the Kubeflow Training Operator's PyTorchJob admits all replicas together and injects rendezvous details, eliminating the race where early ranks initialize before their peers, while the controller fails the whole job if any replica cannot run.

Exam trap

The trap here is treating a plain Kubernetes Job with high parallelism as equivalent to gang scheduling, when it actually offers no atomic admission or rendezvous coordination across ranks.

127
Multi-Selecthard

An administrator is tuning a Kubernetes cluster that runs many small inference pods on NVIDIA GPUs. Utilization is low because each pod reserves a full GPU while using only a fraction of its memory and compute. The administrator wants to share GPUs across pods while preserving memory-level isolation between processes. Which TWO configurations achieve this? (Choose two.)

Select 2 answers
A.Enable the NVIDIA MPS control daemon through the GPU Operator and configure pods to share a GPU with memory limits
B.Configure MIG on the GPUs and expose instances as separate resource types such as nvidia.com/mig-1g.5gb
C.Set the CUDA_VISIBLE_DEVICES environment variable in each pod to a different index of the same physical GPU
D.Enable time-slicing in the device plugin so multiple pods receive the same GPU replica with no memory limit
E.Apply a LimitRange that sets a memory limit on the nvidia.com/gpu resource in the namespace
AnswersA, B

NVIDIA MPS allows multiple CUDA processes to run concurrently on one GPU with hardware-partitioned scheduling and per-client memory limits, so several inference pods can share a device safely. Enabling it via the GPU Operator's ClusterPolicy and setting memory fractions gives the required isolation while raising utilization, which satisfies both goals.

Why this answer

MPS and MIG are the two NVIDIA mechanisms that allow several pods to share a physical GPU while maintaining isolation. MPS enforces per-client memory limits and concurrent execution under a control daemon, while MIG provides hardware-level partitioning with dedicated memory and compute slices advertised as distinct resources. Time-slicing shares a GPU without memory isolation, and environment variables or LimitRanges cannot partition GPU resources.

Exam trap

The trap here is treating time-slicing as sufficient GPU sharing, when it interleaves execution without any memory isolation between processes.

128
Multi-Selecthard

A team is diagnosing a training job that intermittently stalls for several seconds at the start of each epoch. The job uses a distributed data loader and an NVIDIA DGX system with local NVMe. Monitoring shows GPU utilization dropping to near zero during the stalls while host CPU utilization spikes. Which two actions should the AI operations engineer take to identify and mitigate the stall? (Choose two.)

Select 2 answers
A.Increase the batch size until the GPU memory is nearly full to amortize the data loading cost.
B.Use `nvidia-smi dmon` to capture per-second GPU utilization and power draw while the job runs, correlating the drop with the stall window.
C.Set `CUDA_LAUNCH_BLOCKING=1` to serialize kernel launches and make the stall easier to observe.
D.Enable the PyTorch data loader with `num_workers` greater than zero and `pin_memory=True` so that batches are prefetched and staged in page-locked host memory.
E.Move the dataset from local NVMe to a networked file system to rule out local disk contention.
AnswersB, D

`nvidia-smi dmon` provides a lightweight time-series view of utilization, memory, and power. Correlating the utilization dip with the stall timestamp confirms whether the GPU is idle waiting for input rather than compute-bound. This evidence distinguishes data pipeline starvation from other causes such as thermal or power throttling.

Why this answer

The signature of host CPU spikes with idle GPUs is input pipeline starvation. Prefetching with multiple worker processes and pinned memory keeps the device fed, removing the serialized load at epoch boundaries. Concurrently, `nvidia-smi dmon` supplies the time-series evidence that ties the utilization dips to the stall windows, confirming the diagnosis before and after the fix.

Exam trap

The trap here is treating the stall as a GPU or storage throughput problem and attempting to mask it with larger batches instead of addressing the serialized data loader.

129
MCQhard

An administrator is deploying NVIDIA AI Enterprise on a Kubernetes cluster and must decide how GPU workloads should request accelerators. The environment has a mix of full-GPU training jobs and inference services that share a single A100. Which approach correctly allows a pod to consume a specific MIG-backed slice rather than the whole device?

A.Request nvidia.com/gpu: 1 and add the annotation nvidia.com/mig-profile=2g.10gb to the pod metadata.
B.Set the environment variable NVIDIA_MIG_PROFILE=2g.10gb in the container spec and request nvidia.com/gpu: 1.
C.Deploy a separate RuntimeClass named mig-2g.10gb and reference it from the pod's spec.runtimeClassName.
D.Request the specific extended resource, such as nvidia.com/mig-2g.10gb, in the pod's resource limits.
AnswerD

The NVIDIA device plugin advertises each configured MIG profile as its own extended resource name. A pod that requests nvidia.com/mig-2g.10gb in its limits is scheduled onto a node with a free instance of exactly that profile, giving the inference service a dedicated slice while full-GPU training jobs use other devices.

Why this answer

In Kubernetes, GPU allocation is expressed exclusively through extended resource requests handled by the NVIDIA device plugin. When MIG is enabled, each configured profile is advertised under a distinct resource name, so a pod requesting a specific MIG profile is bound to a matching instance. This is what enables a single A100 to serve both whole-device training and sliced inference workloads.

Exam trap

The trap here is believing that annotations or environment variables can steer GPU allocation, when only extended resource requests in the pod spec determine what the device plugin hands out.

130
Multi-Selecthard

Which TWO methods are effective for enforcing GPU resource isolation in a multi-tenant NVIDIA Kubernetes environment?

Select 2 answers
A.Enabling NVIDIA Multi-Instance GPU (MIG) for hardware partitioning.
B.Configuring standard Docker cgroups for GPU memory limits.
C.Applying Kubernetes Taints, Tolerations, and Node Affinity.
D.Implementing standard OS-level priority queuing via 'nice'.
E.Setting a global environment variable for GPU frequency scaling.
AnswersA, C

MIG allows a single GPU to be carved into multiple independent instances, each with its own dedicated memory, cache, and compute cores. This provides strict hardware-enforced isolation, ensuring that one workload cannot interfere with the performance or data security of another tenant running on the same physical chip.

Why this answer

Resource isolation is paramount in multi-tenant environments to prevent noisy neighbor effects where one workload consumes disproportionate GPU cycles. NVIDIA Multi-Instance GPU (MIG) provides hardware-level isolation for partitioning, while Kubernetes-native device plugins with affinity and tolerations provide software-level scheduling control. These mechanisms combined ensure predictable performance and security, preventing cross-tenant interference during intensive training or inference cycles in shared GPU infrastructure clusters.

Exam trap

Candidates tend to pick only software scheduling or only hardware partitioning, forgetting that robust multi-tenant isolation requires a combination of both MIG and Kubernetes controls.

131
MCQmedium

An AI operations team is deploying NVIDIA Base Command Manager to manage a cluster of DGX nodes. They want to ensure that only authorized users can submit jobs and that all job submissions are audited. Which combination of Base Command Manager features should the administrator configure to meet these requirements?

A.Deploy a separate Kubernetes cluster with RBAC policies and use Base Command Manager only for node provisioning.
B.Configure IPMI access controls on each DGX node and enable the Base Command Manager telemetry collector.
C.Enable SELinux enforcing mode on all nodes and configure Base Command Manager to use SSH key-based authentication only.
D.Integrate Base Command Manager with an LDAP or Active Directory identity provider and enable the audit logging feature.
AnswerD

Base Command Manager supports integration with enterprise identity providers such as LDAP or Active Directory for authentication and authorization. Enabling audit logging captures job submission events and user actions. Together, these features enforce authorized access and provide the required audit trail for job submissions in this scenario.

Why this answer

Base Command Manager centralizes cluster administration, including user authentication and job scheduling. Integrating with LDAP or Active Directory ensures only authorized users can authenticate and submit jobs, while audit logging records those submissions for compliance. This combination directly addresses both the access control and auditing requirements in the scenario.

Exam trap

The trap here is assuming that infrastructure hardening measures like SELinux or IPMI controls provide user-level job authorization and audit trails, when those functions require identity integration and audit logging in Base Command Manager.

132
MCQmedium

Which component is responsible for exposing the GPU as a schedulable resource in a Kubernetes cluster?

A.NVIDIA Container Toolkit
B.NVIDIA Device Plugin
C.NVIDIA DCGM Exporter
D.The Kubernetes Scheduler.
AnswerB

The device plugin is the crucial component for Kubernetes integration. It polls the host for GPU status and reports capacity to the kubelet, which then tells the scheduler. This bridge is essential for enabling the 'nvidia.com/gpu' resource type, which allows pods to request GPUs as first-class resources.

Why this answer

The Kubernetes device plugin is the interface that allows the kubelet to communicate with the GPU. It advertises the number of available GPUs on each node to the Kubernetes API server. When a pod requests a GPU, the scheduler uses this information to place the workload on the correct node.

Without this plugin, Kubernetes is unaware of the GPU hardware, making it impossible to manage and allocate GPU resources for containerized workloads.

Exam trap

Candidates often select the NVIDIA Container Toolkit or GPU Operator instead of the specific component that directly advertises resources to the kubelet scheduler.

133
MCQeasy

An AI operations team needs to monitor GPU health and utilization across a fleet of DGX nodes from a single dashboard. They want per-GPU metrics such as power, temperature, utilization, and ECC errors, and they want to retain historical data for capacity planning. Which NVIDIA tool is purpose-built to collect and expose these GPU telemetry metrics for centralized monitoring?

A.NVIDIA DCGM (Data Center GPU Manager) with the DCGM exporter for Prometheus.
B.NVIDIA Base Command Manager's job scheduler logs.
C.NVIDIA Nsight Systems for profiling GPU kernels and collecting timeline traces.
D.NVIDIA CUDA Toolkit's `nvidia-smi` command run manually on each node.
AnswerA

DCGM is the NVIDIA tool designed for data center GPU monitoring and management. It collects health, utilization, power, temperature, and ECC metrics, and the DCGM exporter exposes them to Prometheus for centralized dashboards and long-term retention. This directly matches the requirement for fleet-wide GPU telemetry and historical capacity planning data.

Why this answer

DCGM is NVIDIA's purpose-built data center GPU monitoring and management tool. It gathers health and performance metrics including power, temperature, utilization, and ECC errors, and the DCGM exporter makes them available to Prometheus for centralized dashboards and historical retention. This aligns exactly with the team's need for fleet-wide GPU telemetry and capacity planning data.

Exam trap

The trap here is confusing a profiler or a manual command-line utility with a continuous telemetry pipeline, when only DCGM with its exporter provides centralized, historical GPU health metrics.

134
Multi-Selecthard

A research lab is deploying NVIDIA AI Enterprise on an air-gapped Kubernetes cluster. The cluster has no internet access, and all software must be installed from a local registry. The administrator plans to use the NVIDIA GPU Operator. Which two actions must be performed to ensure a successful deployment in this environment? (Choose two.)

Select 2 answers
A.Install the NVIDIA GPU Operator using the --offline flag.
B.Disable the GPU Operator's driver management and install the driver manually on each node.
C.Configure the GPU Operator to use the private registry by setting the appropriate Helm chart values.
D.Mirror all required NVIDIA container images to a private registry accessible by the cluster.
E.Ensure the cluster nodes have direct access to the NVIDIA licensing server for vGPU.
AnswersC, D

The GPU Operator must be configured to pull images from the private registry instead of the default public registries. This is done by setting Helm values such as operator.repository, driver.repository, toolkit.repository, and others to point to the private registry. Without this configuration, the Operator would attempt to pull from nvcr.io and fail due to no internet access. Thus, this is a required action for air-gapped deployment.

Why this answer

In an air-gapped environment, the GPU Operator cannot reach public registries to pull container images. Therefore, all necessary images must be mirrored to a private registry, and the GPU Operator must be configured to use that registry via Helm values. These two actions ensure that the Operator can deploy its components without internet access.

Other options are either not required or invalid for this scenario.

Exam trap

The trap here is thinking that an offline flag exists or that driver management must be disabled, when the key steps are mirroring images and configuring the private registry.

135
MCQmedium

An administrator notices that GPU utilization on a training cluster hovers around 25 percent even though many jobs are queued. Investigation shows that each job requests a full GPU, but the models are small and alternate between short data-loading phases and brief compute bursts. The administrator wants to increase effective GPU utilization without changing model code. Which action should be taken first?

A.Enable GPU time-slicing or MIG so multiple small jobs can share a device concurrently
B.Move data loading to CPU-only nodes to eliminate the idle phases
C.Add more GPU nodes to the cluster so queued jobs start sooner
D.Increase the batch size of each job so compute bursts last longer
AnswerA

Time-slicing lets several containers share one physical GPU through rapid context switching, and MIG provides isolated slices; both allow small, bursty jobs to occupy a device concurrently instead of idling it during data-loading phases. This raises effective utilization without modifying model code, directly addressing the observed low utilization with queued work.

Why this answer

When small, bursty jobs each hold a full GPU exclusively, the device idles during data-loading windows. Enabling time-slicing or MIG allows multiple jobs to share the device concurrently, filling those idle gaps and raising effective utilization without touching model code or adding hardware.

Exam trap

The trap here is reaching for more GPUs or bigger batches when the real issue is exclusive device occupancy by jobs that spend much of their time idle.

136
MCQmedium

An administrator supports a shared inference cluster where a single A100 GPU must serve several small models concurrently. They configure the NVIDIA device plugin with a time-slicing configuration that advertises multiple replicas of the same physical device. After deployment, users report that one noisy model starves the others and latency spikes unpredictably. Which statement best explains the observed behaviour?

A.Multi-Instance GPU mode must be enabled alongside time-slicing, because the two modes are designed to be combined on the same device.
B.The device plugin failed to register replicas, so the scheduler placed all pods on one logical device instead of distributing them.
C.Time-slicing interleaves work on one physical GPU with no memory or fault isolation, so a heavy workload can dominate device time and memory bandwidth.
D.The pods lack a RuntimeClass reference, so the NVIDIA container runtime never applies the configured replica strategy at container start.
AnswerC

Time-slicing exposes several logical replicas of one physical GPU and lets the driver rotate execution among them, but the replicas share the same memory space and compute engine. There is no partitioning of memory capacity, cache, or bandwidth, so a demanding model can monopolize execution slots and evict or slow others. That matches the reported starvation and unpredictable latency exactly.

Why this answer

Time-slicing creates multiple schedulable replicas of one physical GPU, which improves packing density but provides no hardware-level isolation. All replicas contend for the same memory, caches, and execution bandwidth, so a dominant workload can starve lighter ones and produce erratic tail latency. Where strict isolation or predictable performance is required, hardware partitioning such as Multi-Instance GPU is the appropriate mechanism instead of replica multiplexing.

Exam trap

The trap here is equating multiple advertised GPU replicas with multiple isolated GPUs; time-sliced replicas share one device's memory and compute engines and therefore cannot guarantee fair or predictable service levels.

137
MCQhard

A cluster runs mixed workloads: latency-sensitive inference services and best-effort batch jobs. Administrators observe that batch jobs occasionally occupy all GPUs, causing inference requests to queue and breach service level objectives. They want inference pods to be admitted immediately while allowing batch work to use remaining capacity and be preempted when needed. Which approach should they implement?

A.Enable time-slicing on all GPUs so batch and inference pods share devices.
B.Set a ResourceQuota on the batch namespace limiting total GPU requests.
C.Apply node affinity so inference pods only run on a dedicated subset of nodes.
D.Define priority classes and a preemption policy so inference pods have higher priority than batch jobs.
AnswerD

Kubernetes priority classes let higher-priority pods preempt lower-priority ones when resources are scarce. Assigning inference a high priority and batch a low priority ensures inference is admitted immediately, while batch jobs yield their GPUs when inference arrives. This matches the requirement that batch work uses leftover capacity and is preempted rather than blocking critical services.

Why this answer

The requirement is immediate admission for inference and preemptible use of leftover capacity by batch work. Priority classes with preemption give inference pods the ability to displace lower-priority batch pods when GPUs are scarce, while batch jobs still consume idle capacity during quiet periods. Affinity, quotas, and time-slicing each constrain or share resources but cannot prioritize or preempt running workloads.

Exam trap

The trap here is equating resource sharing or quota limits with prioritization, when only priority and preemption can guarantee immediate admission.

138
MCQmedium

An administrator is preparing a multi-node NVIDIA DGX H100 cluster for a distributed training job using NVIDIA Base Command. The cluster nodes have InfiniBand adapters, but the job's inter-node throughput is far below expectations. The administrator runs `ibstat` and sees that the ports are in the INIT state rather than ACTIVE. Which action should the administrator take first?

A.Enable GPUDirect Storage on each node to bypass the CPU for inter-node communication.
B.Reinstall the NVIDIA GPU driver on all nodes to reset the InfiniBand firmware.
C.Disable ECC memory on the GPUs to increase available bandwidth for inter-node traffic.
D.Verify that the InfiniBand subnet manager is running and that the IPoIB interfaces are configured correctly.
AnswerD

An InfiniBand port remains in INIT when it has not been initialized by a subnet manager, so checking that the subnet manager service is active on the fabric and that IPoIB (or RDMA) interfaces are up is the correct first diagnostic step. Without subnet manager configuration, ports never transition to ACTIVE and inter-node bandwidth collapses regardless of GPU or driver health.

Why this answer

InfiniBand ports stuck in INIT have not been configured by a subnet manager, which is the authoritative service that initializes and activates fabric links. Checking the subnet manager and IPoIB configuration directly addresses why ports never reach ACTIVE. Without an active subnet manager, distributed training traffic cannot flow at expected speeds, so this is the correct first action.

Exam trap

The trap here is assuming that low inter-node throughput is always a GPU or driver problem, when an InfiniBand port in INIT points to a fabric-layer subnet manager or link configuration issue.

139
MCQeasy

An administrator is preparing a Kubernetes cluster for AI workloads and needs to ensure that the NVIDIA GPU Operator can be installed. The cluster nodes have NVIDIA GPUs, and the administrator wants to verify that the nodes are ready. Which command should the administrator run to check if the NVIDIA driver is already loaded on a node?

A.lspci | grep -i nvidia
B.kubectl get nodes -o wide
C.kubectl describe node <node-name>
D.nvidia-smi
AnswerD

nvidia-smi is the NVIDIA System Management Interface command. Running it on a node displays GPU information and driver version, confirming that the driver is loaded and functional. This is the standard way to verify driver installation on a node.

Why this answer

To verify that the NVIDIA driver is loaded on a node, the administrator should run nvidia-smi. This command queries the driver and displays GPU details such as driver version, GPU utilization, and memory usage. It is the definitive check for driver functionality.

Other commands may show GPU hardware or Kubernetes resources but do not confirm driver operation.

Exam trap

The trap here is confusing hardware detection with driver verification; lspci shows the GPU exists, but only nvidia-smi confirms the driver is active.

140
MCQmedium

An administrator is configuring a multi-tenant NVIDIA AI Enterprise environment. Which mechanism is most effective for ensuring hardware-level isolation between concurrent training jobs on a single A100 GPU?

A.Kubernetes namespaces with resource quotas
B.NVIDIA Driver process scheduling priority
C.Multi-Instance GPU (MIG) partitioning
D.Docker container CPU pinning
AnswerC

MIG hardware partitions ensure that each workload receives a dedicated set of compute units and memory buffers. By physically isolating the GPU resources, you guarantee deterministic performance for each tenant, which is necessary when running sensitive or high-throughput AI training models concurrently on a single hardware accelerator.

Why this answer

NVIDIA Multi-Instance GPU (MIG) allows a single physical GPU to be partitioned into multiple isolated instances, each with dedicated memory and compute cores. This is critical in multi-tenant AI environments because it prevents noisy neighbor issues, ensuring one workload's memory usage or compute demand does not degrade the performance of another. Proper resource partitioning is essential for maintaining strict SLAs and security boundaries within shared infrastructure.

Exam trap

Candidates often suggest software-level container limits or time-slicing when the question specifically asks for hardware-level isolation using partitioning features like MIG.

141
MCQmedium

An administrator supports a shared Kubernetes cluster running NVIDIA GPU Operator. Data scientists report that their inference pods remain in Pending state, yet the GPU Operator pods and node feature discovery pods are healthy, and the GPU nodes show no hardware alarms. The administrator confirms that the cluster has a mixture of MIG-capable A100 nodes and non-MIG T4 nodes. Which immediate administrative action is most appropriate to diagnose the scheduling failure?

A.Inspect the pod events with kubectl describe pod and verify the MIG strategy configured in the device plugin, because a mismatch between the node's MIG configuration and the requested resource can prevent GPU resource advertisement.
B.Restart the kubelet on all worker nodes to force the device plugin to re-register GPU resources, because a stale kubelet cache is the most frequent cause of Pending workloads.
C.Delete the pending pods and recreate them with a higher priority class, because the scheduler may be preempting them in favor of system pods.
D.Scale the GPU Operator controller manager to zero replicas and back to one, because the operator may have lost track of node labels after a recent upgrade.
AnswerA

Pod events reveal whether the scheduler rejected the pod due to missing nvidia.com/gpu or nvidia.com/mig-* resources. If the device plugin advertises MIG profiles but the workload requests a full GPU (or vice versa), the pod stays Pending. Checking the MIG strategy and node labels directly addresses this common scheduling mismatch.

Why this answer

The pod events are the fastest way to see whether the scheduler cannot find a suitable node due to resource requests. In a mixed MIG and non-MIG environment, the device plugin advertises different resource names, and a mismatch between the requested GPU resource and what the node exposes leaves pods Pending. Verifying the MIG strategy and node labels directly identifies that mismatch.

Exam trap

The trap here is assuming that healthy GPU Operator pods guarantee that GPU resources are correctly advertised and schedulable, when in fact a MIG strategy mismatch can leave pods Pending without any operator-level failure.

142
MCQmedium

A system administrator is installing NVIDIA AI Enterprise on a Kubernetes cluster that will use Multi-Instance GPU (MIG) on A100 GPUs. The administrator wants to ensure that MIG instances are properly exposed as schedulable resources. Which action must be taken after enabling MIG mode on the GPUs?

A.Manually create Kubernetes custom resources for each MIG instance using the NVIDIA MIG Manager.
B.Set the environment variable NVIDIA_MIG_CONFIG_DEVICES to 'all' on the kubelet and restart the kubelet service.
C.Deploy a separate device plugin for each MIG instance using a DaemonSet with node affinity to the specific GPU.
D.Install the NVIDIA GPU Operator with the MIG strategy set to 'mixed' and configure the device plugin to advertise MIG resources.
AnswerD

The GPU Operator supports MIG by deploying a device plugin that advertises MIG instances as resources. Setting the MIG strategy to 'mixed' allows both MIG and non-MIG GPUs in the cluster. The device plugin then exposes each MIG instance as a schedulable resource, enabling pods to request specific MIG profiles. This is the correct way to integrate MIG with Kubernetes scheduling.

Why this answer

To expose MIG instances as schedulable resources in Kubernetes, the NVIDIA GPU Operator must be installed with the MIG strategy configured (e.g., 'mixed' or 'single'). The Operator's device plugin then discovers and advertises each MIG instance as a resource. This allows pods to request specific MIG profiles via resource limits, enabling efficient scheduling.

Exam trap

The trap here is assuming that enabling MIG mode on the GPU is sufficient, when Kubernetes also requires the device plugin to advertise MIG resources.

143
MCQeasy

An administrator is deploying the NVIDIA GPU Operator on a Kubernetes cluster using Helm. The cluster nodes have NVIDIA GPUs and the administrator wants to ensure that the GPU Operator can automatically label nodes with GPU properties and install the device plugin. Which prerequisite must be met on the cluster nodes before installing the GPU Operator?

A.The nodes must have a supported container runtime such as Docker or containerd installed and configured.
B.The NVIDIA driver must be pre-installed on the nodes.
C.The nodes must have the NVIDIA Container Toolkit installed.
D.The nodes must have the NVIDIA DCGM exporter pre-installed for monitoring.
AnswerA

The GPU Operator relies on a container runtime to deploy its components and to run GPU-accelerated workloads. Kubernetes requires a container runtime like Docker or containerd to be installed and configured on each node. This is a fundamental prerequisite for any Kubernetes cluster, and the GPU Operator assumes it is present. Without a container runtime, the Operator cannot schedule its daemonsets, making this the correct prerequisite.

Why this answer

Before installing the NVIDIA GPU Operator, the Kubernetes cluster must have a container runtime such as Docker or containerd installed and configured on all nodes. This is a basic Kubernetes requirement because the Operator deploys its components as containers. Other components like the NVIDIA driver, container toolkit, and DCGM exporter are managed by the Operator itself and do not need to be pre-installed.

Exam trap

The trap here is assuming that NVIDIA-specific components like the driver or container toolkit must be pre-installed, when the GPU Operator is designed to manage them automatically.

144
MCQeasy

An administrator is installing the NVIDIA GPU Operator on a Kubernetes cluster. They want to verify that the GPU Operator's components are running correctly after installation. Which command should they use to check the status of the GPU Operator pods?

A.kubectl get pods -n gpu-operator
B.kubectl get nodes --show-labels
C.helm list -n gpu-operator
D.nvidia-smi -q
AnswerA

The NVIDIA GPU Operator installs its components in the 'gpu-operator' namespace by default. Running 'kubectl get pods -n gpu-operator' lists all pods in that namespace, allowing the administrator to verify that the operator and its managed components are running. This is the standard way to check the status.

Why this answer

The GPU Operator deploys its components into the 'gpu-operator' namespace. Checking the pods in that namespace with 'kubectl get pods -n gpu-operator' directly shows whether the operator and its managed pods are running. Other commands like nvidia-smi or helm list provide different information and do not replace the need to inspect pod status.

Exam trap

The trap here is confusing deployment verification with host-level GPU queries or Helm release checks, when the direct way to see component status is to list the pods in the operator's namespace.

145
MCQeasy

A cloud operations engineer is deploying the NVIDIA GPU Operator on a managed Kubernetes service where the worker nodes already have the NVIDIA data center driver installed by the cloud provider. The team wants the Operator to manage only the device plugin, container toolkit, and monitoring components. Which Helm value should the engineer set during installation?

A.--set driver.enabled=false
B.--set operator.driver.install=false
C.--set mig.strategy=none
D.--set toolkit.enabled=false
AnswerA

Setting driver.enabled=false tells the GPU Operator not to deploy the driver container, so it uses the preinstalled host driver. The Operator still deploys the device plugin, container toolkit, DCGM exporter, and other components. This is the standard approach on managed services where the provider owns the driver lifecycle and node image updates.

Why this answer

The GPU Operator chart exposes driver.enabled to control whether the driver container is deployed. On managed Kubernetes services the node image already contains a validated NVIDIA driver, so disabling the Operator's driver avoids conflicts and duplicate work while still letting the Operator manage the device plugin, container toolkit, DCGM exporter, and related components. The other values either do not exist or disable components the team needs.

Exam trap

The trap here is inventing or misremembering a Helm key for skipping the driver, when the actual chart value is driver.enabled and the goal is to keep all other GPU software components Operator-managed.

146
MCQmedium

An AI operations engineer is optimizing a real-time inference pipeline on an NVIDIA T4 GPU. The pipeline uses TensorRT and receives requests with variable input sizes. Profiling shows that the engine recompiles for each new input shape, causing latency spikes. Which optimization should the engineer apply to eliminate recompilation while maintaining acceptable accuracy?

A.Enable dynamic shaping in TensorRT by defining an optimization profile with min, opt, and max shapes.
B.Pad all input tensors to a fixed maximum size before inference.
C.Convert the model to use INT8 precision with a calibration dataset to reduce inference time.
D.Increase the workspace size allocated for TensorRT to allow more kernel autotuning.
AnswerA

TensorRT dynamic shaping allows a single engine to handle a range of input dimensions without recompiling. By specifying an optimization profile with min, opt, and max shapes, the engine is built once and can process any input within that range, eliminating recompilation latency spikes. This is the correct approach for variable input sizes while balancing performance and accuracy.

Why this answer

TensorRT dynamic shaping with an optimization profile allows one engine to handle a range of input shapes, eliminating the need to rebuild the engine for each new size. This directly removes the latency spikes from recompilation while maintaining performance across variable inputs. Other options either do not address recompilation or trade off too much efficiency or accuracy.

Exam trap

The trap here is thinking that precision reduction or workspace tuning solves shape variability, when the root cause is engine recompilation for new shapes.

147
MCQmedium

During deployment, an AI model experiences high variance in latency during inference. The system uses a fixed instance count. What is the most likely cause for this performance jitter?

A.The GPU memory is over-provisioned.
B.OS context switching and interrupt handling.
C.The model is too large for the GPU cache.
D.The network switch is experiencing congestion.
AnswerB

Latency variance is frequently caused by the OS scheduling other tasks or handling interrupts on the same cores used by the inference server. This creates 'jitter' as the processor pauses inference tasks to handle housekeeping. Pinning the inference server to dedicated, isolated cores effectively minimizes this contention and stabilizes latency.

Why this answer

Inference latency jitter is often caused by external processes or system interrupts competing for CPU resources that the inference server relies on for data pre-processing or orchestration. By pinning CPU cores to the inference process (CPU affinity) and using isolated cores, the administrator can prevent these context switches and OS interrupts, leading to more consistent and predictable inference response times for production workloads.

Exam trap

Candidates often misattribute inference latency jitter to model complexity or network overhead, ignoring local OS operations like context switching and CPU interrupt handling that disrupt execution consistency.

148
MCQhard

Refer to the exhibit. An engineer is troubleshooting inconsistent training performance across two GPUs in a single DGX node. Why is one GPU reporting a lower clock speed despite being in P0 state?

A.The GPU is misconfigured in the system BIOS.
B.The GPU is experiencing thermal throttling due to improper airflow or fan failure.
C.The CUDA driver is applying a different profile to each GPU.
D.The GPU is running an outdated firmware version.
AnswerB

When a GPU reaches its thermal limit, the firmware automatically reduces the graphics clock to prevent physical damage. A discrepancy in clock speeds between two identical GPUs in the same P0 state indicates that one card is running hotter than the other due to cooling issues.

Why this answer

The P0 power state represents maximum performance, but the actual clock frequency is governed by the hardware's thermal and power monitoring sub-systems. When two identical GPUs show different clock speeds in the same state, it suggests external environmental factors or hardware degradation. This distinction is critical for AI operations to maintain balanced parallel training workloads and ensure distributed training does not bottleneck on the slowest card.

Exam trap

Candidates often assume the GPU is faulty and needs replacement. They fail to consider environmental factors like airflow or fan failure, which trigger thermal throttling even on healthy hardware.

149
MCQhard

When running multi-instance GPU (MIG) workloads, what is the main advantage of assigning specific MIG profiles to different Kubernetes namespaces?

A.It allows the GPU to switch between different CUDA versions per instance.
B.It enhances GPU memory security by isolating workload address spaces.
C.It automatically compresses the model weights for faster loading.
D.It allows the cluster to bypass standard Kubernetes scheduler policies.
AnswerB

MIG profiles provide hardware-level isolation of memory and compute resources. By assigning specific profiles to namespaces, you effectively prevent cross-talk between different workloads' address spaces. This is a critical security and operational feature for multi-tenant environments, ensuring that sensitive data in one inference pod cannot be accessed by another.

Why this answer

MIG profiles allow for fine-grained hardware isolation, ensuring that different workloads—such as small inference tasks and large training runs—do not interfere with each other. By mapping profiles to specific namespaces, administrators can enforce strict hardware separation, preventing 'noisy neighbor' issues where one workload consumes memory bandwidth or compute cycles required by another, thus improving overall cluster stability and predictability.

Exam trap

Candidates often assume MIG is primarily for performance tuning or cost optimization, missing that its fundamental architectural strength in Kubernetes is hardware-level memory isolation between distinct, potentially insecure, user workloads.

150
MCQhard

An administrator is responsible for a large NVIDIA DGX SuperPOD used for multi-node training. They need to ensure that GPU telemetry and health metrics are collected centrally and can trigger alerts when GPUs exceed temperature thresholds. Which component of NVIDIA Base Command Manager (BCM) should they configure to achieve this?

A.NVIDIA Base Command Manager's built-in 'cluster monitoring' dashboard
B.NVIDIA Management Library (NVML) on each node
C.NVIDIA Container Toolkit on each compute node
D.NVIDIA Data Center GPU Manager (DCGM) integrated with BCM
AnswerD

DCGM is designed for centralized GPU telemetry, health monitoring, and policy enforcement in data center environments. BCM integrates with DCGM to collect metrics from all nodes and can forward them to monitoring systems like Prometheus for alerting. Configuring DCGM within BCM provides the required centralized collection and threshold-based alerts for temperature and other metrics.

Why this answer

NVIDIA DCGM is the standard tool for centralized GPU telemetry and health monitoring in data centers. When integrated with Base Command Manager, it collects metrics from all nodes and can be configured with alerting rules for temperature and other thresholds. This provides the required centralized visibility and automated alerts for the DGX SuperPOD.

Exam trap

The trap here is confusing a monitoring dashboard with the actual telemetry collection component; the dashboard displays data but does not collect or alert on its own.

Page 1

Page 2 of 5

Page 3

All pages