Courseiva

CCNA Troubleshooting and Optimization Questions

75 of 86 questions · Page 1/2 · Troubleshooting and Optimization · Answers revealed

1
MCQmedium

During a multi-GPU training job, you notice that one GPU consistently reports lower utilization and longer communication times compared to others. What is the most likely reason for this performance imbalance?

A.The GPU is defective and needs replacement.
B.The GPU is connected via a slower interconnect.
C.The training framework is not using CUDA streams.
D.The ambient server temperature is too high.
AnswerB

If one GPU in a cluster lacks the high-speed NVLink connection enjoyed by the others, it will be limited by the bandwidth of the slower interface (like PCIe). During AllReduce operations, the entire cluster must wait for this 'straggler' to complete, leading to lower total utilization and increased latency.

Why this answer

In multi-GPU systems, uneven load distribution often stems from mismatched NVLink topologies or PCIe lane configurations. If one GPU is connected via a slower PCIe link rather than a high-speed NVLink interconnect, it becomes the bottleneck in collective communication operations like AllReduce. Ensuring symmetric connectivity across all GPUs is essential for predictable performance and preventing the 'straggler' effect in distributed deep learning training.

Exam trap

Candidates often attribute performance imbalances to software bugs or uneven data batching, failing to account for physical hardware topology issues, such as mismatched PCIe lanes or NVLink connectivity.

2
MCQmedium

An AI researcher is using Nsight Systems to profile an application. They notice a large gap in the timeline where neither the CPU nor the GPU is performing significant work. What does this gap most likely represent?

A.The GPU is busy performing background garbage collection.
B.The application is performing blocking synchronization.
C.The profiler has encountered a buffer overflow.
D.The system is entering a power-saving state.
AnswerB

Gaps in a timeline often signify synchronization barriers where the host is waiting for a device to finish, or vice versa. This effectively pauses the entire execution pipeline. By using non-blocking CUDA streams and events, these synchronization gaps can be removed, allowing the system to maintain continuous compute throughput.

Why this answer

Large gaps in profiling timelines often indicate synchronization points where the CPU is waiting for the GPU to finish a task, or vice versa, due to blocking operations. These 'stalls' are common in poorly optimized code that issues many synchronous calls. Identifying these gaps allows the developer to re-structure the code to use asynchronous execution, thereby overlapping compute and data movement to improve performance.

Exam trap

Candidates often assume that a timeline gap means the hardware is broken or under-powered, overlooking software-level synchronization barriers where the CPU and GPU are simply waiting on each other.

3
MCQmedium

What is the most accurate way to verify that a training job is utilizing Tensor Cores?

A.Check the GPU power consumption in nvidia-smi.
B.Monitor the SM Tensor utilization metric via Nsight Compute.
C.Check the memory bandwidth in DCGM.
D.Verify the CUDA driver version is above 500.0.
AnswerB

Nsight Compute provides specific hardware counters for SM Tensor utilization. This allows an engineer to directly verify if the kernels are issuing the specific matrix multiply-accumulate instructions that run on Tensor Cores, confirming that the hardware acceleration is active for the workload.

Why this answer

Tensor Cores are specialized hardware units for mixed-precision matrix operations. Monitoring the SM occupancy and specific instruction usage via Nsight Compute is the definitive way to confirm their engagement. Understanding how to verify this is essential for engineers to ensure that their optimization efforts, such as mixed-precision training, are actually resulting in the intended hardware utilization and performance improvements.

Exam trap

Candidates often assume that high GPU utilization automatically implies Tensor Core usage. They fail to distinguish between general CUDA core compute and specialized matrix operations performed by Tensor Cores.

4
MCQeasy

An AI operations engineer is deploying a model on an NVIDIA GPU and notices that inference latency is higher than expected. The model uses a batch size of 1, and profiling shows that the GPU is idle between kernel launches. Which optimization technique should the engineer use to reduce latency by overlapping data transfer with computation?

A.Enable NVIDIA MPS to allow multiple processes to share the GPU.
B.Switch to a lower-precision data type such as FP16 to reduce memory transfer size.
C.Increase the batch size to improve GPU utilization and reduce per-inference latency.
D.Use CUDA streams to overlap host-to-device memory copies with kernel execution.
AnswerD

CUDA streams allow asynchronous execution of memory copies and kernels, enabling overlap of data transfer with computation. For a batch size of 1, where kernels are short, this overlap can hide transfer latency and reduce overall inference time. The engineer should use separate streams for data transfer and compute, and synchronize appropriately, to achieve the desired latency reduction.

Why this answer

The idle time between kernel launches indicates that data transfers and kernel executions are serialized. Using CUDA streams to overlap host-to-device copies with kernel execution can hide transfer latency and reduce overall inference time. This is the correct technique for overlapping data transfer with computation, especially with small batch sizes where transfer overhead is significant.

Exam trap

The trap here is assuming that increasing batch size or using lower precision will reduce latency, when the issue is serialization of transfers and kernels.

5
MCQhard

A production inference service using TensorRT is showing lower than expected throughput. Profiling shows that the model is spending significant time in "host-to-device" transfers. What is the best optimization?

A.Increase the GPU clock speed via nvidia-smi.
B.Use pinned memory for data transfers.
C.Switch to FP64 precision for higher accuracy.
D.Implement multi-threaded data preprocessing on the GPU.
AnswerB

Pinned memory (page-locked memory) allows for faster transfer rates between the CPU and GPU because it enables the GPU to perform direct memory access without the CPU needing to copy data to a temporary buffer. This significantly reduces the overhead of host-to-device transfers in high-throughput inference pipelines.

Why this answer

Data movement is a primary latency bottleneck in deep learning inference. By using pinned memory (page-locked memory) for host-side buffers, the system can enable faster direct memory access (DMA) transfers between the CPU and the GPU. This minimizes the time spent in data copy operations, directly reducing latency and increasing total throughput, which is essential for meeting strict Service Level Agreements (SLAs) in production AI deployments.

Exam trap

Candidates often suggest optimizing the model architecture or increasing batch size, ignoring the fundamental I/O bottleneck caused by standard pageable memory transfers between the CPU and GPU host-device boundary.

6
Multi-Selecthard

Which TWO of the following actions should be taken to optimize GPU memory usage when encountering Out-of-Memory (OOM) errors during model training?

Select 2 answers
A.Implement gradient checkpointing to trade compute for memory.
B.Increase the number of CPU threads in the data loader.
C.Use mixed-precision training (FP16/BF16) to reduce weight storage.
D.Disable the use of NCCL to reduce inter-node memory overhead.
E.Switch from the Adam optimizer to Stochastic Gradient Descent (SGD).
AnswersA, C

Gradient checkpointing stores only a subset of activations and recomputes the rest during the backward pass. This significantly reduces the memory footprint of the activation graph, allowing for larger models or batch sizes to fit in memory at the cost of additional compute cycles during training.

Why this answer

OOM errors are common in deep learning when the model size, activation memory, and batch size exceed VRAM capacity. Implementing gradient checkpointing and mixed-precision training are standard industry practices to manage memory pressure. Mastery of these techniques is essential for AI operations engineers to maintain system stability and enable the training of large-scale models without requiring immediate hardware upgrades.

Exam trap

Candidates often select 'increasing batch size' or 'adding more hardware' as solutions. These actually exacerbate OOM errors or ignore the requirement to optimize memory within existing hardware constraints.

7
MCQmedium

A team is deploying a large language model for inference using NVIDIA Triton Inference Server on a GPU. They observe that the first inference request has high latency compared to subsequent requests. What is the most likely cause and the appropriate optimization?

A.The GPU is thermal throttling on the first request; set a higher power limit using nvidia-smi -pl to stabilize performance.
B.The model is not using TensorRT; converting it to a TensorRT engine will eliminate first-request latency.
C.The model uses dynamic batching; disable dynamic batching to ensure the first request is processed immediately.
D.The first request triggers model loading and CUDA context initialization; enable model warmup in Triton to pre-load and initialize the model.
AnswerD

Triton's model warmup feature allows the server to run dummy inferences during initialization, loading the model into GPU memory and initializing CUDA contexts. This moves the overhead from the first real request to server startup, reducing first-request latency. It is a standard practice for latency-sensitive deployments. The warmup can be configured with sample inputs to cover typical shapes.

Why this answer

The first inference request often incurs overhead from loading the model into GPU memory, compiling kernels, and initializing CUDA contexts. Triton's model warmup feature performs dummy inferences at startup, effectively pre-warming the model. This shifts the initialization cost away from the first real request, resulting in consistent low latency for all requests.

Exam trap

The trap here is attributing first-request latency to the inference engine or batching strategy, when it is actually caused by lazy initialization that warmup can mitigate.

8
MCQmedium

An AI Operations engineer is managing a multi-node training job using NVIDIA NCCL. The logs indicate frequent 'NCCL WARN' messages related to 'net_ib_init' failures. What is the most likely cause of this issue?

A.Incompatible NCCL versions across the cluster nodes.
B.Misconfigured InfiniBand subnet manager or mismatched firmware.
C.Insufficient system RAM on the worker nodes.
D.Incorrect GPU driver version installed on the login node.
AnswerB

NCCL requires a correctly configured InfiniBand fabric to initialize efficiently. 'net_ib_init' errors almost exclusively result from either improper Subnet Manager configuration, outdated HCA firmware, or physical layer issues that prevent the NCCL communicator from establishing the required high-speed RDMA connections between GPUs in the cluster.

Why this answer

NCCL relies on InfiniBand (IB) for high-speed inter-GPU communication in multi-node setups. Failures in 'net_ib_init' suggest that the network configuration or driver stack is misaligned between nodes. Identifying and resolving these connectivity issues is critical because NCCL is the backbone of distributed training; any failure here leads to synchronized hangs, data corruption, or severe training slowdowns, undermining the benefits of cluster-level resource parallelism.

Exam trap

Candidates often assume 'net_ib_init' failures are application-level coding bugs within the training script, rather than recognizing them as infrastructure-level configuration issues within the InfiniBand fabric.

9
Multi-Selecthard

An AI operations team is troubleshooting a distributed training job on an NVIDIA DGX SuperPOD that uses NCCL for inter-GPU communication. The job intermittently hangs during the all-reduce phase. Which two actions should be taken to diagnose and resolve the issue? (Choose two.)

Select 2 answers
A.Disable the use of InfiniBand and force NCCL to use TCP sockets for communication.
B.Set NCCL_DEBUG=INFO and NCCL_DEBUG_SUBSYS=INIT,COLL to capture detailed NCCL initialization and collective logs.
C.Verify that all nodes have consistent NCCL versions and that the network interfaces used for communication are up and have sufficient bandwidth.
D.Restart the training job with a smaller number of GPUs to see if the hang persists.
E.Increase the batch size to reduce the frequency of all-reduce operations.
AnswersB, C

Enabling NCCL debug logging provides insights into the communication setup and collective operations. It can reveal misconfigurations, such as incorrect network interface selection or topology issues, that cause hangs. This is a standard first step in diagnosing NCCL-related problems, as it shows the chosen algorithm and any errors during initialization or execution.

Why this answer

Intermittent hangs in NCCL all-reduce often stem from configuration or network issues. Enabling detailed NCCL logging helps identify the exact failure point, while verifying consistent NCCL versions and network interface health addresses common root causes. These two actions together provide both diagnostic information and a path to resolution without degrading performance.

Exam trap

The trap here is thinking that reducing batch size or GPU count will solve the hang, but those are workarounds that do not diagnose or fix the underlying communication problem.

10
MCQhard

A production inference service running on NVIDIA T4 GPUs shows that GPU utilization is consistently below 20% while request latency is high. Profiling with Nsight Systems reveals that the model execution time is short but there are frequent gaps between kernels. Which optimization should be applied first to improve GPU utilization?

A.Increase the batch size in the inference server configuration.
B.Switch from FP32 to FP16 precision for the model weights.
C.Increase the number of concurrent model instances on each GPU.
D.Enable CUDA graphs to capture and replay the inference sequence.
AnswerA

Increasing batch size allows more requests to be processed per kernel launch, reducing the relative overhead of kernel launch gaps and improving GPU utilization. With small batches, the GPU sits idle between kernels. Larger batches keep the GPU busy and amortize launch overhead, directly addressing the observed gaps and low utilization. This is a standard first optimization for underutilized inference GPUs.

Why this answer

The low GPU utilization and gaps between kernels indicate that the GPU is not receiving enough work per launch. Increasing the batch size allows more data to be processed per kernel, filling the gaps and raising utilization. This is the most direct and effective first step before considering more complex techniques like CUDA graphs or precision changes.

Exam trap

The trap here is focusing on kernel launch overhead as the primary cause, when the real issue is insufficient work per launch due to small batch sizes.

11
MCQeasy

A data scientist reports that a Jupyter notebook running on a GPU-enabled server is extremely slow when training a small neural network, even though nvidia-smi shows the GPU is idle. The notebook uses TensorFlow. Which is the most likely cause?

A.TensorFlow is not configured to use the GPU; it is running on the CPU.
B.The GPU is being used by another process, causing contention.
C.The Jupyter notebook kernel needs to be restarted to detect the GPU.
D.The neural network is too small to benefit from GPU acceleration, so TensorFlow automatically uses the CPU.
AnswerA

If the GPU is idle while training, TensorFlow is likely defaulting to CPU execution. This can happen if the GPU is not visible to TensorFlow due to missing CUDA libraries, incorrect environment variables, or a CPU-only TensorFlow installation. Checking tf.config.list_physical_devices('GPU') would confirm. This is a common oversight in notebook environments where the kernel may not have GPU access.

Why this answer

The idle GPU during training strongly suggests TensorFlow is not using it. Common causes include missing GPU support in TensorFlow, incorrect CUDA/cuDNN versions, or the process not having access to the GPU. Verifying TensorFlow's device configuration is the first step.

Other options are less likely given the evidence.

Exam trap

The trap here is assuming the GPU is too busy or the model too small, when the real issue is that TensorFlow is not configured to use the GPU at all.

12
MCQmedium

When troubleshooting a NCCL collective communication timeout in a distributed training environment, which component should be the primary focus of initial investigation?

A.The GPU driver version on the master node.
B.The network interface card (NIC) topology and interconnect health.
C.The model weight initialization strategy.
D.The local disk I/O latency for checkpoint saving.
AnswerB

NCCL relies heavily on high-speed interconnects like InfiniBand or RoCE. Timeout errors typically indicate that packets are being dropped or delayed beyond the threshold defined in the NCCL configuration, pointing directly to network-related bottlenecks, faulty hardware, or incorrect fabric configuration between the participating nodes.

Why this answer

NCCL timeouts are frequently caused by network congestion, MTU mismatches, or faulty interconnect cables between nodes. By verifying the physical and logical network path, engineers can isolate whether the issue is at the application layer or the fabric layer. This focus is critical because distributed training is highly sensitive to latency and packet loss across the GPU cluster interconnect.

Exam trap

Candidates often assume NCCL timeouts are purely application bugs in the deep learning framework, leading them to unnecessarily debug model code instead of checking network hardware.

13
MCQeasy

An AI operations engineer is validating a new NVIDIA-certified server before putting it into production. The job runs correctly but the team wants to confirm that the GPUs are operating at the expected clocks and not being limited by power or thermal constraints. Which command provides the most direct evidence of the current power and thermal limits and any throttling reasons?

A.`nvidia-smi --query-gpu=name,driver_version --format=csv`
B.`nvidia-smi topo -m`
C.`nvidia-smi --gpu-reset -i 0`
D.`nvidia-smi -q -d PERFORMANCE`
AnswerD

The performance query section reports the current clocks, the active performance state, and the specific throttle reasons such as power cap, thermal limit, or hardware slowdown. This directly answers whether the GPU is being limited and why, making it the most appropriate single command for this validation step.

Why this answer

The performance query reports clocks, performance state, and explicit throttle reason flags such as power cap, thermal slowdown, and hardware limit. That combination gives direct evidence of whether a GPU is constrained and by what. Other commands either report inventory or topology and cannot answer the throttling question.

Exam trap

The trap here is choosing a command that lists GPU inventory or topology, which looks authoritative but contains no clock, power, or throttle information.

14
MCQmedium

An administrator is optimizing a large model training job to reduce checkpointing time to storage. Which strategy is most effective for minimizing the impact on training throughput?

A.Use faster NVMe drives for checkpoint storage.
B.Implement asynchronous checkpointing.
C.Increase the checkpoint frequency.
D.Compress the model weights before saving.
AnswerB

Asynchronous checkpointing allows the training process to save the model state in the background without blocking the main training loop. By offloading the serialization and I/O to a separate thread, the GPU remains free to continue compute-intensive operations, significantly increasing overall training throughput and reducing total job wall-clock time.

Why this answer

Checkpointing large models involves writing gigabytes of data to storage, which can pause training for significant periods. Using asynchronous checkpointing or offloading the save process to a background thread allows the GPU to continue training while the data is written to persistent storage. This eliminates the I/O wait time, ensuring that the heavy computational resources are not wasted during the periodic saving of model states.

Exam trap

Candidates often suggest faster storage hardware. While helpful, it does not solve the fundamental issue of the GPU being blocked by synchronous I/O operations during the checkpointing process.

15
MCQhard

A multi-tenant NVIDIA GPU-accelerated Kubernetes cluster utilizing NVIDIA AI Enterprise experiences intermittent out-of-memory errors on Triton Inference Server pods despite adequate node memory reservation. Which monitoring and troubleshooting action correctly isolates the root cause?

A.Analyze cAdvisor container memory metrics using kubectl top pods to evaluate resident set size growth over sustained inference workloads.
B.Review the Kubernetes cluster autoscaler logs to determine if node scaling events triggered unexpected pod evictions and restarts.
C.Inspect DCGM-Exporter metrics for DCGM_FI_DEV_FB_FREE and DCGM_FI_DEV_GPU_TEMP alongside Triton server request concurrency logs.
D.Execute systemctl status containerd on the worker node to verify container runtime stability and daemon responsiveness.
AnswerC

Tracking free frame buffer metrics and GPU temperatures via DCGM-Exporter reveals exact memory headroom and potential throttling conditions during peak concurrency. Correlating these metrics with inference request logs confirms whether dynamic batching parameters exceeded available device memory.

Why this answer

Inspect NVIDIA Data Center GPU Manager metrics via Prometheus and DCGM-Exporter to capture real-time device memory consumption patterns. This practice is critical because standard Kubernetes container metrics fail to track internal GPU frame buffer allocations and fragmentation specific to deep learning inference engines.

Exam trap

Many administrators mistakenly rely solely on standard cAdvisor memory metrics exposed by Kubernetes, completely missing GPU-specific memory exhaustion occurring inside the device driver context.

16
MCQmedium

An AI administrator is tasked with monitoring GPU utilization in a multi-user cluster. Which tool provides the most granular real-time visibility into process-level GPU memory usage and compute utilization?

A.The system 'top' utility.
B.The NVIDIA 'nvidia-smi' tool.
C.The 'df' command.
D.The system network logs.
AnswerB

NVIDIA-SMI is the standard interface for querying the status of NVIDIA GPU devices. It displays real-time statistics including per-process memory consumption, duty cycle, and power usage. This granularity is essential for identifying which specific user processes are saturating the GPU or causing resource conflicts in a multi-user environment.

Why this answer

The 'nvidia-smi' utility is the foundational tool for monitoring NVIDIA GPU hardware. It provides direct, real-time feedback on memory usage, compute utilization, and process-level diagnostics. For cluster-wide administration, it is the primary command-line tool used to identify exactly which processes are consuming resources, allowing administrators to troubleshoot contention and manage workloads effectively across the available GPU hardware resources.

Exam trap

Candidates often suggest high-level monitoring dashboards or cloud-native tools. While useful, they lack the immediate, process-level granularity required for troubleshooting specific resource contention on a local DGX node.

17
MCQhard

Refer to the exhibit. An administrator notices that 'user_a' is consistently hitting resource limits despite having sufficient total system GPU memory. Based on the policy JSON, what is the cause?

A.The system has reached the max_concurrent_jobs limit.
B.The user is exceeding the 16GB VRAM limit.
C.The NVIDIA driver is blocking the user's access.
D.The GPU memory is fragmented across the nodes.
AnswerB

The JSON clearly defines a resource quota of '16GB'. When the user's training or inference job requests memory beyond this allocated limit, the enforcement mechanism blocks further allocation. This is a deliberate configuration to ensure fair resource distribution among multiple users in a shared GPU cluster environment.

Why this answer

The policy specifies a hard resource quota of '16GB' for the user. Even if the total system memory is higher, the scheduler enforces this limit per user. When the user's workload attempts to allocate more than 16GB of VRAM, the job will fail or throttle.

This is a common method for preventing single users from starving the entire cluster of shared resources.

Exam trap

Candidates often look at the total cluster capacity rather than the specific user-level quota. They assume the user can access all available VRAM, ignoring the scheduler's hard limit policy.

18
MCQhard

An inference model running on Triton Inference Server is reporting high latency for requests. The model uses a fixed-size batching strategy. What is the most effective way to optimize throughput while maintaining latency targets?

A.Increase the number of instances for the model.
B.Configure dynamic batching with a maximum delay.
C.Disable all logging to reduce CPU overhead.
D.Force the model to run on the CPU.
AnswerB

Dynamic batching allows the server to aggregate multiple requests into a single inference call. Setting a maximum delay ensures that the server waits just long enough to improve throughput without exceeding the latency budget. This is the standard method to optimize GPU inference workloads on Triton Inference Server.

Why this answer

Dynamic batching is a powerful feature in Triton Inference Server that groups individual requests together to saturate the GPU's compute capability. By configuring the 'max_queue_delay_microseconds', the system waits briefly to aggregate requests, significantly increasing throughput. This optimization is crucial for balancing the trade-off between individual request latency and overall system efficiency, ensuring that the GPU is not performing trivial computations for tiny batches.

Exam trap

Candidates often choose static batch resizing or manual request throttling, which fails to automatically adapt to fluctuating incoming request rates and traffic spikes.

19
MCQmedium

Refer to the exhibit. The training job fails with a CUDA OOM error. Given the memory profile, which optimization strategy provides the most immediate relief while maintaining model performance?

A.Increase the batch size to 256.
B.Enable Automatic Mixed Precision (AMP).
C.Disable all CUDA kernels.
D.Replace the GPU with a higher-clocked model.
AnswerB

Mixed precision reduces memory usage by using FP16 or BF16 for most operations, which requires less memory than FP32. This approach allows the training process to utilize significantly less VRAM, providing the necessary headroom to avoid the current OOM error while keeping the model training pipeline functional.

Why this answer

Switching to Mixed Precision (FP16/BF16) effectively halves the memory requirement for model activations and gradients. This is a standard optimization strategy for Transformer models when GPU memory capacity is the primary constraint. By reducing the memory footprint of floating-point operations, researchers can fit larger batch sizes or deeper architectures into the same GPU memory, significantly improving training efficiency and throughput without sacrificing convergence stability.

Exam trap

Candidates often choose complex model architecture changes like gradient checkpointing or model parallelism, which are harder to implement than enabling AMP for immediate memory relief.

20
MCQhard

An administrator notices that a specific containerized training job reports high 'GPU Duty Cycle' but low 'Memory Bandwidth Utilization'. What does this pattern indicate about the workload?

A.The model is suffering from excessive CPU-to-GPU data transfer overhead.
B.The workload is compute-bound, performing heavy arithmetic operations.
C.The system is experiencing PCIe lane bandwidth saturation.
D.The batch size is too small to saturate the GPU compute units.
AnswerB

When the GPU is constantly busy (high duty cycle) but not demanding high amounts of data from VRAM (low bandwidth), it indicates that the kernels are performing a large number of computations relative to the amount of data read, characterizing a compute-bound operation.

Why this answer

High duty cycle combined with low memory bandwidth suggests the model is compute-bound rather than memory-bound. This usually occurs with models that have very high arithmetic intensity, such as small models with many layers. Recognizing this helps in selecting the appropriate hardware, such as focusing on TFLOPS capability rather than HBM bandwidth, to optimize the training speed of the specific neural network architecture.

Exam trap

Candidates often confuse low memory bandwidth with memory-bound workloads, wrongly assuming the GPU lacks sufficient VRAM capacity, whereas it actually indicates that the processor is saturated with intense arithmetic computations.

21
Multi-Selecthard

An AI operations engineer is tuning a real-time inference service on NVIDIA A100 GPUs. Profiling with Nsight Systems shows that the GPU is idle for long periods while waiting for input data, and that host-to-device memory copies are frequent and small. The service uses a fixed batch size of 1 and a custom data loader. Which two changes are most likely to improve GPU utilization and reduce inference latency? (Choose two.)

Select 2 answers
A.Increase the number of CUDA streams used for memory copies and inference kernels.
B.Increase the inference batch size and implement dynamic batching to group requests.
C.Use `cudaMemcpy` instead of `cudaMemcpyAsync` for all host-to-device transfers.
D.Pin the inference process to a single CPU core to reduce context switching.
E.Enable NVIDIA TensorRT with FP16 precision and optimize the model for the target GPU.
AnswersB, E

Larger batches amortize kernel launch and memory copy overhead across more work, keeping the GPU busier. Dynamic batching groups multiple incoming requests into a single inference pass, improving throughput and reducing per-request latency when the service is under load. This directly addresses the idle GPU periods caused by small, frequent transfers and fixed batch size of one.

Why this answer

The GPU idles because it waits on small, frequent data transfers and processes one sample at a time. Increasing batch size with dynamic batching and optimizing the model with TensorRT in FP16 reduce per-inference overhead and better utilize Tensor Cores. Together they increase work per transfer and speed up computation, directly improving utilization and latency.

Exam trap

The trap here is focusing on CPU-side or stream-level tweaks while overlooking that the core inefficiency is the tiny batch size and unoptimized model execution.

22
MCQmedium

An AI engineer observes that a model training job on an NVIDIA DGX system is underutilizing the GPU. The monitoring logs show high CPU wait times and low GPU duty cycles. Which action should the engineer take first to resolve the bottleneck?

A.Increase the GPU clock frequency via NVML.
B.Enable GPUDirect Storage on the local drive.
C.Optimize the data preprocessing pipeline using NVIDIA DALI.
D.Increase the batch size to maximize memory usage.
AnswerC

NVIDIA DALI offloads data augmentation and preprocessing tasks from the CPU to the GPU. By moving these compute-intensive preprocessing steps onto the hardware acceleration engines, the CPU is relieved of the burden, allowing it to prepare batches faster and keeping the GPU fully saturated with training data.

Why this answer

High CPU wait times alongside low GPU utilization suggest an I/O or data preprocessing bottleneck. The CPU cannot feed data to the GPU fast enough, forcing the GPU to idle while waiting for the next batch. Optimizing the data pipeline, such as increasing prefetch buffers or using NVIDIA DALI to move image processing to the GPU, directly addresses the starvation of the compute resources, ensuring efficient utilization.

Exam trap

Candidates often try to optimize the GPU architecture or hyperparameters, missing the fact that the CPU data preprocessing pipeline is the actual bottleneck.

23
MCQhard

Which of the following is the primary indicator of PCIe bus saturation when profiling a training job on an NVIDIA DGX system?

A.High SM occupancy but low memory bandwidth usage.
B.Low GPU duty cycle and high PCIe throughput utilization.
C.High GPU temperature and low clock speeds.
D.Frequent ECC errors in the GPU memory logs.
AnswerB

If the GPU duty cycle is low, it means the GPU is waiting for data. If the PCIe throughput utilization is simultaneously high, it indicates that the bus is fully saturated with data transfers, confirming that the GPU is stalled because the data cannot arrive fast enough.

Why this answer

PCIe bus saturation occurs when the GPU's requirement for data exceeds the transfer capacity of the bus. This leads to GPU starvation, where the compute cores idle while waiting for data. Recognizing the specific correlation between low GPU utilization and high PCIe throughput is essential for identifying I/O-bound bottlenecks and optimizing data pipelines for large-scale training jobs.

Exam trap

Candidates often confuse PCIe saturation with general GPU memory bandwidth bottlenecks. They mistakenly assume high GPU utilization is required to identify a bus issue, ignoring that starvation results in low utilization.

24
MCQmedium

Refer to the exhibit. What is the most effective way to resolve this specific throttling condition?

A.Update the GPU firmware to the latest release.
B.Reinstall the NVIDIA display driver.
C.Increase the power limit for the GPU using nvidia-smi.
D.Replace the power distribution unit (PDU) in the rack.
AnswerC

The 'Sw Power Cap' throttle reason explicitly states that the software power limit is currently restricting the GPU performance. Increasing the power limit via the nvidia-smi tool allows the driver to permit higher clock speeds, effectively removing this constraint and enabling the GPU to perform at full capacity.

Why this answer

The exhibit indicates that the GPU is hitting a software-defined power cap. This is a common situation when a node is configured with a restricted power limit to manage rack-level power consumption. Adjusting the power limit using nvidia-smi ensures the GPU can access the necessary power to reach higher clock frequencies, which is essential for maximizing performance in compute-intensive deep learning tasks.

Exam trap

Candidates often assume the issue is a software bug or driver failure. They overlook the possibility that the GPU is simply hitting a configured power limit enforced by the administrator.

25
MCQmedium

A data scientist reports that a PyTorch training job on an NVIDIA V100 GPU is running slower than expected. The job uses a data loader with num_workers=4. Monitoring shows GPU utilization is around 50%, and CPU usage is high. Which action should an AI operations engineer recommend to improve GPU utilization?

A.Pin memory in the DataLoader to speed up host-to-device transfers.
B.Enable mixed precision training to speed up computations.
C.Increase num_workers in the DataLoader to overlap data preprocessing with GPU computation.
D.Increase the batch size to better utilize the GPU.
AnswerC

When CPU usage is high and GPU utilization is low, the data loading pipeline cannot keep up with the GPU. Increasing num_workers allows more parallel data preprocessing processes, reducing the time the GPU waits for data. This directly addresses the bottleneck and can significantly improve GPU utilization, provided CPU resources are sufficient.

Why this answer

The combination of high CPU usage and low GPU utilization points to a data loading bottleneck where the GPU is starved for data. Increasing the number of DataLoader workers enables more parallel data preprocessing, allowing the GPU to be fed continuously. Other options do not address the CPU-bound pipeline and may not improve the situation.

Exam trap

The trap here is assuming that GPU-side optimizations like mixed precision or larger batches will help, when the real issue is insufficient data preprocessing throughput.

26
Multi-Selecthard

An AI operations team is troubleshooting a distributed training job on a cluster of NVIDIA DGX A100 systems connected via InfiniBand. The job runs but achieves only 40% of expected scaling efficiency. The team suspects communication bottlenecks. Which two actions should they take to confirm and address the issue? (Choose two.)

Select 2 answers
A.Verify that all nodes have identical NCCL versions and environment variables, and that InfiniBand fabric manager is running.
B.Run NCCL_DEBUG=INFO and inspect the logs for warnings about falling back to slower transports or retries.
C.Switch the job to use Ethernet instead of InfiniBand to simplify troubleshooting.
D.Use nvidia-smi to monitor GPU utilization and memory bandwidth on each node during training.
E.Increase the global batch size proportionally to the number of GPUs to improve compute-to-communication ratio.
AnswersA, B

Inconsistent NCCL versions or missing environment variables (e.g., NCCL_IB_HCA, NCCL_SOCKET_IFNAME) can cause suboptimal transport selection or failures. The InfiniBand fabric manager ensures proper link configuration. Ensuring uniformity across nodes is crucial for optimal collective communication performance and is a common fix for scaling inefficiencies.

Why this answer

Scaling inefficiency in distributed training often stems from communication overhead. NCCL debug logs directly show transport selection and errors, while ensuring consistent NCCL versions and InfiniBand fabric health addresses common misconfigurations. Together, these actions confirm and resolve bottlenecks.

The other options either do not diagnose the network or would degrade performance.

Exam trap

The trap here is focusing on GPU-level metrics or hyperparameter changes instead of directly inspecting NCCL communication behavior and InfiniBand fabric health.

27
MCQmedium

Refer to the exhibit. The training job fails with an OOM error. Which optimization strategy will most effectively resolve this while maintaining model convergence?

A.Increase the learning rate to compensate for smaller batches.
B.Implement Gradient Accumulation to simulate a larger batch size.
C.Disable mixed-precision training to reduce memory overhead.
D.Clear the GPU cache using torch.cuda.empty_cache() after every iteration.
AnswerB

Gradient accumulation allows you to simulate a large batch size by breaking it into smaller chunks that fit in GPU memory. You perform multiple forward/backward passes and accumulate the gradients, updating weights only after reaching the target batch size, effectively bypassing the physical memory limitation.

Why this answer

Memory management is central to deep learning stability. When a model exceeds physical VRAM, gradient accumulation allows for larger effective batch sizes without increasing memory footprint. By accumulating gradients over multiple small steps and performing a single weight update, the model effectively sees a larger batch size, maintaining the convergence characteristics of the original design while fitting within the strict physical memory constraints of the GPU hardware.

Exam trap

Candidates frequently try to resolve OOM errors by simply reducing the batch size without realizing it can negatively impact model convergence and training accuracy.

28
Multi-Selectmedium

Which TWO of the following actions are recommended for optimizing NVIDIA GPU utilization during a high-concurrency inference deployment?

Select 2 answers
A.Disable ECC memory on all GPUs to increase raw clock speed.
B.Implement NVIDIA Multi-Process Service (MPS) to allow concurrent kernel execution.
C.Convert models to TensorRT format to leverage layer fusion and precision tuning.
D.Switch the GPU power mode to maximum performance via NVML.
E.Increase the batch size to the maximum allowed by GPU memory.
AnswersB, C

MPS allows multiple processes to share GPU resources more effectively by enabling concurrent kernel execution. This is particularly useful for small inference models where a single process cannot saturate the GPU, leading to higher overall utilization and better performance during high-concurrency periods.

Why this answer

Optimizing inference requires maximizing throughput while maintaining low latency. Using TensorRT for model compilation and Multi-Process Service (MPS) for resource sharing are standard industry practices. These methods ensure that compute resources are efficiently allocated across multiple streams, preventing underutilization of the GPU's tensor cores.

Mastering these tools is essential for AI Operations engineers managing production-grade inference pipelines that must scale effectively under varying loads.

Exam trap

Candidates often suggest increasing the batch size or adding more GPUs. These do not address the efficiency of the individual GPU's execution streams or the model's runtime format optimization.

29
Multi-Selecthard

An AI operations engineer is investigating a training job on an NVIDIA DGX system that intermittently fails with 'uncorrectable ECC error' on a GPU. The job is using NCCL for multi-GPU communication. The engineer needs to identify the appropriate immediate actions to diagnose and mitigate the issue. (Choose two.)

Select 2 answers
A.Increase the NCCL timeout to allow the job to recover from the ECC error.
B.Drain the affected GPU from the scheduler to prevent new jobs from being assigned.
C.Restart the training job immediately to see if the error recurs.
D.Use nvidia-smi --gpu-reset to reset the GPU and clear the error state.
E.Run nvidia-smi -q to check the ECC error counts and identify the affected GPU.
AnswersB, E

Draining the GPU prevents additional jobs from being scheduled on a potentially failing device, reducing the risk of further job failures and data corruption. This is a standard operational mitigation for hardware errors. It allows the administrator to perform maintenance or replacement without impacting other workloads, and it is an appropriate immediate action after identifying the affected GPU.

Why this answer

The immediate actions should include checking ECC error counts with nvidia-smi -q to identify the affected GPU, and draining that GPU from the scheduler to prevent further job failures. These steps diagnose the issue and mitigate risk while preserving the ability to perform maintenance. Restarting, resetting, or adjusting NCCL timeouts do not address the underlying hardware error and may delay proper resolution.

Exam trap

The trap here is treating an uncorrectable ECC error as a software or communication issue and attempting to fix it with job restarts or NCCL tuning.

30
MCQhard

A team runs multi-node training with NCCL over InfiniBand on a cluster of DGX systems. Jobs scale well to four nodes but throughput drops sharply at eight nodes, and `nccl-tests` all-reduce bandwidth falls well below line rate at that size. The fabric uses a fat-tree topology with adaptive routing enabled. Which investigation is most likely to reveal the cause?

A.Check whether NCCL is selecting the correct InfiniBand HCAs and whether the ring or tree algorithm choice matches the fabric's oversubscription ratio.
B.Move the collective operations from NCCL to a TCP socket backend over the management network.
C.Increase the batch size per GPU so that communication is amortized over more compute.
D.Disable adaptive routing on the InfiniBand fabric to force deterministic paths.
AnswerA

Beyond a certain node count, NCCL's default topology detection can pick a ring that traverses oversubscribed uplinks, or it can select the wrong HCA when multiple adapters exist per node. Verifying `NCCL_IB_HCA` and forcing the appropriate algorithm or topology file aligns communication with the physical fabric. This directly explains why bandwidth collapses only at larger scale while small jobs look healthy.

Why this answer

NCCL chooses rings and trees based on detected topology, and on multi-node jobs this can cross oversubscribed uplinks or use a suboptimal adapter when several are present. When scaling stops at a specific node count, the cause is usually that the communication pattern no longer matches the fabric's capacity. Confirming HCA selection and steering the algorithm or topology to respect the oversubscription ratio restores bandwidth and explains why smaller jobs appeared fine.

Exam trap

The trap here is assuming the interconnect is broken or that adaptive routing is at fault, when the collapse at a specific node count usually reflects NCCL's topology and algorithm choices crossing oversubscribed links.

31
MCQmedium

An AI engineer needs to monitor GPU utilization across a large cluster of nodes in real-time. Which NVIDIA tool is the most appropriate for this high-level observability task?

A.nvidia-smi
B.NVIDIA DCGM (Data Center GPU Manager)
C.CUDA Profiler (nsys)
D.nvcc
AnswerB

DCGM is the enterprise-grade tool for managing and monitoring NVIDIA GPU clusters. It provides the necessary APIs to export metrics to monitoring systems like Prometheus and Grafana, allowing for real-time observability across large fleets of GPUs, which is critical for maintaining high availability in production AI environments.

Why this answer

Observability at scale requires tools that can aggregate metrics across multiple nodes. NVIDIA DCGM (Data Center GPU Manager) is specifically architected for this purpose, providing a comprehensive API and service to monitor GPU health, utilization, and power consumption across entire clusters. This is essential for operations teams managing large-scale AI infrastructure to proactively detect performance issues and optimize resource utilization across the fleet.

Exam trap

Candidates often confuse local single-GPU utilities like nvidia-smi with cluster-wide observability tools required for managing multi-node infrastructures.

32
MCQeasy

An AI operations engineer notices that a real-time inference service on an NVIDIA T4 GPU has highly variable latency, with occasional spikes to over 100 ms. The service uses TensorRT and runs in a Docker container. Which action should the engineer take to reduce latency variability?

A.Enable dynamic batching in the inference server configuration.
B.Increase the number of CPU cores allocated to the container.
C.Set the GPU to persistence mode using nvidia-smi -pm 1.
D.Lock the GPU clocks to a fixed frequency using nvidia-smi -lgc.
AnswerD

Locking GPU clocks prevents the GPU from dynamically adjusting frequencies based on power and thermal headroom, which can cause latency variability. A fixed clock ensures consistent performance, reducing spikes. This is especially effective for latency-sensitive inference where predictable execution time is critical. It trades off some power efficiency for stability.

Why this answer

Latency variability in GPU inference often results from dynamic clock adjustments as the GPU balances power and thermal constraints. Locking clocks to a stable frequency eliminates this variability, providing consistent execution times. Other options either add latency (dynamic batching), address different issues (persistence mode), or do not target GPU clock behavior.

Exam trap

The trap here is assuming that persistence mode or CPU allocation affects runtime latency, when the primary cause of variability is often GPU clock throttling.

33
MCQeasy

A data scientist reports that a Jupyter notebook running on a NVIDIA GPU server is extremely slow when executing a deep learning model, even though `nvidia-smi` shows the GPU is idle. The notebook uses TensorFlow. Which action should be taken first to diagnose the issue?

A.Increase the notebook's memory allocation by setting `TF_GPU_ALLOCATOR=cuda_malloc_async`.
B.Check that TensorFlow is configured to use the GPU by running `tf.config.list_physical_devices('GPU')`.
C.Restart the Jupyter kernel and clear the GPU memory with `nvidia-smi --gpu-reset`.
D.Run `nvidia-smi -q -d PERFORMANCE` to check if the GPU is throttled due to thermal or power issues.
AnswerB

If the GPU is idle while the notebook runs slowly, TensorFlow may be defaulting to CPU execution. Verifying that TensorFlow detects the GPU is the first diagnostic step. The function `tf.config.list_physical_devices('GPU')` returns available GPUs; if empty, TensorFlow isn't using the GPU. This directly addresses the symptom of an idle GPU during a supposedly GPU-accelerated workload. It is a simple, non-invasive check that can quickly identify misconfiguration.

Why this answer

An idle GPU during a slow TensorFlow workload strongly suggests that TensorFlow is not using the GPU. The quickest way to confirm is to check if TensorFlow detects the GPU via `tf.config.list_physical_devices('GPU')`. If the list is empty, the issue is likely missing CUDA/cuDNN libraries or a CPU-only TensorFlow installation.

This diagnostic step is non-disruptive and directly targets the root cause.

Exam trap

The trap here is assuming that a slow notebook with an idle GPU indicates a hardware problem, when it often means the framework isn't configured to use the GPU.

34
MCQmedium

An operations engineer is troubleshooting a distributed training job that uses NVIDIA Magnum IO GPUDirect Storage to read training data directly from a local NVMe SSD into GPU memory. The job reports lower than expected I/O bandwidth. `nvidia-smi` shows normal GPU utilization, and the NVMe drive's throughput is well below its peak. Which factor is most likely limiting GPUDirect Storage performance in this scenario?

A.The training job is using too many CPU threads for data preprocessing.
B.The GPU's ECC memory is enabled, reducing available bandwidth.
C.The GPU is not connected to the NVMe drive through a supported PCIe topology.
D.The NVMe drive is formatted with a file system that does not support GPUDirect Storage.
AnswerC

GPUDirect Storage requires a direct data path between the storage device and the GPU, typically over PCIe with peer-to-peer support. If the NVMe drive and GPU are behind different PCIe switches or root complexes without proper peer-to-peer capabilities, data must bounce through host memory, reducing bandwidth. This is a common limitation in multi-socket servers where the GPU and NVMe are on different NUMA nodes.

Why this answer

GPUDirect Storage achieves maximum bandwidth only when the GPU and NVMe drive can communicate directly over PCIe with peer-to-peer support. In many servers, the GPU and drive are on different PCIe root complexes or NUMA nodes, forcing data through host memory and halving effective bandwidth. Checking the PCIe topology with tools like `nvidia-smi topo -m` or `lspci` reveals whether a direct path exists.

Other factors like CPU threads or file system are less likely to cause the specific bandwidth shortfall.

Exam trap

The trap here is assuming that any NVMe drive can achieve full GPUDirect Storage bandwidth regardless of PCIe topology, when peer-to-peer support and NUMA locality are critical.

35
Multi-Selecthard

A team is diagnosing a training job that intermittently stalls for several seconds at the start of each epoch. The job uses a distributed data loader and an NVIDIA DGX system with local NVMe. Monitoring shows GPU utilization dropping to near zero during the stalls while host CPU utilization spikes. Which two actions should the AI operations engineer take to identify and mitigate the stall? (Choose two.)

Select 2 answers
A.Increase the batch size until the GPU memory is nearly full to amortize the data loading cost.
B.Use `nvidia-smi dmon` to capture per-second GPU utilization and power draw while the job runs, correlating the drop with the stall window.
C.Set `CUDA_LAUNCH_BLOCKING=1` to serialize kernel launches and make the stall easier to observe.
D.Enable the PyTorch data loader with `num_workers` greater than zero and `pin_memory=True` so that batches are prefetched and staged in page-locked host memory.
E.Move the dataset from local NVMe to a networked file system to rule out local disk contention.
AnswersB, D

`nvidia-smi dmon` provides a lightweight time-series view of utilization, memory, and power. Correlating the utilization dip with the stall timestamp confirms whether the GPU is idle waiting for input rather than compute-bound. This evidence distinguishes data pipeline starvation from other causes such as thermal or power throttling.

Why this answer

The signature of host CPU spikes with idle GPUs is input pipeline starvation. Prefetching with multiple worker processes and pinned memory keeps the device fed, removing the serialized load at epoch boundaries. Concurrently, `nvidia-smi dmon` supplies the time-series evidence that ties the utilization dips to the stall windows, confirming the diagnosis before and after the fix.

Exam trap

The trap here is treating the stall as a GPU or storage throughput problem and attempting to mask it with larger batches instead of addressing the serialized data loader.

36
MCQmedium

An AI operations engineer is optimizing a real-time inference pipeline on an NVIDIA T4 GPU. The pipeline uses TensorRT and receives requests with variable input sizes. Profiling shows that the engine recompiles for each new input shape, causing latency spikes. Which optimization should the engineer apply to eliminate recompilation while maintaining acceptable accuracy?

A.Enable dynamic shaping in TensorRT by defining an optimization profile with min, opt, and max shapes.
B.Pad all input tensors to a fixed maximum size before inference.
C.Convert the model to use INT8 precision with a calibration dataset to reduce inference time.
D.Increase the workspace size allocated for TensorRT to allow more kernel autotuning.
AnswerA

TensorRT dynamic shaping allows a single engine to handle a range of input dimensions without recompiling. By specifying an optimization profile with min, opt, and max shapes, the engine is built once and can process any input within that range, eliminating recompilation latency spikes. This is the correct approach for variable input sizes while balancing performance and accuracy.

Why this answer

TensorRT dynamic shaping with an optimization profile allows one engine to handle a range of input shapes, eliminating the need to rebuild the engine for each new size. This directly removes the latency spikes from recompilation while maintaining performance across variable inputs. Other options either do not address recompilation or trade off too much efficiency or accuracy.

Exam trap

The trap here is thinking that precision reduction or workspace tuning solves shape variability, when the root cause is engine recompilation for new shapes.

37
MCQmedium

During deployment, an AI model experiences high variance in latency during inference. The system uses a fixed instance count. What is the most likely cause for this performance jitter?

A.The GPU memory is over-provisioned.
B.OS context switching and interrupt handling.
C.The model is too large for the GPU cache.
D.The network switch is experiencing congestion.
AnswerB

Latency variance is frequently caused by the OS scheduling other tasks or handling interrupts on the same cores used by the inference server. This creates 'jitter' as the processor pauses inference tasks to handle housekeeping. Pinning the inference server to dedicated, isolated cores effectively minimizes this contention and stabilizes latency.

Why this answer

Inference latency jitter is often caused by external processes or system interrupts competing for CPU resources that the inference server relies on for data pre-processing or orchestration. By pinning CPU cores to the inference process (CPU affinity) and using isolated cores, the administrator can prevent these context switches and OS interrupts, leading to more consistent and predictable inference response times for production workloads.

Exam trap

Candidates often misattribute inference latency jitter to model complexity or network overhead, ignoring local OS operations like context switching and CPU interrupt handling that disrupt execution consistency.

38
MCQhard

Refer to the exhibit. An engineer is troubleshooting inconsistent training performance across two GPUs in a single DGX node. Why is one GPU reporting a lower clock speed despite being in P0 state?

A.The GPU is misconfigured in the system BIOS.
B.The GPU is experiencing thermal throttling due to improper airflow or fan failure.
C.The CUDA driver is applying a different profile to each GPU.
D.The GPU is running an outdated firmware version.
AnswerB

When a GPU reaches its thermal limit, the firmware automatically reduces the graphics clock to prevent physical damage. A discrepancy in clock speeds between two identical GPUs in the same P0 state indicates that one card is running hotter than the other due to cooling issues.

Why this answer

The P0 power state represents maximum performance, but the actual clock frequency is governed by the hardware's thermal and power monitoring sub-systems. When two identical GPUs show different clock speeds in the same state, it suggests external environmental factors or hardware degradation. This distinction is critical for AI operations to maintain balanced parallel training workloads and ensure distributed training does not bottleneck on the slowest card.

Exam trap

Candidates often assume the GPU is faulty and needs replacement. They fail to consider environmental factors like airflow or fan failure, which trigger thermal throttling even on healthy hardware.

39
MCQeasy

A monitoring system reports that a DGX node's GPUs are running at reduced clocks during a long training job, and `nvidia-smi -q -d PERFORMANCE` shows the throttle reason as 'SW Power Cap'. The job's power draw is at the configured limit. Which action is MOST appropriate to restore higher clocks?

A.Reduce the batch size so the GPU consumes less power and stays below the cap.
B.Raise the GPU power limit with `nvidia-smi -pl` to a value within the GPU's supported maximum and re-run the job.
C.Enable persistence mode with `nvidia-smi -pm 1` to keep the driver loaded and improve clock stability.
D.Lock the GPU clocks to their maximum with `nvidia-smi -lgc` to override the throttle.
AnswerB

A 'SW Power Cap' throttle reason means the GPU is hitting its configured power limit, so clocks are reduced to stay within budget. If the hardware and cooling support it, raising the power limit to a validated value within the GPU's supported range allows higher sustained clocks, directly addressing the throttle cause. This is the correct, targeted remediation for the reported condition.

Why this answer

The throttle reason 'SW Power Cap' indicates the GPU is constrained by its configured power limit. Raising the power limit to a supported value within the GPU's validated maximum allows the device to sustain higher clocks, provided cooling and power delivery can support it. This directly removes the cause of the reduced clocks rather than working around it.

Exam trap

The trap here is confusing a power-cap throttle with a thermal or clock-lock issue and attempting to override clocks instead of adjusting the power budget.

40
MCQhard

A production inference service on NVIDIA GPUs reports that p99 latency spikes every few minutes while p50 remains stable. Metrics show GPU memory utilization near the limit and periodic `cudaMalloc` calls in the application logs. The model and batch size are fixed. Which change is MOST likely to eliminate the latency spikes?

A.Enable CUDA lazy module loading to defer kernel module initialization until first use.
B.Increase the number of model instances per GPU to smooth out request distribution.
C.Lower the GPU clock to reduce power draw and stabilize thermals during inference.
D.Pre-allocate the inference workspace and reuse CUDA memory pools instead of calling cudaMalloc during request handling.
AnswerD

Periodic `cudaMalloc` calls during request handling force synchronous memory allocation and can trigger driver-level operations that stall the GPU pipeline, producing p99 spikes while p50 stays flat. Pre-allocating workspaces and using a memory pool (or CUDA graphs with static allocations) removes allocation from the hot path, which directly addresses the observed pattern and is the correct remediation here.

Why this answer

The correlation between p99 spikes and periodic `cudaMalloc` calls, combined with stable p50, points to allocation occurring inside the request path. Synchronous allocation can serialize GPU work and introduce jitter. Pre-allocating workspaces and using a memory pool or CUDA graphs with static buffers removes the allocator from the hot path, eliminating the periodic stalls while leaving steady-state latency unchanged.

Exam trap

The trap here is attributing p99 spikes to concurrency or thermals, when the log evidence ties them to allocator calls in the request path.

41
MCQmedium

An AI operations engineer is troubleshooting a multi-node NCCL training job on an NVIDIA DGX SuperPOD. The job runs but scales poorly: inter-node bandwidth is roughly half of the expected 200 Gb/s per GPU, while intra-node NVLink traffic is at full rate. Running `nvidia-smi topo -m` shows that GPUs in each node are connected to the NICs through the PCIe switch, but the job sets `NCCL_SOCKET_IFNAME` to the management interface. Which action is the most appropriate to resolve the bottleneck?

A.Enable GPUDirect RDMA by allowing the container access to the host IB verbs devices and the `/dev/infiniband` tree, and set `NCCL_IB_HCA` to the compute fabric adapters.
B.Increase `NCCL_BUFFSIZE` to 16 MB and set `NCCL_NTHREADS` to 8 to raise channel throughput.
C.Set `NCCL_P2P_DISABLE=1` so that all inter-node traffic is routed through the CPUs, avoiding PCIe switch contention.
D.Bind the job to a single NUMA node using `numactl --cpunodebind=0 --membind=0` to reduce cross-socket traffic.
AnswerA

The symptom of full intra-node NVLink but degraded inter-node throughput points to the job falling back to TCP over the management interface. GPUDirect RDMA lets the NIC read/write GPU memory directly over the InfiniBand/RoCE compute fabric, bypassing host memory copies. Exposing `/dev/infiniband` to the container and pinning `NCCL_IB_HCA` restores the intended high-speed path.

Why this answer

When intra-node NVLink performs at full rate but inter-node bandwidth is roughly half, the job is usually not using the compute fabric. GPUDirect RDMA bypasses host memory copies by letting the HCA access GPU memory directly, and `NCCL_IB_HCA` ensures NCCL selects the correct adapters. Exposing `/dev/infiniband` to the container is required for RDMA verbs.

This restores the expected inter-node bandwidth.

Exam trap

The trap here is assuming that NCCL tuning parameters such as buffer size or thread count can compensate for a transport that is falling back to TCP over the management network.

42
MCQmedium

An AI engineer observes that a training job on an NVIDIA DGX H100 system is experiencing significant performance degradation. The GPU utilization is high, but the throughput remains low. Which tool should be used first to identify if the bottleneck is related to data loading or PCIe bandwidth saturation?

A.NVIDIA-SMI
B.NVIDIA DCGM
C.NVIDIA Nsight Systems
D.NVIDIA Nsight Compute
AnswerC

Nsight Systems provides a system-wide view of CPU and GPU activities, allowing engineers to correlate data transfer operations with kernel execution. This visibility is essential for identifying stalls, data starvation, or PCIe bus congestion, which are the most common causes of low throughput despite high GPU utilization.

Why this answer

NVIDIA Nsight Systems is the primary tool for analyzing system-wide performance, including CPU-GPU interactions and data transfer bottlenecks. Identifying whether a bottleneck occurs in the data pipeline or within the hardware interconnects is critical for optimizing training speed. By visualizing timelines, engineers can pinpoint if the GPU is starving for data or if the PCIe bus is congested, enabling targeted remediation for large-scale distributed training clusters.

Exam trap

Candidates often select NVIDIA-SMI or Nsight Compute, failing to realize that Nsight Systems is the correct tool for identifying system-wide bottlenecks between the CPU, data pipeline, and GPU.

43
MCQeasy

A data scientist reports that a Jupyter notebook running on a DGX station cannot allocate GPU memory, even though other users' jobs are running fine. The notebook kernel was started before a system administrator updated the NVIDIA driver and rebooted the node. Which action should the data scientist take to resolve the issue?

A.Run `nvidia-smi --gpu-reset` to reset the GPU and clear any stale contexts.
B.Set the environment variable `CUDA_VISIBLE_DEVICES=0` to force the notebook to use a specific GPU.
C.Restart the Jupyter kernel to reinitialize CUDA and pick up the new driver.
D.Reinstall the NVIDIA driver using the `.run` installer without rebooting.
AnswerC

When the NVIDIA driver is updated and the system reboots, any existing processes that had already initialized CUDA hold references to the old driver. The Jupyter kernel is such a process. Restarting the kernel terminates the old process and starts a new one that loads the updated driver, allowing CUDA calls to succeed and GPU memory to be allocated.

Why this answer

After a driver update and reboot, any process that started before the update holds an outdated CUDA context. The Jupyter kernel is such a process, so it cannot use the new driver. Restarting the kernel creates a fresh process that loads the updated driver and can allocate GPU memory normally, resolving the issue without affecting other users.

Exam trap

The trap here is assuming that a GPU reset or environment variable change is needed, when simply restarting the long-running process that predates the driver update is sufficient.

44
MCQmedium

An operations team is troubleshooting a multi-node NCCL training job on an NVIDIA DGX SuperPOD that stalls at initialization and never begins gradient exchange. Running `nccl-tests` with `all_reduce_perf` on the same nodes fails identically, but single-node `all_reduce_perf` succeeds. Which action should the team take first to isolate the fault?

A.Set NCCL_DEBUG=INFO and NCCL_DEBUG_SUBSYS=INIT,COLL,NET on the launcher and inspect which transport (NET/IB or P2P) is selected and where the hang occurs.
B.Disable GPUDirect RDMA by setting NCCL_NET_GDR_LEVEL=0 across all nodes to force host-memory staging.
C.Increase NCCL_BUFFSIZE to 16 MB on every rank and rerun the job to rule out buffer exhaustion.
D.Reduce the number of ranks per node so each process has a dedicated GPU, then rerun the multi-node test.
AnswerA

NCCL_DEBUG=INFO with NCCL_DEBUG_SUBSYS=INIT,COLL,NET surfaces transport selection, ring/tree topology construction, and per-rank connection progress. Because single-node all_reduce_perf passes but multi-node fails, the fault is almost certainly in the inter-node path, and this logging pinpoints whether NCCL fell back to a broken interface or is stuck establishing connections.

Why this answer

Because single-node all_reduce_perf succeeds while multi-node fails, the failure is in inter-node communication setup rather than GPU-local collectives. Enabling NCCL_DEBUG=INFO with INIT, COLL, and NET subsystems exposes which transport NCCL selects, how rings and trees are built, and where connection establishment stalls, giving the team actionable evidence before any configuration change.

Exam trap

The trap here is assuming a collective hang is a buffer-size or GPU-memory problem, when a job that never starts exchanging gradients is actually stuck in NCCL transport and topology initialization.

45
MCQhard

An AI operations engineer is troubleshooting a multi-GPU training job that fails intermittently with a NCCL error: 'unhandled system error'. The job runs on a DGX-1 with eight V100 GPUs connected via NVLink. Which step should the engineer take first to resolve the issue?

A.Disable NVLink by setting NCCL_P2P_DISABLE=1 to force PCIe communication.
B.Reinstall the NVIDIA driver and CUDA toolkit on all nodes.
C.Reduce the number of GPUs used in the job to four to see if the error disappears.
D.Set NCCL_DEBUG=INFO and reproduce the failure to capture detailed logs.
AnswerD

NCCL debug logs provide detailed information about the communication setup, including which transports are used, any fallbacks, and specific error codes. This is the most direct way to diagnose the cause of an 'unhandled system error', which can stem from hardware, driver, or configuration issues. Capturing logs during failure is the essential first step before attempting fixes.

Why this answer

The first step in troubleshooting any NCCL error should be to gather detailed logs using NCCL_DEBUG=INFO. This provides visibility into the communication path and error specifics, enabling targeted fixes. Other options either mask the issue, involve unnecessary reinstallation, or reduce resources without diagnosis.

Exam trap

The trap here is jumping to hardware workarounds or reinstallations without first collecting diagnostic information that NCCL can provide.

46
MCQhard

An operations engineer is investigating a sudden drop in throughput for a multi-GPU training job on an NVIDIA DGX A100. The job uses PyTorch with DDP. Logs show that one GPU is consistently at 100% utilization while others are below 50%. Which tool and approach should be used to identify the bottleneck?

A.Use nvidia-smi to monitor GPU utilization and memory, then rebalance the workload by reducing the batch size on the overutilized GPU.
B.Use NVIDIA Nsight Systems to profile the job and examine the CUDA API and kernel timeline to identify serialization or synchronization issues.
C.Use NVIDIA Nsight Compute to profile each GPU's kernels and compare execution times to find the slowest kernel.
D.Use nvidia-smi topo -m to check the GPU topology and then adjust the NCCL communication settings to use a different ring order.
AnswerB

Nsight Systems provides a detailed timeline of CPU and GPU activity, showing kernel executions, memory copies, and synchronization points. It can reveal if one GPU is waiting on data or if there is an imbalance in computation. This is the correct tool to pinpoint the bottleneck in a multi-GPU job, as it captures cross-GPU interactions and can highlight stragglers.

Why this answer

Nsight Systems is designed for system-level profiling, capturing CPU and GPU activities across multiple devices. It can show if one GPU is performing more work or if others are idle waiting for synchronization. This makes it the right tool to identify the bottleneck in a DDP job where one GPU is overutilized.

Profiling before making changes is essential.

Exam trap

The trap here is choosing a profiling tool that is too granular, like Nsight Compute, or a monitoring tool that lacks detail, like nvidia-smi, instead of the system-level profiler needed for this multi-GPU imbalance.

47
MCQeasy

An AI operations team is troubleshooting a training job that crashes with a segmentation fault after several hours. The job uses multiple GPUs and NCCL for communication. System logs show no errors, but dmesg reveals repeated 'NVRM: Xid' errors. Which action should be taken first to diagnose the issue?

A.Restart the NVIDIA driver with rmmod and modprobe commands.
B.Check the Xid error code in the NVIDIA documentation to identify the specific GPU fault.
C.Enable NCCL_DEBUG=INFO and re-run the job to capture detailed NCCL logs.
D.Run nvidia-smi -q to check GPU temperature and power usage.
AnswerB

Xid errors are reported by the NVIDIA driver and indicate GPU hardware or driver issues. Each Xid code corresponds to a specific fault, such as a corrupted memory access or a fallen off the bus error. Looking up the code in NVIDIA's documentation provides the exact cause and recommended actions. This is the most direct first step to diagnose the segmentation fault linked to GPU errors.

Why this answer

Xid errors are logged by the NVIDIA driver and each code maps to a specific GPU fault. Interpreting the Xid code is the fastest way to understand whether the segmentation fault is due to a hardware issue, driver bug, or application error. This guides subsequent troubleshooting steps, such as replacing hardware or updating the driver.

Exam trap

The trap here is jumping to application-level debugging tools like NCCL_DEBUG when the system logs already point to a GPU-level fault via Xid errors.

48
MCQmedium

A machine learning engineer is optimizing a recommendation model for inference on an NVIDIA T4 GPU. The model uses dynamic input shapes, and profiling shows that kernel launch overhead is a significant contributor to latency. Which optimization technique should be applied to reduce this overhead?

A.Convert the model to TensorRT with dynamic shapes and enable CUDA graphs during inference.
B.Use mixed precision (FP16) to reduce the number of kernels executed.
C.Enable NVIDIA MPS (Multi-Process Service) to allow concurrent kernel execution.
D.Increase the batch size to amortize kernel launch overhead across more samples.
AnswerA

CUDA graphs capture a sequence of kernel launches and replay them with a single launch, drastically reducing launch overhead. TensorRT supports dynamic shapes and can build engines that use CUDA graphs. This is ideal for models with variable input shapes where launch overhead is high. The combination reduces CPU overhead and improves latency.

Why this answer

CUDA graphs reduce kernel launch overhead by capturing a sequence of operations and replaying them as a single graph. TensorRT with dynamic shapes can leverage CUDA graphs to maintain flexibility while minimizing overhead. This directly addresses the profiling finding, making it the most effective optimization for the described scenario.

Exam trap

The trap here is assuming that mixed precision or larger batch sizes reduce launch overhead, but they target different bottlenecks and do not minimize the number of CPU-GPU interactions.

49
MCQmedium

A production inference service on an NVIDIA A100 GPU experiences a gradual increase in latency over several hours, eventually requiring a pod restart. GPU memory utilization climbs steadily, but the model and batch size are fixed. Which action should an AI operations engineer take first to diagnose the root cause?

A.Use PyTorch's torch.cuda.memory_summary() or TensorFlow's memory profiler to capture allocation snapshots and compare over time.
B.Run nvidia-smi --query-gpu=memory.used --format=csv -l 1 to log memory usage over time and correlate with request rate.
C.Enable CUDA memory leak detection with compute-sanitizer --tool memcheck on the running inference process.
D.Increase the GPU memory limit in the Kubernetes pod spec to prevent the pod from being killed.
AnswerA

Framework-level memory profilers show exactly which tensors or operations are allocating memory and whether those allocations are freed. By taking snapshots at intervals, an engineer can see if memory is retained across inference calls, indicating a leak in the application code or framework caching allocator. This directly identifies the source without disrupting the service.

Why this answer

The steady memory increase with fixed workload strongly suggests a memory leak in the inference application or framework. Framework-specific memory profilers provide allocation-level visibility, allowing engineers to identify unreleased tensors or cached allocations. Aggregate GPU monitoring or sanitizer tools lack the necessary granularity.

Increasing limits merely postpones failure.

Exam trap

The trap here is assuming that nvidia-smi memory monitoring is sufficient to diagnose leaks, when it only shows aggregate usage without per-process or allocation detail.

50
MCQeasy

An AI operations engineer is troubleshooting a model inference service deployed with NVIDIA Triton Inference Server on a GPU. The service occasionally returns incorrect predictions, and the engineer suspects that the input data is not being preprocessed correctly. The model expects input tensors in FP32 format, but the client is sending FP16 data. Which action should the engineer take to resolve the issue?

A.Enable dynamic batching in Triton to automatically convert FP16 inputs to FP32.
B.Set the model to use FP16 precision to match the client's input data type.
C.Increase the instance count to handle the load and reduce the chance of data corruption.
D.Configure the Triton model's input to specify the correct data type (FP32) and ensure the client sends matching data.
AnswerD

Triton validates input data types against the model's configuration. If the model expects FP32 but receives FP16, it may either reject the request or misinterpret the data, leading to incorrect predictions. Setting the correct data type in the model configuration and aligning the client ensures proper preprocessing and accurate inference.

Why this answer

The root cause is a mismatch between the client's data type (FP16) and the model's expected input type (FP32). Triton requires that input tensors match the model's configuration. By setting the correct data type in the model configuration and ensuring the client sends FP32 data, the engineer ensures proper data handling and restores prediction accuracy.

Exam trap

The trap here is assuming that Triton automatically handles data type conversions, when in fact it enforces strict type matching based on the model configuration.

51
MCQmedium

During a training job, the system reports "NCCL WARN" regarding a slow network path. What is the most likely culprit for this performance bottleneck in a multi-node InfiniBand environment?

A.The GPU clock speed is set too low.
B.InfiniBand link speed is negotiated at a lower rate than expected.
C.The batch size is too small.
D.The CPU is running at maximum capacity.
AnswerB

If an InfiniBand link negotiates at a lower rate (e.g., SDR instead of HDR), the network throughput will be severely throttled. This mismatch is a classic cause of "slow path" NCCL warnings, as the collective communication operations are unable to achieve the expected bandwidth required for efficient multi-node training.

Why this answer

In high-performance clusters, the network fabric is often the bottleneck. InfiniBand performance relies on proper subnet manager configuration and accurate link-speed reporting. Identifying "slow paths" using tools like ibdiagnet helps pinpoint physical or configuration issues in the interconnect.

Addressing these issues is vital because, in distributed training, the speed of the cluster is limited by the slowest link in the communication path, severely impacting the overall training efficiency.

Exam trap

Candidates often blame the GPU driver or the training code itself, overlooking that InfiniBand interconnects are physical network layers that can negotiate down to lower speeds due to cabling issues.

52
MCQmedium

A distributed training job on a multi-GPU node is exhibiting poor scaling efficiency: each GPU shows high utilization, but overall throughput increases by only 15% when doubling the number of GPUs. The job uses NCCL for communication. Which diagnostic step is most appropriate to identify the bottleneck?

A.Enable NVIDIA MPS (Multi-Process Service) to allow multiple processes to share the GPUs more efficiently.
B.Check the GPU clock speeds with `nvidia-smi -q -d CLOCK` to ensure they are running at maximum frequency.
C.Increase the batch size per GPU to reduce the frequency of gradient synchronization across GPUs.
D.Run `nvidia-smi topo -m` to inspect the GPU interconnect topology and verify that NCCL is using the fastest available path between GPUs.
AnswerD

Poor scaling with high GPU utilization often indicates communication overhead. `nvidia-smi topo -m` reveals whether GPUs are connected via NVLink or PCIe and whether data traverses slower paths like QPI/UPI. If NCCL cannot leverage NVLink, all-reduce operations become the bottleneck. This command is the standard first step to diagnose interconnect issues and ensure NCCL selects the optimal topology-aware route, directly addressing the scaling inefficiency.

Why this answer

Poor scaling efficiency in distributed training often results from communication bottlenecks. The `nvidia-smi topo -m` command provides a matrix of GPU interconnect paths, helping verify whether NCCL can use NVLink or is forced over slower PCIe/QPI links. If the topology shows suboptimal paths, adjusting NCCL environment variables or hardware configuration can improve scaling.

Thus, inspecting the topology is the most direct diagnostic step.

Exam trap

The trap here is assuming that high GPU utilization guarantees optimal scaling, overlooking that inter-GPU communication overhead can dominate even when GPUs appear busy.

53
MCQmedium

An AI engineer is optimizing a real-time inference pipeline on an NVIDIA A100 GPU. The model uses dynamic input shapes, and profiling shows that the GPU spends significant time on memory copies between host and device. Which optimization should be implemented to reduce this overhead?

A.Enable Unified Memory to automatically manage data migration.
B.Increase the batch size to amortize the cost of memory copies.
C.Use pinned (page-locked) host memory for input and output buffers.
D.Use CUDA streams to overlap memory copies with computation.
AnswerC

Pinned memory allows the GPU to directly access host memory via DMA, avoiding the overhead of staging through pageable memory. This reduces the time spent on host-to-device and device-to-host copies, which is critical for real-time inference with dynamic shapes where copies are frequent. It is a standard optimization to improve transfer efficiency and lower latency.

Why this answer

The overhead from host-device memory copies is best reduced by using pinned memory, which enables faster DMA transfers. This is especially important for real-time inference with dynamic shapes where copies are frequent. While CUDA streams can overlap transfers with compute, they are more effective when combined with pinned memory.

Exam trap

The trap here is assuming that Unified Memory or CUDA streams alone will solve the copy overhead, when the fundamental issue is pageable memory causing slow transfers.

54
MCQmedium

An AI operations engineer is investigating an inference service running on NVIDIA Triton Inference Server. Clients report sporadic 500 errors under peak load. The server logs show occasional 'Failed to allocate memory' messages, and `nvidia-smi` shows VRAM nearly full. The service uses dynamic batching with a maximum batch size of 64 and multiple model instances per GPU. Which change is the most appropriate first step to stabilize the service?

A.Reduce the number of model instances per GPU and lower the maximum dynamic batch size to match the available VRAM headroom.
B.Switch the service to use the CPU execution provider for overflow requests when GPU memory is exhausted.
C.Enable `--strict-model-config=false` so Triton can auto-generate the model configuration and allocate memory more efficiently.
D.Increase the GPU memory pool by setting `CUDA_MPS_ACTIVE_THREAD_PERCENTAGE` to 100.
AnswerA

The allocation failures under peak load with near-full VRAM indicate that the configured instances and maximum batch size exceed the memory budget when concurrent requests spike. Reducing instances and capping the batch size lowers the peak memory footprint per inference, restoring headroom so the allocator can satisfy every request instead of failing.

Why this answer

Triton allocates memory per model instance and per active batch execution. When instances and the maximum dynamic batch size are sized for average load, peak concurrency can exceed VRAM and cause allocation failures. Reducing instances and capping batch size brings the peak footprint within budget, letting the scheduler queue requests instead of failing them, which stabilizes error rates under load.

Exam trap

The trap here is assuming that runtime flags or MPS settings can expand available VRAM, when the real issue is a peak memory footprint larger than the device can satisfy.

55
MCQmedium

A production inference service running on NVIDIA GPUs exhibits periodic latency spikes every few minutes, correlating with CPU-side stalls and low GPU utilization during those intervals. Profiling with Nsight Systems shows large gaps between kernel launches and frequent cudaMalloc/cudaFree calls. Which action best addresses the root cause?

A.Switch the model to use TensorRT with FP16 precision to reduce kernel execution time.
B.Enable CUDA Graph capture for the inference sequence and reuse the captured graph for repeated executions.
C.Increase the number of worker threads on the CPU to better overlap preprocessing with GPU execution.
D.Enable Multi-Process Service (MPS) to allow concurrent kernel execution from multiple processes.
AnswerB

CUDA Graphs capture a sequence of kernel launches and memory operations into a single graph, then replay it with minimal CPU overhead. This eliminates repeated cudaMalloc/cudaFree and launch gaps, smoothing latency spikes. It directly targets the CPU-side stalls and low GPU utilization observed, making it the appropriate optimization for this scenario.

Why this answer

The profile shows CPU-side stalls and frequent cudaMalloc/cudaFree, indicating launch and allocation overhead. CUDA Graphs capture the entire sequence and replay it with minimal CPU involvement, eliminating both the allocation churn and the launch gaps. This directly reduces latency spikes and improves GPU utilization, making it the most effective remedy for the described symptoms.

Exam trap

The trap here is assuming that reducing kernel execution time or adding CPU threads will fix latency spikes caused by launch overhead, rather than addressing the orchestration inefficiency itself.

56
MCQhard

A team is profiling a distributed training job using NVIDIA NCCL for inter-GPU communication on a DGX A100 system. They observe that all-reduce operations are taking longer than expected, and the NCCL debug logs show frequent 'NVLS' (NVLink SHARP) errors. Which action should be taken to resolve the issue?

A.Increase the NCCL buffer size by setting NCCL_BUFFSIZE to a larger value to reduce the number of messages.
B.Disable NVLink SHARP by setting NCCL_NVLS_ENABLE=0 in the environment and restart the job.
C.Update the GPU driver and CUDA toolkit to the latest versions to fix NVLS bugs.
D.Switch the NCCL algorithm to 'Tree' by setting NCCL_ALGO=Tree to avoid NVLS usage.
AnswerB

NVLS (NVLink SHARP) is an in-network reduction feature that can accelerate all-reduce operations, but it requires compatible hardware and software. If errors occur, disabling it forces NCCL to fall back to standard ring or tree algorithms, which are reliable. This resolves the immediate errors and restores communication performance, though it may not achieve peak efficiency. It is a valid troubleshooting step when NVLS is unstable.

Why this answer

NVLS errors indicate that the NVLink SHARP acceleration is failing, likely due to hardware or software incompatibility. Disabling NVLS via the NCCL_NVLS_ENABLE environment variable forces NCCL to use standard algorithms, which are robust and should eliminate the errors. Other options either do not target NVLS or are less direct and potentially disruptive.

Exam trap

The trap here is assuming that increasing buffer size or changing algorithms will fix NVLS errors, when the correct approach is to disable the failing feature.

57
Multi-Selecthard

An AI operations engineer is investigating a training job that exhibits poor scaling efficiency when moving from 8 to 32 GPUs on a DGX SuperPOD. Profiling indicates that the communication time in NCCL all-reduce operations is disproportionately high. Which two actions should be taken to improve scaling efficiency? (Choose two.)

Select 2 answers
A.Check that the InfiniBand fabric is not experiencing congestion or errors by monitoring performance counters.
B.Ensure that the job is using the correct NCCL topology by setting NCCL_TOPO_FILE to a custom topology XML.
C.Increase the NCCL buffer size by setting NCCL_BUFFSIZE to a larger value.
D.Enable NCCL tree algorithm for all-reduce by setting NCCL_ALGO=Tree.
E.Verify that NCCL is using the highest-bandwidth network interface, such as InfiniBand, by checking NCCL_IB_HCA.
AnswersA, E

High communication time can be caused by network congestion or errors on the InfiniBand fabric. Monitoring counters such as port errors, link down events, or congestion can reveal underlying issues. If the fabric is congested, all-reduce operations will slow down, impacting scaling. This is a fundamental check before tuning NCCL parameters. It directly addresses potential network-level bottlenecks.

Why this answer

Poor scaling in NCCL all-reduce often stems from network misconfiguration or fabric issues. Verifying that NCCL uses the correct high-bandwidth HCAs and checking the InfiniBand fabric for congestion or errors are essential first steps. These actions ensure that communication occurs over the fastest available path and that the network is healthy, directly improving scaling efficiency.

Exam trap

The trap here is focusing on NCCL algorithm or buffer tuning before confirming that the underlying network hardware and fabric are correctly utilized and healthy.

58
MCQmedium

If a training job on a multi-node cluster shows a significant performance drop during checkpointing, what is the most likely bottleneck?

A.Network latency between compute nodes.
B.Shared filesystem I/O throughput.
C.GPU memory allocation overhead.
D.The model weight update frequency.
AnswerB

Checkpointing requires writing large amounts of data to disk. If multiple nodes attempt to write to a shared filesystem simultaneously, the aggregate I/O demand can exceed the filesystem's bandwidth, causing the training process to hang while waiting for the write operation to complete.

Why this answer

Checkpointing involves writing large model states to non-volatile storage. If this process is not handled asynchronously or if the underlying storage system is undersized, the training process will stall. Identifying this bottleneck allows engineers to implement optimized I/O strategies, such as using distributed file systems or NVMe-based local storage, to minimize the time spent in the checkpoint phase and maximize overall cluster efficiency.

Exam trap

Candidates often assume checkpoint performance drops are caused by insufficient GPU memory or slow CPU speeds, overlooking the massive I/O bottleneck created by writing large state files simultaneously.

59
MCQeasy

Which component of the NVIDIA AI Enterprise stack is primarily responsible for ensuring the long-term stability and compatibility of drivers and libraries across heterogeneous hardware configurations?

A.NVIDIA CUDA Toolkit
B.NVIDIA AI Enterprise
C.NVIDIA Triton Inference Server
D.NVIDIA Collective Communications Library
AnswerB

NVIDIA AI Enterprise is a cloud-native software suite that provides validated, secure, and supported software. It is specifically designed to ensure that the entire stack, from drivers to containers, is compatible and stable, allowing organizations to maintain production consistency across diverse data center hardware environments over long periods.

Why this answer

NVIDIA AI Enterprise provides a validated and supported software stack that ensures consistency across different hardware generations. Stability is the foundation of AI operations; without a supported, validated stack, minor version mismatches between kernels and libraries can lead to silent failures or system crashes. This support model is essential for enterprises to avoid the 'dependency hell' that often plagues research-grade environments when scaling to production operations.

Exam trap

Candidates often name individual components like CUDA or NCCL, ignoring that NVIDIA AI Enterprise is the overarching suite specifically designed for validated support and version compatibility.

60
MCQmedium

An AI researcher is running a training job using mixed precision (FP16/BF16). The loss function is diverging unexpectedly. What is the most likely culprit?

A.The GPU driver is too old for FP16.
B.The learning rate is too high for mixed precision.
C.The lack of loss scaling for FP16.
D.The batch size is too large.
AnswerC

FP16 has a limited dynamic range. During backpropagation, many small gradient values can be rounded to zero (underflow). Loss scaling multiplies the loss before backpropagation, pushing the gradients into a representable range in FP16, then unscaling them before applying weight updates. This is critical for preventing divergence in mixed precision.

Why this answer

Mixed precision training uses lower precision formats to speed up math and reduce memory usage, but this can lead to numerical instability. Divergence in loss is a classic symptom of 'underflow' or 'overflow' issues occurring in FP16, where small gradients or large weight updates exceed the representable range. Implementing loss scaling is the standard industry technique to preserve the precision of gradients and ensure the training process remains stable.

Exam trap

Candidates often assume divergence is caused by a poor learning rate or bad initialization, overlooking the precision limitations inherent in standard FP16 training.

61
MCQhard

An operations team observes that a distributed training job using NVIDIA Collective Communications Library (NCCL) across eight nodes occasionally hangs during the all-reduce phase. Logs show no errors, and the hang resolves only after a node is manually restarted. Which action is MOST appropriate to diagnose the intermittent hang?

A.Disable InfiniBand and fall back to TCP sockets by setting NCCL_IB_DISABLE=1.
B.Increase the NCCL buffer size with NCCL_BUFFSIZE to reduce the number of messages exchanged.
C.Set NCCL_ALGO=RING to force a single algorithm and eliminate variability.
D.Enable NCCL debug logging with NCCL_DEBUG=INFO and NCCL_DEBUG_SUBSYS=COLL, and set a NCCL watchdog timeout to capture the stalled collective.
AnswerD

Intermittent hangs without errors require visibility into which collective and rank stalled. NCCL_DEBUG=INFO with NCCL_DEBUG_SUBSYS=COLL prints per-collective progress and rank participation, while a watchdog timeout can trigger a dump when progress stops. This combination captures the state at the moment of the hang, making it the correct diagnostic action for this scenario.

Why this answer

Intermittent NCCL hangs without errors are best diagnosed by capturing state at the moment of the stall. Enabling NCCL debug logging for the collective subsystem and configuring a watchdog timeout produces logs that show which rank and collective stopped progressing. This evidence is necessary to distinguish a slow rank, a fabric problem, or a software defect from a simple performance issue.

Exam trap

The trap here is trying to work around the hang by changing transport or algorithm rather than capturing diagnostic data at the point of failure.

62
MCQmedium

An AI operations team is using NVIDIA DCGM to monitor a cluster of GPUs. They want to set up alerts for when GPUs are running at high temperatures for extended periods. Which DCGM feature should they use?

A.DCGM's built-in email alerting system configured via dcgm.conf.
B.DCGM diagnostics with the '-r' flag to run a comprehensive test and report temperature violations.
C.NVIDIA-smi's '--query-gpu=temperature.gpu' with a cron job to check and send alerts.
D.DCGM health checks with custom thresholds and alerting via DCGM exporter to Prometheus.
AnswerD

DCGM provides health checks that can be configured with thresholds for temperature and other metrics. The DCGM exporter can expose these metrics to Prometheus, which supports alerting rules based on sustained conditions. This combination allows for proactive monitoring and alerting on high temperatures over time. It is the standard approach for cluster-wide GPU monitoring.

Why this answer

DCGM health checks allow defining thresholds for GPU metrics, including temperature. When combined with the DCGM exporter and Prometheus, teams can create alerting rules that trigger on sustained high temperatures. This is the recommended method for cluster-wide monitoring.

Other options either misuse DCGM diagnostics, assume non-existent features, or use less scalable methods.

Exam trap

The trap here is confusing DCGM diagnostics (point-in-time testing) with continuous monitoring and alerting, which requires integration with a metrics platform.

63
MCQeasy

An AI operations engineer notices that a training job on an NVIDIA A100 GPU is running slower than expected. Running nvidia-smi shows that the GPU is in 'P0' state but the 'SM Clock' is significantly lower than the maximum boost clock. The job is not memory-bound. Which action should the engineer take first to diagnose the issue?

A.Reinstall the NVIDIA driver to ensure the latest version.
B.Increase the batch size to improve GPU utilization.
C.Switch the job to use mixed precision to reduce compute load.
D.Check the GPU's power consumption and whether it is hitting the power limit.
AnswerD

Low SM clock despite P0 state often indicates power or thermal throttling. Checking power consumption against the configured limit helps determine if the GPU is throttling due to power constraints. This is a common cause of reduced clocks and is a logical first diagnostic step before changing application settings.

Why this answer

When a GPU is in P0 state but SM clock is low, power or thermal throttling is a likely cause. Checking power consumption and limits is a quick, non-invasive diagnostic step. If the GPU is hitting its power limit, the engineer can then adjust power settings or optimize the workload to reduce power draw.

The other actions do not directly diagnose the throttling issue.

Exam trap

The trap here is jumping to application-level optimizations like batch size or precision changes before verifying whether the GPU is simply power-limited or thermally throttled.

64
MCQhard

A team runs a multi-node NCCL all-reduce training job on four DGX H100 nodes connected by an InfiniBand fabric. Scaling efficiency is poor: throughput barely improves beyond two nodes, and `nvidia-smi` shows NIC transmit counters on each GPU's assigned HCA are far below the PCIe link capacity while GPU compute utilization sits at ~55%. The fabric manager logs report all links as up with no symbol errors. Which action should the administrator take first to diagnose the interconnect bottleneck?

A.Switch the job to use the NCCL tree algorithm via NCCL_ALGO=TREE because ring all-reduce cannot scale across four nodes.
B.Enable GPUDirect RDMA by exporting NCCL_NET_GDR_LEVEL=SYS and rerunning the job to confirm whether host-memory staging was the cause.
C.Run `nccl-tests` all_reduce_perf with the same process count and message sizes, then compare its reported bus bandwidth against the theoretical peak for the fabric.
D.Increase the NCCL_BUFFSIZE environment variable to its maximum so each all-reduce chunk transfers more data per operation.
AnswerC

Running nccl-tests all_reduce_perf reproduces the collective in isolation and reports algorithmic and bus bandwidth, so a gap versus theoretical fabric peak localizes the loss to the communication path rather than the model code. Because every link is error-free and HCA counters are low, the fault is likely in topology, ring/tree selection, or placement, and this microbenchmark exposes exactly that before any configuration is changed.

Why this answer

When every fabric link is error-free but HCA counters sit well below link bandwidth, the bottleneck is in how the collective is scheduled or placed, not in raw link health. Reproducing the all-reduce with nccl-tests and comparing achieved bus bandwidth to the fabric's theoretical peak isolates the communication path and reveals whether the gap comes from topology detection, ring construction, or process placement before any tuning is applied.

Exam trap

The trap here is assuming that low NIC counters plus healthy links prove the network is fine, when it actually points to the collective not saturating the fabric at all.

65
MCQhard

A system administrator is troubleshooting a 'CUDA error: invalid device ordinal' when launching a job on a multi-GPU system. What is the most likely cause?

A.The GPU driver is corrupted and requires a reinstall.
B.CUDA_VISIBLE_DEVICES is set to an out-of-range index.
C.The GPU memory is full.
D.The InfiniBand fabric is misconfigured.
AnswerB

If the environment variable CUDA_VISIBLE_DEVICES specifies an index that does not exist on the host, the CUDA runtime will throw an 'invalid device ordinal' error. This is a configuration error where the system believes it has fewer GPUs than the application is trying to access via the environment variable.

Why this answer

The 'invalid device ordinal' error typically indicates that the application is attempting to access a GPU index (e.g., GPU 4) that does not exist or is not visible to the process. This is often caused by environment variables like CUDA_VISIBLE_DEVICES being configured incorrectly, mapping the application to non-existent hardware. Ensuring the logical-to-physical GPU mapping is accurate is vital for correct job execution on shared infrastructure.

Exam trap

Candidates frequently assume this error implies a faulty physical GPU or driver crash, overlooking the simpler and more common configuration error where environment variables restrict access to non-existent indices.

66
Multi-Selecthard

An AI researcher is debugging a multi-node training job using NCCL. Which TWO actions should they take to diagnose potential network-related performance degradation?

Select 2 answers
A.Set NCCL_DEBUG=INFO to monitor communication patterns and connection issues.
B.Increase the number of threads in the PyTorch DataLoader.
C.Run nccl-tests to establish a baseline for collective performance.
D.Switch from NCCL to MPI for all communication operations.
E.Lower the precision of the model to float16.
AnswersA, C

Setting NCCL_DEBUG to INFO provides detailed logs about how NCCL discovers the network topology and establishes peer-to-peer connections between nodes. This is essential for identifying misconfigured interconnects or slow paths that could be hindering collective operation performance during distributed training cycles across multiple GPU instances.

Why this answer

Debugging multi-node communication requires verifying both software configurations and physical network health. NCCL_DEBUG settings provide granular insight into connection establishment and topology detection, while NCCL tests provide baseline performance metrics. Identifying these issues early is essential for scaling models across large clusters, as network overhead can quickly become the dominant factor in training time during distributed synchronized gradient descent.

Exam trap

Candidates often attempt to debug the code logic or model hyperparameters, ignoring the specialized diagnostic tools (NCCL_DEBUG, nccl-tests) specifically designed to isolate communication and network-level issues in distributed training.

67
Multi-Selecthard

An AI operations engineer is optimizing a PyTorch training job on an NVIDIA DGX A100 system. The job uses a data loader with multiple workers, but the engineer observes that GPU utilization fluctuates between 40% and 60%, and `nvidia-smi dmon` shows periods of zero GPU utilization. The engineer suspects that the data input pipeline is the bottleneck. Which TWO actions should the engineer take to improve GPU utilization? (Choose two.)

Select 2 answers
A.Increase the number of DataLoader workers to better overlap data loading with GPU computation.
B.Decrease the batch size to reduce memory usage and allow more frequent updates.
C.Set the CUDA_LAUNCH_BLOCKING environment variable to 1 to ensure synchronous kernel execution.
D.Enable mixed precision training to reduce memory footprint and increase throughput.
E.Enable pinned memory in the DataLoader to speed up host-to-device transfers.
AnswersA, E

Increasing DataLoader workers allows more parallel data preprocessing, reducing the time the GPU waits for data. This directly addresses the bottleneck by improving overlap between CPU data loading and GPU computation, leading to higher and more consistent GPU utilization. It is a standard optimization for input-bound training jobs.

Why this answer

The observed GPU utilization fluctuations and zero-utilization periods indicate the GPU is often waiting for data. Increasing DataLoader workers parallelizes data preprocessing, and enabling pinned memory accelerates host-to-device transfers. Together, these reduce data loading latency and improve overlap with GPU computation, raising utilization.

Other options either do not address the bottleneck or harm performance.

Exam trap

The trap here is focusing on GPU-side optimizations like mixed precision when the bottleneck is actually the CPU data pipeline, which requires adjusting DataLoader parameters.

68
Multi-Selectmedium

A team trains a model inside an NGC PyTorch container on a DGX H100 node. Training starts, but after a few minutes the process dies and `dmesg` shows `Xid 79: GPU has fallen off the bus` on one GPU. The team needs to determine whether the fault is hardware or software before opening an RMA. Which two actions should they take to gather useful evidence? (Choose two.)

Select 2 answers
A.Immediately reboot the node and re-run the training job to see whether it fails again.
B.Run `nvidia-smi -q` and `nvidia-smi -q -d ECC,PAGE_RETIREMENT,ROW_REMAPPER` and capture the output.
C.Reset the GPU with `nvidia-smi --gpu-reset` and continue using the node if training succeeds afterward.
D.Collect `nvidia-bug-report.sh` output and the relevant `dmesg`/`journalctl -k` window around the failure.
E.Delete and recreate the NGC container, then pull the image again from the registry.
AnswersB, D

These queries report ECC errors, pending row remaps, and page retirement state for each GPU. On an H100, Xid 79 is frequently linked to a row-remapping or ECC escalation event, so capturing these counters before any reset preserves the evidence needed to distinguish a hardware fault from a transient software issue. The output is also requested by NVIDIA support when validating an RMA case.

Why this answer

Xid 79 indicates the GPU dropped off the PCIe bus, which can stem from a hardware fault, a degraded link, or a driver escalation after ECC or row-remap events. Before resetting anything, engineers should query ECC and row-remapper state and bundle the full bug report with the matching kernel log. These two sources together show whether the GPU had pending remaps or link errors, enabling an accurate hardware-versus-software determination and a defensible RMA decision.

Exam trap

The trap here is assuming that rebooting or resetting the GPU is a safe first step, when it actually erases the volatile counters and logs that distinguish a failing GPU from a software or fabric issue.

69
MCQhard

A team is running a multi-GPU training job on an NVIDIA DGX A100 system using NCCL for inter-GPU communication. Training throughput is much lower than expected, and the NCCL logs show repeated 'NCCL WARN Call to ibv_reg_mr failed' errors. The job uses a container with host networking. Which action should the AI operations engineer take to resolve the issue?

A.Reduce the batch size per GPU to lower memory pressure and avoid the need for large memory registrations.
B.Increase the NCCL_IB_TIMEOUT environment variable to allow more time for memory registration.
C.Disable InfiniBand by setting NCCL_IB_DISABLE=1 to force NCCL to use TCP sockets.
D.Verify and increase the container's locked memory limit (ulimit -l) and ensure the NVIDIA driver and OFED stack are compatible.
AnswerD

The 'ibv_reg_mr failed' error typically occurs when the process cannot pin enough memory for RDMA, often because the locked memory limit is too low or the OFED stack is incompatible with the NVIDIA driver. In containers, the default ulimit -l may be insufficient. Raising the limit and validating driver/OFED compatibility directly addresses the registration failure, resolving the NCCL warnings and restoring throughput.

Why this answer

The NCCL warning 'ibv_reg_mr failed' points to an inability to register memory regions for InfiniBand, commonly caused by a low locked memory limit or mismatched OFED and NVIDIA drivers. In containerized environments, the default ulimit -l is often too low. Increasing the locked memory limit and ensuring driver compatibility directly fixes the registration failure, restoring proper NCCL communication and training throughput.

Exam trap

The trap here is treating the NCCL warning as a timeout or bandwidth issue and adjusting unrelated variables instead of addressing the memory registration failure.

70
MCQmedium

Refer to the exhibit. During a multi-node training job, communication between nodes fails. What is the most likely cause of this error?

A.The GPU driver version is incompatible with the installed NCCL library.
B.The NCCL_SOCKET_IFNAME environment variable is misconfigured for the cluster network.
C.The model weights are too large to be synchronized across the network.
D.The CUDA visibility is restricted to only one GPU per node.
AnswerB

NCCL uses the NCCL_SOCKET_IFNAME variable to identify the specific network interface for inter-node communication. If this variable is unset or points to an incorrect interface, nodes cannot discover each other, leading to connection refused errors. Correctly specifying the high-speed interconnect is essential for successful multi-node operations.

Why this answer

The error indicates a network connectivity issue during the NCCL initialization phase, specifically failing to route traffic between nodes. NCCL relies on correct interface configuration for inter-node communication. Misconfigured network interfaces or firewall rules between nodes are common failure points in distributed training.

Troubleshooting these network layer issues is vital for ensuring high-performance communication across GPU clusters during large-scale model training.

Exam trap

Test-takers commonly blame the deep learning framework or code implementation rather than checking low-level cluster networking environment variables required for multi-node NCCL communication.

71
MCQmedium

Refer to the exhibit. An AI administrator investigates why a GPU node is performing significantly slower than expected. Based on the output, what is the most likely cause?

A.The GPU driver is outdated and needs a patch.
B.The server's cooling system is failing or airflow is restricted.
C.The power supply unit is malfunctioning.
D.The workload exceeds the GPU memory capacity.
AnswerB

The 'HW Thermal Slowdown' status directly confirms that the GPU has reached an internal temperature threshold and is actively reducing performance to lower heat generation. This indicates a physical cooling deficiency, likely caused by obstructed air intakes, failing server fans, or an inadequate data center ambient temperature environment.

Why this answer

The exhibit shows 'HW Thermal Slowdown' and 'HW Slowdown' are active. This indicates the GPU hardware is actively reducing its clock frequency to prevent damage due to excessive heat. This is a critical performance issue that requires immediate attention to the data center cooling or physical airflow within the server chassis.

Ensuring adequate thermal management is fundamental to maintaining consistent compute performance during heavy training loads.

Exam trap

Candidates often blame software bottlenecks or outdated drivers when unexpected performance drops occur, failing to check hardware thermal and clock throttling status.

72
MCQhard

An inference service runs a 70B parameter model with TensorRT-LLM on a single H100 using in-flight batching. Operators report that time-to-first-token is acceptable, but inter-token latency degrades sharply once concurrent request count exceeds a certain point, and GPU memory utilization sits near 98 percent. Which change most directly addresses the inter-token latency degradation?

A.Disable in-flight batching so each request is processed to completion before the next begins.
B.Increase the maximum batch size so more requests are processed per decode step.
C.Enable FP8 quantization for the KV cache and reduce the maximum batch size in the TensorRT-LLM build configuration.
D.Switch the service to a round-robin load balancer across two replicas of the same model on one GPU.
AnswerC

Near-saturated memory with growing concurrency means the KV cache is crowding out the workspace and forcing the scheduler to admit requests it cannot serve efficiently. FP8 KV cache roughly halves cache footprint, and lowering the maximum batch size keeps the runtime from over-admitting sequences. Together they reduce per-step memory pressure and shorten decode iterations, directly improving inter-token latency under high concurrency.

Why this answer

When GPU memory is nearly exhausted, the runtime cannot hold enough KV cache plus workspace to serve all admitted sequences efficiently, so each decode step stretches and inter-token latency rises with concurrency. Shrinking the KV cache with FP8 quantization and lowering the maximum batch size reduces per-step memory and compute demand, restoring shorter decode iterations. These changes target the resource constraint that actually causes the latency curve to bend upward.

Exam trap

The trap here is treating higher concurrency as a throughput problem to solve by enlarging the batch, when the observed memory saturation means the batch ceiling is already too high for the available KV cache and workspace.

73
MCQhard

An AI operations engineer is investigating intermittent failures in a long-running distributed training job. The job occasionally aborts with a collective timeout error, but no GPU errors, ECC events, or fabric link flaps appear in logs. Which action should the engineer take first to identify the root cause?

A.Restart the job with a reduced global batch size to decrease the time each collective step requires.
B.Enable per-rank NCCL debug logging and correlate the last completed collective across all ranks to find the straggler.
C.Lower the NCCL timeout value so failures surface faster and can be correlated with system events.
D.Switch the collective algorithm from ring to tree to reduce sensitivity to a single slow rank.
AnswerB

Collective timeouts occur when one or more ranks fail to arrive at a collective. Per-rank debug logs show the last operation each rank completed, so comparing them identifies the rank that stalled and the operation it was executing. This directly localizes the fault without disrupting the run, making it the correct first step.

Why this answer

A collective timeout with no hardware errors indicates one rank is arriving late or not at all. Enabling per-rank NCCL debug logging and comparing the last completed collective across ranks isolates the straggler and the operation it was executing, providing direct evidence of the fault without altering the job configuration or losing diagnostic information.

Exam trap

The trap here is treating a collective timeout as a network or GPU hardware fault, when the absence of ECC and link-flap events points instead to a single rank stalling in software or host resources.

74
MCQmedium

An engineer is troubleshooting a CUDA program that terminates unexpectedly. Which tool should be used to detect memory leaks and race conditions in the CUDA kernel code?

A.nvidia-smi
B.NVIDIA Compute Sanitizer
C.nsys (Nsight Systems)
D.nvprof
AnswerB

The NVIDIA Compute Sanitizer is a functional correctness checking tool for CUDA kernels. It can detect memory access errors, race conditions, and various other issues that lead to unexpected program termination, making it the correct choice for debugging problematic kernel code during development or deployment.

Why this answer

Debugging parallel code is inherently difficult due to the non-deterministic nature of race conditions and memory access patterns. The NVIDIA Compute Sanitizer is the essential diagnostic tool for identifying these issues, such as out-of-bounds memory access or race conditions, before they lead to erratic, intermittent crashes in production. Using it early in the development lifecycle saves massive amounts of time by ensuring code correctness and robustness.

Exam trap

Test-takers often recommend standard CPU debuggers or basic logging flags, overlooking the specialized memory and race condition diagnostic requirements unique to parallel CUDA kernel code execution.

75
MCQmedium

An operations team runs a multi-node NCCL all-reduce training job across four DGX nodes connected via InfiniBand. Training throughput is far below the expected linear scaling, and `nvidia-smi` shows GPU utilization oscillating between 20% and 40%. The network fabric is healthy and the GPUs are not thermally throttled. Which diagnostic step is MOST appropriate to identify the bottleneck?

A.Reinstall the CUDA toolkit on all nodes to ensure matching driver and runtime versions.
B.Increase the training batch size by 4x to raise arithmetic intensity and re-measure GPU utilization.
C.Enable ECC memory scrubbing on all GPUs and monitor for corrected errors during training.
D.Run `nccl-tests` (all_reduce_perf) with varying message sizes and collect NCCL debug logs by setting NCCL_DEBUG=INFO to inspect topology and algorithm selection.
AnswerD

Low, oscillating GPU utilization in a multi-node all-reduce job typically points to communication stalls. `nccl-tests` with varying message sizes reveals the achieved bus bandwidth per size, and NCCL_DEBUG=INFO exposes the detected topology, ring/tree algorithm choice, and channel count. This directly isolates whether the collective is the bottleneck and why, making it the correct diagnostic action for this scenario.

Why this answer

The oscillating, low GPU utilization across multiple nodes during an all-reduce strongly suggests the collective communication is stalling. Running `nccl-tests` with multiple message sizes measures achieved bandwidth and highlights which sizes are inefficient, while NCCL_DEBUG=INFO reveals the detected topology, algorithm, and channels. Together these pinpoint whether the bottleneck is the fabric, the ring/tree algorithm, or an unexpected topology detection.

Exam trap

The trap here is assuming low GPU utilization always means a compute problem, when in multi-node all-reduce jobs the bottleneck is often the collective communication path.

Page 1 of 2 · 86 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Troubleshooting and Optimization questions.