NCP-AIO Troubleshooting and Optimization Practice Question
A distributed training job on a multi-GPU node is exhibiting poor scaling efficiency: each GPU shows high utilization, but overall throughput increases by only 15% when doubling the number of GPUs. The job uses NCCL for communication. Which diagnostic step is most appropriate to identify the bottleneck?
⚠ Common exam trap
The trap here is assuming that high GPU utilization guarantees optimal scaling, overlooking that inter-GPU communication overhead can dominate even when GPUs appear busy.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Run `nvidia-smi topo -m` to inspect the GPU interconnect topology and verify that NCCL is using the fastest available path between GPUs.
Poor scaling efficiency in distributed training often results from communication bottlenecks. The `nvidia-smi topo -m` command provides a matrix of GPU interconnect paths, helping verify whether NCCL can use NVLink or is forced over slower PCIe/QPI links. If the topology shows suboptimal paths, adjusting NCCL environment variables or hardware configuration can improve scaling. Thus, inspecting the topology is the most direct diagnostic step.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable NVIDIA MPS (Multi-Process Service) to allow multiple processes to share the GPUs more efficiently.
Why it's wrong here
MPS is designed to improve concurrency for multiple processes sharing a single GPU, not to enhance inter-GPU communication for a single distributed job. In fact, MPS can interfere with NCCL's ability to manage resources. The scaling issue stems from inter-GPU data transfer, so MPS would not address the bottleneck and might introduce additional complexity or contention, making it an inappropriate diagnostic or corrective step.
- ✗
Check the GPU clock speeds with `nvidia-smi -q -d CLOCK` to ensure they are running at maximum frequency.
Why it's wrong here
Clock speed issues typically cause lower per-GPU throughput, but here GPUs are already highly utilized. The problem is scaling across GPUs, not individual GPU performance. Monitoring clocks would not reveal communication bottlenecks between GPUs. While thermal throttling can reduce performance, the scenario states high utilization, making clock speed a less likely culprit for poor scaling efficiency.
- ✗
Increase the batch size per GPU to reduce the frequency of gradient synchronization across GPUs.
Why it's wrong here
While increasing batch size can reduce synchronization frequency, it changes the training dynamics and may affect convergence. It does not diagnose the root cause of poor scaling. The symptom of high utilization with low scaling points to communication overhead, not insufficient work per GPU. Blindly increasing batch size could mask the issue or lead to out-of-memory errors, and it fails to confirm whether the interconnect is the actual bottleneck.
- ✓
Run `nvidia-smi topo -m` to inspect the GPU interconnect topology and verify that NCCL is using the fastest available path between GPUs.
Why this is correct
Poor scaling with high GPU utilization often indicates communication overhead. `nvidia-smi topo -m` reveals whether GPUs are connected via NVLink or PCIe and whether data traverses slower paths like QPI/UPI. If NCCL cannot leverage NVLink, all-reduce operations become the bottleneck. This command is the standard first step to diagnose interconnect issues and ensure NCCL selects the optimal topology-aware route, directly addressing the scaling inefficiency.
About these practice questions
One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.