NCP-AIO Troubleshooting and Optimization Practice Question
An AI operations team is troubleshooting a distributed training job on a cluster of NVIDIA DGX A100 systems connected via InfiniBand. The job runs but achieves only 40% of expected scaling efficiency. The team suspects communication bottlenecks. Which two actions should they take to confirm and address the issue? (Choose two.)
⚠ Common exam trap
The trap here is focusing on GPU-level metrics or hyperparameter changes instead of directly inspecting NCCL communication behavior and InfiniBand fabric health.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Verify that all nodes have identical NCCL versions and environment variables, and that InfiniBand fabric manager is running.
Scaling inefficiency in distributed training often stems from communication overhead. NCCL debug logs directly show transport selection and errors, while ensuring consistent NCCL versions and InfiniBand fabric health addresses common misconfigurations. Together, these actions confirm and resolve bottlenecks. The other options either do not diagnose the network or would degrade performance.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Verify that all nodes have identical NCCL versions and environment variables, and that InfiniBand fabric manager is running.
Why this is correct
Inconsistent NCCL versions or missing environment variables (e.g., NCCL_IB_HCA, NCCL_SOCKET_IFNAME) can cause suboptimal transport selection or failures. The InfiniBand fabric manager ensures proper link configuration. Ensuring uniformity across nodes is crucial for optimal collective communication performance and is a common fix for scaling inefficiencies.
- ✓
Run NCCL_DEBUG=INFO and inspect the logs for warnings about falling back to slower transports or retries.
Why this is correct
NCCL debug logs reveal the transport used (e.g., InfiniBand vs. TCP) and any errors or retries. If NCCL falls back to TCP or shows excessive retransmissions, it indicates a communication bottleneck. This is a primary diagnostic step to confirm whether the interconnect is the limiting factor and to guide corrective actions such as driver or topology fixes.
- ✗
Switch the job to use Ethernet instead of InfiniBand to simplify troubleshooting.
Why it's wrong here
Ethernet typically offers lower bandwidth and higher latency than InfiniBand in HPC clusters, which would likely worsen scaling efficiency. Switching to Ethernet is not a troubleshooting step but a downgrade. The goal is to identify and fix InfiniBand issues, not to avoid them by using an inferior interconnect.
- ✗
Use nvidia-smi to monitor GPU utilization and memory bandwidth on each node during training.
Why it's wrong here
nvidia-smi provides GPU utilization and memory throughput but does not directly measure interconnect performance or NCCL communication efficiency. While low GPU utilization could hint at communication stalls, it is not specific enough to confirm an InfiniBand bottleneck. This action is more about overall GPU health than network diagnostics.
- ✗
Increase the global batch size proportionally to the number of GPUs to improve compute-to-communication ratio.
Why it's wrong here
Increasing batch size can improve efficiency by amortizing communication overhead, but it does not diagnose the root cause and may harm convergence. In this scenario, the team first needs to confirm the bottleneck. Blindly changing hyperparameters could mask the issue without resolving underlying network problems, and it may not be appropriate for all models.
Visual reference
About these practice questions
One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.