NCP-AIO Troubleshooting and Optimization Practice Question
A team runs multi-node training with NCCL over InfiniBand on a cluster of DGX systems. Jobs scale well to four nodes but throughput drops sharply at eight nodes, and `nccl-tests` all-reduce bandwidth falls well below line rate at that size. The fabric uses a fat-tree topology with adaptive routing enabled. Which investigation is most likely to reveal the cause?
⚠ Common exam trap
The trap here is assuming the interconnect is broken or that adaptive routing is at fault, when the collapse at a specific node count usually reflects NCCL's topology and algorithm choices crossing oversubscribed links.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Check whether NCCL is selecting the correct InfiniBand HCAs and whether the ring or tree algorithm choice matches the fabric's oversubscription ratio.
NCCL chooses rings and trees based on detected topology, and on multi-node jobs this can cross oversubscribed uplinks or use a suboptimal adapter when several are present. When scaling stops at a specific node count, the cause is usually that the communication pattern no longer matches the fabric's capacity. Confirming HCA selection and steering the algorithm or topology to respect the oversubscription ratio restores bandwidth and explains why smaller jobs appeared fine.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Check whether NCCL is selecting the correct InfiniBand HCAs and whether the ring or tree algorithm choice matches the fabric's oversubscription ratio.
Why this is correct
Beyond a certain node count, NCCL's default topology detection can pick a ring that traverses oversubscribed uplinks, or it can select the wrong HCA when multiple adapters exist per node. Verifying `NCCL_IB_HCA` and forcing the appropriate algorithm or topology file aligns communication with the physical fabric. This directly explains why bandwidth collapses only at larger scale while small jobs look healthy.
- ✗
Move the collective operations from NCCL to a TCP socket backend over the management network.
Why it's wrong here
The management network offers far lower bandwidth and higher latency than InfiniBand, so this would reduce training throughput dramatically. It also abandons the high-speed interconnect that the cluster was built around. The investigation should identify why the existing InfiniBand path underperforms, not replace it with a slower transport.
- ✗
Increase the batch size per GPU so that communication is amortized over more compute.
Why it's wrong here
Larger batches reduce the frequency of all-reduce operations but do not raise the achievable bandwidth of a single collective. The nccl-tests result shows the transport itself is underperforming at eight nodes, independent of how often it is invoked. Amortization hides the symptom at the application level while leaving the fabric or algorithm misconfiguration unresolved.
- ✗
Disable adaptive routing on the InfiniBand fabric to force deterministic paths.
Why it's wrong here
Adaptive routing generally improves utilization on fat-tree fabrics by spreading traffic across available paths. Disabling it can create congestion on fixed routes and degrade collective performance further. The described symptom, bandwidth falling below line rate only at larger scale, is more consistent with a topology or HCA selection mismatch than with adaptive routing being harmful.
About these practice questions
Courseiva writes every NCP-AIO question from scratch — 309 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.