NCP-AIO Troubleshooting and Optimization Practice Question
A team runs a multi-node NCCL all-reduce training job on four DGX H100 nodes connected by an InfiniBand fabric. Scaling efficiency is poor: throughput barely improves beyond two nodes, and `nvidia-smi` shows NIC transmit counters on each GPU's assigned HCA are far below the PCIe link capacity while GPU compute utilization sits at ~55%. The fabric manager logs report all links as up with no symbol errors. Which action should the administrator take first to diagnose the interconnect bottleneck?
⚠ Common exam trap
The trap here is assuming that low NIC counters plus healthy links prove the network is fine, when it actually points to the collective not saturating the fabric at all.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Run `nccl-tests` all_reduce_perf with the same process count and message sizes, then compare its reported bus bandwidth against the theoretical peak for the fabric.
When every fabric link is error-free but HCA counters sit well below link bandwidth, the bottleneck is in how the collective is scheduled or placed, not in raw link health. Reproducing the all-reduce with nccl-tests and comparing achieved bus bandwidth to the fabric's theoretical peak isolates the communication path and reveals whether the gap comes from topology detection, ring construction, or process placement before any tuning is applied.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Switch the job to use the NCCL tree algorithm via NCCL_ALGO=TREE because ring all-reduce cannot scale across four nodes.
Why it's wrong here
Ring all-reduce scales well across multiple nodes; it is the default for good reason, and tree is generally preferred for smaller messages or when the ring is constrained by a slow link. Forcing tree without evidence of a ring-specific fault masks symptoms rather than diagnosing them, and a misconfigured tree can perform worse for the large payloads typical in distributed training.
- ✗
Enable GPUDirect RDMA by exporting NCCL_NET_GDR_LEVEL=SYS and rerunning the job to confirm whether host-memory staging was the cause.
Why it's wrong here
GPUDirect RDMA level tuning is a remediation, not a diagnostic step, and on a healthy DGX H100 with current drivers it is normally enabled by default. Setting the level blindly cannot explain why throughput plateaus at two nodes, and changing transport behavior before measuring the collective makes it impossible to attribute any improvement to a specific root cause.
- ✓
Run `nccl-tests` all_reduce_perf with the same process count and message sizes, then compare its reported bus bandwidth against the theoretical peak for the fabric.
Why this is correct
Running nccl-tests all_reduce_perf reproduces the collective in isolation and reports algorithmic and bus bandwidth, so a gap versus theoretical fabric peak localizes the loss to the communication path rather than the model code. Because every link is error-free and HCA counters are low, the fault is likely in topology, ring/tree selection, or placement, and this microbenchmark exposes exactly that before any configuration is changed.
- ✗
Increase the NCCL_BUFFSIZE environment variable to its maximum so each all-reduce chunk transfers more data per operation.
Why it's wrong here
NCCL_BUFFSIZE enlarges the per-channel buffer used for chunked transfers, which can help small-message latency but does nothing when the fabric is already underutilized at large message sizes. With HCA counters far below link capacity, the problem is not insufficient buffer depth; raising it consumes more memory per channel and may even reduce the number of channels NCCL can open, worsening the observed scaling.
About these practice questions
Courseiva writes every NCP-AIO question from scratch — 309 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.