NCP-AIO Administration Practice Question
An administrator is troubleshooting a multi-node NVIDIA GPU training job that intermittently hangs during the all-reduce phase. The job uses NCCL over InfiniBand. Logs show that some ranks time out while others complete. The administrator suspects a network fabric issue. Which action should the administrator take first to isolate whether the problem is in the InfiniBand fabric or in the NCCL configuration?
⚠ Common exam trap
The trap here is jumping to hardware replacement or timeout adjustments before verifying whether NCCL is correctly configured and using the expected transport.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Run the NCCL tests (nccl-tests) with the same topology and environment variables, and enable NCCL debug logging to capture detailed transport and topology information.
Using nccl-tests with debug logging is a controlled way to reproduce the collective communication and observe transport selection, topology detection, and timeouts. It helps determine whether the issue is in NCCL configuration or the InfiniBand fabric. Replacing hardware, forcing TCP, or increasing timeouts are premature and do not isolate the root cause.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Disable InfiniBand and force NCCL to use TCP sockets instead, because TCP is more reliable for collective operations.
Why it's wrong here
Forcing TCP sockets would bypass the InfiniBand fabric and likely reduce performance significantly. It does not diagnose the root cause and may mask a configuration issue. TCP is not inherently more reliable for high-performance collectives; InfiniBand is preferred for low latency and high bandwidth.
- ✗
Immediately replace all InfiniBand cables and transceivers on the affected nodes, because intermittent hangs during all-reduce almost always indicate physical link errors.
Why it's wrong here
Replacing cables and transceivers is disruptive and premature without evidence of physical layer errors. The administrator should first determine whether the issue is in NCCL configuration or the fabric. Physical replacement does not address configuration mismatches and may not resolve the hang.
- ✓
Run the NCCL tests (nccl-tests) with the same topology and environment variables, and enable NCCL debug logging to capture detailed transport and topology information.
Why this is correct
Running nccl-tests with debug logging reproduces the collective communication pattern outside the training job. It reveals which transport (InfiniBand, RoCE, or sockets) is used, whether topology detection fails, and where timeouts occur. This isolates NCCL configuration issues from fabric problems before deeper network diagnostics.
- ✗
Increase the NCCL timeout value in the training script and restart the job, because the hangs are likely due to slow network convergence.
Why it's wrong here
Increasing the timeout may delay the symptom but does not identify the cause. If the fabric is healthy but NCCL is misconfigured, a longer timeout may not help. This action postpones diagnosis and could allow the job to hang longer, wasting resources.
About these practice questions
One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.