NCP-AIO Troubleshooting and Optimization Practice Question
An operations team observes that a distributed training job using NVIDIA Collective Communications Library (NCCL) across eight nodes occasionally hangs during the all-reduce phase. Logs show no errors, and the hang resolves only after a node is manually restarted. Which action is MOST appropriate to diagnose the intermittent hang?
⚠ Common exam trap
The trap here is trying to work around the hang by changing transport or algorithm rather than capturing diagnostic data at the point of failure.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable NCCL debug logging with NCCL_DEBUG=INFO and NCCL_DEBUG_SUBSYS=COLL, and set a NCCL watchdog timeout to capture the stalled collective.
Intermittent NCCL hangs without errors are best diagnosed by capturing state at the moment of the stall. Enabling NCCL debug logging for the collective subsystem and configuring a watchdog timeout produces logs that show which rank and collective stopped progressing. This evidence is necessary to distinguish a slow rank, a fabric problem, or a software defect from a simple performance issue.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Disable InfiniBand and fall back to TCP sockets by setting NCCL_IB_DISABLE=1.
Why it's wrong here
Falling back to TCP changes the transport and may avoid the hang, but it does not diagnose the root cause and typically reduces performance. It could also hide a fabric or driver issue that would recur elsewhere. Because the objective is to identify the cause of the intermittent hang, disabling the primary transport is a workaround rather than a diagnostic step and is not appropriate.
- ✗
Increase the NCCL buffer size with NCCL_BUFFSIZE to reduce the number of messages exchanged.
Why it's wrong here
NCCL_BUFFSIZE affects chunking within collectives but does not expose or resolve intermittent hangs. Larger buffers can change timing and potentially mask the issue without revealing the stalled rank or collective. Since the goal is diagnosis, not tuning, this parameter change does not provide the needed visibility into where the hang occurs and is not the appropriate action.
- ✗
Set NCCL_ALGO=RING to force a single algorithm and eliminate variability.
Why it's wrong here
Forcing a specific algorithm removes algorithm variability but does not provide diagnostic information about why a collective stalls. The hang could occur in any algorithm depending on the underlying cause, such as a slow rank or fabric issue. Without logs or a watchdog dump, this change would not reveal the stalled component and is therefore not the correct diagnostic action.
- ✓
Enable NCCL debug logging with NCCL_DEBUG=INFO and NCCL_DEBUG_SUBSYS=COLL, and set a NCCL watchdog timeout to capture the stalled collective.
Why this is correct
Intermittent hangs without errors require visibility into which collective and rank stalled. NCCL_DEBUG=INFO with NCCL_DEBUG_SUBSYS=COLL prints per-collective progress and rank participation, while a watchdog timeout can trigger a dump when progress stops. This combination captures the state at the moment of the hang, making it the correct diagnostic action for this scenario.
Visual reference
About these practice questions
Courseiva writes every NCP-AIO question from scratch — 309 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.