Courseiva

NCP-AIO Troubleshooting and Optimization Practice Question

An operations team runs a multi-node NCCL all-reduce training job across four DGX nodes connected via InfiniBand. Training throughput is far below the expected linear scaling, and `nvidia-smi` shows GPU utilization oscillating between 20% and 40%. The network fabric is healthy and the GPUs are not thermally throttled. Which diagnostic step is MOST appropriate to identify the bottleneck?

⚠ Common exam trap

The trap here is assuming low GPU utilization always means a compute problem, when in multi-node all-reduce jobs the bottleneck is often the collective communication path.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Run `nccl-tests` (all_reduce_perf) with varying message sizes and collect NCCL debug logs by setting NCCL_DEBUG=INFO to inspect topology and algorithm selection.

The oscillating, low GPU utilization across multiple nodes during an all-reduce strongly suggests the collective communication is stalling. Running `nccl-tests` with multiple message sizes measures achieved bandwidth and highlights which sizes are inefficient, while NCCL_DEBUG=INFO reveals the detected topology, algorithm, and channels. Together these pinpoint whether the bottleneck is the fabric, the ring/tree algorithm, or an unexpected topology detection.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Reinstall the CUDA toolkit on all nodes to ensure matching driver and runtime versions.

    Why it's wrong here

    Version mismatches would typically cause hard failures or clear error messages, not a consistent pattern of low but functional throughput with oscillating utilization. Reinstalling the toolkit is disruptive and does not provide any diagnostic signal about the collective communication path. Without evidence of a version conflict, this action is unwarranted and misdirected for the described symptoms.

  • ✗

    Increase the training batch size by 4x to raise arithmetic intensity and re-measure GPU utilization.

    Why it's wrong here

    Changing the batch size alters the workload rather than diagnosing it, and can even worsen the imbalance if the collective is already the bottleneck. It does not reveal why NCCL communication is slow or whether the topology was detected correctly. This is a tuning action, not a diagnostic step, and would mask rather than identify the root cause of the oscillating utilization.

  • ✗

    Enable ECC memory scrubbing on all GPUs and monitor for corrected errors during training.

    Why it's wrong here

    ECC errors, when present, usually manifest as XID errors, job aborts, or performance degradation tied to memory traffic, not as a clean oscillation between 20% and 40% utilization synchronized across nodes. Enabling scrubbing adds overhead without explaining the all-reduce stalls. This action addresses memory integrity, which is not implicated by the described network-bound symptom pattern.

  • ✓

    Run `nccl-tests` (all_reduce_perf) with varying message sizes and collect NCCL debug logs by setting NCCL_DEBUG=INFO to inspect topology and algorithm selection.

    Why this is correct

    Low, oscillating GPU utilization in a multi-node all-reduce job typically points to communication stalls. `nccl-tests` with varying message sizes reveals the achieved bus bandwidth per size, and NCCL_DEBUG=INFO exposes the detected topology, ring/tree algorithm choice, and channel count. This directly isolates whether the collective is the bottleneck and why, making it the correct diagnostic action for this scenario.

About these practice questions

One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.