NCP-AIO Troubleshooting and Optimization Practice Question
An AI operations engineer is troubleshooting a multi-node NCCL training job on an NVIDIA DGX SuperPOD. The job runs but scales poorly: inter-node bandwidth is roughly half of the expected 200 Gb/s per GPU, while intra-node NVLink traffic is at full rate. Running `nvidia-smi topo -m` shows that GPUs in each node are connected to the NICs through the PCIe switch, but the job sets `NCCL_SOCKET_IFNAME` to the management interface. Which action is the most appropriate to resolve the bottleneck?
⚠ Common exam trap
The trap here is assuming that NCCL tuning parameters such as buffer size or thread count can compensate for a transport that is falling back to TCP over the management network.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable GPUDirect RDMA by allowing the container access to the host IB verbs devices and the `/dev/infiniband` tree, and set `NCCL_IB_HCA` to the compute fabric adapters.
When intra-node NVLink performs at full rate but inter-node bandwidth is roughly half, the job is usually not using the compute fabric. GPUDirect RDMA bypasses host memory copies by letting the HCA access GPU memory directly, and `NCCL_IB_HCA` ensures NCCL selects the correct adapters. Exposing `/dev/infiniband` to the container is required for RDMA verbs. This restores the expected inter-node bandwidth.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable GPUDirect RDMA by allowing the container access to the host IB verbs devices and the `/dev/infiniband` tree, and set `NCCL_IB_HCA` to the compute fabric adapters.
Why this is correct
The symptom of full intra-node NVLink but degraded inter-node throughput points to the job falling back to TCP over the management interface. GPUDirect RDMA lets the NIC read/write GPU memory directly over the InfiniBand/RoCE compute fabric, bypassing host memory copies. Exposing `/dev/infiniband` to the container and pinning `NCCL_IB_HCA` restores the intended high-speed path.
- ✗
Increase `NCCL_BUFFSIZE` to 16 MB and set `NCCL_NTHREADS` to 8 to raise channel throughput.
Why it's wrong here
Tuning buffer size and thread counts can improve overlap, but it does not change the transport. Since the job is constrained to the management interface, the fabric bandwidth ceiling remains. Raising these values would yield marginal gains at best and could even increase memory pressure without addressing the missing RDMA path.
- ✗
Set `NCCL_P2P_DISABLE=1` so that all inter-node traffic is routed through the CPUs, avoiding PCIe switch contention.
Why it's wrong here
Disabling peer-to-peer would force traffic through host memory, which is the slow path the scenario is already suffering from. It does not restore the compute fabric and would likely worsen latency. This setting is used for debugging topology issues, not for recovering InfiniBand bandwidth.
- ✗
Bind the job to a single NUMA node using `numactl --cpunodebind=0 --membind=0` to reduce cross-socket traffic.
Why it's wrong here
NUMA binding can help CPU-side affinity and memory locality, but it cannot fix a NIC that is not being used. The bottleneck is transport selection, not CPU locality. Binding to one socket may even restrict access to NICs attached to the other socket and lower throughput further.
Visual reference
About these practice questions
Courseiva writes every NCP-AIO question from scratch — 309 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.