Courseiva
Workload Management →hardMultiple Choice

NCP-AIO Workload Management Practice Question

A research organization runs an NVIDIA DGX SuperPOD with a Kubernetes cluster managed by the NVIDIA GPU Operator and Network Operator. A distributed training job using PyTorch DDP across 32 nodes stalls at initialization, and the administrator suspects the collective communication library is not selecting the high-speed fabric. Which configuration should the administrator verify first to ensure NCCL uses the correct network interface and topology?

⚠ Common exam trap

The trap here is assuming that adding GPUs per pod or adjusting CUDA_VISIBLE_DEVICES will fix a distributed hang, when the real cause is NCCL's network interface and fabric selection.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Confirm that NCCL_IB_DISABLE is set to 0, NCCL_SOCKET_IFNAME matches the high-speed fabric interface, and NCCL_TOPO_FILE or the topology-aware plugin is loaded on each node.

NCCL chooses transports and interfaces based on environment variables and detected topology. When NCCL_IB_DISABLE is set incorrectly or NCCL_SOCKET_IFNAME points to the wrong interface, collectives fall back to TCP over the management network or fail to connect, causing distributed jobs to hang at initialization. Verifying these variables and the topology file on every node is the correct first diagnostic step.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the pod's nvidia.com/gpu limit to 8 so each node exposes all GPUs to the training process.

    Why it's wrong here

    GPU count per pod affects local parallelism but does not influence which network fabric NCCL selects for inter-node communication. If the interface variables are wrong, adding more GPUs per pod will not fix the collective hang and may even increase the number of stalled ranks. The network path selection is governed by NCCL environment variables, not GPU limits.

  • ✗

    Set the CUDA_VISIBLE_DEVICES variable to list all GPUs and restart the training job.

    Why it's wrong here

    CUDA_VISIBLE_DEVICES controls which GPUs a process can see locally; it has no effect on inter-node communication or fabric selection. A hang during distributed initialization is typically caused by NCCL failing to establish connections over the intended interface, not by GPU visibility. Adjusting this variable will not resolve a collective communication stall.

  • ✗

    Enable the NVIDIA MIG feature on all nodes so each rank gets an isolated GPU slice for communication.

    Why it's wrong here

    MIG isolates GPU resources within a single device but does not alter how NCCL selects network interfaces across nodes. Enabling MIG could actually reduce available bandwidth per rank and complicate topology discovery. The stall at initialization points to fabric selection, so MIG configuration is irrelevant to resolving the communication failure.

  • ✓

    Confirm that NCCL_IB_DISABLE is set to 0, NCCL_SOCKET_IFNAME matches the high-speed fabric interface, and NCCL_TOPO_FILE or the topology-aware plugin is loaded on each node.

    Why this is correct

    NCCL relies on these environment variables and topology data to select InfiniBand or RoCE interfaces and to build the correct ring or tree topology across nodes. If NCCL_IB_DISABLE is set to 1, or NCCL_SOCKET_IFNAME points at the management interface, NCCL falls back to slower paths and initialization can stall. Verifying these values is the primary diagnostic step.

About these practice questions

This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.