NCP-AIO Workload Management Practice Question
A research team submits a multi-node training job using a `Job` with eight pods, each requesting one GPU. The cluster has eight GPU nodes, each with one A100. The administrator observes that all eight pods are spread one per node and the job runs, but throughput is far below expectations and NCCL logs show repeated fallback from GPUDirect RDMA to socket transport. Which action most directly addresses the root cause?
⚠ Common exam trap
The trap here is tuning NCCL environment variables or job topology when the actual gap is that the RDMA and GPUDirect components were never deployed on the nodes.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy the NVIDIA Network Operator to install and configure the RDMA stack and GPUDirect RDMA on the GPU nodes.
The symptom points to NCCL abandoning RDMA and using the slower socket transport for inter-node collectives. That path depends on the high-speed fabric drivers, RDMA devices, and GPUDirect RDMA support being installed and configured on every GPU node, which is precisely what the NVIDIA Network Operator provisions in tandem with the GPU Operator. Changing rank counts, disabling peer-to-peer, or collapsing topology onto one node does not restore the missing RDMA capability.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set `NCCL_P2P_DISABLE=1` in the job's environment to force peer-to-peer transfers over PCIe.
Why it's wrong here
Disabling peer-to-peer access removes an optimization for intra-node GPU communication; it does nothing to repair the inter-node fabric path that is failing here. Since the pods are spread one per node, the relevant path is network, not PCIe, so this variable addresses a different layer and would not restore RDMA transport.
- ✓
Deploy the NVIDIA Network Operator to install and configure the RDMA stack and GPUDirect RDMA on the GPU nodes.
Why this is correct
NCCL falling back to socket transport indicates the RDMA path is unavailable, typically because the high-speed fabric drivers, RDMA devices, and GPUDirect RDMA support are not provisioned on the nodes. The Network Operator deploys and configures those components alongside the GPU Operator, restoring the RDMA transport that NCCL prefers for multi-node collectives.
- ✗
Add a node affinity rule forcing all eight pods onto a single node to enable NVLink communication.
Why it's wrong here
A single A100 per node cannot host eight GPU-requesting pods, so this affinity rule would leave most pods Pending and the job would never start. Even on an eight-GPU node, consolidating ranks changes the topology but does not provision the missing RDMA stack that the inter-node collectives require.
- ✗
Increase the number of replicas in the Job so more GPUs participate in the collective.
Why it's wrong here
Adding replicas does not fix a transport fallback; it would simply add more participants communicating over the already-degraded socket path. The bottleneck is the interconnect path NCCL is using, not the number of ranks, so scaling out would likely worsen the symptom while consuming additional GPUs.
About these practice questions
This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.