NCP-GENL GPU Acceleration and Optimization Practice Question
A team is training a large language model on 8 NVIDIA A100 GPUs using PyTorch's DistributedDataParallel (DDP). Profiling shows that all GPUs are frequently idle, waiting for gradient synchronization. The network interconnect between nodes is a 100 Gb Ethernet with TCP/IP, and the model has 13 billion parameters. What is the most effective optimization to reduce the idle time?
⚠ Common exam trap
The trap here is assuming that increasing batch size or compressing gradients will fully resolve communication stalls, when the underlying high-latency interconnect is the true limiting factor.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Replace the Ethernet interconnect with NVIDIA Mellanox InfiniBand and enable NCCL over RDMA.
The idle time is caused by slow gradient synchronization over TCP/IP Ethernet. InfiniBand with RDMA and NCCL provides the necessary bandwidth and low latency to reduce all-reduce time, directly addressing the bottleneck. Other options either do not target communication or are secondary optimizations that do not resolve the fundamental interconnect limitation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the batch size per GPU to keep the GPUs busy during communication.
Why it's wrong here
Increasing batch size may improve compute utilization but does not address the communication bottleneck. Larger batches can even increase gradient size and prolong synchronization. The GPUs remain idle waiting for all-reduce to complete; the root cause is slow interconnect, not insufficient work.
- ✓
Replace the Ethernet interconnect with NVIDIA Mellanox InfiniBand and enable NCCL over RDMA.
Why this is correct
InfiniBand with RDMA provides higher bandwidth and lower latency than TCP/IP over Ethernet, reducing the all-reduce time for gradient synchronization. For a 13B parameter model, gradient tensors are large, and the communication overhead dominates. NCCL over RDMA bypasses the CPU and kernel network stack, significantly improving throughput and lowering idle time.
- ✗
Enable gradient compression using FP16 all-reduce to halve communication volume.
Why it's wrong here
Gradient compression with FP16 can reduce communication volume, but it may affect convergence and requires careful scaling. However, the primary bottleneck is the high latency of TCP/IP over Ethernet, not just volume. Without RDMA, latency remains high, and the idle time persists. Thus, this is secondary to upgrading the interconnect.
- ✗
Use NVIDIA GPUDirect Storage to accelerate data loading from NVMe drives.
Why it's wrong here
GPUDirect Storage accelerates data transfer between storage and GPU memory, which is irrelevant to gradient synchronization across GPUs. The bottleneck is inter-GPU communication, not data loading. This would not reduce the idle time caused by waiting for all-reduce.
Visual reference
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.