NCP-GENL GPU Acceleration and Optimization Practice Question
A team is training a large language model using NVIDIA DGX A100 nodes with 8 GPUs per node. They observe that GPU utilization is high on all GPUs, but the training throughput scales poorly when adding more nodes. Profiling shows that the communication time during all-reduce operations increases significantly with node count. Which of the following optimizations is most likely to improve scaling efficiency?
⚠ Common exam trap
The trap here is assuming that gradient accumulation or CPU thread tuning addresses communication overhead, when the bottleneck is specifically inter-node all-reduce performance.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure NCCL to use InfiniBand with GPUDirect RDMA and enable adaptive routing.
The poor scaling with increasing node count points to inter-node communication overhead, specifically during all-reduce operations. Leveraging InfiniBand with GPUDirect RDMA and adaptive routing optimizes the network path, reducing latency and improving bandwidth. This directly targets the communication bottleneck, enabling better scaling efficiency for distributed LLM training across multiple DGX nodes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use gradient accumulation to increase the effective batch size.
Why it's wrong here
Gradient accumulation increases the effective batch size without requiring more memory, but it does not directly address inter-node communication overhead. While it can improve convergence stability, it does not reduce the time spent in all-reduce operations. In fact, it may increase the number of all-reduce calls if not carefully tuned, exacerbating the scaling issue.
- ✗
Increase the number of CPU threads dedicated to data loading.
Why it's wrong here
Increasing CPU threads for data loading can help if the input pipeline is a bottleneck, but here profiling indicates that communication time during all-reduce is the primary issue. Data loading optimizations would not reduce inter-node communication overhead. Therefore, this change would not improve scaling efficiency in this scenario.
- ✗
Enable NCCL's tree algorithm for all-reduce operations.
Why it's wrong here
NCCL's tree algorithm is typically used for small message sizes and can reduce latency but may not be optimal for large messages common in LLM training. For large all-reduce operations, ring or hierarchical algorithms often provide better bandwidth utilization. Enabling tree could actually worsen performance if message sizes are large, making it an ineffective choice for this scenario.
- ✓
Configure NCCL to use InfiniBand with GPUDirect RDMA and enable adaptive routing.
Why this is correct
Using InfiniBand with GPUDirect RDMA allows GPUs to communicate directly with the network adapter, bypassing host memory and reducing latency. Adaptive routing dynamically selects less congested paths, improving bandwidth utilization across multiple nodes. This combination significantly reduces communication overhead in large-scale all-reduce operations, directly addressing the poor scaling observed when adding nodes.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.