NCP-GENL GPU Acceleration and Optimization Practice Question
A team is training a 13B-parameter LLM on 8 NVIDIA A100 GPUs using NVIDIA NeMo. They observe that the all-reduce communication during data-parallel training consumes nearly 40% of each iteration. Which of the following changes is most likely to reduce this communication overhead while preserving convergence?
⚠ Common exam trap
The trap here is assuming that any parallelism change (like tensor parallelism) will reduce communication, when it often increases it due to more frequent synchronization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use NVIDIA NCCL with the ring algorithm and overlap communication with computation via gradient bucketing.
The communication bottleneck in data-parallel training is best mitigated by using an efficient collective algorithm and hiding latency. NCCL's ring algorithm excels with large messages on high-bandwidth links, and gradient bucketing enables overlap with compute. This preserves the data-parallel semantics and convergence while reducing wall-clock time spent in all-reduce.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Switch from data parallelism to tensor parallelism across all 8 GPUs.
Why it's wrong here
Tensor parallelism splits individual layers across GPUs, which introduces frequent all-reduce operations within each layer and often increases communication volume. It is typically used when a model does not fit on a single GPU, not to reduce communication in an already-fitting data-parallel setup. For this scenario, it would likely worsen the bottleneck rather than alleviate it.
- ✗
Enable gradient accumulation with a larger micro-batch size and use NCCL with tree algorithm.
Why it's wrong here
Gradient accumulation reduces the number of all-reduce calls by increasing the effective batch size, but it also increases memory pressure and may not reduce total communication volume. The NCCL tree algorithm can help for small messages but is not a guaranteed fix for large all-reduce. This combination is plausible but not the most direct solution here.
- ✓
Use NVIDIA NCCL with the ring algorithm and overlap communication with computation via gradient bucketing.
Why this is correct
NCCL's ring algorithm is optimized for large messages and high-bandwidth interconnects like NVLink, making it efficient for all-reduce in data-parallel training. Overlapping communication with computation using gradient bucketing hides latency behind backpropagation. Together, these reduce the effective communication overhead without changing model convergence, directly addressing the observed bottleneck.
- ✗
Reduce the number of GPUs to 4 and increase the per-GPU batch size.
Why it's wrong here
Reducing the number of GPUs lowers total compute throughput and may require a smaller global batch size to fit memory, potentially harming convergence. While it reduces the number of participants in all-reduce, it does not address the underlying inefficiency of the communication pattern and sacrifices scalability. This is a step backward for a large model.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.