NCP-GENL Model Optimization Practice Question
A team is deploying a large language model using NVIDIA TensorRT-LLM on a multi-GPU node with NVLink. They want to minimize inter-GPU communication overhead during inference. Which parallelism strategy should they use to achieve this?
⚠ Common exam trap
The trap here is assuming that pipeline parallelism minimizes communication because it reduces frequency, but it introduces pipeline bubbles and does not leverage NVLink as effectively as tensor parallelism.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Tensor parallelism (TP)
Tensor parallelism splits layers across GPUs and uses all-reduce for activations. On a system with NVLink, the high bandwidth and low latency of NVLink make TP efficient, minimizing communication overhead. Pipeline parallelism introduces bubbles, data parallelism replicates the model, and expert parallelism is for MoE models and incurs all-to-all communication.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Pipeline parallelism (PP)
Why it's wrong here
Pipeline parallelism divides layers into stages across GPUs, reducing communication frequency but introducing pipeline bubbles and requiring careful micro-batch scheduling. While it can be efficient, it does not inherently minimize communication overhead as effectively as TP on NVLink, and it may increase latency due to bubbles.
- ✗
Expert parallelism (EP) in a Mixture-of-Experts model
Why it's wrong here
Expert parallelism distributes experts across GPUs, requiring all-to-all communication to route tokens to the appropriate experts. This can introduce significant communication overhead, especially with many experts. While NVLink helps, EP is not the best choice for a dense LLM without MoE architecture, and it does not minimize overhead compared to TP.
- ✗
Data parallelism (DP)
Why it's wrong here
Data parallelism replicates the entire model on each GPU and splits the batch. It requires no inter-GPU communication during inference if each GPU processes separate requests, but it does not reduce the model's memory footprint per GPU. The goal is to minimize communication overhead, but DP does not address model size and may not be feasible for large models.
- ✓
Tensor parallelism (TP)
Why this is correct
Tensor parallelism splits individual layers across GPUs, requiring frequent all-reduce operations for activations. NVLink provides high bandwidth and low latency, making TP efficient. By using TP, the team can minimize communication overhead compared to other strategies that might use slower interconnects or require more synchronization, thus achieving the goal.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.