Courseiva
Model Optimization →hardMultiple Choice

NCP-GENL Model Optimization Practice Question

A team is deploying a large language model using NVIDIA TensorRT-LLM on a multi-GPU node with NVLink. They want to minimize inter-GPU communication overhead during inference. Which parallelism strategy should they use to achieve this?

⚠ Common exam trap

The trap here is assuming that pipeline parallelism minimizes communication because it reduces frequency, but it introduces pipeline bubbles and does not leverage NVLink as effectively as tensor parallelism.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Tensor parallelism (TP)

Tensor parallelism splits layers across GPUs and uses all-reduce for activations. On a system with NVLink, the high bandwidth and low latency of NVLink make TP efficient, minimizing communication overhead. Pipeline parallelism introduces bubbles, data parallelism replicates the model, and expert parallelism is for MoE models and incurs all-to-all communication.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Pipeline parallelism (PP)

    Why it's wrong here

    Pipeline parallelism divides layers into stages across GPUs, reducing communication frequency but introducing pipeline bubbles and requiring careful micro-batch scheduling. While it can be efficient, it does not inherently minimize communication overhead as effectively as TP on NVLink, and it may increase latency due to bubbles.

  • ✗

    Expert parallelism (EP) in a Mixture-of-Experts model

    Why it's wrong here

    Expert parallelism distributes experts across GPUs, requiring all-to-all communication to route tokens to the appropriate experts. This can introduce significant communication overhead, especially with many experts. While NVLink helps, EP is not the best choice for a dense LLM without MoE architecture, and it does not minimize overhead compared to TP.

  • ✗

    Data parallelism (DP)

    Why it's wrong here

    Data parallelism replicates the entire model on each GPU and splits the batch. It requires no inter-GPU communication during inference if each GPU processes separate requests, but it does not reduce the model's memory footprint per GPU. The goal is to minimize communication overhead, but DP does not address model size and may not be feasible for large models.

  • ✓

    Tensor parallelism (TP)

    Why this is correct

    Tensor parallelism splits individual layers across GPUs, requiring frequent all-reduce operations for activations. NVLink provides high bandwidth and low latency, making TP efficient. By using TP, the team can minimize communication overhead compared to other strategies that might use slower interconnects or require more synchronization, thus achieving the goal.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.