Courseiva
Model Deployment →hardMultiple Choice

NCP-GENL Model Deployment Practice Question

A team is deploying a 70B-parameter LLM across four NVIDIA H100 GPUs using NVIDIA TensorRT-LLM with tensor parallelism. They observe that inference works but throughput is lower than expected, and profiling shows significant inter-GPU communication overhead. Which optimization should they apply first to reduce communication overhead?

⚠ Common exam trap

The trap here is thinking that pipeline parallelism eliminates inter-GPU communication, when it actually introduces different communication and pipeline bubbles.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable NVLink and ensure the GPUs are connected via NVSwitch for peer-to-peer communication.

Tensor parallelism relies on frequent all-reduce operations, so the interconnect bandwidth is critical. NVLink with NVSwitch provides the necessary high-speed peer-to-peer communication to minimize overhead. Without it, PCIe becomes the bottleneck. Other options either do not address the communication pattern or are infeasible for a 70B model on four GPUs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Enable NVLink and ensure the GPUs are connected via NVSwitch for peer-to-peer communication.

    Why this is correct

    Tensor parallelism requires frequent all-reduce operations between GPUs. NVLink with NVSwitch provides high-bandwidth, low-latency peer-to-peer communication, which is essential to reduce the overhead. Without NVLink, communication over PCIe becomes a bottleneck. Ensuring NVLink is enabled and the topology uses NVSwitch is the first and most impactful optimization for multi-GPU tensor parallelism.

  • ✗

    Reduce the tensor parallel size to 2 and run two independent replicas.

    Why it's wrong here

    Reducing tensor parallel size to 2 means each GPU holds more parameters, likely causing out-of-memory for a 70B model. Running two replicas would require eight GPUs, not four. This does not address the communication overhead and may be infeasible. Tensor parallel size must match the model sharding needs and available GPU memory.

  • ✗

    Increase the number of attention heads to improve parallelism.

    Why it's wrong here

    The number of attention heads is a property of the model architecture and cannot be arbitrarily increased without retraining. Changing it would alter the model and likely degrade accuracy. It does not reduce communication overhead; it may increase it. The scenario requires an infrastructure or runtime optimization, not a model architecture change.

  • ✗

    Switch from tensor parallelism to pipeline parallelism to eliminate inter-GPU communication.

    Why it's wrong here

    Pipeline parallelism still requires communication of activations between stages and introduces pipeline bubbles, which can reduce throughput. It does not eliminate inter-GPU communication; it changes the pattern. For a 70B model on four GPUs, tensor parallelism is often preferred for latency, and pipeline parallelism alone may not solve the communication bottleneck.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.