Courseiva
Model Deployment →mediumMultiple Select

NCP-GENL Model Deployment Practice Question

A team is optimizing an NVIDIA TensorRT-LLM deployment of a 70B model on multiple GPUs. They want to reduce inter-GPU communication overhead and improve throughput. Which two techniques should they consider? (Choose two.)

⚠ Common exam trap

The trap here is treating memory-management knobs like KV cache block size as if they also reduce inter-GPU communication overhead.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable pipeline parallelism to split layers into stages.

Tensor parallelism benefits from NVLink's high bandwidth to reduce all-reduce overhead, while pipeline parallelism lowers communication frequency by staging layers across GPUs. Together they address the communication bottleneck for large multi-GPU models. KV cache block size, disabling in-flight batching, and FP32 precision do not reduce inter-GPU communication and can hurt throughput.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Enable pipeline parallelism to split layers into stages.

    Why this is correct

    Pipeline parallelism assigns different layer stages to different GPUs, reducing the frequency of cross-GPU communication compared with tensor parallelism. With micro-batching, stages can overlap and keep GPUs busy. This lowers communication overhead per token and can improve throughput when combined with appropriate batch scheduling.

  • ✗

    Disable in-flight batching to simplify scheduling.

    Why it's wrong here

    Disabling in-flight batching reduces the number of concurrent sequences processed together, which typically lowers throughput. It does not reduce inter-GPU communication and may worsen GPU utilization. This change works against the goal of improving throughput in a multi-GPU 70B deployment.

  • ✗

    Use FP32 precision for all weights to improve numerical stability.

    Why it's wrong here

    FP32 weights double memory usage and increase data transfer volume between GPUs, raising communication overhead. It does not improve throughput and can make the deployment infeasible for a 70B model. Lower precision formats like FP16 or BF16 are preferred to reduce both memory and communication costs.

  • ✓

    Use tensor parallelism with NVLink-connected GPUs.

    Why this is correct

    Tensor parallelism splits layers across GPUs and requires frequent all-reduce communication. NVLink provides high bandwidth and low latency between GPUs, reducing the communication bottleneck. On systems without NVLink, this overhead can dominate, so pairing tensor parallelism with NVLink is a standard optimization for large models like a 70B deployment.

  • ✗

    Increase the KV cache block size to 256 tokens.

    Why it's wrong here

    KV cache block size affects memory paging granularity for attention, not inter-GPU communication. Larger blocks may reduce fragmentation but do not change how often GPUs exchange activations. This setting is orthogonal to the communication overhead the team wants to reduce.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.