NCP-GENL Model Deployment Practice Question
A team is deploying a 70B parameter LLM with NVIDIA Triton Inference Server across four NVIDIA H100 GPUs. They are using TensorRT-LLM and need to fit the model within the combined GPU memory while maintaining high throughput. Which two techniques should they use? (Choose two.)
⚠ Common exam trap
Candidates often confuse data parallelism, which replicates the full model per GPU, with model parallelism, which actually splits weights to fit a large model across devices.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Pipeline parallelism, which assigns entire transformer layers to different GPUs and passes activations between stages.
Tensor parallelism and pipeline parallelism are the two model-parallel techniques that split a 70B model across multiple GPUs, allowing the weights to fit in aggregate memory on four H100s. Tensor parallelism shards layers and attention heads, while pipeline parallelism assigns layer groups to stages; together they enable large-model deployment with high throughput.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Pipeline parallelism, which assigns entire transformer layers to different GPUs and passes activations between stages.
Why this is correct
Pipeline parallelism assigns groups of layers to different GPUs, so each GPU stores only a portion of the model. Combined with tensor parallelism, it allows very large models such as 70B parameters to fit across four H100 GPUs. It is commonly used together with tensor parallelism to balance memory and communication overhead in multi-GPU TensorRT-LLM deployments.
- ✗
Increasing the maximum batch size to the largest value the backend accepts, so memory is fully utilized.
Why it's wrong here
Raising the maximum batch size increases KV cache and activation memory, which can cause out-of-memory errors rather than helping the model fit. It does not distribute weights across GPUs and therefore does not address the core requirement of fitting a 70B model on four H100s. Batch size tuning is a throughput optimization after the model already fits.
- ✗
Data parallelism, which replicates the full model on every GPU and splits incoming requests across replicas.
Why it's wrong here
Data parallelism replicates the entire model on each GPU, which requires each GPU to hold the full 70B weights. That does not solve the problem of fitting the model across four GPUs and would exceed single-GPU memory. It is useful for scaling throughput once the model already fits per GPU, but it is the wrong choice when aggregate memory is the constraint.
- ✗
Enabling FP32 precision for all weights and activations to maximize numerical stability across GPUs.
Why it's wrong here
FP32 precision uses the most memory per parameter and would make it harder to fit a 70B model across four GPUs. It does not provide a distribution mechanism across GPUs and conflicts with the goal of fitting the model in aggregate memory. Reduced precision such as FP16, BF16, or FP8 is typically used instead to lower memory footprint.
- ✓
Tensor parallelism, which shards model layers and attention heads across the GPUs so each GPU holds a fraction of the weights.
Why this is correct
Tensor parallelism splits individual layers and attention heads across multiple GPUs, distributing weight and compute load. For a 70B model that does not fit on one GPU, this is a primary technique to fit the model in aggregate memory while enabling high throughput. It requires inter-GPU communication each layer, so NVLink or high-bandwidth interconnect is important for performance.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.