Courseiva
Model Deployment →hardMultiple Select

NCP-GENL Model Deployment Practice Question

A platform team is deploying a 70B-parameter LLM with NVIDIA TensorRT-LLM across four 80 GB H100 GPUs and needs to serve long-context requests efficiently. They are deciding how to combine parallelism and memory techniques in the build and runtime configuration. (Choose two.)

⚠ Common exam trap

The trap here is treating pipeline parallelism or Unified Memory as easy wins, when pipeline bubbles and host-device paging both hurt latency-sensitive long-context serving.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use tensor parallelism across the four GPUs so each GPU holds a shard of every layer's weights

Tensor parallelism shards each layer across the four H100s so a 70B model fits and serves with low latency over NVLink, while a properly sized paged KV cache prevents memory over-reservation for long contexts. Together they address both weight distribution and the dominant dynamic memory consumer in long-context LLM serving.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use tensor parallelism across the four GPUs so each GPU holds a shard of every layer's weights

    Why this is correct

    Tensor parallelism splits each layer's weight matrices across GPUs, allowing a 70B model that would not fit on one 80 GB H100 to be served across four. With NVLink-connected H100s, the all-reduce communication overhead is manageable, and it is the standard way to serve very large models with low latency.

  • ✓

    Configure the paged KV cache with a block size and max sequence length sized for the target context window

    Why this is correct

    Long-context requests make the KV cache the dominant memory consumer. Sizing the paged KV cache block size and max sequence length to the real context window prevents over-reservation and OOM while still supporting the required sequence lengths. This is essential alongside tensor parallelism for efficient long-context serving.

  • ✗

    Reduce the number of GPUs to one and enable FP8 quantization to fit the model on a single device

    Why it's wrong here

    Even with FP8 quantization, a 70B model plus KV cache for long contexts will not fit comfortably on a single 80 GB H100 while maintaining quality and throughput. Removing GPUs also eliminates the parallelism needed for low-latency serving. This choice conflicts with the stated goal of deploying across four H100s.

  • ✗

    Rely on CUDA Unified Memory to transparently page weights and KV cache between host and device

    Why it's wrong here

    Unified Memory can silently migrate pages over PCIe, causing unpredictable latency spikes during inference. For latency-sensitive LLM serving, explicit placement of weights and KV cache in device memory is required. Relying on Unified Memory would undermine the performance goals and is not the recommended approach for TensorRT-LLM deployments.

  • ✗

    Enable pipeline parallelism with a single micro-batch to minimize inter-GPU communication

    Why it's wrong here

    Pipeline parallelism with a single micro-batch leaves GPUs idle because only one stage is active at a time, creating pipeline bubbles and poor utilization. It is typically used with many micro-batches to fill the pipeline. For low-latency long-context serving on four tightly coupled H100s, this configuration is counterproductive.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.