Courseiva
Model Optimization →hardMultiple Choice

NCP-GENL Model Optimization Practice Question

A team has built a TensorRT-LLM engine for a 70B model on four NVIDIA H100 GPUs using tensor parallelism. They now need to serve the same model on a single H100 for a development environment, accepting higher latency. What is the most appropriate approach?

⚠ Common exam trap

The trap here is assuming a multi-GPU engine can be restricted at runtime to fewer devices, when parallelism is compiled into the plan and requires a rebuild.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Rebuild the engine with tensor parallelism set to one and apply weight quantization so the model fits in a single GPU's memory.

Tensor parallelism is a build-time property, so changing the number of GPUs requires rebuilding the engine with tensor parallel size one. Because a 70B model in FP16 exceeds a single H100's memory, weight quantization is also needed to fit. The other options either attempt to reuse an incompatible engine, misapply parallelism concepts, or increase precision, none of which yield a working single-GPU deployment.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use pipeline parallelism with a single stage to split the model across the GPU's SMs automatically.

    Why it's wrong here

    Pipeline parallelism partitions layers across devices, not within a single GPU's streaming multiprocessors. A single-stage pipeline on one GPU provides no memory benefit and does not reduce the model's footprint. This option confuses intra-GPU execution with inter-GPU model partitioning and would not make a 70B model fit in one H100.

  • ✗

    Load the existing four-GPU engine on a single H100 and set the runtime to use only one device.

    Why it's wrong here

    TensorRT engines are compiled for a specific parallel configuration; a four-way tensor-parallel plan cannot execute on one GPU. The runtime cannot simply mask devices because the engine's communication and layer partitioning are baked in. Attempting this would fail at load time or produce incorrect results, so it is not a viable path for single-GPU development.

  • ✗

    Convert the engine to a TensorRT plan file with FP32 precision so it can use unified memory on the H100.

    Why it's wrong here

    FP32 doubles weight memory compared with FP16, making a 70B model even less likely to fit on one H100. Unified memory does not eliminate the capacity limit and would introduce severe performance penalties through host-device transfers. This approach moves in the wrong direction for memory reduction and is not a standard TensorRT-LLM deployment strategy.

  • ✓

    Rebuild the engine with tensor parallelism set to one and apply weight quantization so the model fits in a single GPU's memory.

    Why this is correct

    Tensor parallelism is fixed at engine build time, so moving from four GPUs to one requires rebuilding the engine with a tensor parallel size of one. A 70B FP16 model needs roughly 140GB, which exceeds a single H100's 80GB, so weight quantization is necessary to fit. This approach produces a valid single-GPU engine while accepting the expected latency increase.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.