NCP-GENL Model Optimization Practice Question
A developer is using NVIDIA TensorRT-LLM to optimize a GPT-based model for inference. They want to reduce the model's memory footprint and improve throughput without retraining. Which two techniques can be applied during the TensorRT-LLM build process to achieve these goals? (Choose two.)
⚠ Common exam trap
The trap here is assuming that any memory-reduction technique like pruning is suitable, but pruning often requires retraining and is not a standard TensorRT-LLM build option.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enabling FP8 precision for both weights and activations on Hopper GPUs
Weight-only quantization and FP8 precision are both build-time optimizations in TensorRT-LLM that reduce memory footprint and improve throughput without retraining. Weight-only quantization lowers weight precision, while FP8 leverages Hopper GPU capabilities for both weights and activations. The other options either require retraining or increase memory usage.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enabling FP8 precision for both weights and activations on Hopper GPUs
Why this is correct
FP8 precision on Hopper GPUs (e.g., H100) reduces memory usage and increases throughput by using 8-bit floating point for weights and activations. TensorRT-LLM supports FP8 quantization, which can be applied during build without retraining. This directly addresses memory footprint and throughput, making it a correct choice.
- ✗
Pruning the model by removing entire attention heads
Why it's wrong here
Pruning attention heads typically requires retraining or fine-tuning to recover accuracy, and it is not a standard build-time optimization in TensorRT-LLM. While pruning can reduce model size, it alters the architecture and may not be supported directly in the TensorRT-LLM build process without custom modifications.
- ✓
Weight-only quantization (e.g., INT8 or INT4) for linear layers
Why this is correct
Weight-only quantization reduces the memory footprint by storing weights in lower precision (e.g., INT8 or INT4) while keeping activations in higher precision. This decreases model size and can improve throughput due to reduced memory bandwidth. It does not require retraining and is supported in TensorRT-LLM, making it a valid technique to meet the objectives.
- ✗
Knowledge distillation from a larger teacher model
Why it's wrong here
Knowledge distillation involves training a smaller student model to mimic a larger teacher, which requires retraining and is not a build-time optimization. TensorRT-LLM focuses on inference optimization without retraining. While distillation can reduce model size, it is not applicable here because the developer wants to avoid retraining.
- ✗
Using a larger batch size to amortize memory overhead
Why it's wrong here
Increasing batch size can improve throughput but also increases memory usage for activations and KV cache, potentially exacerbating memory footprint. It does not reduce the model's static memory footprint. The goal is to reduce memory footprint, so this technique is counterproductive and not a build-time optimization for memory reduction.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.