NCP-GENL Model Optimization Practice Question
An engineer is optimizing an LLM for inference with NVIDIA TensorRT-LLM and wants to reduce both memory footprint and latency without retraining the model. Which two techniques should they apply? (Choose two.)
⚠ Common exam trap
The trap here is selecting quantization-aware training because it sounds like a precision optimization, when it violates the no-retraining requirement.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Quantize model weights to INT8 or FP8 using post-training quantization with calibration.
Post-training quantization reduces weight memory and enables faster low-precision kernels, while paged KV cache manages generation-time memory efficiently and improves batching. Together they lower footprint and latency without retraining. The rejected options either require retraining, increase model size, or break the architecture, so they do not fit the no-retraining constraint or the optimization goals.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Quantize model weights to INT8 or FP8 using post-training quantization with calibration.
Why this is correct
Post-training weight quantization reduces the bytes per parameter, cutting memory footprint and enabling faster low-precision GEMM kernels on supported GPUs. It requires only calibration data, not retraining, which matches the constraint. This directly lowers both memory and latency, making it a primary optimization for production LLM inference in TensorRT-LLM.
- ✓
Enable paged KV cache with block-based memory management.
Why this is correct
Paged KV cache allocates memory in fixed-size blocks, reducing fragmentation and allowing more concurrent sequences to share the cache efficiently. It lowers peak memory pressure and improves batching, which in turn raises GPU utilization and reduces latency. It works alongside quantization and does not require retraining, so it is a core TensorRT-LLM optimization.
- ✗
Disable attention and rely solely on feed-forward layers to speed up decoding.
Why it's wrong here
Attention is essential to transformer language models; removing it would destroy the model's ability to use context and produce coherent output. This is not a supported optimization and would require retraining a fundamentally different architecture. It does not reduce memory or latency in a meaningful, production-valid way for an existing LLM.
- ✗
Increase the number of transformer layers to improve parallelism across the GPU.
Why it's wrong here
Adding transformer layers increases model size, memory consumption, and compute, which is the opposite of the stated optimization goals. It would also change the model architecture and require retraining. This option confuses model capacity scaling with inference optimization and would degrade both memory and latency rather than improve them.
- ✗
Retrain the model with quantization-aware training to learn optimal scaling factors.
Why it's wrong here
Quantization-aware training can improve accuracy but explicitly requires retraining, which the scenario forbids. While it is a valid technique in other contexts, it does not satisfy the no-retraining constraint here. Post-training quantization and paged KV cache deliver memory and latency benefits without modifying model weights through training.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.