NCP-GENL Model Optimization Practice Question
An engineer is using NVIDIA TensorRT-LLM to optimize an LLM for inference. They want to reduce the memory footprint of the KV cache during long-context generation. Which TWO techniques are supported by TensorRT-LLM to achieve this? (Choose two.)
⚠ Common exam trap
Watch out — candidates often confuse model-level optimizations like pruning or sliding window attention with runtime KV cache optimizations that TensorRT-LLM directly supports.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable paged KV cache with block sharing.
TensorRT-LLM provides native support for FP8 KV cache quantization and paged KV cache with block sharing. FP8 quantization halves cache memory, while paged cache with block sharing reduces duplication across sequences. Both are configuration-time features that do not require model changes. The other options involve model modifications or misapply TensorRT features, making them incorrect for this scenario.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use TensorRT's INT8 calibration for the KV cache.
Why it's wrong here
TensorRT's INT8 calibration is for model weights and activations, not for the KV cache. TensorRT-LLM handles KV cache quantization separately, and INT8 calibration does not apply to the cache. Thus, this is not a supported technique for reducing KV cache memory.
- ✓
Enable paged KV cache with block sharing.
Why this is correct
TensorRT-LLM's paged KV cache divides the cache into blocks and allows sharing of identical blocks across sequences, reducing memory when multiple requests share prefixes. This is a core feature for efficient memory management in long-context scenarios and is enabled by default in recent versions.
- ✗
Apply structured pruning to remove attention heads.
Why it's wrong here
Structured pruning of attention heads reduces model weights, not the KV cache memory. While it can lower overall memory, it does not target the KV cache specifically and requires retraining or fine-tuning. TensorRT-LLM does not provide built-in pruning for KV cache reduction during inference.
- ✓
Use FP8 KV cache quantization.
Why this is correct
TensorRT-LLM supports quantizing the KV cache to FP8, which halves its memory footprint compared to FP16. This is done by setting the kv_cache_dtype to 'fp8' in the build configuration. It is a native feature and does not require model retraining, making it a direct way to reduce memory during long-context generation.
- ✗
Enable sliding window attention in the model architecture.
Why it's wrong here
Sliding window attention is a model architecture modification that limits the attention span, reducing KV cache size, but it requires changing the model and retraining. TensorRT-LLM does not automatically apply it; it must be part of the model definition. Therefore, it is not a TensorRT-LLM supported technique per se.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.