NCP-GENL GPU Acceleration and Optimization Practice Question
An engineer is optimizing a large language model for inference on NVIDIA GPUs using TensorRT-LLM. They want to reduce the memory footprint of the KV cache to support longer context lengths and more concurrent requests. Which two techniques should they implement? (Choose two.)
⚠ Common exam trap
Candidates often confuse batch size increases with memory savings, when larger batches actually increase total KV cache memory.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable paged KV cache with block-based memory allocation.
INT8 quantization and paged KV cache are both designed to reduce memory footprint. Quantization lowers the bit-width of stored keys/values, while paging eliminates fragmentation and allows more efficient allocation. Together, they enable longer contexts and higher concurrency without increasing GPU memory.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use FP32 precision for the KV cache to avoid accuracy loss.
Why it's wrong here
FP32 precision doubles the memory footprint compared to FP16, making it the opposite of what is needed. While it may offer higher accuracy, it severely limits context length and concurrency. For memory optimization, lower precision is preferred, not higher.
- ✗
Enable multi-head attention with larger head dimension.
Why it's wrong here
Increasing head dimension increases the size of the KV cache per token, as each head stores its own keys and values. This would increase memory usage, not reduce it. The head dimension is typically fixed by the model architecture and not a tunable optimization for memory footprint.
- ✓
Enable paged KV cache with block-based memory allocation.
Why this is correct
Paged KV cache divides the cache into fixed-size blocks, eliminating fragmentation and allowing non-contiguous storage. This enables more efficient memory utilization and supports a larger number of concurrent sequences. It is a core feature of TensorRT-LLM for high-throughput serving, directly addressing memory footprint and scalability for long contexts.
- ✓
Use INT8 quantization for the KV cache.
Why this is correct
INT8 quantization reduces the KV cache memory footprint by storing keys and values in 8-bit integers instead of 16-bit floats, cutting memory usage in half. This allows longer context lengths and more concurrent sequences. TensorRT-LLM supports INT8 KV cache quantization with minimal accuracy loss when calibrated properly, making it a key optimization for memory-constrained inference.
- ✗
Increase the batch size to improve memory reuse.
Why it's wrong here
Increasing batch size increases the number of concurrent sequences, which actually increases total KV cache memory usage. While it can improve GPU utilization, it does not reduce per-sequence memory footprint and may lead to out-of-memory errors. This is contrary to the goal of reducing memory footprint for longer contexts.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.