NCP-GENL GPU Acceleration and Optimization Practice Question
A company is running a large language model inference service on NVIDIA GPUs. They observe that GPU memory is nearly full, limiting the batch size and thus throughput. The model weights are stored in FP16, and the KV cache consumes a significant portion of memory. Which technique can reduce memory usage while maintaining model accuracy and enabling larger batch sizes?
⚠ Common exam trap
Many candidates confuse training-time memory optimizations like gradient checkpointing with inference-time techniques, or assuming that increasing precision helps when the goal is to reduce memory.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Quantize the KV cache to INT8.
The memory bottleneck is largely due to the KV cache. Quantizing it to INT8 halves its memory footprint, freeing space for larger batch sizes and improving throughput. This technique is supported in frameworks like TensorRT-LLM and maintains accuracy with proper calibration. Other options either do not apply to inference or increase memory usage.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Quantize the KV cache to INT8.
Why this is correct
Quantizing the KV cache to INT8 reduces its memory footprint by half compared to FP16, allowing larger batch sizes. Modern techniques like NVIDIA's TensorRT-LLM support INT8 KV cache quantization with minimal accuracy loss. This directly addresses the memory bottleneck while preserving model accuracy, making it an effective solution for increasing throughput.
- ✗
Increase the number of GPUs to distribute the model.
Why it's wrong here
Distributing the model across GPUs can reduce per-GPU memory usage, but it introduces communication overhead and may not improve throughput linearly. Moreover, it requires additional hardware and complexity. The question asks for a technique to reduce memory usage on existing GPUs; scaling out is a hardware solution, not a software optimization for the given constraint.
- ✗
Convert the model to FP32 to improve numerical stability.
Why it's wrong here
Converting to FP32 doubles memory usage for weights and activations, worsening the memory bottleneck. It does not help with the KV cache memory issue. While FP32 can improve accuracy, the model is already in FP16 and presumably accurate enough. This change would reduce batch size further, contradicting the goal.
- ✗
Use gradient checkpointing during inference.
Why it's wrong here
Gradient checkpointing is a training technique to reduce memory by recomputing activations during backward pass. It is not applicable to inference, where no backward pass occurs. Applying it during inference would add unnecessary overhead without reducing memory for weights or KV cache. Thus, it is not a valid solution for this scenario.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.