NCP-GENL GPU Acceleration and Optimization Practice Question
An engineer is optimizing a large language model for inference on NVIDIA GPUs and wants to reduce memory usage to fit a larger model or increase batch size. Which two techniques are most effective for reducing GPU memory consumption during inference? (Choose two.)
⚠ Common exam trap
A common mix-up: candidates confuse training-time memory optimizations, like activation checkpointing, with inference-time memory reductions, such as weight quantization and paged attention.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Quantize model weights to 8-bit integers using TensorRT or similar tools.
Quantizing weights to INT8 reduces memory footprint by storing weights in lower precision, while paged attention optimizes KV cache memory by reducing fragmentation. Both are effective for fitting larger models or increasing batch size during inference. Activation checkpointing and micro-batching are training or throughput techniques, and CPU offloading introduces latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable activation checkpointing to recompute activations during the backward pass.
Why it's wrong here
Activation checkpointing is a training technique that trades compute for memory by recomputing activations. During inference, there is no backward pass, so it does not apply. It does not reduce memory usage for inference workloads and is irrelevant in this context.
- ✗
Increase the number of micro-batches to overlap computation and memory transfers.
Why it's wrong here
Micro-batching is used in pipeline parallelism to improve utilization, but it increases memory usage because multiple micro-batches require additional activation storage. It does not reduce memory consumption; in fact, it can increase it. Therefore, it is not a memory-reduction technique.
- ✓
Quantize model weights to 8-bit integers using TensorRT or similar tools.
Why this is correct
Quantizing weights to 8-bit integers reduces the memory footprint by up to 4x compared to FP32. This allows larger models to fit in GPU memory and enables larger batch sizes. TensorRT supports INT8 quantization with calibration to maintain accuracy, making it a standard technique for memory reduction during inference.
- ✓
Use a paged attention mechanism to manage the KV cache efficiently.
Why this is correct
Paged attention, as implemented in NVIDIA TensorRT-LLM, stores the KV cache in non-contiguous blocks, reducing fragmentation and allowing more efficient memory usage. This can significantly lower memory consumption for long sequences or large batches, enabling larger models or batch sizes.
- ✗
Store the model weights on the CPU and transfer them to the GPU on demand.
Why it's wrong here
Transferring weights from CPU to GPU on demand introduces high latency due to PCIe bandwidth limits and does not reduce GPU memory usage if weights are cached. It also complicates memory management and is not a practical solution for reducing memory footprint during inference.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.