NCP-GENL GPU Acceleration and Optimization Practice Question
A team is optimizing a large language model for inference on NVIDIA GPUs. They want to reduce the memory footprint of the model to fit on a single GPU with limited VRAM. Which two techniques are most appropriate for reducing memory usage during inference? (Choose two.)
⚠ Common exam trap
The trap here is assuming that any optimization that improves inference efficiency also reduces memory; techniques like CUDA graphs and larger batches do not lower VRAM footprint.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Quantization of weights to INT8 or FP8.
Quantization and pruning directly reduce the memory required to store and execute the model. Quantization lowers the precision of weights, while pruning removes unnecessary parameters. Both techniques decrease VRAM usage and can be combined for greater effect. The other options either increase memory usage or address different bottlenecks like launch overhead.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Quantization of weights to INT8 or FP8.
Why this is correct
Quantizing weights to lower precision reduces the memory required to store the model parameters, often by 2x or 4x. This directly lowers VRAM usage and can also speed up inference on supported hardware. It is a standard technique for fitting large models on constrained GPUs, though it may require calibration to maintain accuracy.
- ✗
Using NVIDIA TensorRT with layer fusion and kernel auto-tuning.
Why it's wrong here
TensorRT optimizes the execution graph and can reduce memory overhead by fusing layers and eliminating intermediate buffers, but the primary memory savings come from precision reduction and model compression. Layer fusion alone typically yields modest memory reductions compared to quantization or pruning, and it does not shrink the stored weights.
- ✗
Enabling CUDA graphs to capture the inference step.
Why it's wrong here
CUDA graphs reduce launch overhead and CPU utilization, but they do not reduce the memory footprint of the model. They capture the same operations and memory allocations, so VRAM usage remains essentially unchanged. They are an optimization for latency and overhead, not for memory reduction.
- ✓
Pruning or sparsifying the model weights.
Why this is correct
Pruning removes redundant weights, reducing the number of parameters and thus memory footprint. Structured sparsity can be exploited by NVIDIA sparse tensor cores for additional speedups. When combined with quantization, pruning can significantly lower VRAM usage, though it may require retraining to recover accuracy.
- ✗
Increasing the batch size to improve utilization.
Why it's wrong here
Increasing batch size raises memory usage because more activations must be stored simultaneously. It does not reduce the model's memory footprint; it exacerbates VRAM pressure. While it can improve throughput, it is counterproductive when the goal is to fit within limited memory.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.