NCP-GENL GPU Acceleration and Optimization Practice Question
A team is serving a 13B-parameter LLM on a single NVIDIA A100 80GB GPU. During generation, they observe that the GPU compute utilization stays below 20% while memory bandwidth utilization is near saturation. They want to improve throughput without changing the model architecture. Which optimization is most appropriate?
⚠ Common exam trap
The trap here is assuming that low compute utilization always means the GPU needs more parallel work, when it can instead indicate a memory-bandwidth bottleneck.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply 4-bit weight-only quantization to reduce memory traffic during token generation.
The symptoms indicate a memory-bandwidth-bound workload, common in autoregressive LLM decoding. Reducing weight precision via 4-bit quantization decreases the volume of data read from GPU memory per token, directly increasing throughput. Other options either do not target memory bandwidth or introduce new bottlenecks.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Apply 4-bit weight-only quantization to reduce memory traffic during token generation.
Why this is correct
The workload is memory-bandwidth bound, as shown by low compute utilization and saturated memory bandwidth. Reducing weight precision to 4-bit lowers the bytes transferred per token, directly alleviating the bottleneck and increasing throughput without changing the architecture. This is a standard optimization for memory-bound LLM inference.
- ✗
Enable multi-GPU tensor parallelism across two A100 GPUs to split the model.
Why it's wrong here
Tensor parallelism splits layers across GPUs and adds communication overhead. While it can reduce per-GPU memory footprint, it introduces inter-GPU synchronization that can worsen latency and does not resolve the underlying memory-bandwidth saturation on a single GPU. It is not the right first step for a memory-bound single-GPU scenario.
- ✗
Increase the batch size to improve arithmetic intensity and better utilize tensor cores.
Why it's wrong here
Increasing batch size raises arithmetic intensity and can improve tensor core utilization, but in this scenario memory bandwidth is already saturated. Larger batches would further pressure memory bandwidth and may not improve throughput; they could even increase latency. This does not address the actual bottleneck.
- ✗
Increase the number of CPU threads used for token sampling to speed up generation.
Why it's wrong here
Token sampling is a lightweight operation compared to the forward pass of a 13B model. CPU-side sampling is not the bottleneck; the GPU memory bandwidth is. Adding CPU threads will not meaningfully improve throughput and may even add overhead if data transfers become less efficient.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.