NCP-GENL Model Optimization Practice Question
A team is deploying a 13B-parameter chatbot on a single NVIDIA A10G GPU (24 GB VRAM). The model's weights are stored in FP16, and the runtime runs out of memory during KV cache allocation under concurrent user sessions. They must keep answer quality essentially unchanged while maximizing concurrent sessions. Which optimization should they apply first?
⚠ Common exam trap
The trap here is assuming that any precision reduction degrades quality unacceptably, when weight-only INT8 with FP16 activations is specifically designed to preserve accuracy while cutting the dominant memory consumer.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply INT8 weight-only quantization to the linear layers with NVIDIA TensorRT-LLM, keeping activations in FP16.
The bottleneck is persistent weight memory (FP16 weights dominate VRAM), which leaves too little room for paged KV cache blocks under concurrency. Weight-only INT8 quantization in TensorRT-LLM roughly halves the linear-layer weight footprint while keeping activations in FP16, preserving quality. That reclaimed VRAM goes directly to KV cache, raising the number of simultaneous sessions without retraining or a new GPU.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Convert the checkpoint to TF32 and rebuild the TensorRT-LLM engine with builder optimization level 5.
Why it's wrong here
TF32 is a 19-bit compute format used for matmul inputs on Ampere and later; storing weights in TF32 does not reduce their footprint versus FP16 and can even increase it. Builder optimization level 5 explores more kernels and tactics for speed, but it cannot overcome a hard VRAM shortfall during KV cache allocation.
- ✓
Apply INT8 weight-only quantization to the linear layers with NVIDIA TensorRT-LLM, keeping activations in FP16.
Why this is correct
Weight-only INT8 quantization halves the memory consumed by the model's linear-layer weights (roughly 26 GB FP16 becomes ~13 GB INT8) while leaving activation precision untouched, which preserves output quality closely. The freed VRAM is then available for KV cache blocks, directly lifting the concurrent session ceiling on the 24 GB A10G.
- ✗
Increase the paged KV cache block size from 16 to 128 tokens per block and disable block reuse.
Why it's wrong here
Larger KV cache blocks reduce the number of block-table entries but do not reduce bytes per token stored; total KV memory is unchanged. Disabling block reuse also prevents freed sequences from returning blocks to the pool, so memory pressure under concurrent sessions gets worse, not better. This does not address the weight footprint that caused the OOM.
- ✗
Enable CUDA graph capture for the decoder and set the maximum batch size to the peak observed concurrency.
Why it's wrong here
CUDA graphs reduce launch overhead and jitter but allocate the same tensors the non-graph path would; they do not shrink weights or KV cache. Pinning max batch size to peak concurrency actually reserves more KV cache up front, making the out-of-memory condition on the A10G more likely rather than resolving it.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.