NCP-GENL Model Optimization Practice Question
An engineer is deploying a 13B-parameter LLM with TensorRT-LLM on a single NVIDIA A100 40GB GPU. The FP16 engine requires 26GB for weights, but during generation the KV cache grows beyond remaining memory, causing out-of-memory errors. The team wants to maximize concurrent requests without retraining. Which optimization should they apply first?
⚠ Common exam trap
The trap here is assuming that weight quantization is the only lever for memory reduction, when KV cache quantization is often the more targeted fix for generation-time OOM.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Reduce the KV cache memory footprint by quantizing the KV cache to INT8 while keeping model weights in FP16.
The KV cache grows linearly with sequence length and batch size, and it is the dominant memory consumer during LLM generation after weights are loaded. Quantizing only the KV cache to INT8 halves its footprint without retraining and without changing weight precision, directly relieving the OOM. This keeps the model on the existing A100 while allowing more concurrent requests, which matches the stated goal.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Reduce the KV cache memory footprint by quantizing the KV cache to INT8 while keeping model weights in FP16.
Why this is correct
Quantizing the KV cache to INT8 halves its per-token memory, directly addressing the growth that causes OOM during generation while leaving weights untouched, so no retraining is needed. TensorRT-LLM supports INT8 KV cache with FP16 weights, and it preserves accuracy better than aggressively quantizing weights. This is the least invasive change that increases concurrent request capacity on the constrained A100.
- ✗
Increase the batch size and sequence length limits so the scheduler can pack more requests into the same memory.
Why it's wrong here
Raising batch size and sequence length limits increases KV cache memory consumption rather than reducing it, which would worsen the OOM condition. The scheduler cannot pack more requests when memory is already exhausted. This option mistakes throughput tuning for memory optimization; it does not lower the per-token or per-request KV cache footprint that is causing the failure.
- ✗
Enable tensor parallelism across two A100 GPUs to split both weights and KV cache across devices.
Why it's wrong here
Tensor parallelism would distribute memory across GPUs, but the scenario states deployment on a single A100 40GB GPU, so adding a second GPU changes the hardware premise. Even if allowed, it introduces inter-GPU communication overhead and does not reduce per-token KV cache size. The question asks for an optimization applicable to the existing single-GPU deployment, making this impractical here.
- ✗
Convert the model weights from FP16 to FP8 using post-training quantization to free memory for the KV cache.
Why it's wrong here
FP8 weight quantization does reduce weight memory, but on an A100 GPU FP8 tensor cores are not available, so the engine cannot execute FP8 matmuls efficiently and may fall back or fail to build. The scenario specifically targets KV cache growth during generation, so freeing weight memory is indirect and platform-inappropriate. This does not solve the dynamic KV cache pressure as cleanly as quantizing the cache itself.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.