NCP-GENL Model Optimization Practice Question
A team is building a TensorRT-LLM engine for a 7B model that must serve both single-turn short prompts and long multi-turn conversations with a shared system prompt. They want to maximize reuse of computation across requests without changing model weights. Which TWO techniques should they enable? (Choose two.)
⚠ Common exam trap
The trap here is conflating weight memory with KV cache memory, so an engineer reaches for weight quantization when the reuse problem is really about activations and cache block lifecycle.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable paged KV cache with block reuse so freed sequence blocks return to the pool for other requests.
Sharing a system prompt and multi-turn history means many requests share long identical prefixes, and mixed short and long prompts create variable cache pressure. KV cache reuse skips recomputing cached prefixes, while the paged KV cache with block reuse keeps memory defragmented and available to new sequences. Together they cut redundant prefill and raise concurrency without touching model weights.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable paged KV cache with block reuse so freed sequence blocks return to the pool for other requests.
Why this is correct
The paged KV cache allocates fixed-size blocks and, with block reuse, returns blocks from finished sequences to a shared pool rather than fragmenting memory. Mixed short and long prompts then coexist efficiently, raising the effective batch size and preventing allocation failures without any change to the underlying model weights.
- ✗
Apply INT4 weight-only quantization to the attention projection matrices to shrink the cache footprint.
Why it's wrong here
Quantizing attention projection weights reduces model weight memory, not KV cache memory, since cached keys and values are activations computed at runtime. It also alters numerics and requires accuracy validation. The scenario asks for computation reuse across requests without changing weights, which weight quantization does not provide.
- ✓
Enable KV cache reuse so identical prompt prefixes share cached key/value blocks.
Why this is correct
TensorRT-LLM's KV cache reuse lets a new request whose prompt prefix matches an existing cached sequence skip recomputing those tokens' key/value states. For multi-turn chats and a shared system prompt, this eliminates redundant prefill work and directly cuts time-to-first-token while leaving model weights untouched.
- ✗
Set the builder to use strongly typed engines so the plugin graph is fully deterministic.
Why it's wrong here
Strongly typed mode controls whether TensorRT may use lower-precision tactics at runtime; it constrains precision choices for determinism. It has no bearing on reusing prompt-prefix computation or on KV cache block lifecycle, so it does not address the goal of sharing work across requests with differing prompt lengths.
- ✗
Increase the beam width so multiple candidate continuations are evaluated per request.
Why it's wrong here
Beam search expands the number of sequences decoded per request, multiplying KV cache consumption and compute rather than reusing it. It improves output quality for some tasks but directly conflicts with serving many concurrent short and long prompts, and it does nothing to share prefix computation across users.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.