NCP-GENL Model Optimization Practice Question
An engineer is using TensorRT-LLM to serve a model that occasionally receives prompts far longer than the typical 512 tokens, up to 8K tokens. With the default engine settings, requests near 8K fail with a cache capacity error while short requests succeed. Which configuration change most directly resolves this without rebuilding for a single worst-case shape?
⚠ Common exam trap
The trap here is reaching for memory-saving features like KV cache quantization or prefill chunking when the real constraint is that the engine was built with a maximum sequence length too small for the incoming prompts.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set the maximum input length and maximum sequence length to accommodate 8K tokens and size the KV cache pool accordingly.
The failure is a capacity limit tied to the engine's configured maximum input and sequence lengths and the KV cache pool sized for them. Raising those limits to cover 8K tokens and sizing the pool to match lets long prompts reserve the blocks they need, while paged allocation means short requests still consume only a few blocks, so a single worst-case build shape is unnecessary.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Set the maximum input length and maximum sequence length to accommodate 8K tokens and size the KV cache pool accordingly.
Why this is correct
TensorRT-LLM engines have fixed maximum input and sequence length settings, and the KV cache pool must be large enough to hold the longest sequence the engine was built for. Raising these limits and sizing the pool for 8K tokens lets long prompts allocate their blocks, while shorter requests continue to use only the blocks they need.
- ✗
Enable INT8 KV cache quantization so each token's cached keys and values consume fewer bytes.
Why it's wrong here
Quantizing the KV cache does shrink bytes per token and can help fit longer contexts, but it introduces accuracy risk and still requires the engine's maximum sequence length to permit 8K tokens. If the built-in limit is below 8K, requests are rejected before cache quantization can help.
- ✗
Enable chunked context prefill so long prompts are processed in multiple smaller prefill passes.
Why it's wrong here
Chunked context prefill reduces peak activation memory during prefill for very long prompts and improves scheduling, but it does not raise the engine's maximum sequence length or enlarge the KV cache pool. The capacity error is about how many token blocks the sequence may hold, which prefill chunking alone does not change.
- ✗
Reduce the maximum batch size to 1 so all KV cache blocks are available to a single long request.
Why it's wrong here
Lowering max batch size frees pool capacity for one sequence but cripples throughput and still fails if the per-sequence block requirement exceeds the pool sized for the original maximum sequence length. The capacity error stems from the sequence-length limit and pool sizing, not from concurrent batch occupancy.
Visual reference
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.