NCP-GENL Model Optimization Practice Question
An engineer is using NVIDIA TensorRT-LLM's in-flight batching to serve a mix of short and very long prompts. They observe that GPU utilization drops and latency for short requests spikes whenever a long prompt is admitted. Which mechanism should they tune to prevent long sequences from monopolizing the batch?
⚠ Common exam trap
The trap here is assuming that more GPUs or larger batches fix latency fairness, when the real constraint is per-iteration token budget allocation within the scheduler.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Adjust the maximum number of tokens processed per iteration and the KV cache block allocation policy to limit how much context a single request can consume.
In-flight batching improves throughput by mixing prefill and decode work, but a very long prompt can consume the entire per-iteration token budget, starving shorter sequences. Limiting tokens per iteration and controlling KV cache block allocation prevents any single request from monopolizing the batch. This restores GPU utilization and keeps short-request latency stable while still allowing long prompts to complete progressively.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Adjust the maximum number of tokens processed per iteration and the KV cache block allocation policy to limit how much context a single request can consume.
Why this is correct
In-flight batching processes a token budget per iteration; if one long prompt consumes most of that budget, short requests stall. Tuning the per-iteration token limit and KV cache block allocation constrains how much context a single sequence can occupy, allowing the scheduler to interleave short and long requests fairly. This directly targets the observed utilization drop and latency spike without changing model precision or hardware.
- ✗
Enable tensor parallelism so the long prompt's attention computation is split across multiple GPUs.
Why it's wrong here
Tensor parallelism distributes model layers across GPUs but does not reduce the token budget a single long sequence consumes within an iteration. The bottleneck here is scheduling fairness, not per-GPU compute capacity. Adding GPUs increases hardware cost and communication overhead while leaving the long prompt's dominance of the batch unchanged, so it fails to address the short-request latency spike.
- ✗
Switch from in-flight batching to a static batching scheme that groups requests by similar prompt length.
Why it's wrong here
Static batching groups requests but introduces head-of-line blocking and idle time while waiting for batches to fill, which typically worsens latency for short requests. It also discards the dynamic interleaving benefit that in-flight batching provides. The scenario needs finer control over token allocation within the existing scheduler, not a regression to a less responsive batching strategy.
- ✗
Increase the maximum batch size so more short requests can be admitted alongside the long prompt.
Why it's wrong here
Raising batch size does not help if the per-iteration token budget is already saturated by the long sequence. The scheduler may still be unable to admit additional requests without exceeding the token limit, and larger batches can increase memory pressure. This option misdiagnoses the problem as insufficient concurrency rather than unfair token allocation, so it would not resolve the latency spike.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.