Courseiva
Model Optimization →hardMultiple Choice

NCP-GENL Model Optimization Practice Question

An engineer is using NVIDIA TensorRT-LLM's in-flight batching to serve a mix of short and very long prompts. They observe that GPU utilization drops and latency for short requests spikes whenever a long prompt is admitted. Which mechanism should they tune to prevent long sequences from monopolizing the batch?

⚠ Common exam trap

The trap here is assuming that more GPUs or larger batches fix latency fairness, when the real constraint is per-iteration token budget allocation within the scheduler.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Adjust the maximum number of tokens processed per iteration and the KV cache block allocation policy to limit how much context a single request can consume.

In-flight batching improves throughput by mixing prefill and decode work, but a very long prompt can consume the entire per-iteration token budget, starving shorter sequences. Limiting tokens per iteration and controlling KV cache block allocation prevents any single request from monopolizing the batch. This restores GPU utilization and keeps short-request latency stable while still allowing long prompts to complete progressively.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Adjust the maximum number of tokens processed per iteration and the KV cache block allocation policy to limit how much context a single request can consume.

    Why this is correct

    In-flight batching processes a token budget per iteration; if one long prompt consumes most of that budget, short requests stall. Tuning the per-iteration token limit and KV cache block allocation constrains how much context a single sequence can occupy, allowing the scheduler to interleave short and long requests fairly. This directly targets the observed utilization drop and latency spike without changing model precision or hardware.

  • ✗

    Enable tensor parallelism so the long prompt's attention computation is split across multiple GPUs.

    Why it's wrong here

    Tensor parallelism distributes model layers across GPUs but does not reduce the token budget a single long sequence consumes within an iteration. The bottleneck here is scheduling fairness, not per-GPU compute capacity. Adding GPUs increases hardware cost and communication overhead while leaving the long prompt's dominance of the batch unchanged, so it fails to address the short-request latency spike.

  • ✗

    Switch from in-flight batching to a static batching scheme that groups requests by similar prompt length.

    Why it's wrong here

    Static batching groups requests but introduces head-of-line blocking and idle time while waiting for batches to fill, which typically worsens latency for short requests. It also discards the dynamic interleaving benefit that in-flight batching provides. The scenario needs finer control over token allocation within the existing scheduler, not a regression to a less responsive batching strategy.

  • ✗

    Increase the maximum batch size so more short requests can be admitted alongside the long prompt.

    Why it's wrong here

    Raising batch size does not help if the per-iteration token budget is already saturated by the long sequence. The scheduler may still be unable to admit additional requests without exceeding the token limit, and larger batches can increase memory pressure. This option misdiagnoses the problem as insufficient concurrency rather than unfair token allocation, so it would not resolve the latency spike.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.