NCP-GENL GPU Acceleration and Optimization Practice Question
An engineer is deploying a large language model using NVIDIA TensorRT-LLM on an A100 GPU. They want to maximize throughput for a chatbot workload with variable-length inputs and outputs. Which of the following techniques should they implement to achieve the highest throughput while maintaining acceptable latency?
⚠ Common exam trap
The trap here is assuming that increasing batch size or GPUs alone will maximize throughput, when dynamic batching and memory management are the real enablers.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable in-flight batching (continuous batching) and use paged KV cache with block-based memory management.
In-flight batching and paged KV cache are key optimizations in TensorRT-LLM for LLM serving. In-flight batching allows sequences to join and leave the batch at each decoding step, avoiding padding and keeping the GPU busy. Paged KV cache manages memory in blocks, enabling more sequences to run concurrently. This combination delivers the highest throughput for variable-length chatbot workloads.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable in-flight batching (continuous batching) and use paged KV cache with block-based memory management.
Why this is correct
In-flight batching dynamically adds and removes sequences at each iteration, eliminating padding and improving GPU utilization. Paged KV cache stores attention keys and values in non-contiguous blocks, reducing memory fragmentation and allowing more concurrent sequences. Together, they maximize throughput for variable-length chatbot traffic while keeping latency low, as implemented in TensorRT-LLM.
- ✗
Use static batching with a fixed batch size of 32 and pad all sequences to the maximum length.
Why it's wrong here
Static batching with padding wastes compute on padding tokens, especially with variable-length sequences. It also cannot adapt to fluctuating request rates, leading to either underutilization or increased latency. While simple, it does not maximize throughput for a chatbot workload where sequence lengths vary widely. This approach is inefficient and outdated for LLM serving.
- ✗
Use FP32 precision to ensure accuracy and rely on CUDA graphs to reduce launch overhead.
Why it's wrong here
FP32 doubles memory footprint and reduces computational throughput compared to FP16 or INT8, directly limiting the number of concurrent sequences and overall throughput. CUDA graphs reduce kernel launch overhead but do not solve batching inefficiency. For a chatbot, FP32 is unnecessary and counterproductive for maximizing throughput.
- ✗
Increase the number of GPUs and use tensor parallelism with a fixed batch size per GPU.
Why it's wrong here
Tensor parallelism can reduce per-GPU memory and compute load, but it introduces communication overhead and does not inherently improve throughput for variable-length sequences. Without dynamic batching, GPUs may still be underutilized due to padding and idle time between requests. This adds complexity and cost without addressing the core inefficiency.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.