NCP-GENL Model Optimization Practice Question
An engineer is tuning a TensorRT-LLM deployment of a 7B model for a latency-sensitive API. Profiling shows that time per output token is higher than expected and that many small kernels run back to back with gaps between them. Which TWO changes are most likely to reduce the per-token latency by cutting kernel launch overhead and redundant memory traffic? (Choose two.)
⚠ Common exam trap
The trap here is chasing sequence length or precision settings when the profile clearly shows launch-bound behavior that graph capture and kernel fusion are designed to eliminate.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable CUDA Graphs so the decoding iteration is captured and replayed as a single graph launch.
The profile shows many short kernels with idle gaps, which is the classic signature of launch overhead and unfused elementwise work. CUDA Graphs collapse the decode iteration into a single replayable launch, and fused attention and GEMM plugins merge operations that would otherwise round-trip intermediates through HBM. Together they reduce both launch count and memory traffic, which are the two costs identified in the profile.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable a debug synchronization after every kernel so profiling timings are accurate.
Why it's wrong here
Forcing synchronization after each kernel serializes execution and prevents the GPU from overlapping work, which inflates measured latency dramatically. It is a diagnostic technique at best and a performance anti-pattern in production. It cannot reduce per-token time and would make the observed gaps and total latency worse.
- ✓
Enable CUDA Graphs so the decoding iteration is captured and replayed as a single graph launch.
Why this is correct
CUDA Graphs record the whole sequence of kernels in a decoding step and replay them with one launch, removing per-kernel launch latency and the CPU-side gaps visible in the profile. For a latency-sensitive decode loop with many small kernels, this directly reduces time per output token without changing numerics or requiring retraining.
- ✓
Use TensorRT-LLM fused multi-head attention and fused GEMM plugins instead of separate elementwise kernels.
Why this is correct
Fused attention and fused GEMM plugins combine several operations, such as scaling, masking, softmax, and the projection, into single kernels. That removes intermediate tensors that would otherwise be written to and read from HBM, cutting both launch count and memory traffic, which matches the small-kernel, gap-riddled profile described.
- ✗
Raise the maximum sequence length in the engine build to accommodate the longest possible prompt.
Why it's wrong here
Increasing max sequence length grows KV cache and workspace allocations and can force the builder to choose less efficient kernels for the larger shapes. It does not reduce the number of small kernels or their launch gaps; if anything, over-provisioning sequence length wastes memory and can slightly increase per-token cost for typical requests.
- ✗
Switch the model weights to FP32 to avoid any dequantization kernels in the decode path.
Why it's wrong here
FP32 doubles weight memory and HBM traffic per token compared with FP16 or quantized formats, and it does not eliminate the small-kernel pattern. Decode latency is bandwidth bound, so larger weight reads make each token slower. This change moves in the opposite direction from the goal of reducing per-token latency.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.