NCP-GENL GPU Acceleration and Optimization Practice Question
A developer is using NVIDIA TensorRT to optimize a BERT-based model for inference. They notice that the engine performs poorly on variable-length input sequences because it was built with a single optimization profile for a fixed sequence length. What should they do to improve performance across different sequence lengths?
⚠ Common exam trap
The trap here is thinking that precision reduction or launch overhead reduction can compensate for a mismatched optimization profile, when the core issue is kernel selection for varying shapes.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Rebuild the engine with multiple optimization profiles covering the range of sequence lengths.
TensorRT optimization profiles allow the engine to tune kernels for specific input shape ranges. When input lengths vary, a single profile cannot cover all cases efficiently. Creating multiple profiles, each with appropriate min/opt/max dimensions, enables the engine to choose the best kernels for each length, improving overall performance.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Pad all input sequences to the maximum length supported by the model.
Why it's wrong here
Padding to a fixed maximum length wastes computation on padding tokens and increases latency for shorter sequences. It does not leverage TensorRT's dynamic shape capabilities and can significantly reduce throughput, especially when sequence lengths vary widely.
- ✓
Rebuild the engine with multiple optimization profiles covering the range of sequence lengths.
Why this is correct
TensorRT optimization profiles define the min, opt, and max shapes for dynamic input dimensions. Building multiple profiles allows the engine to select the best kernel configurations for different sequence lengths, improving performance across the range. This is the recommended approach for variable-length inputs.
- ✗
Use CUDA graphs to capture and replay the inference execution.
Why it's wrong here
CUDA graphs reduce kernel launch overhead but do not change how TensorRT selects kernels for different input shapes. If the engine is not optimized for the actual sequence lengths, CUDA graphs will replay the same suboptimal kernels, yielding limited benefit.
- ✗
Enable FP16 precision to reduce memory usage and speed up kernels.
Why it's wrong here
FP16 precision can improve performance, but it does not address the mismatch between the engine's fixed optimization profile and the actual variable sequence lengths. Without proper profiles, the engine may still use suboptimal kernels for shapes outside the optimized range.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.