NCP-GENL Model Optimization Practice Question
An engineer is using NVIDIA TensorRT to optimize a Transformer model for inference on an NVIDIA A100 GPU. They want to maximize throughput while ensuring that the model runs correctly with varying input sequence lengths. Which TensorRT feature should they configure to allow the engine to handle different input shapes at runtime?
⚠ Common exam trap
Watch out — candidates often confuse quantization (INT8 calibration) with shape flexibility, but calibration only affects precision, not the ability to accept different input dimensions.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Dynamic shapes with optimization profiles
Dynamic shapes with optimization profiles enable a single TensorRT engine to handle inputs of different sizes by defining shape ranges. This allows the engine to optimize for the actual input shape at runtime, which is crucial for Transformer models with variable sequence lengths. Static shapes, multiple engines, or INT8 calibration do not provide this flexibility.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Static shapes with fixed dimensions
Why it's wrong here
Static shapes require the input dimensions to be fixed at build time. If the actual input shape differs, inference fails. This does not support varying sequence lengths, so it is not suitable for the engineer's requirement. Static shapes are simpler but inflexible for dynamic workloads.
- ✓
Dynamic shapes with optimization profiles
Why this is correct
Dynamic shapes with optimization profiles allow a TensorRT engine to accept input tensors of different dimensions at runtime. The profiles define minimum, optimal, and maximum shapes for each input, enabling the engine to select the best kernel for the actual shape. This is essential for handling varying sequence lengths in Transformer models without rebuilding the engine.
- ✗
INT8 calibration with a representative dataset
Why it's wrong here
INT8 calibration is a quantization technique to reduce precision and improve performance, but it does not address input shape flexibility. Calibration is orthogonal to dynamic shapes. The engineer needs to handle varying sequence lengths, which requires dynamic shape support, not calibration.
- ✗
Multiple engines for each possible input length
Why it's wrong here
Building separate engines for each input length is impractical and resource-intensive. It would require predicting all possible lengths and managing multiple engines, increasing complexity and memory usage. TensorRT's dynamic shapes feature is designed to avoid this by using a single engine with profiles.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.