NCP-GENL Model Optimization Practice Question
An engineer is optimizing a BERT-like model for inference using NVIDIA TensorRT. They want to reduce latency further by using lower precision without significant accuracy loss. Which TensorRT precision mode should they choose to enable INT8 inference while maintaining accuracy through calibration?
⚠ Common exam trap
It's easy for candidates to confuse TF32 with INT8, as TF32 is often mentioned for Tensor Core acceleration but is not an 8-bit integer format and does not use calibration.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
INT8
INT8 precision mode in TensorRT allows inference using 8-bit integers, reducing latency and memory footprint. To preserve accuracy, a calibrator computes scaling factors from a calibration dataset. The other precisions (FP32, FP16, TF32) do not provide INT8 inference and do not use calibration for quantization.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
FP16
Why it's wrong here
FP16 uses half-precision floating point, offering speedup over FP32 with minimal accuracy loss. However, it does not provide the same level of performance as INT8 for many models. The question asks for INT8 inference with calibration, so FP16 is not the correct choice. FP16 does not require calibration and is a different optimization path.
- ✗
FP32
Why it's wrong here
FP32 is the default full-precision mode. It provides the highest accuracy but does not reduce latency or memory usage compared to lower precisions. The engineer specifically wants INT8 inference, so FP32 is not appropriate. Using FP32 would not leverage the speed benefits of Tensor Cores for INT8 operations.
- ✗
TF32
Why it's wrong here
TF32 is a 19-bit format used for Tensor Core operations in FP32 mode, providing a balance between speed and accuracy on Ampere and later GPUs. It is not INT8 and does not involve calibration. TF32 is automatically used when FP32 is selected and does not reduce precision to 8-bit integers, so it does not meet the requirement.
- ✓
INT8
Why this is correct
INT8 precision mode in TensorRT enables 8-bit integer inference, which significantly reduces latency and memory usage. To maintain accuracy, TensorRT uses a calibrator to determine scaling factors from a representative dataset. This matches the engineer's goal of using INT8 with calibration to minimize accuracy loss, making it the correct choice.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.