Courseiva
Model Optimization →easyMultiple Choice

NCP-GENL Model Optimization Practice Question

A team is preparing a Llama-based chatbot for production and wants to reduce GPU memory and latency without retraining. They decide to apply post-training quantization. Which TensorRT-LLM workflow correctly produces an INT8 or FP8 quantized engine from an existing FP16 checkpoint?

⚠ Common exam trap

The trap here is assuming quantization can be applied to an already-built engine or selected dynamically at runtime, when it must be baked into the checkpoint and engine build.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Run the quantization toolkit to produce a quantized checkpoint, then build the TensorRT-LLM engine with the appropriate quantization flags.

Post-training quantization in TensorRT-LLM is a two-stage process: first produce a quantized checkpoint with calibration or scaling data, then build the engine with the matching quantization flags so the builder selects INT8 or FP8 kernels. This avoids retraining and yields memory and latency improvements. The other options either attempt unsupported post-build modification, rely on runtime precision switching that does not exist, or bypass the required engine compilation step.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Build the FP16 engine first, then apply a post-build conversion script that rewrites the engine plan to INT8 weights.

    Why it's wrong here

    TensorRT engines are compiled, hardware-specific plans; they cannot be rewritten after the fact to change precision. Quantization decisions affect kernel selection, tactics, and calibration, all of which are baked in during the build. There is no supported post-build script that converts an FP16 plan to INT8, so this approach would produce an invalid or non-functional engine.

  • ✓

    Run the quantization toolkit to produce a quantized checkpoint, then build the TensorRT-LLM engine with the appropriate quantization flags.

    Why this is correct

    TensorRT-LLM expects a quantized checkpoint and quantization metadata before the engine build, so the supported path is to quantize the model with the provided toolkit and then pass the quantization mode during engine construction. This preserves accuracy through calibrated scaling factors and lets the builder select INT8 or FP8 kernels. It requires no retraining and matches the stated goal of reducing memory and latency in production.

  • ✗

    Convert the checkpoint to ONNX with INT8 operators and load it directly into the TensorRT-LLM runtime without an engine build.

    Why it's wrong here

    TensorRT-LLM does not execute ONNX graphs directly at runtime; it requires a compiled TensorRT engine. While ONNX may appear in some conversion paths, INT8 operators in ONNX do not automatically become optimized TensorRT-LLM kernels. Skipping the engine build omits kernel autotuning and quantization calibration, so this would not yield a valid or performant production deployment.

  • ✗

    Enable automatic mixed precision in the runtime and let the inference server choose INT8 kernels dynamically at request time.

    Why it's wrong here

    TensorRT-LLM does not perform dynamic precision selection at request time; precision and kernels are fixed when the engine is built. Automatic mixed precision is a training concept, not an inference runtime feature here. Relying on the server to choose INT8 kernels dynamically would not reduce the model's memory footprint or deliver the expected latency gains, and it is not how TensorRT-LLM operates.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.