Courseiva
Model Optimization →easyMultiple Choice

NCP-GENL Model Optimization Practice Question

A developer has a fine-tuned Llama-family model in Hugging Face format and wants to run it with NVIDIA TensorRT-LLM on an H100. Which artifact must be produced before the runtime can execute the model?

⚠ Common exam trap

Many exam-takers confuse an interchange or community format such as ONNX or GGUF with the compiled TensorRT engine that the TensorRT-LLM runtime actually loads and executes.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

A TensorRT engine built from the model's network definition for the target GPU architecture.

TensorRT-LLM executes a serialized TensorRT engine that is compiled for a specific GPU architecture, precision, and parallelism configuration. After converting the Hugging Face checkpoint into the TensorRT-LLM checkpoint format, the developer must build the engine for the H100. Only that engine can be loaded by the runtime, so it is the required artifact before any inference can occur.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    A TensorRT engine built from the model's network definition for the target GPU architecture.

    Why this is correct

    TensorRT-LLM compiles the model graph into a serialized TensorRT engine that is specialized for the target GPU compute capability, chosen precision, and parallelism layout. The runtime loads and executes this engine; without it there is nothing for the executor to run. Building the engine is therefore the mandatory step between a Hugging Face checkpoint and inference.

  • ✗

    An ONNX graph exported with dynamic axes and a matching runtime configuration JSON.

    Why it's wrong here

    ONNX can be an intermediate representation in some TensorRT pipelines, and TensorRT-LLM uses an internal network definition rather than requiring a user-supplied ONNX file. Producing an ONNX graph alone does not yield something the TensorRT-LLM runtime can execute; the engine must still be built for the target GPU.

  • ✗

    A PyTorch TorchScript trace of the full forward pass saved as a .pt file.

    Why it's wrong here

    TorchScript is a PyTorch deployment format and is not consumed by the TensorRT-LLM runtime. While some workflows use tracing internally during conversion, the artifact the runtime loads is a TensorRT engine, not a TorchScript module. Relying on a .pt file would leave the model unable to execute on the TensorRT-LLM executor.

  • ✗

    A quantized GGUF file generated with a community conversion script.

    Why it's wrong here

    GGUF is a format used by llama.cpp-style runtimes, not by NVIDIA TensorRT-LLM. Supplying a GGUF file would not be recognized by the TensorRT-LLM build or runtime path. The developer needs to convert the Hugging Face checkpoint into a TensorRT-LLM checkpoint and then compile a TensorRT engine for the H100.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.