NCA-GENL Software Development Practice Question
A developer has a working TensorRT-LLM engine and wants to expose it through NVIDIA Triton Inference Server so that multiple client applications can call it over HTTP and gRPC with a stable interface. Which Triton feature should they configure to serve the TensorRT-LLM engine as a backend?
⚠ Common exam trap
The trap here is reaching for a custom or Python wrapper when an official Triton backend already targets TensorRT-LLM engines.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The TensorRT-LLM backend, declaring the engine in the model repository with a config.pbtxt that sets the backend and engine directory.
Triton's TensorRT-LLM backend is designed specifically to load TensorRT-LLM engines from the model repository and serve them over the standard HTTP and gRPC endpoints. Declaring the engine directory and backend in config.pbtxt gives clients a stable interface while retaining optimized in-flight batching and KV cache management. Custom backends and ensembles add complexity without providing the required serving capability.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
An ensemble model that chains a preprocessing model, the TensorRT-LLM engine, and a postprocessing model.
Why it's wrong here
Ensembles compose multiple models into one pipeline and are useful for RAG or multi-step flows, but they do not by themselves expose a TensorRT-LLM engine as a servable backend. The engine still needs to be declared as a model with the correct backend and parameters. Ensembling is an architectural overlay, not the mechanism that makes the engine callable.
- ✗
A Python backend model that loads the engine and implements the inference logic in a custom script.
Why it's wrong here
A Python backend can wrap inference code, but it bypasses the purpose-built TensorRT-LLM backend and typically loses optimized in-flight batching and KV cache management. It is a valid escape hatch for custom logic, not the standard way to expose a TensorRT-LLM engine. For a stable, optimized HTTP and gRPC interface, the dedicated backend is the appropriate choice.
- ✗
A custom C++ backend compiled against the TensorRT-LLM libraries and registered with Triton at startup.
Why it's wrong here
Writing a custom C++ backend duplicates functionality already provided by the official TensorRT-LLM backend and adds maintenance burden. It is unnecessary for exposing an existing engine over HTTP and gRPC. Unless the team needs behavior the official backend cannot provide, this approach increases risk without benefit and forfeits upstream optimizations and support.
- ✓
The TensorRT-LLM backend, declaring the engine in the model repository with a config.pbtxt that sets the backend and engine directory.
Why this is correct
Triton's TensorRT-LLM backend is purpose-built to load TensorRT-LLM engines and expose them through the standard HTTP and gRPC inference protocols. Placing the engine in the model repository and pointing config.pbtxt at the backend and engine directory gives clients a stable interface plus optimized batching and KV cache handling. This is the intended production path for serving TensorRT-LLM engines.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.