NCP-GENL Model Deployment Practice Question
An engineer is preparing an ensemble in NVIDIA Triton Inference Server that chains a Python preprocessing model with a TensorRT-LLM backend model. The preprocessing model must run on CPU while the LLM must run on GPU, and the ensemble must expose a single HTTP endpoint. Which configuration is required to make the ensemble execute correctly?
⚠ Common exam trap
The trap here is reaching for custom code or instance_group tuning when Triton's built-in ensemble scheduling already solves multi-model chaining.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Define each step as a separate Triton model and reference them in a platform 'ensemble' model's ensemble_scheduling block
Triton ensembles use a model with platform 'ensemble' and an ensemble_scheduling block that lists each step with model name, version, and tensor mappings. This enables chaining a CPU Python preprocessing model with a GPU TensorRT-LLM model behind a single endpoint, which is exactly the requirement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Write a custom backend that loads both models and orchestrates them in C++
Why it's wrong here
A custom backend is possible but far more complex than needed and duplicates orchestration that Triton already provides. The ensemble feature exists precisely to chain existing models without writing new backend code. For this scenario it adds maintenance burden and risk without any functional advantage over ensembles.
- ✗
Merge both models into a single TensorRT-LLM engine using a custom plugin
Why it's wrong here
TensorRT-LLM engines contain only the compiled language model graph; they cannot host arbitrary Python preprocessing logic. Building a custom plugin would be a heavy, unsupported path and would not provide the simple chaining the team needs. It fails to satisfy the requirement of running preprocessing on CPU with a single endpoint.
- ✗
Use Triton's 'instance_group' to pin both models to the same GPU and rely on default routing
Why it's wrong here
instance_group controls where a single model's instances execute; it does not chain two models or map tensors between them. Without an ensemble or business logic script, Triton has no way to pass the preprocessing output into the LLM input. This approach does not create the required pipeline.
- ✓
Define each step as a separate Triton model and reference them in a platform 'ensemble' model's ensemble_scheduling block
Why this is correct
Triton ensembles are defined by a model with platform 'ensemble', which contains an ensemble_scheduling section listing each step, its model name, version, and input/output mappings. This lets Triton route tensors between the CPU Python model and the GPU TensorRT-LLM model while exposing one endpoint.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.