Courseiva
Model Deployment →easyMultiple Choice

NCP-GENL Model Deployment Practice Question

An engineer is preparing an ensemble in NVIDIA Triton Inference Server that chains a Python preprocessing model with a TensorRT-LLM backend model. The preprocessing model must run on CPU while the LLM must run on GPU, and the ensemble must expose a single HTTP endpoint. Which configuration is required to make the ensemble execute correctly?

⚠ Common exam trap

The trap here is reaching for custom code or instance_group tuning when Triton's built-in ensemble scheduling already solves multi-model chaining.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Define each step as a separate Triton model and reference them in a platform 'ensemble' model's ensemble_scheduling block

Triton ensembles use a model with platform 'ensemble' and an ensemble_scheduling block that lists each step with model name, version, and tensor mappings. This enables chaining a CPU Python preprocessing model with a GPU TensorRT-LLM model behind a single endpoint, which is exactly the requirement.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Write a custom backend that loads both models and orchestrates them in C++

    Why it's wrong here

    A custom backend is possible but far more complex than needed and duplicates orchestration that Triton already provides. The ensemble feature exists precisely to chain existing models without writing new backend code. For this scenario it adds maintenance burden and risk without any functional advantage over ensembles.

  • ✗

    Merge both models into a single TensorRT-LLM engine using a custom plugin

    Why it's wrong here

    TensorRT-LLM engines contain only the compiled language model graph; they cannot host arbitrary Python preprocessing logic. Building a custom plugin would be a heavy, unsupported path and would not provide the simple chaining the team needs. It fails to satisfy the requirement of running preprocessing on CPU with a single endpoint.

  • ✗

    Use Triton's 'instance_group' to pin both models to the same GPU and rely on default routing

    Why it's wrong here

    instance_group controls where a single model's instances execute; it does not chain two models or map tensors between them. Without an ensemble or business logic script, Triton has no way to pass the preprocessing output into the LLM input. This approach does not create the required pipeline.

  • ✓

    Define each step as a separate Triton model and reference them in a platform 'ensemble' model's ensemble_scheduling block

    Why this is correct

    Triton ensembles are defined by a model with platform 'ensemble', which contains an ensemble_scheduling section listing each step, its model name, version, and input/output mappings. This lets Triton route tensors between the CPU Python model and the GPU TensorRT-LLM model while exposing one endpoint.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.