An engineer is preparing an ensemble in NVIDIA Triton Inference Server that chains a Python preprocessing model with a TensorRT-LLM backend model. The preprocessing model must run on CPU while the LLM must run on GPU, and the ensemble must expose a single HTTP endpoint. Which configuration is required to make the ensemble execute correctly?
Triton ensembles are defined by a model with platform 'ensemble', which contains an ensemble_scheduling section listing each step, its model name, version, and input/output mappings. This lets Triton route tensors between the CPU Python model and the GPU TensorRT-LLM model while exposing one endpoint.
Why this answer
Triton ensembles use a model with platform 'ensemble' and an ensemble_scheduling block that lists each step with model name, version, and tensor mappings. This enables chaining a CPU Python preprocessing model with a GPU TensorRT-LLM model behind a single endpoint, which is exactly the requirement.
Exam trap
The trap here is reaching for custom code or instance_group tuning when Triton's built-in ensemble scheduling already solves multi-model chaining.