NCP-GENL Model Deployment Practice Question
A media analytics company runs a TensorRT-LLM optimized GPT-J model on a single NVIDIA A100 80GB GPU using NVIDIA Triton Inference Server. During peak hours, request concurrency rises sharply and the team observes that the GPU is idle for long periods while waiting on host-side tokenization and detokenization. Profiling shows that CPU preprocessing and postprocessing dominate request latency. The team wants to reduce end-to-end latency without changing model weights or adding GPUs. Which Triton feature should they use?
⚠ Common exam trap
The trap here is assuming that generic batching or ensemble orchestration will fix a host-side preprocessing bottleneck, when the actual fix is moving tokenization into the TensorRT-LLM backend and using in-flight batching.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The TensorRT-LLM backend's in-flight batching and built-in tokenizer, which perform tokenization and detokenization inside the backend and keep the GPU busy with continuous batching.
The TensorRT-LLM Triton backend provides an integrated tokenizer and in-flight batching, which eliminate the host-side tokenization and detokenization stalls observed in the profile. In-flight batching also continuously admits new requests into the running batch, improving GPU utilization and reducing end-to-end latency without changing model weights or adding GPUs.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Dynamic batching, because it groups requests on the server and reduces the number of model executions.
Why it's wrong here
Dynamic batching reduces the number of forward passes by grouping requests, which helps throughput and can improve GPU utilization. However, the profiling in this scenario shows host-side tokenization and detokenization are the bottleneck, so batching does not remove that CPU work. It may even increase queueing delay if the batch timeout is too long, so it does not directly address the stated latency cause.
- ✓
The TensorRT-LLM backend's in-flight batching and built-in tokenizer, which perform tokenization and detokenization inside the backend and keep the GPU busy with continuous batching.
Why this is correct
The TensorRT-LLM backend for Triton includes an integrated tokenizer and supports in-flight batching, which moves tokenization into the backend and schedules new requests into ongoing GPU work. This directly targets the host-side bottleneck described in the profile and improves latency by keeping the GPU continuously utilized without adding hardware or changing weights.
- ✗
Sequence batching with a custom scheduler, because it keeps tokenization on the host and overlaps it with GPU compute across requests.
Why it's wrong here
Sequence batching is designed for stateful models such as those with recurrent state or KV cache management across requests. It can improve throughput for conversational workloads, but it does not move tokenization off the host. The scenario identifies host-side preprocessing and postprocessing as the dominant cost, so sequence batching alone would not eliminate that CPU bottleneck.
- ✗
A Triton ensemble that places a Python model for tokenization before the TensorRT-LLM model and a Python model for detokenization after it.
Why it's wrong here
An ensemble chains models so intermediate tensors stay in Triton and avoid network round trips, but the tokenization and detokenization still execute on the host CPU. The ensemble reduces client-side orchestration overhead, not the CPU work itself. Since profiling shows host preprocessing and postprocessing dominate, this approach would not materially reduce the identified latency source.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.