NCP-GENL Model Deployment Practice Question
An inference team runs a TensorRT-LLM model behind NVIDIA Triton Inference Server with the tensorrtllm_backend. They notice that when clients send requests with widely varying prompt lengths, throughput drops and some requests wait much longer than others. They want Triton to group compatible requests dynamically to improve GPU utilization without changing the model build. Which Triton feature should they configure?
⚠ Common exam trap
Many exam-takers confuse sequence batching, which maintains state for ordered requests, with dynamic batching, which merges independent concurrent requests.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Dynamic batching with a preferred batch size and max queue delay in the model's config.pbtxt
Triton's dynamic batching groups independent requests arriving within max_queue_delay into a single batch, using preferred_batch_size to target efficient batch shapes. This improves GPU utilization under varying prompt lengths without rebuilding the TensorRT-LLM engine, directly addressing the throughput and latency variance described.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Model warmup with sample inputs to pre-allocate memory at load time
Why it's wrong here
Warmup runs sample inferences at model load to trigger memory allocation and CUDA graph capture, reducing first-request latency. It does not affect steady-state batching behavior or throughput under mixed prompt lengths. It is unrelated to the dynamic grouping of concurrent requests the team needs.
- ✓
Dynamic batching with a preferred batch size and max queue delay in the model's config.pbtxt
Why this is correct
Dynamic batching lets Triton combine independent requests arriving within a time window into a single batch, improving GPU utilization when request sizes vary. Setting preferred_batch_size and max_queue_delay controls the tradeoff between latency and throughput, and it requires no model rebuild.
- ✗
Rate limiting with a max queue size to reject excess requests
Why it's wrong here
Rate limiting and queue size caps protect the server from overload but do not improve GPU utilization or merge compatible requests. They would instead shed load, which is the opposite of the goal of increasing throughput for mixed-length prompts. This setting does not address the batching behavior needed.
- ✗
Sequence batching with a sequence ID to maintain state across requests
Why it's wrong here
Sequence batching is designed for stateful models that process ordered requests over time, such as streaming or recurrent workloads. It does not combine independent, unrelated client requests to raise throughput. Using it here would not address the varying prompt lengths or the idle GPU cycles between requests.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.