Courseiva
Model Deployment →hardMultiple Choice

NCP-GENL Model Deployment Practice Question

An inference team runs a TensorRT-LLM model behind NVIDIA Triton Inference Server with the tensorrtllm_backend. They notice that when clients send requests with widely varying prompt lengths, throughput drops and some requests wait much longer than others. They want Triton to group compatible requests dynamically to improve GPU utilization without changing the model build. Which Triton feature should they configure?

⚠ Common exam trap

Many exam-takers confuse sequence batching, which maintains state for ordered requests, with dynamic batching, which merges independent concurrent requests.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Dynamic batching with a preferred batch size and max queue delay in the model's config.pbtxt

Triton's dynamic batching groups independent requests arriving within max_queue_delay into a single batch, using preferred_batch_size to target efficient batch shapes. This improves GPU utilization under varying prompt lengths without rebuilding the TensorRT-LLM engine, directly addressing the throughput and latency variance described.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Model warmup with sample inputs to pre-allocate memory at load time

    Why it's wrong here

    Warmup runs sample inferences at model load to trigger memory allocation and CUDA graph capture, reducing first-request latency. It does not affect steady-state batching behavior or throughput under mixed prompt lengths. It is unrelated to the dynamic grouping of concurrent requests the team needs.

  • ✓

    Dynamic batching with a preferred batch size and max queue delay in the model's config.pbtxt

    Why this is correct

    Dynamic batching lets Triton combine independent requests arriving within a time window into a single batch, improving GPU utilization when request sizes vary. Setting preferred_batch_size and max_queue_delay controls the tradeoff between latency and throughput, and it requires no model rebuild.

  • ✗

    Rate limiting with a max queue size to reject excess requests

    Why it's wrong here

    Rate limiting and queue size caps protect the server from overload but do not improve GPU utilization or merge compatible requests. They would instead shed load, which is the opposite of the goal of increasing throughput for mixed-length prompts. This setting does not address the batching behavior needed.

  • ✗

    Sequence batching with a sequence ID to maintain state across requests

    Why it's wrong here

    Sequence batching is designed for stateful models that process ordered requests over time, such as streaming or recurrent workloads. It does not combine independent, unrelated client requests to raise throughput. Using it here would not address the varying prompt lengths or the idle GPU cycles between requests.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.