NCA-GENL Software Development Practice Question
A developer is using the NVIDIA Triton Inference Server to deploy a TensorRT-LLM optimized model. They want to send a request with multiple prompts to be processed in a single inference call. Which Triton feature should they use?
⚠ Common exam trap
It's easy for candidates to confuse server-side dynamic batching with client-side request batching; only the latter lets you put multiple prompts into one request.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Request batching
Triton Inference Server allows clients to send a single request containing multiple input elements, which are then processed together as a batch. This is often referred to as request batching or client-side batching. For TensorRT-LLM models, the model's input tensor can have a batch dimension, so the client can populate it with multiple prompts. This reduces the number of network round trips and can improve GPU utilization.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Ensemble models
Why it's wrong here
Ensemble models allow chaining multiple models together in a pipeline, where the output of one model becomes the input to another. This is useful for multi-step inference but does not provide a way to send multiple prompts in a single request. It addresses a different use case, so it is not the correct answer.
- ✓
Request batching
Why this is correct
Triton Inference Server supports request batching, where a single client request can contain multiple input items (e.g., multiple prompts) that are processed together in one inference call. This is exactly what the developer needs to submit multiple prompts efficiently. The server will handle the batching internally if the model supports it.
- ✗
Sequence batching
Why it's wrong here
Sequence batching is used for stateful models that need to maintain state across multiple inference requests, such as in recurrent neural networks or conversational models. It does not enable sending multiple prompts in a single request. Therefore, it does not meet the requirement of processing multiple prompts in one call.
- ✗
Dynamic batching
Why it's wrong here
Dynamic batching automatically groups individual inference requests into batches on the server side to improve throughput, but it does not allow a client to submit multiple prompts in one request. The client still sends separate requests. This feature is about server-side optimization, not about client-side multi-prompt submission, so it is not the correct choice.
About these practice questions
Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.