NCP-GENL Model Deployment Practice Question
An engineer needs to expose an LLM served by NVIDIA Triton Inference Server to a web application over HTTP with token streaming. Which Triton feature should they enable?
⚠ Common exam trap
Many candidates confuse throughput optimizations such as dynamic batching with the response-streaming capability required for token-by-token delivery.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Decoupled mode with streaming responses in the model backend.
Triton's decoupled mode lets a backend emit multiple responses for one request, which is the mechanism used for LLM token streaming. Configuring the model for decoupled transactions and using a streaming-capable client over HTTP or gRPC delivers tokens incrementally. Other Triton features like batching, ensembles, or instance groups improve throughput or composition but do not stream partial outputs.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Model ensembles combining preprocessing and postprocessing models.
Why it's wrong here
Ensembles chain multiple models into a pipeline but do not enable streaming of partial outputs to clients. They are useful for composing preprocessing, inference, and postprocessing steps, yet the response is still returned as a complete result. This does not satisfy the token-streaming requirement for a web application.
- ✗
Dynamic batching in the model configuration.
Why it's wrong here
Dynamic batching groups multiple inference requests into a single batch to improve throughput. It does not provide token-by-token streaming over HTTP. While it improves server efficiency, it operates at the request scheduling layer and does not affect how responses are delivered to the client incrementally.
- ✗
Instance groups with multiple GPU instances per model.
Why it's wrong here
Instance groups control how many model instances run on available GPUs to scale throughput. They do not change the response delivery mechanism and cannot produce incremental token streams. This setting affects concurrency and resource allocation, not the streaming behavior needed for a chat-style web interface.
- ✓
Decoupled mode with streaming responses in the model backend.
Why this is correct
Decoupled mode allows a model backend to return multiple responses for a single request, which is essential for token streaming in LLMs. Triton's HTTP and gRPC endpoints support streaming when the model is configured for decoupled transactions. This lets the web application receive partial outputs as tokens are generated, improving perceived latency.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.