Courseiva
Model Deployment →easyMultiple Choice

NCP-GENL Model Deployment Practice Question

An engineer needs to expose an LLM served by NVIDIA Triton Inference Server to a web application over HTTP with token streaming. Which Triton feature should they enable?

⚠ Common exam trap

Many candidates confuse throughput optimizations such as dynamic batching with the response-streaming capability required for token-by-token delivery.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Decoupled mode with streaming responses in the model backend.

Triton's decoupled mode lets a backend emit multiple responses for one request, which is the mechanism used for LLM token streaming. Configuring the model for decoupled transactions and using a streaming-capable client over HTTP or gRPC delivers tokens incrementally. Other Triton features like batching, ensembles, or instance groups improve throughput or composition but do not stream partial outputs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Model ensembles combining preprocessing and postprocessing models.

    Why it's wrong here

    Ensembles chain multiple models into a pipeline but do not enable streaming of partial outputs to clients. They are useful for composing preprocessing, inference, and postprocessing steps, yet the response is still returned as a complete result. This does not satisfy the token-streaming requirement for a web application.

  • ✗

    Dynamic batching in the model configuration.

    Why it's wrong here

    Dynamic batching groups multiple inference requests into a single batch to improve throughput. It does not provide token-by-token streaming over HTTP. While it improves server efficiency, it operates at the request scheduling layer and does not affect how responses are delivered to the client incrementally.

  • ✗

    Instance groups with multiple GPU instances per model.

    Why it's wrong here

    Instance groups control how many model instances run on available GPUs to scale throughput. They do not change the response delivery mechanism and cannot produce incremental token streams. This setting affects concurrency and resource allocation, not the streaming behavior needed for a chat-style web interface.

  • ✓

    Decoupled mode with streaming responses in the model backend.

    Why this is correct

    Decoupled mode allows a model backend to return multiple responses for a single request, which is essential for token streaming in LLMs. Triton's HTTP and gRPC endpoints support streaming when the model is configured for decoupled transactions. This lets the web application receive partial outputs as tokens are generated, improving perceived latency.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.