Courseiva
Software Development →hardMultiple Choice

NCA-GENL Software Development Practice Question

A developer is using NVIDIA Triton Inference Server to deploy a TensorRT-LLM optimized model. The model must support multiple concurrent users with low latency. The developer notices that latency spikes when many requests arrive simultaneously. Which Triton feature should be configured to improve throughput while maintaining acceptable latency?

⚠ Common exam trap

It's easy for candidates to confuse model warmup or instance groups with dynamic batching, which is the primary mechanism to handle concurrent request throughput in Triton.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Dynamic batching with a `max_batch_size` and `preferred_batch_size` in the model configuration.

Triton's dynamic batching is essential for handling concurrent requests efficiently. By configuring `max_batch_size` and `preferred_batch_size`, the server can group multiple inference requests into a single batch, reducing the number of forward passes and improving GPU utilization. This leads to higher throughput and lower latency under load, which is critical for serving TensorRT-LLM models to multiple users. Other features like warmup or multiple instances address different aspects but not the core batching need.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Sequence batching to maintain state across multiple inference requests for stateful models.

    Why it's wrong here

    Sequence batching is designed for stateful models like recurrent neural networks, where state must be preserved across requests. TensorRT-LLM models for generative AI are typically stateless per request, and sequence batching would add unnecessary complexity. It does not improve throughput for independent concurrent requests; dynamic batching is the appropriate feature for that scenario.

  • ✗

    Model warmup with sample inputs to preload CUDA kernels and reduce cold-start latency.

    Why it's wrong here

    Model warmup helps reduce initial latency when the model is first loaded, but it does not address latency spikes under concurrent load. Warmup ensures that CUDA kernels are compiled and cached, but once the model is running, dynamic batching is needed to handle multiple requests efficiently. Warmup alone would not improve throughput for simultaneous requests.

  • ✓

    Dynamic batching with a `max_batch_size` and `preferred_batch_size` in the model configuration.

    Why this is correct

    Triton's dynamic batching automatically groups incoming inference requests into batches to improve GPU utilization and throughput. By setting `max_batch_size` and `preferred_batch_size`, the server can form batches that fit within latency constraints. For TensorRT-LLM models, this is crucial for handling concurrent users efficiently, as it reduces the number of forward passes and amortizes overhead, directly addressing latency spikes under load.

  • ✗

    Instance groups with multiple model instances per GPU to increase parallelism.

    Why it's wrong here

    Multiple model instances can increase parallelism, but for large TensorRT-LLM models, they may not fit in GPU memory simultaneously. Moreover, without dynamic batching, each instance processes one request at a time, which may not fully utilize the GPU. Dynamic batching is more effective for maximizing throughput with large language models because it combines requests into a single batch, leveraging the GPU's parallel processing capabilities.

About these practice questions

Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.