Courseiva

NCP-GENL Production Monitoring and Reliability Practice Question

A production LLM service on NVIDIA Triton Inference Server uses dynamic batching. During peak load, the 99th percentile latency increases significantly, but GPU utilization remains at 60%. Which configuration change is most likely to improve latency while maintaining throughput?

⚠ Common exam trap

The trap here is assuming that increasing batch size always improves performance, overlooking the impact of batch timeout on latency.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Decrease the maximum batch size and adjust the batch timeout to reduce waiting

When GPU utilization is low but latency is high, the bottleneck is often the batching mechanism waiting to accumulate requests. Reducing the maximum batch size and batch timeout allows requests to be processed sooner, lowering queue time and P99 latency. Since the GPU has spare capacity, this change can improve latency without sacrificing throughput.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Decrease the maximum batch size and adjust the batch timeout to reduce waiting

    Why this is correct

    With GPU utilization at 60%, the GPU has spare capacity. High P99 latency during peak load suggests requests are waiting too long for batches to fill. Reducing the maximum batch size and lowering the batch timeout (e.g., from 100 microseconds to 50) allows smaller batches to be processed more quickly, reducing queue time. This can improve latency while still maintaining acceptable throughput because the GPU is not fully utilized.

  • ✗

    Enable model instances to increase parallelism on the same GPU

    Why it's wrong here

    Enabling multiple model instances can increase throughput by allowing concurrent execution, but it may not reduce latency if the bottleneck is batch formation delay. With GPU utilization at 60%, adding instances might increase utilization but could also introduce contention. The primary issue is likely excessive waiting for batches, so adjusting batching parameters is more directly effective.

  • ✗

    Switch to a larger GPU with more memory bandwidth

    Why it's wrong here

    A larger GPU with more memory bandwidth could help if the model is memory-bound, but GPU utilization at 60% indicates the GPU is not fully utilized. The latency spike is more likely due to batching delays rather than insufficient hardware resources. Upgrading hardware is costly and may not address the root cause, which is the batching configuration.

  • ✗

    Increase the maximum batch size in the dynamic batching configuration

    Why it's wrong here

    Increasing the maximum batch size can improve throughput but may actually worsen latency because larger batches take longer to process, increasing per-request wait time. With GPU utilization at 60%, the GPU is not saturated, so larger batches are unlikely to help and could exacerbate the latency issue by delaying individual requests. The bottleneck may be elsewhere, such as batch formation time or memory bandwidth.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.