Courseiva

NCP-GENL Production Monitoring and Reliability Practice Question

An LLM inference service on NVIDIA Triton Inference Server is experiencing a gradual increase in P99 latency over several hours without a corresponding increase in request rate. GPU utilization remains stable at around 60%. Which action should the team take first to diagnose the root cause?

⚠ Common exam trap

The trap here is assuming that latency increases always require scaling, without considering resource leaks that don't affect utilization.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Check for memory leaks in the model or Triton server process.

The gradual increase in P99 latency with stable GPU utilization and request rate suggests a resource leak, such as memory fragmentation or a leak in the model or Triton process. Checking for memory leaks is the most direct first step to diagnose the issue. Scaling or profiling may be premature and could mask the root cause. Reviewing input data is unlikely given the stable utilization.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the number of model instances to handle the load.

    Why it's wrong here

    Increasing model instances might temporarily reduce latency, but it does not diagnose the root cause. Since GPU utilization is stable at 60%, the system is not resource-starved. Adding instances could waste resources and mask the underlying issue. The team should first investigate why latency is increasing before scaling.

  • ✓

    Check for memory leaks in the model or Triton server process.

    Why this is correct

    A gradual increase in latency without increased load or GPU utilization suggests a resource leak, such as memory fragmentation or a memory leak in the model or server. Over time, this can cause more frequent garbage collection or swapping, increasing latency. Checking for memory leaks is a logical first step to identify the root cause before applying fixes.

  • ✗

    Profile the inference pipeline with NVIDIA Nsight Systems to identify bottlenecks.

    Why it's wrong here

    Profiling with Nsight Systems is useful for identifying performance bottlenecks, but it is a more intensive step. Given the gradual latency increase, a memory leak is a more likely cause. Profiling might not reveal a leak unless specifically looking for memory allocation patterns. It is better to first check for leaks, which is quicker and more targeted.

  • ✗

    Review the model's input data for changes in sequence length.

    Why it's wrong here

    Changes in input sequence length can affect latency, but the scenario states no increase in request rate and stable GPU utilization. If sequence lengths increased, GPU utilization would likely rise. Therefore, this is less likely. The gradual increase over hours points to a systemic issue like a leak rather than input variation.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.