Courseiva

NCP-GENL Production Monitoring and Reliability Practice Question

A production LLM inference service on NVIDIA Triton Inference Server is experiencing intermittent slowdowns. The team wants a single metric that best indicates whether the GPU is the bottleneck during these periods. Which metric should they monitor?

⚠ Common exam trap

The trap here is assuming that host CPU or network metrics are sufficient to diagnose GPU-bound LLM inference slowdowns.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

GPU utilization (SM utilization)

GPU utilization (SM utilization) is the most direct indicator of whether the GPU's compute units are saturated. During LLM inference, if slowdowns coincide with near-100% SM utilization, the GPU is the bottleneck. Other metrics like CPU, network, or disk I/O are less likely to explain intermittent latency spikes when the GPU is the primary compute resource.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    GPU utilization (SM utilization)

    Why this is correct

    GPU utilization, specifically SM utilization, directly measures the percentage of time the GPU's streaming multiprocessors are active. In LLM inference, high SM utilization sustained near 100% strongly suggests the GPU is the bottleneck. Monitoring this metric during slowdowns helps determine if the GPU is saturated by compute or memory-bound operations, guiding optimization such as batching or model quantization.

  • ✗

    Disk I/O operations per second on the model repository

    Why it's wrong here

    Disk I/O on the model repository is relevant during model loading or swapping, not during steady-state inference. Once models are loaded into GPU memory, disk activity is minimal. Monitoring disk I/O would not reveal GPU bottlenecks causing intermittent slowdowns. This metric is more suited for detecting storage performance issues during model initialization or updates.

  • ✗

    Network throughput between clients and the server

    Why it's wrong here

    Network throughput measures data transfer rates, which matters for large payloads but is rarely the primary bottleneck for LLM inference since inputs and outputs are small text tokens. Slowdowns due to GPU compute would not be captured by network metrics. Relying on this metric could lead to unnecessary network upgrades while the GPU remains the true constraint.

  • ✗

    CPU utilization of the inference server host

    Why it's wrong here

    CPU utilization indicates host processing load, including request preprocessing, tokenization, and orchestration. While high CPU usage can delay request handling, it does not reflect GPU compute saturation. In LLM inference, the GPU typically dominates latency, so CPU metrics alone would mislead the team into optimizing the wrong component, missing the actual GPU bottleneck.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.