NCP-GENL Production Monitoring and Reliability Practice Question
A production LLM inference service on NVIDIA Triton Inference Server is experiencing intermittent slowdowns. The team wants a single metric that best indicates whether the GPU is the bottleneck during these periods. Which metric should they monitor?
⚠ Common exam trap
The trap here is assuming that host CPU or network metrics are sufficient to diagnose GPU-bound LLM inference slowdowns.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
GPU utilization (SM utilization)
GPU utilization (SM utilization) is the most direct indicator of whether the GPU's compute units are saturated. During LLM inference, if slowdowns coincide with near-100% SM utilization, the GPU is the bottleneck. Other metrics like CPU, network, or disk I/O are less likely to explain intermittent latency spikes when the GPU is the primary compute resource.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
GPU utilization (SM utilization)
Why this is correct
GPU utilization, specifically SM utilization, directly measures the percentage of time the GPU's streaming multiprocessors are active. In LLM inference, high SM utilization sustained near 100% strongly suggests the GPU is the bottleneck. Monitoring this metric during slowdowns helps determine if the GPU is saturated by compute or memory-bound operations, guiding optimization such as batching or model quantization.
- ✗
Disk I/O operations per second on the model repository
Why it's wrong here
Disk I/O on the model repository is relevant during model loading or swapping, not during steady-state inference. Once models are loaded into GPU memory, disk activity is minimal. Monitoring disk I/O would not reveal GPU bottlenecks causing intermittent slowdowns. This metric is more suited for detecting storage performance issues during model initialization or updates.
- ✗
Network throughput between clients and the server
Why it's wrong here
Network throughput measures data transfer rates, which matters for large payloads but is rarely the primary bottleneck for LLM inference since inputs and outputs are small text tokens. Slowdowns due to GPU compute would not be captured by network metrics. Relying on this metric could lead to unnecessary network upgrades while the GPU remains the true constraint.
- ✗
CPU utilization of the inference server host
Why it's wrong here
CPU utilization indicates host processing load, including request preprocessing, tokenization, and orchestration. While high CPU usage can delay request handling, it does not reflect GPU compute saturation. In LLM inference, the GPU typically dominates latency, so CPU metrics alone would mislead the team into optimizing the wrong component, missing the actual GPU bottleneck.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.