NCP-GENL Production Monitoring and Reliability Practice Question
An LLM inference service on NVIDIA Triton Inference Server uses TensorRT-LLM as the backend. During peak load, you notice that the first token latency (time to first token) is high, but subsequent tokens are generated quickly. Which metric should you monitor to diagnose this issue?
⚠ Common exam trap
The trap here is focusing on GPU utilization or throughput instead of the specific latency metric that reflects the prefill phase.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Time to first token (TTFT)
Time to first token (TTFT) is the key metric for diagnosing prefill latency in LLM inference. Since the issue is high TTFT but fast subsequent tokens, monitoring TTFT allows you to isolate the prefill phase and optimize it, for example by reducing batch size or using more efficient attention kernels.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Time to first token (TTFT)
Why this is correct
TTFT directly measures the latency from request submission to the first generated token. In LLM inference, the prefill phase (processing the input prompt) dominates TTFT. Monitoring TTFT helps identify if the prefill is slow due to large batch sizes, long prompts, or insufficient compute, allowing targeted optimization.
- ✗
GPU utilization
Why it's wrong here
GPU utilization measures how busy the GPU is overall, but it does not distinguish between the prefill and decode phases. High time to first token could occur even with moderate GPU utilization, so this metric alone would not pinpoint the cause.
- ✗
Requests per second (RPS)
Why it's wrong here
RPS indicates overall throughput but does not provide insight into latency phases. A high RPS could coexist with high TTFT if the system is overloaded. RPS alone would not help diagnose the specific issue of slow first token generation.
- ✗
Inter-token latency (ITL)
Why it's wrong here
ITL measures the time between consecutive tokens during generation. The scenario states that subsequent tokens are generated quickly, so ITL is likely normal. Focusing on ITL would not address the high TTFT, which is the reported problem.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.