Courseiva

NCP-GENL Production Monitoring and Reliability Practice Question

An LLM inference service on NVIDIA Triton Inference Server uses TensorRT-LLM as the backend. During peak load, you notice that the first token latency (time to first token) is high, but subsequent tokens are generated quickly. Which metric should you monitor to diagnose this issue?

⚠ Common exam trap

The trap here is focusing on GPU utilization or throughput instead of the specific latency metric that reflects the prefill phase.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Time to first token (TTFT)

Time to first token (TTFT) is the key metric for diagnosing prefill latency in LLM inference. Since the issue is high TTFT but fast subsequent tokens, monitoring TTFT allows you to isolate the prefill phase and optimize it, for example by reducing batch size or using more efficient attention kernels.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Time to first token (TTFT)

    Why this is correct

    TTFT directly measures the latency from request submission to the first generated token. In LLM inference, the prefill phase (processing the input prompt) dominates TTFT. Monitoring TTFT helps identify if the prefill is slow due to large batch sizes, long prompts, or insufficient compute, allowing targeted optimization.

  • ✗

    GPU utilization

    Why it's wrong here

    GPU utilization measures how busy the GPU is overall, but it does not distinguish between the prefill and decode phases. High time to first token could occur even with moderate GPU utilization, so this metric alone would not pinpoint the cause.

  • ✗

    Requests per second (RPS)

    Why it's wrong here

    RPS indicates overall throughput but does not provide insight into latency phases. A high RPS could coexist with high TTFT if the system is overloaded. RPS alone would not help diagnose the specific issue of slow first token generation.

  • ✗

    Inter-token latency (ITL)

    Why it's wrong here

    ITL measures the time between consecutive tokens during generation. The scenario states that subsequent tokens are generated quickly, so ITL is likely normal. Focusing on ITL would not address the high TTFT, which is the reported problem.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.