Courseiva
Evaluation →hardMultiple Choice

NCP-GENL Evaluation Practice Question

A team is evaluating a large language model (LLM) deployed with NVIDIA Triton Inference Server for a real-time question-answering system. They observe that the model's responses are accurate but latency spikes during peak load, causing timeouts. They need to evaluate the model's performance under high concurrency to identify the maximum throughput while maintaining a 95th percentile latency below 200 ms. Which evaluation approach should they use?

⚠ Common exam trap

Watch out — candidates often confuse theoretical compute limits or single-request latency with actual concurrent performance, which requires load testing with multiple simultaneous requests.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use NVIDIA Triton's Performance Analyzer with concurrency sweep and latency constraints.

Evaluating performance under high concurrency requires simulating multiple simultaneous requests and measuring latency percentiles. NVIDIA Triton's Performance Analyzer is designed for this, allowing concurrency sweeps and latency constraints to find the maximum throughput meeting the SLA. Other options either lack concurrency or measure irrelevant metrics, so they cannot identify the operating point that satisfies the 95th percentile latency requirement.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Measure GPU utilization with NVIDIA Management Library (NVML) during a single long-running inference.

    Why it's wrong here

    GPU utilization from NVML indicates how busy the GPU is but does not provide latency or throughput metrics under concurrent load. A single long-running inference does not simulate multiple simultaneous requests, so it cannot reveal latency spikes or maximum throughput. Thus, this method fails to evaluate the system's behavior under high concurrency.

  • ✗

    Run a one-off inference with a batch size of 1 and measure latency.

    Why it's wrong here

    A single inference with batch size 1 does not simulate concurrent load and will not reveal latency spikes under high concurrency. It only measures baseline latency without contention for GPU resources. Therefore, it cannot identify the maximum throughput that meets the latency SLA. This approach is inadequate for the scenario's requirement to evaluate under peak load.

  • ✗

    Compute the model's FLOPS and compare against the GPU's peak theoretical FLOPS.

    Why it's wrong here

    FLOPS comparison gives a theoretical upper bound on compute but ignores memory bandwidth, kernel launch overhead, and concurrency effects. It does not measure actual latency or throughput under real workloads. Therefore, it cannot determine the maximum throughput that meets a specific latency SLA. This approach is too abstract for the scenario's operational evaluation.

  • ✓

    Use NVIDIA Triton's Performance Analyzer with concurrency sweep and latency constraints.

    Why this is correct

    NVIDIA Triton's Performance Analyzer can simulate concurrent requests by sweeping concurrency levels and measuring latency percentiles. It supports setting a latency constraint (e.g., 95th percentile < 200 ms) and automatically finds the maximum throughput that satisfies it. This directly addresses the need to evaluate performance under high concurrency and identify the optimal operating point.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.