Courseiva

NCA-GENL Data Analysis and Visualization Practice Question

A team is preparing a dashboard to monitor an LLM inference service in production. They want visualizations that surface latency problems and resource saturation before users are affected. Which two visualizations are most appropriate for this goal? (Choose two.)

⚠ Common exam trap

The trap here is choosing charts that describe what users are sending rather than charts that measure how the service is responding and whether its resources are saturating.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

A gauge or time-series chart of GPU memory utilization and KV cache occupancy

Effective production monitoring pairs a latency view with a resource view. Percentile time series expose tail latency changes that averages conceal, while GPU memory and KV cache occupancy show the underlying pressure that causes those changes. Together they let the team act on leading indicators. Traffic-share pies, token word clouds, and input-output scatter plots describe workload characteristics rather than service health, so they do not support preemptive detection of degradation.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    A gauge or time-series chart of GPU memory utilization and KV cache occupancy

    Why this is correct

    GPU memory and KV cache occupancy are leading indicators of inference saturation. As concurrent sequences grow, the KV cache expands and can approach memory limits, forcing request queuing or preemption that spikes latency. Tracking these resources alongside traffic lets the team see pressure building before it manifests as timeouts, making this visualization directly relevant to preemptive monitoring.

  • ✗

    A word cloud of the most frequent prompt tokens received today

    Why it's wrong here

    A word cloud summarizes prompt vocabulary, which is useful for content analysis but irrelevant to latency and resource saturation. It provides no temporal signal and cannot indicate whether the service is slowing down or running out of memory. Including it would add visual noise without helping the team detect performance degradation before users are impacted.

  • ✓

    A time-series chart of request latency percentiles (p50, p95, p99) over the last 24 hours

    Why this is correct

    Percentile time series reveal the tail behavior that averages hide. A rising p99 while p50 stays flat indicates a subset of requests is degrading, often before users broadly notice. Plotting p50, p95, and p99 together over a rolling window shows both typical and worst-case latency, enabling the team to detect saturation trends and correlate spikes with traffic or deployment events.

  • ✗

    A scatter plot of prompt length versus generated response length for a random sample

    Why it's wrong here

    This scatter plot describes input-output relationships, not service health. It might inform capacity planning over long horizons, but it does not update with production load and cannot show current latency or memory pressure. It fails to provide the real-time leading indicators needed to catch saturation before users experience slow or failed requests.

  • ✗

    A pie chart of total requests grouped by model version for the current day

    Why it's wrong here

    A pie chart of request share by model version shows traffic distribution, not latency or saturation. It cannot reveal slow requests, queue buildup, or resource pressure. While useful for rollout tracking, it does not surface the degradation signals the team wants to catch before users are affected, so it is not among the most appropriate choices for this monitoring goal.

About these practice questions

This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.