NCA-GENL Data Analysis and Visualization Practice Question
A developer is profiling an LLM inference endpoint on an NVIDIA L40S and observes that time-to-first-token (TTFT) is stable but inter-token latency spikes periodically. They want to determine whether the spikes align with KV cache growth or with batch-size changes. Which visualization strategy best isolates the cause?
⚠ Common exam trap
The trap here is reaching for a profiling tool that explains a single request when the symptom is periodic across many requests.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Plot a time series of inter-token latency with KV cache size and active batch size overlaid on the same time axis.
The goal is to determine whether periodic latency spikes align temporally with KV cache growth or batch-size changes. Only a shared time axis can establish that alignment. Plotting inter-token latency, KV cache size, and active batch size together makes coincident events visible immediately. Summary statistics, length histograms, and single-request flame graphs all discard the cross-request temporal relationship required to isolate the cause.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Generate a histogram of token lengths for all requests served during the profiling window.
Why it's wrong here
A token-length histogram describes the workload mix but has no time axis, so it cannot show whether a specific latency spike coincided with a particular request or cache state. Long requests may correlate with spikes, but the chart cannot demonstrate alignment. It is a workload characterization, not a latency diagnosis.
- ✓
Plot a time series of inter-token latency with KV cache size and active batch size overlaid on the same time axis.
Why this is correct
Overlaying inter-token latency, KV cache size, and active batch size on one time axis lets the developer see whether each latency spike coincides with a cache growth event or a batch-size change. Temporal alignment is the key diagnostic, and this single view provides it without needing to correlate separate charts by eye.
- ✗
Compute the mean and standard deviation of inter-token latency and report them as a bar chart.
Why it's wrong here
Summary statistics collapse the time dimension, so periodic spikes are averaged away and cannot be aligned with any cause. A bar chart of mean and standard deviation tells you variability exists but not when it occurs or what accompanied it. This loses exactly the temporal information needed to isolate the cause.
- ✗
Render a flame graph of GPU kernel execution for a single representative request.
Why it's wrong here
A flame graph reveals where time is spent inside one request but discards wall-clock alignment across requests, so periodic spikes that occur every N requests would not be visible. It also does not show KV cache size or batch size. This tool answers a per-request question, not a periodic-correlation question.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.