NCA-GENL Data Analysis and Visualization Practice Question
Exhibit
{
"model_name": "llama-3-8b",
"metrics": {
"throughput_tokens_per_sec": 4500,
"latency_ms": 120,
"gpu_utilization_avg": 95,
"kv_cache_usage_percent": 98
},
"status": "WARNING: KV cache fragmentation high"
}Refer to the exhibit. The monitoring JSON indicates high KV cache fragmentation. Which visualization best helps developers diagnose if this is caused by heterogeneous request lengths in the workload?
⚠ Common exam trap
Candidates often choose a 'memory usage over time' plot, which shows that memory is high but fails to explain the root cause (heterogeneous request lengths) of the fragmentation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A histogram of request input/output sequence lengths
KV cache fragmentation occurs when varying sequence lengths lead to non-contiguous memory allocations. By plotting a histogram of 'input sequence lengths' versus 'output sequence lengths', developers can see the variance in request sizes. High variance indicates a need for paged attention or continuous batching optimization, which allows the engine to handle variable lengths efficiently without wasting memory on fragmented cache blocks, directly addressing the performance degradation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
A line chart of GPU memory temperature
Why it's wrong here
GPU temperature is a hardware telemetry metric that does not correlate with KV cache fragmentation. While thermal throttling can affect performance, it is irrelevant to the specific software-level memory management issue indicated by the KV cache fragmentation warning, which is caused by software-side request handling patterns.
- ✓
A histogram of request input/output sequence lengths
Why this is correct
Visualizing the distribution of sequence lengths is the standard way to diagnose KV cache fragmentation. If the histogram shows a wide range of lengths, the memory manager is struggling to fit blocks efficiently. This confirms that the workload requires strategies like PagedAttention to minimize memory waste.
- ✗
A heat map of individual GPU core activity
Why it's wrong here
Core activity heat maps are for profiling compute-bound tasks. KV cache management is a memory-bound process. Monitoring core usage will not explain why the cache is fragmented, as the fragmentation happens in memory address space, not in the execution cycles of individual CUDA streaming multiprocessors.
- ✗
A scatter plot of token generation probabilities
Why it's wrong here
Token generation probabilities describe the model's confidence and output quality, not the efficiency of the underlying hardware memory allocation. This visualization is useful for evaluating model accuracy but provides no insight into the technical bottlenecks related to KV cache management and fragmentation.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.