Courseiva

NCA-GENL Data Analysis and Visualization Practice Question

Exhibit

{
  "model_name": "llama-3-8b",
  "metrics": {
    "throughput_tokens_per_sec": 4500,
    "latency_ms": 120,
    "gpu_utilization_avg": 95,
    "kv_cache_usage_percent": 98
  },
  "status": "WARNING: KV cache fragmentation high"
}

Refer to the exhibit. The monitoring JSON indicates high KV cache fragmentation. Which visualization best helps developers diagnose if this is caused by heterogeneous request lengths in the workload?

⚠ Common exam trap

Candidates often choose a 'memory usage over time' plot, which shows that memory is high but fails to explain the root cause (heterogeneous request lengths) of the fragmentation.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

A histogram of request input/output sequence lengths

KV cache fragmentation occurs when varying sequence lengths lead to non-contiguous memory allocations. By plotting a histogram of 'input sequence lengths' versus 'output sequence lengths', developers can see the variance in request sizes. High variance indicates a need for paged attention or continuous batching optimization, which allows the engine to handle variable lengths efficiently without wasting memory on fragmented cache blocks, directly addressing the performance degradation.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    A line chart of GPU memory temperature

    Why it's wrong here

    GPU temperature is a hardware telemetry metric that does not correlate with KV cache fragmentation. While thermal throttling can affect performance, it is irrelevant to the specific software-level memory management issue indicated by the KV cache fragmentation warning, which is caused by software-side request handling patterns.

  • ✓

    A histogram of request input/output sequence lengths

    Why this is correct

    Visualizing the distribution of sequence lengths is the standard way to diagnose KV cache fragmentation. If the histogram shows a wide range of lengths, the memory manager is struggling to fit blocks efficiently. This confirms that the workload requires strategies like PagedAttention to minimize memory waste.

  • ✗

    A heat map of individual GPU core activity

    Why it's wrong here

    Core activity heat maps are for profiling compute-bound tasks. KV cache management is a memory-bound process. Monitoring core usage will not explain why the cache is fragmented, as the fragmentation happens in memory address space, not in the execution cycles of individual CUDA streaming multiprocessors.

  • ✗

    A scatter plot of token generation probabilities

    Why it's wrong here

    Token generation probabilities describe the model's confidence and output quality, not the efficiency of the underlying hardware memory allocation. This visualization is useful for evaluating model accuracy but provides no insight into the technical bottlenecks related to KV cache management and fragmentation.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.