NCA-GENL Data Analysis and Visualization Practice Question
A team is building a dashboard to monitor an LLM training run and wants to detect data-quality problems early. They have access to per-batch training loss, per-batch gradient norm, input sequence length statistics, and token frequency counts. Which two visualizations are most appropriate for surfacing data-quality issues rather than hardware or throughput issues? (Choose two.)
⚠ Common exam trap
The trap here is treating infrastructure metrics like GPU utilization as proxies for data quality, when they only measure hardware efficiency.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A line chart of per-batch training loss over steps
Per-batch training loss and input sequence length histograms both reflect the model's interaction with the data. Loss spikes or shifts can indicate corrupted or mislabeled batches, while sequence-length histograms reveal truncation or padding anomalies. GPU utilization, memory usage, and network throughput are infrastructure metrics that describe hardware behavior, not data correctness, so they do not surface data-quality issues.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
A line chart of per-batch training loss over steps
Why this is correct
Per-batch training loss over steps reveals sudden jumps or sustained elevations that often correspond to corrupted, mislabeled, or out-of-distribution batches. A data-quality problem typically manifests as a spike or shift in loss that is not explained by learning-rate changes. This chart is directly tied to model behavior on the data, making it one of the most useful early indicators of data issues.
- ✓
A histogram of input sequence lengths
Why this is correct
A histogram of input sequence lengths exposes truncation, padding anomalies, or a shift toward very short or very long documents. If a data pipeline change accidentally truncates examples, the histogram will show a pile-up at the maximum length. This is a data-quality signal that is independent of hardware performance and helps catch preprocessing bugs before they degrade model quality.
- ✗
A line chart of network throughput between nodes
Why it's wrong here
Network throughput reflects communication efficiency in distributed training. It is useful for diagnosing slow collectives or interconnect bottlenecks, but it is unrelated to the semantic or structural quality of the training data. High throughput can coexist with corrupted examples, so this chart does not help detect data-quality problems.
- ✗
A line chart of GPU utilization over time
Why it's wrong here
GPU utilization measures hardware efficiency and can reveal data-loading stalls, but it does not indicate whether the data content is correct or well-formed. A pipeline can feed perfectly valid data at full utilization while still containing label errors or distribution shifts. This metric belongs to throughput monitoring, not data-quality diagnosis.
- ✗
A bar chart of memory usage per GPU
Why it's wrong here
Memory usage per GPU is a capacity and stability metric. It helps detect fragmentation or out-of-memory risk, but it says nothing about whether the training examples are correctly labeled or properly preprocessed. Memory pressure can occur with perfectly clean data, so this chart does not serve the data-quality monitoring goal.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.