NCA-GENL Data Analysis and Visualization Practice Question
A data scientist is analyzing the output of a Llama 3 8B model on a summarization task. The token-level log-probabilities are extracted, and the goal is to visualize how confident the model is in each generated token across the summary. Which visualization is most appropriate for showing the per-token probability distribution and identifying tokens where the model is uncertain?
⚠ Common exam trap
Candidates often confuse embedding-space visualizations like t-SNE with output-probability visualizations, when only the latter directly shows model confidence per token.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A line chart plotting the log-probability of each generated token against its position in the output sequence.
The scenario asks for per-token confidence across a generated summary. A line chart of log-probability by token position preserves the sequential order and directly displays the model's confidence at each step. The other options either aggregate away position information or use dimensionality reduction that does not represent output probabilities.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
A bar chart of the top-5 predicted tokens for the entire summary, aggregated across all positions.
Why it's wrong here
Aggregating top-5 predictions over the whole summary destroys the per-position information needed to see where uncertainty occurs. A bar chart of aggregated counts shows which tokens were frequently predicted, not how confident the model was at each step. It cannot reveal token-level uncertainty across the summary.
- ✗
A t-SNE scatter plot of the hidden states of all generated tokens, colored by token ID.
Why it's wrong here
t-SNE visualizes high-dimensional embeddings in 2D and preserves local neighborhood structure, but it does not directly represent the model's output probability for each token. Colored by token ID, it would show clustering of hidden states, not confidence. It cannot answer which tokens the model is uncertain about during generation.
- ✓
A line chart plotting the log-probability of each generated token against its position in the output sequence.
Why this is correct
Plotting log-probability versus token position directly shows per-token confidence and highlights dips where the model is uncertain, which is exactly the goal. This line chart is a standard way to inspect generation quality token by token, and it scales to long sequences without losing individual token information.
- ✗
A confusion matrix comparing predicted tokens to ground-truth tokens for the entire summary.
Why it's wrong here
A confusion matrix is designed for classification with a fixed, finite label set, not for open-ended text generation where the label space is the whole vocabulary. It would collapse token identities into an enormous matrix and obscure the sequential confidence pattern. It does not show per-token probability or position, so it fails this scenario.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.