NCA-GENL · domain
Data Analysis and Visualization
This domain covers interpreting LLM evaluation outputs, training diagnostics, and benchmark comparisons. Questions present exhibits such as loss curves, gradient-norm plots, or RAG benchmark distributions and ask you to choose the correct visualization or explain the observed behavior. Expect scenario-based items tied to fine-tuning on NVIDIA DGX systems and NVIDIA NIM or NeMo workflows.
Focused practice
Practice Data Analysis and Visualization questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Data Analysis and Visualization
Be able to read loss curves, gradient-norm diagnostics, and benchmark exhibits, then pick the visualization matching the comparison. The key is distinguishing transient training spikes from real divergence and choosing distribution-aware plots over summary statistics.
Reading training loss curves to distinguish benign spikes from divergence or data issues
Selecting visualizations that compare distributions across two LLM architectures
Using per-example gradient norms to flag outlier training examples
Interpreting RAG benchmark evaluation exhibits for model-update decisions
Watch out for
Common Data Analysis and Visualization exam traps
- ▸Treating a localized loss spike with immediate recovery as catastrophic divergence instead of a transient batch or data artifact
- ▸Choosing a bar chart of means when the question asks to compare full distributions across architectures
- ▸Ignoring gradient-norm outliers that indicate mislabeled or anomalous examples corrupting fine-tuning
Question index
All Data Analysis and Visualization questions (65)
Click any question to see the full explanation, or start a practice session above.
A data scientist is monitoring a fine-tuning job on a DGX system. The training loss graph shows a sharp, localized spike followed by an immediate return to the previous trend. What is the most likely cause?
Medium2During the evaluation of a Large Language Model, you notice that the model consistently predicts the most frequent tokens regardless of the context. Which visualization would most clearly illustrate this phenomenon of 'probability collapse'?
Easy3A data scientist has embedded 200,000 LLM training documents with a sentence-transformer and wants to visualize the embedding space to inspect semantic clusters. Running UMAP on the full set is too slow, so they first reduce dimensions with PCA. Which approach best preserves local cluster structure for the final visualization?
Hard4Refer to the exhibit. The monitoring JSON indicates high KV cache fragmentation. Which visualization best helps developers diagnose if this is caused by heterogeneous request lengths in the workload?
Hard5A data scientist has a table of 500 LLM evaluation runs, each with a numerical faithfulness score from 0 to 1 and a categorical model version label. They want a compact view comparing the score distributions across model versions, including medians and spread, in a single figure. Which visualization should they choose?
Easy6You are analyzing the output of an LLM inference endpoint that returns a JSON payload containing a top-k token probability distribution for a single generated step. Which visualization most directly communicates the model's confidence ranking across the returned tokens?
Medium7You are analyzing the quality of a synthetic data generation pipeline for an LLM. You want to ensure the synthetic data does not suffer from 'mode collapse' compared to the real-world dataset. Which visualization technique is most effective for comparing the diversity of the two datasets?
Medium8A data analyst is preparing a dashboard for an LLM inference service. The service logs per-request latency in milliseconds. The team wants a single chart that lets an on-call engineer quickly see both the typical latency and how often requests exceed the service-level objective of 500 ms. Which visualization best supports that goal?
Easy9A team is fine-tuning an NVIDIA NIM-hosted Llama 3 8B model and wants a single visualization that tracks per-step training loss, learning rate, and GPU memory utilization together, so they can correlate a mid-run loss spike with resource pressure. They need a framework that integrates natively with the NVIDIA NeMo training stack and requires minimal custom plotting code. Which visualization approach best meets these requirements?
Medium10A data scientist is analyzing a large corpus of LLM training documents and wants to visualize which topics appear together across documents. After computing TF-IDF vectors, they apply non-negative matrix factorization (NMF) to reduce dimensionality. Which visualization best shows the relationships between the discovered topics and the documents?
Medium11While reviewing training logs from a multi-node NVIDIA DGX cluster running data-parallel fine-tuning, you plot per-step gradient norm alongside loss. The gradient norm shows sharp periodic spikes every N steps that align with evaluation checkpoints. Which action should you take to determine whether the spikes are an artifact of the evaluation loop or a genuine optimization problem?
Hard12You have generated 512-dimensional embeddings for 200,000 documents using an NVIDIA NeMo embedding model and want to inspect whether semantically similar documents cluster together. Which technique should you apply first to project these embeddings into two dimensions for visual inspection?
Easy13A data scientist is preparing a visualization of an LLM evaluation suite that covers several task types with different score ranges. The audience includes both engineers and non-technical stakeholders. Which two practices best ensure the visualization is accurate and interpretable? (Choose two.)
Medium14In the context of analyzing LLM output safety, what does a 'confusion matrix' help identify?
Medium15You are preparing a quarterly report on an LLM's inference latency for stakeholders. The raw data contains 50,000 individual request latencies in milliseconds. You need a single visualization that shows the full distribution shape, including any long tail of slow requests, without losing information to binning. Which visualization should you use?
Easy16A data scientist is analyzing 2 million document embeddings from a RAG corpus on a single NVIDIA GPU. A full pairwise cosine similarity matrix would require roughly 16 TB of memory, which is infeasible. They need to identify near-duplicate documents and visualize cluster density without materializing the full matrix. Which approach is most appropriate?
Hard17A data scientist observes that the model's loss plateaus early during fine-tuning. Which visualization would best help diagnose if the model is suffering from 'catastrophic forgetting'?
Medium18A data scientist is analyzing the output of a Llama 3 8B model on a summarization task. The token-level log-probabilities are extracted, and the goal is to visualize how confident the model is in each generated token across the summary. Which visualization is most appropriate for showing the per-token probability distribution and identifying tokens where the model is uncertain?
Medium19You are building a dashboard to monitor an LLM inference service deployed with NVIDIA Triton Inference Server. Stakeholders want to detect quality degradation and latency regressions before users complain. Which two metrics should be tracked continuously to surface these issues earliest? (Choose two.)
Medium20A team is analyzing the latency of an LLM inference service. They have per-request latency data for 10,000 requests and want to visualize the distribution to identify whether there is a long tail that could violate a service-level objective. Which visualization is most appropriate for this purpose?
Hard21When evaluating LLM output quality using human-in-the-loop data, which THREE metrics or techniques are most effective for detecting systemic hallucinations?
Medium22A data scientist is preparing a visualization to compare the performance of three different LLM fine-tuning runs on a summarization benchmark. They want to show both the central tendency and the variability of ROUGE-L scores across multiple evaluation samples. Which two visualizations are most appropriate for this goal? (Choose two.)
Medium23Which THREE of the following are considered best practices for visualizing LLM evaluation results to key stakeholders?
Medium24Which approach is most effective for visualizing 'attention heads' in a Transformer model to debug why the model ignores specific information?
Medium25An engineer needs to track validation loss, learning rate, and GPU utilization together over training steps for a fine-tuning run, and wants the ability to compare multiple runs side by side in a web dashboard. Which approach best meets this need?
Easy26Which visualization tool is most suitable for tracking the gradient norm evolution during the training of a large language model to detect vanishing or exploding gradients?
Easy27Refer to the exhibit. The model is failing with an OOM at layer 42 during training. What visualization would most likely point to the cause of the memory fragmentation?
Hard28You are comparing the inference throughput of an LLM served with two different batching strategies across a range of request arrival rates. You want a single visualization that shows both the median throughput and the variability at each arrival rate. Which visualization is most appropriate?
Medium29A data scientist is analyzing token length distribution across a 12-million-document pretraining corpus destined for an NVIDIA NCA-GENL pipeline. The histogram is heavily right-skewed with a long tail beyond 8,192 tokens. Which visualization should be produced NEXT to decide a safe max_sequence_length without discarding most of the corpus?
Medium30You need to compare the performance of two different LLMs on a set of benchmark tasks. Which visualization technique is most appropriate for a side-by-side comparison of multiple performance metrics (e.g., accuracy, latency, and truthfulness)?
Medium31During a fine-tuning run you observe that the training loss decreases smoothly, but validation loss begins rising after epoch 3. You want a single visualization that makes this divergence and the resulting overfitting point immediately obvious to reviewers. Which plot should you produce?
Hard32A developer is profiling an LLM inference endpoint on an NVIDIA L40S and observes that time-to-first-token (TTFT) is stable but inter-token latency spikes periodically. They want to determine whether the spikes align with KV cache growth or with batch-size changes. Which visualization strategy best isolates the cause?
Medium33A data scientist is analyzing token-level loss values produced by an LLM evaluation run on a summarization dataset. Losses are stored as a list of floats, and most values cluster around 2.1, but a few exceed 9.0. The team wants a visualization that shows the shape of the loss distribution, including those extreme values, without hiding them through bin aggregation. Which visualization is most appropriate?
Medium34Refer to the exhibit. How should a data scientist interpret this evaluation result regarding the recent model update?
Hard35A data scientist wants to track training loss, learning rate, and GPU utilization side by side across thousands of steps in an interactive dashboard that supports comparing multiple runs. Which tool is designed specifically for this experiment-tracking and interactive visualization workflow?
Easy36Which TWO of the following visualization techniques are most effective for identifying latent patterns in high-dimensional embedding spaces during LLM evaluation?
Medium37A team has 1,024-dimensional document embeddings from a retrieval corpus and needs an interactive visualization to explore semantic neighborhoods for debugging retrieval failures. They want to preserve both global structure and local neighborhoods as faithfully as possible while keeping the tool responsive during pan and zoom. Which approach best fits?
Hard38You are performing a comparative analysis of two different LLM architectures by visualizing their performance on a RAG (Retrieval-Augmented Generation) benchmark. Which visualization is best for comparing the distributions of answer accuracy scores?
Medium39You are analyzing a dataset of 50,000 LLM training samples and want to visualize how sample lengths are distributed to decide on a maximum sequence length cutoff. The lengths range from 10 to 8,000 tokens with a long right tail. Which visualization should you use to best reveal the shape, central tendency, and outliers of this single continuous variable?
Medium40A team is building a dashboard to monitor an LLM evaluation pipeline that scores model outputs against a reference dataset. They want the dashboard to support rapid diagnosis when a new model checkpoint regresses. Which TWO visualization practices best support that goal? (Choose two.)
Hard41A team is comparing two LLM checkpoints on a summarization benchmark. They want a single visualization that shows, for each evaluation metric, both the mean score and the spread across the benchmark's document categories, while making it easy to see whether the two checkpoints overlap. Which visualization best fits this requirement?
Hard42A data scientist is profiling an LLM inference service on NVIDIA GPUs and has collected per-request latency samples. The distribution has a long right tail caused by a small number of requests that queue behind large batches. Which pair of summary statistics BEST communicates both the typical experience and the tail pain to the engineering team?
Medium43A data scientist is preparing an exploratory report on a large corpus of prompt-completion pairs used to fine-tune an LLM. They want to visualize the distribution of a single numerical feature, prompt token count, to check for skew before choosing a tokenization budget. Which visualization is most appropriate?
Easy44A machine learning engineer is monitoring an LLM inference service deployed on NVIDIA GPUs. They want a real-time dashboard that shows GPU utilization, memory usage, and request latency, and they need to set alerts when thresholds are exceeded. Which NVIDIA tool is purpose-built for this monitoring and alerting?
Easy45A team is comparing two LLM fine-tuning runs on the same dataset. Run A used a cosine learning-rate schedule, and Run B used a constant learning rate. They plot validation loss versus training step for both runs on the same axes. Run A's curve is smooth, while Run B's curve shows a sharp upward spike around step 800 and then recovers. The team wants to determine whether the spike in Run B indicates a data-order artifact or a genuine optimization instability. Which additional visualization is most useful for that diagnosis?
Hard46You are analyzing embedding quality for a retrieval-augmented generation system. You have 1,000 document embeddings of 4,096 dimensions and want to inspect whether semantically similar documents form visible clusters. Which dimensionality-reduction approach is most appropriate before plotting in two dimensions?
Hard47During a RAG evaluation, a data scientist computes cosine similarity between 40,000 query embeddings and 40,000 retrieved-chunk embeddings using an NVIDIA-accelerated pipeline. They then reduce the 4,096-dimensional vectors with t-SNE to 2D for a scatter plot, but the plot shows no separation between relevant and irrelevant retrievals. What is the MOST likely reason the visualization fails to reveal the retrieval quality signal?
Hard48A team is preparing a stakeholder report on an LLM evaluation run. They must show how the model's accuracy on a question-answering benchmark changes as the temperature parameter is swept from 0.0 to 1.0 in steps of 0.1. Which visualization is MOST appropriate for this single-variable sweep?
Easy49A team is analyzing an LLM evaluation dataset with thousands of prompts and multiple scoring dimensions such as correctness, fluency, and safety. They want a single visualization that reveals how these dimensions correlate and whether any prompts score unusually on several dimensions at once. Which visualization is most suitable?
Medium50You are analyzing token frequency distribution across a 50 GB pretraining corpus before fine-tuning an NVIDIA NIM-deployed Llama model. The raw frequency histogram is heavily right-skewed, making it impossible to compare low-frequency tokens. Which transformation should you apply to the x-axis to make the distribution easier to compare across the full vocabulary?
Medium51A team is evaluating an LLM-based summarization service and wants a visualization that shows how the distribution of generated summary lengths compares to the reference summaries across 5,000 test articles. They want to see whether the model systematically produces shorter or longer outputs. Which visualization is best suited?
Easy52A team is fine-tuning an LLM and wants to detect whether individual training examples are causing unusually large gradient updates. They plan to visualize per-example gradient norms alongside other diagnostics. Which two visualizations are most appropriate for identifying these influential examples? (Choose two.)
Hard53Refer to the exhibit. What is the primary risk indicated by the provided logs for this training job?
Hard54Refer to the exhibit. A monitoring script outputs this JSON for an LLM inference service. What does the 'p99' metric represent in this context?
Hard55You are conducting an error analysis on an LLM's performance. Which THREE visualizations are most effective for identifying where the model struggles with factual accuracy in a RAG (Retrieval-Augmented Generation) pipeline?
Medium56You are performing exploratory data analysis on a massive dataset for an LLM training pipeline. You need to visualize the distribution of token frequencies in a corpus of 10 billion tokens. Which visualization technique is most effective for identifying long-tail patterns in power-law distributions typical of natural language data?
Medium57When evaluating a generative model, why is it important to visualize the distribution of output sequence lengths?
Medium58A team is preparing a dashboard to monitor an LLM inference service in production. They want visualizations that surface latency problems and resource saturation before users are affected. Which two visualizations are most appropriate for this goal? (Choose two.)
Medium59A team is evaluating a retrieval-augmented generation pipeline. They have a dataset of 500 queries, each with a retrieved context and a generated answer. The goal is to visualize how often the generated answer is faithful to the retrieved context versus hallucinated, and to compare this across three different retriever configurations. Which visualization best supports this comparison?
Hard60A team is building a dashboard to monitor an LLM training run and wants to detect data-quality problems early. They have access to per-batch training loss, per-batch gradient norm, input sequence length statistics, and token frequency counts. Which two visualizations are most appropriate for surfacing data-quality issues rather than hardware or throughput issues? (Choose two.)
Medium61You are monitoring a production LLM inference service on an NVIDIA GPU. The service's request latency distribution is heavily right-skewed, and a small fraction of requests take far longer than the rest. You need a visualization that shows the full distribution shape, including the median and the extreme tail, to decide whether the GPU is under-provisioned. Which visualization should you use?
Medium62You are analyzing token probability distributions from an LLM inference service to detect hallucination risk. You need a single visualization that shows, for one generated response, how the model's confidence evolved token-by-token and where it suddenly dropped. Which visualization is most appropriate?
Medium63You are performing a bias audit on a fine-tuned chat model. You need to visualize the model's responses to sensitive prompts across various demographic categories. Which visualization is most effective for identifying systemic bias?
Medium64A data scientist is building a dashboard to detect data drift in the input distribution of a production LLM endpoint. They have access to daily embedding vectors of incoming prompts and to the model's output token statistics. Which two visualizations are MOST appropriate for surfacing prompt-distribution drift over time? (Choose two.)
Hard65You are refining a dataset for a domain-specific LLM using NVIDIA NeMo. You want to visualize the similarity of documents to ensure your training set covers the required technical domains effectively. Which tool and visualization combination is best suited for this?
MediumOther domains
All NCA-GENL exam domains
Frequently asked questions
- What does the Data Analysis and Visualization domain cover on the NCA-GENL exam?
- Be able to read loss curves, gradient-norm diagnostics, and benchmark exhibits, then pick the visualization matching the comparison. The key is distinguishing transient training spikes from real divergence and choosing distribution-aware plots over summary statistics.
- How many questions are in this domain?
- This page lists all 65 Data Analysis and Visualization questions in the NCA-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Data Analysis and Visualization questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.