Courseiva

NCA-GENL · domain

Experimentation

The Experimentation domain covers how to design, run, and interpret LLM training and inference experiments on NVIDIA platforms. Questions are scenario-based: you diagnose side effects of changing hyperparameters like sequence length or batch size, interpret validation loss variance across repeated runs, and apply reproducibility controls such as fixed seeds and NeMo configuration management.

82 questions18 easy36 medium28 hard

Focused practice

Practice Experimentation questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Experimentation

A candidate must be able to design controlled LLM experiments, interpret loss variance across repeated runs, and manage memory when scaling sequence length or batch size. The single most important thing: isolate one variable at a time and use seeds plus repeated runs to separate real effects from noise.

Managing memory and compute trade-offs when increasing sequence length in transformer training

Interpreting validation loss variance across repeated NeMo fine-tuning runs with identical configurations

Using fixed random seeds to make training runs reproducible and comparable

Diagnosing OOM errors from larger batch sizes on NVIDIA DGX systems despite GPU utilization headroom

Watch out for

Common Experimentation exam traps

  • ▸Assuming a fixed seed guarantees identical results across different GPU counts, libraries, or NeMo versions, when nondeterminism can still arise.
  • ▸Treating validation loss differences across repeated runs as real model improvements rather than run-to-run noise from initialization and data ordering.
  • ▸Increasing batch size to fix OOM errors, when larger batches consume more memory and can worsen the problem.

Question index

All Experimentation questions (82)

Click any question to see the full explanation, or start a practice session above.

1

When fine-tuning a model on a new dataset, why is it important to keep a portion of the original pre-training data in the fine-tuning mix?

Medium
2

Refer to the exhibit. An engineer is monitoring a large model training job. Based on the sudden latency spike at step 502, what is the most likely cause during the experimentation phase?

Hard
3

An ML engineer is comparing two fine-tuned variants of the same base LLM on an internal question-answering benchmark. Variant A scores higher on exact-match accuracy, while Variant B produces answers that human reviewers rate as more helpful and better formatted. The team must choose one variant to ship to production. Which evaluation approach is most appropriate for making this decision?

Easy
4

You are running an experiment comparing two different fine-tuning methods (Full Fine-tuning vs. PEFT). Why is it crucial to keep the dataset and evaluation benchmark identical?

Medium
5

A team runs a NeMo fine-tuning experiment and observes that validation loss decreases for several epochs and then steadily rises while training loss keeps falling. They want to confirm whether the checkpoint from the best validation epoch is genuinely better than the final checkpoint. Which action provides the strongest evidence?

Hard
6

An ML engineer is using NVIDIA NeMo to evaluate a retrieval-augmented generation pipeline. They want to measure whether adding a reranker improves answer faithfulness, but they must ensure the experiment is reproducible and comparable across runs. Which practice best supports a valid comparison between the pipeline with and without the reranker?

Hard
7

In the context of NVIDIA NeMo, which THREE actions are part of a robust experiment tracking workflow for fine-tuning?

Hard
8

An AI engineer at an automotive enterprise is running an LLM experimentation pipeline using NeMo. The primary objective is to evaluate how different prompt engineering strategies affect the model's spatial reasoning capabilities across diverse spatial datasets. Which foundational workflow step should be prioritized to ensure reproducible experimental results?

Medium
9

Which of the following describes the purpose of a 'Validation Set' during the model experimentation cycle?

Medium
10

A data scientist wants to determine how sensitive an LLM's summarization quality is to the temperature sampling parameter. They plan a sweep across several temperature values. Which experimental approach gives the clearest signal about temperature's effect?

Easy
11

A research team is running a hyperparameter sweep over learning rate and warmup steps for a NeMo fine-tuning job. They notice that runs with identical hyperparameters produce different final validation losses across repeated executions. Which two changes would most directly improve the reproducibility of these experiments? (Choose two.)

Medium
12

Refer to the exhibit. You are experimenting with a model and find the validation loss is increasing while training loss decreases. Which parameter should you adjust first?

Hard
13

A research team is comparing two LoRA fine-tuning runs of the same Llama-based model in NeMo. Run 1 uses rank 8 and alpha 16; Run 2 uses rank 64 and alpha 16, with all other hyperparameters identical. Run 2 achieves lower training loss but worse accuracy on a held-out evaluation set. Which conclusion is most defensible from this experiment?

Hard
14

A data scientist is experimenting with an LLM for a text generation task using NVIDIA NeMo. They want to measure how the model's output diversity changes when adjusting the temperature parameter. They plan to generate 100 samples for each temperature setting and compute the distinct-n metric. Which experimental design principle are they applying?

Easy
15

A researcher is running a hyperparameter sweep over learning rate and batch size for an LLM fine-tune. To keep the experiment tractable, they want to prune unpromising trials early. Which approach best supports early stopping of poorly performing trials while preserving statistical validity?

Hard
16

During an experiment, the researcher decides to increase the model's sequence length. What is the most significant side effect they must manage?

Medium
17

When fine-tuning a model for domain-specific tasks, which THREE metrics should you monitor during the training phase to ensure the experiment is progressing healthily?

Medium
18

When experimenting with synthetic data generation to improve model performance, what is the most important risk to monitor?

Medium
19

A researcher is using NVIDIA's 'TensorRT-LLM' to optimize an LLM. During the experimentation phase, they observe the model's accuracy drops significantly after quantization. What is the most appropriate next step?

Hard
20

An engineer is evaluating different prompting strategies (Zero-shot, Few-shot, Chain-of-Thought) for an RAG pipeline. Which TWO metrics are most effective for quantifying the quality of the generative output during this experimentation?

Medium
21

During an LLM experimentation phase using NVIDIA NeMo, an ML engineer needs to systematically track hyperparameters, dataset lineage, and evaluation artifacts to meet strict auditing standards. Which TWO actions should the engineer take to achieve comprehensive experiment tracking?

Hard
22

An AI researcher is fine-tuning a Llama-3 model using NeMo Framework and notices high GPU memory usage during training. Which experimentation technique is most effective for reducing memory footprint without sacrificing model quality?

Medium
23

An engineer is designing an experiment to measure how prompt phrasing affects the output quality of a deployed LLM. They will test several prompt templates against a fixed evaluation set and want the comparison to be valid. Which two practices are required? (Choose two.)

Medium
24

Refer to the exhibit. You are running a multi-node distributed fine-tuning experiment and receive this error. What does this indicate about your experimentation environment?

Hard
25

A data science team is running a controlled experiment with NVIDIA NeMo to compare two fine-tuning recipes for a 7B-parameter LLM: one with a constant learning rate and one with a cosine decay schedule. They notice the evaluation loss curves diverge significantly after step 500, but they cannot tell whether the difference is caused by the learning-rate schedule or by random seed variance. Which experimental change should they make to isolate the effect of the schedule?

Medium
26

A team is running an A/B experiment comparing two prompt templates for a customer-facing LLM assistant. After one week, template A shows a 2% higher task-completion rate with a p-value of 0.04. The team lead wants to declare A the winner immediately. Which consideration is most important before making that decision?

Hard
27

In the experimentation loop, what is the role of a 'validation split' during model fine-tuning?

Medium
28

During a fine-tuning experiment in NVIDIA NeMo, validation loss begins to rise after epoch 4 while training loss continues to fall. The team wants to determine the earliest epoch at which the model still generalizes well. Which experimental action is most appropriate?

Hard
29

In an experiment comparing different fine-tuning methods (LoRA vs. Full Fine-tuning), which metric is most useful for determining the efficiency of the experimentation process itself?

Medium
30

A team is fine-tuning a NeMo Megatron GPT model on an internal corpus and observes that validation loss begins rising after epoch three while training loss continues to fall. They want to detect this condition automatically during future experiments without manually watching the curves. Which NeMo callback or mechanism should they configure to stop training when validation loss stops improving?

Medium
31

A team's NeMo fine-tuning experiment runs on a fixed compute budget and they must choose how to allocate it between searching hyperparameters and training the final model. Their hyperparameter search space is large and each trial is expensive. Which allocation strategy best balances finding a strong configuration against producing a well-trained final model?

Hard
32

An ML engineer is setting up an experiment log for a fine-tuning run and wants to record the metadata necessary to reproduce the resulting model later. Which two items are most essential to capture for reproducibility? (Choose two.)

Medium
33

When conducting an experiment to tune the 'Top-P' (Nucleus Sampling) parameter for a text generation task, what is the primary goal of the researcher?

Easy
34

An ML team is running an ablation study with NVIDIA NeMo to determine which components of their LLM pipeline contribute most to answer quality. They remove one component at a time and re-evaluate. After several runs, they notice that removing the retrieval component causes a large drop in quality, but removing the reranker causes almost no change. What is the most reasonable interpretation of this result?

Hard
35

A data scientist is running a fine-tuning experiment with NVIDIA NeMo on a single A100 GPU. They want to establish a repeatable baseline before sweeping any hyperparameters, so that a later run can be compared fairly. Which practice best supports this goal?

Easy
36

An enterprise AI researcher is conducting ablation studies on a large language model using NVIDIA NeMo. To ensure the experimentation results are scientifically valid and statistically sound, which THREE practices must be enforced during the study?

Hard
37

A team has fine-tuned a small LLM with NeMo and now wants to quantify how much the fine-tuning improved performance on a domain question-answering task relative to the base model. They have a curated set of 500 question-answer pairs that were never used during training. What is the most appropriate next step?

Easy
38

You are experimenting with RAG and notice the model is frequently ignoring the provided context. Which of the following is the most likely culprit to investigate first?

Medium
39

A researcher is running an ablation study in which they vary the number of attention heads in a NeMo Megatron GPT model while holding parameter count, dataset, and learning rate fixed. After the first run, they change tensor parallel size and pipeline parallel size to fit larger variants on the available GPUs. A colleague argues this invalidates the comparison. Which statement best explains the scientific concern?

Hard
40

When experimenting with model quantization (e.g., INT8 or FP8), what is the most important trade-off to monitor?

Medium
41

An engineer at a customer-support automation company is experimenting with top-p sampling values for a NeMo-served LLM. They want to quantify how output diversity changes across settings without relying on human judgment alone. Which evaluation approach best supports this experiment?

Medium
42

Which of the following is the primary goal of the 'Experimentation' phase in an LLM project?

Easy
43

An engineer is setting up an automated experiment sweep over temperature and top-p for a NeMo-served LLM, and wants the results to be comparable and reproducible. Which two practices should be applied? (Choose two.)

Medium
44

Which of the following is a primary objective of 'Ablation Studies' in LLM experimentation?

Easy
45

An engineer is evaluating a RAG-based assistant and wants to isolate whether retrieval quality or the generator is responsible for wrong answers. They build a small labeled set of questions with known correct passages and known correct answers. Which experimental design most cleanly separates the contribution of the retriever from that of the generator?

Hard
46

A data science team is fine-tuning a Llama 3 8B model on a proprietary customer-support corpus using NVIDIA NeMo. They need to run dozens of experiments with different learning rates and batch sizes. Because the dataset contains personally identifiable information, they cannot send any telemetry to an external tracking server, but they still need to compare runs later and reproduce the best configuration. Which approach best satisfies both the reproducibility and data-privacy requirements?

Medium
47

An AI researcher is designing an experiment to compare two prompt templates for a customer-support LLM using NVIDIA NeMo. To ensure the comparison is fair and reproducible, which two practices should they follow? (Choose two.)

Medium
48

During an ablation study on a retrieval-augmented LLM in NeMo, an engineer removes the reranking stage and observes that answer accuracy drops by 12 points, but latency improves by 40 percent. A stakeholder asks whether reranking should be kept. Which experimental next step best supports a defensible recommendation?

Hard
49

What is the primary benefit of tracking experiments using a centralized experiment management platform (e.g., Weights & Biases, MLflow)?

Easy
50

An AI engineer at a financial services company is running an LLM experimentation pipeline using NVIDIA NeMo. The primary objective is to evaluate how different tokenizers affect the accuracy of a named entity recognition (NER) task on financial documents. The engineer has already fixed the model architecture, the training dataset, and the hyperparameters. To ensure the experiment isolates the effect of the tokenizer, which action should the engineer take next?

Medium
51

A team is comparing two fine-tuned variants of the same base LLM for a customer-support summarization task. Variant A was trained on 10,000 examples and variant B on 2,000 examples drawn from the same distribution. To make a fair comparison of their generalization, which evaluation practice should the team adopt?

Medium
52

A data scientist is preparing to perform hyperparameter tuning for a Retrieval-Augmented Generation (RAG) system. Which TWO parameters should be prioritized for experimentation to improve retrieval accuracy?

Medium
53

A research team is using NVIDIA NeMo to experiment with a large language model for a summarization task. They observe that the model's ROUGE scores vary significantly across different runs even when using the same hyperparameters and dataset. They suspect that non-deterministic operations in the training pipeline are causing this variance. Which step should they take to improve reproducibility of their experiment results?

Hard
54

Refer to the exhibit. The experiment fails with an Out-of-Memory (OOM) error during the second epoch. Given the configuration, which change is most effective for immediate stabilization?

Hard
55

Which of the following best describes the role of 'AB testing' in Generative AI experimentation?

Easy
56

Which TWO factors should be considered when evaluating the cost-benefit of an LLM experimentation strategy?

Medium
57

A research team is designing an experiment to measure how prompt phrasing affects the factuality of an LLM in a retrieval-augmented question-answering pipeline. Which two design choices are necessary to attribute observed factuality differences to the prompt rather than to other pipeline components? (Choose two.)

Hard
58

An ML engineer is comparing two fine-tuned variants of the same base LLM on an internal question-answering benchmark. Variant 1 scores higher on exact-match accuracy, but Variant 2 produces answers that human reviewers judge more helpful and better grounded. The engineer must decide which variant to promote. Which evaluation approach is most appropriate?

Medium
59

You are designing an experiment to measure how quantization (FP16 versus INT8) affects inference latency and answer quality for an LLM deployed with NVIDIA TensorRT-LLM. Which two practices are required for the comparison to be valid? (Choose two.)

Medium
60

A team fine-tuning a NeMo large language model runs the same training configuration three times and obtains validation loss values of 2.14, 2.31, and 2.09 at the end of the same number of steps. They need subsequent runs to produce tightly clustered, comparable numbers so hyperparameter comparisons are meaningful. Which change most directly addresses this problem?

Medium
61

You are running a NeMo fine-tuning experiment where validation loss decreases for the first three epochs, then rises steadily while training loss keeps falling. You want to confirm overfitting and select the most appropriate intervention. Which experiment action should you take first?

Medium
62

An engineer must decide how to split a labeled dataset of 50,000 customer support conversations before fine-tuning a NeMo LLM for intent classification. The goal is an honest estimate of how the tuned model will behave on never-before-seen tickets once deployed. Which splitting approach best supports that goal?

Easy
63

An ML engineer at a healthcare analytics company is starting a fine-tuning experiment on a Llama 2 7B model using NVIDIA NeMo Framework. Before launching the training job, the engineer wants a single immutable record that captures the exact model checkpoint, dataset version, hyperparameters, and evaluation scores so that any later run can be traced back to it. Which component of the NVIDIA NeMo experimentation workflow should the engineer use to store that record?

Easy
64

A researcher is experimenting with prompt-tuning and finds that the model output is repetitive. They decide to adjust the sampling hyperparameters. Which combination of changes is most likely to increase the diversity of the output?

Medium
65

During LLM experimentation, what is the primary purpose of maintaining a consistent 'seed' value across different runs?

Easy
66

An AI researcher is testing a new LLM architecture on an NVIDIA DGX system. They observe that increasing the batch size leads to memory OOM errors despite available GPU utilization headroom. Which experimentation strategy should be employed first to isolate the bottleneck?

Medium
67

During an ablation study, a team removes the instruction-tuning stage from their NeMo pipeline and observes that the model still answers factual questions but frequently ignores the requested output format. They want to attribute this change in behavior to the removed stage rather than to noise. Which experimental design element is most important for supporting that attribution?

Medium
68

A research team is comparing three fine-tuning recipes for a NeMo LLM and wants the comparison to be defensible in a later review. Which two practices most improve the credibility of the reported comparison? (Choose two.)

Hard
69

Refer to the exhibit. The experiment shows the model is failing to converge and exhibits loss spikes. Which adjustment to the configuration is most likely to stabilize the training process?

Hard
70

A team is running an LLM fine-tuning experiment using NVIDIA NeMo and wants to track how the validation loss changes over training. They need a reliable way to detect overfitting early. Which metric should they monitor most directly during the experiment?

Easy
71

A data scientist is preparing an LLM fine-tuning experiment on NVIDIA NeMo and wants every run to be reproducible weeks later. The team's experiment tracker currently logs only the final validation loss. Which additional item is most important to record so a run can be reproduced exactly?

Easy
72

A team is designing a controlled experiment to measure whether increasing LoRA rank improves instruction-following accuracy on a held-out benchmark. They want the comparison to be scientifically valid. Which experimental design choice best supports a valid conclusion?

Hard
73

An AI engineer is conducting an experiment to compare two different fine-tuning approaches for a large language model using NVIDIA NeMo: full fine-tuning versus parameter-efficient fine-tuning (PEFT) with LoRA. The engineer wants to determine which approach yields better performance on a downstream question-answering task while minimizing computational cost. Which metric should the engineer prioritize to evaluate the trade-off between performance and cost?

Medium
74

A team is evaluating an LLM for a customer-support summarization task. They want to compare three prompt templates. Which experimental design most directly isolates the effect of the prompt template?

Easy
75

Which THREE factors influence the reproducibility of an LLM experiment?

Hard
76

A research team is running a hyperparameter sweep for LoRA fine-tuning of a 13B-parameter LLM on NVIDIA GPUs. They notice that runs with identical configurations sometimes produce noticeably different evaluation scores, and the variance is larger than the differences between some of the hyperparameter settings being compared. Which action best addresses this problem?

Hard
77

When conducting A/B testing on an LLM-based application, which metric is the most reliable indicator of user satisfaction regarding response quality?

Easy
78

A research team is running a series of controlled LLM fine-tuning experiments on NVIDIA DGX systems using NeMo Framework to compare two learning-rate schedules. They want the comparison to be scientifically valid and repeatable by another engineer next quarter. Which two practices are required to make the experiments reproducible? (Choose two.)

Medium
79

Which of the following describes the 'Stop-Loss' technique in the context of LLM experimentation?

Easy
80

An ML engineer is running a hyperparameter sweep with NVIDIA NeMo and notices that runs with identical configurations sometimes produce slightly different final loss values. Which cause should the engineer investigate first?

Medium
81

When designing an experiment to evaluate the performance of an LLM on a downstream classification task, which THREE factors should be controlled to ensure the results are comparable across different model sizes?

Hard
82

During an ablation study on a retrieval-augmented LLM pipeline, the team removes the reranking stage and observes a large drop in answer accuracy on their benchmark. Before concluding that reranking is essential, which additional experiment is most important to run?

Hard

Frequently asked questions

What does the Experimentation domain cover on the NCA-GENL exam?
A candidate must be able to design controlled LLM experiments, interpret loss variance across repeated runs, and manage memory when scaling sequence length or batch size. The single most important thing: isolate one variable at a time and use seeds plus repeated runs to separate real effects from noise.
How many questions are in this domain?
This page lists all 82 Experimentation questions in the NCA-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Experimentation questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
nvidia-nca-genl NVIDIA-NCA-GENL nca experimentation Practice Questions