NCA-GENL · domain
Experimentation
The Experimentation domain covers how to design, run, and interpret LLM training and inference experiments on NVIDIA platforms. Questions are scenario-based: you diagnose side effects of changing hyperparameters like sequence length or batch size, interpret validation loss variance across repeated runs, and apply reproducibility controls such as fixed seeds and NeMo configuration management.
Focused practice
Practice Experimentation questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Experimentation
A candidate must be able to design controlled LLM experiments, interpret loss variance across repeated runs, and manage memory when scaling sequence length or batch size. The single most important thing: isolate one variable at a time and use seeds plus repeated runs to separate real effects from noise.
Managing memory and compute trade-offs when increasing sequence length in transformer training
Interpreting validation loss variance across repeated NeMo fine-tuning runs with identical configurations
Using fixed random seeds to make training runs reproducible and comparable
Diagnosing OOM errors from larger batch sizes on NVIDIA DGX systems despite GPU utilization headroom
Watch out for
Common Experimentation exam traps
- ▸Assuming a fixed seed guarantees identical results across different GPU counts, libraries, or NeMo versions, when nondeterminism can still arise.
- ▸Treating validation loss differences across repeated runs as real model improvements rather than run-to-run noise from initialization and data ordering.
- ▸Increasing batch size to fix OOM errors, when larger batches consume more memory and can worsen the problem.
Question index
All Experimentation questions (82)
Click any question to see the full explanation, or start a practice session above.
When fine-tuning a model on a new dataset, why is it important to keep a portion of the original pre-training data in the fine-tuning mix?
Medium2Refer to the exhibit. An engineer is monitoring a large model training job. Based on the sudden latency spike at step 502, what is the most likely cause during the experimentation phase?
Hard3An ML engineer is comparing two fine-tuned variants of the same base LLM on an internal question-answering benchmark. Variant A scores higher on exact-match accuracy, while Variant B produces answers that human reviewers rate as more helpful and better formatted. The team must choose one variant to ship to production. Which evaluation approach is most appropriate for making this decision?
Easy4You are running an experiment comparing two different fine-tuning methods (Full Fine-tuning vs. PEFT). Why is it crucial to keep the dataset and evaluation benchmark identical?
Medium5A team runs a NeMo fine-tuning experiment and observes that validation loss decreases for several epochs and then steadily rises while training loss keeps falling. They want to confirm whether the checkpoint from the best validation epoch is genuinely better than the final checkpoint. Which action provides the strongest evidence?
Hard6An ML engineer is using NVIDIA NeMo to evaluate a retrieval-augmented generation pipeline. They want to measure whether adding a reranker improves answer faithfulness, but they must ensure the experiment is reproducible and comparable across runs. Which practice best supports a valid comparison between the pipeline with and without the reranker?
Hard7In the context of NVIDIA NeMo, which THREE actions are part of a robust experiment tracking workflow for fine-tuning?
Hard8An AI engineer at an automotive enterprise is running an LLM experimentation pipeline using NeMo. The primary objective is to evaluate how different prompt engineering strategies affect the model's spatial reasoning capabilities across diverse spatial datasets. Which foundational workflow step should be prioritized to ensure reproducible experimental results?
Medium9Which of the following describes the purpose of a 'Validation Set' during the model experimentation cycle?
Medium10A data scientist wants to determine how sensitive an LLM's summarization quality is to the temperature sampling parameter. They plan a sweep across several temperature values. Which experimental approach gives the clearest signal about temperature's effect?
Easy11A research team is running a hyperparameter sweep over learning rate and warmup steps for a NeMo fine-tuning job. They notice that runs with identical hyperparameters produce different final validation losses across repeated executions. Which two changes would most directly improve the reproducibility of these experiments? (Choose two.)
Medium12Refer to the exhibit. You are experimenting with a model and find the validation loss is increasing while training loss decreases. Which parameter should you adjust first?
Hard13A research team is comparing two LoRA fine-tuning runs of the same Llama-based model in NeMo. Run 1 uses rank 8 and alpha 16; Run 2 uses rank 64 and alpha 16, with all other hyperparameters identical. Run 2 achieves lower training loss but worse accuracy on a held-out evaluation set. Which conclusion is most defensible from this experiment?
Hard14A data scientist is experimenting with an LLM for a text generation task using NVIDIA NeMo. They want to measure how the model's output diversity changes when adjusting the temperature parameter. They plan to generate 100 samples for each temperature setting and compute the distinct-n metric. Which experimental design principle are they applying?
Easy15A researcher is running a hyperparameter sweep over learning rate and batch size for an LLM fine-tune. To keep the experiment tractable, they want to prune unpromising trials early. Which approach best supports early stopping of poorly performing trials while preserving statistical validity?
Hard16During an experiment, the researcher decides to increase the model's sequence length. What is the most significant side effect they must manage?
Medium17When fine-tuning a model for domain-specific tasks, which THREE metrics should you monitor during the training phase to ensure the experiment is progressing healthily?
Medium18When experimenting with synthetic data generation to improve model performance, what is the most important risk to monitor?
Medium19A researcher is using NVIDIA's 'TensorRT-LLM' to optimize an LLM. During the experimentation phase, they observe the model's accuracy drops significantly after quantization. What is the most appropriate next step?
Hard20An engineer is evaluating different prompting strategies (Zero-shot, Few-shot, Chain-of-Thought) for an RAG pipeline. Which TWO metrics are most effective for quantifying the quality of the generative output during this experimentation?
Medium21During an LLM experimentation phase using NVIDIA NeMo, an ML engineer needs to systematically track hyperparameters, dataset lineage, and evaluation artifacts to meet strict auditing standards. Which TWO actions should the engineer take to achieve comprehensive experiment tracking?
Hard22An AI researcher is fine-tuning a Llama-3 model using NeMo Framework and notices high GPU memory usage during training. Which experimentation technique is most effective for reducing memory footprint without sacrificing model quality?
Medium23An engineer is designing an experiment to measure how prompt phrasing affects the output quality of a deployed LLM. They will test several prompt templates against a fixed evaluation set and want the comparison to be valid. Which two practices are required? (Choose two.)
Medium24Refer to the exhibit. You are running a multi-node distributed fine-tuning experiment and receive this error. What does this indicate about your experimentation environment?
Hard25A data science team is running a controlled experiment with NVIDIA NeMo to compare two fine-tuning recipes for a 7B-parameter LLM: one with a constant learning rate and one with a cosine decay schedule. They notice the evaluation loss curves diverge significantly after step 500, but they cannot tell whether the difference is caused by the learning-rate schedule or by random seed variance. Which experimental change should they make to isolate the effect of the schedule?
Medium26A team is running an A/B experiment comparing two prompt templates for a customer-facing LLM assistant. After one week, template A shows a 2% higher task-completion rate with a p-value of 0.04. The team lead wants to declare A the winner immediately. Which consideration is most important before making that decision?
Hard27In the experimentation loop, what is the role of a 'validation split' during model fine-tuning?
Medium28During a fine-tuning experiment in NVIDIA NeMo, validation loss begins to rise after epoch 4 while training loss continues to fall. The team wants to determine the earliest epoch at which the model still generalizes well. Which experimental action is most appropriate?
Hard29In an experiment comparing different fine-tuning methods (LoRA vs. Full Fine-tuning), which metric is most useful for determining the efficiency of the experimentation process itself?
Medium30A team is fine-tuning a NeMo Megatron GPT model on an internal corpus and observes that validation loss begins rising after epoch three while training loss continues to fall. They want to detect this condition automatically during future experiments without manually watching the curves. Which NeMo callback or mechanism should they configure to stop training when validation loss stops improving?
Medium31A team's NeMo fine-tuning experiment runs on a fixed compute budget and they must choose how to allocate it between searching hyperparameters and training the final model. Their hyperparameter search space is large and each trial is expensive. Which allocation strategy best balances finding a strong configuration against producing a well-trained final model?
Hard32An ML engineer is setting up an experiment log for a fine-tuning run and wants to record the metadata necessary to reproduce the resulting model later. Which two items are most essential to capture for reproducibility? (Choose two.)
Medium33When conducting an experiment to tune the 'Top-P' (Nucleus Sampling) parameter for a text generation task, what is the primary goal of the researcher?
Easy34An ML team is running an ablation study with NVIDIA NeMo to determine which components of their LLM pipeline contribute most to answer quality. They remove one component at a time and re-evaluate. After several runs, they notice that removing the retrieval component causes a large drop in quality, but removing the reranker causes almost no change. What is the most reasonable interpretation of this result?
Hard35A data scientist is running a fine-tuning experiment with NVIDIA NeMo on a single A100 GPU. They want to establish a repeatable baseline before sweeping any hyperparameters, so that a later run can be compared fairly. Which practice best supports this goal?
Easy36An enterprise AI researcher is conducting ablation studies on a large language model using NVIDIA NeMo. To ensure the experimentation results are scientifically valid and statistically sound, which THREE practices must be enforced during the study?
Hard37A team has fine-tuned a small LLM with NeMo and now wants to quantify how much the fine-tuning improved performance on a domain question-answering task relative to the base model. They have a curated set of 500 question-answer pairs that were never used during training. What is the most appropriate next step?
Easy38You are experimenting with RAG and notice the model is frequently ignoring the provided context. Which of the following is the most likely culprit to investigate first?
Medium39A researcher is running an ablation study in which they vary the number of attention heads in a NeMo Megatron GPT model while holding parameter count, dataset, and learning rate fixed. After the first run, they change tensor parallel size and pipeline parallel size to fit larger variants on the available GPUs. A colleague argues this invalidates the comparison. Which statement best explains the scientific concern?
Hard40When experimenting with model quantization (e.g., INT8 or FP8), what is the most important trade-off to monitor?
Medium41An engineer at a customer-support automation company is experimenting with top-p sampling values for a NeMo-served LLM. They want to quantify how output diversity changes across settings without relying on human judgment alone. Which evaluation approach best supports this experiment?
Medium42Which of the following is the primary goal of the 'Experimentation' phase in an LLM project?
Easy43An engineer is setting up an automated experiment sweep over temperature and top-p for a NeMo-served LLM, and wants the results to be comparable and reproducible. Which two practices should be applied? (Choose two.)
Medium44Which of the following is a primary objective of 'Ablation Studies' in LLM experimentation?
Easy45An engineer is evaluating a RAG-based assistant and wants to isolate whether retrieval quality or the generator is responsible for wrong answers. They build a small labeled set of questions with known correct passages and known correct answers. Which experimental design most cleanly separates the contribution of the retriever from that of the generator?
Hard46A data science team is fine-tuning a Llama 3 8B model on a proprietary customer-support corpus using NVIDIA NeMo. They need to run dozens of experiments with different learning rates and batch sizes. Because the dataset contains personally identifiable information, they cannot send any telemetry to an external tracking server, but they still need to compare runs later and reproduce the best configuration. Which approach best satisfies both the reproducibility and data-privacy requirements?
Medium47An AI researcher is designing an experiment to compare two prompt templates for a customer-support LLM using NVIDIA NeMo. To ensure the comparison is fair and reproducible, which two practices should they follow? (Choose two.)
Medium48During an ablation study on a retrieval-augmented LLM in NeMo, an engineer removes the reranking stage and observes that answer accuracy drops by 12 points, but latency improves by 40 percent. A stakeholder asks whether reranking should be kept. Which experimental next step best supports a defensible recommendation?
Hard49What is the primary benefit of tracking experiments using a centralized experiment management platform (e.g., Weights & Biases, MLflow)?
Easy50An AI engineer at a financial services company is running an LLM experimentation pipeline using NVIDIA NeMo. The primary objective is to evaluate how different tokenizers affect the accuracy of a named entity recognition (NER) task on financial documents. The engineer has already fixed the model architecture, the training dataset, and the hyperparameters. To ensure the experiment isolates the effect of the tokenizer, which action should the engineer take next?
Medium51A team is comparing two fine-tuned variants of the same base LLM for a customer-support summarization task. Variant A was trained on 10,000 examples and variant B on 2,000 examples drawn from the same distribution. To make a fair comparison of their generalization, which evaluation practice should the team adopt?
Medium52A data scientist is preparing to perform hyperparameter tuning for a Retrieval-Augmented Generation (RAG) system. Which TWO parameters should be prioritized for experimentation to improve retrieval accuracy?
Medium53A research team is using NVIDIA NeMo to experiment with a large language model for a summarization task. They observe that the model's ROUGE scores vary significantly across different runs even when using the same hyperparameters and dataset. They suspect that non-deterministic operations in the training pipeline are causing this variance. Which step should they take to improve reproducibility of their experiment results?
Hard54Refer to the exhibit. The experiment fails with an Out-of-Memory (OOM) error during the second epoch. Given the configuration, which change is most effective for immediate stabilization?
Hard55Which of the following best describes the role of 'AB testing' in Generative AI experimentation?
Easy56Which TWO factors should be considered when evaluating the cost-benefit of an LLM experimentation strategy?
Medium57A research team is designing an experiment to measure how prompt phrasing affects the factuality of an LLM in a retrieval-augmented question-answering pipeline. Which two design choices are necessary to attribute observed factuality differences to the prompt rather than to other pipeline components? (Choose two.)
Hard58An ML engineer is comparing two fine-tuned variants of the same base LLM on an internal question-answering benchmark. Variant 1 scores higher on exact-match accuracy, but Variant 2 produces answers that human reviewers judge more helpful and better grounded. The engineer must decide which variant to promote. Which evaluation approach is most appropriate?
Medium59You are designing an experiment to measure how quantization (FP16 versus INT8) affects inference latency and answer quality for an LLM deployed with NVIDIA TensorRT-LLM. Which two practices are required for the comparison to be valid? (Choose two.)
Medium60A team fine-tuning a NeMo large language model runs the same training configuration three times and obtains validation loss values of 2.14, 2.31, and 2.09 at the end of the same number of steps. They need subsequent runs to produce tightly clustered, comparable numbers so hyperparameter comparisons are meaningful. Which change most directly addresses this problem?
Medium61You are running a NeMo fine-tuning experiment where validation loss decreases for the first three epochs, then rises steadily while training loss keeps falling. You want to confirm overfitting and select the most appropriate intervention. Which experiment action should you take first?
Medium62An engineer must decide how to split a labeled dataset of 50,000 customer support conversations before fine-tuning a NeMo LLM for intent classification. The goal is an honest estimate of how the tuned model will behave on never-before-seen tickets once deployed. Which splitting approach best supports that goal?
Easy63An ML engineer at a healthcare analytics company is starting a fine-tuning experiment on a Llama 2 7B model using NVIDIA NeMo Framework. Before launching the training job, the engineer wants a single immutable record that captures the exact model checkpoint, dataset version, hyperparameters, and evaluation scores so that any later run can be traced back to it. Which component of the NVIDIA NeMo experimentation workflow should the engineer use to store that record?
Easy64A researcher is experimenting with prompt-tuning and finds that the model output is repetitive. They decide to adjust the sampling hyperparameters. Which combination of changes is most likely to increase the diversity of the output?
Medium65During LLM experimentation, what is the primary purpose of maintaining a consistent 'seed' value across different runs?
Easy66An AI researcher is testing a new LLM architecture on an NVIDIA DGX system. They observe that increasing the batch size leads to memory OOM errors despite available GPU utilization headroom. Which experimentation strategy should be employed first to isolate the bottleneck?
Medium67During an ablation study, a team removes the instruction-tuning stage from their NeMo pipeline and observes that the model still answers factual questions but frequently ignores the requested output format. They want to attribute this change in behavior to the removed stage rather than to noise. Which experimental design element is most important for supporting that attribution?
Medium68A research team is comparing three fine-tuning recipes for a NeMo LLM and wants the comparison to be defensible in a later review. Which two practices most improve the credibility of the reported comparison? (Choose two.)
Hard69Refer to the exhibit. The experiment shows the model is failing to converge and exhibits loss spikes. Which adjustment to the configuration is most likely to stabilize the training process?
Hard70A team is running an LLM fine-tuning experiment using NVIDIA NeMo and wants to track how the validation loss changes over training. They need a reliable way to detect overfitting early. Which metric should they monitor most directly during the experiment?
Easy71A data scientist is preparing an LLM fine-tuning experiment on NVIDIA NeMo and wants every run to be reproducible weeks later. The team's experiment tracker currently logs only the final validation loss. Which additional item is most important to record so a run can be reproduced exactly?
Easy72A team is designing a controlled experiment to measure whether increasing LoRA rank improves instruction-following accuracy on a held-out benchmark. They want the comparison to be scientifically valid. Which experimental design choice best supports a valid conclusion?
Hard73An AI engineer is conducting an experiment to compare two different fine-tuning approaches for a large language model using NVIDIA NeMo: full fine-tuning versus parameter-efficient fine-tuning (PEFT) with LoRA. The engineer wants to determine which approach yields better performance on a downstream question-answering task while minimizing computational cost. Which metric should the engineer prioritize to evaluate the trade-off between performance and cost?
Medium74A team is evaluating an LLM for a customer-support summarization task. They want to compare three prompt templates. Which experimental design most directly isolates the effect of the prompt template?
Easy75Which THREE factors influence the reproducibility of an LLM experiment?
Hard76A research team is running a hyperparameter sweep for LoRA fine-tuning of a 13B-parameter LLM on NVIDIA GPUs. They notice that runs with identical configurations sometimes produce noticeably different evaluation scores, and the variance is larger than the differences between some of the hyperparameter settings being compared. Which action best addresses this problem?
Hard77When conducting A/B testing on an LLM-based application, which metric is the most reliable indicator of user satisfaction regarding response quality?
Easy78A research team is running a series of controlled LLM fine-tuning experiments on NVIDIA DGX systems using NeMo Framework to compare two learning-rate schedules. They want the comparison to be scientifically valid and repeatable by another engineer next quarter. Which two practices are required to make the experiments reproducible? (Choose two.)
Medium79Which of the following describes the 'Stop-Loss' technique in the context of LLM experimentation?
Easy80An ML engineer is running a hyperparameter sweep with NVIDIA NeMo and notices that runs with identical configurations sometimes produce slightly different final loss values. Which cause should the engineer investigate first?
Medium81When designing an experiment to evaluate the performance of an LLM on a downstream classification task, which THREE factors should be controlled to ensure the results are comparable across different model sizes?
Hard82During an ablation study on a retrieval-augmented LLM pipeline, the team removes the reranking stage and observes a large drop in answer accuracy on their benchmark. Before concluding that reranking is essential, which additional experiment is most important to run?
HardOther domains
All NCA-GENL exam domains
Frequently asked questions
- What does the Experimentation domain cover on the NCA-GENL exam?
- A candidate must be able to design controlled LLM experiments, interpret loss variance across repeated runs, and manage memory when scaling sequence length or batch size. The single most important thing: isolate one variable at a time and use seeds plus repeated runs to separate real effects from noise.
- How many questions are in this domain?
- This page lists all 82 Experimentation questions in the NCA-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Experimentation questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.