Courseiva

CCNA Experimentation Questions

75 of 82 questions · Page 1/2 · Experimentation · Answers revealed

1
MCQmedium

When fine-tuning a model on a new dataset, why is it important to keep a portion of the original pre-training data in the fine-tuning mix?

A.To increase the total number of training epochs
B.To prevent catastrophic forgetting
C.To reduce the computational time of fine-tuning
D.To improve the hardware utilization rates
AnswerB

Mixing a small fraction of original pre-training data ensures the model remains anchored to its general knowledge base. Without this 'replay' technique, the model tends to overwrite its pre-trained weights with specific new patterns, resulting in a loss of general-purpose capabilities that were present before the fine-tuning process started.

Why this answer

Retaining pre-training data during fine-tuning prevents 'catastrophic forgetting,' where the model loses its general knowledge while adapting to new tasks. This practice is essential for maintaining the model's capabilities in reasoning, coding, or language fluency. For NVIDIA NCA-GENL standards, understanding how to preserve base model utility while specializing for specific domains is a critical skill for successful model lifecycle management.

Exam trap

Candidates often assume fine-tuning is only about learning new data and ignore the risk of losing existing capabilities. They forget the model might 'forget' how to perform basic tasks.

2
MCQhard

Refer to the exhibit. An engineer is monitoring a large model training job. Based on the sudden latency spike at step 502, what is the most likely cause during the experimentation phase?

A.The model has reached a local minimum in the loss function.
B.A periodic checkpointing operation triggered at step 502.
C.The learning rate scheduler reduced the step size.
D.The model architecture was automatically reconfigured.
AnswerB

Periodic I/O operations such as saving model weights to disk or synchronizing distributed state often cause transient latency spikes. Because the latency doubled at step 502, it is highly indicative of a blocking I/O operation or a synchronization barrier that is occurring at regular training intervals.

Why this answer

The sudden latency jump suggests a periodic operation like checkpointing, data logging, or a hardware-level thermal throttling event. In LLM training, frequent checkpoints or massive synchronization steps are primary suspects for sudden, brief stalls. Identifying these spikes early allows developers to tune checkpoint frequency or optimize I/O paths, ensuring that experiments maintain consistent performance and avoid unnecessary overhead during training.

Exam trap

Examinees often attribute sudden latency spikes to complex network bottlenecks or gradient explosions, overlooking routine systems maintenance tasks like periodic model checkpointing.

3
MCQeasy

An ML engineer is comparing two fine-tuned variants of the same base LLM on an internal question-answering benchmark. Variant A scores higher on exact-match accuracy, while Variant B produces answers that human reviewers rate as more helpful and better formatted. The team must choose one variant to ship to production. Which evaluation approach is most appropriate for making this decision?

A.Combine automated metrics with a structured human-preference evaluation, and choose based on the product's primary success criterion.
B.Average the exact-match score and a normalized human-preference score into a single number and pick the higher value.
C.Ship Variant A because exact-match accuracy is an objective metric and human ratings are subjective.
D.Ship Variant B because human preference is the only metric that matters for generative models.
AnswerA

Generative LLM evaluation typically pairs task-specific automated metrics with human or model-based preference judgments, because neither alone captures the full quality picture. Weighting them according to the product's success criterion, such as helpfulness for a user-facing assistant, yields a defensible shipping decision. This approach uses both signals the scenario provides.

Why this answer

For generative LLM tasks, automated metrics and human preference capture different quality dimensions. The right decision combines both and weights them by the product's primary success criterion. Shipping on accuracy alone ignores user-perceived helpfulness, shipping on preference alone ignores objective regressions, and averaging without justified weights hides the real trade-off.

Exam trap

The trap here is treating exact-match accuracy as inherently superior because it is numeric, when generative quality often depends on human-perceived helpfulness.

4
MCQmedium

You are running an experiment comparing two different fine-tuning methods (Full Fine-tuning vs. PEFT). Why is it crucial to keep the dataset and evaluation benchmark identical?

A.To minimize the computational cost of the experiment.
B.To ensure that performance differences reflect the tuning method.
C.To prevent the model from crashing due to memory issues.
D.To speed up the training time of the PEFT approach.
AnswerB

Keeping evaluation benchmarks constant ensures that differences in model performance can be attributed to the tuning method (Full vs. PEFT). This isolates the independent variable, allowing the researcher to draw clear, evidence-based conclusions about which method is better for their specific application requirements.

Why this answer

Scientific rigor demands that all variables except the independent variable be controlled. If the dataset or evaluation criteria differ, it becomes impossible to determine if the performance variance is caused by the fine-tuning method or by the inherent differences in the evaluation content. Consistent evaluation acts as a 'control' for the experiment, ensuring that the findings regarding the effectiveness of the fine-tuning technique are statistically valid and replicable.

Exam trap

Students frequently think changing multiple experimental variables speeds up optimization, overlooking the fundamental scientific requirement of controlling variables to isolate fine-tuning method impacts.

5
MCQhard

A team runs a NeMo fine-tuning experiment and observes that validation loss decreases for several epochs and then steadily rises while training loss keeps falling. They want to confirm whether the checkpoint from the best validation epoch is genuinely better than the final checkpoint. Which action provides the strongest evidence?

A.Compare the training loss values of the two checkpoints
B.Evaluate both checkpoints on a separate held-out test set and compare task metrics
C.Continue training the final checkpoint for more epochs and recheck validation loss
D.Pick whichever checkpoint has the lower validation loss
AnswerB

Validation loss guided model selection, so it can be optimistically biased for the checkpoint chosen by that same metric. A separate held-out test set that played no role in selecting the checkpoint gives an unbiased estimate of generalization. Comparing task metrics on this set directly answers whether the earlier checkpoint is truly better rather than an artifact of selection.

Why this answer

When a checkpoint is selected using validation loss, that metric becomes optimistically biased for the chosen model. The strongest evidence comes from an untouched test set that was not involved in any selection decision. Evaluating both checkpoints there with task-relevant metrics yields an unbiased comparison, whereas training loss and repeated validation checks cannot settle the question.

Exam trap

The trap here is reusing the validation metric that drove checkpoint selection as if it were an independent judge, which inflates confidence in the selected model.

6
MCQhard

An ML engineer is using NVIDIA NeMo to evaluate a retrieval-augmented generation pipeline. They want to measure whether adding a reranker improves answer faithfulness, but they must ensure the experiment is reproducible and comparable across runs. Which practice best supports a valid comparison between the pipeline with and without the reranker?

A.Increase the temperature of the LLM in the reranker variant to encourage more diverse answers and better faithfulness.
B.Evaluate the reranker variant on a larger and more diverse dataset than the baseline so the results are more statistically significant.
C.Run the baseline and reranker variants on different GPU models to test robustness across hardware configurations.
D.Use the same evaluation dataset, same base LLM checkpoint, and same decoding parameters for both pipeline variants, changing only the reranker component.
AnswerD

A valid A/B comparison requires that only the component under test—the reranker—differs between conditions. Keeping the evaluation dataset, base checkpoint, and decoding parameters identical ensures that any change in faithfulness scores is attributable to the reranker rather than to data, model, or sampling differences. This is the controlled-variable principle applied to RAG pipeline experimentation.

Why this answer

To attribute a change in faithfulness to the reranker, every other element of the RAG pipeline must be held constant. Using the same evaluation set, base checkpoint, and decoding parameters for both variants ensures the only difference is the presence or absence of the reranker. This controlled design yields a clean, reproducible comparison and avoids confounding variables that would invalidate the result.

Exam trap

The trap here is thinking that a larger dataset or more compute makes an experiment more valid, when the real requirement is holding all non-tested variables constant across conditions.

7
Multi-Selecthard

In the context of NVIDIA NeMo, which THREE actions are part of a robust experiment tracking workflow for fine-tuning?

Select 3 answers
A.Logging hyperparameters for every run
B.Deleting logs to conserve disk space
C.Saving model checkpoints periodically
D.Automating metric collection via tools like W&B
E.Manually calculating gradients during training
AnswersA, C, D

Logging hyperparameters is necessary to reproduce experiments. Without knowing the exact settings like learning rate, optimizer parameters, and batch size, it is impossible to verify why a specific run achieved its results, making it difficult to improve performance iteratively or justify the configuration to stakeholders in a professional environment.

Why this answer

Robust experiment tracking is the foundation of reproducibility in machine learning. By logging configurations, monitoring metrics in real-time using tools like Weights & Biases or TensorBoard, and saving versioned checkpoints, researchers can compare results across iterations. These actions are vital for ensuring that performance gains are attributable to specific hyperparameter changes rather than random chance or environmental variations during training.

Exam trap

Candidates often select manual tracking approaches or assume that logging only the final model output is sufficient, ignoring the crucial need for continuous metric collection, hyperparameters, and versioned checkpoints during iterative workflows.

8
MCQmedium

An AI engineer at an automotive enterprise is running an LLM experimentation pipeline using NeMo. The primary objective is to evaluate how different prompt engineering strategies affect the model's spatial reasoning capabilities across diverse spatial datasets. Which foundational workflow step should be prioritized to ensure reproducible experimental results?

A.Maximizing GPU batch size to accelerate throughput during the prompt evaluation phase.
B.Utilizing mixed precision FP16 computations to reduce memory footprint across nodes.
C.Fixing random seeds across all libraries and versioning the evaluation datasets.
D.Upgrading the cluster interconnect fabric to InfiniBand for faster inter-GPU communication.
AnswerC

Controlling stochasticity through explicit random seed initialization in frameworks like PyTorch and NeMo ensures that generation behavior remains identical across runs. Coupled with strict dataset versioning, this guarantees that performance deltas are driven exclusively by prompt variations.

Why this answer

Establishing strict deterministic seeds and dataset versioning is foundational for reliable LLM experimentation. Without fixed random seeds and tracked dataset states, variability in generation outputs prevents meaningful comparison between distinct prompt strategies, undermining the scientific validity of the enterprise validation process.

Exam trap

Candidates often focus on model architecture or hardware specs, forgetting that reproducibility in AI experiments is primarily achieved through controlling random seeds and data versions.

9
MCQmedium

Which of the following describes the purpose of a 'Validation Set' during the model experimentation cycle?

A.It is used for the final performance evaluation of the model.
B.It provides a mechanism to tune hyperparameters iteratively.
C.It replaces the training set to reduce compute usage.
D.It acts as a buffer to store temporary model weights.
AnswerB

The validation set serves as an independent benchmark for comparing different model versions and hyperparameter settings. By evaluating on this set during the training process, the researcher can make informed decisions about which architectural or parameter changes actually improve the model's ability to generalize to new data.

Why this answer

The validation set is used during training to monitor performance and tune hyperparameters without leaking information from the test set. It acts as an objective checkpoint for model improvement, allowing the researcher to stop training if overfitting occurs. Proper use of this set is a hallmark of robust experimentation, ensuring that the final model generalizes to unseen data in real-world production environments.

Exam trap

Test-takers frequently confuse the validation set with the test set, mistakenly believing validation data is used for final unbiased model evaluation rather than iterative hyperparameter tuning.

10
MCQeasy

A data scientist wants to determine how sensitive an LLM's summarization quality is to the temperature sampling parameter. They plan a sweep across several temperature values. Which experimental approach gives the clearest signal about temperature's effect?

A.Use a different evaluation metric for each temperature value to capture more aspects of quality.
B.Change temperature and the prompt template together to explore the joint space faster.
C.Test only the lowest and highest temperature values to save compute.
D.Vary temperature while keeping the model, prompts, and evaluation metric constant.
AnswerD

Temperature is the factor under test, so isolating it by holding the model, prompts, and metric fixed lets any quality change be attributed to sampling temperature. This is the direct way to measure sensitivity. The sweep then reveals whether quality is stable, improves, or degrades across the temperature range.

Why this answer

To measure sensitivity to a single parameter, that parameter must be the only thing changing. Fixing the model, prompts, and evaluation metric ensures any quality difference across the sweep comes from temperature. This controlled setup produces a curve that shows whether summarization quality is robust or fragile with respect to the sampling temperature setting.

Exam trap

The trap here is treating a faster joint sweep of temperature and prompt as equivalent, when it removes the ability to attribute results to temperature alone.

11
Multi-Selectmedium

A research team is running a hyperparameter sweep over learning rate and warmup steps for a NeMo fine-tuning job. They notice that runs with identical hyperparameters produce different final validation losses across repeated executions. Which two changes would most directly improve the reproducibility of these experiments? (Choose two.)

Select 2 answers
A.Log the loss curve to TensorBoard so the team can visually compare runs.
B.Enable deterministic kernels and disable nondeterministic cuDNN algorithms in the framework settings.
C.Set a fixed random seed for data shuffling, weight initialization, and dropout in the training configuration.
D.Increase the number of GPUs per run so each experiment finishes faster.
E.Reduce the learning rate by a factor of ten across all sweep configurations.
AnswersB, C

Certain GPU kernels, especially in cuDNN, are nondeterministic by design for performance reasons, causing small numerical differences that accumulate over training steps. Enabling deterministic algorithms removes this source of variance, making repeated runs with the same seed produce identical or near-identical results. This is a standard reproducibility control.

Why this answer

Run-to-run variance in identical configurations usually comes from uncontrolled randomness in data order, initialization, dropout, and nondeterministic GPU kernels. Fixing the random seed and enabling deterministic kernels address both the algorithmic and hardware-level sources. Adding GPUs, logging, or lowering the learning rate change the experiment or its observability but do not make repeated runs converge.

Exam trap

The trap here is assuming that more logging or more hardware improves reproducibility, when the actual cause is uncontrolled stochasticity in the training pipeline.

12
MCQhard

Refer to the exhibit. You are experimenting with a model and find the validation loss is increasing while training loss decreases. Which parameter should you adjust first?

A.Change activation function to ReLU
B.Increase dropout
C.Change optimizer to SGD
D.Decrease weight decay
AnswerB

Increasing dropout is a direct way to regularize the model. By randomly dropping neurons during training, you ensure the model doesn't over-rely on any single path, helping it generalize better to unseen data. This is the most appropriate first step when validation loss shows signs of training-time overfitting.

Why this answer

The divergence between training loss and validation loss is a classic sign of overfitting. Increasing the 'dropout' value is a highly effective way to introduce noise during training, preventing the model from relying on specific neuron activations. This forces more robust feature learning and is a standard first-line defense in the experimentation process to bridge the gap between training and validation performance.

Exam trap

Candidates frequently try to fix overfitting by adjusting the learning rate or adding more training epochs, rather than directly applying regularization techniques like dropout.

13
MCQhard

A research team is comparing two LoRA fine-tuning runs of the same Llama-based model in NeMo. Run 1 uses rank 8 and alpha 16; Run 2 uses rank 64 and alpha 16, with all other hyperparameters identical. Run 2 achieves lower training loss but worse accuracy on a held-out evaluation set. Which conclusion is most defensible from this experiment?

A.The two runs cannot be compared because LoRA rank changes the base model architecture.
B.Higher LoRA rank always improves downstream accuracy, so the evaluation set must be mislabeled.
C.The alpha value should be doubled for Run 2 to restore the intended scaling ratio.
D.The rank-64 run overfits the training data, so the lower training loss does not translate to better generalization.
AnswerD

Increasing LoRA rank raises the number of trainable parameters in the adapter, giving the model more capacity to fit the training set. When training loss drops but held-out accuracy worsens, the extra capacity has been used to memorize rather than generalize. The rank-8 adapter is more constrained and therefore generalizes better on this evaluation set. The experiment supports the overfitting interpretation rather than a labeling error.

Why this answer

The pattern of lower training loss with worse held-out accuracy indicates that the higher-rank adapter used its additional capacity to fit training-specific noise. LoRA rank sets the dimensionality of the update matrices, so rank 64 has more trainable parameters than rank 8. The defensible conclusion is that the larger adapter overfit, and the more constrained adapter generalized better on this evaluation set.

Exam trap

The trap here is treating lower training loss as proof of a better model, when a widening gap between training loss and held-out accuracy signals overfitting instead.

14
MCQeasy

A data scientist is experimenting with an LLM for a text generation task using NVIDIA NeMo. They want to measure how the model's output diversity changes when adjusting the temperature parameter. They plan to generate 100 samples for each temperature setting and compute the distinct-n metric. Which experimental design principle are they applying?

A.Ablation study
B.Cross-validation
C.Hyperparameter optimization
D.Controlled experiment
AnswerD

By varying only the temperature while keeping other factors constant, the scientist is conducting a controlled experiment. This design isolates the effect of temperature on output diversity, allowing valid conclusions. Generating multiple samples and computing distinct-n provides a quantitative measure, which is characteristic of controlled experimentation in LLM evaluation.

Why this answer

The scientist is manipulating a single independent variable (temperature) while holding other factors constant, and measuring its effect on a dependent variable (output diversity via distinct-n). This systematic approach is a controlled experiment, which allows for causal inference about the relationship between temperature and diversity. It is a fundamental design principle in LLM experimentation.

Exam trap

The trap here is confusing controlled experimentation with hyperparameter optimization; the former seeks to understand effects, while the latter seeks to find optimal values.

15
MCQhard

A researcher is running a hyperparameter sweep over learning rate and batch size for an LLM fine-tune. To keep the experiment tractable, they want to prune unpromising trials early. Which approach best supports early stopping of poorly performing trials while preserving statistical validity?

A.Stop any trial whose training loss has not decreased in the last 100 steps and discard its results.
B.Run all trials to full length but evaluate them on a smaller validation set to save time.
C.Eliminate trials whose first-epoch loss is above the median of all trials, without further evaluation.
D.Use a successive halving or Hyperband-style scheduler that allocates small budgets to many trials and promotes only the best performers to larger budgets.
AnswerD

Successive halving and Hyperband evaluate many configurations with small resource budgets, then promote the top performers to larger budgets. This concentrates compute on promising trials while still exploring broadly early on. It is a principled early-stopping strategy that preserves the ability to identify strong hyperparameter regions.

Why this answer

Successive halving and Hyperband explicitly trade exploration for exploitation by giving many configurations a small budget and progressively promoting the best ones. This preserves the chance to discover strong hyperparameters while avoiding full-length runs for clearly weak trials. Compared with fixed patience or first-epoch thresholds, the promotion structure is more robust to differences in loss-curve shape across hyperparameters.

Exam trap

The trap here is treating early stopping as a fixed loss-plateau rule rather than a budget-allocation strategy across many trials.

16
MCQmedium

During an experiment, the researcher decides to increase the model's sequence length. What is the most significant side effect they must manage?

A.The model will automatically converge in fewer training steps.
B.The memory requirement for the attention matrix will grow quadratically.
C.The learning rate must be increased to maintain stability.
D.The vocabulary size must be increased to accommodate new tokens.
AnswerB

Standard self-attention mechanisms require storing a $N \times N$ matrix where $N$ is the sequence length. Doubling the sequence length quadruples the memory required for the attention scores. This quadratic growth is a primary constraint that researchers must mitigate during experimentation through techniques like attention optimization or memory-efficient kernels.

Why this answer

Increasing sequence length in transformers typically leads to a quadratic increase in memory usage for the attention mechanism. This necessitates strategies like FlashAttention or model parallelism to prevent OOM errors. Understanding this trade-off is fundamental to the experimentation process, as it dictates the physical constraints and architectural choices available when building models that handle longer inputs for complex reasoning tasks.

Exam trap

Test-takers often assume memory scales linearly with sequence length, forgetting the quadratic complexity inherent in the transformer's self-attention mechanism.

17
Multi-Selectmedium

When fine-tuning a model for domain-specific tasks, which THREE metrics should you monitor during the training phase to ensure the experiment is progressing healthily?

Select 3 answers
A.Training Loss
B.Validation Loss
C.Model disk usage
D.Gradient Norm
E.Total number of GPUs used
AnswersA, B, D

Training loss is the primary indicator that the model is learning from the provided data. If it fails to decrease, the learning rate might be too low, or the model architecture might be misconfigured, providing immediate feedback during the early stages of the training experiment.

Why this answer

Monitoring training dynamics is crucial for detecting issues early. Training loss confirms the optimizer is reducing error, while validation loss catches overfitting before it becomes permanent. Gradient norm tracking is the most reliable way to identify instability—such as the infamous 'exploding gradients'—which is common in complex LLM training.

Together, these metrics form the dashboard necessary for a data scientist to make informed decisions about when to stop, adjust, or continue their experiments.

Exam trap

Candidates often select output-based metrics like BLEU or ROUGE instead of training-specific diagnostic metrics. These output metrics are for evaluation, not for monitoring the health of the training process itself.

18
MCQmedium

When experimenting with synthetic data generation to improve model performance, what is the most important risk to monitor?

A.The increase in training time.
B.The introduction of model bias or hallucinations.
C.The file format of the training data.
D.The cost of disk storage.
AnswerB

Synthetic data generated by an LLM is prone to hallucination and biases. If this data is used for fine-tuning, the student model may amplify these errors, leading to degraded performance. Monitoring for quality and factuality in the synthetic dataset is critical to prevent the model from learning incorrect patterns.

Why this answer

Synthetic data can inadvertently contain artifacts or biases present in the teacher model. Over time, 'model collapse'—where the model learns from its own generated noise—can occur, leading to a degradation in performance. Regular evaluation on a held-out, human-verified dataset is essential to ensure that the synthetic data is actually improving the model's ability to reason, rather than just forcing it to mirror the stylistic flaws of the synthetic data source.

Exam trap

Candidates often focus solely on the speed and cost advantages of synthetic data generation while ignoring the gradual accumulation of compounding model biases and hallucinations.

19
MCQhard

A researcher is using NVIDIA's 'TensorRT-LLM' to optimize an LLM. During the experimentation phase, they observe the model's accuracy drops significantly after quantization. What is the most appropriate next step?

A.Revert to full FP32 training without any optimization.
B.Increase the model's hidden dimension size.
C.Use a representative calibration dataset for quantization.
D.Switch to a smaller base model architecture.
AnswerC

Quantization parameters are often determined by the distribution of activation values. Using a representative calibration dataset allows the algorithm to estimate the optimal scale and zero-point values more accurately, which significantly reduces the performance degradation that typically occurs when models are converted to lower precision.

Why this answer

Post-training quantization often introduces errors that degrade model accuracy. To mitigate this, techniques like 'Quantization-Aware Training' (QAT) or using a calibration dataset are essential. These methods help the model adapt to the lower precision format during or after the process.

Mastering these techniques is critical for delivering high-performance, resource-efficient models that maintain their accuracy in production environments.

Exam trap

Candidates often suggest re-training the whole model or changing the architecture. They overlook the standard, less compute-intensive solution of using calibration data for quantization adjustment.

20
Multi-Selectmedium

An engineer is evaluating different prompting strategies (Zero-shot, Few-shot, Chain-of-Thought) for an RAG pipeline. Which TWO metrics are most effective for quantifying the quality of the generative output during this experimentation?

Select 2 answers
A.Faithfulness score
B.Average token generation speed
C.Answer Relevance score
D.Number of model parameters
E.Total GPU memory consumption
AnswersA, C

Faithfulness measures whether the generated response is derived strictly from the retrieved context. This is critical in RAG experimentation to ensure the model does not hallucinate information outside the provided documents, maintaining accuracy and reliability for business applications where fact-based responses are required for user trust.

Why this answer

Quantitative evaluation is essential to move beyond subjective intuition in LLM experiments. Faithfulness (grounding in context) and Answer Relevance (usefulness to the query) provide distinct, measurable dimensions of performance. By measuring these, developers can iterate on prompts with empirical data, ensuring that changes to the prompt template actually improve system utility rather than just altering the verbosity or style of the generated response.

Exam trap

Candidates often choose general performance metrics like BLEU or ROUGE instead of RAG-specific metrics like Faithfulness and Answer Relevance, which specifically measure grounding against the retrieved context.

21
Multi-Selecthard

During an LLM experimentation phase using NVIDIA NeMo, an ML engineer needs to systematically track hyperparameters, dataset lineage, and evaluation artifacts to meet strict auditing standards. Which TWO actions should the engineer take to achieve comprehensive experiment tracking?

Select 2 answers
A.Integrate an experiment tracking framework like MLflow to log hyperparameters and metrics automatically.
B.Rely solely on manual terminal output logs stored in local temporary directories.
C.Disable all logging mechanisms to maximize GPU execution speed and prevent I/O bottlenecks.
D.Maintain a version-controlled artifact repository containing dataset snapshots and configuration files.
E.Store all intermediate checkpoints in volatile system memory without persisting to disk.
AnswersA, D

Experiment tracking tools capture essential metadata including learning rates, batch sizes, and validation metrics during every training or evaluation run. This integration enables engineers to visualize performance curves, compare trial results, and maintain a centralized registry of successful configurations.

Why this answer

Effective experiment tracking in enterprise generative AI requires systematic logging of both code configurations and artifact lineage. Integrating MLflow or NeMo-native logging captures hyperparameters, while artifact stores maintain dataset versions and model weights, ensuring full traceability and regulatory compliance during model evaluation.

Exam trap

Candidates often select only one action, such as just logging metrics. They overlook the necessity of version-controlling the actual dataset snapshots, which is required for full reproducibility and auditing.

22
MCQmedium

An AI researcher is fine-tuning a Llama-3 model using NeMo Framework and notices high GPU memory usage during training. Which experimentation technique is most effective for reducing memory footprint without sacrificing model quality?

A.Increase the learning rate significantly
B.Reduce the batch size to one
C.Enable gradient checkpointing
D.Switch to a smaller model architecture
AnswerC

Gradient checkpointing saves memory by discarding intermediate activations and recomputing them during the backward pass. This allows for training larger models or using larger batch sizes on constrained hardware. It is the industry-standard experimentation approach for managing memory overhead without compromising the mathematical integrity of the training process.

Why this answer

Gradient checkpointing is a standard technique in large model experimentation that trades computation time for memory efficiency. By storing only a subset of activations during the forward pass and recomputing others during the backward pass, it enables training larger models or larger batch sizes within the same VRAM constraints. This is critical for scaling experiments when hardware resources are restricted during initial prototyping phases.

Exam trap

Candidates frequently suggest reducing model size or quantization immediately, missing that gradient checkpointing specifically saves activation memory without altering weights.

23
Multi-Selectmedium

An engineer is designing an experiment to measure how prompt phrasing affects the output quality of a deployed LLM. They will test several prompt templates against a fixed evaluation set and want the comparison to be valid. Which two practices are required? (Choose two.)

Select 2 answers
A.Retrain the model after each prompt template so it adapts to the new phrasing.
B.Use a different evaluation dataset for each prompt template so the results are not correlated.
C.Increase the temperature for each successive template to explore more diverse outputs.
D.Hold the model, decoding parameters, and evaluation dataset constant across all prompt templates.
E.Score every template's outputs with the same rubric and scoring procedure.
AnswersD, E

Controlling model, decoding settings, and evaluation data ensures the only difference between runs is the prompt template, which is the variable under study. If any of these drift, an observed quality change could be caused by the model or sampling rather than the phrasing, invalidating the comparison the experiment is meant to support.

Why this answer

A valid prompt experiment isolates the prompt template as the only variable. Keeping the model, decoding parameters, and evaluation dataset fixed means any quality difference can be attributed to phrasing, and scoring all outputs with the same rubric ensures the measurement itself does not shift. Changing datasets, retraining, or varying temperature introduces confounds that make the comparison meaningless.

Exam trap

The trap here is assuming that varying other factors like temperature or dataset adds useful coverage, when in a controlled prompt comparison every factor except the prompt must stay constant.

24
MCQhard

Refer to the exhibit. You are running a multi-node distributed fine-tuning experiment and receive this error. What does this indicate about your experimentation environment?

A.The model weights are corrupted.
B.The learning rate is too high.
C.There is an issue with the cluster interconnect.
D.The batch size is too large.
AnswerC

NCCL relies on stable network communication to synchronize gradients across distributed ranks. A 'Connection reset by peer' indicates that the network path between nodes was interrupted. This is a common infrastructure error in multi-node experimentation, highlighting a failure in the hardware networking fabric rather than the training script.

Why this answer

This NCCL (NVIDIA Collective Communications Library) error signifies a breakdown in inter-node communication. In distributed training, nodes must synchronize gradients; a reset connection implies that the network fabric is failing to maintain the persistent connections required for this synchronization. This is a critical infrastructure issue that prevents the experiment from proceeding, requiring a check of the interconnects, such as InfiniBand or Ethernet switches, before any further experimentation can occur.

Exam trap

Candidates often misidentify this as a software bug or a code-level exception. They attempt to debug the model architecture or hyperparameters instead of addressing the underlying physical network infrastructure and connectivity issues.

25
MCQmedium

A data science team is running a controlled experiment with NVIDIA NeMo to compare two fine-tuning recipes for a 7B-parameter LLM: one with a constant learning rate and one with a cosine decay schedule. They notice the evaluation loss curves diverge significantly after step 500, but they cannot tell whether the difference is caused by the learning-rate schedule or by random seed variance. Which experimental change should they make to isolate the effect of the schedule?

A.Run both recipes with a fixed random seed and identical data ordering, then compare the loss curves.
B.Switch both recipes to a warmup-stable-decay schedule and compare the final evaluation loss instead of the curves.
C.Increase the batch size for both recipes until the loss curves become smoother and easier to compare visually.
D.Run each recipe on a different GPU type to see whether hardware differences explain the divergence.
AnswerA

Holding the random seed and data ordering constant removes the confounding effect of initialization and batch-order variance, so any remaining divergence between the constant and cosine schedules can be attributed to the learning-rate schedule itself. This is the core principle of a controlled experiment: change one factor at a time while controlling all others, including seeds and data shuffling.

Why this answer

A controlled experiment requires isolating the independent variable—here the learning-rate schedule—while holding all other factors constant. Random seed and data ordering are common sources of run-to-run variance in LLM fine-tuning. Fixing them across both recipes ensures that observed differences in evaluation loss are caused by the schedule rather than initialization or batch order, making the comparison valid.

Exam trap

The trap here is assuming that smoother curves or more data automatically make an experiment conclusive, when the real issue is uncontrolled seed and data-order variance confounding the comparison.

26
MCQhard

A team is running an A/B experiment comparing two prompt templates for a customer-facing LLM assistant. After one week, template A shows a 2% higher task-completion rate with a p-value of 0.04. The team lead wants to declare A the winner immediately. Which consideration is most important before making that decision?

A.Whether the two prompt templates were written by different engineers, since authorship affects output quality.
B.Whether the experiment was stopped early or the sample size was fixed in advance, because repeatedly checking and stopping inflates false-positive rates.
C.Whether the assistant uses a temperature greater than zero, since sampling randomness invalidates A/B tests.
D.Whether the p-value was computed with a one-tailed or two-tailed test, since one-tailed tests are always invalid.
AnswerB

Peeking at results and stopping as soon as significance appears inflates the Type I error rate well above the nominal threshold. A p-value of 0.04 from an optional-stopping procedure may not reflect a true 4% false-positive probability. Before declaring a winner, the team must confirm the analysis plan was pre-registered or apply a sequential-testing correction.

Why this answer

A p-value of 0.04 is fragile if the team monitored results continuously and stopped at the first significant reading, because optional stopping inflates the false-positive rate. The critical check is whether the sample size and stopping rule were fixed in advance or whether a sequential correction applies. Test direction, author identity, and sampling temperature are secondary concerns.

Exam trap

The trap here is treating a p-value just below 0.05 as a definitive result without questioning whether the data-collection process allowed the team to stop at a favorable moment.

27
MCQmedium

In the experimentation loop, what is the role of a 'validation split' during model fine-tuning?

A.To increase the total amount of training data
B.To provide an unbiased evaluation of generalization
C.To speed up the backpropagation process
D.To serve as the final test set for deployment
AnswerB

The validation set allows the researcher to see how the model performs on data it has not seen during the optimization process. This is the only way to detect overfitting or poor generalization, ensuring that the model's performance improvements are real and not just the result of memorizing the training set.

Why this answer

A validation split is used to monitor performance on unseen data during the training process, providing a metric for generalization. Unlike the training set, which the model directly optimizes, the validation set acts as an objective check. This prevents developers from making decisions based on overfitting, ensuring that the model maintains its utility on real-world data and identifying when to stop the training process.

Exam trap

Candidates often mistake the validation split for a way to improve training speed or accuracy, rather than understanding its primary purpose as an objective metric for evaluating model generalization.

28
MCQhard

During a fine-tuning experiment in NVIDIA NeMo, validation loss begins to rise after epoch 4 while training loss continues to fall. The team wants to determine the earliest epoch at which the model still generalizes well. Which experimental action is most appropriate?

A.Train for more epochs to let validation loss eventually decrease again.
B.Increase the learning rate so the model escapes the overfitting region faster.
C.Reduce the size of the validation set so the measured validation loss becomes less noisy.
D.Enable checkpoint saving at every epoch and select the checkpoint with the lowest validation loss for downstream evaluation.
AnswerD

Rising validation loss while training loss falls is the classic signature of overfitting. Saving checkpoints each epoch and selecting the one with minimum validation loss captures the model at its best generalization point. This directly identifies the earliest epoch that still generalizes well, which is the team's stated goal.

Why this answer

The divergence between falling training loss and rising validation loss indicates the model is memorizing training data. The practical remedy in an experimentation context is to checkpoint each epoch and choose the model with the lowest validation loss, which corresponds to the point of best generalization. This yields the earliest useful epoch without altering the training dynamics.

Exam trap

The trap here is thinking that continuing to train will eventually bring validation loss back down.

29
MCQmedium

In an experiment comparing different fine-tuning methods (LoRA vs. Full Fine-tuning), which metric is most useful for determining the efficiency of the experimentation process itself?

A.The total number of parameters in the model.
B.The final training loss at the end of the experiment.
C.The compute hours required to reach a target validation score.
D.The frequency of the model checkpointing during the run.
AnswerC

Compute hours normalized by performance targets are the standard measure for comparing the efficiency of different training methodologies. This metric allows researchers to quantify the trade-off between the reduced resource demands of parameter-efficient methods like LoRA and the potential quality gains of full fine-tuning approaches.

Why this answer

When comparing fine-tuning techniques, efficiency metrics like 'Time-to-convergence' or 'Compute-efficiency-per-epoch' are vital. These metrics quantify the resource cost of achieving a target accuracy, allowing researchers to choose the most cost-effective approach for their specific hardware. This is essential in an industrial setting where GPU hours and time-to-market are significant constraints for project viability.

Exam trap

Candidates often select 'accuracy' or 'loss' as the efficiency metric. These measure model quality, not the efficiency of the *process* of experimentation itself.

30
MCQmedium

A team is fine-tuning a NeMo Megatron GPT model on an internal corpus and observes that validation loss begins rising after epoch three while training loss continues to fall. They want to detect this condition automatically during future experiments without manually watching the curves. Which NeMo callback or mechanism should they configure to stop training when validation loss stops improving?

A.NeMo EarlyStopping callback monitoring validation loss
B.TensorRT-LLM quantization calibration pass
C.NeMo ModelCheckpoint with save_top_k set to a negative value
D.NVIDIA Nsight Compute kernel replay
AnswerA

The NeMo EarlyStopping callback watches a monitored metric such as validation loss and halts training when no improvement is seen for a configured patience period. This directly addresses the observed divergence between training and validation loss by ending the run automatically, saving compute and preventing further overfitting in future experiments.

Why this answer

The EarlyStopping callback in NeMo Framework is designed to monitor a validation metric and stop training after a configurable patience window without improvement. That matches the scenario of validation loss rising while training loss falls, letting the team end runs automatically. Profilers, quantization passes, and checkpoint retention settings do not influence when training stops, so they cannot address the overfitting signal.

Exam trap

The trap here is conflating checkpoint-saving behavior with training termination, assuming that a checkpoint configuration can also stop an overfitting run.

31
MCQhard

A team's NeMo fine-tuning experiment runs on a fixed compute budget and they must choose how to allocate it between searching hyperparameters and training the final model. Their hyperparameter search space is large and each trial is expensive. Which allocation strategy best balances finding a strong configuration against producing a well-trained final model?

A.Spend the entire budget on a dense grid search over every hyperparameter combination.
B.Run many short trials with identical settings to reduce measurement noise, then pick any configuration.
C.Train one configuration to completion and skip hyperparameter search entirely.
D.Use a budget-aware search such as successive halving or Bayesian optimization, then spend the remaining budget training the best configuration to completion.
AnswerD

Budget-aware search allocates few resources to clearly poor trials and progressively more to promising ones, which is efficient when trials are expensive. Reserving budget to fully train the winning configuration ensures the final model is not under-trained, balancing exploration against the quality of the delivered artifact.

Why this answer

With a fixed budget and expensive trials, the efficient path is to let early results prune weak candidates and concentrate resources on promising ones, then commit remaining budget to fully training the winner. This avoids both the waste of exhaustive grid search and the risk of delivering an under-trained final model.

Exam trap

The trap here is treating hyperparameter search and final training as separate unlimited activities, when in a fixed budget every trial spent searching is budget unavailable for producing the final model.

32
Multi-Selectmedium

An ML engineer is setting up an experiment log for a fine-tuning run and wants to record the metadata necessary to reproduce the resulting model later. Which two items are most essential to capture for reproducibility? (Choose two.)

Select 2 answers
A.The names of the engineers who reviewed the training logs.
B.The random seed used for data shuffling and initialization.
C.The wall-clock duration of each training epoch.
D.The GPU model and driver version used during training.
E.The exact dataset version or snapshot identifier used for training and validation.
AnswersB, E

The random seed controls data ordering and parameter initialization, both of which materially affect the trained result. Without recording it, a later rerun may produce a different model even with identical code and data. Capturing the seed is therefore one of the minimum metadata items needed for a reproducible fine-tuning experiment.

Why this answer

Reproducibility requires capturing the inputs that determine the trained weights: the randomness source and the exact training data. The seed governs shuffling and initialization, and the dataset version pins the examples used. Together they let another engineer rerun the same configuration and obtain a comparable model, which is the practical definition of a reproducible experiment.

Exam trap

The trap here is confusing operational telemetry such as epoch duration or reviewer names with the configuration metadata that actually determines model weights.

33
MCQeasy

When conducting an experiment to tune the 'Top-P' (Nucleus Sampling) parameter for a text generation task, what is the primary goal of the researcher?

A.To increase the training speed of the model.
B.To control the randomness and diversity of model outputs.
C.To reduce the physical VRAM footprint of the model.
D.To change the number of hidden layers in the model.
AnswerB

Top-P sampling limits the sampling pool to the smallest set of tokens whose cumulative probability exceeds the threshold P. By adjusting this, researchers can control how 'narrow' or 'broad' the model's choices are, effectively balancing the trade-off between repetitive, safe outputs and creative, diverse text.

Why this answer

Top-P sampling allows the model to select from a dynamic subset of the probability mass, which helps balance diversity and coherence. During experimentation, researchers adjust this value to find the 'sweet spot' that minimizes hallucinations while maintaining output creativity. This parameter tuning is a foundational practice for optimizing model behavior for specific use cases like creative writing or technical documentation generation.

Exam trap

Students often confuse Top-P with Top-K or temperature, mistakenly thinking it restricts the exact number of top tokens rather than dynamically adjusting based on cumulative probability mass.

34
MCQhard

An ML team is running an ablation study with NVIDIA NeMo to determine which components of their LLM pipeline contribute most to answer quality. They remove one component at a time and re-evaluate. After several runs, they notice that removing the retrieval component causes a large drop in quality, but removing the reranker causes almost no change. What is the most reasonable interpretation of this result?

A.Retrieval and the reranker are equally important, but the reranker's effect is masked by the retriever.
B.Retrieval is a critical contributor to quality in this pipeline, while the reranker adds little measurable value under the current evaluation setup.
C.The reranker is broken and must be replaced with a different model before any conclusion can be drawn.
D.The evaluation metric is too noisy to detect the reranker's effect, so the experiment should be discarded.
AnswerB

In an ablation study, the size of the performance drop when a component is removed indicates that component's contribution. A large drop from removing retrieval shows it is essential; a negligible drop from removing the reranker suggests it is not improving quality on this evaluation set. This is exactly the kind of insight ablation studies are designed to produce.

Why this answer

Ablation studies estimate each component's marginal contribution by removing it and measuring the performance change. A large drop when retrieval is removed indicates it is essential; a negligible drop when the reranker is removed indicates it adds little value under the current evaluation. This supports decisions such as simplifying the pipeline or re-evaluating the reranker with a harder test set.

Exam trap

The trap here is treating a small ablation effect as proof that a component is defective, when it may simply be redundant or under-stressed by the current evaluation.

35
MCQeasy

A data scientist is running a fine-tuning experiment with NVIDIA NeMo on a single A100 GPU. They want to establish a repeatable baseline before sweeping any hyperparameters, so that a later run can be compared fairly. Which practice best supports this goal?

A.Increase the batch size to the maximum the GPU memory allows so the baseline finishes quickly.
B.Fix a random seed in the NeMo training configuration and record the exact dataset version, model checkpoint, and NeMo container tag used.
C.Enable automatic mixed precision and Tensor Core acceleration to reduce training time.
D.Run the experiment three times and report the best validation loss observed.
AnswerB

Pinning the seed plus capturing dataset version, starting checkpoint, and container tag makes the baseline reproducible; re-running with the same inputs yields a comparable result. Without this record, later sweeps cannot be attributed to the hyperparameter change, because data or environment drift would confound the comparison.

Why this answer

A trustworthy baseline requires controlling every input that affects the result, then recording those inputs so the run can be repeated. Fixing the random seed, freezing the dataset version, naming the starting checkpoint, and noting the NeMo container tag together make the run reproducible and comparable. Performance tweaks and best-of-N reporting change or obscure the result rather than establishing a stable reference point.

Exam trap

The trap here is assuming that making the run faster or averaging several runs automatically makes it a valid baseline, when reproducibility actually depends on controlling and recording the inputs.

36
Multi-Selecthard

An enterprise AI researcher is conducting ablation studies on a large language model using NVIDIA NeMo. To ensure the experimentation results are scientifically valid and statistically sound, which THREE practices must be enforced during the study?

Select 3 answers
A.Vary multiple architectural hyperparameters simultaneously in every single experimental iteration.
B.Evaluate across multiple random initialization seeds to account for stochastic variance in training.
C.Keep all non-target hyperparameters strictly constant while isolating the variable under study.
D.Apply appropriate statistical significance testing to validate performance differences between variants.
E.Discard all experimental runs that fail to meet performance expectations without logging the failure.
AnswersB, C, D

Deep learning models exhibit sensitivity to initial weight distributions and data shuffling order. Running evaluations across multiple distinct random seeds ensures that reported performance improvements reflect genuine architectural gains rather than fortunate random initialization artifacts.

Why this answer

Rigorous ablation studies require isolating individual components while holding all other experimental variables constant. Enforcing multiple random seeds accounts for initialization variance, controlling confounding hyperparameters prevents attribution errors, and applying rigorous statistical testing confirms whether observed performance deltas are statistically significant.

Exam trap

Candidates often overlook the random initialization seed. By failing to repeat experiments with different seeds, they risk attributing performance changes to their variable rather than to stochastic training noise.

37
MCQeasy

A team has fine-tuned a small LLM with NeMo and now wants to quantify how much the fine-tuning improved performance on a domain question-answering task relative to the base model. They have a curated set of 500 question-answer pairs that were never used during training. What is the most appropriate next step?

A.Evaluate both the base model and the fine-tuned model on the held-out 500-question set using the same metric and decoding settings.
B.Run the fine-tuned model on the original training questions and compare its answers to the stored labels.
C.Fine-tune the base model a second time with a different seed and compare the two training curves.
D.Report the final training loss of the fine-tuned model as the measure of improvement.
AnswerA

A held-out set that was excluded from training provides an unbiased estimate of generalization. Running both the base and fine-tuned models with identical prompts, decoding parameters, and scoring metric isolates the effect of fine-tuning. Comparing the two scores directly answers how much the fine-tuning improved domain question answering, which is the stated goal.

Why this answer

Quantifying the benefit of fine-tuning requires an unbiased comparison between the adapted model and its base counterpart on data neither has seen in training. Using the same held-out questions, prompts, decoding settings, and metric for both models yields a clean estimate of the improvement. Training loss, training-set accuracy, or a second training run do not measure generalization gain.

Exam trap

The trap here is using training loss or training-set accuracy as evidence of improvement, when only held-out evaluation isolates the generalization benefit of fine-tuning.

38
MCQmedium

You are experimenting with RAG and notice the model is frequently ignoring the provided context. Which of the following is the most likely culprit to investigate first?

A.The GPU clock speed.
B.The prompt instruction strength.
C.The number of training epochs.
D.The total number of parameters in the model.
AnswerB

The prompt acts as the steering mechanism for the model. If it is too vague, the model defaults to internal knowledge. Strengthening the instruction to explicitly state that the answer must be derived solely from the provided context is the most efficient first step in iterative prompt experimentation.

Why this answer

LLMs often exhibit 'pre-training bias,' where they prioritize their internal knowledge base over the provided context. If the context is ignored, the retrieval quality or the prompt's instruction is usually insufficient to override this behavior. Experimentation should focus on prompt engineering—specifically strengthening the system instructions to enforce context usage—before attempting more complex architectural changes to the retrieval process or the model itself.

Exam trap

Candidates often jump to retraining the model or changing the vector database, ignoring the simpler, more effective fix of adjusting the system instructions to force context adherence.

39
MCQhard

A researcher is running an ablation study in which they vary the number of attention heads in a NeMo Megatron GPT model while holding parameter count, dataset, and learning rate fixed. After the first run, they change tensor parallel size and pipeline parallel size to fit larger variants on the available GPUs. A colleague argues this invalidates the comparison. Which statement best explains the scientific concern?

A.Attention head count is not a valid ablation variable because it is determined by hidden size and cannot be varied independently.
B.Different parallelism configurations change numerical reduction order and kernel selection, introducing variance unrelated to the attention-head variable.
C.Changing parallelism alters the optimizer state sharding, so the effective learning rate changes per parameter.
D.Pipeline parallelism forces different micro-batch sizes, which always changes the global batch size and therefore the loss landscape.
AnswerB

Tensor and pipeline parallelism alter how partial sums are reduced across GPUs and which kernels execute, producing small numerical differences and sometimes different convergence behavior. In an ablation isolating attention heads, those parallelism-induced effects become confounds. The colleague is right because the observed result could reflect parallelism rather than the variable under study, weakening the causal claim.

Why this answer

An ablation study aims to attribute observed differences to one variable. When tensor and pipeline parallel sizes change between runs, the reduction order, kernel selection, and communication patterns change too, introducing numerical and convergence variance that is unrelated to attention head count. That makes the comparison confounded, even though optimizer sharding and micro-batching can be handled without altering the effective update or global batch size.

Exam trap

The trap here is assuming that any configuration change between ablation runs is harmless, when parallelism settings can silently introduce confounds into the comparison.

40
MCQmedium

When experimenting with model quantization (e.g., INT8 or FP8), what is the most important trade-off to monitor?

A.Power consumption versus disk space
B.Inference speed versus accuracy degradation
C.Training time versus model size
D.GPU clock speed versus CPU utilization
AnswerB

Quantization is a classic trade-off between throughput and output quality. As precision drops, latency improves, but the model may lose nuance or become prone to errors. Successfully implementing quantization requires quantifying exactly how much accuracy is sacrificed for the specific speed gains achieved in the target deployment environment.

Why this answer

Quantization reduces memory footprint and increases inference speed by reducing the precision of weights. However, the trade-off is often a small decrease in model accuracy. Monitoring this 'accuracy-vs-efficiency' curve is the primary task during quantization experimentation.

Engineers must ensure the degradation remains within acceptable business tolerances for the specific application, ensuring that speed gains do not come at the cost of correctness.

Exam trap

Candidates often focus solely on the speed gain, ignoring the potential for accuracy loss. They forget that an optimized model is useless if it no longer provides correct answers.

41
MCQmedium

An engineer at a customer-support automation company is experimenting with top-p sampling values for a NeMo-served LLM. They want to quantify how output diversity changes across settings without relying on human judgment alone. Which evaluation approach best supports this experiment?

A.Measure end-to-end inference latency with NVIDIA Triton Inference Server metrics.
B.Compute distinct-n and self-BLEU across a fixed prompt set for each top-p value.
C.Compare training loss curves from the original fine-tuning job.
D.Track GPU utilization and memory bandwidth during generation with Nsight Systems.
AnswerB

Distinct-n measures lexical variety and self-BLEU measures similarity among generated outputs, both computed automatically over a fixed prompt set. Together they quantify diversity changes as top-p varies, giving the engineer an objective, repeatable signal. This matches the goal of measuring output diversity without depending solely on subjective human ratings.

Why this answer

To quantify output diversity across top-p settings, the engineer needs metrics that capture lexical variety and redundancy in generated text. Distinct-n and self-BLEU are computed automatically over a fixed prompt set and directly reflect diversity changes. Latency, GPU profiling, and training loss describe performance or training behavior, not the diversity of inference-time outputs, so they cannot answer the experiment's question.

Exam trap

The trap here is reaching for infrastructure metrics like latency or GPU utilization when the experiment is about the semantic diversity of generated text.

42
MCQeasy

Which of the following is the primary goal of the 'Experimentation' phase in an LLM project?

A.Deploying the model to production
B.Determining optimal configurations for a task
C.Purchasing new hardware for the team
D.Writing marketing copy for the product
AnswerB

The objective of experimentation is to systematically test configurations to find those that yield the best performance. Whether it's hyperparameter tuning for fine-tuning or prompt testing for RAG, this phase provides the data-driven evidence needed to select the model setup that best balances quality with resource constraints.

Why this answer

The experimentation phase focuses on hypothesis testing, hyperparameter tuning, and prompt engineering to find the optimal configuration for a specific task. By isolating variables, researchers can identify the best balance between accuracy, latency, and cost. This phase is essential for moving from a general-purpose model to a specialized, reliable solution, ensuring the project meets defined success criteria before proceeding to full-scale deployment.

Exam trap

Students often select final deployment or dataset collection as the primary goal, overlooking that the experimentation phase is specifically about finding optimal configurations.

43
Multi-Selectmedium

An engineer is setting up an automated experiment sweep over temperature and top-p for a NeMo-served LLM, and wants the results to be comparable and reproducible. Which two practices should be applied? (Choose two.)

Select 2 answers
A.Use the same prompt set and maximum token count across all temperature and top-p combinations.
B.Allow each trial to choose its own prompt template so the model can show its best behavior.
C.Fix the random seed used by the generation sampler for every trial in the sweep.
D.Increase the model's parameter count for trials with higher temperature to compensate for randomness.
E.Disable logging of per-trial parameters to reduce storage overhead during the sweep.
AnswersA, C

Holding the input prompts and generation length constant isolates the sampling parameters as the only variables. If prompt sets or output lengths vary between trials, observed quality differences may stem from the inputs rather than from temperature or top-p. Controlled inputs are required for a valid comparison of decoding settings.

Why this answer

Comparable decoding sweeps require controlling both stochasticity and inputs. Fixing the sampler seed makes each configuration repeatable, while holding prompts and maximum token count constant ensures the only varying factors are temperature and top-p. Changing the model, the prompt template, or removing parameter logs introduces confounds or destroys traceability, so those practices undermine the experiment.

Exam trap

The trap here is believing that more randomness or more model capacity makes a sweep more informative, when uncontrolled variation actually prevents attributing results to the sampling parameters under test.

44
MCQeasy

Which of the following is a primary objective of 'Ablation Studies' in LLM experimentation?

A.To increase the total number of parameters in the model.
B.To determine the contribution of individual components.
C.To accelerate the training speed by using less data.
D.To debug and fix errors in the model's training code.
AnswerB

Ablation studies isolate specific parts of the model or training pipeline to measure their impact on the final performance metrics. By disabling one feature at a time, researchers can quantify the 'value add' of each component, which is crucial for architectural refinement and resource optimization.

Why this answer

Ablation studies systematically remove components (like layers, heads, or data sources) to measure their specific contribution to model performance. This process is essential for understanding the model's architecture and optimizing it by removing redundant or inefficient parts. For NVIDIA-certified professionals, ablation studies provide the empirical evidence needed to defend design choices and justify model architectural simplifications in complex projects.

Exam trap

Candidates often confuse ablation studies with fine-tuning or quantization, falsely believing they are meant to improve overall model accuracy rather than isolate the impact of specific architectural components.

45
MCQhard

An engineer is evaluating a RAG-based assistant and wants to isolate whether retrieval quality or the generator is responsible for wrong answers. They build a small labeled set of questions with known correct passages and known correct answers. Which experimental design most cleanly separates the contribution of the retriever from that of the generator?

A.Measure retriever recall against the known correct passages, then feed the known correct passages to the generator and measure answer accuracy separately.
B.Compare the end-to-end accuracy of the RAG system against the same generator without retrieval.
C.Increase the retriever top-k until end-to-end accuracy stops improving, then freeze the retriever and tune the generator.
D.Run the full pipeline and measure end-to-end answer accuracy, then retrain the generator on the failures.
AnswerA

By scoring retrieval against ground-truth passages and separately scoring generation given oracle passages, the design isolates each component's contribution. Retrieval recall shows whether the right context is found, and oracle-context accuracy shows whether the generator can use correct context. This controlled decomposition is the standard way to attribute failures in RAG systems.

Why this answer

Isolating retriever versus generator contributions requires scoring each stage against ground truth. Retriever recall measured against known correct passages reveals retrieval quality, and feeding those oracle passages to the generator reveals generation quality independent of retrieval. End-to-end comparisons and top-k tuning mix the two stages and cannot attribute failures.

Exam trap

The trap here is using end-to-end accuracy as if it were a diagnostic, when it aggregates two independent failure sources into one number.

46
MCQmedium

A data science team is fine-tuning a Llama 3 8B model on a proprietary customer-support corpus using NVIDIA NeMo. They need to run dozens of experiments with different learning rates and batch sizes. Because the dataset contains personally identifiable information, they cannot send any telemetry to an external tracking server, but they still need to compare runs later and reproduce the best configuration. Which approach best satisfies both the reproducibility and data-privacy requirements?

A.Enable Weights & Biases integration inside the NeMo experiment manager and configure the project to log only aggregate metrics.
B.Disable all logging and rely on the saved .nemo checkpoint files, since the checkpoint embeds the full training configuration.
C.Run each experiment in a separate container and manually copy the console output into a spreadsheet after each run.
D.Use the built-in NeMo experiment manager with a local file store, logging configs, metrics, and checkpoints to a shared on-premises directory.
AnswerD

NeMo's experiment manager supports a local file store backend that writes configuration, metrics, and checkpoints to a directory you control. This keeps all PII-adjacent metadata on premises while still capturing the full run configuration needed to reproduce the best experiment. It is the only option that satisfies both constraints simultaneously without additional infrastructure.

Why this answer

The team needs reproducibility without external telemetry. NeMo's experiment manager with a local file store writes configurations, metrics, and checkpoints to a path the team controls, so nothing leaves the secure environment. External SaaS trackers violate the privacy constraint, checkpoints alone lack the metrics needed for comparison, and manual logging is unreliable and incomplete.

Exam trap

The trap here is assuming that a cloud tracking service can be made compliant simply by logging fewer fields, when the requirement is that no telemetry leaves the environment at all.

47
Multi-Selectmedium

An AI researcher is designing an experiment to compare two prompt templates for a customer-support LLM using NVIDIA NeMo. To ensure the comparison is fair and reproducible, which two practices should they follow? (Choose two.)

Select 2 answers
A.Vary the temperature for each template to explore a wider range of outputs.
B.Fine-tune the LLM separately for each prompt template before comparing them.
C.Use a different evaluation metric for each template to capture their unique strengths.
D.Use the same underlying LLM checkpoint and decoding parameters for both prompt templates.
E.Evaluate both prompt templates on the same held-out set of customer queries with identical scoring criteria.
AnswersD, E

Holding the model checkpoint and decoding parameters constant ensures that any difference in output quality is attributable to the prompt template rather than to model weights or sampling settings. This is essential for a controlled comparison, because changing the checkpoint or temperature would introduce confounding variables that make the results uninterpretable.

Why this answer

A fair prompt-template comparison requires isolating the template as the only independent variable. Keeping the model checkpoint, decoding parameters, evaluation dataset, and scoring criteria identical across both conditions ensures that observed differences are caused by the prompt itself. This controlled design is the foundation of reproducible LLM experimentation and supports confident deployment decisions.

Exam trap

The trap here is assuming that changing multiple factors at once—such as fine-tuning per template or varying temperature—provides a richer comparison, when it actually destroys the ability to attribute results to the prompt.

48
MCQhard

During an ablation study on a retrieval-augmented LLM in NeMo, an engineer removes the reranking stage and observes that answer accuracy drops by 12 points, but latency improves by 40 percent. A stakeholder asks whether reranking should be kept. Which experimental next step best supports a defensible recommendation?

A.Keep reranking because accuracy is always more important than latency in production systems.
B.Measure the accuracy-latency trade-off across several reranker sizes and retrieval depths, then compare against the application's latency budget and quality target.
C.Increase the number of retrieved documents while keeping the reranker disabled to recover the lost accuracy.
D.Remove the retrieval stage entirely to see whether the model's parametric knowledge is sufficient.
AnswerB

The single ablation shows a trade-off but not the shape of the curve or whether a cheaper configuration can retain most of the accuracy gain. Sweeping reranker sizes and retrieval depths reveals intermediate operating points, and mapping them to the latency budget and quality target turns the data into a concrete recommendation. This is the experiment that supports a defensible decision rather than a binary choice.

Why this answer

A single ablation establishes that reranking trades latency for accuracy but does not show whether a middle ground exists. Sweeping reranker sizes and retrieval depths produces a trade-off curve, and evaluating those points against the application's latency budget and quality target converts measurements into a recommendation. Asserting a priority, removing retrieval, or compensating with more documents does not answer the stakeholder's question.

Exam trap

The trap here is treating a two-point ablation as sufficient evidence for a binary keep-or-remove decision, when the useful answer is often an intermediate configuration found by sweeping cost and quality.

49
MCQeasy

What is the primary benefit of tracking experiments using a centralized experiment management platform (e.g., Weights & Biases, MLflow)?

A.It automatically scales GPU resources based on workload.
B.It enforces strict security policies on data access.
C.It ensures reproducibility by logging parameters and code state.
D.It increases the training speed of the LLM model.
AnswerC

Experiment trackers store the specific configuration, hyperparameters, code commits, and environment details for every run. This creates an audit trail that allows any researcher to replicate a previous experiment exactly, which is essential for ensuring scientific integrity and building upon successful results in a structured team environment.

Why this answer

Centralized tracking provides a historical log of all runs, parameters, code versions, and results. This traceability is critical for reproducibility, allowing researchers to compare outcomes across months or even different team members. In an enterprise NVIDIA environment, this documentation prevents redundant experimentation and ensures that the best-performing models are easily identified and promoted for deployment into production.

Exam trap

Candidates often assume the primary benefit is 'model performance improvement,' when the actual primary benefit of experiment tracking is the ability to reproduce results through logged parameters and code state.

50
MCQmedium

An AI engineer at a financial services company is running an LLM experimentation pipeline using NVIDIA NeMo. The primary objective is to evaluate how different tokenizers affect the accuracy of a named entity recognition (NER) task on financial documents. The engineer has already fixed the model architecture, the training dataset, and the hyperparameters. To ensure the experiment isolates the effect of the tokenizer, which action should the engineer take next?

A.Replace the tokenizer with a different one while keeping all other configurations unchanged, then compare NER accuracy.
B.Increase the training dataset size to improve NER accuracy before changing the tokenizer.
C.Fine-tune the model on a general-domain corpus before evaluating on financial documents.
D.Retrain the model with the same tokenizer but different random seeds to measure variance.
AnswerA

This directly isolates the tokenizer as the independent variable. By holding the model architecture, dataset, and hyperparameters constant, any change in NER accuracy can be attributed to the tokenizer. This controlled comparison is essential for valid experimentation and aligns with the goal of evaluating tokenizer impact on financial NER.

Why this answer

To isolate the effect of the tokenizer, the engineer must change only the tokenizer while keeping all other factors constant. This controlled approach ensures that any observed difference in NER accuracy is due to tokenization rather than confounding variables like model architecture or dataset size. Such isolation is fundamental to valid experimentation and enables clear conclusions about tokenizer impact.

Exam trap

The trap here is assuming that improving overall accuracy through data scaling or fine-tuning will reveal tokenizer effects, when in fact those changes introduce confounding variables.

51
MCQmedium

A team is comparing two fine-tuned variants of the same base LLM for a customer-support summarization task. Variant A was trained on 10,000 examples and variant B on 2,000 examples drawn from the same distribution. To make a fair comparison of their generalization, which evaluation practice should the team adopt?

A.Evaluate both variants on the same held-out test set that neither saw during training
B.Report the training loss of each variant at the final step
C.Compare the number of parameters each variant updated during fine-tuning
D.Let each variant be scored on its own randomly split test set
AnswerA

Holding the evaluation data constant isolates the effect of the training-set size from the effect of the evaluation data. If each variant were scored on a different test set, differences in difficulty or domain mix would confound the comparison. A shared, untouched held-out set gives both models the same challenge, so observed differences can be attributed to training choices rather than evaluation noise.

Why this answer

Fair comparison requires controlling the evaluation data. When both variants are scored on the same held-out set that neither encountered during training, differences in their scores reflect the training choices rather than variation in test difficulty. Training loss and parameter counts are process metrics that do not measure generalization to unseen customer-support summaries.

Exam trap

The trap here is treating training loss as a proxy for quality, when it mostly reflects how much data the model memorized rather than how well it generalizes.

52
Multi-Selectmedium

A data scientist is preparing to perform hyperparameter tuning for a Retrieval-Augmented Generation (RAG) system. Which TWO parameters should be prioritized for experimentation to improve retrieval accuracy?

Select 2 answers
A.Chunk size
B.Model quantization bit-width
C.Similarity search top-k
D.System prompt length
E.GPU clock speed
AnswersA, C

Chunk size directly determines how much context is captured in each vector index entry. Too small, and the model lacks enough information; too large, and the content becomes noisy, leading to irrelevant retrieval. Finding the optimal balance through experimentation is essential for ensuring high-quality context retrieval for the LLM.

Why this answer

Retrieval accuracy in RAG systems is heavily influenced by the chunking strategy and the similarity search configuration. Experimenting with these two components allows developers to optimize the context provided to the LLM. Properly tuned retrieval ensures that the model has the most relevant information, which is the most critical factor in reducing hallucinations and improving the factual grounding of generated responses.

Exam trap

Candidates often choose parameters related to model architecture or training rather than retrieval. They forget that RAG performance is primarily driven by how data is fetched and segmented.

53
MCQhard

A research team is using NVIDIA NeMo to experiment with a large language model for a summarization task. They observe that the model's ROUGE scores vary significantly across different runs even when using the same hyperparameters and dataset. They suspect that non-deterministic operations in the training pipeline are causing this variance. Which step should they take to improve reproducibility of their experiment results?

A.Run the experiment multiple times and average the ROUGE scores to report a mean value.
B.Enable deterministic algorithms and set a fixed random seed in the NeMo training configuration.
C.Use a different optimizer with adaptive learning rates to stabilize training.
D.Increase the batch size to reduce the number of gradient updates per epoch.
AnswerB

Enabling deterministic algorithms and fixing the random seed ensures that operations like weight initialization, dropout, and data shuffling produce identical results across runs. This directly addresses the observed variance by eliminating non-deterministic sources. NeMo supports these settings, making it the correct approach to achieve reproducible experiments in summarization tasks.

Why this answer

To achieve reproducibility, the team must eliminate non-deterministic sources. Setting a fixed random seed and enabling deterministic algorithms ensures that random number generation, data shuffling, and GPU operations are consistent across runs. This directly reduces variance in ROUGE scores and allows other researchers to replicate the experiment exactly, which is a cornerstone of rigorous LLM experimentation.

Exam trap

The trap here is thinking that statistical averaging or hyperparameter tuning can fix irreproducibility, when the real issue is non-deterministic operations that must be explicitly controlled.

54
MCQhard

Refer to the exhibit. The experiment fails with an Out-of-Memory (OOM) error during the second epoch. Given the configuration, which change is most effective for immediate stabilization?

A.Switch precision to FP32.
B.Increase batch_size to 64.
C.Reduce batch_size to 16.
D.Increase max_seq_len to 8192.
AnswerC

Reducing the batch size decreases the memory demand for forward and backward passes. This allows the process to fit within the GPU constraints, enabling the experiment to continue and finish. This trade-off is standard in LLM experimentation when balancing hardware limitations against model size and sequence length requirements.

Why this answer

The OOM error indicates the memory footprint exceeds the available VRAM, likely due to activation growth or KV cache accumulation. Reducing the batch size is the most direct way to lower memory consumption per iteration. Experimentation requires iterative adjustment of hyperparameters; scaling back memory-intensive settings allows the job to complete successfully, providing a baseline from which performance optimizations like gradient accumulation can then be systematically applied to restore throughput.

Exam trap

Candidates often suggest increasing memory or changing hardware types. They fail to realize that batch size is the most direct control variable for memory consumption in a standard training loop.

55
MCQeasy

Which of the following best describes the role of 'AB testing' in Generative AI experimentation?

A.It is used to calculate the model's perplexity.
B.It validates performance improvements with real user feedback.
C.It automates the hyperparameter tuning process.
D.It eliminates the need for any offline evaluation.
AnswerB

AB testing allows for the assessment of model performance based on real-world usage patterns. Since user satisfaction is often subjective, comparing two variants in a live environment provides the most accurate reflection of how effectively the generative system meets the needs of the actual target audience.

Why this answer

AB testing is a controlled method for comparing two variations of a generative system—such as different prompt templates or model versions—using real-world user interactions. By splitting traffic, researchers can gather empirical evidence on which variation yields superior user metrics. This is the gold standard for validating whether experimental improvements in a lab environment translate into actual value for the end-users in a production setting.

Exam trap

Candidates often confuse AB testing with model evaluation or benchmarking, focusing on internal metrics rather than the specific goal of capturing real user behavior and empirical preference in production.

56
Multi-Selectmedium

Which TWO factors should be considered when evaluating the cost-benefit of an LLM experimentation strategy?

Select 2 answers
A.Total GPU hours per experiment
B.The color scheme of the monitoring dashboard
C.Expected improvement in evaluation metrics
D.The popularity of the LLM framework used
E.The number of research papers published by the team
AnswersA, C

GPU hours represent the primary financial cost of LLM experimentation. Since large-scale training is expensive, researchers must track usage to ensure that the budget is spent on high-probability improvements. Projects must balance the need for rigorous testing with the reality of cloud compute costs in a professional enterprise environment.

Why this answer

Evaluating the cost-benefit of LLM experimentation requires balancing the financial cost of compute with the performance gains observed. In professional environments, experiments must be scoped to maximize meaningful insights while minimizing wasted GPU cycles. This involves prioritizing experiments that are likely to yield the highest impact on model quality relative to the resources consumed by the training runs.

Exam trap

Candidates frequently focus solely on the financial cost of GPU hours, failing to realize that the value of an experiment is only realized when measured against the expected improvement in model metrics.

57
Multi-Selecthard

A research team is designing an experiment to measure how prompt phrasing affects the factuality of an LLM in a retrieval-augmented question-answering pipeline. Which two design choices are necessary to attribute observed factuality differences to the prompt rather than to other pipeline components? (Choose two.)

Select 2 answers
A.Hold the retrieved context, model checkpoint, and decoding parameters constant across prompt variants
B.Evaluate each prompt variant on a different dataset tailored to its strengths
C.Increase the temperature for variants that produce shorter answers to equalize response length
D.Use a fixed, representative evaluation set of questions with reference answers scored by the same rubric
E.Allow the retrieval index to be rebuilt with different embedding models for each prompt variant
AnswersA, D

If the retrieved passages, model weights, or sampling settings change between prompt variants, any factuality difference could come from those factors rather than the prompt. Fixing them makes the prompt the only manipulated variable, which is the definition of a controlled experiment. This isolation is what allows the team to draw a causal conclusion about phrasing.

Why this answer

Attributing factuality differences to prompt phrasing requires isolating the prompt as the only changed variable. Holding retrieval, model, and decoding settings constant removes competing explanations, while a shared evaluation set and rubric make scores comparable. Altering the retrieval index, temperature, or datasets per variant introduces confounds that make any observed improvement impossible to credit to the prompt itself.

Exam trap

The trap here is optimizing each variant's surrounding pipeline to make it look best, which destroys the control needed to attribute results to the prompt.

58
MCQmedium

An ML engineer is comparing two fine-tuned variants of the same base LLM on an internal question-answering benchmark. Variant 1 scores higher on exact-match accuracy, but Variant 2 produces answers that human reviewers judge more helpful and better grounded. The engineer must decide which variant to promote. Which evaluation approach is most appropriate?

A.Promote Variant 1 because exact-match accuracy is an objective, reproducible metric that removes human subjectivity.
B.Promote Variant 2 without further measurement because human reviewers are the ultimate authority on answer quality.
C.Combine the automated metric with a structured human or LLM-judge rubric that scores helpfulness and grounding, then decide using both.
D.Retrain both variants on the benchmark questions until exact-match accuracy converges, then promote the higher scorer.
AnswerC

Open-ended QA quality is multi-dimensional, so pairing a reproducible automated metric with a rubric-based judgment captures both correctness and the qualities reviewers valued. Using both signals lets the engineer weigh benchmark accuracy against real helpfulness and grounding, producing a decision that reflects actual user experience rather than a single narrow number.

Why this answer

When automated and human signals disagree, the right move is usually to measure both dimensions rather than discard one. A rubric-based human or LLM-judge evaluation for helpfulness and grounding complements exact-match accuracy, which is reproducible but blind to paraphrase and partial credit. Deciding with both signals yields a promotion choice that reflects real user value while retaining an objective regression check.

Exam trap

The trap here is treating the decision as a choice between objective and subjective evaluation, when the two measure different quality dimensions and should be combined.

59
Multi-Selectmedium

You are designing an experiment to measure how quantization (FP16 versus INT8) affects inference latency and answer quality for an LLM deployed with NVIDIA TensorRT-LLM. Which two practices are required for the comparison to be valid? (Choose two.)

Select 2 answers
A.Benchmark both configurations on the same hardware and under the same batch size and concurrency conditions.
B.Retrain the model after quantization so the weights adapt to the lower precision.
C.Use the same input prompts and generation parameters, such as max tokens and temperature, for both precision configurations.
D.Measure latency only on the first generated token, since subsequent tokens are less affected by precision.
E.Apply the same random seed to both configurations so token sampling is identical.
AnswersA, C

Latency depends heavily on GPU model, batch size, and concurrent request load. Measuring FP16 and INT8 on different hardware or load levels would make the results incomparable. Keeping hardware and serving conditions identical isolates precision as the variable under test.

Why this answer

A valid quantization comparison must control everything except precision. Using identical prompts and generation parameters ensures the workload is the same, and benchmarking on identical hardware under identical batch and concurrency conditions ensures the environment is the same. Together these practices let the team attribute measured differences in latency and quality to FP16 versus INT8 rather than to confounds.

Exam trap

The trap here is assuming that a shared random seed makes FP16 and INT8 outputs directly comparable token by token.

60
MCQmedium

A team fine-tuning a NeMo large language model runs the same training configuration three times and obtains validation loss values of 2.14, 2.31, and 2.09 at the end of the same number of steps. They need subsequent runs to produce tightly clustered, comparable numbers so hyperparameter comparisons are meaningful. Which change most directly addresses this problem?

A.Add early stopping based on validation loss plateau detection.
B.Increase the global batch size while holding the learning rate constant.
C.Set a fixed random seed and enable deterministic behavior for the training run.
D.Switch the optimizer from AdamW to SGD with momentum.
AnswerC

A fixed seed makes weight initialization, dropout masks, and data shuffling identical across runs, and enabling deterministic kernels removes nondeterministic GPU operations, so repeated runs converge to nearly identical loss values. This directly targets the run-to-run variance observed and makes hyperparameter comparisons valid.

Why this answer

Run-to-run variance in loss comes from uncontrolled randomness: weight initialization, dropout sampling, data shuffling, and nondeterministic GPU kernels. Pinning a seed and enabling deterministic execution makes those sources identical across repeated runs, so differences in outcomes can be attributed to the hyperparameters under test rather than to chance.

Exam trap

The trap here is assuming that a larger batch or a different optimizer removes variance, when only controlling randomness and nondeterministic kernels makes repeated runs comparable.

61
MCQmedium

You are running a NeMo fine-tuning experiment where validation loss decreases for the first three epochs, then rises steadily while training loss keeps falling. You want to confirm overfitting and select the most appropriate intervention. Which experiment action should you take first?

A.Switch the optimizer from Adam to SGD without changing any other hyperparameter.
B.Reduce the size of the validation set to lower evaluation noise.
C.Increase the number of training epochs and re-run to see if validation loss recovers.
D.Add regularization such as dropout or weight decay, then re-run the experiment and compare validation curves.
AnswerD

A widening gap between falling training loss and rising validation loss is the classic overfitting signature. Introducing dropout or weight decay constrains model capacity and is the standard first intervention. Re-running with the same data split and seed lets you isolate the regularization effect, confirming whether overfitting is the cause rather than a data or learning-rate artifact.

Why this answer

Rising validation loss alongside falling training loss indicates the model is fitting training-specific patterns that do not generalize. Adding regularization such as dropout or weight decay directly limits effective capacity and is a controlled, reversible change. Re-running with the same split and seed isolates the intervention, so the resulting validation curve tells you whether overfitting was indeed the dominant cause.

Exam trap

The trap here is assuming more training will eventually close the gap, when a rising validation curve signals that additional epochs will only deepen the overfitting.

62
MCQeasy

An engineer must decide how to split a labeled dataset of 50,000 customer support conversations before fine-tuning a NeMo LLM for intent classification. The goal is an honest estimate of how the tuned model will behave on never-before-seen tickets once deployed. Which splitting approach best supports that goal?

A.Randomly assign 70 percent to training, 15 percent to validation, and 15 percent to a test set that is touched only once at the end.
B.Use all 50,000 conversations for training and rely on the training loss curve to judge generalization.
C.Tune hyperparameters repeatedly against the test split until accuracy is maximized.
D.Split the data by conversation length, putting the longest 15 percent in the test set.
AnswerA

Holding out a test split that is never used for tuning gives an unbiased estimate of generalization, while the validation split supports hyperparameter decisions. Random assignment keeps class proportions approximately stable in each split, so the final test measurement reflects expected production behavior on unseen tickets.

Why this answer

An untouched holdout test set is the only split that yields an unbiased generalization estimate, because it never influences training or hyperparameter choices. Pairing it with a validation split lets the team tune decisions without contaminating the final measurement, and random assignment keeps class balance comparable across splits.

Exam trap

The trap here is treating the test set as another tuning signal, which silently converts it into a validation set and inflates the reported accuracy.

63
MCQeasy

An ML engineer at a healthcare analytics company is starting a fine-tuning experiment on a Llama 2 7B model using NVIDIA NeMo Framework. Before launching the training job, the engineer wants a single immutable record that captures the exact model checkpoint, dataset version, hyperparameters, and evaluation scores so that any later run can be traced back to it. Which component of the NVIDIA NeMo experimentation workflow should the engineer use to store that record?

A.NVIDIA Triton Inference Server model repository
B.NVIDIA TensorRT-LLM build configuration
C.NVIDIA Nsight Systems profiling report
D.NeMo Experiment Manager
AnswerD

NeMo Experiment Manager is the component that logs and organizes experiment metadata such as model checkpoints, dataset versions, hyperparameters, and evaluation metrics, giving the engineer a single traceable record for the healthcare fine-tuning run. It is designed precisely for the reproducibility and auditability goals described, so it fits the scenario without requiring external tooling.

Why this answer

The requirement is a traceable record combining model checkpoint, dataset version, hyperparameters, and evaluation metrics. NeMo Experiment Manager is built to log and organize exactly that metadata during fine-tuning, enabling reproducibility and comparison across runs. Inference servers, inference compilers, and profilers serve different purposes and do not persist the training experiment lineage needed here.

Exam trap

The trap here is assuming that any NVIDIA tool that touches the model lifecycle also records experiment metadata, when only Experiment Manager is designed for that tracking role.

64
MCQmedium

A researcher is experimenting with prompt-tuning and finds that the model output is repetitive. They decide to adjust the sampling hyperparameters. Which combination of changes is most likely to increase the diversity of the output?

A.Decrease temperature and decrease top-p.
B.Increase temperature and increase top-p.
C.Set temperature to 0 and top-p to 1.
D.Increase frequency penalty and decrease top-p.
AnswerB

Higher temperature flattens the probability distribution, increasing the chance of picking diverse tokens. Increasing top-p broadens the cumulative probability mass considered during sampling. Together, these settings allow the model to select from a wider vocabulary, effectively reducing repetitive output patterns during the experimentation phase of model evaluation.

Why this answer

Sampling parameters control the trade-off between coherence and creativity. Increasing temperature shifts the probability distribution, allowing for less likely tokens to be selected, while increasing top-p (nucleus sampling) expands the set of tokens considered. Balancing these allows researchers to explore the model's creative range during experimentation, ensuring the output is varied enough to be useful while maintaining sufficient logical coherence for the application's specific requirements.

Exam trap

Test-takers frequently confuse parameter directions, accidentally suggesting decreases in temperature or top-p when trying to fix repetitive and deterministic model outputs.

65
MCQeasy

During LLM experimentation, what is the primary purpose of maintaining a consistent 'seed' value across different runs?

A.To increase the training speed of the model.
B.To ensure reproducibility of experimental results.
C.To reduce the VRAM usage during training.
D.To prevent overfitting on the training set.
AnswerB

Reproducibility is essential to verify that improvements in model performance are due to hyperparameter changes rather than luck in initialization. A fixed seed allows researchers to compare models fairly, ensuring that observed differences are statistically significant and attributable to specific architectural or configuration decisions made during experimentation.

Why this answer

Consistency is the cornerstone of empirical science. By fixing the random seed, researchers ensure that weight initialization and data shuffling occur identically across runs. This allows them to isolate the impact of specific hyperparameter changes, such as learning rate or batch size, without the confounding variable of stochastic randomness, which is vital for reproducible and meaningful comparative analysis in complex machine learning workflows.

Exam trap

Candidates sometimes believe fixed seed values improve model accuracy or convergence speed, rather than strictly ensuring experimental reproducibility and controlling stochastic randomness.

66
MCQmedium

An AI researcher is testing a new LLM architecture on an NVIDIA DGX system. They observe that increasing the batch size leads to memory OOM errors despite available GPU utilization headroom. Which experimentation strategy should be employed first to isolate the bottleneck?

A.Increase the learning rate to accelerate convergence speed.
B.Switch to a different optimizer like SGD instead of Adam.
C.Implement gradient checkpointing to trade compute for memory.
D.Upgrade the NVIDIA driver version on the DGX host.
AnswerC

Gradient checkpointing is a standard experimentation technique to reduce memory usage by discarding intermediate activations during the forward pass and recomputing them during backpropagation. This effectively trades a small amount of additional compute time for a significant reduction in peak GPU memory usage, resolving the OOM error.

Why this answer

Memory allocation during deep learning training is often impacted by activation storage rather than just model weights. By systematically reducing the batch size or implementing gradient checkpointing, the researcher can determine if the OOM is due to peak activation memory usage. This experimentation phase is critical for optimizing hardware utilization and ensuring stable training cycles in high-throughput enterprise environments.

Exam trap

Candidates often assume Out-Of-Memory errors during batch size increases are caused by model weights, ignoring that activation memory during the forward pass scales with batch size and sequence length.

67
MCQmedium

During an ablation study, a team removes the instruction-tuning stage from their NeMo pipeline and observes that the model still answers factual questions but frequently ignores the requested output format. They want to attribute this change in behavior to the removed stage rather than to noise. Which experimental design element is most important for supporting that attribution?

A.Increase the ablated model's training steps so it receives more optimization than the full pipeline.
B.Retrain the ablated model several times with different random seeds and report the highest format-compliance score.
C.Run both the full pipeline and the ablated pipeline with all other variables held constant, then compare on the same evaluation set.
D.Evaluate the ablated model with a different, more format-focused prompt template than the full pipeline.
AnswerC

Holding every other factor constant and evaluating both variants on identical data isolates the instruction-tuning stage as the only difference between conditions. Any systematic change in format compliance can then be attributed to that stage rather than to confounding changes in data, hyperparameters, or evaluation setup.

Why this answer

An ablation is only interpretable when the manipulated component is the sole difference between conditions. Keeping data, hyperparameters, training budget, and evaluation protocol identical, and scoring both variants on the same held-out set, lets the team attribute the observed format-compliance change to removing instruction tuning instead of to unrelated variation.

Exam trap

The trap here is adding extra training or a different prompt to the ablated variant, which introduces a second difference and destroys the ability to attribute the outcome to the removed stage.

68
Multi-Selecthard

A research team is comparing three fine-tuning recipes for a NeMo LLM and wants the comparison to be defensible in a later review. Which two practices most improve the credibility of the reported comparison? (Choose two.)

Select 2 answers
A.Tune each recipe's learning rate separately until its score is maximized.
B.Reuse the evaluation set during training as an additional validation signal.
C.Report only the single best run for each recipe to keep the summary concise.
D.Evaluate every recipe on the same held-out evaluation set with identical decoding settings.
E.Log the exact dataset version, tokenizer, base checkpoint, and hyperparameters for every run.
AnswersD, E

Using one fixed evaluation set with identical decoding parameters ensures the recipes are judged on the same inputs under the same conditions, which isolates the effect of the recipe itself. If evaluation data or decoding settings differ between recipes, score differences may reflect the measurement setup rather than the training method.

Why this answer

Credible comparisons rest on two pillars: full provenance of each run so results are reproducible, and a shared, uncontaminated evaluation protocol so scores are measured identically. Provenance lets reviewers attribute differences to the recipes, while a fixed held-out set with consistent decoding settings ensures the numbers being compared actually measure the same thing.

Exam trap

The trap here is believing that a higher reported score proves a better recipe, when unequal tuning effort or a contaminated evaluation set can produce that score without any real advantage.

69
MCQhard

Refer to the exhibit. The experiment shows the model is failing to converge and exhibits loss spikes. Which adjustment to the configuration is most likely to stabilize the training process?

A.Increase the weight decay to 0.5.
B.Change mixed_precision to bf16.
C.Reduce the warmup_steps to 0.
D.Switch the optimizer to standard SGD.
AnswerB

FP16 has a limited dynamic range that often leads to overflow/underflow issues during large-scale model training, resulting in loss spikes. BF16 provides the same dynamic range as FP32, making it significantly more stable for training, especially when using modern NVIDIA hardware that supports it natively.

Why this answer

Loss spikes in mixed-precision training are often caused by the limited dynamic range of FP16. Lowering the gradient clip value or switching to BF16 (if hardware permits) are common remedies to maintain stability. By analyzing the configuration during experimentation, researchers can identify these hyperparameter sensitivities and prevent training failures, ensuring more robust and efficient model development cycles.

Exam trap

Test-takers often attempt to resolve mixed-precision loss spikes by increasing batch size or changing optimizers, ignoring the limited dynamic range constraints inherent to FP16.

70
MCQeasy

A team is running an LLM fine-tuning experiment using NVIDIA NeMo and wants to track how the validation loss changes over training. They need a reliable way to detect overfitting early. Which metric should they monitor most directly during the experiment?

A.Training loss computed on the same data batches used for gradient updates.
B.GPU utilization percentage reported by the NeMo training logs.
C.The number of tokens processed per second during training.
D.Validation loss computed on a held-out dataset at regular intervals during training.
AnswerD

Validation loss on a held-out set directly measures how well the model generalizes to unseen data. When validation loss begins to rise while training loss continues to fall, that divergence is the classic signal of overfitting. Monitoring it at regular intervals allows the team to stop training or adjust regularization before the model degrades further.

Why this answer

Overfitting is characterized by a growing gap between training and validation performance. Validation loss on a held-out set is the most direct indicator: when it stops improving and begins to rise while training loss keeps falling, the model is memorizing rather than generalizing. Monitoring it during training lets the team intervene early with regularization or early stopping.

Exam trap

The trap here is confusing operational metrics like GPU utilization or throughput with model-quality metrics, when only held-out validation loss reveals generalization behavior.

71
MCQeasy

A data scientist is preparing an LLM fine-tuning experiment on NVIDIA NeMo and wants every run to be reproducible weeks later. The team's experiment tracker currently logs only the final validation loss. Which additional item is most important to record so a run can be reproduced exactly?

A.The wall-clock duration of the final epoch only
B.The random seed, framework version, and full training configuration
C.A screenshot of the loss curve rendered in the dashboard
D.The names of the engineers who launched each training job
AnswerB

Reproducibility requires capturing the stochastic and environmental inputs: the random seed used for data shuffling and initialization, the exact NeMo and CUDA versions, and the complete training configuration such as learning rate, batch size, and precision. Without these, rerunning the job can produce different weights and metrics even with identical data, making comparisons across experiments unreliable.

Why this answer

Exact reproduction of an LLM experiment depends on controlling every stochastic and environmental input. The seed governs initialization and data ordering, the framework and CUDA versions govern kernel behavior and numerics, and the full configuration governs optimization. Logging only a final metric captures an outcome but discards the recipe, so the run cannot be recreated reliably.

Exam trap

The trap here is assuming that saving final metrics or dashboards is equivalent to experiment tracking, when reproducibility actually requires the seed and full configuration inputs.

72
MCQhard

A team is designing a controlled experiment to measure whether increasing LoRA rank improves instruction-following accuracy on a held-out benchmark. They want the comparison to be scientifically valid. Which experimental design choice best supports a valid conclusion?

A.Change LoRA rank and the base model simultaneously so the experiment covers more configurations in one run.
B.Report only the best accuracy observed across several random seeds for each rank.
C.Evaluate each rank on a freshly sampled held-out set to reduce any dataset bias.
D.Hold the base model, dataset, prompt template, and evaluation harness fixed while varying only LoRA rank.
AnswerD

Controlling all other factors isolates LoRA rank as the single independent variable, so any accuracy difference can be attributed to it. Fixed data, prompts, and evaluation harness also ensure the held-out benchmark measures the same capability across runs. This is the core requirement for a valid controlled comparison in LLM experimentation.

Why this answer

A scientifically valid comparison changes one independent variable while holding everything else constant. Here, LoRA rank is the factor under test, so the base model, training data, prompt template, and evaluation harness must remain identical across runs. Using a shared held-out benchmark and reporting aggregate accuracy across seeds ensures observed differences reflect rank rather than confounding factors or evaluation noise.

Exam trap

The trap here is believing that changing multiple factors at once is more efficient, when it actually destroys the ability to attribute the result to LoRA rank.

73
MCQmedium

An AI engineer is conducting an experiment to compare two different fine-tuning approaches for a large language model using NVIDIA NeMo: full fine-tuning versus parameter-efficient fine-tuning (PEFT) with LoRA. The engineer wants to determine which approach yields better performance on a downstream question-answering task while minimizing computational cost. Which metric should the engineer prioritize to evaluate the trade-off between performance and cost?

A.Training loss convergence rate
B.Validation accuracy per GPU-hour
C.Number of trainable parameters
D.Inference latency on a CPU
AnswerB

Validation accuracy per GPU-hour quantifies the trade-off between model performance and computational expense. It measures how much accuracy is gained for each unit of GPU time, directly addressing the goal of minimizing cost while maximizing performance. This metric allows the engineer to compare full fine-tuning and LoRA by showing which approach delivers better accuracy more efficiently.

Why this answer

Validation accuracy per GPU-hour effectively combines the two objectives: it measures the model's performance on the question-answering task (validation accuracy) and normalizes it by the computational cost (GPU-hours). This allows the engineer to compare full fine-tuning and LoRA on a level playing field, identifying which method provides the best accuracy for the resources invested. It directly supports the goal of minimizing cost while maximizing performance.

Exam trap

The trap here is focusing solely on parameter count or training speed, which are incomplete because they ignore either the performance or the cost side of the trade-off.

74
MCQeasy

A team is evaluating an LLM for a customer-support summarization task. They want to compare three prompt templates. Which experimental design most directly isolates the effect of the prompt template?

A.Use a different model for each template to see which combination performs best overall.
B.Use the same model, decoding parameters, and evaluation dataset for all three templates, changing only the template text.
C.Vary temperature and top-p across templates so each template is tested under its own best decoding settings.
D.Evaluate each template on a different dataset to cover more customer scenarios.
AnswerB

Holding model, decoding parameters, and dataset constant while varying only the template text isolates the template as the independent variable. Any measured difference can then be attributed to the template rather than to confounds. This is the core principle of a controlled experiment and directly answers the team's question.

Why this answer

A controlled comparison requires that only the variable of interest changes between conditions. By fixing the model, decoding parameters, and evaluation data and varying only the prompt template text, the team ensures that observed differences in summary quality are attributable to the template. This makes the experiment reproducible and the conclusions defensible.

Exam trap

The trap here is believing that testing each prompt with its own tuned decoding settings gives a fairer comparison.

75
Multi-Selecthard

Which THREE factors influence the reproducibility of an LLM experiment?

Select 3 answers
A.Random seed initialization
B.The ambient temperature of the datacenter
C.Version of the deep learning framework
D.Data preprocessing and splitting strategy
E.The number of hours the researcher works
AnswersA, C, D

In deep learning, random seeds affect initialization, data shuffling, and dropout masks. Setting a fixed seed ensures that the stochastic elements of the training run are deterministic, which is essential for verifying that performance improvements are due to algorithmic changes rather than the random state of the model parameters.

Why this answer

Reproducibility requires strict control over the experimental environment. Because LLM training involves non-deterministic factors, controlling the random seeds, the exact data splits, and the library versions is paramount. Professional experiment tracking requires the ability to recreate identical results, which is only possible when every component of the pipeline is version-controlled and explicitly defined in the experiment's configuration manifest.

Exam trap

Candidates frequently forget that hardware configurations or infrastructure metrics alone do not dictate LLM reproducibility without locking down random seeds and framework versions.

Page 1 of 2 · 82 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Experimentation questions.