Courseiva

CCNA Nca Experimentation Questions

7 of 82 questions · Page 2/2 · Nca Experimentation topic · Answers revealed

76
MCQhard

A research team is running a hyperparameter sweep for LoRA fine-tuning of a 13B-parameter LLM on NVIDIA GPUs. They notice that runs with identical configurations sometimes produce noticeably different evaluation scores, and the variance is larger than the differences between some of the hyperparameter settings being compared. Which action best addresses this problem?

A.Report only the single best score achieved by each configuration, since the best result shows the model's potential.
B.Increase the learning rate for all runs so the model converges faster and the noise is averaged out.
C.Run each configuration with multiple fixed random seeds and compare mean evaluation scores with their confidence intervals.
D.Reduce the number of sweep configurations so fewer comparisons are affected by variance.
AnswerC

Repeating each configuration across several fixed seeds converts a single noisy observation into a distribution, and comparing means with confidence intervals reveals whether a hyperparameter effect exceeds run-to-run noise. This directly addresses the reported problem, where variance is larger than the differences between settings, by quantifying uncertainty before drawing conclusions.

Why this answer

When run-to-run variance exceeds the differences between configurations, single-run comparisons cannot support a conclusion. Repeating each configuration with several fixed seeds and reporting mean scores with confidence intervals turns noise into a measurable uncertainty, so the team can tell whether a hyperparameter effect is real. This is the standard remedy for high-variance sweeps and prevents selecting configurations on the basis of lucky seeds.

Exam trap

The trap here is believing that more runs or a best-of-N summary solves variance, when only replication with controlled seeds and uncertainty reporting makes the comparison valid.

77
MCQeasy

When conducting A/B testing on an LLM-based application, which metric is the most reliable indicator of user satisfaction regarding response quality?

A.Inference latency
B.Tokens per second
C.LLM-as-a-Judge score
D.GPU memory utilization
AnswerC

Using a stronger model to evaluate the outputs of the model under test provides a consistent, scalable quality metric. It mimics human evaluation, allowing for rapid experimentation cycles where qualitative performance can be measured against specific criteria like reasoning, tone, and accuracy in a reproducible and automated manner.

Why this answer

Human-in-the-loop evaluation or automated model-based grading (LLM-as-a-Judge) are the gold standards for response quality. Unlike latency or throughput, which measure system performance, quality metrics quantify the utility and accuracy of the generated text. In AI experimentation, balancing technical performance with qualitative user feedback is essential to ensure that system optimizations do not inadvertently degrade the helpfulness of the model output.

Exam trap

Many candidates choose infrastructure metrics like latency or throughput to measure response quality, failing to realize these only reflect system performance rather than semantic correctness.

78
Multi-Selectmedium

A research team is running a series of controlled LLM fine-tuning experiments on NVIDIA DGX systems using NeMo Framework to compare two learning-rate schedules. They want the comparison to be scientifically valid and repeatable by another engineer next quarter. Which two practices are required to make the experiments reproducible? (Choose two.)

Select 2 answers
A.Deploy the resulting checkpoints to NVIDIA Triton Inference Server for latency benchmarking.
B.Increase the number of GPUs in every run so that throughput is maximized.
C.Record the exact dataset version, tokenizer configuration, and NeMo container image tag used.
D.Enable mixed precision training to reduce memory usage on each DGX node.
E.Seed all random number generators, including data shuffling, dropout, and weight initialization.
AnswersC, E

Capturing dataset version, tokenizer settings, and container image tag pins the environment so a later engineer can rebuild the identical setup. Tokenizer changes alter token counts and therefore effective sequence lengths, while a different container may ship different library versions. Without these records, the learning-rate comparison cannot be faithfully repeated next quarter.

Why this answer

Reproducibility in a controlled fine-tuning comparison requires controlling randomness and pinning the environment. Seeding all stochastic operations keeps data order, dropout, and initialization identical, while recording dataset version, tokenizer configuration, and container image tag lets another engineer rebuild the same setup. Performance-oriented choices such as GPU count, mixed precision, and inference benchmarking do not establish repeatability of the training comparison.

Exam trap

The trap here is confusing performance optimizations, such as adding GPUs or enabling mixed precision, with the controls that actually make a training experiment reproducible.

79
MCQeasy

Which of the following describes the 'Stop-Loss' technique in the context of LLM experimentation?

A.Stopping training after a fixed number of days
B.Terminating training if loss does not improve
C.Increasing the learning rate when loss is high
D.Saving a checkpoint every 100 iterations
AnswerB

This is the definition of a stop-loss strategy. By setting a patience threshold for metric improvement, you can programmatically kill experiments that have stalled, effectively managing cloud spend and compute availability. It ensures that GPU resources are utilized only for runs that demonstrate potential for reaching the desired performance targets.

Why this answer

Stop-loss is an essential technique for managing compute costs. By monitoring training metrics in real-time, the system automatically halts experiments that show no sign of convergence or diverge prematurely. This prevents the waste of expensive GPU resources on doomed runs, allowing researchers to reallocate capacity to more promising experiments and improving the overall efficiency of the research team's pipeline.

Exam trap

Candidates often confuse 'Stop-Loss' with 'Early Stopping' or 'Gradient Clipping'. They fail to distinguish between a general training strategy and a specific cost-saving experimentation technique.

80
MCQmedium

An ML engineer is running a hyperparameter sweep with NVIDIA NeMo and notices that runs with identical configurations sometimes produce slightly different final loss values. Which cause should the engineer investigate first?

A.The checkpoint saving interval is too frequent
B.Nondeterministic GPU operations and unseeded randomness in the data pipeline
C.The number of GPUs is too high for the batch size being used
D.The validation set is too small to measure loss accurately
AnswerB

Identical configurations can still diverge when nondeterministic CUDA kernels, unseeded data shuffling, or uncontrolled worker threads introduce variation. Atomic operations and certain reduction kernels do not guarantee identical ordering across runs, and if the random seed is not fixed or the dataloader workers reseed independently, each run sees a different data order, producing small but real loss differences.

Why this answer

Run-to-run variation with identical configuration points to uncontrolled randomness and nondeterministic computation. Unseeded data shuffling changes which samples appear in which batch, and nondeterministic GPU kernels can accumulate floating-point differences. Together these produce slightly different loss trajectories.

Validation set size and checkpoint intervals affect measurement and I/O, not the training computation itself.

Exam trap

The trap here is blaming hardware quantity or dataset size when the real culprit is uncontrolled randomness and nondeterministic kernels in the training computation.

81
Multi-Selecthard

When designing an experiment to evaluate the performance of an LLM on a downstream classification task, which THREE factors should be controlled to ensure the results are comparable across different model sizes?

Select 3 answers
A.Decoding strategy (temperature, top-p)
B.The prompt template used
C.The learning rate used during fine-tuning
D.The validation/test dataset
E.The hardware GPU generation (e.g., A100 vs H100)
AnswersA, B, D

The decoding strategy directly influences the stochastic nature of the output. If one model uses greedy search and another uses high-temperature sampling, the performance differences are skewed by the sampling method. Consistency here is essential to isolate the model's architectural capacity from its generation behavior during the experiment.

Why this answer

To conduct valid cross-model comparisons, one must neutralize external variables. Inference-time settings like decoding parameters, the specific prompt template, and the evaluation dataset must remain constant. If these vary, the results reflect differences in the evaluation environment rather than the intrinsic capabilities of the models being tested, rendering the experimentation data inconclusive for determining which model size is truly optimal for the specific classification application.

Exam trap

Candidates often overlook inference-time parameters like temperature and top-p, assuming that only the dataset and model weights need to remain identical during comparative evaluations.

82
MCQhard

During an ablation study on a retrieval-augmented LLM pipeline, the team removes the reranking stage and observes a large drop in answer accuracy on their benchmark. Before concluding that reranking is essential, which additional experiment is most important to run?

A.Run a control condition that keeps the pipeline identical except for a neutral change unrelated to reranking.
B.Increase the number of retrieved documents to compensate for the missing reranker.
C.Replace the benchmark with a larger one to increase statistical power.
D.Re-run the ablation with a different random seed to see if the accuracy drop persists.
AnswerA

A control isolates the effect of the intervention from other differences between the ablated and baseline runs. If a neutral modification produces little change while removing reranking produces a large drop, the conclusion that reranking drives accuracy is much stronger. This is the key validity check before attributing the effect to the removed component.

Why this answer

An ablation shows that a change in the pipeline correlates with an accuracy drop, but correlation is not causation unless other differences are ruled out. A control condition with a neutral modification tests whether the pipeline is sensitive to arbitrary changes. If the control shows little effect while removing reranking causes a large drop, the causal claim about reranking is substantially strengthened.

Exam trap

The trap here is jumping from an observed accuracy drop to a causal claim about reranking without checking whether the comparison is confounded by other pipeline differences.

← PreviousPage 2 of 2 · 82 questions total

Ready to test yourself?

Try a timed practice session using only Nca Experimentation questions.