Courseiva
Experimentation →mediumMultiple Choice

NCA-GENL Experimentation Practice Question

A team fine-tuning a NeMo large language model runs the same training configuration three times and obtains validation loss values of 2.14, 2.31, and 2.09 at the end of the same number of steps. They need subsequent runs to produce tightly clustered, comparable numbers so hyperparameter comparisons are meaningful. Which change most directly addresses this problem?

⚠ Common exam trap

The trap here is assuming that a larger batch or a different optimizer removes variance, when only controlling randomness and nondeterministic kernels makes repeated runs comparable.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set a fixed random seed and enable deterministic behavior for the training run.

Run-to-run variance in loss comes from uncontrolled randomness: weight initialization, dropout sampling, data shuffling, and nondeterministic GPU kernels. Pinning a seed and enabling deterministic execution makes those sources identical across repeated runs, so differences in outcomes can be attributed to the hyperparameters under test rather than to chance.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Add early stopping based on validation loss plateau detection.

    Why it's wrong here

    Early stopping halts training when validation loss stops improving, which can shorten runs but does not make them reproducible. Because each run starts from different random initialization and sees differently shuffled data, the stopping point and the final loss value still vary, leaving the comparison problem unsolved.

  • ✗

    Increase the global batch size while holding the learning rate constant.

    Why it's wrong here

    A larger batch reduces gradient noise and can shift the final loss, but it does not remove the underlying nondeterminism from random initialization, dropout sampling, and data-order shuffling. The three runs would still diverge because their random draws differ, so the variance the team is seeing would largely remain.

  • ✓

    Set a fixed random seed and enable deterministic behavior for the training run.

    Why this is correct

    A fixed seed makes weight initialization, dropout masks, and data shuffling identical across runs, and enabling deterministic kernels removes nondeterministic GPU operations, so repeated runs converge to nearly identical loss values. This directly targets the run-to-run variance observed and makes hyperparameter comparisons valid.

  • ✗

    Switch the optimizer from AdamW to SGD with momentum.

    Why it's wrong here

    Changing the optimizer alters the optimization trajectory and may change absolute loss values, but it does nothing to eliminate randomness from initialization, dropout, or data ordering. Repeated runs would still produce scattered results, so this does not establish the run-to-run comparability the team requires.

About these practice questions

Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.