NCA-GENL Experimentation Practice Question
An ML engineer is running a hyperparameter sweep with NVIDIA NeMo and notices that runs with identical configurations sometimes produce slightly different final loss values. Which cause should the engineer investigate first?
⚠ Common exam trap
The trap here is blaming hardware quantity or dataset size when the real culprit is uncontrolled randomness and nondeterministic kernels in the training computation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Nondeterministic GPU operations and unseeded randomness in the data pipeline
Run-to-run variation with identical configuration points to uncontrolled randomness and nondeterministic computation. Unseeded data shuffling changes which samples appear in which batch, and nondeterministic GPU kernels can accumulate floating-point differences. Together these produce slightly different loss trajectories. Validation set size and checkpoint intervals affect measurement and I/O, not the training computation itself.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The checkpoint saving interval is too frequent
Why it's wrong here
Checkpoint frequency affects storage and I/O overhead, not the numerical trajectory of training. Saving more or fewer checkpoints does not alter gradients, data order, or kernel behavior. It is a plausible operational knob but irrelevant to why two identical jobs end at slightly different losses, so it should not be the first area investigated.
- ✓
Nondeterministic GPU operations and unseeded randomness in the data pipeline
Why this is correct
Identical configurations can still diverge when nondeterministic CUDA kernels, unseeded data shuffling, or uncontrolled worker threads introduce variation. Atomic operations and certain reduction kernels do not guarantee identical ordering across runs, and if the random seed is not fixed or the dataloader workers reseed independently, each run sees a different data order, producing small but real loss differences.
- ✗
The number of GPUs is too high for the batch size being used
Why it's wrong here
More GPUs than needed can reduce per-device work, but it does not by itself cause run-to-run variance when the configuration is otherwise identical. If anything, changing the GPU count changes the effective global batch size, which would be visible in the configuration. The symptom described is variation with the same configuration, so hardware count is not the first suspect.
- ✗
The validation set is too small to measure loss accurately
Why it's wrong here
A small validation set makes the reported metric noisy across different evaluations, but here the same runs are being repeated. It explains variance in validation scores, not in final training loss across identical runs. The engineer should first address sources that change the computation itself, such as kernel nondeterminism and unseeded shuffling.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.