NCA-GENL Experimentation Practice Question
Exhibit
config_json: { "optimizer": "adamw", "mixed_precision": "fp16", "grad_clip": 1.0, "warmup_steps": 500, "weight_decay": 0.1 }Refer to the exhibit. The experiment shows the model is failing to converge and exhibits loss spikes. Which adjustment to the configuration is most likely to stabilize the training process?
⚠ Common exam trap
Test-takers often attempt to resolve mixed-precision loss spikes by increasing batch size or changing optimizers, ignoring the limited dynamic range constraints inherent to FP16.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Change mixed_precision to bf16.
Loss spikes in mixed-precision training are often caused by the limited dynamic range of FP16. Lowering the gradient clip value or switching to BF16 (if hardware permits) are common remedies to maintain stability. By analyzing the configuration during experimentation, researchers can identify these hyperparameter sensitivities and prevent training failures, ensuring more robust and efficient model development cycles.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the weight decay to 0.5.
Why it's wrong here
Increasing weight decay too aggressively can restrict the model's ability to learn complex patterns in the data, potentially exacerbating underfitting. While it helps with regularization, it is not a primary solution for the numerical instability or loss spikes typically associated with FP16 training configurations.
- ✓
Change mixed_precision to bf16.
Why this is correct
FP16 has a limited dynamic range that often leads to overflow/underflow issues during large-scale model training, resulting in loss spikes. BF16 provides the same dynamic range as FP32, making it significantly more stable for training, especially when using modern NVIDIA hardware that supports it natively.
- ✗
Reduce the warmup_steps to 0.
Why it's wrong here
Warmup steps are crucial for stabilizing the optimizer state during the initial phase of training. Removing or reducing warmup steps will likely increase the probability of divergence as the model encounters large gradient updates before the optimizer has a chance to adapt, rather than solving the instability.
- ✗
Switch the optimizer to standard SGD.
Why it's wrong here
Switching from AdamW to standard SGD would likely result in slower convergence without addressing the root cause of the loss spikes, which is usually related to floating-point precision issues. AdamW is specifically designed to handle weight decay better than standard SGD, making it more robust for LLMs.
About these practice questions
Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.