Courseiva
Experimentation →hardMultiple Choice

NCA-GENL Experimentation Practice Question

Exhibit

config_json: { "optimizer": "adamw", "mixed_precision": "fp16", "grad_clip": 1.0, "warmup_steps": 500, "weight_decay": 0.1 }

Refer to the exhibit. The experiment shows the model is failing to converge and exhibits loss spikes. Which adjustment to the configuration is most likely to stabilize the training process?

⚠ Common exam trap

Test-takers often attempt to resolve mixed-precision loss spikes by increasing batch size or changing optimizers, ignoring the limited dynamic range constraints inherent to FP16.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Change mixed_precision to bf16.

Loss spikes in mixed-precision training are often caused by the limited dynamic range of FP16. Lowering the gradient clip value or switching to BF16 (if hardware permits) are common remedies to maintain stability. By analyzing the configuration during experimentation, researchers can identify these hyperparameter sensitivities and prevent training failures, ensuring more robust and efficient model development cycles.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the weight decay to 0.5.

    Why it's wrong here

    Increasing weight decay too aggressively can restrict the model's ability to learn complex patterns in the data, potentially exacerbating underfitting. While it helps with regularization, it is not a primary solution for the numerical instability or loss spikes typically associated with FP16 training configurations.

  • ✓

    Change mixed_precision to bf16.

    Why this is correct

    FP16 has a limited dynamic range that often leads to overflow/underflow issues during large-scale model training, resulting in loss spikes. BF16 provides the same dynamic range as FP32, making it significantly more stable for training, especially when using modern NVIDIA hardware that supports it natively.

  • ✗

    Reduce the warmup_steps to 0.

    Why it's wrong here

    Warmup steps are crucial for stabilizing the optimizer state during the initial phase of training. Removing or reducing warmup steps will likely increase the probability of divergence as the model encounters large gradient updates before the optimizer has a chance to adapt, rather than solving the instability.

  • ✗

    Switch the optimizer to standard SGD.

    Why it's wrong here

    Switching from AdamW to standard SGD would likely result in slower convergence without addressing the root cause of the loss spikes, which is usually related to floating-point precision issues. AdamW is specifically designed to handle weight decay better than standard SGD, making it more robust for LLMs.

About these practice questions

Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.