NCA-GENL Core Machine Learning and AI Knowledge Practice Question
An ML engineer is training a transformer-based language model on a single NVIDIA A100 GPU. They observe that the training loss decreases initially but then becomes NaN after a few hundred steps. The learning rate is 1e-4, and mixed precision with FP16 is enabled. Which action is most likely to stabilize training while preserving the benefits of mixed precision?
⚠ Common exam trap
The trap here is assuming that any change to training hyperparameters or precision will fix NaNs, when the specific cause in mixed precision is often gradient underflow that requires loss scaling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable dynamic loss scaling to automatically adjust the loss scale factor during training.
Dynamic loss scaling is the standard technique to prevent underflow and overflow in FP16 mixed precision training. It scales the loss up to keep gradients in representable range and automatically backs off when NaNs occur. This stabilizes training while retaining the performance and memory advantages of mixed precision, unlike switching to FP32 or altering hyperparameters.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Switch to FP32 training for all layers.
Why it's wrong here
Switching entirely to FP32 would eliminate the numerical instability from FP16 overflow, but it also removes the memory and speed benefits of mixed precision. The goal is to preserve mixed precision, so this is an overly broad solution that reduces training efficiency and may not be necessary if dynamic loss scaling can resolve the issue.
- ✗
Reduce the batch size to 1 to minimize memory usage.
Why it's wrong here
Reducing batch size might lower memory pressure but does not fix FP16 numerical overflow. In fact, smaller batches can introduce more gradient noise, potentially worsening instability. The NaN issue is related to loss scaling, not batch size, so this change would not reliably stabilize training.
- ✓
Enable dynamic loss scaling to automatically adjust the loss scale factor during training.
Why this is correct
Dynamic loss scaling multiplies the loss by a large factor to prevent small gradients from underflowing in FP16, and it automatically reduces the scale when overflows (NaNs) are detected. This directly addresses the NaN issue while keeping FP16 computation, thus preserving mixed precision benefits and stabilizing training.
- ✗
Increase the learning rate to 1e-3 to escape the NaN region faster.
Why it's wrong here
Increasing the learning rate would likely exacerbate the instability, causing more frequent NaN occurrences due to larger gradient updates. The problem is numerical overflow in FP16, not a suboptimal learning rate. A higher learning rate would not address the root cause and would probably make training diverge further.
About these practice questions
Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.