NCA-GENL Core Machine Learning and AI Knowledge Practice Question
Exhibit
2023-10-27 10:00:01 [INFO] Training Epoch 5/10 completed 2023-10-27 10:00:05 [WARNING] Gradient norm 452.1 exceeds threshold 1.0 2023-10-27 10:00:06 [ERROR] Divergence detected: Loss is NaN
Refer to the exhibit. Which technique is most appropriate to prevent this specific training failure?
⚠ Common exam trap
Candidates often confuse exploding gradients with vanishing gradients, incorrectly suggesting that increasing the learning rate or adding deeper layers will resolve numerical instability.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply gradient clipping during the backpropagation step.
The exhibit shows a classic case of the exploding gradient problem, where large weight updates lead to numerical instability and NaN loss values. This is common in deep architectures, particularly RNNs or deep Transformers. Gradient clipping is the standard solution to constrain the norm of the gradients during backpropagation, ensuring that updates remain within a numerically stable range, thereby preventing the model parameters from reaching infinity or NaN states.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the learning rate to bypass the threshold.
Why it's wrong here
Increasing the learning rate will exacerbate the exploding gradient problem. Larger steps will result in even bigger gradient values, accelerating the divergence of the loss function. This approach is counterproductive when the model is already reporting NaN values due to unstable updates and excessively large parameter changes during training.
- ✓
Apply gradient clipping during the backpropagation step.
Why this is correct
Gradient clipping scales the gradient vector if its norm exceeds a defined threshold, ensuring the updates remain within a stable range. This prevents extreme weight changes that cause numerical overflow, effectively stopping the NaN divergence observed in the logs while allowing the training process to continue successfully.
- ✗
Remove the normalization layers to increase model capacity.
Why it's wrong here
Removing normalization layers often makes training more unstable, not less. Normalization helps stabilize the distribution of activations, which generally prevents gradients from growing uncontrollably. Removing these layers would likely make the model even more susceptible to the gradient explosions seen in the provided error log output.
- ✗
Change the optimizer from Adam to SGD without momentum.
Why it's wrong here
Changing the optimizer does not address the underlying issue of gradients exceeding numerical limits. While SGD without momentum might converge slower, it does not provide the safety mechanism required to handle exploding gradients. Gradient clipping remains the targeted solution regardless of which specific gradient-based optimizer is being utilized.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.