NCA-GENL Core Machine Learning and AI Knowledge Practice Question
Exhibit
2023-10-27 10:00:01 INFO: Training iteration 5000/10000 2023-10-27 10:00:05 WARNING: Gradient norm 500.2 exceeds threshold 1.0 2023-10-27 10:00:06 ERROR: Model weights contain NaN values
Refer to the exhibit. Which technique is most effective for preventing the reported NaN error during model training?
⚠ Common exam trap
Candidates often try to resolve NaN errors by lowering the learning rate or increasing batch size, which are indirect fixes, instead of applying gradient clipping to explicitly prevent exploding gradients.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Gradient Clipping
The exhibit shows a classic case of exploding gradients, where the gradient norm spikes and leads to non-finite weight values. Gradient clipping is the standard industry practice to mitigate this, as it caps the gradient magnitude during backpropagation. Implementing this ensures stability in deep networks, preventing training collapses that waste expensive GPU compute resources and time in large-scale AI development cycles.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Gradient Clipping
Why this is correct
Gradient clipping involves re-scaling gradients if their norm exceeds a pre-defined threshold. By capping the update magnitude, it prevents the weights from exploding into non-finite numbers when the loss surface is steep. This is a critical stability measure for training large transformers and deep networks on high-performance compute clusters like NVIDIA DGX.
- ✗
Using a larger batch size
Why it's wrong here
Increasing batch size does not inherently prevent exploding gradients caused by the optimization landscape. While larger batches might provide smoother gradient estimates, they do not provide a hard limit on the magnitude of the backpropagated updates. NaN values will still occur if the learning rate or model initialization is poorly configured.
- ✗
Enabling Layer Normalization
Why it's wrong here
Layer normalization stabilizes hidden state distributions, which helps training convergence but is not a direct remedy for exploding gradients. When gradients are actively exploding, normalization may slightly shift the distribution, but it does not constrain the magnitude of updates during the backward pass to prevent non-finite weight values from occurring.
- ✗
Changing the optimizer to SGD
Why it's wrong here
Switching to Stochastic Gradient Descent (SGD) does not resolve exploding gradients if the architecture's learning rate is too high or the weights are poorly initialized. SGD still computes large updates on steep loss surfaces. The fundamental issue requires constraining the gradient magnitude itself rather than merely changing the update algorithm used during training.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.