Courseiva

NCA-GENL Core Machine Learning and AI Knowledge Practice Question

A team is pretraining a large language model on a cluster of NVIDIA GPUs. They observe that the model's training loss decreases steadily for the first few epochs but then suddenly spikes and eventually becomes NaN. They suspect this is due to exploding gradients. Which technique is most appropriate to address this issue?

⚠ Common exam trap

The trap here is assuming that any change to the optimizer or batch size will fix exploding gradients, when the direct and reliable solution is to clip gradients.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply gradient clipping to limit the magnitude of gradients during backpropagation.

Exploding gradients cause large parameter updates that can destabilize training, leading to loss spikes and NaN values. Gradient clipping is a standard technique that caps gradient norms or values before the optimizer step, preventing such destructive updates. The other options either exacerbate the problem or do not directly target the root cause, making gradient clipping the most appropriate remedy for this scenario.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Apply gradient clipping to limit the magnitude of gradients during backpropagation.

    Why this is correct

    Gradient clipping caps the norm or value of gradients before the optimizer step, preventing excessively large updates that destabilize training. In this scenario, the sudden loss spike and NaN indicate exploding gradients, so clipping directly mitigates the problem. It is a standard, low-overhead intervention that preserves the model architecture and data pipeline, making it the most appropriate immediate fix.

  • ✗

    Reduce the batch size to decrease the variance of gradient estimates.

    Why it's wrong here

    Reducing batch size increases the variance of gradient estimates, which can make training noisier and potentially worsen instability. While smaller batches sometimes help generalization, they do not reliably prevent exploding gradients. The observed loss spike is better addressed by controlling gradient magnitude directly, not by altering batch statistics.

  • ✗

    Increase the learning rate to help the model escape the spiking region.

    Why it's wrong here

    Increasing the learning rate amplifies update steps, which worsens exploding gradients and accelerates divergence to NaN. The scenario already shows instability, so a larger learning rate would likely cause earlier failure. The correct approach is to stabilize updates, not to make them more aggressive. Thus, this option is counterproductive for the described symptom.

  • ✗

    Switch from Adam to stochastic gradient descent (SGD) without momentum.

    Why it's wrong here

    Changing the optimizer might alter dynamics, but SGD without momentum does not inherently prevent exploding gradients; it can still produce large updates if gradients are large. Moreover, such a switch could slow convergence and require extensive retuning. It does not directly address the root cause of gradient explosion, so it is not the most appropriate fix in this situation.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.