Courseiva

NCA-GENL Core Machine Learning and AI Knowledge Practice Question

A machine learning engineer is training a large language model on a cluster of NVIDIA GPUs. During training, she observes that the loss occasionally spikes to NaN, causing the training to fail. She suspects that the issue is related to the numerical precision of the computations. Which technique is most appropriate to mitigate this issue while maintaining training stability?

⚠ Common exam trap

The trap here is assuming that reducing batch size or increasing learning rate can fix NaN losses, when the root cause is often gradient explosion that gradient clipping directly addresses.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use gradient clipping

Gradient clipping is a widely used technique to prevent exploding gradients, which are a common cause of NaN losses in deep learning, especially when training large models with mixed precision. By capping gradient norms, it ensures that parameter updates remain stable, allowing training to proceed without numerical overflow.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Switch to double precision

    Why it's wrong here

    Switching to double precision (FP64) would increase numerical range and reduce overflow risk, but it is computationally expensive and typically not feasible for large language models due to memory and speed constraints. It is overkill when gradient clipping can effectively mitigate the issue while maintaining mixed precision training.

  • ✗

    Increase the learning rate

    Why it's wrong here

    Increasing the learning rate would likely exacerbate the problem by causing larger weight updates, which can lead to more frequent NaN losses. It does not address numerical precision issues. In this scenario, a higher learning rate could destabilize training further, making NaN spikes more common and preventing convergence.

  • ✗

    Reduce the batch size

    Why it's wrong here

    Reducing the batch size changes the gradient estimation but does not directly prevent numerical overflow. Smaller batches may introduce more noise, potentially worsening stability. While it can sometimes help with memory issues, it is not the primary solution for NaN losses caused by precision problems in this scenario.

  • ✓

    Use gradient clipping

    Why this is correct

    Gradient clipping limits the magnitude of gradients during backpropagation, preventing explosive gradients that can cause numerical overflow and NaN losses. In training large language models, especially with mixed precision, gradient clipping is a standard technique to maintain stability. It directly addresses the symptom of loss spikes by keeping updates within a manageable range.

About these practice questions

This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.