Courseiva
Fine-Tuning →mediumMultiple Choice

NCP-GENL Fine-Tuning Practice Question

A team is fine-tuning a Llama 2 7B model with NVIDIA NeMo Framework on a single A100 80GB GPU. They observe that training loss decreases initially but then diverges, and the model outputs become repetitive and incoherent. The team used a learning rate of 5e-5 with AdamW and no warm-up. Which change is most likely to stabilize training and improve convergence?

⚠ Common exam trap

The trap here is assuming that a higher learning rate always speeds up convergence, when in fine-tuning it often causes divergence and degraded output quality.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Reduce the learning rate to 1e-5 and add a linear warm-up over the first 10% of training steps.

The model diverges due to a learning rate that is too high for fine-tuning without warm-up. Reducing the learning rate to 1e-5 and adding a warm-up phase allows the optimizer to take smaller, more controlled steps initially, preventing large updates that disrupt pretrained weights. This combination is a well-established best practice in NVIDIA NeMo and other frameworks for stable fine-tuning of LLMs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the learning rate to 1e-4 to escape local minima.

    Why it's wrong here

    Increasing the learning rate would likely exacerbate divergence, as the current rate is already too high for stable fine-tuning. A higher rate can cause parameter updates to overshoot optimal values, leading to worse loss and incoherent outputs. This change does not address the root cause of instability, which is an inappropriately high learning rate without warm-up.

  • ✗

    Switch the optimizer to SGD with momentum 0.9 and keep the learning rate at 5e-5.

    Why it's wrong here

    SGD with momentum may not converge as efficiently as AdamW for fine-tuning LLMs, and keeping the same high learning rate could still cause instability. SGD often requires more tuning and may not handle the sparse gradients typical in language models as well as AdamW. This change does not directly mitigate the divergence caused by the high learning rate.

  • ✓

    Reduce the learning rate to 1e-5 and add a linear warm-up over the first 10% of training steps.

    Why this is correct

    A lower learning rate (e.g., 1e-5) combined with warm-up prevents large initial updates that can destabilize training. Warm-up gradually increases the learning rate, allowing the model to adapt smoothly. This is a standard practice for fine-tuning large language models, especially when full fine-tuning or using AdamW, and directly addresses the observed divergence and repetitive outputs.

  • ✗

    Increase the batch size to 64 and keep all other hyperparameters unchanged.

    Why it's wrong here

    Increasing batch size can reduce gradient noise, but it does not address the fundamental issue of an excessively high learning rate. With a larger batch, the effective learning rate per sample might decrease slightly, but the optimizer step size remains the same, so divergence could still occur. This change alone is unlikely to stabilize training or fix repetitive outputs.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.