NCP-GENL Fine-Tuning Practice Question
A team is fine-tuning a Llama 2 7B model with NVIDIA NeMo Framework on a single A100 80GB GPU. They observe that training loss decreases initially but then diverges, and the model outputs become repetitive and incoherent. The team used a learning rate of 5e-5 with AdamW and no warm-up. Which change is most likely to stabilize training and improve convergence?
⚠ Common exam trap
The trap here is assuming that a higher learning rate always speeds up convergence, when in fine-tuning it often causes divergence and degraded output quality.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Reduce the learning rate to 1e-5 and add a linear warm-up over the first 10% of training steps.
The model diverges due to a learning rate that is too high for fine-tuning without warm-up. Reducing the learning rate to 1e-5 and adding a warm-up phase allows the optimizer to take smaller, more controlled steps initially, preventing large updates that disrupt pretrained weights. This combination is a well-established best practice in NVIDIA NeMo and other frameworks for stable fine-tuning of LLMs.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the learning rate to 1e-4 to escape local minima.
Why it's wrong here
Increasing the learning rate would likely exacerbate divergence, as the current rate is already too high for stable fine-tuning. A higher rate can cause parameter updates to overshoot optimal values, leading to worse loss and incoherent outputs. This change does not address the root cause of instability, which is an inappropriately high learning rate without warm-up.
- ✗
Switch the optimizer to SGD with momentum 0.9 and keep the learning rate at 5e-5.
Why it's wrong here
SGD with momentum may not converge as efficiently as AdamW for fine-tuning LLMs, and keeping the same high learning rate could still cause instability. SGD often requires more tuning and may not handle the sparse gradients typical in language models as well as AdamW. This change does not directly mitigate the divergence caused by the high learning rate.
- ✓
Reduce the learning rate to 1e-5 and add a linear warm-up over the first 10% of training steps.
Why this is correct
A lower learning rate (e.g., 1e-5) combined with warm-up prevents large initial updates that can destabilize training. Warm-up gradually increases the learning rate, allowing the model to adapt smoothly. This is a standard practice for fine-tuning large language models, especially when full fine-tuning or using AdamW, and directly addresses the observed divergence and repetitive outputs.
- ✗
Increase the batch size to 64 and keep all other hyperparameters unchanged.
Why it's wrong here
Increasing batch size can reduce gradient noise, but it does not address the fundamental issue of an excessively high learning rate. With a larger batch, the effective learning rate per sample might decrease slightly, but the optimizer step size remains the same, so divergence could still occur. This change alone is unlikely to stabilize training or fix repetitive outputs.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.