NCP-GENL Fine-Tuning Practice Question
A team is fine-tuning a 13B-parameter Llama model with NVIDIA NeMo Framework on a single A100 80GB GPU. They apply LoRA adapters to the attention projection layers, but the adapters are producing negligible changes to model behavior even after several epochs, and the loss curve stays flat. They confirm the dataset is clean and the tokenizer is correct. Which LoRA configuration issue is the most likely cause?
⚠ Common exam trap
The trap here is assuming any LoRA misconfiguration equally explains a flat loss, when the alpha-to-rank scaling ratio specifically controls how much the adapter can influence the forward pass.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The LoRA alpha scaling value is set extremely low relative to the rank, so the adapter's effective update magnitude is near zero.
LoRA updates are scaled by alpha divided by rank, so an alpha that is very small relative to rank shrinks the adapter's effective contribution and can leave the loss nearly unchanged. Correcting the alpha-to-rank ratio restores meaningful adapter influence, allowing the fine-tune to actually shift model behavior on the target task.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The optimizer was configured with a momentum value of zero, preventing the adapter weights from accumulating updates.
Why it's wrong here
Momentum of zero simply means plain SGD updates without velocity; weights would still update on every step. A flat loss is not explained by lack of momentum alone, especially with a reasonable learning rate. This is a plausible-sounding but incorrect cause for stalled adapter learning.
- ✗
The LoRA target modules were set to the embedding layer rather than the attention projection layers.
Why it's wrong here
Applying LoRA to embeddings would still produce some learning signal, though it is less effective for task adaptation. The question states adapters were applied to attention projection layers, so this scenario does not match. Even if it did, it would not fully explain a completely flat loss with negligible behavior change.
- ✗
The base model weights were accidentally left trainable, so gradients are flowing into the full model instead of the adapters.
Why it's wrong here
If base weights were trainable, the loss would typically still decrease, just with higher memory and slower steps. A flat loss with no behavioral change points to adapter updates being too small, not to extra trainable parameters. This misconfiguration would more likely cause memory pressure than a stalled loss curve.
- ✓
The LoRA alpha scaling value is set extremely low relative to the rank, so the adapter's effective update magnitude is near zero.
Why this is correct
The LoRA scaling factor is alpha divided by rank. If alpha is tiny compared to rank (for example alpha=1 with rank=64), the adapter's contribution to the forward pass is scaled to a negligible magnitude, so the frozen base weights dominate and the loss barely moves. Raising alpha or lowering rank restores a meaningful update magnitude.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.