NCP-GENL LLM Architecture Practice Question
A team is fine-tuning a pretrained decoder-only model on a small domain-specific dataset. They observe that the model quickly overfits and loses general language ability. They want to update only a small number of additional parameters while keeping the base weights frozen. Which approach should they use?
⚠ Common exam trap
The trap here is thinking that a low learning rate or extra dropout makes full fine-tuning parameter-efficient, when the base weights are still being updated.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Low-Rank Adaptation (LoRA), which injects trainable low-rank matrices into existing layers while freezing the base weights.
Parameter-efficient fine-tuning methods such as LoRA freeze the pretrained weights and train only small injected matrices, which reduces optimizer memory and limits drift from the base model. This is well suited to small domain datasets where full fine-tuning would overfit and degrade general language ability.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Full fine-tuning with a very low learning rate and early stopping.
Why it's wrong here
Full fine-tuning updates all base weights, which is exactly what the team wants to avoid to preserve general ability and reduce memory. Even with a low learning rate, catastrophic forgetting can occur on small datasets, so this does not meet the parameter-efficient requirement.
- ✗
Increasing the dropout rate in the feed-forward layers during fine-tuning.
Why it's wrong here
Dropout can regularize training and reduce overfitting to some degree, but it does not change which parameters are updated. The base weights are still fully trained, so memory usage and the risk of losing general language ability remain.
- ✗
Freezing the embedding layer only and training all transformer blocks.
Why it's wrong here
Freezing just the embeddings leaves the vast majority of parameters trainable, so optimizer memory and forgetting risk remain high. This is not parameter-efficient and does not satisfy the goal of updating only a small number of additional parameters.
- ✓
Low-Rank Adaptation (LoRA), which injects trainable low-rank matrices into existing layers while freezing the base weights.
Why this is correct
LoRA adds small trainable rank-decomposition matrices to selected weight matrices and leaves the pretrained weights frozen. This updates far fewer parameters, reduces optimizer memory, and is less prone to catastrophic forgetting, which matches the team's constraints.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.