NCP-GENL Fine-Tuning Practice Question
When fine-tuning a Large Language Model using Low-Rank Adaptation (LoRA), which architectural component is primarily modified to reduce computational overhead while maintaining performance?
⚠ Common exam trap
Candidates often mistakenly believe LoRA modifies the entire transformer block, failing to identify that it specifically targets the attention weight projection matrices to optimize memory and computational efficiency.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The attention weight projection matrices
LoRA freezes the pre-trained model weights and injects trainable rank decomposition matrices into the transformer architecture layers. By targeting specifically the attention query, key, and value projection matrices, practitioners can achieve high performance with a fraction of the trainable parameters. This approach is critical for memory-constrained environments, allowing fine-tuning on consumer-grade NVIDIA GPUs while avoiding the massive memory requirements associated with full parameter updates.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The entire embedding layer matrix
Why it's wrong here
Modifying the entire embedding layer increases the number of trainable parameters significantly, which contradicts the primary goal of LoRA. Efficient fine-tuning relies on low-rank updates rather than full parameter re-training, as updating embeddings would require substantial memory and compute resources that LoRA specifically aims to bypass during training.
- ✗
The feed-forward network activation functions
Why it's wrong here
While feed-forward layers are part of the transformer, LoRA specifically targets the weight matrices of the attention mechanism. Modifying activation functions would not provide the rank decomposition benefit necessary to reduce parameter space, nor would it align with the standard implementation of LoRA as defined in current research literature.
- ✓
The attention weight projection matrices
Why this is correct
LoRA injects trainable low-rank matrices into the attention mechanism's query, key, and value projections. By adapting these specific components, the model learns to capture domain-specific patterns without updating the billions of frozen parameters, drastically reducing VRAM consumption and making model fine-tuning feasible on limited NVIDIA hardware resources.
- ✗
The entire decoder hidden state output
Why it's wrong here
Updating the decoder hidden states directly would lead to catastrophic forgetting and would not constitute a parameter-efficient fine-tuning method. LoRA functions by freezing the original weights and introducing adapters, whereas modifying the output directly would interfere with the fundamental inference mechanics of the pre-trained transformer model architecture.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.