Courseiva
Fine-Tuning →mediumMultiple Choice

NCP-GENL Fine-Tuning Practice Question

When fine-tuning a Large Language Model using Low-Rank Adaptation (LoRA), which architectural component is primarily modified to reduce computational overhead while maintaining performance?

⚠ Common exam trap

Candidates often mistakenly believe LoRA modifies the entire transformer block, failing to identify that it specifically targets the attention weight projection matrices to optimize memory and computational efficiency.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The attention weight projection matrices

LoRA freezes the pre-trained model weights and injects trainable rank decomposition matrices into the transformer architecture layers. By targeting specifically the attention query, key, and value projection matrices, practitioners can achieve high performance with a fraction of the trainable parameters. This approach is critical for memory-constrained environments, allowing fine-tuning on consumer-grade NVIDIA GPUs while avoiding the massive memory requirements associated with full parameter updates.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The entire embedding layer matrix

    Why it's wrong here

    Modifying the entire embedding layer increases the number of trainable parameters significantly, which contradicts the primary goal of LoRA. Efficient fine-tuning relies on low-rank updates rather than full parameter re-training, as updating embeddings would require substantial memory and compute resources that LoRA specifically aims to bypass during training.

  • ✗

    The feed-forward network activation functions

    Why it's wrong here

    While feed-forward layers are part of the transformer, LoRA specifically targets the weight matrices of the attention mechanism. Modifying activation functions would not provide the rank decomposition benefit necessary to reduce parameter space, nor would it align with the standard implementation of LoRA as defined in current research literature.

  • ✓

    The attention weight projection matrices

    Why this is correct

    LoRA injects trainable low-rank matrices into the attention mechanism's query, key, and value projections. By adapting these specific components, the model learns to capture domain-specific patterns without updating the billions of frozen parameters, drastically reducing VRAM consumption and making model fine-tuning feasible on limited NVIDIA hardware resources.

  • ✗

    The entire decoder hidden state output

    Why it's wrong here

    Updating the decoder hidden states directly would lead to catastrophic forgetting and would not constitute a parameter-efficient fine-tuning method. LoRA functions by freezing the original weights and introducing adapters, whereas modifying the output directly would interfere with the fundamental inference mechanics of the pre-trained transformer model architecture.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.