Courseiva
Software Development →easyMultiple Choice

NCA-GENL Software Development Practice Question

When fine-tuning a Large Language Model using PEFT (Parameter-Efficient Fine-Tuning) techniques like LoRA, what is the primary technical advantage being leveraged?

⚠ Common exam trap

Students frequently assume PEFT updates all model weights with a smaller learning rate, misunderstanding that base weights remain completely frozen.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Weight matrices are frozen and low-rank adaptors are trained.

LoRA injects trainable rank decomposition matrices into the transformer layers while keeping pre-trained weights frozen. This approach significantly reduces the number of parameters requiring gradient updates, which saves memory and compute. This is essential for developers working with limited GPU hardware, allowing them to adapt massive models to specific tasks without full-parameter fine-tuning, which would be computationally prohibitive for most enterprise-grade infrastructure deployments.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Full parameter updates are performed on every layer.

    Why it's wrong here

    Full parameter updates characterize traditional fine-tuning, not PEFT or LoRA. This process is extremely resource-intensive and often requires massive GPU clusters, making it impractical for quick adaptation tasks where efficiency and hardware footprint are primary constraints for the software development lifecycle.

  • ✓

    Weight matrices are frozen and low-rank adaptors are trained.

    Why this is correct

    LoRA freezes the original model weights and adds small, trainable rank decomposition matrices to transformer layers. This drastically reduces the total count of trainable parameters, leading to much lower VRAM usage during the training process while maintaining high performance on the target downstream tasks.

  • ✗

    The entire model is converted to INT8 precision during training.

    Why it's wrong here

    While quantization can be used alongside LoRA, the core benefit of LoRA is the reduction of trainable parameters through rank decomposition. Quantization is a separate optimization technique that reduces the memory footprint of the weights themselves, not the count of weights that are updated during optimization.

  • ✗

    All model activations are stored in CPU memory.

    Why it's wrong here

    Storing activations in CPU memory would create a massive performance bottleneck due to the slow PCIe bus transfer rates between the GPU and CPU. Efficient fine-tuning keeps the active tensors on the GPU to maintain high throughput and minimize the training time of the adapted model.

About these practice questions

This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.