NCA-GENL Core Machine Learning and AI Knowledge Practice Question
An AI researcher is fine-tuning a large language model and wants to minimize GPU memory consumption during training without altering the model's primary weight representations or introducing quantization error during inference. Which technique provides this capability by decomposing weight matrices into low-rank trainable adaptation matrices?
⚠ Common exam trap
Candidates frequently confuse LoRA with post-training quantization methods like INT4 or INT8. While quantization compresses weights for inference, LoRA is an efficient parameter-efficient fine-tuning technique that preserves original precision weight matrices.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Low-Rank Adaptation (LoRA) freezing pre-trained weights and training injected rank decomposition matrices.
Low-Rank Adaptation (LoRA) freezes the original pre-trained model weights and injects trainable rank decomposition matrices into the architecture. This drastically reduces the number of trainable parameters and optimizer states in GPU memory during backpropagation, while allowing weights to be merged back into the original matrices for zero-latency deployment.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Post-Training Quantization (PTQ) converting weights from FP16 to INT8 integer representations.
Why it's wrong here
PTQ alters the primary weight representation itself, converting FP16 to INT8, which introduces exactly the quantization error the scenario forbids. It is tempting because it genuinely reduces memory and speeds inference, and would be correct when deploying a frozen model where some accuracy loss is acceptable and no further training is needed.
- ✓
Low-Rank Adaptation (LoRA) freezing pre-trained weights and training injected rank decomposition matrices.
Why this is correct
LoRA freezes the pre-trained weights and injects trainable rank decomposition matrices, so only these small adapters update during fine-tuning. This cuts GPU memory use substantially while leaving the original weight representations untouched, and because the base weights stay full precision, no quantization error is introduced at inference.
- ✗
Full-Parameter Fine-Tuning updating all weights across every transformer layer simultaneously using FP32 precision.
Why it's wrong here
Updating all weights in FP32 maximises optimiser-state and activation memory, the opposite of the requirement. It is tempting as the baseline for maximum quality, but LoRA freezes base weights and trains low-rank matrices, cutting memory without quantising inference weights.
- ✗
Knowledge Distillation transferring learned representations from a large teacher model to a smaller student network architecture.
Why it's wrong here
Distillation trains a smaller student model from a teacher's outputs, changing the deployed architecture and requiring full retraining; it never decomposes existing weight matrices into low-rank adapters. It is tempting because it does cut inference cost and model size, and would be right when the goal is a permanently smaller production model rather than memory-efficient fine-tuning.
About these practice questions
Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.