Courseiva

NCA-GENL Core Machine Learning and AI Knowledge Practice Question

An AI researcher is fine-tuning a large language model and wants to minimize GPU memory consumption during training without altering the model's primary weight representations or introducing quantization error during inference. Which technique provides this capability by decomposing weight matrices into low-rank trainable adaptation matrices?

⚠ Common exam trap

Candidates frequently confuse LoRA with post-training quantization methods like INT4 or INT8. While quantization compresses weights for inference, LoRA is an efficient parameter-efficient fine-tuning technique that preserves original precision weight matrices.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Low-Rank Adaptation (LoRA) freezing pre-trained weights and training injected rank decomposition matrices.

Low-Rank Adaptation (LoRA) freezes the original pre-trained model weights and injects trainable rank decomposition matrices into the architecture. This drastically reduces the number of trainable parameters and optimizer states in GPU memory during backpropagation, while allowing weights to be merged back into the original matrices for zero-latency deployment.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Post-Training Quantization (PTQ) converting weights from FP16 to INT8 integer representations.

    Why it's wrong here

    PTQ alters the primary weight representation itself, converting FP16 to INT8, which introduces exactly the quantization error the scenario forbids. It is tempting because it genuinely reduces memory and speeds inference, and would be correct when deploying a frozen model where some accuracy loss is acceptable and no further training is needed.

  • ✓

    Low-Rank Adaptation (LoRA) freezing pre-trained weights and training injected rank decomposition matrices.

    Why this is correct

    LoRA freezes the pre-trained weights and injects trainable rank decomposition matrices, so only these small adapters update during fine-tuning. This cuts GPU memory use substantially while leaving the original weight representations untouched, and because the base weights stay full precision, no quantization error is introduced at inference.

  • ✗

    Full-Parameter Fine-Tuning updating all weights across every transformer layer simultaneously using FP32 precision.

    Why it's wrong here

    Updating all weights in FP32 maximises optimiser-state and activation memory, the opposite of the requirement. It is tempting as the baseline for maximum quality, but LoRA freezes base weights and trains low-rank matrices, cutting memory without quantising inference weights.

  • ✗

    Knowledge Distillation transferring learned representations from a large teacher model to a smaller student network architecture.

    Why it's wrong here

    Distillation trains a smaller student model from a teacher's outputs, changing the deployed architecture and requiring full retraining; it never decomposes existing weight matrices into low-rank adapters. It is tempting because it does cut inference cost and model size, and would be right when the goal is a permanently smaller production model rather than memory-efficient fine-tuning.

About these practice questions

Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.