easyMultiple Choice
Generative AI Leader Practice Question: The key advantage of using adapter-based…
What is the key advantage of using adapter-based fine-tuning methods like LoRA compared to full fine-tuning of a large language model?
⚠ Common exam trap
Google often tests the misconception that parameter-efficient methods like LoRA improve inference speed, when in reality they primarily reduce memory during training and do not accelerate inference.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
LoRA significantly reduces the number of trainable parameters, making fine-tuning more memory-efficient
LoRA (Low-Rank Adaptation) injects trainable low-rank matrices into the transformer layers while keeping the original model weights frozen. This drastically reduces the number of trainable parameters (often by 10,000x), which lowers GPU memory requirements for storing optimizer states and gradients during training, making fine-tuning feasible on consumer hardware.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
LoRA significantly reduces the number of trainable parameters, making fine-tuning more memory-efficient
Why this is correct
LoRA freezes the base weights and injects small trainable low-rank matrices into attention layers, cutting trainable parameters by orders of magnitude. This slashes optimiser and gradient memory, so fine-tuning fits on far cheaper GPUs than full fine-tuning, which updates every weight.
- ✗
LoRA is faster at inference time compared to the fully fine-tuned model
Why it's wrong here
LoRA adds adapter matrices that are merged into base weights or evaluated alongside them, so inference cost matches the fully fine-tuned model rather than beating it. Its benefit is training-time: far fewer trainable parameters and smaller optimiser state. It is tempting because LoRA reduces training resource use, but that saving does not transfer to serving latency.
- ✗
LoRA eliminates the need for a base model
Why it's wrong here
LoRA freezes the pretrained weights and trains small low-rank adapter matrices, so the base model is still required at inference; it reduces trainable parameters and memory, not the base model itself. It would be the right choice when GPU memory or storage for full weight updates is constrained.
- ✗
LoRA enables training on a larger dataset than full fine-tuning
Why it's wrong here
LoRA does not raise the dataset ceiling; dataset size is bounded by memory and training budget regardless of method. Its advantage is training only small low-rank adapter matrices, leaving base weights frozen. It is tempting because LoRA does cut GPU memory and checkpoint size, which can let you fit more data per device — but that is a memory effect, not a larger-dataset capability.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.