A company is fine-tuning an LLM for a domain-specific task using LoRA. They have limited GPU memory and need to reduce memory footprint without sacrificing fine-tuning quality. Which approach should they consider?
QLoRA quantises the frozen base model to 4-bit and trains low-rank adapters, cutting GPU memory substantially while preserving fine-tuning quality. This directly satisfies the stem's limited-memory constraint, unlike standard LoRA, which still loads full-precision base weights.
Why this answer
QLoRA combines 4-bit NormalFloat quantization of the base model with LoRA adapters, drastically reducing GPU memory usage while preserving fine-tuning quality through techniques like double quantization and paged optimizers. This directly addresses the constraint of limited GPU memory without sacrificing the model's ability to learn domain-specific tasks effectively.
Exam trap
CompTIA often tests the misconception that increasing model capacity (e.g., higher LoRA rank or full fine-tuning) always improves quality, when in fact memory-constrained environments require efficient techniques like QLoRA that balance resource usage and performance.
How to eliminate wrong answers
Option B is wrong because increasing batch size increases GPU memory consumption, which is counterproductive when memory is limited. Option C is wrong because fine-tuning all layers of the base model requires full gradient storage and optimizer states for every parameter, dramatically increasing memory footprint and defeating the purpose of memory reduction. Option D is wrong because increasing the rank of LoRA adapters increases the number of trainable parameters and their associated optimizer states, raising memory usage without guaranteeing improved fine-tuning quality.