AI0-001 Implementing AI Solutions Practice Question
A company is fine-tuning a large language model using PEFT (Parameter-Efficient Fine-Tuning) to reduce GPU memory usage. They have limited hardware and need to fine-tune a 70B parameter model on a single GPU with 24 GB VRAM. Which technique is MOST suitable?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
QLoRA (Quantization-aware LoRA) with 4-bit quantization
QLoRA combines quantization (4-bit) and LoRA to fine-tune very large models on limited hardware, achieving significant memory reduction while maintaining performance.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Full fine-tuning with gradient checkpointing
Why it's wrong here
Full fine-tuning with gradient checkpointing still updates all 70B parameters, so optimiser states and gradients alone exceed 24 GB VRAM. Gradient checkpointing only trades compute for activation memory, not parameter memory. It would suit scenarios where the whole model fits and activations dominate, but PEFT's low-rank adapters are required here.
- ✓
QLoRA (Quantization-aware LoRA) with 4-bit quantization
Why this is correct
QLoRA quantises the frozen base weights to 4-bit NF4 and trains only small LoRA adapters, cutting memory enough to fine-tune a 70B model on a single 24 GB GPU. Plain LoRA or full fine-tuning cannot fit within that VRAM budget.
- ✗
Instruction tuning with a smaller 7B model
Why it's wrong here
Instruction tuning with a 7B model changes the base model rather than the fine-tuning method, so it abandons the stated 70B target entirely. It is tempting because instruction tuning genuinely adapts a model to follow prompts, and a 7B model would fit 24 GB VRAM — but that solves a different scenario, not PEFT on 70B parameters.
- ✗
LoRA (Low-Rank Adaptation) alone
Why it's wrong here
LoRA alone still trains adapters across every layer, so a 70B model's optimiser states and activations exceed 24 GB. It is tempting because LoRA is the standard PEFT method, and it would be correct on a larger GPU or with aggressive quantisation such as QLoRA.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.