MLA-C01 ML Model Development Practice Question
A company is fine-tuning a large language model using LoRA on SageMaker. They want to reduce GPU memory usage during training. Which configuration change would help?
⚠ Common exam trap
The trap is choosing gradient accumulation because it is a common memory-related technique — but MLA-C01 tests that only quantization (QLoRA) reduces the base model's memory footprint, while accumulation and batch/sequence changes affect activation memory differently.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use QLoRA (quantized LoRA) with 4-bit quantization
QLoRA extends LoRA by quantizing the frozen base model weights to 4-bit (typically NF4) while keeping LoRA adapters in higher precision, dramatically reducing GPU memory required for fine-tuning. This lets you train larger models on smaller GPUs with minimal accuracy loss, directly addressing the goal of reducing memory usage. It is the standard memory-optimization technique for LoRA fine-tuning on SageMaker.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use QLoRA (quantized LoRA) with 4-bit quantization
Why this is correct
QLoRA quantises the frozen base model weights to 4-bit, so they occupy roughly a quarter of the memory that 16-bit weights require, while LoRA adapters remain trainable in higher precision. This directly satisfies the stem's constraint of reducing GPU memory during fine-tuning, with minimal accuracy loss.
- ✗
Enable gradient accumulation
Why it's wrong here
Gradient accumulation reduces the number of optimiser updates by summing gradients over several micro-batches, but each micro-batch still holds its full activation memory, so peak GPU memory is unchanged. It is the correct choice when the goal is a larger effective batch size than GPU memory would otherwise permit.
- ✗
Increase the sequence length
Why it's wrong here
Increasing sequence length raises activation and attention memory roughly quadratically, so peak GPU usage grows rather than falls. It is tempting because longer sequences can improve model quality on long documents, making it the correct change when the objective is handling longer inputs, not conserving memory.
- ✗
Increase the batch size
Why it's wrong here
Larger batches increase activation memory, raising GPU usage rather than reducing it; gradient checkpointing or lower precision would cut it. Batch size tuning targets throughput and convergence stability, so it would be the right lever when training time, not memory, is the bottleneck.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.