Courseiva
ML Model Development →hardMultiple Choice

MLA-C01 ML Model Development Practice Question

A company is fine-tuning a large language model using LoRA on SageMaker. They want to reduce GPU memory usage during training. Which configuration change would help?

⚠ Common exam trap

The trap is choosing gradient accumulation because it is a common memory-related technique — but MLA-C01 tests that only quantization (QLoRA) reduces the base model's memory footprint, while accumulation and batch/sequence changes affect activation memory differently.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use QLoRA (quantized LoRA) with 4-bit quantization

QLoRA extends LoRA by quantizing the frozen base model weights to 4-bit (typically NF4) while keeping LoRA adapters in higher precision, dramatically reducing GPU memory required for fine-tuning. This lets you train larger models on smaller GPUs with minimal accuracy loss, directly addressing the goal of reducing memory usage. It is the standard memory-optimization technique for LoRA fine-tuning on SageMaker.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use QLoRA (quantized LoRA) with 4-bit quantization

    Why this is correct

    QLoRA quantises the frozen base model weights to 4-bit, so they occupy roughly a quarter of the memory that 16-bit weights require, while LoRA adapters remain trainable in higher precision. This directly satisfies the stem's constraint of reducing GPU memory during fine-tuning, with minimal accuracy loss.

  • ✗

    Enable gradient accumulation

    Why it's wrong here

    Gradient accumulation reduces the number of optimiser updates by summing gradients over several micro-batches, but each micro-batch still holds its full activation memory, so peak GPU memory is unchanged. It is the correct choice when the goal is a larger effective batch size than GPU memory would otherwise permit.

  • ✗

    Increase the sequence length

    Why it's wrong here

    Increasing sequence length raises activation and attention memory roughly quadratically, so peak GPU usage grows rather than falls. It is tempting because longer sequences can improve model quality on long documents, making it the correct change when the objective is handling longer inputs, not conserving memory.

  • ✗

    Increase the batch size

    Why it's wrong here

    Larger batches increase activation memory, raising GPU usage rather than reducing it; gradient checkpointing or lower precision would cut it. Batch size tuning targets throughput and convergence stability, so it would be the right lever when training time, not memory, is the bottleneck.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.