Courseiva
ML Model Development →hardMultiple Choice

MLA-C01 ML Model Development Practice Question

An ML engineer is fine-tuning a large language model using LoRA on SageMaker. The training is converging slowly, and GPU utilization is low. The engineer suspects the bottleneck is data loading. Which action should the engineer take to improve GPU utilization?

⚠ Common exam trap

The trap is treating GPU underutilization as a compute or memory problem — candidates reach for batch size or parallelism changes, but low GPU utilization with slow convergence almost always points to an input pipeline bottleneck that Pipe mode is designed to solve.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use SageMaker Pipe mode to stream data from S3 directly to the training instances

SageMaker Pipe mode streams training data directly from Amazon S3 to the training container over a high-throughput channel, bypassing the local disk and the download-then-read pattern of File mode. This removes the I/O bottleneck that starves the GPU when the dataset is large or when many small files cause slow reads. With faster data delivery, the GPU spends less time idle waiting for batches, raising utilization and speeding convergence.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the batch size to maximize GPU memory usage

    Why it's wrong here

    Larger batches increase memory pressure and step time without addressing the input pipeline, so GPUs still stall waiting on data. It is tempting because batch size tuning often improves throughput, but that applies when compute, not data loading, is the bottleneck; here the loader must be accelerated.

  • ✗

    Enable checkpointing and use spot instances

    Why it's wrong here

    Checkpointing and spot instances address training resilience and cost, not input pipeline throughput; they leave the data loading bottleneck untouched, so GPU utilisation stays low. Spot instances suit interruptible, cost-sensitive jobs, and checkpointing preserves progress across interruptions — neither feeds data to the GPU faster.

  • ✓

    Use SageMaker Pipe mode to stream data from S3 directly to the training instances

    Why this is correct

    Pipe mode streams training data directly from Amazon S3 to the instances, removing the download-and-store step that starves the GPU. This raises input throughput and GPU utilisation, addressing the data-loading bottleneck the engineer suspects during LoRA fine-tuning.

  • ✗

    Reduce model parallelism to decrease communication overhead

    Why it's wrong here

    Reducing model parallelism lowers inter-GPU communication, which matters when GPUs stall on collective operations, not on data loading. With low GPU utilisation blamed on the input pipeline, cutting parallelism changes nothing; it would help a communication-bound job where all-reduce or pipeline bubbles dominate step time.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.