Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A data engineer is preparing a large training dataset stored in Amazon S3 as many small Parquet files, and a SageMaker training job that reads directly from S3 is spending most of its time on the input channel rather than on model computation. The engineer needs to improve the input throughput without changing the model code or the training algorithm. Which action should the engineer take?

⚠ Common exam trap

The trap here is assuming that any SageMaker input-mode or acceleration feature will fix slow data loading, when the real issue is per-object overhead from many small files rather than raw bandwidth.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the SageMaker File System Input with Amazon FSx for Lustre linked to the S3 bucket.

Feeding a training job from a high-performance shared file system removes the request-per-small-object overhead that dominates when many tiny Parquet files are read directly from S3. Amazon FSx for Lustre linked to the S3 bucket presents the data as files that the training container mounts and reads at high throughput, improving the input channel without altering model code or the training algorithm, which is exactly what the scenario requires.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable SageMaker Training Compiler and set the framework to accelerated mode.

    Why it's wrong here

    Training Compiler rewrites the model graph to speed up computation on GPU instances; it does not improve how training data is read from S3. Since the bottleneck is the input channel rather than model computation, this would leave the actual problem untouched and add unnecessary configuration and compatibility constraints to the training job.

  • ✓

    Use the SageMaker File System Input with Amazon FSx for Lustre linked to the S3 bucket.

    Why this is correct

    FSx for Lustre linked to the S3 bucket exposes the dataset as a high-throughput, low-latency POSIX file system that SageMaker training jobs can mount, so the many small Parquet files are read far faster than repeated S3 GET requests. It improves input throughput without changing model code or the training algorithm, satisfying the requirement.

  • ✗

    Increase the number of records per S3 GET request by enabling S3 Transfer Acceleration on the bucket.

    Why it's wrong here

    S3 Transfer Acceleration speeds up transfers over long geographic distances by routing through edge locations; it does not batch multiple objects into one GET and does nothing for the per-object request overhead caused by many small files. The training job is presumably in the same Region, so this would not resolve the input channel bottleneck.

  • ✗

    Convert the output to TFRecord format and use Pipe mode with the SageMaker TensorFlow estimator.

    Why it's wrong here

    Pipe mode streams data through a Linux FIFO, which can reduce startup time, but TFRecord is a format suited to TensorFlow training jobs, not a general fix for a framework-agnostic input bottleneck. It also requires changing how the algorithm consumes records, which the engineer was told not to do, so it does not address the stated constraint.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.