Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A data scientist has a 200 GB Parquet dataset in Amazon S3 that will be used to train a SageMaker model. The training script reads the data with the SageMaker training toolkit's File mode, and the job currently spends a long time downloading before training begins. The team wants to reduce startup time without changing the training algorithm. Which change should the data scientist make?

⚠ Common exam trap

The trap here is conflating capacity with latency, adding disk space when the real delay is waiting for a complete copy of the data.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Switch the input channel to FastFile mode so the training container streams data instead of downloading the full dataset first.

FastFile mode removes the blocking full download by presenting S3 objects as a streamed filesystem inside the container, so the training process starts as soon as it reads the first bytes. This directly shortens the pre-training wait for a large Parquet dataset and requires no change to the training algorithm or data format.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Switch the input channel to FastFile mode so the training container streams data instead of downloading the full dataset first.

    Why this is correct

    SageMaker FastFile mode exposes S3 objects to the training container through a POSIX-like interface that streams bytes on demand, so the job can begin reading immediately without a full upfront download. For large Parquet datasets this substantially reduces the time before training starts while leaving the algorithm unchanged.

  • ✗

    Enable Pipe mode and reformat the Parquet files into the recordIO-protobuf format.

    Why it's wrong here

    Pipe mode streams data directly from S3, but it expects supported formats such as recordIO-protobuf or CSV, and Parquet is not natively supported by the built-in pipe-mode readers. Reformatting 200 GB adds a significant conversion step and forces changes to the data pipeline, which conflicts with the goal of a low-effort startup improvement.

  • ✗

    Add more instances to the training cluster and enable distributed data parallel.

    Why it's wrong here

    Distributed training spreads computation across instances, and each worker would still download its portion of the dataset under File mode before training starts. This increases cost and complexity without removing the upfront download, and the request is specifically to shorten startup while keeping the algorithm as is.

  • ✗

    Increase the volume size of the training instance so the full dataset fits in local storage.

    Why it's wrong here

    A larger attached volume gives the download more room, but File mode still copies the entire dataset before the script runs, so startup time is governed by download duration rather than disk capacity. It addresses a storage shortage, not the latency the team is trying to reduce.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.