MLA-C01 Data Preparation for Machine Learning Practice Question
A data scientist has a 200 GB Parquet dataset in Amazon S3 that will be used to train a SageMaker model. The training script reads the data with the SageMaker training toolkit's File mode, and the job currently spends a long time downloading before training begins. The team wants to reduce startup time without changing the training algorithm. Which change should the data scientist make?
⚠ Common exam trap
The trap here is conflating capacity with latency, adding disk space when the real delay is waiting for a complete copy of the data.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Switch the input channel to FastFile mode so the training container streams data instead of downloading the full dataset first.
FastFile mode removes the blocking full download by presenting S3 objects as a streamed filesystem inside the container, so the training process starts as soon as it reads the first bytes. This directly shortens the pre-training wait for a large Parquet dataset and requires no change to the training algorithm or data format.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Switch the input channel to FastFile mode so the training container streams data instead of downloading the full dataset first.
Why this is correct
SageMaker FastFile mode exposes S3 objects to the training container through a POSIX-like interface that streams bytes on demand, so the job can begin reading immediately without a full upfront download. For large Parquet datasets this substantially reduces the time before training starts while leaving the algorithm unchanged.
- ✗
Enable Pipe mode and reformat the Parquet files into the recordIO-protobuf format.
Why it's wrong here
Pipe mode streams data directly from S3, but it expects supported formats such as recordIO-protobuf or CSV, and Parquet is not natively supported by the built-in pipe-mode readers. Reformatting 200 GB adds a significant conversion step and forces changes to the data pipeline, which conflicts with the goal of a low-effort startup improvement.
- ✗
Add more instances to the training cluster and enable distributed data parallel.
Why it's wrong here
Distributed training spreads computation across instances, and each worker would still download its portion of the dataset under File mode before training starts. This increases cost and complexity without removing the upfront download, and the request is specifically to shorten startup while keeping the algorithm as is.
- ✗
Increase the volume size of the training instance so the full dataset fits in local storage.
Why it's wrong here
A larger attached volume gives the download more room, but File mode still copies the entire dataset before the script runs, so startup time is governed by download duration rather than disk capacity. It addresses a storage shortage, not the latency the team is trying to reduce.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.