MLA-C01 ML Model Development Practice Question
A machine learning team at a bank is training a binary classification model using SageMaker's built-in XGBoost algorithm on a dataset with 20 million rows and 300 features. They need to reduce training time while maintaining model accuracy. The data is stored in Amazon S3 as CSV files. Which approach should they take to speed up training?
⚠ Common exam trap
The trap here is assuming that simply adding more instances or preprocessing data will speed up training, when the bottleneck is often data loading from S3.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use SageMaker Pipe mode with the RecordIO protobuf format instead of File mode with CSV.
SageMaker's built-in XGBoost can consume data in Pipe mode with RecordIO protobuf, which streams data directly from S3 and reduces I/O overhead, leading to faster training on large datasets. File mode with CSV requires downloading all data to disk, which is slower. Distributed training options for XGBoost are limited and not configured via parameter_distribution. Thus, switching to Pipe mode with RecordIO is the recommended approach.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use SageMaker Processing to preprocess the data into a single large CSV file and then train with File mode.
Why it's wrong here
Combining data into a single CSV file does not address the I/O bottleneck; File mode still downloads the entire file to disk before training. Preprocessing adds overhead and does not reduce training time. This approach is inefficient for large datasets and does not leverage SageMaker's optimized data streaming capabilities.
- ✓
Use SageMaker Pipe mode with the RecordIO protobuf format instead of File mode with CSV.
Why this is correct
Pipe mode streams data directly from S3 to the training container, eliminating the need to download the full dataset to disk. RecordIO protobuf is a compact binary format that reduces I/O overhead and allows efficient shuffling. For large datasets with many features, this significantly speeds up training and reduces disk usage, while maintaining accuracy by feeding all data.
- ✗
Increase the number of instances in the training cluster and enable distributed training with the parameter_distribution parameter set to 'fully'.
Why it's wrong here
SageMaker's built-in XGBoost supports distributed training only in certain modes, and using parameter_distribution='fully' is not a valid option for the built-in algorithm. Increasing instances can help but requires proper configuration; simply adding instances without correct distribution may not scale linearly and can increase cost without guaranteed speedup.
- ✗
Convert the CSV files to TFRecord format and use File mode to load the data.
Why it's wrong here
TFRecord is a format designed for TensorFlow, not for SageMaker's built-in XGBoost algorithm. Using File mode still downloads the entire dataset to each instance's disk, which is slow for large datasets. This approach would not reduce training time and may even cause compatibility issues, as XGBoost expects CSV or RecordIO protobuf.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.