Courseiva
ML Model Development →mediumMultiple Choice

MLA-C01 ML Model Development Practice Question

A machine learning team at a bank is training a binary classification model using SageMaker's built-in XGBoost algorithm on a dataset with 20 million rows and 300 features. They need to reduce training time while maintaining model accuracy. The data is stored in Amazon S3 as CSV files. Which approach should they take to speed up training?

⚠ Common exam trap

The trap here is assuming that simply adding more instances or preprocessing data will speed up training, when the bottleneck is often data loading from S3.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use SageMaker Pipe mode with the RecordIO protobuf format instead of File mode with CSV.

SageMaker's built-in XGBoost can consume data in Pipe mode with RecordIO protobuf, which streams data directly from S3 and reduces I/O overhead, leading to faster training on large datasets. File mode with CSV requires downloading all data to disk, which is slower. Distributed training options for XGBoost are limited and not configured via parameter_distribution. Thus, switching to Pipe mode with RecordIO is the recommended approach.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use SageMaker Processing to preprocess the data into a single large CSV file and then train with File mode.

    Why it's wrong here

    Combining data into a single CSV file does not address the I/O bottleneck; File mode still downloads the entire file to disk before training. Preprocessing adds overhead and does not reduce training time. This approach is inefficient for large datasets and does not leverage SageMaker's optimized data streaming capabilities.

  • ✓

    Use SageMaker Pipe mode with the RecordIO protobuf format instead of File mode with CSV.

    Why this is correct

    Pipe mode streams data directly from S3 to the training container, eliminating the need to download the full dataset to disk. RecordIO protobuf is a compact binary format that reduces I/O overhead and allows efficient shuffling. For large datasets with many features, this significantly speeds up training and reduces disk usage, while maintaining accuracy by feeding all data.

  • ✗

    Increase the number of instances in the training cluster and enable distributed training with the parameter_distribution parameter set to 'fully'.

    Why it's wrong here

    SageMaker's built-in XGBoost supports distributed training only in certain modes, and using parameter_distribution='fully' is not a valid option for the built-in algorithm. Increasing instances can help but requires proper configuration; simply adding instances without correct distribution may not scale linearly and can increase cost without guaranteed speedup.

  • ✗

    Convert the CSV files to TFRecord format and use File mode to load the data.

    Why it's wrong here

    TFRecord is a format designed for TensorFlow, not for SageMaker's built-in XGBoost algorithm. Using File mode still downloads the entire dataset to each instance's disk, which is slow for large datasets. This approach would not reduce training time and may even cause compatibility issues, as XGBoost expects CSV or RecordIO protobuf.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.