MLA-C01 Data Preparation for Machine Learning Practice Question
A data engineer is using Amazon SageMaker Processing to run a data preprocessing script on a dataset with 500 million rows. The script runs out of memory on a single ml.r5.24xlarge instance. The engineer needs to modify the processing job to handle the dataset size. Which approach is most cost-effective and scalable?
⚠ Common exam trap
AWS often tests the misconception that increasing instance size or using swap space is the primary solution for memory issues, whereas the correct approach is to distribute the workload horizontally using SageMaker's built-in data sharding feature.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure the Processing job with multiple instances and use ShardedByS3Key for data splitting.
SageMaker Processing with ShardedByS3Key splits the input dataset by S3 object boundaries across multiple instances, allowing distributed processing of the 500 million rows without exceeding memory on any single instance. This approach is cost-effective as it uses multiple smaller instances (e.g., ml.r5.xlarge) rather than a single oversized instance, and scales linearly with data size.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Configure the Processing job with multiple instances and use ShardedByS3Key for data splitting.
Why this is correct
ShardedByS3Key distributes objects across multiple instances, so each processes a subset in parallel and memory per instance stays bounded. Scaling horizontally on smaller instances costs less than one oversized ml.r5.24xlarge and handles 500 million rows.
- ✗
Write the script to process data in chunks and write intermediate results to local ephemeral storage.
Why it's wrong here
Chunking to local ephemeral storage still executes on one instance, so it neither scales horizontally nor avoids the single-node memory ceiling. It is tempting because chunked processing is genuinely correct when data exceeds memory but the dataset fits on one node's disk and parallelism is unnecessary.
- ✗
Increase the instance type to a larger one like ml.p3dn.24xlarge with more memory.
Why it's wrong here
A single larger instance still processes the full dataset on one node, and ml.p3dn.24xlarge is GPU-accelerated, so its cost is not justified for preprocessing. It is tempting because vertical scaling is correct when memory demand is modest and the workload genuinely benefits from GPU acceleration.
- ✗
Reduce the number of instances to one and increase the volume size for swap space.
Why it's wrong here
Swap on a single instance's volume is orders of magnitude slower than RAM, and reducing to one instance removes the distributed parallelism SageMaker Processing provides. It is tempting because adding swap is a real remedy for marginal memory shortfalls on a single machine, not for 500 million rows.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.