Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A data engineer is using Amazon SageMaker Processing to run a data preprocessing script on a dataset with 500 million rows. The script runs out of memory on a single ml.r5.24xlarge instance. The engineer needs to modify the processing job to handle the dataset size. Which approach is most cost-effective and scalable?

⚠ Common exam trap

AWS often tests the misconception that increasing instance size or using swap space is the primary solution for memory issues, whereas the correct approach is to distribute the workload horizontally using SageMaker's built-in data sharding feature.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure the Processing job with multiple instances and use ShardedByS3Key for data splitting.

SageMaker Processing with ShardedByS3Key splits the input dataset by S3 object boundaries across multiple instances, allowing distributed processing of the 500 million rows without exceeding memory on any single instance. This approach is cost-effective as it uses multiple smaller instances (e.g., ml.r5.xlarge) rather than a single oversized instance, and scales linearly with data size.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Configure the Processing job with multiple instances and use ShardedByS3Key for data splitting.

    Why this is correct

    ShardedByS3Key distributes objects across multiple instances, so each processes a subset in parallel and memory per instance stays bounded. Scaling horizontally on smaller instances costs less than one oversized ml.r5.24xlarge and handles 500 million rows.

  • ✗

    Write the script to process data in chunks and write intermediate results to local ephemeral storage.

    Why it's wrong here

    Chunking to local ephemeral storage still executes on one instance, so it neither scales horizontally nor avoids the single-node memory ceiling. It is tempting because chunked processing is genuinely correct when data exceeds memory but the dataset fits on one node's disk and parallelism is unnecessary.

  • ✗

    Increase the instance type to a larger one like ml.p3dn.24xlarge with more memory.

    Why it's wrong here

    A single larger instance still processes the full dataset on one node, and ml.p3dn.24xlarge is GPU-accelerated, so its cost is not justified for preprocessing. It is tempting because vertical scaling is correct when memory demand is modest and the workload genuinely benefits from GPU acceleration.

  • ✗

    Reduce the number of instances to one and increase the volume size for swap space.

    Why it's wrong here

    Swap on a single instance's volume is orders of magnitude slower than RAM, and reducing to one instance removes the distributed parallelism SageMaker Processing provides. It is tempting because adding swap is a real remedy for marginal memory shortfalls on a single machine, not for 500 million rows.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.