MLA-C01 Data Preparation for Machine Learning Practice Question
A data scientist is preparing a CSV dataset in Amazon S3 for a SageMaker training job. Several rows contain missing values in numeric feature columns, and the chosen algorithm cannot handle NaNs. The scientist wants a repeatable, code-based transformation that runs inside a SageMaker Processing job before training. Which step is the MOST appropriate?
⚠ Common exam trap
The trap here is treating missing-value handling as an algorithm setting or a storage-lifecycle concern, when it is really a preprocessing step that belongs in a scripted Processing job.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Write a preprocessing script that uses pandas to impute or drop missing values and run it in a SageMaker Processing job with a scikit-learn container.
Running a pandas-based cleaning script in a SageMaker Processing job makes the imputation or row removal reproducible and code-driven, and the job reads from and writes to S3 so the cleaned dataset is available to training. This is the canonical pattern for repeatable preprocessing that must run before a training job and be re-executed as data changes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use an S3 Lifecycle rule to expire objects containing missing values so only clean data remains in the bucket.
Why it's wrong here
S3 Lifecycle rules operate on object age, storage class, or tags, not on the contents of CSV rows. They cannot inspect values inside a file, so they cannot selectively remove rows with missing features. This option misapplies a storage-management feature to a data-quality problem and would not produce a clean training dataset.
- ✗
Open the CSV in a SageMaker notebook, manually edit the missing cells, and save the file back to S3.
Why it's wrong here
Manual edits are not repeatable, are not captured in code, and do not scale beyond small files. They also risk introducing inconsistent values and leave no auditable record of the transformation. For a production ML pipeline, this approach fails the requirement for a repeatable, code-based transformation that can be rerun as new data arrives.
- ✗
Configure the SageMaker training job's input channel with a content type that tells the algorithm to ignore NaN values automatically.
Why it's wrong here
SageMaker input channels pass data to the algorithm but do not provide a generic NaN-handling switch. Whether missing values are tolerated depends on the algorithm and framework, and many built-in algorithms reject them outright. Relying on a content type to silently drop or ignore missing values is not a supported or reliable mechanism for this preprocessing requirement.
- ✓
Write a preprocessing script that uses pandas to impute or drop missing values and run it in a SageMaker Processing job with a scikit-learn container.
Why this is correct
SageMaker Processing jobs run a containerized script against data in S3 and write outputs back to S3, which makes the transformation repeatable, versionable, and independent of notebook state. Using pandas for imputation or row removal inside the scikit-learn container is a standard, well-supported approach that produces a clean dataset the training job can consume directly.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.