MLA-C01 Data Preparation for Machine Learning Practice Question
A data scientist is preparing a large dataset for training a machine learning model. The dataset contains missing values in several columns. Which approach is the MOST efficient for handling missing values in a large dataset using AWS services?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Amazon SageMaker Data Wrangler to impute missing values using built-in transforms.
Amazon SageMaker Data Wrangler provides a visual interface and built-in transforms for handling missing values efficiently at scale, without writing custom code. Glue ETL is more code-heavy, and imputation with pandas is not scalable for large datasets. Removing all rows with missing values is not always optimal and may not be efficient.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use AWS Glue ETL to write a custom Python script that imputes missing values with the mean.
Why it's wrong here
A custom Python script in AWS Glue imputes values but requires hand-written code and offers no built-in ML-based missing-value handling. It is tempting because Glue is serverless and scales across large datasets, making it the right choice for bespoke transformation logic.
- ✓
Use Amazon SageMaker Data Wrangler to impute missing values using built-in transforms.
Why this is correct
Data Wrangler provides built-in imputation transforms that run as scalable Spark processing, avoiding custom code for a large dataset. This satisfies the efficiency constraint by handling missing values across many columns in one visual flow, with results exportable directly to SageMaker training.
- ✗
Use pandas in a SageMaker notebook to impute missing values with the median.
Why it's wrong here
Pandas loads the dataset into a single notebook's memory, which cannot scale to large data and offers no distributed processing. It is tempting because pandas is the familiar tool for quick median imputation on small, single-node datasets.
- ✗
Remove all rows with missing values from the dataset.
Why it's wrong here
Dropping every row containing a null discards substantial training data and biases the model toward complete cases, which is unacceptable at scale. It is tempting because listwise deletion is trivial to implement and valid when missingness is rare and completely at random.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.