easyMultiple Choice
MLA-C01 Practice Question: A data scientist is working on a binary…
A data scientist is working on a binary classification problem and wants to use AWS Glue for data preparation. The dataset has missing values in several numeric columns. Which imputation strategy is MOST appropriate for the scientist to apply in AWS Glue ETL?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the FillMissingValues transform to replace missing values with the mean of each column
AWS Glue ETL (PySpark) supports the `Imputer` transformer which can impute missing numeric values using the mean or median of the column. This is a built-in, straightforward approach.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use the FillMissingValues transform to replace missing values with the mean of each column
Why this is correct
FillMissingValues (or Imputer) with mean is a standard imputation strategy for numeric features.
- ✗
Use a machine learning model to predict missing values
Why it's wrong here
While using a machine learning model to predict missing values is a powerful imputation technique, it's overkill for the described scenario. AWS Glue ETL's primary purpose is data transformation and preparation, not complex model training for imputation. This approach is tempting because it offers sophisticated imputation, ideal for situations where relationships between variables are intricate and simple statistical methods would be insufficient, but it introduces unnecessary computational overhead and complexity for a standard missing value problem within an ETL pipeline.
- ✗
Drop all rows with missing values using the Drop transform
Why it's wrong here
Dropping rows discards potentially useful data; imputation is preferred unless missingness is very high.
- ✗
Set missing values to zero
Why it's wrong here
Setting to zero may introduce bias and distort distributions; mean/median is generally better.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.