Courseiva
easyMultiple Choice

MLA-C01 Practice Question: A data scientist is working on a binary…

A data scientist is working on a binary classification problem and wants to use AWS Glue for data preparation. The dataset has missing values in several numeric columns. Which imputation strategy is MOST appropriate for the scientist to apply in AWS Glue ETL?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the FillMissingValues transform to replace missing values with the mean of each column

AWS Glue ETL (PySpark) supports the `Imputer` transformer which can impute missing numeric values using the mean or median of the column. This is a built-in, straightforward approach.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use the FillMissingValues transform to replace missing values with the mean of each column

    Why this is correct

    FillMissingValues (or Imputer) with mean is a standard imputation strategy for numeric features.

  • ✗

    Use a machine learning model to predict missing values

    Why it's wrong here

    While using a machine learning model to predict missing values is a powerful imputation technique, it's overkill for the described scenario. AWS Glue ETL's primary purpose is data transformation and preparation, not complex model training for imputation. This approach is tempting because it offers sophisticated imputation, ideal for situations where relationships between variables are intricate and simple statistical methods would be insufficient, but it introduces unnecessary computational overhead and complexity for a standard missing value problem within an ETL pipeline.

  • ✗

    Drop all rows with missing values using the Drop transform

    Why it's wrong here

    Dropping rows discards potentially useful data; imputation is preferred unless missingness is very high.

  • ✗

    Set missing values to zero

    Why it's wrong here

    Setting to zero may introduce bias and distort distributions; mean/median is generally better.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.