Courseiva
Exploratory Data AnalysismediumMultiple ChoiceObjective-mapped

MLS-C01 Exploratory Data Analysis Practice Question

A data scientist is working with a dataset that has missing values in 30% of rows for a categorical feature 'city'. Which EDA step should be performed before deciding on imputation?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Check if missingness is related to other features or random

Before deciding on imputation for the 'city' feature, the first exploratory data analysis (EDA) step is to investigate the pattern of missingness. Option A is correct because you must determine whether the missing data are Missing Completely at Random (MCAR), Missing at Random (MAR), or Missing Not at Random (MNAR). This involves checking if missingness in 'city' is related to other features or is random. Understanding the missing mechanism informs the appropriate imputation strategy. Option B (impute with mode) is an imputation method, not a diagnostic step; applying it without prior analysis risks introducing bias. Option C (drop rows) may be valid only if missingness is MCAR and the amount of data loss is acceptable, but it should not be the first step. Option D (label encoding) transforms categorical data and does not address missing values.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Check if missingness is related to other features or random

    Why this is correct

    Before deciding on imputation, you must investigate the pattern of missingness to determine if it is MCAR, MAR, or MNAR. This involves checking if missingness in 'city' is related to other features or random. Understanding the missing mechanism informs the appropriate imputation strategy.

  • Impute missing values with the mode of the column

    Why it's wrong here

    Imputing missing values with the mode is an imputation method, not a diagnostic step. Applying it without prior analysis of the missingness pattern risks introducing bias and may lead to incorrect conclusions.

  • Drop all rows with missing values

    Why it's wrong here

    This is wrong because dropping rows with missing values is a data cleaning action, not an initial EDA step. It may be considered after analyzing the missingness pattern, but is not the first step.

  • Encode the city feature using label encoding

    Why it's wrong here

    This is wrong because label encoding is a categorical data transformation, not related to handling missing values. It does not address the missing data issue.

About these practice questions

Courseiva writes every MLS-C01 question from scratch — 1,672 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.