Courseiva
Question 506 of 1,672
Exploratory Data AnalysishardMultiple SelectObjective-mapped

Missing Data Mechanisms — MCAR, MAR, MNAR

Which TWO statements about handling missing data during exploratory data analysis are correct? (Select TWO.)

Quick Answer

The correct answer is that understanding the missing data mechanism—whether MCAR, MAR, or MNAR—is critical for choosing an appropriate imputation strategy. This is because each mechanism implies a different relationship between the missingness and the data values: MCAR (Missing Completely at Random) means the missingness is unrelated to any data, MAR (Missing at Random) means it depends on observed data but not the missing values themselves, and MNAR (Missing Not at Random) means the missingness is related to the unobserved values. On the AWS Certified Machine Learning Specialty MLS-C01 exam, this concept tests your ability to avoid common pitfalls like using listwise deletion on non-MCAR data or mean imputation that reduces variance. A frequent trap is assuming mean imputation is always safe, but it distorts distributions under MAR or MNAR. Memory tip: MCAR is the only mechanism where you can safely delete rows without bias—think “MCAR = Missing Completely at Random = safe to drop.”

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Visualizing the pattern of missingness can help determine if data is missing at random.

Options B and C are correct. Visualizing the pattern of missingness helps determine if data is missing at random, which is a key EDA step. Understanding the missing data mechanism (MCAR, MAR, MNAR) is important for selecting an appropriate imputation strategy. Option A is incorrect because missing values should be addressed during EDA, not deferred to model training. Option D is incorrect because listwise deletion can introduce bias if data is not MCAR. Option E is incorrect because mean imputation reduces the variance of the imputed variable.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Missing values can be ignored during EDA and handled during model training.

    Why it's wrong here

    EDA should address missing data to inform preprocessing.

  • Visualizing the pattern of missingness can help determine if data is missing at random.

    Why this is correct

    Missingness patterns inform assumptions about missing data mechanisms.

  • Understanding the missing data mechanism (MCAR, MAR, MNAR) is important for choosing an imputation strategy.

    Why this is correct

    The mechanism affects the validity of imputation methods.

  • Listwise deletion (removing rows with missing values) is always safe and unbiased.

    Why it's wrong here

    Listwise deletion can bias results if data is not MCAR.

  • Imputing missing values with the mean preserves the original variance.

    Why it's wrong here

    Mean imputation reduces variance and can distort relationships.

About these practice questions

Courseiva creates original exam-style practice questions with explanations and wrong-answer analysis. It does not publish real exam questions, exam dumps, or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on MLS-C01

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. Which TWO statements about handling missing data during EDA are correct? (Select TWO.)

medium
  • A.Dropping columns with >50% missing values is always recommended.
  • B.Mean imputation preserves the variance of the original distribution.
  • C.If data are missing completely at random (MCAR), listwise deletion yields unbiased estimates.
  • D.Multiple imputation (MICE) is always the safest method regardless of missing data mechanism.
  • E.Imputing with the median is more robust to outliers than imputing with the mean.

Why C: Options C and E are correct. Option C is correct because when data are Missing Completely at Random (MCAR), the missingness is independent of both observed and unobserved data, so listwise deletion (removing rows with missing values) does not introduce bias; the remaining sample is still a random subsample. Option E is correct because the median is not influenced by extreme values, making it a more robust imputation method compared to the mean, which can be skewed by outliers. Option A is incorrect because dropping columns with >50% missing values is not always recommended; it depends on the importance of the variable and the analysis goals. Option B is incorrect because mean imputation reduces the variance of the imputed variable, as it forces imputed values to the center. Option D is incorrect because Multiple Imputation by Chained Equations (MICE) is not always the safest; it assumes data are Missing at Random (MAR) and can be complex or inappropriate for other missingness mechanisms.

Last reviewed: Jun 20, 2026

Question Discussion

Share a tip, memory trick, or ask about the reasoning behind this question. Do not post real exam questions, leaked content, braindumps, or copyrighted exam material. Comments are moderated and may be removed without notice.

Loading comments…

Sign in to join the discussion.

This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.