Question 506 of 1,672
Missing Data Mechanisms — MCAR, MAR, MNAR
Which TWO statements about handling missing data during exploratory data analysis are correct? (Select TWO.)
Quick Answer
The correct answer is that understanding the missing data mechanism—whether MCAR, MAR, or MNAR—is critical for choosing an appropriate imputation strategy. This is because each mechanism implies a different relationship between the missingness and the data values: MCAR (Missing Completely at Random) means the missingness is unrelated to any data, MAR (Missing at Random) means it depends on observed data but not the missing values themselves, and MNAR (Missing Not at Random) means the missingness is related to the unobserved values. On the AWS Certified Machine Learning Specialty MLS-C01 exam, this concept tests your ability to avoid common pitfalls like using listwise deletion on non-MCAR data or mean imputation that reduces variance. A frequent trap is assuming mean imputation is always safe, but it distorts distributions under MAR or MNAR. Memory tip: MCAR is the only mechanism where you can safely delete rows without bias—think “MCAR = Missing Completely at Random = safe to drop.”
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Visualizing the pattern of missingness can help determine if data is missing at random.
Options B and C are correct. Visualizing the pattern of missingness helps determine if data is missing at random, which is a key EDA step. Understanding the missing data mechanism (MCAR, MAR, MNAR) is important for selecting an appropriate imputation strategy. Option A is incorrect because missing values should be addressed during EDA, not deferred to model training. Option D is incorrect because listwise deletion can introduce bias if data is not MCAR. Option E is incorrect because mean imputation reduces the variance of the imputed variable.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Missing values can be ignored during EDA and handled during model training.
Why it's wrong here
EDA should address missing data to inform preprocessing.
- ✓
Visualizing the pattern of missingness can help determine if data is missing at random.
Why this is correct
Missingness patterns inform assumptions about missing data mechanisms.
- ✓
Understanding the missing data mechanism (MCAR, MAR, MNAR) is important for choosing an imputation strategy.
Why this is correct
The mechanism affects the validity of imputation methods.
- ✗
Listwise deletion (removing rows with missing values) is always safe and unbiased.
Why it's wrong here
Listwise deletion can bias results if data is not MCAR.
- ✗
Imputing missing values with the mean preserves the original variance.
Why it's wrong here
Mean imputation reduces variance and can distort relationships.
About these practice questions
Courseiva creates original exam-style practice questions with explanations and wrong-answer analysis. It does not publish real exam questions, exam dumps, or protected exam content. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on MLS-C01
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. Which TWO statements about handling missing data during EDA are correct? (Select TWO.)
medium- A.Dropping columns with >50% missing values is always recommended.
- B.Mean imputation preserves the variance of the original distribution.
- ✓ C.If data are missing completely at random (MCAR), listwise deletion yields unbiased estimates.
- D.Multiple imputation (MICE) is always the safest method regardless of missing data mechanism.
- ✓ E.Imputing with the median is more robust to outliers than imputing with the mean.
Why C: Options C and E are correct. Option C is correct because when data are Missing Completely at Random (MCAR), the missingness is independent of both observed and unobserved data, so listwise deletion (removing rows with missing values) does not introduce bias; the remaining sample is still a random subsample. Option E is correct because the median is not influenced by extreme values, making it a more robust imputation method compared to the mean, which can be skewed by outliers. Option A is incorrect because dropping columns with >50% missing values is not always recommended; it depends on the importance of the variable and the analysis goals. Option B is incorrect because mean imputation reduces the variance of the imputed variable, as it forces imputed values to the center. Option D is incorrect because Multiple Imputation by Chained Equations (MICE) is not always the safest; it assumes data are Missing at Random (MAR) and can be complex or inappropriate for other missingness mechanisms.
Last reviewed: Jun 20, 2026
This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.
Question Discussion
Share a tip, memory trick, or ask about the reasoning behind this question. Do not post real exam questions, leaked content, braindumps, or copyrighted exam material. Comments are moderated and may be removed without notice.
Sign in to join the discussion.