DA0-002 Data Analysis Practice Question
A data analyst is cleaning a dataset and finds that a numeric field has several missing values. The variable is normally distributed. Which imputation method is most appropriate?
⚠ Common exam trap
The trap is that candidates might choose median imputation thinking it's always more robust, but for a normal distribution, mean imputation is the standard; the question specifies normal distribution to guide you to the mean.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Mean imputation
For a normally distributed numeric variable, mean imputation is the most appropriate method because the mean is the central tendency that best represents the typical value in a symmetric distribution. Replacing missing values with the mean preserves the overall mean of the variable and is statistically sound when the data is normally distributed. Other methods like median or mode are better for skewed or categorical data, respectively.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Median imputation
Why it's wrong here
Median imputation suits skewed numeric data, where the mean is pulled by outliers; for a normally distributed variable the mean equals the median, so it discards the distribution's actual centre. It is tempting because median imputation is the standard robust choice when a numeric field is heavily skewed.
- ✓
Mean imputation
Why this is correct
For a normally distributed variable, the mean preserves the central tendency and keeps the overall distribution shape intact, unlike median or mode imputation. Since the data is symmetric, the mean is the best unbiased estimate for replacing missing numeric entries, satisfying the normality constraint in the stem.
- ✗
Mode imputation
Why it's wrong here
Mode imputation targets categorical frequency, not central tendency of a continuous normal distribution; for numeric data it can return a value that never occurs and distorts variance. It is tempting because mode is the standard choice for imputing missing categorical or nominal fields, where the most frequent category is the sensible fill.
- ✗
Forward-fill
Why it's wrong here
Forward-fill carries the last observed value forward, which distorts a normally distributed variable's mean and variance and ignores its distribution. It suits time-series or sequential data where the previous value persists, whereas a normal distribution calls for mean or regression imputation.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.