DA0-002 Data Analysis Practice Question
A dataset has missing values in the 'age' column. The distribution of age is approximately normal with few outliers. Which imputation method is most appropriate?
⚠ Common exam trap
DA0-002 often tests the assumption that mean imputation is always appropriate, but the trap is failing to recognize that it is only suitable for continuous, normally distributed data without outliers; mode is for categorical, forward-fill for time-series.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Mean imputation
Mean imputation is appropriate when the data is approximately normally distributed and has few outliers, because the mean is a representative measure of central tendency for symmetric distributions. It preserves the overall mean of the variable and is simple to implement. Forward-fill and mode imputation are better for time-series or categorical data, respectively, and deleting rows can introduce bias and reduce sample size.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Mean imputation
Why this is correct
Age is approximately normal with few outliers, so the mean is a stable, representative estimate and preserves the sample mean. Mean imputation satisfies this distributional constraint better than median or mode, which suit skewed data or categorical fields respectively.
- ✗
Forward-fill
Why it's wrong here
Forward-fill propagates the last observed value downward, which suits time-series or ordered sequential data, not a roughly normal age distribution. For normally distributed numeric data with few outliers, mean imputation preserves the central tendency. Forward-fill would distort the distribution and introduce order-dependent bias.
- ✗
Delete all rows with missing data
Why it's wrong here
Deleting every row with a missing age discards otherwise valid records, reducing sample size and potentially biasing results if missingness is non-random. Mean imputation retains those rows. Listwise deletion is the right choice only when missing data are few and demonstrably missing completely at random.
- ✗
Mode imputation
Why it's wrong here
Mode imputation replaces missing values with the most frequent value, which suits categorical variables, not a continuous, approximately normal age column. For numeric normal data, mean imputation is appropriate. Mode would collapse many missing ages onto one repeated value, distorting variance and the distribution's shape.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.