Courseiva
Data Analysis →mediumMultiple Choice

DA0-002 Data Analysis Practice Question

A dataset has missing values in the 'age' column. The distribution of age is approximately normal with few outliers. Which imputation method is most appropriate?

⚠ Common exam trap

DA0-002 often tests the assumption that mean imputation is always appropriate, but the trap is failing to recognize that it is only suitable for continuous, normally distributed data without outliers; mode is for categorical, forward-fill for time-series.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Mean imputation

Mean imputation is appropriate when the data is approximately normally distributed and has few outliers, because the mean is a representative measure of central tendency for symmetric distributions. It preserves the overall mean of the variable and is simple to implement. Forward-fill and mode imputation are better for time-series or categorical data, respectively, and deleting rows can introduce bias and reduce sample size.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Mean imputation

    Why this is correct

    Age is approximately normal with few outliers, so the mean is a stable, representative estimate and preserves the sample mean. Mean imputation satisfies this distributional constraint better than median or mode, which suit skewed data or categorical fields respectively.

  • ✗

    Forward-fill

    Why it's wrong here

    Forward-fill propagates the last observed value downward, which suits time-series or ordered sequential data, not a roughly normal age distribution. For normally distributed numeric data with few outliers, mean imputation preserves the central tendency. Forward-fill would distort the distribution and introduce order-dependent bias.

  • ✗

    Delete all rows with missing data

    Why it's wrong here

    Deleting every row with a missing age discards otherwise valid records, reducing sample size and potentially biasing results if missingness is non-random. Mean imputation retains those rows. Listwise deletion is the right choice only when missing data are few and demonstrably missing completely at random.

  • ✗

    Mode imputation

    Why it's wrong here

    Mode imputation replaces missing values with the most frequent value, which suits categorical variables, not a continuous, approximately normal age column. For numeric normal data, mean imputation is appropriate. Mode would collapse many missing ages onto one repeated value, distorting variance and the distribution's shape.

About these practice questions

Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.