DA0-002 Data Analysis Practice Question
A data analyst is preparing a dataset for a machine learning model to predict customer churn. The dataset contains missing values in several columns. The analyst wants to handle missing data in a way that preserves as much information as possible and avoids introducing bias. Which TWO of the following techniques are appropriate? (Choose two.)
⚠ Common exam trap
The trap here is assuming that any imputation method is fine, but simple methods like mean imputation or zero replacement can bias the model, while listwise deletion reduces sample size and may introduce bias.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
K-nearest neighbors (KNN) imputation
Multiple imputation by chained equations and KNN imputation are both advanced techniques that preserve data structure and reduce bias. MICE models each variable with missing data using other variables and creates multiple imputations to account for uncertainty. KNN imputation uses similarity between observations to estimate missing values. Both are superior to mean imputation, listwise deletion, or zero replacement, which either discard information or introduce bias.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Listwise deletion
Why it's wrong here
Listwise deletion removes any row with missing values, which can drastically reduce sample size and introduce bias if missingness is not completely random. It discards potentially valuable information from other columns in those rows. For churn prediction, losing customers with missing data could skew the model. Therefore, it does not preserve information and is not appropriate.
- ✓
K-nearest neighbors (KNN) imputation
Why this is correct
KNN imputation estimates missing values by finding the most similar complete observations and averaging their values. It leverages correlations between features and can preserve local data structure. This method is non-parametric and can handle both numerical and categorical data with appropriate distance metrics. It is suitable for churn prediction as it maintains information and reduces bias compared to simple imputation.
- ✗
Mean imputation for numerical features
Why it's wrong here
Mean imputation replaces missing values with the mean of the observed values. While simple, it reduces variance and can distort relationships with other variables, potentially introducing bias. It also ignores the uncertainty of missing data. For a churn prediction model, this could weaken the model's ability to detect patterns. Therefore, it is not an appropriate technique when the goal is to preserve information and avoid bias.
- ✗
Replacing missing values with zero
Why it's wrong here
Replacing missing values with zero assumes that missingness equates to a zero value, which is rarely true and can severely bias the model. For example, a missing age would become zero, an impossible value. This distorts distributions and relationships. It does not preserve information and introduces systematic error. Thus, it is not an appropriate technique for handling missing data in churn prediction.
- ✓
Multiple imputation by chained equations (MICE)
Why this is correct
Multiple imputation by chained equations creates several imputed datasets by modeling each feature with missing values as a function of other features. It accounts for the uncertainty of missing data and produces unbiased estimates under the missing at random assumption. This method preserves relationships and variability, making it suitable for churn prediction. It is a robust approach that avoids the bias of single imputation methods.
Go deeper
Related to this question
About these practice questions
This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.