DA0-002 Data Analysis Practice Question
A data analyst is preparing a dataset for analysis and needs to address data quality issues. Which TWO of the following are common data cleaning tasks?
⚠ Common exam trap
The trap is mixing analysis activities (hypothesis testing, regression, correlation) with cleaning activities — candidates must distinguish preprocessing/data-quality remediation from downstream statistical modeling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Imputing missing values
Option B (Imputing missing values) is correct because missing data is a classic data quality problem, and imputation—filling gaps using methods like mean, median, mode, or model-based estimates—is a standard data cleaning step that makes the dataset complete and usable for analysis. Option E (Deduplicating records) is correct because duplicate rows or records distort counts, aggregates, and model results, so identifying and removing or merging duplicates is a core data cleaning task. The unmarked options do not belong because hypothesis testing (A), building a regression model (C), and calculating correlation coefficients (D) are all downstream analytical or statistical modeling activities performed on already-cleaned data, not cleaning operations themselves.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Performing hypothesis testing
Why it's wrong here
Hypothesis testing is inferential statistics used to assess evidence about a population parameter, not a data cleaning activity. It belongs in the analysis phase once the dataset is clean, for example when comparing group means or validating whether an observed effect is statistically significant.
- ✓
Imputing missing values
Why this is correct
Imputing missing values replaces nulls with substituted estimates such as mean, median or model-predicted figures, satisfying the stem's data quality remediation goal. It preserves row counts for analysis rather than discarding incomplete records, directly addressing missingness as a cleaning task.
- ✗
Building a regression model
Why it's wrong here
Regression modelling is an analysis technique for predicting a continuous target, not a cleaning step; it is performed after cleaning. It would be the right choice when the analyst needs to quantify relationships between variables or forecast values, not when remedying missing, duplicated or malformed records.
- ✗
Calculating correlation coefficients
Why it's wrong here
Correlation coefficients measure the strength of association between numeric variables during exploratory analysis, not data cleaning. Calculating them is appropriate when investigating relationships or selecting features, whereas cleaning addresses missing values, duplicates, outliers and inconsistent formats before such analysis.
- ✓
Deduplicating records
Why this is correct
Deduplicating records identifies and removes or merges repeated rows caused by duplicate ingestion or matching errors, satisfying the data quality requirement. It ensures each entity appears once, preventing inflated counts and skewed aggregates during analysis.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.