DA0-002 Data Analysis Practice Question
A data team is preparing data for a clustering analysis. Which THREE of the following steps are commonly part of data cleaning?
⚠ Common exam trap
Test-takers frequently confuse data cleaning with data exploration or modeling — candidates see 'calculating the mean' and think it's part of preparation because it's a common early step, but it doesn't clean anything.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Removing duplicate records
Option A (Removing duplicate records) is correct because duplicate rows distort distance calculations in clustering, causing the same observation to be counted multiple times and biasing cluster centroids, so deduplication is a standard data-cleaning step. Option B (Imputing missing values) is correct because clustering algorithms such as k-means cannot handle nulls, so missing entries must be filled via mean/median/mode imputation, k-NN, or similar methods before analysis. Option E (Capping outliers at the 5th and 95th percentiles) is correct because winsorizing extreme values limits their disproportionate influence on distance metrics and centroid placement, which is a recognized cleaning technique. Option C (Calculating the mean) is not a cleaning step but a descriptive statistic or profiling operation, and Option D (Training a regression model) is a modeling task, not data preparation, so neither belongs to data cleaning.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Removing duplicate records
Why this is correct
Removing duplicate records eliminates redundant observations that would otherwise distort distance calculations, causing clustering algorithms to over-weight repeated points. This directly satisfies the stem's data-cleaning requirement by ensuring each entity contributes once, preventing artificial density concentrations that skew centroid placement and cluster assignment.
- ✓
Imputing missing values
Why this is correct
Imputing missing values replaces nulls with substituted estimates such as mean, median or model predictions. Clustering algorithms compute distances between records, so absent values would distort or break those calculations; imputation restores complete feature vectors, making it a standard data cleaning step.
- ✗
Calculating the mean
Why it's wrong here
Calculating the mean is a summary statistic, not a cleaning operation; cleaning handles missing values, duplicates, outliers, and inconsistent formats. It is tempting because means are used in imputation, but computing a mean alone neither detects nor corrects data quality problems before clustering.
- ✗
Training a regression model
Why it's wrong here
Training a regression model is a modelling step performed after cleaning, not part of it. It is tempting because both occur in the analytics pipeline, but cleaning addresses missing values, outliers, duplicates, and formatting; regression belongs to supervised learning, whereas clustering is unsupervised.
- ✓
Capping outliers at the 5th and 95th percentiles
Why this is correct
Capping outliers at the 5th and 95th percentiles winsorises extreme values, limiting their leverage on distance metrics. Because clustering is sensitive to scale and outliers, this transformation is a recognised cleaning step that keeps records while curbing distortion.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.