Courseiva
Data Analysis →mediumMultiple Select

DA0-002 Data Analysis Practice Question

A data team is preparing data for a clustering analysis. Which THREE of the following steps are commonly part of data cleaning?

⚠ Common exam trap

Test-takers frequently confuse data cleaning with data exploration or modeling — candidates see 'calculating the mean' and think it's part of preparation because it's a common early step, but it doesn't clean anything.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Removing duplicate records

Option A (Removing duplicate records) is correct because duplicate rows distort distance calculations in clustering, causing the same observation to be counted multiple times and biasing cluster centroids, so deduplication is a standard data-cleaning step. Option B (Imputing missing values) is correct because clustering algorithms such as k-means cannot handle nulls, so missing entries must be filled via mean/median/mode imputation, k-NN, or similar methods before analysis. Option E (Capping outliers at the 5th and 95th percentiles) is correct because winsorizing extreme values limits their disproportionate influence on distance metrics and centroid placement, which is a recognized cleaning technique. Option C (Calculating the mean) is not a cleaning step but a descriptive statistic or profiling operation, and Option D (Training a regression model) is a modeling task, not data preparation, so neither belongs to data cleaning.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Removing duplicate records

    Why this is correct

    Removing duplicate records eliminates redundant observations that would otherwise distort distance calculations, causing clustering algorithms to over-weight repeated points. This directly satisfies the stem's data-cleaning requirement by ensuring each entity contributes once, preventing artificial density concentrations that skew centroid placement and cluster assignment.

  • ✓

    Imputing missing values

    Why this is correct

    Imputing missing values replaces nulls with substituted estimates such as mean, median or model predictions. Clustering algorithms compute distances between records, so absent values would distort or break those calculations; imputation restores complete feature vectors, making it a standard data cleaning step.

  • ✗

    Calculating the mean

    Why it's wrong here

    Calculating the mean is a summary statistic, not a cleaning operation; cleaning handles missing values, duplicates, outliers, and inconsistent formats. It is tempting because means are used in imputation, but computing a mean alone neither detects nor corrects data quality problems before clustering.

  • ✗

    Training a regression model

    Why it's wrong here

    Training a regression model is a modelling step performed after cleaning, not part of it. It is tempting because both occur in the analytics pipeline, but cleaning addresses missing values, outliers, duplicates, and formatting; regression belongs to supervised learning, whereas clustering is unsupervised.

  • ✓

    Capping outliers at the 5th and 95th percentiles

    Why this is correct

    Capping outliers at the 5th and 95th percentiles winsorises extreme values, limiting their leverage on distance metrics. Because clustering is sensitive to scale and outliers, this transformation is a recognised cleaning step that keeps records while curbing distortion.

About these practice questions

Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.