Courseiva
Data Analysis →mediumMultiple Choice

DA0-002 Data Analysis Practice Question

A data scientist is preparing data for a K-means clustering algorithm. The dataset contains features measured in different units (e.g., income in dollars and age in years). Which preprocessing step is most critical before running K-means?

⚠ Common exam trap

The trap is thinking that removing outliers or encoding categorical variables is the most critical step, but the question specifically highlights features in different units, which directly points to scaling. Candidates might also confuse feature selection with preprocessing necessity.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Standardize or normalize the features

K-means clustering uses Euclidean distance to measure similarity between data points. If features are on different scales (e.g., income in dollars vs. age in years), the feature with the larger range will dominate the distance calculation, leading to biased clusters. Standardizing (z-score normalization) or normalizing (min-max scaling) the features ensures that all features contribute equally to the distance metric, which is critical for K-means to produce meaningful clusters.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Remove outliers

    Why it's wrong here

    Removing outliers cleans extreme values but does not reconcile dollars against years, so distance calculations remain dominated by income magnitude. It is tempting because outliers distort centroids, yet scaling is the critical step for mixed units; outlier removal would be correct when extreme values skew cluster assignments after scaling.

  • ✗

    Encode categorical variables

    Why it's wrong here

    Encoding categorical variables addresses non-numeric data, whereas K-means requires numeric input; it does nothing about differing measurement scales. It is tempting because encoding is a common preprocessing step, but the stem's income-versus-age units demand feature scaling, which is the critical step here.

  • ✓

    Standardize or normalize the features

    Why this is correct

    K-means relies on Euclidean distance, so features in different units let larger-scale variables such as income dominate the distance calculation. Standardising or normalising puts every feature on a comparable scale, satisfying the requirement for meaningful cluster assignment.

  • ✗

    Perform feature selection

    Why it's wrong here

    Feature selection reduces dimensionality but leaves the remaining features on incompatible scales, so Euclidean distance still lets income dominate age. It is tempting because fewer features can aid clustering, but scaling is what equalises unit contributions; feature selection would be correct if irrelevant or redundant columns needed removing.

About these practice questions

Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.