DA0-002 Data Analysis Practice Question
A data scientist is preparing data for a K-means clustering algorithm. The dataset contains features measured in different units (e.g., income in dollars and age in years). Which preprocessing step is most critical before running K-means?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Standardize or normalize the features
K-means is sensitive to the scale of features because it uses Euclidean distance. Min-max normalization or standardization ensures all features contribute equally.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Remove outliers
Why it's wrong here
Outliers can affect clustering but scaling is more critical for distance calculation.
- ✗
Encode categorical variables
Why it's wrong here
While encoding is needed, scaling numerical features is more directly critical for K-means.
- ✓
Standardize or normalize the features
Why this is correct
Scaling ensures equal weighting; both min-max and Z-score are common.
- ✗
Perform feature selection
Why it's wrong here
Feature selection may help but scaling is essential for distance-based algorithms.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DA0-002 question from scratch — 986 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.