Courseiva
Data Analysis →easyMultiple Choice

DA0-002 Data Analysis Practice Question

A marketing analyst wants to segment customers based on their purchase history, including total spent, number of transactions, and average order value. The analyst runs k-means clustering with k=5 on the raw data but notices that the cluster assignments change significantly every time the algorithm is executed. What should the analyst do first to obtain consistent and meaningful clusters?

⚠ Common exam trap

Watch out — candidates often think the instability is due to the choice of k or the algorithm itself, rather than recognizing that k-means is sensitive to feature scaling and random initialization, which are the first things to address for consistency.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Normalize the features and set a fixed random seed for the initial centroids.

The instability in cluster assignments is caused by the algorithm's sensitivity to the scale of features and the random initialization of centroids. Normalizing the features ensures that each variable contributes equally to the distance calculations, while setting a fixed random seed makes the initial centroid selection deterministic, leading to reproducible results.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Normalize the features and set a fixed random seed for the initial centroids.

    Why this is correct

    Features on different scales distort Euclidean distance, so k-means centroids shift with each random initialisation. Normalising the features and fixing the random seed stabilises initial centroid selection, satisfying the requirement for consistent, meaningful cluster assignments across runs.

  • ✗

    Switch to hierarchical clustering, which does not require specifying k.

    Why it's wrong here

    Hierarchical clustering also depends on distance computations, so unscaled features with differing ranges still distort the dendrogram and yield unstable groupings. It suits small datasets where the cluster count is unknown; here the instability stems from unstandardised variables, which scaling addresses first.

  • ✗

    Increase the number of clusters to k=10 to capture more detail.

    Why it's wrong here

    Adding clusters does not address the instability, which stems from unscaled features and random initialisation causing different centroids each run. Increasing k is tempting to capture finer segments, but standardising the three variables and setting a fixed random seed first yields reproducible, meaningful clusters.

  • ✗

    Use principal component analysis (PCA) to reduce the number of features to two.

    Why it's wrong here

    PCA reduces dimensionality but retains the unscaled variance structure, so total spent dominates the components and cluster assignments remain unstable. It suits correlated high-dimensional data for visualisation or noise reduction; the immediate cause here is unstandardised features, which scaling fixes.

About these practice questions

One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.