Courseiva
Data Analysis →hardMultiple Choice

DA0-002 Data Analysis Practice Question

A data scientist is analyzing a dataset with multiple features and wants to apply k-means clustering to segment customers. She chooses k = 4 based on the elbow method. During the iteration process, which of the following correctly describes a step in the k-means algorithm?

⚠ Common exam trap

The trap is mixing up k-means with other algorithms (e.g., PCA for initialization) or using median instead of mean. Candidates might also confuse the update step with using medians (as in k-medians).

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Assign each point to the nearest centroid based on Euclidean distance, then update centroids as the mean of points in each cluster.

In the k-means algorithm, after initializing centroids, each data point is assigned to the nearest centroid based on Euclidean distance, and then centroids are recomputed as the mean of all points in the cluster. This iterative process continues until convergence. Option D accurately describes this step.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Compute the covariance matrix and use principal components to initialize centroids.

    Why it's wrong here

    Principal component analysis is a dimensionality-reduction technique, not a k-means initialisation step; covariance and eigenvectors play no role in the algorithm. It is tempting because PCA is often applied before clustering, but k-means iterations only assign points to nearest centroids and recompute their means.

  • ✗

    Use hierarchical clustering to determine initial centroids.

    Why it's wrong here

    K-means initialises centroids by random selection or k-means++ seeding, not hierarchical clustering, which is a separate algorithm. It is tempting because hierarchical methods can suggest k, but within k-means iterations the steps are assignment of points to nearest centroids and recomputation of centroid means.

  • ✗

    Randomly assign centroids and then compute distances to the cluster medians.

    Why it's wrong here

    K-means recomputes centroids as the mean of assigned points, not medians; medians belong to k-medoids. The step is tempting because initialisation does place centroids randomly, but the iteration step named is wrong, so the option misstates the algorithm's actual update mechanism.

  • ✓

    Assign each point to the nearest centroid based on Euclidean distance, then update centroids as the mean of points in each cluster.

    Why this is correct

    K-means alternates two steps: each point is assigned to the cluster whose centroid is nearest by Euclidean distance, then every centroid is recomputed as the mean of its assigned points. Iteration repeats until assignments stabilise, satisfying the k=4 segmentation.

About these practice questions

One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.