Courseiva
Data Analysis →mediumMultiple Choice

DA0-002 Data Analysis Practice Question

A retail company wants to identify customer segments based on purchase history and demographics. Which technique is most appropriate for this task?

⚠ Common exam trap

The trap is confusing supervised classification (logistic regression) with unsupervised clustering (K-means) — candidates who see 'segments' and think 'categories' pick logistic regression, but segmentation discovers groups rather than predicting predefined labels.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

K-means clustering

K-means clustering is an unsupervised learning algorithm that partitions observations into k groups based on similarity across multiple features, making it ideal for segmenting customers by purchase history and demographics. It doesn't require labeled outcomes, which matches the exploratory nature of customer segmentation. The algorithm iteratively assigns points to the nearest centroid and recalculates centroids until convergence.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Linear regression

    Why it's wrong here

    Linear regression predicts a continuous numeric outcome from input variables; it cannot assign observations to discrete segments. It is tempting because it is a familiar modelling technique, and it would be the correct choice when forecasting a continuous quantity such as next month's spend.

  • ✓

    K-means clustering

    Why this is correct

    K-means clustering partitions unlabelled data into k groups by minimising within-cluster variance, directly satisfying the requirement to identify customer segments from purchase history and demographics without predefined labels. Unlike classification, it discovers natural groupings, making it appropriate for exploratory segmentation of retail customers.

  • ✗

    Chi-square test

    Why it's wrong here

    The chi-square test assesses association between categorical variables; it produces a significance value, not customer groupings. It is tempting because it handles demographic categories, and it would be correct when testing whether two categorical variables, such as region and product preference, are independent.

  • ✗

    Logistic regression

    Why it's wrong here

    Logistic regression predicts a binary or categorical outcome from labelled training data; it cannot discover unknown segments without predefined classes. It is tempting because it classifies, and it would be correct when predicting a known label such as whether a customer will churn.

About these practice questions

One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.