DA0-002 Data Analysis Practice Question
A retail company wants to identify customer segments based on purchase history and demographics. Which technique is most appropriate for this task?
⚠ Common exam trap
The trap is confusing supervised classification (logistic regression) with unsupervised clustering (K-means) — candidates who see 'segments' and think 'categories' pick logistic regression, but segmentation discovers groups rather than predicting predefined labels.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
K-means clustering
K-means clustering is an unsupervised learning algorithm that partitions observations into k groups based on similarity across multiple features, making it ideal for segmenting customers by purchase history and demographics. It doesn't require labeled outcomes, which matches the exploratory nature of customer segmentation. The algorithm iteratively assigns points to the nearest centroid and recalculates centroids until convergence.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Linear regression
Why it's wrong here
Linear regression predicts a continuous numeric outcome from input variables; it cannot assign observations to discrete segments. It is tempting because it is a familiar modelling technique, and it would be the correct choice when forecasting a continuous quantity such as next month's spend.
- ✓
K-means clustering
Why this is correct
K-means clustering partitions unlabelled data into k groups by minimising within-cluster variance, directly satisfying the requirement to identify customer segments from purchase history and demographics without predefined labels. Unlike classification, it discovers natural groupings, making it appropriate for exploratory segmentation of retail customers.
- ✗
Chi-square test
Why it's wrong here
The chi-square test assesses association between categorical variables; it produces a significance value, not customer groupings. It is tempting because it handles demographic categories, and it would be correct when testing whether two categorical variables, such as region and product preference, are independent.
- ✗
Logistic regression
Why it's wrong here
Logistic regression predicts a binary or categorical outcome from labelled training data; it cannot discover unknown segments without predefined classes. It is tempting because it classifies, and it would be correct when predicting a known label such as whether a customer will churn.
About these practice questions
One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.