DA0-002 Data Analysis Practice Question
A marketing team wants to segment customers into groups based on purchasing behavior without prior labels. Which algorithm should the data analyst use?
⚠ Common exam trap
A common mix-up: candidates confuse unsupervised clustering (K-means) with supervised classification (K-nearest neighbors) because both involve 'K' and grouping, but KNN requires labeled data and predicts labels, while K-means discovers inherent structures without labels.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
K-means clustering
K-means clustering is the correct choice because it is an unsupervised learning algorithm that groups unlabeled data into clusters based on feature similarity. Since the marketing team has no prior labels for customer segments, K-means can partition customers by purchasing behavior patterns, such as frequency and monetary value, without needing predefined categories.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
K-means clustering
Why this is correct
K-means clustering partitions unlabelled data into k groups by minimising within-cluster variance, directly satisfying the stem's requirement for segmentation without prior labels. Unlike supervised methods such as logistic regression or decision trees, it needs no target variable, making it the appropriate choice for discovering behavioural customer segments.
- ✗
K-nearest neighbors
Why it's wrong here
K-nearest neighbours is supervised classification or regression; it requires labelled training examples to assign a class, so it cannot discover groups without prior labels. It is tempting because it groups similar points, but it would only be correct once segment labels already exist for training.
- ✗
Linear regression
Why it's wrong here
Linear regression predicts a continuous numeric target from labelled input features, so it cannot assign customers to unlabelled groups. It is tempting because it is a core supervised technique, and it would be correct if the team had a known numeric outcome such as predicted spend to model.
- ✗
Decision tree
Why it's wrong here
A decision tree is a supervised algorithm that splits labelled data to predict a known target class or value, so it cannot segment customers without prior labels. It is tempting because its structure resembles clustering, but it would be correct only when a predefined outcome label is available for training.
About these practice questions
One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.