Courseiva
Data Analysis →mediumMultiple Choice

DA0-002 Data Analysis Practice Question

A marketing team wants to segment customers into groups based on purchasing behavior without prior labels. Which algorithm should the data analyst use?

⚠ Common exam trap

A common mix-up: candidates confuse unsupervised clustering (K-means) with supervised classification (K-nearest neighbors) because both involve 'K' and grouping, but KNN requires labeled data and predicts labels, while K-means discovers inherent structures without labels.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

K-means clustering

K-means clustering is the correct choice because it is an unsupervised learning algorithm that groups unlabeled data into clusters based on feature similarity. Since the marketing team has no prior labels for customer segments, K-means can partition customers by purchasing behavior patterns, such as frequency and monetary value, without needing predefined categories.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    K-means clustering

    Why this is correct

    K-means clustering partitions unlabelled data into k groups by minimising within-cluster variance, directly satisfying the stem's requirement for segmentation without prior labels. Unlike supervised methods such as logistic regression or decision trees, it needs no target variable, making it the appropriate choice for discovering behavioural customer segments.

  • ✗

    K-nearest neighbors

    Why it's wrong here

    K-nearest neighbours is supervised classification or regression; it requires labelled training examples to assign a class, so it cannot discover groups without prior labels. It is tempting because it groups similar points, but it would only be correct once segment labels already exist for training.

  • ✗

    Linear regression

    Why it's wrong here

    Linear regression predicts a continuous numeric target from labelled input features, so it cannot assign customers to unlabelled groups. It is tempting because it is a core supervised technique, and it would be correct if the team had a known numeric outcome such as predicted spend to model.

  • ✗

    Decision tree

    Why it's wrong here

    A decision tree is a supervised algorithm that splits labelled data to predict a known target class or value, so it cannot segment customers without prior labels. It is tempting because its structure resembles clustering, but it would be correct only when a predefined outcome label is available for training.

About these practice questions

One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.