Courseiva
Data Analysis →mediumMultiple Choice

DA0-002 Data Analysis Practice Question

A company’s marketing team wants to segment customers based on purchase history, demographics, and website behavior. The data includes both numeric and categorical variables. Which clustering algorithm is best suited for handling mixed data types?

⚠ Common exam trap

A common mix-up: candidates assume K-means or DBSCAN can handle mixed data by simply encoding categorical variables, but they overlook that Euclidean distance on encoded data distorts the geometry and fails to preserve the natural dissimilarity structure of categorical variables.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Hierarchical clustering with Gower distance

Hierarchical clustering with Gower distance is best suited for mixed data types because Gower distance computes a dissimilarity measure that handles both numeric and categorical variables by normalizing numeric differences and using a simple matching coefficient for categorical ones. This allows the algorithm to create a distance matrix that equally weights all variable types, making it ideal for segmenting customers with purchase history, demographics, and website behavior data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Hierarchical clustering with Gower distance

    Why this is correct

    Hierarchical clustering with Gower distance computes pairwise dissimilarity across numeric and categorical attributes simultaneously, so mixed-type records can be segmented without arbitrary encoding. It satisfies the stem's mixed data constraint, unlike k-means, which relies on Euclidean distance and requires numeric, scaled input.

  • ✗

    K-modes clustering

    Why it's wrong here

    K-modes handles only categorical variables via modal dissimilarity; it discards the numeric purchase-history and behavioural measures the scenario requires. It would be correct for segmenting purely categorical survey responses, but here it cannot process the numeric half of the dataset.

  • ✗

    DBSCAN with Euclidean distance

    Why it's wrong here

    Euclidean distance assumes continuous, ordered features, so it cannot quantify categorical values such as demographics or page categories without arbitrary encoding that distorts cluster structure. DBSCAN suits spatial, density-based outlier detection over purely numeric coordinates, not mixed-type customer segmentation.

  • ✗

    K-means clustering

    Why it's wrong here

    K-means computes means and Euclidean distances, which are undefined for categorical variables; encoding them creates false ordering and meaningless centroids. It would be correct for clustering purely numeric data such as continuous spend or session duration, not mixed demographic and behavioural attributes.

About these practice questions

Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.