A data analyst is working with a dataset that includes a categorical variable 'product_category' with 50 unique values. The analyst wants to reduce dimensionality before clustering. Which technique should the analyst use?
Multiple correspondence analysis is specifically designed to reduce dimensionality of categorical data by transforming categories into a lower-dimensional numerical space. It captures associations between categories and can handle variables with many levels. This makes it ideal for the analyst's goal of reducing the 50 product categories before clustering. MCA preserves the categorical structure while enabling the use of distance-based clustering algorithms.
Why this answer
Multiple correspondence analysis is a dimensionality reduction technique tailored for categorical data. It converts categories into numerical dimensions that capture the underlying structure, making it suitable for clustering. Unlike one-hot encoding, it reduces rather than expands the feature space.
PCA and factor analysis are designed for continuous data, so they are not directly applicable to a categorical variable with many levels.
Exam trap
The trap here is assuming PCA can be applied to any data after encoding, but PCA on one-hot encoded data may not effectively reduce dimensionality and can lose interpretability.