hardMultiple Select
MLA-C01 Practice Question: A data scientist is preparing a dataset for a…
A data scientist is preparing a dataset for a multi-class classification problem. The dataset contains a categorical feature with 50,000 unique values (high cardinality). The scientist wants to reduce dimensionality while preserving predictive information. Which TWO approaches are appropriate? (Choose 2)
⚠ Common exam trap
MLA-C01 often tests the misconception that one-hot or label encoding is suitable for high-cardinality features, when in fact they either explode dimensionality or impose false ordinality, making target and count encoding the correct choices.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Target encoding
Target encoding (Option A) is appropriate because it replaces each of the 50,000 categories with a statistic derived from the target variable (e.g., the mean target value per category), collapsing the feature into a single numeric column and thereby reducing dimensionality while retaining predictive signal about the classes. Count encoding (Option C) is also appropriate because it maps each category to its frequency in the dataset, producing one numeric column that captures useful information (rare vs. common categories) without expanding the feature space. One-hot encoding (Option B) is not suitable here since it would create roughly 50,000 sparse binary columns, which is the opposite of dimensionality reduction. Ordinal encoding based on alphabetical order (Option D) imposes an arbitrary, meaningless ordering on the categories and does not reduce dimensionality in a way that preserves predictive information. Label encoding (Option E) similarly assigns arbitrary integer IDs to categories, introducing a false ordinal relationship and offering no principled preservation of predictive information.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Target encoding
Why this is correct
Target encoding replaces each of the 50,000 categories with a statistic derived from the target, collapsing the feature to a single numeric column. This reduces dimensionality while retaining predictive signal, though regularisation or cross-fitting is needed to limit target leakage.
- ✗
One-hot encoding
Why it's wrong here
One-hot encoding expands the single feature into 50,000 sparse binary columns, increasing dimensionality rather than reducing it. It is tempting because it avoids false ordinal relationships and suits low-cardinality categoricals, but it would be correct for features with a handful of distinct values, not tens of thousands.
- ✓
Count encoding (frequency encoding)
Why this is correct
Count encoding replaces each category with its frequency in the dataset, producing one numeric column instead of 50,000 one-hot dimensions. It preserves information about category prevalence, which can correlate with the target, and handles unseen categories gracefully at inference.
- ✗
Ordinal encoding based on alphabetical order
Why it's wrong here
Alphabetical ordinal encoding imposes an arbitrary order on 50,000 categories, fabricating relationships the model will treat as meaningful and preserving full cardinality. It is tempting because it is compact and simple, but it would be correct for genuinely ordered features such as size bands, not unordered high-cardinality identifiers.
- ✗
Label encoding
Why it's wrong here
Label encoding assigns each of the 50,000 categories an arbitrary integer, implying ordinal relationships and retaining full cardinality without reducing dimensionality. It is tempting because it is memory-efficient and works for tree models on moderate cardinality, but it would be correct for ordered categories or low-cardinality features.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.