Courseiva
hardMultiple Select

MLA-C01 Practice Question: A data scientist is preparing a dataset for a…

A data scientist is preparing a dataset for a multi-class classification problem. The dataset contains a categorical feature with 50,000 unique values (high cardinality). The scientist wants to reduce dimensionality while preserving predictive information. Which TWO approaches are appropriate? (Choose 2)

⚠ Common exam trap

MLA-C01 often tests the misconception that one-hot or label encoding is suitable for high-cardinality features, when in fact they either explode dimensionality or impose false ordinality, making target and count encoding the correct choices.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Target encoding

Target encoding (Option A) is appropriate because it replaces each of the 50,000 categories with a statistic derived from the target variable (e.g., the mean target value per category), collapsing the feature into a single numeric column and thereby reducing dimensionality while retaining predictive signal about the classes. Count encoding (Option C) is also appropriate because it maps each category to its frequency in the dataset, producing one numeric column that captures useful information (rare vs. common categories) without expanding the feature space. One-hot encoding (Option B) is not suitable here since it would create roughly 50,000 sparse binary columns, which is the opposite of dimensionality reduction. Ordinal encoding based on alphabetical order (Option D) imposes an arbitrary, meaningless ordering on the categories and does not reduce dimensionality in a way that preserves predictive information. Label encoding (Option E) similarly assigns arbitrary integer IDs to categories, introducing a false ordinal relationship and offering no principled preservation of predictive information.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Target encoding

    Why this is correct

    Target encoding replaces each of the 50,000 categories with a statistic derived from the target, collapsing the feature to a single numeric column. This reduces dimensionality while retaining predictive signal, though regularisation or cross-fitting is needed to limit target leakage.

  • ✗

    One-hot encoding

    Why it's wrong here

    One-hot encoding expands the single feature into 50,000 sparse binary columns, increasing dimensionality rather than reducing it. It is tempting because it avoids false ordinal relationships and suits low-cardinality categoricals, but it would be correct for features with a handful of distinct values, not tens of thousands.

  • ✓

    Count encoding (frequency encoding)

    Why this is correct

    Count encoding replaces each category with its frequency in the dataset, producing one numeric column instead of 50,000 one-hot dimensions. It preserves information about category prevalence, which can correlate with the target, and handles unseen categories gracefully at inference.

  • ✗

    Ordinal encoding based on alphabetical order

    Why it's wrong here

    Alphabetical ordinal encoding imposes an arbitrary order on 50,000 categories, fabricating relationships the model will treat as meaningful and preserving full cardinality. It is tempting because it is compact and simple, but it would be correct for genuinely ordered features such as size bands, not unordered high-cardinality identifiers.

  • ✗

    Label encoding

    Why it's wrong here

    Label encoding assigns each of the 50,000 categories an arbitrary integer, implying ordinal relationships and retaining full cardinality without reducing dimensionality. It is tempting because it is memory-efficient and works for tree models on moderate cardinality, but it would be correct for ordered categories or low-cardinality features.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.