Courseiva
mediumMultiple Choice

MLA-C01 Practice Question: A machine learning team is building a fraud…

A machine learning team is building a fraud detection model. They have a dataset with a categorical feature 'merchant_id' that has over 10,000 unique values. Which feature engineering technique should they apply to 'merchant_id' to reduce dimensionality while retaining predictive power?

⚠ Common exam trap

MLA-C01 often tests the high-cardinality categorical scenario, and the trap is choosing one-hot encoding because it is the 'standard' approach, ignoring that 10,000+ categories make it computationally and statistically infeasible.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Target encoding

Target encoding replaces each merchant_id category with a statistic derived from the target variable (e.g., mean fraud rate per merchant), collapsing 10,000+ categories into a single numeric column while preserving the category's relationship to the label. This dramatically reduces dimensionality compared to one-hot encoding, which would create over 10,000 sparse columns. Because the encoded value carries predictive signal about fraud, target encoding retains useful information that label or ordinal encoding would lose.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    One-hot encoding

    Why it's wrong here

    One-hot encoding creates a separate binary column per merchant_id, expanding 10,000 categories into 10,000 sparse features — the opposite of dimensionality reduction. It is tempting because it avoids imposing false order on nominal categories, and would be the right choice for a low-cardinality nominal feature such as payment_type.

  • ✓

    Target encoding

    Why this is correct

    Target encoding replaces each merchant_id with a statistic of the target computed from that category, collapsing 10,000 levels into one numeric column. This retains each merchant's fraud signal while eliminating the dimensionality one-hot encoding would create.

  • ✗

    Ordinal encoding

    Why it's wrong here

    Ordinal encoding assigns each merchant_id an integer rank, implying a false ordering between unrelated merchants that tree and linear models may exploit incorrectly. It suits genuinely ordered categories such as 'low', 'medium', 'high', not 10,000 unordered identifiers needing dimensionality reduction.

  • ✗

    Label encoding

    Why it's wrong here

    Label encoding maps each merchant_id to an arbitrary integer, producing 10,000 distinct values with no dimensionality reduction and injecting spurious ordinal relationships. It is appropriate for binary or low-cardinality ordered categories, or for tree models where raw integer splits remain meaningful.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.