mediumMultiple Choice
MLA-C01 Practice Question: A machine learning team is building a fraud…
A machine learning team is building a fraud detection model. They have a dataset with a categorical feature 'merchant_id' that has over 10,000 unique values. Which feature engineering technique should they apply to 'merchant_id' to reduce dimensionality while retaining predictive power?
⚠ Common exam trap
MLA-C01 often tests the high-cardinality categorical scenario, and the trap is choosing one-hot encoding because it is the 'standard' approach, ignoring that 10,000+ categories make it computationally and statistically infeasible.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Target encoding
Target encoding replaces each merchant_id category with a statistic derived from the target variable (e.g., mean fraud rate per merchant), collapsing 10,000+ categories into a single numeric column while preserving the category's relationship to the label. This dramatically reduces dimensionality compared to one-hot encoding, which would create over 10,000 sparse columns. Because the encoded value carries predictive signal about fraud, target encoding retains useful information that label or ordinal encoding would lose.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
One-hot encoding
Why it's wrong here
One-hot encoding creates a separate binary column per merchant_id, expanding 10,000 categories into 10,000 sparse features — the opposite of dimensionality reduction. It is tempting because it avoids imposing false order on nominal categories, and would be the right choice for a low-cardinality nominal feature such as payment_type.
- ✓
Target encoding
Why this is correct
Target encoding replaces each merchant_id with a statistic of the target computed from that category, collapsing 10,000 levels into one numeric column. This retains each merchant's fraud signal while eliminating the dimensionality one-hot encoding would create.
- ✗
Ordinal encoding
Why it's wrong here
Ordinal encoding assigns each merchant_id an integer rank, implying a false ordering between unrelated merchants that tree and linear models may exploit incorrectly. It suits genuinely ordered categories such as 'low', 'medium', 'high', not 10,000 unordered identifiers needing dimensionality reduction.
- ✗
Label encoding
Why it's wrong here
Label encoding maps each merchant_id to an arbitrary integer, producing 10,000 distinct values with no dimensionality reduction and injecting spurious ordinal relationships. It is appropriate for binary or low-cardinality ordered categories, or for tree models where raw integer splits remain meaningful.
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.