easyMultiple Choice
MLA-C01 Practice Question: A data engineer wants to transform a categorical…
A data engineer wants to transform a categorical feature with 1,000 possible values into numerical features for a linear model. Which feature engineering technique is most appropriate for this high-cardinality feature?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Target encoding
Target encoding replaces each category with the mean of the target variable for that category, which handles high cardinality without exploding dimensionality. One-hot encoding creates 1,000 columns, which is problematic for linear models.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
One-hot encoding
Why it's wrong here
One-hot encoding creates 1,000 sparse binary columns, causing the curse of dimensionality and unstable coefficients in a linear model. It is tempting because one-hot encoding is the standard correct technique when a categorical feature has low cardinality, such as fewer than ten distinct values.
- ✗
Ordinal encoding
Why it's wrong here
Ordinal encoding imposes an arbitrary numeric order on 1,000 unordered categories, so the linear model treats the assigned integers as meaningful magnitudes. It is tempting because ordinal encoding is correct when categories have a genuine inherent ranking, such as severity levels low, medium and high.
- ✓
Target encoding
Why this is correct
Target encoding replaces each of the 1,000 categories with a statistic derived from the target, producing a single numerical column rather than 1,000 sparse dummy columns. This directly addresses the high-cardinality constraint while remaining suitable for a linear model.
- ✗
Label encoding
Why it's wrong here
Label encoding assigns each of the 1,000 categories an arbitrary integer, which a linear model interprets as a continuous numeric relationship, distorting predictions. It is tempting because label encoding is correct for tree-based models, which split on individual encoded values without assuming ordinal distance.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.