hardMultiple Choice
MLA-C01 Practice Question: Using Amazon SageMaker Data Wrangler to prepare a…
A company is using Amazon SageMaker Data Wrangler to prepare a dataset with over 200 features. The dataset includes a categorical feature with more than 10,000 unique values (high cardinality). The ML engineer wants to transform this feature into a numeric representation suitable for a linear model without increasing dimensionality too much. Which built-in transform in Data Wrangler should the engineer use?
⚠ Common exam trap
AWS often tests the misconception that high-cardinality categorical features must be one-hot encoded, but the trap here is that one-hot encoding explodes dimensionality, while target encoding provides a compact, target-informed numeric representation suitable for linear models.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Target encoding
Target encoding is the correct choice because it replaces each category with the mean of the target variable for that category, producing a single numeric column that captures predictive signal without exploding dimensionality. This is ideal for high-cardinality features (e.g., >10,000 unique values) when used with linear models, as it avoids the sparsity and multicollinearity issues of one-hot encoding while retaining correlation with the target.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Ordinal encoding
Why it's wrong here
Ordinal encoding assigns arbitrary integer ranks to 10,000 unordered categories, implying a false numeric ordering that a linear model would treat as meaningful magnitude. It suits genuinely ordered variables such as low/medium/high. For high-cardinality nominal features, target or frequency encoding preserves cardinality without fabricating order.
- ✓
Target encoding
Why this is correct
Target encoding replaces each category with the mean of the target variable for that category, collapsing 10,000+ unique values into a single numeric column. This satisfies the stem's constraint of avoiding dimensionality growth, unlike one-hot encoding, while producing the numeric representation a linear model requires.
- ✗
One-hot encoding
Why it's wrong here
One-hot encoding creates a separate binary column per distinct category, so 10,000 unique values expand into 10,000 sparse dimensions, directly violating the requirement to avoid increasing dimensionality. It works well for low-cardinality nominal features, typically under a few dozen categories, where each level can be represented independently.
- ✗
Frequency encoding
Why it's wrong here
Frequency encoding replaces each category with its occurrence count, which is numeric and compact, but it is not a built-in SageMaker Data Wrangler transform; the question asks specifically for a built-in option. It would suit high-cardinality features when implemented manually or via custom code, not through the native transform list.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.