Courseiva
easyMultiple Choice

MLA-C01 Practice Question: A data engineer wants to transform a categorical…

A data engineer wants to transform a categorical feature with 1,000 possible values into numerical features for a linear model. Which feature engineering technique is most appropriate for this high-cardinality feature?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Target encoding

Target encoding replaces each category with the mean of the target variable for that category, which handles high cardinality without exploding dimensionality. One-hot encoding creates 1,000 columns, which is problematic for linear models.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    One-hot encoding

    Why it's wrong here

    One-hot encoding creates 1,000 sparse binary columns, causing the curse of dimensionality and unstable coefficients in a linear model. It is tempting because one-hot encoding is the standard correct technique when a categorical feature has low cardinality, such as fewer than ten distinct values.

  • ✗

    Ordinal encoding

    Why it's wrong here

    Ordinal encoding imposes an arbitrary numeric order on 1,000 unordered categories, so the linear model treats the assigned integers as meaningful magnitudes. It is tempting because ordinal encoding is correct when categories have a genuine inherent ranking, such as severity levels low, medium and high.

  • ✓

    Target encoding

    Why this is correct

    Target encoding replaces each of the 1,000 categories with a statistic derived from the target, producing a single numerical column rather than 1,000 sparse dummy columns. This directly addresses the high-cardinality constraint while remaining suitable for a linear model.

  • ✗

    Label encoding

    Why it's wrong here

    Label encoding assigns each of the 1,000 categories an arbitrary integer, which a linear model interprets as a continuous numeric relationship, distorting predictions. It is tempting because label encoding is correct for tree-based models, which split on individual encoded values without assuming ordinal distance.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.