Courseiva
ML Model Development →mediumMultiple Choice

MLA-C01 ML Model Development Practice Question

A machine learning engineer is preparing a training dataset in Amazon SageMaker for a binary classification model. The dataset is stored as a single CSV file in Amazon S3 and contains 12 categorical features with high cardinality (thousands of unique values each). The engineer wants to avoid the curse of dimensionality and reduce training time while preserving predictive power. Which preprocessing approach should be used with the SageMaker built-in XGBoost algorithm?

⚠ Common exam trap

The trap here is assuming that one-hot encoding is always the default for categorical features, overlooking the dimensionality explosion with high-cardinality data.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use target encoding (mean encoding) for the categorical features, computed within cross-validation folds to prevent leakage.

High-cardinality categorical features require an encoding that avoids exponential feature expansion. Target encoding summarizes each category by its relationship to the target, reducing dimensionality and training time. Performing it within cross-validation folds prevents data leakage, ensuring that validation metrics remain honest. This approach is compatible with the SageMaker built-in XGBoost algorithm, which accepts numerical features and can benefit from the reduced feature space.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use target encoding (mean encoding) for the categorical features, computed within cross-validation folds to prevent leakage.

    Why this is correct

    Target encoding replaces each category with a statistic (e.g., mean of the target) and is well-suited for high-cardinality features. Computing it within cross-validation folds prevents target leakage, which would otherwise inflate validation performance. This reduces dimensionality and training time while retaining predictive signal, directly addressing the engineer's goals without creating thousands of sparse columns.

  • ✗

    Convert the categorical features to integer codes using LabelEncoder and pass them directly to XGBoost.

    Why it's wrong here

    Label encoding assigns arbitrary integers to categories, implying an ordinal relationship that does not exist. XGBoost may split on these integers and create nonsensical thresholds, harming model performance. While it is simple and fast, it fails to capture category semantics and can mislead the tree splits, making it a poor choice for high-cardinality nominal features.

  • ✗

    Hash the categorical features into a fixed number of buckets using a hashing trick, then one-hot encode the hashed values.

    Why it's wrong here

    The hashing trick reduces dimensionality but introduces collisions that can degrade performance, and one-hot encoding the hashed values still creates many sparse columns. It does not inherently preserve predictive power and adds complexity. For this scenario, target encoding is a more direct and interpretable solution that avoids collisions and excessive dimensionality.

  • ✗

    Apply one-hot encoding to all categorical features before training.

    Why it's wrong here

    One-hot encoding would create tens of thousands of sparse binary columns, dramatically increasing dimensionality and memory usage. For high-cardinality categorical features, this leads to the curse of dimensionality, slower training, and potential overfitting. The scenario explicitly asks to avoid this outcome, so one-hot encoding is inappropriate despite being a common default for low-cardinality features.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.