Courseiva
hardMultiple Choice

MLA-C01 Practice Question: Building a fraud detection model on credit card…

A company is building a fraud detection model on credit card transactions. The dataset contains a column 'merchant_id' with 50,000 unique values, many with low frequency. The team wants to avoid overfitting while preserving predictive signal. Which feature engineering approach is most appropriate?

⚠ Common exam trap

MLA-C01 often tests the trade-off between preserving signal and avoiding overfitting with high-cardinality features, tricking candidates into choosing one-hot encoding or dropping the feature instead of using target encoding with smoothing.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply target encoding with smoothing based on global mean

Target encoding with smoothing replaces each category with a blend of the category's mean target value and the global mean, weighted by frequency. This preserves predictive signal from high-frequency merchants while reducing overfitting for low-frequency ones by shrinking their encoded values toward the global mean. It is a standard technique for high-cardinality categorical features.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Drop the 'merchant_id' column to avoid overfitting

    Why it's wrong here

    Dropping merchant_id discards the predictive signal it carries, since merchant behaviour is a strong fraud indicator; overfitting should instead be handled by frequency encoding, grouping rare merchants, or regularisation. Dropping is tempting when cardinality is extreme, and would suit an identifier with no predictive value.

  • ✗

    Apply label encoding and treat it as a numeric feature

    Why it's wrong here

    Label encoding imposes an arbitrary ordinal relationship on merchant IDs, so the model treats ID 40,000 as numerically greater than ID 200, fabricating signal that does not exist. It suits genuinely ordered categories such as severity ratings, not nominal identifiers.

  • ✓

    Apply target encoding with smoothing based on global mean

    Why this is correct

    Target encoding replaces each of the 50,000 merchant_id values with a statistic, avoiding the high dimensionality of one-hot encoding. Smoothing blends each category's mean toward the global mean, so low-frequency merchants are not overfitted while their predictive signal is retained.

  • ✗

    One-hot encode the 'merchant_id' column

    Why it's wrong here

    One-hot encoding creates 50,000 sparse binary columns, so rare merchants get their own weights and the model memorises them, worsening overfitting. It suits low-cardinality categoricals where every level has ample observations, not this high-cardinality column.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.