Courseiva
hardMultiple Choice

MLA-C01 Practice Question: A machine learning practitioner is building a…

A machine learning practitioner is building a binary classifier with severe class imbalance (1:1000). They want to use SMOTE for oversampling. What is a potential drawback of applying SMOTE on the entire dataset before splitting into training and test sets?

⚠ Common exam trap

MLA-C01 often tests the misconception that resampling is a harmless preprocessing step, when in fact applying SMOTE before the train/test split is a classic data-leakage trap that inflates validation metrics.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

It causes data leakage, making validation metrics overly optimistic

When SMOTE is applied before the train/test split, synthetic minority samples are generated using information from the entire dataset, including the observations that will later become the test set. This means the test set is no longer independent of the training data, so the model has effectively 'seen' information from the test distribution. The result is data leakage that inflates validation metrics (precision, recall, F1) and produces a model that appears far better than it will perform in production.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    SMOTE increases the risk of overfitting to the minority class

    Why it's wrong here

    Overfitting to the minority class can occur whenever SMOTE is applied, regardless of when the split happens. It is tempting because oversampling does inflate minority influence, but the pre-split drawback is that synthetic points derived from test-set neighbours leak test information into training.

  • ✗

    SMOTE cannot be applied to categorical features

    Why it's wrong here

    SMOTE's k-nearest-neighbour interpolation operates on numeric feature space, so categorical handling is a separate preprocessing concern, not the pre-split issue. It is tempting because encoding matters, but the question targets leakage caused by synthesising before the train-test split.

  • ✓

    It causes data leakage, making validation metrics overly optimistic

    Why this is correct

    SMOTE synthesises new minority-class samples from nearest neighbours. Applying it before splitting lets synthetic points derived from test-set instances appear in training, so the model effectively sees test data. Validation metrics then look overly optimistic, overstating real-world performance.

  • ✗

    SMOTE generates synthetic samples that may not be realistic

    Why it's wrong here

    Synthetic minority samples being unrealistic is a general property of SMOTE interpolation, not a consequence of ordering it before the split. It is tempting because unrealistic samples do harm models, but the question asks specifically about pre-split application, whose defect is leakage.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.