hardMultiple Choice
MLA-C01 Practice Question: A machine learning practitioner is building a…
A machine learning practitioner is building a binary classifier with severe class imbalance (1:1000). They want to use SMOTE for oversampling. What is a potential drawback of applying SMOTE on the entire dataset before splitting into training and test sets?
⚠ Common exam trap
MLA-C01 often tests the misconception that resampling is a harmless preprocessing step, when in fact applying SMOTE before the train/test split is a classic data-leakage trap that inflates validation metrics.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
It causes data leakage, making validation metrics overly optimistic
When SMOTE is applied before the train/test split, synthetic minority samples are generated using information from the entire dataset, including the observations that will later become the test set. This means the test set is no longer independent of the training data, so the model has effectively 'seen' information from the test distribution. The result is data leakage that inflates validation metrics (precision, recall, F1) and produces a model that appears far better than it will perform in production.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
SMOTE increases the risk of overfitting to the minority class
Why it's wrong here
Overfitting to the minority class can occur whenever SMOTE is applied, regardless of when the split happens. It is tempting because oversampling does inflate minority influence, but the pre-split drawback is that synthetic points derived from test-set neighbours leak test information into training.
- ✗
SMOTE cannot be applied to categorical features
Why it's wrong here
SMOTE's k-nearest-neighbour interpolation operates on numeric feature space, so categorical handling is a separate preprocessing concern, not the pre-split issue. It is tempting because encoding matters, but the question targets leakage caused by synthesising before the train-test split.
- ✓
It causes data leakage, making validation metrics overly optimistic
Why this is correct
SMOTE synthesises new minority-class samples from nearest neighbours. Applying it before splitting lets synthetic points derived from test-set instances appear in training, so the model effectively sees test data. Validation metrics then look overly optimistic, overstating real-world performance.
- ✗
SMOTE generates synthetic samples that may not be realistic
Why it's wrong here
Synthetic minority samples being unrealistic is a general property of SMOTE interpolation, not a consequence of ordering it before the split. It is tempting because unrealistic samples do harm models, but the question asks specifically about pre-split application, whose defect is leakage.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.