SMOTE for Class Imbalance in AWS ML Engineer Associate
A company is building a fraud detection model on an imbalanced dataset (99% legitimate, 1% fraudulent). To improve recall on the minority class, they want to resample data. Which combination of techniques should they use?
⚠ Common exam trap
MLA-C01 often tests whether candidates apply resampling before the train/test split — the trap is forgetting that SMOTE or oversampling on the full dataset leaks synthetic information into the test set and invalidates evaluation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
SMOTE on training set only
SMOTE (Synthetic Minority Over-sampling Technique) generates synthetic minority-class samples by interpolating between existing minority instances. Applying SMOTE only to the training set prevents synthetic samples from leaking into the test set, which would inflate evaluation metrics and produce an overly optimistic model. This is the correct resampling approach for improving recall on an imbalanced fraud dataset.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
SMOTE on entire dataset before train/test split
Why it's wrong here
Applying SMOTE before the train/test split leaks synthetic minority samples into the test set, inflating recall estimates because near-duplicates of training rows appear in evaluation. It is tempting because SMOTE genuinely balances class distributions, and it would be correct if applied only to the training partition after splitting.
- ✗
Random oversampling of minority class before train/test split
Why it's wrong here
Duplicating minority rows before splitting places identical copies in both training and test sets, so the model is evaluated on data it has already seen. Random oversampling is a legitimate balancing technique, and it would be correct if confined to the training set after the split.
- ✗
Random undersampling of majority class
Why it's wrong here
Random undersampling alone discards majority-class examples, losing information and risking underfitting; the question asks for a combination. It would be correct paired with an oversampling method such as SMOTE, which together balance the classes while retaining data.
- ✓
SMOTE on training set only
Why this is correct
SMOTE synthesises new minority-class examples by interpolating between existing fraudulent cases, directly raising recall on the 1% class. Applying it only to the training set preserves the genuine 99:1 distribution in validation and test data, preventing the inflated performance estimates that leakage from resampled holdout data would cause.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.