MLA-C01 Data Preparation for Machine Learning Practice Question
A data scientist is preparing a dataset for training a binary classification model. The dataset has 100,000 rows and 50 features. The target variable is imbalanced, with only 5% positive cases. Which technique should the data scientist apply to address the class imbalance BEFORE training?
⚠ Common exam trap
AWS often tests whether candidates confuse data preprocessing techniques (scaling, encoding, dimensionality reduction) with methods that directly modify the class distribution, leading them to pick a plausible but irrelevant option like PCA or scaling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Random oversampling of the minority class
Random oversampling of the minority class (Option B) directly addresses the class imbalance by duplicating examples from the positive class until the class distribution is more balanced. This prevents the binary classification model from being biased toward the majority class, which is critical when only 5% of the 100,000 rows are positive cases. Oversampling is applied before training to ensure the model sees sufficient minority examples during learning.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Principal Component Analysis (PCA) dimensionality reduction
Why it's wrong here
PCA reduces feature dimensionality by projecting onto principal components; it leaves the 5% positive prevalence untouched, so imbalance persists. PCA is appropriate for cutting redundant correlated features or speeding training, not for correcting skewed class distributions.
- ✓
Random oversampling of the minority class
Why this is correct
Random oversampling duplicates minority-class rows, raising the 5% positive rate toward parity so the algorithm no longer biases toward the majority class. It satisfies the pre-training constraint by rebalancing class distribution before the model sees the data, unlike threshold tuning applied afterwards.
- ✗
Standard scaling of numerical features
Why it's wrong here
Standard scaling normalises feature magnitudes; it leaves the 5% positive rate untouched, so the classifier still sees skewed priors. It is tempting because scaling is a genuine preprocessing step, but it belongs to feature preparation, not to resampling techniques such as class weights or SMOTE that alter the target distribution.
- ✗
One-hot encoding of categorical variables
Why it's wrong here
One-hot encoding converts categorical features into binary columns; it does not change the 5% positive rate, so the classifier still sees skewed classes. It is the correct preprocessing step when nominal features must be represented numerically, not when addressing class imbalance.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.