easyMultiple Choice
MLA-C01 Practice Question: A data scientist discovers that a dataset for…
A data scientist discovers that a dataset for binary classification contains 95% negative samples and 5% positive samples. Which technique is MOST appropriate to address the class imbalance?
⚠ Common exam trap
MLA-C01 often tests the misconception that any resampling works — candidates must distinguish SMOTE's synthetic interpolation from naive duplication (RandomOverSampler) and from destructive undersampling that throws away majority data.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use SMOTE to generate synthetic samples for the minority class
SMOTE (Synthetic Minority Over-sampling Technique) generates new, synthetic minority-class samples by interpolating between existing minority neighbors, which increases the minority representation without simply duplicating rows. This helps the model learn the minority decision boundary and is the standard, most appropriate technique for a 95/5 imbalance. It preserves all majority information while enriching the minority class.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use SMOTE to generate synthetic samples for the minority class
Why this is correct
SMOTE generates synthetic minority-class samples by interpolating between existing positive instances, rebalancing the 95:5 distribution. This gives the classifier enough positive examples to learn the minority pattern, unlike plain resampling that discards data or duplicates points.
- ✗
Remove all samples from the majority class
Why it's wrong here
Deleting every majority-class sample discards the 95% of genuine negatives the model needs, leaving a single-class dataset that cannot train a binary classifier. It is tempting because undersampling the majority class is a recognised balancing technique, but that approach removes only a subset, preserving enough negatives to learn from.
- ✗
Increase the learning rate of the model
Why it's wrong here
Learning rate controls the size of gradient-descent weight updates; it does not change the 95:5 class ratio, so the model still favours the majority class. It is tempting because tuning learning rate does improve convergence, which is correct when training is unstable or slow, not for imbalance.
- ✗
Downsample the majority class to 5% of the original size
Why it's wrong here
Downsampling the majority class to 5% discards roughly 95% of negative records, destroying information and leaving a tiny dataset that biases the model. It is tempting because undersampling does balance class ratios, which suits very large datasets where majority examples are genuinely redundant.
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.