mediumMultiple Choice
MLA-C01 Practice Question: A machine learning engineer is preparing a…
A machine learning engineer is preparing a dataset for a binary classification model. The dataset has 10,000 samples with a 1:100 class imbalance. The engineer needs to balance the classes before training. Which technique would create a balanced dataset without discarding majority class samples and without generating synthetic data?
⚠ Common exam trap
MLA-C01 often tests the distinction between techniques that balance data (oversampling, undersampling, SMOTE) and techniques that adjust learning (class weights), causing candidates to pick cost-sensitive learning when the question explicitly requires a balanced dataset.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Random oversampling of the minority class
Random oversampling of the minority class duplicates existing minority samples until the class distribution is balanced, which increases the minority count without discarding any majority samples and without creating synthetic data. This satisfies all three constraints: balanced dataset, no majority samples removed, and no synthetic generation. It is a straightforward resampling technique commonly used as a baseline before considering more advanced methods.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cost-sensitive learning with class weights
Why it's wrong here
Class weights alter the loss function's penalty per class; they do not create a balanced dataset, leaving the 1:100 sample ratio untouched. It is tempting because it addresses imbalance without discarding or synthesising rows, and would be correct where the requirement is to bias training toward the minority class rather than rebalance the data itself.
- ✓
Random oversampling of the minority class
Why this is correct
Random oversampling duplicates existing minority class records until both classes are equally represented, satisfying the stem's constraints: no majority samples are discarded and no synthetic data is generated. It differs from SMOTE, which interpolates new synthetic points, and from undersampling, which removes majority data.
- ✗
Synthetic Minority Over-sampling Technique (SMOTE)
Why it's wrong here
SMOTE fabricates new minority-class points by interpolating between existing neighbours, so it generates synthetic data — explicitly excluded by the requirement. It is tempting because it balances classes without discarding majority samples, and would be correct where synthetic minority examples are acceptable and the aim is to avoid information loss from undersampling.
- ✗
Random undersampling of the majority class
Why it's wrong here
Random undersampling discards majority-class samples, directly violating the stem's requirement to retain all 10,000 records. It is tempting because it genuinely balances classes and is computationally cheap, making it a valid choice when the majority class is enormous and losing data is acceptable. Here, the explicit no-discard constraint rules it out.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.