easyMultiple Choice
MLA-C01 Practice Question: A data scientist is preparing a dataset for…
A data scientist is preparing a dataset for binary classification. The dataset has a target variable with 90% of samples belonging to class 0 and 10% to class 1. Which data splitting strategy should the scientist use to ensure that the training and test sets maintain the same class proportion as the original dataset?
⚠ Common exam trap
MLA-C01 often tests the difference between splitting strategies and resampling techniques — candidates pick k-fold cross-validation because it sounds rigorous, but the question specifically asks for preserving class proportions, which is the definition of stratified sampling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Stratified sampling
Stratified sampling preserves the class proportion of the original dataset in each split, which is essential for imbalanced binary classification (90/10). By sampling within each class stratum, the training and test sets both retain roughly 90% class 0 and 10% class 1, preventing a split where the minority class is underrepresented or absent.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Time-series split
Why it's wrong here
Time-series split partitions data chronologically with expanding or rolling windows, preserving temporal order rather than class proportions, so the 90/10 imbalance is not maintained. It is tempting because it is the correct choice when observations are sequential and future data must not leak into training.
- ✗
Simple random split
Why it's wrong here
Simple random splitting assigns rows independently, so with a 90/10 imbalance the test set's class ratio can drift substantially from the original. It is tempting because it is the correct choice for large, balanced datasets where random sampling already approximates the population distribution.
- ✗
k-fold cross-validation
Why it's wrong here
k-fold cross-validation partitions data into folds for repeated training and evaluation; it does not by itself preserve the 90/10 class ratio unless stratified, and it is not a single train/test split. It is tempting because it is the correct choice when maximising use of limited data for model evaluation.
- ✓
Stratified sampling
Why this is correct
Stratified sampling preserves the original class ratio in each split by sampling within each class separately, so both training and test sets retain roughly 90% class 0 and 10% class 1. Random splitting could otherwise skew the minority class, satisfying the stem's proportion constraint.
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.