Courseiva
easyMultiple Choice

MLA-C01 Practice Question: A data scientist is preparing a dataset for…

A data scientist is preparing a dataset for binary classification. The dataset has a target variable with 90% of samples belonging to class 0 and 10% to class 1. Which data splitting strategy should the scientist use to ensure that the training and test sets maintain the same class proportion as the original dataset?

⚠ Common exam trap

MLA-C01 often tests the difference between splitting strategies and resampling techniques — candidates pick k-fold cross-validation because it sounds rigorous, but the question specifically asks for preserving class proportions, which is the definition of stratified sampling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Stratified sampling

Stratified sampling preserves the class proportion of the original dataset in each split, which is essential for imbalanced binary classification (90/10). By sampling within each class stratum, the training and test sets both retain roughly 90% class 0 and 10% class 1, preventing a split where the minority class is underrepresented or absent.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Time-series split

    Why it's wrong here

    Time-series split partitions data chronologically with expanding or rolling windows, preserving temporal order rather than class proportions, so the 90/10 imbalance is not maintained. It is tempting because it is the correct choice when observations are sequential and future data must not leak into training.

  • ✗

    Simple random split

    Why it's wrong here

    Simple random splitting assigns rows independently, so with a 90/10 imbalance the test set's class ratio can drift substantially from the original. It is tempting because it is the correct choice for large, balanced datasets where random sampling already approximates the population distribution.

  • ✗

    k-fold cross-validation

    Why it's wrong here

    k-fold cross-validation partitions data into folds for repeated training and evaluation; it does not by itself preserve the 90/10 class ratio unless stratified, and it is not a single train/test split. It is tempting because it is the correct choice when maximising use of limited data for model evaluation.

  • ✓

    Stratified sampling

    Why this is correct

    Stratified sampling preserves the original class ratio in each split by sampling within each class separately, so both training and test sets retain roughly 90% class 0 and 10% class 1. Random splitting could otherwise skew the minority class, satisfying the stem's proportion constraint.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.