Courseiva
easyMultiple Choice

MLA-C01 Practice Question: A data scientist needs to split a time-series…

A data scientist needs to split a time-series dataset into training and testing sets while avoiding data leakage from future values. Which splitting technique should the data scientist use?

⚠ Common exam trap

MLA-C01 often tests the misconception that standard cross-validation or random splits are suitable for time-series data, but they cause data leakage; candidates must recognize that temporal order must be preserved.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Time-series split (walk-forward validation)

Time-series split (walk-forward validation) is the correct technique because it preserves the temporal order of observations, ensuring that training data always precedes test data. This prevents data leakage from future values, which is critical for time-series forecasting. Random shuffling or standard cross-validation would mix past and future data, leading to overly optimistic performance estimates.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Random shuffle followed by a 80/20 split

    Why it's wrong here

    Random shuffling mixes future observations into the training set, leaking information the model would not have at prediction time. It is tempting because it is the default split for independent samples, and would be correct for non-sequential data where observations have no temporal ordering.

  • ✗

    K-fold cross-validation with shuffled folds

    Why it's wrong here

    Shuffled K-fold cross-validation trains on folds containing later timestamps than the validation fold, leaking future values into training. It is tempting because it maximises data use and gives stable estimates, and would be correct for i.i.d. data without temporal dependence.

  • ✗

    Stratified sampling based on the target variable

    Why it's wrong here

    Stratified sampling preserves target class proportions but still selects rows randomly across the whole timeline, so future values enter training. It is tempting because it balances imbalanced classes, and would be correct for classification tasks on non-temporal data with skewed labels.

  • ✓

    Time-series split (walk-forward validation)

    Why this is correct

    Time-series split (walk-forward validation) preserves chronological order, training only on past observations and testing on subsequent ones. This directly prevents future values leaking into training, satisfying the stem's no-leakage constraint. Random or stratified splitting would mix future data into training, invalidating the evaluation of a temporal model.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.