easyMultiple Choice
MLA-C01 Practice Question: A data scientist needs to split a time-series…
A data scientist needs to split a time-series dataset into training and testing sets while avoiding data leakage from future values. Which splitting technique should the data scientist use?
⚠ Common exam trap
MLA-C01 often tests the misconception that standard cross-validation or random splits are suitable for time-series data, but they cause data leakage; candidates must recognize that temporal order must be preserved.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Time-series split (walk-forward validation)
Time-series split (walk-forward validation) is the correct technique because it preserves the temporal order of observations, ensuring that training data always precedes test data. This prevents data leakage from future values, which is critical for time-series forecasting. Random shuffling or standard cross-validation would mix past and future data, leading to overly optimistic performance estimates.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Random shuffle followed by a 80/20 split
Why it's wrong here
Random shuffling mixes future observations into the training set, leaking information the model would not have at prediction time. It is tempting because it is the default split for independent samples, and would be correct for non-sequential data where observations have no temporal ordering.
- ✗
K-fold cross-validation with shuffled folds
Why it's wrong here
Shuffled K-fold cross-validation trains on folds containing later timestamps than the validation fold, leaking future values into training. It is tempting because it maximises data use and gives stable estimates, and would be correct for i.i.d. data without temporal dependence.
- ✗
Stratified sampling based on the target variable
Why it's wrong here
Stratified sampling preserves target class proportions but still selects rows randomly across the whole timeline, so future values enter training. It is tempting because it balances imbalanced classes, and would be correct for classification tasks on non-temporal data with skewed labels.
- ✓
Time-series split (walk-forward validation)
Why this is correct
Time-series split (walk-forward validation) preserves chronological order, training only on past observations and testing on subsequent ones. This directly prevents future values leaking into training, satisfying the stem's no-leakage constraint. Random or stratified splitting would mix future data into training, invalidating the evaluation of a temporal model.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.