mediumMultiple Choice
MLA-C01 Practice Question: A team needs to split a time-series dataset for a…
A team needs to split a time-series dataset for a forecasting model. They want to avoid data leakage and evaluate model performance on future unseen data. Which data splitting strategy should they use?
⚠ Common exam trap
AWS often tests the misconception that standard cross-validation techniques like k-fold or random holdout are universally applicable, but the trap here is that they fail for time-series data because they ignore temporal dependencies, leading to data leakage and invalid performance metrics.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Walk-forward validation
Walk-forward validation is the correct strategy for time-series forecasting because it preserves the temporal order of data, training on past observations and testing on future ones in sequential steps. This avoids data leakage by ensuring that no future information is used to predict past events, which is critical for evaluating model performance on unseen future data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Holdout validation with random sampling
Why it's wrong here
Random sampling for the holdout draws rows from all periods, so the training set contains future observations relative to the test set. It is tempting because a random holdout suits i.i.d. data, whereas forecasting requires the test period to follow the training period chronologically.
- ✓
Walk-forward validation
Why this is correct
Walk-forward validation repeatedly trains on past data and tests on the immediately following period, preserving chronological order. This mirrors real forecasting conditions and prevents future observations leaking into training, satisfying the requirement to evaluate on genuinely unseen future data.
- ✗
K-fold cross-validation
Why it's wrong here
K-fold cross-validation shuffles or folds observations across the whole series, letting later timestamps train on earlier ones and leaking future information. It is tempting because k-fold suits i.i.d. tabular data where every row is exchangeable, not ordered time series.
- ✗
Random stratified split
Why it's wrong here
A random stratified split samples rows across the entire timeline, so training data includes observations recorded after test points, leaking future information. It is tempting because stratification preserves class proportions in imbalanced i.i.d. datasets, not chronological ordering.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.