mediumMultiple Choice
MLA-C01 Practice Question: A machine learning engineer needs to split a…
A machine learning engineer needs to split a time-series dataset for a forecasting model. The data spans 3 years of daily sales. Which splitting strategy should they use to avoid look-ahead bias?
⚠ Common exam trap
Watch out — candidates often default to k-fold cross-validation or random splits because they are standard for non-temporal data, failing to recognize that time-series data requires strict temporal ordering to avoid look-ahead bias.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Walk-forward validation (time-series split)
Walk-forward validation (time-series split) is the correct strategy because it preserves the temporal order of the data, training on past observations and testing on future observations sequentially. This avoids look-ahead bias, where future information leaks into the training set, which would invalidate the forecasting model's performance metrics.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
k-fold cross-validation with shuffling
Why it's wrong here
Shuffling before folding mixes future observations into earlier training folds, so the model trains on data that postdates its validation points — exactly the look-ahead bias to avoid. K-fold with shuffling suits i.i.d. tabular data, whereas time-series forecasting needs forward-chaining splits that preserve temporal order.
- ✗
Random train-test split with 80/20 ratio
Why it's wrong here
A random 80/20 split assigns days to train or test irrespective of date, so later sales figures train a model evaluated on earlier ones, leaking future information. Random splitting is fine for independent observations, but ordered forecasting data requires the test set to follow the training period chronologically.
- ✗
Stratified sampling based on sales volume
Why it's wrong here
Stratified sampling partitions rows by a class label, preserving class proportions; it ignores chronological order, so future days can land in the training set and leak information backwards. It is the right choice for imbalanced classification, but forecasting requires a chronological split such as rolling-origin evaluation.
- ✓
Walk-forward validation (time-series split)
Why this is correct
Walk-forward validation trains on earlier data and tests on later, chronologically subsequent data, preserving temporal order. This prevents look-ahead bias because the model never sees future observations during training, unlike random k-fold splitting, which would leak future values into past windows.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.