hardMultiple Choice
MLA-C01 Practice Question: A machine learning engineer is building a…
A machine learning engineer is building a time-series forecasting model to predict daily sales for the next 30 days. The dataset spans two years of daily sales data. To evaluate model performance, the engineer needs to simulate a realistic forecasting scenario where the model is trained on past data and tested on future data without leakage. Which data splitting strategy should they use?
⚠ Common exam trap
MLA-C01 often tests whether candidates recognize that standard cross-validation techniques leak future information in time-series problems, tempting them to pick k-fold or random hold-out for its familiarity.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Walk-forward validation with an expanding window
Walk-forward validation with an expanding window respects the temporal order of time-series data by training on all data up to time t and testing on the next period, then expanding the training set forward. This simulates the real forecasting scenario where the model only ever sees past data when predicting the future, and it prevents leakage that random splits would introduce. For a 30-day forecast horizon on two years of daily data, this is the standard approach.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Hold-out validation using a random 80/20 split
Why it's wrong here
A random 80/20 split shuffles observations across time, so the model trains on future days and tests on past ones, leaking temporal information. Hold-out splitting suits independent, identically distributed data; time-series forecasting requires chronological partitioning with the test set strictly after the training set.
- ✓
Walk-forward validation with an expanding window
Why this is correct
Walk-forward validation with an expanding window trains on all past observations and tests on the subsequent future period, then rolls forward, mirroring real forecasting and preventing the temporal leakage that random or shuffled splits would introduce.
- ✗
Stratified sampling based on sales volume
Why it's wrong here
Stratified sampling preserves class proportions but still assigns individual days randomly across folds, mixing past and future observations and leaking future information into training. Stratification is correct for imbalanced classification, where preserving label ratios across splits matters more than temporal order.
- ✗
k-fold cross-validation with random shuffling
Why it's wrong here
k-fold cross-validation with random shuffling interleaves past and future days across every fold, so training always sees data from the test period. k-fold suits stationary, exchangeable datasets; forecasting requires forward-chaining splits where each validation window follows its training window in time.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.