Courseiva
hardMultiple Choice

MLA-C01 Practice Question: An ML engineer is preparing a time-series dataset…

An ML engineer is preparing a time-series dataset for a forecasting model that predicts daily sales for the next 30 days. The dataset contains 3 years of daily sales data. Which data splitting strategy should the engineer use to evaluate the model's performance on future data?

⚠ Common exam trap

MLA-C01 often tests whether candidates recognize that standard cross-validation techniques (k-fold, stratified, leave-one-out) leak future information in time-series problems, tempting them to pick a familiar but temporally invalid split.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Walk-forward validation (time-series split)

Walk-forward validation (time-series split) preserves the temporal order of observations by training on past data and testing on subsequent future periods, which mirrors how a forecasting model is actually used in production. Because daily sales data has trend, seasonality, and autocorrelation, random or stratified splits leak future information into training and produce optimistically biased metrics. Walk-forward validation with a 30-day horizon directly simulates predicting the next 30 days from prior history.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Leave-one-out cross-validation

    Why it's wrong here

    Leave-one-out trains on later data to predict earlier points, leaking future information into evaluation, so reported error understates true forecasting performance. It is tempting because leave-one-out maximises training data for small, independent, identically distributed datasets where temporal order is irrelevant.

  • ✗

    Random 80/20 train-test split

    Why it's wrong here

    Random splitting shuffles chronological observations, so training rows can postdate test rows, leaking future information and inflating accuracy on the 30-day horizon. It is tempting because random splits suit independent, identically distributed tabular data, where ordering carries no meaning.

  • ✗

    Stratified k-fold cross-validation

    Why it's wrong here

    Stratified k-fold preserves class proportions but still shuffles time order across folds, so the model trains on later dates and tests on earlier ones. It is tempting for imbalanced classification, where preserving label ratios per fold genuinely improves evaluation.

  • ✓

    Walk-forward validation (time-series split)

    Why this is correct

    Walk-forward validation trains on all data up to a cutoff and tests on the immediately following window, then rolls the cutoff forward. This preserves chronological order, so the model is always evaluated on genuinely future observations rather than leaking later sales into training.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.