Courseiva
mediumMultiple Choice

MLA-C01 Practice Question: A machine learning engineer needs to split a…

A machine learning engineer needs to split a time-series dataset for a forecasting model. The data spans 3 years of daily sales. Which splitting strategy should they use to avoid look-ahead bias?

⚠ Common exam trap

Watch out — candidates often default to k-fold cross-validation or random splits because they are standard for non-temporal data, failing to recognize that time-series data requires strict temporal ordering to avoid look-ahead bias.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Walk-forward validation (time-series split)

Walk-forward validation (time-series split) is the correct strategy because it preserves the temporal order of the data, training on past observations and testing on future observations sequentially. This avoids look-ahead bias, where future information leaks into the training set, which would invalidate the forecasting model's performance metrics.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    k-fold cross-validation with shuffling

    Why it's wrong here

    Shuffling before folding mixes future observations into earlier training folds, so the model trains on data that postdates its validation points — exactly the look-ahead bias to avoid. K-fold with shuffling suits i.i.d. tabular data, whereas time-series forecasting needs forward-chaining splits that preserve temporal order.

  • ✗

    Random train-test split with 80/20 ratio

    Why it's wrong here

    A random 80/20 split assigns days to train or test irrespective of date, so later sales figures train a model evaluated on earlier ones, leaking future information. Random splitting is fine for independent observations, but ordered forecasting data requires the test set to follow the training period chronologically.

  • ✗

    Stratified sampling based on sales volume

    Why it's wrong here

    Stratified sampling partitions rows by a class label, preserving class proportions; it ignores chronological order, so future days can land in the training set and leak information backwards. It is the right choice for imbalanced classification, but forecasting requires a chronological split such as rolling-origin evaluation.

  • ✓

    Walk-forward validation (time-series split)

    Why this is correct

    Walk-forward validation trains on earlier data and tests on later, chronologically subsequent data, preserving temporal order. This prevents look-ahead bias because the model never sees future observations during training, unlike random k-fold splitting, which would leak future values into past windows.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.