Courseiva
mediumMultiple Choice

MLA-C01 Practice Question: A team needs to split a time-series dataset for a…

A team needs to split a time-series dataset for a forecasting model. They want to avoid data leakage and evaluate model performance on future unseen data. Which data splitting strategy should they use?

⚠ Common exam trap

AWS often tests the misconception that standard cross-validation techniques like k-fold or random holdout are universally applicable, but the trap here is that they fail for time-series data because they ignore temporal dependencies, leading to data leakage and invalid performance metrics.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Walk-forward validation

Walk-forward validation is the correct strategy for time-series forecasting because it preserves the temporal order of data, training on past observations and testing on future ones in sequential steps. This avoids data leakage by ensuring that no future information is used to predict past events, which is critical for evaluating model performance on unseen future data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Holdout validation with random sampling

    Why it's wrong here

    Random sampling for the holdout draws rows from all periods, so the training set contains future observations relative to the test set. It is tempting because a random holdout suits i.i.d. data, whereas forecasting requires the test period to follow the training period chronologically.

  • ✓

    Walk-forward validation

    Why this is correct

    Walk-forward validation repeatedly trains on past data and tests on the immediately following period, preserving chronological order. This mirrors real forecasting conditions and prevents future observations leaking into training, satisfying the requirement to evaluate on genuinely unseen future data.

  • ✗

    K-fold cross-validation

    Why it's wrong here

    K-fold cross-validation shuffles or folds observations across the whole series, letting later timestamps train on earlier ones and leaking future information. It is tempting because k-fold suits i.i.d. tabular data where every row is exchangeable, not ordered time series.

  • ✗

    Random stratified split

    Why it's wrong here

    A random stratified split samples rows across the entire timeline, so training data includes observations recorded after test points, leaking future information. It is tempting because stratification preserves class proportions in imbalanced i.i.d. datasets, not chronological ordering.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.