Courseiva
easyMultiple Choice

MLA-C01 Practice Question: A data engineer needs to split a time-series…

A data engineer needs to split a time-series dataset into training and validation sets for a forecasting model. Which split method should be used to avoid data leakage?

⚠ Common exam trap

For time-series data in AWS, using random splitting or cross-validation (e.g., with SageMaker) ignores temporal order and leads to data leakage. Always use a temporal split to preserve chronological order.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Temporal split where training uses data up to a cutoff date and validation uses later data.

Time-series data has temporal dependencies, and random splits or k-fold cross-validation with shuffling would cause data leakage by allowing future information to influence training. A temporal split ensures that the model is trained only on past data and validated on future data, preserving the chronological order and preventing leakage.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use k-fold cross-validation with random shuffling.

    Why it's wrong here

    Random shuffling before folding mixes future observations into earlier training folds, so the model learns from data it should not yet see. It is tempting because k-fold cross-validation is standard for independent observations, but time-series data requires forward-chaining splits that preserve chronological order.

  • ✗

    Use feature importance scores to weight the splitting process.

    Why it's wrong here

    Feature importance scores rank predictor influence on the target; they do not partition rows chronologically, so future observations still enter training. It is tempting because importance guides feature selection, but that is a modelling step, not a splitting mechanism, and it leaves the temporal ordering that causes leakage untouched.

  • ✗

    Random split with 80% training and 20% validation.

    Why it's wrong here

    Randomly assigning 80% of rows to training places later timestamps before earlier ones, leaking future information into the model. It is tempting because random splitting is the default for i.i.d. data, but forecasting requires a chronological cutoff, such as training on earlier periods and validating on later ones.

  • ✓

    Temporal split where training uses data up to a cutoff date and validation uses later data.

    Why this is correct

    A temporal split preserves chronological order, training on earlier observations and validating on later ones, which mirrors real forecasting conditions. Random or stratified splits leak future information into training because adjacent time points correlate strongly. This satisfies the stem's requirement to avoid data leakage when validating a time-series forecasting model.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.