Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

An ML engineer needs to split a dataset into training, validation, and test sets. The dataset has a time-based column that should not be leaked. Which split method is most appropriate?

⚠ Common exam trap

AWS often tests the concept of data leakage by presenting random or stratified splits as viable options, trapping candidates who overlook the time-based column and assume standard splitting methods are always safe.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Temporal split based on date

A temporal split ensures that the time-based column is not leaked by preserving the chronological order of the data. This method uses the date column to assign earlier records to the training set and later records to the validation and test sets, preventing future information from influencing the model during training.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Stratified split based on target

    Why it's wrong here

    Stratifying preserves the target class ratio but still shuffles rows across time, so future data leaks into training. It is tempting for imbalanced classification on i.i.d. data, where preserving class proportions in each split prevents skewed evaluation.

  • ✓

    Temporal split based on date

    Why this is correct

    A temporal split partitions rows by date, so training uses earlier records and validation/test use later ones. This preserves chronological order and prevents future information leaking into training, directly satisfying the stem's constraint that the time-based column must not be leaked. Random or stratified splits would mix periods and leak future data.

  • ✗

    Random split with 70/20/10

    Why it's wrong here

    Random splitting shuffles rows, so future observations can land in the training set and leak temporal information into the model. It is tempting because it is the default for i.i.d. data, where shuffling is harmless and gives balanced, representative partitions.

  • ✗

    K-fold cross-validation

    Why it's wrong here

    K-fold cross-validation reuses every row across folds, so later timestamps train folds that validate on earlier ones, leaking future information. It is tempting for small i.i.d. datasets, where it maximises training data and gives a robust performance estimate.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.