Courseiva

PDE Preparing and Using Data for Analysis Practice Question

You need to split a time-series dataset into training and evaluation sets for a forecasting model. The data is ordered by timestamp. Which splitting technique should you use?

⚠ Common exam trap

PDE often tests the misconception that standard cross-validation techniques (like k-fold or random split) are universally applicable, but for time-series data, they cause data leakage and invalidate the evaluation.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Sequential split where training data precedes evaluation data in time.

For time-series forecasting, the training set must only contain data that occurs before the evaluation set to prevent data leakage from the future into the past. A sequential split preserves the temporal order, ensuring the model is trained on historical data and evaluated on subsequent, unseen future data. This mimics real-world forecasting where you predict future values based on past observations. Random or stratified splits violate this principle by allowing future data points to influence training, leading to overly optimistic performance estimates.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Sequential split where training data precedes evaluation data in time.

    Why this is correct

    A sequential split keeps all training observations earlier in time than evaluation observations, preserving temporal order. This satisfies the forecasting constraint, since random splitting would leak future information into training and produce misleadingly optimistic evaluation metrics.

  • ✗

    Use k-fold cross-validation with random folds.

    Why it's wrong here

    Random k-fold folds mix past and future timestamps across every fold, leaking future data into training. It is tempting because k-fold maximises training data and gives variance estimates, and would be correct for stationary independent samples, not ordered time series.

  • ✗

    Stratified split based on the target variable.

    Why it's wrong here

    Stratification preserves class proportions, not chronological order, so future timestamps still enter training. It is tempting because stratified splits prevent class imbalance in classification, and would be correct for imbalanced independent data rather than a forecasting series.

  • ✗

    Random split with 80% training, 20% evaluation.

    Why it's wrong here

    Random splitting shuffles timestamps, so training rows postdate evaluation rows, leaking future information and inflating accuracy. It is tempting because random splits are standard for independent observations, and would be correct for non-temporal tabular data where row order carries no meaning.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.