PDE Preparing and Using Data for Analysis Practice Question
You need to split a time-series dataset into training and evaluation sets for a forecasting model. The data is ordered by timestamp. Which splitting technique should you use?
⚠ Common exam trap
PDE often tests the misconception that standard cross-validation techniques (like k-fold or random split) are universally applicable, but for time-series data, they cause data leakage and invalidate the evaluation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Sequential split where training data precedes evaluation data in time.
For time-series forecasting, the training set must only contain data that occurs before the evaluation set to prevent data leakage from the future into the past. A sequential split preserves the temporal order, ensuring the model is trained on historical data and evaluated on subsequent, unseen future data. This mimics real-world forecasting where you predict future values based on past observations. Random or stratified splits violate this principle by allowing future data points to influence training, leading to overly optimistic performance estimates.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Sequential split where training data precedes evaluation data in time.
Why this is correct
A sequential split keeps all training observations earlier in time than evaluation observations, preserving temporal order. This satisfies the forecasting constraint, since random splitting would leak future information into training and produce misleadingly optimistic evaluation metrics.
- ✗
Use k-fold cross-validation with random folds.
Why it's wrong here
Random k-fold folds mix past and future timestamps across every fold, leaking future data into training. It is tempting because k-fold maximises training data and gives variance estimates, and would be correct for stationary independent samples, not ordered time series.
- ✗
Stratified split based on the target variable.
Why it's wrong here
Stratification preserves class proportions, not chronological order, so future timestamps still enter training. It is tempting because stratified splits prevent class imbalance in classification, and would be correct for imbalanced independent data rather than a forecasting series.
- ✗
Random split with 80% training, 20% evaluation.
Why it's wrong here
Random splitting shuffles timestamps, so training rows postdate evaluation rows, leaking future information and inflating accuracy. It is tempting because random splits are standard for independent observations, and would be correct for non-temporal tabular data where row order carries no meaning.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.