Courseiva

PDE Preparing and Using Data for Analysis Practice Question

You are designing a data pipeline for ML training with Vertex AI. You need to split time-series data into train/validation/test sets without leaking future data. Which THREE practices should you follow?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a sliding window validation approach for hyperparameter tuning.

Option A is correct because a sliding-window (rolling-origin) validation scheme trains on past data and validates on the immediately following window, which respects temporal order and prevents future leakage during hyperparameter tuning. Option B is correct because keeping all data points from the same time period (e.g., same day or timestamp bucket) in a single split avoids boundary leakage, where records from one period appear in both train and test. Option E is correct because using an explicit date column to define split boundaries (for example, train on data before 2024-01-01, validate on Q1 2024, test on Q2 2024) enforces a strict chronological cutoff. Option C is not appropriate because Looker is a BI/analytics and dashboarding tool, not a mechanism for producing leakage-safe ML dataset splits. Option D is wrong because random row assignment mixes past and future observations across splits, which is exactly the temporal leakage the scenario must avoid.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use a sliding window validation approach for hyperparameter tuning.

    Why this is correct

    Sliding-window validation trains on an expanding or rolling past window and validates on the immediately following period, so every fold respects chronological order. This directly prevents future leakage during hyperparameter tuning, satisfying the stem's requirement to split time-series data without exposing the model to later observations.

  • ✓

    Ensure that all data points for a given time period are in the same split.

    Why this is correct

    Grouping all records from one timestamp or period into a single split prevents the same moment appearing in both training and validation. This preserves temporal integrity, since leakage often arises when contemporaneous rows straddle splits, directly satisfying the stem's constraint against future data contaminating training.

  • ✗

    Use Looker to generate the splits automatically.

    Why it's wrong here

    Looker is a business intelligence and dashboarding product with no role in generating ML dataset splits, so it cannot enforce temporal ordering. It is tempting because it sits within Google Cloud's analytics portfolio, and it would be correct if the requirement were visualising pipeline metrics rather than partitioning time-series data.

  • ✗

    Randomly assign rows to each split to ensure statistical distribution.

    Why it's wrong here

    Random row assignment lets future observations enter training, leaking information across the temporal boundary and inflating validation scores. It is tempting because random splitting is standard for independent, identically distributed tabular data, and it would be correct if the records had no time ordering.

  • ✓

    Use a date column to define the split boundaries.

    Why this is correct

    A date column provides the explicit chronological ordering needed to place split boundaries at points in time rather than at random rows. This guarantees training data precedes validation and test data, satisfying the stem's requirement to avoid leaking future observations into earlier splits.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.