Courseiva
hardMultiple Choice

PDE Practice Question: Building a continuous training pipeline that…

A company is building a continuous training pipeline that retrains a model daily using new data from a feature store. The training data must include features computed up to the timestamp of each training run. Which architecture should be used to ensure time-consistent feature values without label leakage?

⚠ Common exam trap

Google Cloud often tests the misconception that simply using the most recent data or a snapshot is sufficient for time-consistency, but the key requirement is to retrieve features as of the exact training timestamp to prevent label leakage, which only point-in-time lookup guarantees.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Vertex AI Feature Store with point-in-time lookup enabled to retrieve features as of the training timestamp.

Vertex AI Feature Store's point-in-time lookup retrieves the exact feature values as they existed at the specified training timestamp, ensuring time-consistency and preventing label leakage. This mechanism avoids using future data that would not have been available at the time of prediction, which is critical for realistic model evaluation and production performance.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Train on a fixed window of the most recent features without considering timestamps.

    Why it's wrong here

    A fixed recent window ignores each training run's timestamp, mixing features computed after earlier labels and causing label leakage. It is tempting because fixed-window training is simple to schedule and suits problems where feature values are static and time ordering is irrelevant.

  • ✓

    Use Vertex AI Feature Store with point-in-time lookup enabled to retrieve features as of the training timestamp.

    Why this is correct

    Point-in-time lookup retrieves feature values as they existed at each training timestamp, preventing future data from leaking into training rows. This satisfies the time-consistent feature requirement, ensuring daily retraining uses only information available at that run's cutoff.

  • ✗

    Store all features in a Cloud SQL database and perform a join at training time.

    Why it's wrong here

    A Cloud SQL join returns current feature values, not the values as of each training run's timestamp, so later-updated features leak future information into historical rows. It is tempting because Cloud SQL is a familiar relational store for joining tables; it would suit static reference data, not time-travelled feature retrieval.

  • ✗

    Use Pub/Sub to stream new features into Cloud Storage and train on the latest snapshot.

    Why it's wrong here

    Streaming features into storage and training on the latest snapshot discards event timestamps, so features computed after each label's time leak future information. It is tempting because snapshot training is straightforward for batch pipelines where temporal correctness is not required.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.