PMLE Practice Question: Collaborating Within and Across Teams to Manage Data and Models
A team is training a model using historical data and wants to avoid data leakage when joining feature values from a feature store. The features include time-varying data like user activity counts. Which retrieval method should they use when creating a training dataset?
⚠ Common exam trap
Many candidates confuse 'latest feature values' with 'correct feature values'—candidates often assume that using the most recent data is always best, but in training, it causes data leakage and inflated offline metrics that fail in production.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use point-in-time correct retrieval with timestamp matching
Point-in-time correct retrieval with timestamp matching ensures that for each training row, the feature values used are the ones that were actually available at the time of the label event, preventing future information from leaking into the training set. This is critical for time-varying features like user activity counts, where using the latest value would introduce look-ahead bias. By joining on entity ID and event timestamp, the feature store returns the feature value as of that timestamp, mimicking the production inference environment.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Retrieve the latest feature values for each entity
Why it's wrong here
Retrieving latest values ignores event timestamps, so training rows receive feature values recorded after the label, leaking future information. Point-in-time retrieval is required. Latest-value retrieval is correct for online inference, where only current feature state matters and no label exists to leak into.
- ✗
Aggregate features over all historical data
Why it's wrong here
Aggregating across all history folds post-label events into each row, so training features encode outcomes the model could not know at prediction time. Point-in-time retrieval uses only values preceding each label timestamp. Whole-history aggregation is valid for descriptive reporting or batch analytics where no temporal prediction boundary exists.
- ✗
Use random sampling of feature values
Why it's wrong here
Random sampling discards the temporal relationship between each label's timestamp and its feature values, so rows can pair a label with activity counts recorded later, leaking future data. Point-in-time joins align values to the label timestamp. Random sampling suits exploratory profiling or unbiased holdout selection, not time-varying training sets.
- ✓
Use point-in-time correct retrieval with timestamp matching
Why this is correct
Point-in-time correct retrieval with timestamp matching returns feature values as they existed at each training example's timestamp, preventing leakage from future user activity counts. Naive latest-value retrieval would expose post-event data, inflating offline metrics relative to production.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
3 more ways this is tested on PMLE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A company uses Vertex AI Feature Store for feature engineering. They need to ensure point-in-time correctness to avoid data leakage during training. Which feature retrieval method should they use?
hard- A.Use the `get_features` API without specifying a timestamp.
- B.Use BigQuery to manually join features with a sliding window.
- ✓ C.Use the offline store with point-in-time join using the `feature_view` with a timestamp column.
- D.Use the online store to retrieve the latest feature values.
Why C: Point-in-time correctness requires retrieving feature values as they existed at the timestamp of each training example, which is exactly what the offline store's point-in-time join does when a feature_view is configured with an event/timestamp column. Vertex AI Feature Store uses this timestamp column to perform an as-of join, preventing future data from leaking into the training row. The online store and timestamp-less get_features calls return only the latest values, which is the classic source of label leakage.
Variation 2. A data scientist needs to retrieve training data from Vertex AI Feature Store that exactly matches the feature values as they were at a specific historical timestamp to avoid label leakage. Which feature view configuration should they use?
medium- ✓ A.Enable point-in-time retrieval on the feature view.
- B.Use the offline store without point-in-time and rely on data ordering.
- C.Use the online store with a timestamp filter.
- D.Create a new feature view with only historical data.
Why A: Point-in-time retrieval is a feature of Vertex AI Feature Store that returns feature values as of a specified timestamp.
Variation 3. A team is building a fraud detection model that requires joining real-time transaction features with historical user features. They need to ensure that the training data does not use future information (data leakage). Which Vertex AI Feature Store capability should they use?
medium- A.Online store serving with Bigtable
- B.Feature store time travel
- ✓ C.Point-in-time correct join
- D.Feature monitoring for drift
Why C: Point-in-time correct joins in Vertex AI Feature Store ensure that when training examples are generated, each row uses only feature values that were valid as of the event timestamp of the label — preventing future data from leaking into training. This is the specific capability designed to avoid label leakage in time-series or event-driven ML. Time travel and online serving do not by themselves guarantee temporal correctness of the join.
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.