PMLE Collaborating to manage data and models Practice Question
A retail company uses Vertex AI AutoML to train a product recommendation model. They have a dataset of past purchases stored in BigQuery. The data science team wants to iteratively train and improve the model. They need to track which dataset version was used for each model and preserve the exact data for reproducibility. They currently export data to CSV files and store them in Cloud Storage. However, the dataset is updated daily, and they want to ensure that models are trained on a consistent snapshot. What should they do?
⚠ Common exam trap
Google Cloud often tests the misconception that exporting to CSV or using Vertex AI Dataset is sufficient for versioning, when in fact BigQuery snapshots provide the native, scalable, and auditable mechanism for point-in-time data consistency without data duplication.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use BigQuery snapshots to capture a versioned dataset and reference the snapshot in the training pipeline.
BigQuery snapshots provide a consistent, versioned view of the dataset at a specific point in time, ensuring reproducibility without duplicating data. By referencing the snapshot in the Vertex AI training pipeline, the team can train models on the exact same data snapshot, even as the source table is updated daily. This approach avoids the overhead of exporting to CSV and Cloud Storage while maintaining data integrity and lineage.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Vertex AI Dataset service to create a dataset and export it to BigQuery.
Why it's wrong here
Vertex AI Datasets cannot be exported into BigQuery; the service manages data for training within Vertex AI, and the direction of movement is wrong. Creating a managed Vertex AI Dataset would be correct when data already resides in Cloud Storage or BigQuery and needs governed versioning for AutoML.
- ✓
Use BigQuery snapshots to capture a versioned dataset and reference the snapshot in the training pipeline.
Why this is correct
BigQuery snapshots preserve table data at a point in time and are immutable, so each training run references a fixed snapshot rather than the daily-changing source table. This guarantees a consistent dataset version and reproducibility without manual CSV exports.
- ✗
Train the model directly on the BigQuery table and let AutoML handle versioning.
Why it's wrong here
Training directly on the live BigQuery table reads whatever rows exist at run time, so the daily updates change the data between runs and no immutable snapshot is preserved for reproducibility. BigQuery direct training suits stable or append-only tables where version pinning is unnecessary.
- ✗
Export the data to a timestamped CSV file and store it in Cloud Storage before each training run.
Why it's wrong here
Timestamped CSV exports do capture a snapshot, but the stem states CSV export is the current approach causing the problem, and manual files give no managed lineage linking each model to its dataset version. This suits ad-hoc experiments, not iterative AutoML training requiring tracked dataset versions.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.