PMLE Practice Question: Collaborating Within and Across Teams to Manage Data and Models
Your team trains models in a shared Vertex AI project. A data engineer accidentally overwrites a BigQuery training table that three production pipelines depend on, and nobody can tell which pipeline used which version of the data. You need to make dataset versions immutable and traceable so that any training run can be reproduced. What should you do?
⚠ Common exam trap
The trap here is assuming that restricting write permissions or copying data on a schedule provides reproducibility, when only an immutable version identifier recorded per run actually ties a training job to its exact input data.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable BigQuery table snapshots and record the snapshot ID in each pipeline's run metadata.
Immutable dataset versions are needed so any training run can be reproduced exactly, even after the source table changes. BigQuery table snapshots provide durable, read-only point-in-time copies, and storing the snapshot identifier alongside the pipeline run metadata creates the audit trail the team is missing. Access control, nightly copies, and time travel all fail to combine immutability with explicit per-run traceability.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Copy the table nightly into a Cloud Storage bucket using a scheduled Dataproc job.
Why it's wrong here
Nightly copies create coarse, time-based duplicates that may not align with when a pipeline actually read the data, and the copy is still mutable if overwritten. There is no linkage between a specific copy and a pipeline execution, so lineage remains unclear. This adds storage cost and operational overhead without delivering immutable, per-run reproducibility.
- ✓
Enable BigQuery table snapshots and record the snapshot ID in each pipeline's run metadata.
Why this is correct
BigQuery table snapshots create immutable, point-in-time copies of a table that persist independently of later writes, so a training job can always be reproduced from the exact snapshot it consumed. Recording the snapshot ID in pipeline metadata ties the artifact to the run, giving the traceability the team currently lacks without duplicating data or blocking the engineer's writes.
- ✗
Enable BigQuery time travel and query the table with a FOR SYSTEM_TIME AS OF clause during retraining.
Why it's wrong here
Time travel lets you read a table as of a past timestamp, but the retention window is limited and the timestamp is not recorded anywhere by default, so older runs become irreproducible. It also does not prevent overwrites or give an explicit version identifier tied to a pipeline execution. It is useful for recovery, not for durable dataset versioning and lineage.
- ✗
Grant the data engineer only bigquery.dataViewer on the table so they can no longer modify it.
Why it's wrong here
Revoking write access stops future accidental overwrites but does nothing for reproducibility of past or future runs. It also does not make any version immutable or traceable, and it blocks legitimate data refreshes. The scenario requires versioned, reproducible datasets, which IAM restriction alone cannot provide, so this does not solve the described problem.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.