PMLE Collaborating to manage data and models Practice Question
You are a machine learning engineer at a retail company. Your team uses Vertex AI Pipelines to train a model that predicts customer churn. The pipeline reads training data from a BigQuery table that is updated daily by an external marketing analytics team. You need to ensure that every pipeline run uses a consistent snapshot of the data and that you can reproduce any past run for auditing. What should you do?
⚠ Common exam trap
The trap here is assuming that querying a table with a date filter or exporting data provides a consistent snapshot, when only a native BigQuery snapshot guarantees immutability.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Before each pipeline run, create a BigQuery table snapshot of the source table and configure the pipeline to read from that snapshot.
Using BigQuery table snapshots ensures that each pipeline run reads an immutable, point-in-time copy of the source data. This guarantees consistency across runs and enables exact reproducibility for audits, because the snapshot preserves the data as it existed when created. Other methods either do not freeze the data or introduce complexity without the same guarantees.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Configure the pipeline's training component to query the BigQuery table using a parameter that specifies the current date, and log that date in the pipeline run metadata.
Why it's wrong here
Parameterizing the date does not guarantee a consistent snapshot because the underlying table can be updated while the pipeline runs, and the date alone does not capture the exact table state. Logging the date in metadata helps with tracking but does not make the data immutable, so past runs may not be reproducible if the table changes.
- ✓
Before each pipeline run, create a BigQuery table snapshot of the source table and configure the pipeline to read from that snapshot.
Why this is correct
BigQuery table snapshots are immutable and preserve the table's data at the time of creation. By reading from a snapshot, each pipeline run uses a consistent, reproducible dataset, and the snapshot can be referenced later for auditing. This directly addresses both consistency and reproducibility without altering the source table.
- ✗
Set up a Cloud Scheduler job that copies the BigQuery table to a Cloud Storage bucket as CSV files, and have the pipeline read from that bucket.
Why it's wrong here
Exporting to Cloud Storage creates a copy, but it is not an atomic snapshot and may not capture a consistent point-in-time view if the table is updated during export. It also adds unnecessary steps and does not provide the same reproducibility guarantees as a native BigQuery snapshot.
- ✗
Use BigQuery's streaming inserts to push data into a new table that the pipeline reads, ensuring that the data is frozen at the time of insertion.
Why it's wrong here
Streaming inserts do not freeze data; they append rows to a table that can still be modified. This approach adds complexity and does not provide an immutable snapshot, so later changes could still affect the data read by the pipeline. It also does not guarantee a consistent point-in-time view.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.