PMLE Automating and Orchestrating ML Pipelines Practice Question
A team is using Vertex AI Pipelines to orchestrate a machine learning workflow. They want to ensure that the pipeline can be reproduced with the same results even if the underlying data changes. Which practice should they follow?
⚠ Common exam trap
The trap here is focusing only on code versioning or hardware consistency, while overlooking the critical role of data versioning and random seed control in achieving reproducible ML results.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a fixed random seed in the training component and snapshot the training data.
Reproducibility in ML pipelines requires controlling both the code and the data. A fixed random seed makes training deterministic, and snapshotting the training data ensures the same input is used across runs. Without data versioning, changes in the source data would lead to different results. Therefore, combining a fixed seed with data snapshots is the correct approach.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use a fixed random seed in the training component and snapshot the training data.
Why this is correct
Reproducibility requires controlling both code and data. A fixed random seed ensures deterministic behavior in stochastic processes like weight initialization. Snapshotting the training data (e.g., using a BigQuery snapshot or copying to a versioned Cloud Storage bucket) ensures the same data is used. Together, these practices enable reproducible results even if the source data changes later.
- ✗
Run the pipeline on a dedicated Vertex AI training cluster with the same machine type and accelerators.
Why it's wrong here
Consistent hardware can reduce variability due to nondeterministic GPU operations, but it does not control data changes. Even with identical hardware, if the input data differs, the model output will differ. Hardware consistency is a minor factor compared to data and seed control, and it does not ensure reproducibility by itself.
- ✗
Version all pipeline components and pin their dependencies to specific versions.
Why it's wrong here
Versioning components and dependencies is important for reproducibility, but it does not guarantee the same results if the data changes. The data itself must be versioned or snapshotted. Pinning dependencies ensures code consistency, but data changes can still alter outcomes. Therefore, this alone is insufficient.
- ✗
Store the pipeline definition in a Git repository and use the same pipeline template for all runs.
Why it's wrong here
Storing the pipeline definition in Git provides version control for the pipeline code, but it does not address data variability. Using the same template ensures the same workflow structure, but if the data changes, the results will differ. Data versioning is missing, so reproducibility is not guaranteed.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.