mediumMultiple Select
ML Experiment Reproducibility in Vertex AI
A machine learning team is collaborating on a project using Vertex AI Experiments to track model training runs. They want to ensure that all team members can reproduce any experiment by using the same code, data, and environment. Which THREE actions should the team take?
Quick Answer
The answer is to store training code in a Cloud Source Repository with tags linked to experiment IDs, record the dataset path and version in experiment parameters, and log the complete environment specification including framework versions. These three actions are correct because they create a full lineage for each Vertex AI experiment, capturing the three critical pillars of reproducibility: code, data, and environment. By tagging code commits to specific experiment IDs, you ensure any team member can checkout the exact source; recording dataset parameters locks in the data source; and logging environment details prevents dependency drift. On the Google Professional Machine Learning Engineer exam, this question tests your understanding of Vertex AI Experiments’ built-in lineage tracking, often appearing as a multi-select scenario where distractors include vague actions like “save the model artifact” without versioning or “use default runtime” without pinning dependencies. A common trap is forgetting that environment specification must be explicit, not implicit. Memory tip: think “Code, Data, Env” as the three legs of the reproducibility stool—if any leg is missing, the experiment falls over.
⚠ Common exam trap
Google Cloud often tests the distinction between actions that enable reproducibility versus actions that improve model performance or access control, so candidates mistakenly select hyperparameter tuning or service account sharing as reproducibility measures.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Store the training code in a Cloud Source Repository and tag commits with the experiment ID.
Option A is correct because storing the training code in Cloud Source Repositories and tagging commits with the experiment ID creates an immutable, versioned link between the exact code revision and the Vertex AI Experiment run, which is essential for reproducibility. Option B is correct because packaging the training environment in a custom container image pushed to Artifact Registry with a fixed (immutable) tag guarantees every team member runs the identical dependencies, libraries, and runtime, eliminating environment drift. Option C is correct because logging the dataset path and version as experiment parameters in Vertex AI Experiments captures the exact data snapshot used, so the same inputs can be retrieved and reused. Option D is not appropriate because sharing a service account key is an insecure credential-management anti-pattern and does not itself contribute to code, data, or environment reproducibility. Option E is not appropriate because hyperparameter tuning optimizes model performance and does not by itself ensure that code, data, and environment are pinned and reproducible.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Store the training code in a Cloud Source Repository and tag commits with the experiment ID.
Why this is correct
Tagging commits in Cloud Source Repositories binds each Vertex AI experiment to an immutable code revision. This satisfies reproducibility by letting any team member check out the exact commit that produced a given run, rather than relying on a mutable branch.
- ✓
Build a custom container image for training and push it to Artifact Registry with a fixed tag.
Why this is correct
A custom container pushed to Artifact Registry with a fixed tag freezes the exact libraries, dependencies and runtime used for training. This satisfies the environment reproducibility requirement, since the same immutable image can be pulled and rerun by any team member.
- ✓
Record the path and version of the training dataset in the experiment parameters.
Why this is correct
Recording the dataset path and version as experiment parameters captures which exact data generation trained the model. This satisfies the data reproducibility requirement, since teammates can retrieve the identical dataset rather than a bucket whose contents may have changed.
- ✗
Share a service account key with all team members so they can access the same resources.
Why it's wrong here
A shared service account key is a long-lived credential that cannot attribute runs to individuals and is not a reproducibility mechanism for code, data or environment. It is tempting because it grants uniform resource access quickly, which suits short-lived demos, but Vertex AI Experiments reproducibility requires versioned datasets, containers and tracked parameters.
- ✗
Use Vertex AI's hyperparameter tuning job to automatically find the best parameters.
Why it's wrong here
Hyperparameter tuning searches for optimal parameters within a single experiment; it neither versions the code, data or environment nor records them for replay. It is tempting because tuning jobs appear in Vertex AI Experiments, but reproducibility demands logged artefacts and parameters, not automated parameter search.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on PMLE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A data science team is collaborating on a project to build a churn prediction model. They use Vertex AI Workbench instances for development. Each data scientist has their own instance with a persistent disk. They share code via a GitHub repository. They want to ensure that the model training is reproducible across different team members' environments. Currently, they manually install Python packages in their instances, and they have noticed that the model metrics differ slightly between runs on different instances. Which of the following is the best action to ensure reproducibility?
medium- A.Standardize the instance machine type and ensure all have the same number of CPUs.
- B.Use Cloud Functions to run the training code instead.
- ✓ C.Use Vertex AI Experiments with a fixed environment by specifying a prebuilt container.
- D.Create a custom Docker image with all dependencies and use it in Vertex AI Training jobs.
- E.Ask all team members to use the same Python virtual environment and install packages from a requirements.txt file.
Why C: Vertex AI Experiments with a prebuilt container ensures a fixed, reproducible environment by pinning the exact OS, Python version, and all dependencies. This eliminates the variability introduced by manual package installations and differing instance configurations, directly addressing the team's issue of inconsistent model metrics across runs.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.