hardMultiple Select
PMLE Practice Question: Which TWO actions should be taken to ensure…
Which TWO actions should be taken to ensure reproducibility of ML experiments when collaborating across teams on Vertex AI?
⚠ Common exam trap
A common mix-up: candidates think 'always use random seeds' is a safe blanket rule, but in practice, seeds must be explicitly set and logged per run, and some operations (e.g., certain GPU kernels) are inherently non-deterministic, making this option an oversimplification that is not a guaranteed action for reproducibility.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Lock dependency versions in a container image used for training
Option A is correct because locking dependency versions inside a container image used for training guarantees that every team member executes the same libraries and framework versions, eliminating environment drift as a source of non-reproducible results on Vertex AI. Option C is correct because versioning datasets with DVC or tracking them through Vertex AI ML Metadata pins the exact data lineage and artifact versions, so any experiment can be re-run against the identical input data. Options B, D, and E do not ensure reproducibility: real-time notebook sharing (B) is a collaboration feature, not a versioning mechanism; letting each team use its own environment (D) introduces dependency divergence; and always using random seeds (E) is misstated, since reproducibility requires fixing a consistent seed value, not merely using random seeds.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Lock dependency versions in a container image used for training
Why this is correct
Locking dependency versions inside a container image guarantees identical library and runtime environments across teams, eliminating drift from differing package versions. This directly satisfies the reproducibility constraint by ensuring every training run executes against the same immutable image, so results can be recreated regardless of which team member or machine triggers the pipeline.
- ✗
Share notebooks via Colab Enterprise with real-time editing
Why it's wrong here
Real-time collaborative editing changes notebook state unpredictably and records no version history, so runs cannot be reproduced. It is tempting because shared notebooks enable live pair work, and it is correct when the goal is interactive brainstorming rather than auditable, repeatable experiment execution.
- ✓
Version control datasets using DVC or Vertex AI ML Metadata
Why this is correct
Versioning datasets with DVC or Vertex AI ML Metadata records immutable dataset snapshots and lineage, so collaborators can rerun training against the exact same data revision rather than a mutable source. This directly satisfies the reproducibility constraint by pinning the data artefact, complementing code and environment versioning across teams.
- ✗
Allow each team to use their own preferred environment
Why it's wrong here
Per-team environments introduce divergent library and dependency versions, so runs cannot be reproduced by colleagues. It is tempting because it respects each team's existing tooling and works fine when teams operate independently and never need to rerun each other's experiments.
- ✗
Always use random seeds for all random operations
Why it's wrong here
Fixed seeds alone do not guarantee reproducibility: hardware non-determinism, library versions and parallel execution still produce divergent results across teams. It is tempting because seeding is genuinely necessary for deterministic random operations, and it is correct when the same code runs on identical, pinned environments.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.