Courseiva
mediumMultiple Choice

PMLE Practice Question: An ML engineer is scaling a prototype to…

An ML engineer is scaling a prototype to production using Vertex AI Pipelines. The pipeline includes data validation, preprocessing, training, and deployment steps. They want to ensure that the pipeline can be reproduced and audited. What is the best practice?

⚠ Common exam trap

It's easy for candidates to confuse partial reproducibility measures (pinning requirements.txt, using fixed Docker tags) with full pipeline reproducibility, which requires a versioned, declarative pipeline definition plus automatic artifact lineage.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Define the pipeline using Kubeflow Pipelines SDK and run it on Vertex AI Pipelines.

Defining the pipeline with the Kubeflow Pipelines (KFP) SDK and running it on Vertex AI Pipelines produces a versioned, declarative pipeline specification (a compiled YAML/JSON IR) that captures each step, its container image, inputs, outputs, and parameters. Vertex AI Pipelines stores this as a reproducible pipeline run with lineage tracking, so any execution can be audited and re-run deterministically. This is the canonical Google-recommended approach for reproducible, auditable ML workflows on Vertex AI.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Define the pipeline using Kubeflow Pipelines SDK and run it on Vertex AI Pipelines.

    Why this is correct

    Defining the pipeline with the Kubeflow Pipelines SDK gives a versioned, declarative specification that Vertex AI Pipelines executes and records, so each run is reproducible and auditable. Ad hoc scripts or manual steps would not provide this traceability.

  • ✗

    Use a Docker container with fixed tags and manually record runs.

    Why it's wrong here

    Fixed Docker tags are mutable, so the same tag can point to different images over time, and manual run recording leaves no lineage linking artefacts to executions. It is tempting because containers do package dependencies, and this would suit a small local prototype without audit requirements.

  • ✗

    Store all data and models in a single Cloud Storage bucket with no versioning.

    Why it's wrong here

    A single bucket without versioning overwrites objects, so prior data and model artefacts cannot be retrieved or traced to a pipeline run. It is tempting because one bucket is simple to configure, and it would be adequate for disposable scratch data with no reproducibility or audit obligations.

  • ✗

    Pin all library versions in a requirements.txt file.

    Why it's wrong here

    Pinning library versions fixes only Python dependencies; it does not capture pipeline component definitions, container images, parameters, or run lineage needed for auditing. It is tempting because it is standard Python practise, and it would be correct for reproducing a single training script's environment.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. An ML team is moving from a prototype Jupyter notebook to a production training pipeline. They want to ensure reproducibility. Which approach should they take?

easy
  • A.Use interactive parameter tuning.
  • ✓ B.Use a container with fixed dependencies and record hyperparameters.
  • C.Export the notebook's output model directly.
  • D.Save the notebook as a .py file.

Why B: Using a container with fixed dependencies and recording hyperparameters ensures that the training environment and configuration are captured, enabling exact reproduction. Option A is wrong because interactive parameter tuning is not reproducible—it introduces manual adjustments. Option C is wrong because exporting the notebook's output model directly lacks environment tracking and hyperparameter records. Option D is wrong because saving the notebook as a .py file does not capture the full environment or dependencies.

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.