easyMultiple Select
Best Practices for Building ML Pipelines on Vertex AI
Which TWO options are best practices for building ML pipelines on Vertex AI?
Quick Answer
The answer is leveraging Vertex ML Metadata to track artifact lineage and using custom container components for pipeline steps. These are best practices for building ML pipelines on Vertex AI because they ensure reproducibility and modularity: Vertex ML Metadata automatically records the inputs, outputs, and parameters of each pipeline run, creating a complete lineage graph that helps debug model drift and trace data transformations, while custom containers encapsulate dependencies and libraries, allowing each step to execute consistently regardless of the environment. On the Google Professional Machine Learning Engineer exam, this tests your understanding of MLOps fundamentals—specifically how to design pipelines that are auditable and maintainable at scale. A common trap is selecting manual logging or monolithic steps, which violate the principles of automation and separation of concerns. Memory tip: think “Metadata for memory, Containers for consistency.”
⚠ Common exam trap
Google Cloud often tests the misconception that serverless functions like Cloud Functions are suitable for ML pipeline steps, but the trap is that ML steps require persistent state, longer timeouts, and specialized hardware, which Cloud Functions cannot provide.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use custom container components to encapsulate reusable logic
Option C is correct because custom container components let you package arbitrary dependencies, libraries, and code into a portable, versioned image that Vertex AI Pipelines can reuse across multiple pipelines, which is a recommended way to encapsulate and share reusable logic. Option E is correct because Vertex ML Metadata automatically records artifacts, executions, and their lineage for Vertex AI Pipelines runs, enabling reproducibility, auditing, and comparison of models and datasets across experiments. Option A is not a best practice because Cloud Functions are event-driven, short-lived, and not designed to run heavy or stateful ML pipeline steps; Vertex AI Pipelines components (or Vertex AI Training jobs) should execute those steps. Option B is wrong because hardcoding parameters in component definitions reduces reusability and makes experimentation and CI/CD harder; parameters should be passed at pipeline submission time via PipelineJob runtime parameters. Option D is wrong because training and serving often require different compute profiles (for example, GPUs for training versus optimized CPUs or different accelerators for serving), and consistency is achieved through containerized artifacts and pinned dependencies, not by forcing identical compute environments.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Cloud Functions to execute individual pipeline steps
Why it's wrong here
Cloud Functions cannot orchestrate multi-step ML workflows with artefact lineage, caching or conditional execution; Vertex AI Pipelines uses Kubeflow-based components with a managed orchestrator tracking inputs and outputs. Cloud Functions suit lightweight event-driven glue, such as triggering a pipeline when a file lands in Cloud Storage, not executing pipeline steps themselves.
- ✗
Hardcode pipeline parameters in the component definitions
Why it's wrong here
Hardcoding parameters into component definitions prevents reuse across environments and forces pipeline recompilation for every change; Vertex AI expects runtime parameters passed via PipelineJob to let one compiled template serve dev, test and prod. It tempts because quick prototypes hardcode values, but that suits only throwaway notebooks, not governed pipelines.
- ✓
Use custom container components to encapsulate reusable logic
Why this is correct
Custom container components package code, dependencies and runtime into a portable image, letting teams reuse identical logic across pipelines without rewriting or re-resolving environments. This encapsulation satisfies the reusability best-practise requirement for Vertex AI Pipelines components.
- ✗
Always use the same compute environment for training and serving to ensure consistency
Why it's wrong here
Serving requires low-latency, autoscaling inference hardware, while training benefits from accelerators and larger batches; forcing one environment sacrifices one workload's needs. It is tempting because identical environments genuinely suit reproducibility when debugging training-serving skew in a small, fixed model.
- ✓
Leverage Vertex ML Metadata to track artifact lineage
Why this is correct
Vertex ML Metadata records parameters, metrics and artifacts across pipeline runs, giving automatic lineage tracking. This satisfies the best-practice requirement for reproducible, auditable ML pipelines by letting you trace which dataset and container produced each model.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on PMLE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. An ML team is using Vertex AI Pipelines to automate model training and deployment. They want to reuse components across multiple pipelines. What is the best practice for managing component code?
medium- A.Define components inline in the pipeline definition
- B.Embed component code in Cloud Composer DAGs
- C.Copy the component definitions into each pipeline's YAML file
- D.Use Cloud Functions to define components
- ✓ E.Store components as container images in Artifact Registry and reference them from pipelines
Why E: Vertex AI Pipelines natively supports reusable components by packaging them as container images stored in Artifact Registry. This allows teams to version, share, and reference components across multiple pipelines without duplicating code, ensuring consistency and reducing maintenance overhead. Container images encapsulate the component's runtime environment and logic, making them portable and independently deployable.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.