PMLE Collaborating to manage data and models Practice Question
A data engineering team uses Dataflow for preprocessing and wants to integrate with Vertex AI Pipelines. They need to pass the preprocessed data location to the training step. What is the best practice?
⚠ Common exam trap
Candidates often confuse metadata services (Data Catalog) or messaging systems (Pub/Sub) with pipeline parameter passing, overlooking that Vertex AI Pipelines uses Kubeflow Pipelines' built-in component I/O for deterministic, graph-based data flow.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use PipelineParam to pass the output path
PipelineParam is the native mechanism in Vertex AI Pipelines (Kubeflow Pipelines SDK) to pass runtime outputs—such as a Cloud Storage path—between components. It creates a dependency graph that ensures the training step receives the exact output path from the preprocessing step, enabling dynamic, reproducible pipelines without hardcoding.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Store the path in Data Catalog
Why it's wrong here
Data Catalog stores metadata for discovery and governance, not runtime parameter values passed between pipeline steps. It is tempting because it is a genuine GCP service, and would be correct for cataloguing and searching datasets rather than passing execution parameters.
- ✗
Use Cloud Pub/Sub
Why it's wrong here
Cloud Pub/Sub transports streaming messages between producers and subscribers; it does not record a pipeline artefact location for a downstream training step to consume. It is tempting because Pub/Sub genuinely decouples Dataflow output from consumers in event-driven architectures, where asynchronous notification of new data arrivals is the actual requirement.
- ✓
Use PipelineParam to pass the output path
Why this is correct
PipelineParam passes the Dataflow output path as a runtime parameter between pipeline components, satisfying the requirement to hand the preprocessed data location to the training step. Vertex AI Pipelines resolves this value at execution, decoupling the training component from hard-coded paths and enabling reuse across runs.
- ✗
Write the output to a fixed Cloud Storage path and hardcode it in the pipeline
Why it's wrong here
Hardcoding a fixed path prevents the training step from consuming each run's actual output, breaking pipeline reproducibility. It is tempting because fixed paths are simple for one-off jobs, and would suit a static dataset that never changes between executions.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.