PMLE Automating and Orchestrating ML Pipelines Practice Question
A machine learning engineer needs to pass a large dataset between two components in a Vertex AI pipeline. What is the recommended way to pass this data?
⚠ Common exam trap
Test-takers frequently assume passing a Cloud Storage bucket name (Option D) is sufficient, but they miss that artifacts provide automatic metadata tracking, type safety, and integration with Vertex AI's lineage system, which is required for production ML pipelines.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Store the dataset as a Dataset artifact and pass the artifact between components.
In Vertex AI Pipelines, the recommended way to pass large datasets between components is to use a `Dataset` artifact. Artifacts are metadata references that point to the underlying data stored in Cloud Storage, enabling efficient, scalable, and type-safe data passing without serialization overhead or size limits. This approach leverages the Kubeflow Pipelines SDK's artifact tracking, which automatically handles lineage and versioning.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Store the dataset as a Dataset artifact and pass the artifact between components.
Why this is correct
Dataset artifacts are references to Cloud Storage or BigQuery URIs, so components exchange a lightweight metadata pointer rather than the bytes themselves. This satisfies the stem's large-dataset constraint, since passing raw data through component inputs would exceed the metadata limits.
- ✗
Write the dataset to a temporary BigQuery table and pass the table name.
Why it's wrong here
Writing a large dataset to a temporary BigQuery table and passing the table name introduces unnecessary latency and cost from BigQuery storage and query overhead, whereas Vertex AI pipelines natively support passing large data objects via cloud storage URIs without an intermediate database. This option is tempting because BigQuery is often used for large-scale analytics and temporary tables are a common pattern for sharing data between SQL-based jobs, making it a correct choice when the downstream component requires SQL querying rather than direct file access.
- ✗
Serialize the dataset to a string and pass it as a pipeline parameter.
Why it's wrong here
Pipeline parameters are stored in the pipeline spec and passed through the orchestrator, so large datasets exceed size limits and bloat metadata. It is tempting because parameters are the standard way to pass small values between components, and would be correct for a short configuration string or scalar hyperparameter.
- ✗
Use a Cloud Storage bucket and pass the bucket name as a parameter.
Why it's wrong here
Passing a bucket name as a parameter only transmits a string reference, so Vertex AI cannot track the dataset as an artefact or lineage between components. Cloud Storage suits persisting large data outside pipeline metadata, but the recommended mechanism is a Dataset artefact output that downstream components consume directly.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.