Courseiva

PMLE Automating and Orchestrating ML Pipelines Practice Question

A team is using Vertex AI Pipelines to orchestrate a multi-step ML workflow. They need to pass a large dataset (several terabytes) between two components: a preprocessing component and a training component. The preprocessing component outputs a preprocessed dataset that the training component consumes. The team wants to minimize data transfer time and cost. What is the most efficient way to pass the data between these components?

⚠ Common exam trap

The trap here is thinking that the pipeline must pass the actual data between components, when it should only pass references to data stored in Cloud Storage.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Write the preprocessed data to a Cloud Storage bucket and pass the bucket URI as an output parameter to the training component.

For large datasets, the best practice is to write data to Cloud Storage and pass the URI as a parameter. This avoids moving data through the pipeline's control plane and leverages Cloud Storage's scalability and high throughput. The training component can then read directly from the bucket.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Mount a Filestore instance as a shared volume to both components and write the data there.

    Why it's wrong here

    While Filestore can provide shared storage, it is not the native way to pass data between Vertex AI Pipeline components. Using Cloud Storage is more cost-effective and scalable for large datasets. Filestore requires additional setup and may not offer the same level of integration.

  • ✗

    Use a `Dataset` artifact to pass the data, and configure the artifact to store the data in the pipeline's metadata store.

    Why it's wrong here

    Vertex AI Pipelines metadata store is not designed to store large datasets. Artifacts in the metadata store are references to data, not the data itself. Storing terabytes of data in the metadata store is impractical and would cause performance issues.

  • ✗

    Have the preprocessing component output the data as a base64-encoded string parameter, and pass it to the training component.

    Why it's wrong here

    Base64-encoding a multi-terabyte dataset into a parameter is infeasible due to size limits on parameters. Vertex AI Pipelines parameters are meant for small values, not large datasets. This approach would also incur significant overhead and likely fail.

  • ✓

    Write the preprocessed data to a Cloud Storage bucket and pass the bucket URI as an output parameter to the training component.

    Why this is correct

    Passing a Cloud Storage URI as an output parameter is the standard and most efficient way to share large datasets between components in Vertex AI Pipelines. The training component can directly read from Cloud Storage, avoiding unnecessary data movement and leveraging Google's high-bandwidth network.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.