Courseiva

PMLE Automating and Orchestrating ML Pipelines Practice Question

A machine learning engineer is building a pipeline with Vertex AI Pipelines and wants to pass a large dataset between components without copying it to the container's memory. What is the best practice for passing data between pipeline components?

⚠ Common exam trap

A common mistake in this exam is to think that large data must be passed as in-memory artifacts or serialized parameters, when the correct Vertex AI Pipelines pattern is to pass a Cloud Storage URI and let components read data lazily from GCS.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Cloud Storage URIs (gs://) to point to the data location.

Vertex AI Pipelines natively supports passing Cloud Storage URIs (gs://) as artifact references between components, allowing components to read the dataset directly from GCS without copying it into container memory. This avoids memory limits and enables efficient handling of large datasets by leveraging GCS's scalable object storage.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Mount an NFS volume to all containers and share data via the filesystem.

    Why it's wrong here

    Vertex AI Pipelines passes data between components by writing artifacts to Cloud Storage and passing URIs, so an NFS mount adds infrastructure that the platform does not manage or track as lineage. NFS suits shared filesystem access across VMs or on-premises clusters, not managed pipeline artifact passing.

  • ✓

    Use Cloud Storage URIs (gs://) to point to the data location.

    Why this is correct

    Passing gs:// URIs means components exchange only small string references, so the large dataset stays in Cloud Storage and each container reads it directly. This satisfies the constraint of avoiding copying data into container memory between pipeline components.

  • ✗

    Serialize the dataset to JSON and include it as a pipeline parameter.

    Why it's wrong here

    Pipeline parameters are stored in the pipeline spec and passed as scalar values, so embedding a large dataset as JSON bloats the spec and forces serialisation through the orchestrator. It is tempting because parameters are the standard mechanism for small configuration values, but Vertex AI Pipelines uses artifact outputs and inputs for datasets, passing only references between components.

  • ✗

    Use the importer component to load the data into the pipeline as an in-memory artifact.

    Why it's wrong here

    The importer component registers an existing artifact's URI into the pipeline metadata; it does not load data into memory, and in-memory transfer is exactly what the stem forbids. Importer suits bringing externally produced artifacts under pipeline tracking, not moving large datasets between components.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.