Courseiva

PMLE Automating and Orchestrating ML Pipelines Practice Question

You are using KFP SDK v2 to define a pipeline. You need to pass a large dataset between components. What is the best practice for passing data?

⚠ Common exam trap

Google often tests the misconception that temporary directories are shared between containers in a pod, but in KFP each component runs in its own container with isolated storage, making Cloud Storage the correct choice for durable, cross-component data sharing.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Write the data to Cloud Storage and pass the GCS URI as an artifact.

In KFP SDK v2, passing large datasets between components is best done by writing the data to Cloud Storage and passing the GCS URI as an artifact. This approach leverages KFP's built-in artifact tracking, ensures data persistence across container restarts, and avoids memory or disk limitations of ephemeral containers. The artifact is automatically serialized and passed as an input/output parameter, enabling efficient, scalable data exchange.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use the component's temporary directory to share data between containers.

    Why it's wrong here

    A component's temporary directory is container-local and destroyed when that container exits, so a downstream component cannot read it; KFP v2 requires data to be written to an artefact location the orchestrator passes on. Temp directories suit scratch files used only within a single component's own execution.

  • ✗

    Pass the data as a serialized Python object in memory.

    Why it's wrong here

    Serialising a large dataset into memory forces every component to materialise it, and KFP's in-memory passing is capped by the orchestrator's artefact size limits, so large payloads fail or exhaust memory. It suits small configuration values or scalars, where writing artefacts to storage would add needless I/O.

  • ✓

    Write the data to Cloud Storage and pass the GCS URI as an artifact.

    Why this is correct

    Writing data to Cloud Storage and passing the GCS URI as an artifact avoids embedding large payloads in pipeline metadata, satisfying the large-dataset constraint. KFP passes lightweight artifact references between components rather than the data itself.

  • ✗

    Store the data in a BigQuery table and pass the table reference.

    Why it's wrong here

    Passing a BigQuery table reference moves only a small string between components, so the dataset is not transferred through the pipeline at all. KFP artefacts are designed for large payloads, with the SDK handling storage and lineage. A table reference suits data already resident and queried in BigQuery, not arbitrary component outputs.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

2 more ways this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company is using Vertex AI Pipelines to automate model retraining. They have a component that creates a BigQuery table with training data. To ensure idempotency, the component should check if the table already exists and recreate it if necessary. What is the best practice for passing data between pipeline components?

medium
  • A.Pass data in-memory as Python objects between components.
  • B.Use BigQuery table names as component outputs and inputs.
  • C.Use Cloud SQL to store intermediate results and pass connection strings.
  • ✓ D.Store data as artifacts in Cloud Storage and pass the GCS URI between components.

Why D: Vertex AI Pipelines is designed to pass data between components via Cloud Storage artifacts. By storing the BigQuery table metadata or training data as a file in Cloud Storage and passing the GCS URI as an artifact, the pipeline ensures idempotency and decouples components. This approach aligns with Kubeflow Pipelines' artifact-based I/O model, where each component's outputs are materialized as URIs rather than in-memory objects.

Variation 2. A machine learning engineer is building a pipeline with Vertex AI Pipelines and wants to pass a large dataset between components without copying it to the container's memory. What is the best practice for passing data between pipeline components?

easy
  • A.Mount an NFS volume to all containers and share data via the filesystem.
  • ✓ B.Use Cloud Storage URIs (gs://) to point to the data location.
  • C.Serialize the dataset to JSON and include it as a pipeline parameter.
  • D.Use the importer component to load the data into the pipeline as an in-memory artifact.

Why B: Vertex AI Pipelines natively supports passing Cloud Storage URIs (gs://) as artifact references between components, allowing components to read the dataset directly from GCS without copying it into container memory. This avoids memory limits and enables efficient handling of large datasets by leveraging GCS's scalable object storage.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.