Courseiva

PMLE Automating and Orchestrating ML Pipelines Practice Question

Which of the following is a best practice when designing idempotent pipeline components in Vertex AI?

⚠ Common exam trap

The Google PMLE exam often tests the misconception that idempotency is about avoiding duplication of output names or using timestamps for uniqueness, when in fact idempotency requires that repeated executions produce the same result without side effects, which is achieved by using immutable, deterministic storage like Cloud Storage URIs rather than mutable state or time-dependent writes.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Pass data through Cloud Storage URIs rather than in-memory.

Passing data through Cloud Storage URIs ensures that component outputs are stored persistently and can be retrieved by downstream components, even if the original component instance is terminated or scaled down. This aligns with the principle of idempotency because the same input will always produce the same output stored at the same URI, and re-running the component will not cause side effects or data loss. In contrast, in-memory data is ephemeral and tied to a specific runtime instance, breaking idempotency across retries or parallel executions.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use global variables to share state between components.

    Why it's wrong here

    Global variables lead to non-determinism and break idempotency.

  • ✓

    Pass data through Cloud Storage URIs rather than in-memory.

    Why this is correct

    Cloud Storage URIs persist data outside component memory, so a retried or rerun component reads the same immutable input and writes the same output, satisfying idempotency. In-memory passing loses state between container executions and cannot guarantee reproducible reruns.

  • ✗

    Write component outputs to a database with timestamps.

    Why it's wrong here

    Timestamps cause non-determinism.

  • ✗

    Use the same output name for all runs to avoid duplication.

    Why it's wrong here

    Output names should be unique per run.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. You are tasked with building a robust ML pipeline that must be idempotent and handle data skew between training and serving. Which three practices should you implement?

hard
  • ✓ A.Store intermediate data in Cloud Storage with unique run IDs.
  • B.Pass large datasets between components as serialized in-memory objects.
  • ✓ C.Monitor feature distributions in training data vs. serving data to detect skew.
  • D.Use the same random seed for every run to ensure reproducibility.
  • ✓ E.Ensure each component produces deterministic outputs given the same inputs.

Why A: Option A is correct because writing intermediate artifacts to Cloud Storage under a unique run ID (e.g., gs://bucket/run_id/...) isolates each pipeline execution, so reruns don't overwrite or collide with prior outputs — a key requirement for idempotency. Option C is correct because comparing training feature distributions against serving feature distributions (e.g., via statistics like mean, variance, or histogram distance) is the standard way to detect training/serving skew and trigger remediation. Option E is correct because deterministic components that yield identical outputs for identical inputs make reruns safe and reproducible, which is the essence of an idempotent pipeline. Option B is wrong because passing large datasets as serialized in-memory objects is fragile, memory-bound, and non-idempotent; components should exchange data via durable storage references instead. Option D is wrong because a fixed random seed only aids reproducibility of stochastic steps and does nothing to guarantee idempotency or address training/serving skew.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.