PMLE Automating and Orchestrating ML Pipelines Practice Question
You are tasked with building a robust ML pipeline that must be idempotent and handle data skew between training and serving. Which three practices should you implement?
⚠ Common exam trap
PMLE often tests the difference between reproducibility (same seed) and idempotency (safe reruns) — candidates pick the seed option because it 'sounds like best practice' but it does not address pipeline idempotency or skew.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Store intermediate data in Cloud Storage with unique run IDs.
Option A is correct because writing intermediate artifacts to Cloud Storage under a unique run ID (e.g., gs://bucket/run_id/...) isolates each pipeline execution, so reruns don't overwrite or collide with prior outputs — a key requirement for idempotency. Option C is correct because comparing training feature distributions against serving feature distributions (e.g., via statistics like mean, variance, or histogram distance) is the standard way to detect training/serving skew and trigger remediation. Option E is correct because deterministic components that yield identical outputs for identical inputs make reruns safe and reproducible, which is the essence of an idempotent pipeline. Option B is wrong because passing large datasets as serialized in-memory objects is fragile, memory-bound, and non-idempotent; components should exchange data via durable storage references instead. Option D is wrong because a fixed random seed only aids reproducibility of stochastic steps and does nothing to guarantee idempotency or address training/serving skew.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Store intermediate data in Cloud Storage with unique run IDs.
Why this is correct
Unique run IDs in Cloud Storage paths make each pipeline execution write to distinct locations, so reruns neither overwrite nor reuse stale intermediates. This directly satisfies the idempotency requirement, since identical inputs produce isolated, reproducible outputs rather than colliding with prior runs.
- ✗
Pass large datasets between components as serialized in-memory objects.
Why it's wrong here
Passing serialised in-memory objects between components couples steps and breaks idempotency, since reruns cannot reliably reconstruct state, and it does nothing for train/serve skew. In-memory passing suits single-process prototyping; production pipelines should persist artefacts to durable storage with versioned references.
- ✓
Monitor feature distributions in training data vs. serving data to detect skew.
Why this is correct
Comparing feature distributions between training and serving data detects training-serving skew, the second constraint in the stem. Monitoring statistical drift in production inputs against the training baseline surfaces mismatches early, letting you retrain or fix feature engineering before predictions degrade.
- ✗
Use the same random seed for every run to ensure reproducibility.
Why it's wrong here
A fixed random seed gives reproducible runs, but idempotency means re-running a pipeline yields identical outputs regardless of prior state, and skew requires consistent train/serve transformations. Seeding addresses neither; it belongs in experiment tracking or model reproducibility work, not pipeline idempotency.
- ✓
Ensure each component produces deterministic outputs given the same inputs.
Why this is correct
Deterministic outputs from identical inputs are the core property of idempotency: rerunning a component yields the same result without side effects. This satisfies the idempotency constraint directly, ensuring pipeline retries and backfills produce consistent artifacts rather than divergent ones.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.