You are tasked with building a robust ML pipeline that must be idempotent and handle data skew between training and serving. Which three practices should you implement?
Unique run IDs in Cloud Storage paths make each pipeline execution write to distinct locations, so reruns neither overwrite nor reuse stale intermediates. This directly satisfies the idempotency requirement, since identical inputs produce isolated, reproducible outputs rather than colliding with prior runs.
Why this answer
Option A is correct because writing intermediate artifacts to Cloud Storage under a unique run ID (e.g., gs://bucket/run_id/...) isolates each pipeline execution, so reruns don't overwrite or collide with prior outputs — a key requirement for idempotency. Option C is correct because comparing training feature distributions against serving feature distributions (e.g., via statistics like mean, variance, or histogram distance) is the standard way to detect training/serving skew and trigger remediation. Option E is correct because deterministic components that yield identical outputs for identical inputs make reruns safe and reproducible, which is the essence of an idempotent pipeline.
Option B is wrong because passing large datasets as serialized in-memory objects is fragile, memory-bound, and non-idempotent; components should exchange data via durable storage references instead. Option D is wrong because a fixed random seed only aids reproducibility of stochastic steps and does nothing to guarantee idempotency or address training/serving skew.
Exam trap
PMLE often tests the difference between reproducibility (same seed) and idempotency (safe reruns) — candidates pick the seed option because it 'sounds like best practice' but it does not address pipeline idempotency or skew.