Courseiva

PMLE Automating and Orchestrating ML Pipelines Practice Question

A pipeline includes a component that produces a model artifact. The team wants to automatically detect skew between the training data distribution and the serving data distribution. Which three best practices should they implement? (Choose three.)

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Compare statistics using a dedicated component and alert on threshold exceedance

Option A is correct because skew detection requires a dedicated comparison component that evaluates training versus serving statistics and raises an alert when a configured threshold (e.g., a divergence metric like KL divergence or PSI) is exceeded. Option C is correct because you must compute statistics on the serving data itself, typically via a component that ingests the live inference data and emits a statistics artifact for comparison. Option E is correct because the training data statistics should be persisted and passed as a Dataset artifact so the comparison component can consume them reproducibly across pipeline runs. Option B is wrong because in-memory data passing does not persist artifacts and is unsuitable for cross-component, cross-run statistical comparison. Option D is wrong because disabling caching does not improve statistical correctness and only adds unnecessary recomputation cost.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Compare statistics using a dedicated component and alert on threshold exceedance

    Why this is correct

    A dedicated comparison component computes the statistical distance between training and serving distributions and raises alerts when thresholds are exceeded, satisfying the automatic detection requirement. Isolating this logic keeps skew monitoring reproducible and lets thresholds be tuned without retraining the model.

  • ✗

    Use in-memory data passing for efficiency

    Why it's wrong here

    In-memory passing moves data between steps of one pipeline execution; it cannot compare training data against live serving traffic, which arrives outside that execution. It is tempting because it avoids S3 round-trips and speeds up pipelines, and it would be correct for passing intermediate results between steps in a single run.

  • ✓

    Compute serving data statistics using a component

    Why this is correct

    Serving data statistics must be computed before any comparison can occur, since skew is defined as divergence between training and live distributions. A component that calculates these statistics on recent serving traffic supplies the input the detection step requires.

  • ✗

    Disable caching to ensure fresh statistics

    Why it's wrong here

    Caching affects whether a component re-executes, not whether training and serving distributions are compared; skew detection needs statistics computed from both datasets. Disabling caching is tempting when stale pipeline outputs cause confusion, and it would be the right move for forcing recomputation of an unchanged component's outputs.

  • ✓

    Pass training data statistics as a Dataset artifact

    Why this is correct

    Passing training data statistics as a Dataset artifact persists the baseline distribution across pipeline runs, so serving statistics can be compared against a stable reference. This satisfies the skew-detection requirement by making the training distribution available as a versioned input.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.