mediumMultiple Select
PMLE Practice Question: Which TWO metrics should you monitor to detect…
Which TWO metrics should you monitor to detect data drift in a batch prediction pipeline?
⚠ Common exam trap
Google Cloud often tests the distinction between monitoring for data drift (input distribution changes) versus monitoring for model performance degradation (accuracy), leading candidates to incorrectly select accuracy as a drift metric when it is actually a downstream effect.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Feature distribution drift (e.g., KS test)
Feature distribution drift (C) is correct because data drift is detected by statistically comparing the distribution of each input feature in recent batches against the training/reference distribution, typically using tests like the Kolmogorov-Smirnov (KS) test or Population Stability Index (PSI). Prediction distribution drift (D) is correct because shifts in the distribution of model outputs (e.g., changing class proportions or score histograms) are a direct, label-free signal that the input data relationship has changed in a batch pipeline. Model accuracy on recent labeled data (A) measures performance degradation, not drift itself, and labels are usually unavailable or delayed in batch prediction. Model prediction latency (B) is an operational performance metric unrelated to distributional change, and training data size (E) is a static dataset property, not a monitoring signal for drift.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Model accuracy on recent labeled data
Why it's wrong here
Accuracy needs ground-truth labels, which arrive late or never in a live batch pipeline, so drift goes undetected until scoring is already wrong. It is tempting because accuracy is the definitive quality measure, and it would be the right metric in an offline evaluation with a labelled holdout set.
- ✗
Model prediction latency
Why it's wrong here
Latency measures serving speed, not whether input feature distributions have shifted away from training data. It is tempting because latency monitoring is standard for pipeline health, and it would be correct when diagnosing throughput or infrastructure performance rather than drift.
- ✓
Feature distribution drift (e.g., KS test)
Why this is correct
Feature distribution drift compares each input feature's distribution between training and serving data using tests such as Kolmogorov-Smirnov, directly detecting changes in the input data itself, which is the definition of data drift in a batch prediction pipeline.
- ✓
Prediction distribution drift
Why this is correct
Prediction distribution drift tracks shifts in the model's output distribution over time, revealing drift that may not appear in individual features; monitoring it alongside feature drift satisfies the requirement to detect data drift in batch prediction.
- ✗
Training data size
Why it's wrong here
Training data size is fixed at build time and does not change as new batches arrive, so it cannot signal drift in incoming data. It is tempting because dataset volume feels like a data-quality indicator, and it would be relevant when auditing whether retraining has enough historical samples.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.