Courseiva
hardMultiple Select

PDE Practice Question: Which TWO are common causes of prediction bias in…

Which TWO are common causes of prediction bias in a deployed machine learning model in production?

⚠ Common exam trap

Google Cloud often tests the distinction between training-time issues (like overfitting) and production-time causes (like data drift and training-serving skew), so candidates mistakenly select overfitting as a production bias cause.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Data drift between training and serving data distributions.

Option B is correct because data drift occurs when the statistical distribution of the input features or target changes between the training environment and the live serving environment, causing the model's learned mappings to become stale and its predictions to be systematically biased. Option E is correct because training-serving skew arises when feature engineering logic differs between the training pipeline and the production inference path (for example, different imputation, normalization, or aggregation code), so the model receives inputs at serving time that do not match what it learned from, producing biased predictions. Option A is not a cause of bias — high accuracy is generally desirable and does not by itself indicate or produce prediction bias. Option C is not correct in this context because overfitting primarily harms generalization and variance rather than being a canonical cause of prediction bias in production, and it is a training-time issue rather than the drift/skew mechanisms described. Option D is not correct because low latency is a performance characteristic of the serving system and has no direct causal relationship with prediction bias.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Model accuracy is too high.

    Why it's wrong here

    High accuracy does not cause prediction bias; bias arises from unrepresentative sampling, skewed labels or proxy features, and can persist even when aggregate accuracy looks strong. It is tempting because accuracy is the headline metric teams monitor, yet a model can be accurate overall while systematically misclassifying an underrepresented subgroup.

  • ✓

    Data drift between training and serving data distributions.

    Why this is correct

    Data drift means serving inputs diverge statistically from the training distribution, so learned relationships no longer hold and predictions skew systematically. This satisfies the bias-cause constraint because the model extrapolates beyond its training support, producing skewed outputs without any code change.

  • ✗

    Model is overfitted to training data.

    Why it's wrong here

    Overfitting is a variance problem: the model memorises training noise and generalises poorly to new data, which is distinct from bias, a systematic skew from unrepresentative data or flawed labels. It is tempting because both degrade production performance, but overfitting shows as unstable, high-variance errors rather than consistent directional skew.

  • ✗

    Low latency predictions.

    Why it's wrong here

    Low latency is an inference-serving performance characteristic and has no causal link to prediction bias, which stems from data sampling, labelling or feature construction. It is tempting because latency matters in production deployment, but optimising response time neither introduces nor removes systematic skew in predicted outcomes.

  • ✓

    Training-serving skew due to differences in feature engineering.

    Why this is correct

    Training-serving skew arises when feature engineering logic differs between the training pipeline and the serving path, so the model receives inputs with a different distribution than it learned. This mismatch directly satisfies the stem's constraint of a deployed production model, producing systematically biased predictions.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.