Courseiva
easyMultiple Select

PMLE Practice Question: An ML team wants to monitor their recommendation…

An ML team wants to monitor their recommendation model for fairness. Which TWO metrics should they track to detect potential bias? (Select TWO.)

⚠ Common exam trap

Google Cloud often tests the misconception that overall accuracy or group-specific recall alone is sufficient for fairness monitoring, when in fact comparative metrics across groups are required to detect bias.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Pair-wise fairness metrics such as equal opportunity difference.

Option A is correct because pair-wise fairness metrics such as equal opportunity difference directly compare model performance (e.g., true positive rates) between groups defined by a sensitive attribute, which is the standard way to quantify disparate impact and bias. Option E is correct because examining the prediction distribution (e.g., top-K recommendations) across sensitive attribute groups reveals whether the model systematically under- or over-represents certain groups in its outputs, a key bias signal for recommender systems. Option B is not sufficient because tracking recall for only the minority group provides no comparison baseline, so it cannot by itself detect bias relative to other groups. Option C is not appropriate because overall accuracy on the test set can mask large disparities between groups and is not a fairness metric. Option D is not appropriate because average prediction confidence per request measures model certainty, not equitable treatment or outcomes across sensitive groups.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Pair-wise fairness metrics such as equal opportunity difference.

    Why this is correct

    Pair-wise fairness metrics such as equal opportunity difference compare true positive rates between sensitive attribute groups, directly quantifying disparate error rates. This satisfies the requirement to detect bias by revealing whether one group receives systematically worse outcomes than another.

  • ✗

    Recall for the minority group only.

    Why it's wrong here

    Recall for the minority group alone gives no comparison baseline, so a low value cannot be attributed to bias rather than genuine class difficulty. It is tempting because minority-group performance is the usual place bias surfaces, but fairness requires the same metric computed for the majority group to expose a disparity.

  • ✗

    Overall accuracy on the test set.

    Why it's wrong here

    Overall accuracy aggregates all users into one figure, so it cannot reveal disparate error rates between demographic groups. It is tempting because accuracy is the default quality metric and a drop would suggest problems, but fairness monitoring requires disaggregated metrics such as true positive rate or demographic parity across groups.

  • ✗

    Average prediction confidence per request.

    Why it's wrong here

    Average prediction confidence measures model certainty, not outcomes across groups, so identical confidence can coexist with biased predictions. It is tempting because confidence monitoring is standard for detecting drift and calibration issues, but bias detection requires comparing outcome metrics between demographic slices rather than aggregate confidence.

  • ✓

    Prediction distribution (e.g., top-K recommendations) across different sensitive attribute groups.

    Why this is correct

    Comparing prediction distributions, such as top-K recommendation sets, across sensitive attribute groups reveals whether the model systematically recommends different items to different groups. This satisfies the fairness monitoring requirement by exposing representational disparity in outputs rather than only error-rate gaps.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.