Courseiva
hardMultiple Select

PMLE Practice Question: An e-commerce company uses a recommendation model…

An e-commerce company uses a recommendation model that suggests products based on user browsing history. The model was trained on data from the past year and has high accuracy on the test set. However, after deployment, the click-through rate (CTR) on recommendations is much lower than expected. Which three steps should the data scientist take to diagnose and improve the model? (Choose THREE)

⚠ Common exam trap

Google Cloud often tests the misconception that high offline accuracy guarantees online success, ignoring that offline metrics can be misleading due to distribution shift, feedback loops, or mismatched optimization objectives (e.g., accuracy vs. CTR).

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set up an A/B experiment comparing the model's recommendations against a baseline

Option B is correct because an A/B experiment comparing the deployed model against a baseline (e.g., current production recommender or popularity-based recommendations) directly measures real-world CTR impact and isolates whether the model itself is underperforming versus other factors like UI placement or latency. Option C is correct because retraining on the most recent three months of data addresses temporal concept drift — user browsing behavior and product trends change quickly in e-commerce, so a model trained on a year-old distribution can have stale item embeddings and relevance signals even with high offline accuracy. Option D is correct because comparing the distribution of predictions (and input feature distributions) against the training set detects data drift and covariate shift, which commonly explain why offline test accuracy fails to translate into online CTR. Option A is not the best step because offline holdout evaluation was already performed and showed high accuracy, so repeating it won't reveal the deployment-time mismatch. Option E is not appropriate because adding two-year-old data would likely worsen drift by reinforcing outdated patterns rather than capturing recent trends.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Run offline evaluation on a holdout dataset to confirm accuracy

    Why it's wrong here

    Offline accuracy is already high; the issue is online performance, which requires online metrics.

  • ✓

    Set up an A/B experiment comparing the model's recommendations against a baseline

    Why this is correct

    An A/B experiment isolates whether the low CTR stems from the model itself or from external factors such as placement or latency, by comparing live recommendations against a baseline under identical conditions. This directly tests the deployment gap the stem describes, where offline accuracy failed to translate into online engagement.

  • ✓

    Retrain the model on the most recent three months of data to capture recent trends

    Why this is correct

    Training on the past year means the model cannot capture recent shifts in browsing behaviour, so retraining on the latest three months addresses temporal drift in user preferences. This directly targets the mismatch between stale training data and current online behaviour that suppresses CTR.

  • ✓

    Check the distribution of predictions versus the training set to detect drift

    Why this is correct

    Comparing the prediction distribution against the training set reveals covariate or concept drift, showing whether live inputs differ from what the model learned. This diagnoses the train-serving skew behind the CTR shortfall without altering the model prematurely.

  • ✗

    Increase the training dataset size by including data from two years ago

    Why it's wrong here

    Older data may be less relevant and could dilute recent patterns.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.