Courseiva

PMLE · domain

Monitoring ML Solutions

This domain covers keeping deployed models healthy on Vertex AI: detecting training/serving skew and drift, monitoring endpoints, logging predictions, and measuring fairness across subgroups. Questions are scenario-based, asking you to pick the correct Vertex AI or Google Cloud service to monitor, alert on, or react to model degradation.

63 questions17 easy24 medium22 hard

Focused practice

Practice Monitoring ML Solutions questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Monitoring ML Solutions

Be able to select the right Vertex AI service for each monitoring need: Model Monitoring for drift and skew, Model Evaluation for fairness and accuracy, and request-response logging for auditing. The key skill is knowing which signal each service produces and how alerts trigger retraining.

Choosing Vertex AI Model Monitoring to detect feature drift and training-serving skew on Endpoints

Using Cloud Logging and BigQuery export to capture Endpoint prediction request/response payloads

Configuring Cloud Monitoring alerting policies on drift metrics to trigger retraining pipelines

Evaluating model fairness with subgroup performance metrics via Vertex AI Model Evaluation

Watch out for

Common Monitoring ML Solutions exam traps

  • ▸Confusing Model Monitoring with Model Evaluation: drift detection is monitoring, while accuracy/fairness against ground truth labels is evaluation
  • ▸Assuming prediction logging to BigQuery is automatic; you must explicitly enable request-response logging on the Endpoint
  • ▸Trying to trigger retraining directly from a Monitoring alert instead of routing it through Pub/Sub or Cloud Functions to a pipeline

Question index

All Monitoring ML Solutions questions (63)

Click any question to see the full explanation, or start a practice session above.

1

A data scientist has deployed a model on Vertex AI Endpoints and wants to monitor the model's predictions for any drift over time. Which Vertex AI service should they use?

Easy
2

A team has deployed a model to a Vertex AI Endpoint and wants to monitor the model's performance in production. They need to track the number of prediction requests and the average latency. Which Google Cloud service should they use to collect and visualize these metrics?

Easy
3

An MLOps team has deployed a model on Vertex AI Endpoints and wants to monitor for skew between training and serving data distributions. Which Vertex AI service should they use?

Easy
4

An ML engineer wants to monitor the performance of a Vertex AI Endpoint. Which TWO metrics are available in Cloud Monitoring for Vertex AI Endpoints? (Choose 2)

Easy
5

An ML engineer manages a Vertex AI Endpoint serving a recommendation model. The team wants to detect when the distribution of a specific numerical feature, average session duration, shifts significantly from its training distribution. They have configured Vertex AI Model Monitoring with a training dataset baseline and a monitoring frequency of one hour. After a week, no drift alerts have fired even though the feature's daily mean has visibly moved. What is the most likely cause?

Hard
6

A data scientist has deployed a model with Vertex AI Endpoints and enabled request/response logging to BigQuery. They want to compute a confusion matrix over time to monitor model quality. What should they do?

Medium
7

A company wants to track the cost of their Vertex AI prediction endpoint. They use a custom machine type with 1 n1-standard-4 (4 vCPU, 15 GB memory) and 1 NVIDIA T4 GPU. The endpoint is configured for automatic scaling with min=1, max=5 replicas. Which cost monitoring approach should they use?

Medium
8

A fraud detection model is deployed to a Vertex AI Endpoint and configured with Vertex AI Model Monitoring for feature drift. The team wants the drift monitor to compare live production traffic against the exact statistics captured from the training dataset, so that alerts reflect deviation from the model's original data distribution rather than from recent traffic. Which configuration should they use?

Medium
9

An ML team has set up automated retraining triggered by Cloud Monitoring alerts. When a feature drift alert fires, a Cloud Function publishes to Pub/Sub, which triggers a Vertex AI Pipeline. However, the retraining pipeline is failing because the training data is not updated. What is the most likely cause?

Hard
10

A credit-risk team runs a tabular model on a Vertex AI Endpoint. They configured Vertex AI Model Monitoring with a training dataset and skew detection using the default threshold. After a week, they receive alerts that many features have high training-serving skew, but the model's business metrics (approval rate, default rate) are unchanged. They suspect the alerts are false positives due to a recent change in an upstream data pipeline that shifted feature distributions. What should they do to reduce these false alerts while still monitoring for real skew?

Medium
11

You manage a Vertex AI Model Monitoring job on an Endpoint that serves an image classification model. The monitoring job reports feature skew for the input feature 'brightness' but no prediction drift. You want to determine whether the skew is caused by a change in the distribution of incoming images compared to the training data. Which monitoring configuration should you inspect first?

Medium
12

An ML engineer needs to monitor the online prediction latency of a Vertex AI Endpoint. Which metrics should they look at in Cloud Monitoring?

Easy
13

A company uses Vertex AI Model Monitoring on an Endpoint that serves a regression model. They configure monitoring for both feature skew and prediction drift with a 10% threshold. After a week, they receive an alert that prediction drift exceeds the threshold, but feature skew remains below threshold. They want to understand what this indicates about the model's performance. What should they conclude?

Hard
14

An ML engineer is configuring Vertex AI Model Monitoring for drift detection on a deployed endpoint. Which TWO settings directly affect the frequency and accuracy of drift detection? (Choose 2)

Medium
15

A company wants to automatically retrain their model when data drift is detected. Which THREE components are needed to implement this pipeline?

Medium
16

A model deployed on a Vertex AI Endpoint uses an image model with XRAI explainability. The team notices that the prediction distributions are shifting over time. They want to monitor prediction drift. However, the explainability feature is not enabled. What must the engineer do to enable monitoring prediction drift?

Hard
17

You are configuring Vertex AI Model Monitoring for a deployed model on a Vertex AI Endpoint. The model uses a mix of numerical and categorical features. You want to ensure that the monitoring job effectively detects drift while minimizing false alerts. Which two actions should you take? (Choose two.)

Medium
18

Your team has deployed a tabular model to a Vertex AI Endpoint and enabled Vertex AI Model Monitoring with training-serving skew detection. You configured the monitoring job to run hourly and store statistics in a Cloud Storage bucket. After the first run, you notice that the job did not produce any drift metrics. You want to determine the cause of the missing metrics. What should you do first?

Medium
19

You are using Vertex AI Model Monitoring to detect prediction drift on a deployed model that serves online predictions. You want to ensure that the monitoring job can correctly compute drift metrics for numerical features. Which two configurations are required for the monitoring job to compute drift for a numerical feature? (Choose two.)

Hard
20

An ML engineer is monitoring a deployed model on Vertex AI Endpoints and wants to detect anomalies in the distribution of a categorical feature 'product_category' which has 50 possible values. The engineer configures Vertex AI Model Monitoring with training-serving skew detection. After a week, they receive an alert that the feature is skewed. Upon investigation, they find that a new product category was introduced in the serving data that was not present in the training data. What is the most likely reason for the alert?

Hard
21

A team has deployed a model on a Vertex AI Endpoint and enabled Vertex AI Model Monitoring for feature skew. They notice that the skew metric for a categorical feature with high cardinality is consistently high, even though the feature's distribution appears stable to the team. What is the most likely cause of this high skew metric?

Hard
22

A financial institution has deployed a fraud detection model on a Vertex AI Endpoint. The model uses both numerical and categorical features. They have enabled Vertex AI Model Monitoring with training data and configured drift thresholds. After several weeks, they notice that the model's recall has dropped significantly, but the overall prediction distribution remains stable. They suspect that a specific subgroup of transactions is being misclassified. Which approach should they use to identify the subgroup and diagnose the issue?

Hard
23

An ML engineer is configuring Vertex AI Model Monitoring for a deployed model that receives a mix of numerical and categorical features. The engineer wants to ensure that the monitoring job can detect both data drift and training-serving skew. Which two configurations are required to enable both types of detection? (Choose two.)

Hard
24

An ML engineer has deployed a tabular binary classification model to a Vertex AI Endpoint. After enabling Vertex AI Model Monitoring with training-serving skew detection, the engineer notices that the feature 'customer_age' is flagged as skewed. The feature values at serving time appear to be shifted upward by about 10 years compared to the training data. Which of the following is the most likely root cause?

Medium
25

A company has a model serving predictions on Vertex AI Endpoints and wants to monitor for prediction drift. They enable Vertex AI Model Monitoring but also need to see a confusion matrix over time. How should they set up the confusion matrix monitoring?

Medium
26

An organisation wants to monitor fairness of their loan approval model across demographic subgroups. They have predictions stored in BigQuery along with ground truth. Which GCP service can evaluate model performance for each subgroup and identify disparities?

Medium
27

An ML engineer is setting up Vertex AI Model Monitoring for a deployed model on a Vertex AI Endpoint. They want to receive alerts when either feature skew or prediction drift exceeds a threshold. Which two configurations are required to enable these alerts? (Choose two.)

Medium
28

An ML team is using Population Stability Index (PSI) to monitor feature drift on a Vertex AI Endpoint. The PSI value for a feature is 0.25, which exceeds the alert threshold of 0.2. The feature has high SHAP importance. The team wants to automatically retrain the model. What is the correct end-to-end setup?

Hard
29

A team is monitoring a model and observes that the error rate (prediction failures) has increased. They have enabled request/response logging on the Vertex AI Endpoint. How can they set up a metric and alert for prediction error rate?

Medium
30

A company wants to monitor the cost of their Vertex AI prediction endpoint. They are charged per hour per replica and per request for GPU instances. Which approach should they use to track these costs?

Easy
31

A company has deployed a model for image classification and wants to monitor for feature drift using XRAI attributions. However, they notice that the XRAI attribution maps are too large and are causing high latency in the monitoring pipeline. What is the most effective way to reduce the overhead of explainability monitoring for image models?

Hard
32

A company has deployed a model to a Vertex AI Endpoint and wants to receive an alert when the model's prediction latency exceeds a threshold. They have configured Cloud Monitoring to track the endpoint's latency metrics. They now need to create a notification channel to send alerts to their on-call team. Which Cloud Monitoring resource should they use to define the condition that triggers the alert?

Easy
33

You are an MLOps engineer at a retail company. Your team has deployed a demand forecasting model to a Vertex AI Endpoint. The model uses 12 numerical features and outputs a single numeric value representing predicted units sold. You have configured Vertex AI Model Monitoring with a training dataset that includes the full feature schema and prediction distribution. After a week, you observe that the feature 'promotion_flag' has a Jensen-Shannon divergence of 0.15, while all other features remain below 0.05. The model's prediction distribution has also shifted. Which action should you take first to diagnose the cause of the drift?

Medium
34

An ML engineer needs to track the costs incurred by Vertex AI prediction endpoints. Which tool should they use to set budget alerts and monitor spending?

Easy
35

A team is using Vertex AI Model Monitoring to detect prediction drift on a deployed model. They have configured the monitoring job to run every 6 hours. After a week, they notice that the drift metric has been consistently high but no alerts have been triggered. They have verified that the alerting policy in Cloud Monitoring is correctly configured. What is the most likely cause?

Medium
36

A data scientist has deployed a classification model on a Vertex AI Endpoint and wants to monitor for feature drift in the serving data compared to the training data. Which Vertex AI service should be used?

Easy
37

An ML engineer is troubleshooting why a Vertex AI Endpoint is returning high prediction latency. They have enabled request/response logging and see that some requests take >1 second while most are fast. Which THREE actions should they take to diagnose the issue?

Hard
38

An ML engineer is using Vertex AI Model Monitoring on a deployed model that predicts customer churn. The model's input features include a mix of numerical and categorical features. The engineer wants to detect changes in the distribution of a specific categorical feature, 'contract_type', which has 20 possible values. Which approach should be used to effectively monitor drift for this feature?

Hard
39

A company wants to monitor fairness of a model by evaluating performance metrics across demographic subgroups. They have ground truth labels stored in BigQuery. Which Vertex AI service should they use?

Easy
40

A company has deployed a model to a Vertex AI Endpoint and wants to receive an email notification whenever Vertex AI Model Monitoring detects feature drift above a configured threshold. They have already set up the monitoring configuration with a training dataset baseline. What should they do next to enable email alerts?

Easy
41

An MLOps engineer needs to collect ground truth labels for a deployed classification model to compare predictions against actuals. Where should the engineer store the ground truth data to enable Vertex AI model quality monitoring?

Easy
42

A media company uses a Vertex AI Endpoint to serve a video recommendation model. They have enabled Vertex AI Model Monitoring for prediction drift. After a major news event, they observe a significant increase in prediction drift alerts, but the model's recommendations remain relevant and user engagement is stable. They want to reduce unnecessary alerts without losing the ability to detect true model degradation. What should they do?

Hard
43

An ML team wants to automatically retrain a model when data drift is detected. They have set up a Cloud Monitoring alert on drift. What service should they use to trigger a retraining pipeline in response to the alert?

Easy
44

A healthcare company has deployed a diagnostic model on Vertex AI Endpoints. They use Vertex AI Model Monitoring to detect drift in features and predictions. The model's input features include patient age, which is a numerical feature. The monitoring job has been running for a month and has generated several alerts for age drift. However, the model's performance has not degraded. The team wants to reduce false alerts without missing critical drifts. What should they do?

Medium
45

A data science team is configuring Vertex AI Model Monitoring for a deployed model. They want to detect both feature skew and feature drift. Which TWO configurations must they set?

Medium
46

You have deployed a model to a Vertex AI Endpoint and enabled Vertex AI Model Monitoring with a monitoring frequency of every 24 hours. The model serves predictions with a feature called 'transaction_amount' that has a skewed distribution. After several days, you notice that the drift metrics for this feature are consistently below the threshold, but you suspect that the feature distribution has actually shifted. You want to improve the sensitivity of drift detection for this feature. What should you do?

Hard
47

A media streaming company uses a recommendation model deployed on Vertex AI Endpoints. The model predicts whether a user will click on a recommended item. They have set up Vertex AI Model Monitoring with a training dataset and configured drift detection for both features and predictions. Recently, they observed that the prediction drift metric (Jensen-Shannon divergence) for the 'click' prediction has exceeded the threshold, but feature drifts are within normal ranges. What is the most likely cause of this prediction drift?

Hard
48

A logistics company has a Vertex AI Model Monitoring job that detects feature drift on a deployed route optimization model. The model uses 50 features, and the monitoring job is configured to monitor all features. The MLOps team notices that the monitoring job is incurring high costs and taking a long time to complete. They want to reduce monitoring overhead while still detecting significant drift on the most important features. What should they do?

Hard
49

A retail company has a model deployed on a Vertex AI Endpoint that predicts customer lifetime value. The ML team wants to monitor the model for feature drift without setting up a full Vertex AI Model Monitoring job. They need a lightweight solution that compares live prediction requests to a reference distribution and sends alerts when drift exceeds a threshold. What should they do?

Easy
50

A retail company has deployed a demand forecasting model on a Vertex AI Endpoint. The model uses 20 numeric features. The MLOps team wants Vertex AI Model Monitoring to detect training-serving skew for each feature and receive alerts when skew exceeds a threshold. They have enabled Model Monitoring for the endpoint and configured a monitoring frequency of every 24 hours. However, after several days, no skew metrics appear in the Vertex AI console. What is the most likely cause?

Medium
51

An ML engineer is monitoring a Vertex AI Endpoint and notices a spike in 5xx error rates. Which TWO metrics should they examine to diagnose the issue? (Choose 2)

Easy
52

An ML engineer has set up Vertex AI Model Monitoring on an endpoint with a sampling rate of 0.1 (10%). They notice that the monitoring job runs hourly but the reported drift metrics seem inconsistent. What is the most likely cause?

Medium
53

A team wants to collect ground truth labels for their model deployed on Vertex AI Endpoint to perform model quality monitoring. They have a process that generates actual outcomes within 24 hours of prediction. What is the recommended approach for storing these labels?

Medium
54

A team uses Vertex AI Pipelines for continuous training triggered by model drift. They want to monitor the pipeline execution cost and optimize resource usage. Which THREE metrics should they track? (Choose 3)

Hard
55

An ML engineer manages a Vertex AI Endpoint serving a fraud detection model. Compliance requires that every prediction be logged with its input features for audit, but the team also wants to minimize storage costs. They decide to enable request-response logging on the endpoint. Which configuration should they use to meet both requirements?

Hard
56

A team deployed a model to a Vertex AI Endpoint and enabled Vertex AI Model Monitoring for skew detection. They notice that training-serving skew metrics are only produced when the endpoint receives traffic, but they want to ensure the skew is computed correctly even during periods of low traffic. Which configuration should they adjust to ensure skew detection remains statistically valid without generating excessive false positives?

Medium
57

An MLOps engineer is setting up monitoring for a deployed model on Vertex AI Endpoints. Which TWO actions are required to enable Vertex AI Model Monitoring for feature skew and drift? (Choose two.)

Medium
58

A company is experiencing high prediction costs on Vertex AI Endpoints. They want to monitor and optimize costs. Which THREE actions should they take? (Choose 3)

Hard
59

A company wants to log all prediction requests and responses from a Vertex AI Endpoint to BigQuery for auditing and debugging. How can they achieve this?

Easy
60

A team is using Vertex AI Explainability with a deployed model. They need to generate explanations for image classification predictions. Which explanation method should they configure in the ExplanationSpec?

Hard
61

An ML engineer needs to monitor the error rate of prediction jobs on a Vertex AI Endpoint. Where can they view the number of failed prediction requests over time?

Easy
62

Your team is using Vertex AI Pipelines to train a model weekly. You want to monitor the pipeline for failures and receive a notification when a pipeline run fails. You have configured the pipeline to send logs to Cloud Logging. What should you do to receive an alert on pipeline failure?

Medium
63

An ML engineer is troubleshooting a Vertex AI Model Monitoring setup on a deployed model. The monitoring configuration uses a training dataset baseline and monitors several numerical features. After several days, the engineer notices that drift scores are being computed, but no alerts have fired even though one feature's distribution has shifted dramatically. The monitoring configuration specifies a drift threshold of 0.3, and the observed drift score for that feature is 0.45. What is the most likely explanation?

Hard

Frequently asked questions

What does the Monitoring ML Solutions domain cover on the PMLE exam?
Be able to select the right Vertex AI service for each monitoring need: Model Monitoring for drift and skew, Model Evaluation for fairness and accuracy, and request-response logging for auditing. The key skill is knowing which signal each service produces and how alerts trigger retraining.
How many questions are in this domain?
This page lists all 63 Monitoring ML Solutions questions in the PMLE question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Monitoring ML Solutions questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
google-pmle GOOGLE-PMLE pmle monitoring Practice Questions