easyMultiple Choice
PMLE Practice Question: An ML team is using Vertex AI Online Prediction…
An ML team is using Vertex AI Online Prediction and wants to receive alerts when the 99th percentile latency exceeds 500ms for more than 5 minutes. What is the best practice to set up this alert in Cloud Monitoring?
⚠ Common exam trap
Google Cloud often tests the misconception that you must create custom metrics or use log-based solutions for percentile-based alerting, when in fact Cloud Monitoring's distribution metrics and percentile aligners handle this natively.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the 'aiplatform.googleapis.com/prediction/online_prediction_latencies' metric with a metric threshold condition set to 500ms and a percentile aligner of 99.
Cloud Monitoring provides a pre-built metric, `aiplatform.googleapis.com/prediction/online_prediction_latencies`, which directly captures prediction latency. By applying a percentile aligner of 99 and a metric threshold condition of 500ms, you can alert when the 99th percentile latency exceeds 500ms for the specified duration, without needing custom instrumentation or external processing.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Create a custom metric from the prediction container that emits latency percentiles, then set an alert on that metric.
Why it's wrong here
Vertex AI already exposes prediction latency percentiles as built-in metrics, so emitting a duplicate custom metric from the container adds instrumentation overhead without new signal. Custom metrics are warranted when the platform exposes no relevant metric, such as bespoke business KPIs.
- ✓
Use the 'aiplatform.googleapis.com/prediction/online_prediction_latencies' metric with a metric threshold condition set to 500ms and a percentile aligner of 99.
Why this is correct
The online_prediction_latencies metric exposes serving latency histograms; applying a 99th percentile aligner isolates tail latency, and a 500ms threshold condition with a five-minute duration triggers only on sustained breaches, matching the stem's alerting requirement precisely.
- ✗
Use a log-based metric to parse latency from Cloud Logging and alert when the average exceeds 500ms.
Why it's wrong here
A log-based metric computes an average across matched entries, which cannot express a 99th percentile, so the 500ms threshold would be evaluated against the wrong statistic. Log-based metrics suit counting discrete events, such as error occurrences, not distribution-based latency alerting.
- ✗
Export prediction latency logs to BigQuery and run a scheduled query to check the 99th percentile, then trigger a Cloud Function to send an alert.
Why it's wrong here
Cloud Monitoring natively evaluates percentile distributions with alerting policies, so exporting to BigQuery adds query latency and cannot meet the five-minute sustained threshold. Scheduled queries suit retrospective analytics or long-horizon reporting, not real-time latency alerting.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.