PMLE Serving and Scaling Models Practice Question
Your team has deployed a model on Vertex AI endpoints. You need to monitor the prediction latency to ensure it meets a 99th percentile SLO of 500ms. You want to set up an alert if the latency exceeds this threshold. Which metric should you use?
⚠ Common exam trap
Google Cloud often tests the distinction between tail latency (percentiles) and central tendency (average) or extreme values (maximum), trapping candidates who confuse SLO monitoring with simple failure counts or averages.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The 99th percentile of the `prediction/online/response_latencies` metric.
The `prediction/online/response_latencies` metric in Vertex AI provides a distribution of latency values, allowing you to query the 99th percentile directly. This aligns with the SLO requirement to monitor the tail latency, not the average or maximum, ensuring that the worst-case performance for 1% of requests stays under 500ms.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
The 99th percentile of the `prediction/online/response_latencies` metric.
Why this is correct
The prediction/online/response_latencies metric records server-side latency for online predictions, and its 99th percentile aligns exactly with the stem's 500ms SLO threshold. Alerting on that percentile detects tail latency affecting the slowest 1% of requests, which averages or medians would obscure.
- ✗
The number of prediction requests that timeout.
Why it's wrong here
Timeouts count requests exceeding a fixed duration, not the latency distribution, so they cannot verify a 99th percentile threshold of 500ms. It is tempting because timeouts indicate slowness, but percentile SLOs require a latency histogram metric such as prediction latency.
- ✗
Average prediction latency from the endpoint's logs.
Why it's wrong here
Averages mask tail behaviour, so a mean latency well under 500ms can coexist with a 99th percentile far above it; the SLO is defined on p99, not the mean. Averages suit capacity trending or cost reporting, where typical request cost matters rather than worst-case user experience.
- ✗
The maximum prediction latency from the endpoint's monitoring dashboard.
Why it's wrong here
A single maximum sample is an outlier, not a percentile, so it cannot evidence a 99th percentile SLO and will trigger on one anomalous request. Maximum latency belongs in spike or worst-case diagnostics, not in an alert whose threshold is defined as p99.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.