easyMultiple Choice
PMLE Practice Question: Set up monitoring for a Vertex AI model that…
You need to set up monitoring for a Vertex AI model that serves predictions in real-time. The model is expected to have a latency SLA of under 100ms. Which metric should you configure an alert on to ensure the SLA is met?
⚠ Common exam trap
Google Cloud often tests the misconception that median (p50) latency is sufficient for SLAs, but the trap is that SLAs require tail-latency guarantees (p99 or p999) to catch performance outliers that violate the threshold.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
p99 latency of prediction requests
P99 latency measures the worst-case latency experienced by 99% of requests, which is the standard metric for enforcing a strict SLA like under 100ms. Monitoring p99 ensures that even the slowest 1% of requests do not violate the threshold, providing a robust guarantee for real-time predictions.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
p50 latency of prediction requests
Why it's wrong here
A p50 alert tracks the median request, so half of requests could exceed 100ms while the median stays below it, leaving the SLA breached undetected. p50 suits typical-user experience monitoring; an SLA ceiling requires a high percentile such as p95 or p99, which captures tail latency.
- ✗
Prediction drift score
Why it's wrong here
Prediction drift measures distributional change between training and serving inputs, not request duration, so it cannot detect a 100ms latency SLA breach. Drift monitoring is the right choice for flagging model staleness or data skew over time, not for real-time latency alerting.
- ✓
p99 latency of prediction requests
Why this is correct
A latency SLA is a tail-latency commitment, so p99 captures the slowest 1% of prediction requests that breach the 100ms threshold. Alerting on p99 detects SLA violations that average latency would hide behind fast responses.
- ✗
Number of prediction requests per second
Why it's wrong here
Request throughput measures load, not response time, so it cannot confirm the 100ms SLA. Alert on prediction latency percentiles from Vertex AI online prediction monitoring. Throughput alerting is correct for capacity planning or quota exhaustion, where volume rather than responsiveness is the concern.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.