Courseiva
Serving and Scaling Models →mediumMultiple Choice

PMLE Serving and Scaling Models Practice Question

Your team has deployed a model on Vertex AI endpoints. You need to monitor the prediction latency to ensure it meets a 99th percentile SLO of 500ms. You want to set up an alert if the latency exceeds this threshold. Which metric should you use?

⚠ Common exam trap

Google Cloud often tests the distinction between tail latency (percentiles) and central tendency (average) or extreme values (maximum), trapping candidates who confuse SLO monitoring with simple failure counts or averages.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The 99th percentile of the `prediction/online/response_latencies` metric.

The `prediction/online/response_latencies` metric in Vertex AI provides a distribution of latency values, allowing you to query the 99th percentile directly. This aligns with the SLO requirement to monitor the tail latency, not the average or maximum, ensuring that the worst-case performance for 1% of requests stays under 500ms.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    The 99th percentile of the `prediction/online/response_latencies` metric.

    Why this is correct

    The prediction/online/response_latencies metric records server-side latency for online predictions, and its 99th percentile aligns exactly with the stem's 500ms SLO threshold. Alerting on that percentile detects tail latency affecting the slowest 1% of requests, which averages or medians would obscure.

  • ✗

    The number of prediction requests that timeout.

    Why it's wrong here

    Timeouts count requests exceeding a fixed duration, not the latency distribution, so they cannot verify a 99th percentile threshold of 500ms. It is tempting because timeouts indicate slowness, but percentile SLOs require a latency histogram metric such as prediction latency.

  • ✗

    Average prediction latency from the endpoint's logs.

    Why it's wrong here

    Averages mask tail behaviour, so a mean latency well under 500ms can coexist with a 99th percentile far above it; the SLO is defined on p99, not the mean. Averages suit capacity trending or cost reporting, where typical request cost matters rather than worst-case user experience.

  • ✗

    The maximum prediction latency from the endpoint's monitoring dashboard.

    Why it's wrong here

    A single maximum sample is an outlier, not a percentile, so it cannot evidence a 99th percentile SLO and will trigger on one anomalous request. Maximum latency belongs in spike or worst-case diagnostics, not in an alert whose threshold is defined as p99.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.