easyMultiple Choice
PMLE Practice Question: An ML engineer is monitoring a Vertex AI Feature…
An ML engineer is monitoring a Vertex AI Feature Store used for online serving. Which metrics are most important to track for ensuring low-latency online serving?
⚠ Common exam trap
It's easy for candidates to confuse metrics for offline batch operations (like export duration) with those for online serving, or assume that storage-level metrics (like utilization) are sufficient for performance monitoring, when in fact only retrieval latency and error rate directly reflect the serving quality.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Feature value retrieval latency (p99) and error rate.
For online serving, the primary concern is the latency and reliability of feature value retrieval at inference time. The p99 retrieval latency directly measures the worst-case delay experienced by users, while the error rate captures failures that could cause serving disruptions. Other metrics like storage utilization or batch export duration are relevant for offline or batch pipelines, not real-time serving.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Number of feature stores and feature values.
Why it's wrong here
Counting feature stores and values describes inventory, not serving latency. Online serving depends on online store read latency, request QPS and cache hit rate, which this metric never captures. It is tempting because feature-count metrics do indicate scale and storage growth, making them the right choice for capacity planning or quota monitoring rather than latency assurance.
- ✗
Storage utilization and write throughput to the feature store.
Why it's wrong here
Storage utilisation and write throughput describe ingestion and capacity, not serving latency, so they miss the online path entirely. They are tempting because they matter for the offline store and bulk ingestion monitoring, but low-latency online serving requires tracking serving latency, request count and error rates.
- ✗
Batch export duration and number of exported features.
Why it's wrong here
Batch export duration measures offline materialisation to BigQuery or GCS, not the online serving path, so it cannot reveal online latency. It is tempting because export jobs genuinely matter when diagnosing stale feature data or scheduling offline training pipelines, where throughput and feature counts are the metrics worth watching.
- ✓
Feature value retrieval latency (p99) and error rate.
Why this is correct
Online serving latency is dominated by feature retrieval, so p99 retrieval latency and error rate directly expose slow or failing lookups that breach the low-latency constraint; aggregate CPU or storage metrics cannot reveal per-request serving delays.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.